{
  "id": 501828,
  "title": "Any suggestion for a GPU Cloud?",
  "url": "/competitions/birdclef-2024/discussion/501828",
  "author_name": "",
  "post_date": "2024-05-10T22:58:03.614679700Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>This year, unfortunately, I don't have access to a high-performance GPU PC. I attempted to utilize Google Colab Pro+ for pretraining, but as the dataset size is too large to load directly into RAM, I linked Google Drive to Colab. However, this approach led to significant delays when loading the dataset. Can you recommend a GPU cloud service that would be suitable for my needs and also affordable?</p>",
  "messages": [
    {
      "id": "2806141",
      "postDate": "05/10/2024 22:58:03",
      "content": "<p>This year, unfortunately, I don't have access to a high-performance GPU PC. I attempted to utilize Google Colab Pro+ for pretraining, but as the dataset size is too large to load directly into RAM, I linked Google Drive to Colab. However, this approach led to significant delays when loading the dataset. Can you recommend a GPU cloud service that would be suitable for my needs and also affordable?</p>",
      "rawMarkdown": "This year, unfortunately, I don't have access to a high-performance GPU PC. I attempted to utilize Google Colab Pro+ for pretraining, but as the dataset size is too large to load directly into RAM, I linked Google Drive to Colab. However, this approach led to significant delays when loading the dataset. Can you recommend a GPU cloud service that would be suitable for my needs and also affordable?",
      "votes": null
    },
    {
      "id": "2806143",
      "postDate": "05/10/2024 23:03:12",
      "content": "<p>You might try placing the data in a cloud bucket; I think this may be significantly faster for colab to access than Drive.</p>\n<p>You can test this by iterating over xeno-canto mp3's in our public GCS bucket; I'll be interested to know what you observe for read speeds compared to Drive:<br>\n<code>gs://chirp-public-bucket/xeno-canto/audio-data/*/*.mp3</code></p>",
      "rawMarkdown": "You might try placing the data in a cloud bucket; I think this may be significantly faster for colab to access than Drive.\n\nYou can test this by iterating over xeno-canto mp3's in our public GCS bucket; I'll be interested to know what you observe for read speeds compared to Drive:\n```gs://chirp-public-bucket/xeno-canto/audio-data/*/*.mp3```",
      "votes": null
    },
    {
      "id": "2806153",
      "postDate": "05/10/2024 23:28:11",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "2806183",
      "postDate": "05/11/2024 00:09:33",
      "content": "<p>(Just a smol caution: the crawl in that bucket is rather out of date, and won't match the Kaggle datasets. Meaning to get around to updating it eventually.)</p>",
      "rawMarkdown": "(Just a smol caution: the crawl in that bucket is rather out of date, and won't match the Kaggle datasets. Meaning to get around to updating it eventually.)",
      "votes": null
    },
    {
      "id": "2806303",
      "postDate": "05/11/2024 03:01:38",
      "content": "<p>I have also been training using Colab.<br>\nLike you, I experienced delays when loading datasets from Google Drive.<br>\nTo address this, I moved the dataset (about 23.43 GB) from Google Drive to Colab's Disk, which allowed for faster processing.</p>\n<pre><code>!tar -xf drivebirdclef2024/dataset.tar\n# using dataset as dataset directory\n</code></pre>\n<p>GPU instance has 201.2 GB disk space. so about 100GB is available for dataset.<br>\nI think you can place it on the disk if it's at most about 100GB</p>\n<p>I hope this can be helpful to you.</p>",
      "rawMarkdown": "I have also been training using Colab.\nLike you, I experienced delays when loading datasets from Google Drive.\nTo address this, I moved the dataset (about 23.43 GB) from Google Drive to Colab's Disk, which allowed for faster processing.\n\n\n```\n!tar -xf /content/drive/MyDrive/birdclef2024/dataset.tar\n# using /content/dataset as dataset directory\n```\n\nGPU instance has 201.2 GB disk space. so about 100GB is available for dataset.\nI think you can place it on the disk if it's at most about 100GB\n\nI hope this can be helpful to you.",
      "votes": null
    },
    {
      "id": "2806403",
      "postDate": "05/11/2024 04:41:19",
      "content": "<p>Thank you for your response. I've just begun the competition, and for the first stage, I need to pretrain on previous years' datasets, which are approximately 100 GB in size. Your strategy works well for handling this year's dataset size. </p>",
      "rawMarkdown": "Thank you for your response. I've just begun the competition, and for the first stage, I need to pretrain on previous years' datasets, which are approximately 100 GB in size. Your strategy works well for handling this year's dataset size.",
      "votes": null
    },
    {
      "id": "2806438",
      "postDate": "05/11/2024 05:12:41",
      "content": "<p>Thanks for letting me know.</p>",
      "rawMarkdown": "Thanks for letting me know.",
      "votes": null
    },
    {
      "id": "2807893",
      "postDate": "05/11/2024 23:57:00",
      "content": "<p>I've grappled with this idea already, and ended up concluding that none of the current cloud options were worth the setup faf at the moment.  Sorry that's not the answer you're hoping for.</p>\n<p>But one possibly more useful observation is that maybe you don't need a 'high performance' GPU as much as you realise.  My old Dell G7 laptop on Ubuntu, with an NVIDIA 1060 is easily out-performing Kaggle's own platform in terms of speed (but not quite as much memory as I'd like).   I'm taking under 10 minutes per epoch, it was much longer on Kaggle.  I think Kaggle is mostly bottlenecked by data retrieval not GPU speed, probably just like Colab.</p>\n<p>Also my 1060 is on Ubuntu is easily beating my 3060 on Windows just due to Windows issues.   So I think for this comp, fancy GPU's aren't essential.</p>",
      "rawMarkdown": "I've grappled with this idea already, and ended up concluding that none of the current cloud options were worth the setup faf at the moment.  Sorry that's not the answer you're hoping for.\n\nBut one possibly more useful observation is that maybe you don't need a 'high performance' GPU as much as you realise.  My old Dell G7 laptop on Ubuntu, with an NVIDIA 1060 is easily out-performing Kaggle's own platform in terms of speed (but not quite as much memory as I'd like).   I'm taking under 10 minutes per epoch, it was much longer on Kaggle.  I think Kaggle is mostly bottlenecked by data retrieval not GPU speed, probably just like Colab.\n\nAlso my 1060 is on Ubuntu is easily beating my 3060 on Windows just due to Windows issues.   So I think for this comp, fancy GPU's aren't essential.",
      "votes": null
    },
    {
      "id": "2808071",
      "postDate": "05/12/2024 04:28:35",
      "content": "<p>what are the problems with windows?</p>",
      "rawMarkdown": "what are the problems with windows?",
      "votes": null
    },
    {
      "id": "2809793",
      "postDate": "05/13/2024 02:05:00",
      "content": "<p>The biggest problem I have with Windows comes when using multi-CPU processing for your data loading prior to sending off to the GPU.  For example if you set num_workers to anything other than 0 with Pytorch Lightning your IDE will crash.  It is because of the way windows spawns a new instance of your whole script, including the bit that does the spawning, so it becomes 2,4,8…, whereas Linux will just take the variables and states it needs for the forked process.   In a .py script there is an easy enough solution, protecting the critical line with an if <strong>name</strong> == 'main': so it doesn't get re-run when spawned,   but it still slows down the overall process because of the spawning.</p>\n<p>It is an even bigger pain in the bum to work with multi-cpu processing in a Jupyter Notebook environment on Windows.  I've wasted a fair bit of time over this, because my work computer needs to be in Windows.   There are work-arounds, but none of them are very appealing.</p>\n<p>Since this particular comp requires quite a lot of pre-processing, especially if you're working directly with sound files rather than spectrograms, I think the CPU part is quite critical.  If you're stuck with a windows machine, then maybe consider debugging in a notebook, or on Kaggle, but then converting to a .py script before long training runs?</p>",
      "rawMarkdown": "The biggest problem I have with Windows comes when using multi-CPU processing for your data loading prior to sending off to the GPU.  For example if you set num_workers to anything other than 0 with Pytorch Lightning your IDE will crash.  It is because of the way windows spawns a new instance of your whole script, including the bit that does the spawning, so it becomes 2,4,8..., whereas Linux will just take the variables and states it needs for the forked process.   In a .py script there is an easy enough solution, protecting the critical line with an if __name__ == 'main': so it doesn't get re-run when spawned,   but it still slows down the overall process because of the spawning.\n\nIt is an even bigger pain in the bum to work with multi-cpu processing in a Jupyter Notebook environment on Windows.  I've wasted a fair bit of time over this, because my work computer needs to be in Windows.   There are work-arounds, but none of them are very appealing.\n\nSince this particular comp requires quite a lot of pre-processing, especially if you're working directly with sound files rather than spectrograms, I think the CPU part is quite critical.  If you're stuck with a windows machine, then maybe consider debugging in a notebook, or on Kaggle, but then converting to a .py script before long training runs?",
      "votes": null
    },
    {
      "id": "2812012",
      "postDate": "05/14/2024 04:10:05",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tjamali\" target=\"_blank\">@tjamali</a>, I had the similar question as you several days ago. Did some research on some popular cloud gpu platforms on the market and tried them out myself. </p>\n<p>Currently <strong><em>vast.ai</em></strong> works for me the best, but also see <strong><em>paperspace</em></strong> and <strong><em>runpod</em></strong> can be good choices depending on your use cases.</p>\n<p>Had a post in General (<a href=\"https://www.kaggle.com/discussions/general/494678\" target=\"_blank\">https://www.kaggle.com/discussions/general/494678</a>) where I shared my experience on using those cloud gpu platforms. If you find something better I would also be happy to know. Thanks!</p>",
      "rawMarkdown": "Hey @tjamali, I had the similar question as you several days ago. Did some research on some popular cloud gpu platforms on the market and tried them out myself. \n\nCurrently ***vast.ai*** works for me the best, but also see ***paperspace*** and ***runpod*** can be good choices depending on your use cases.\n\nHad a post in General (https://www.kaggle.com/discussions/general/494678) where I shared my experience on using those cloud gpu platforms. If you find something better I would also be happy to know. Thanks!",
      "votes": null
    },
    {
      "id": "2812940",
      "postDate": "05/14/2024 13:39:13",
      "content": "<p>Hey! Thanks for your reply. I haven't tried these services yet. Currently, I'm working with Hyperstack, and it has been good so far. I went with it because of the price. The only issue I faced was that it wasn't clear how to connect using SSH the first time, and you need to install CUDA and some other things on your instance virtual machine. The price is reasonable, and compared to vast.ai, as you mentioned in your post, your machine is always available.</p>",
      "rawMarkdown": "Hey! Thanks for your reply. I haven't tried these services yet. Currently, I'm working with Hyperstack, and it has been good so far. I went with it because of the price. The only issue I faced was that it wasn't clear how to connect using SSH the first time, and you need to install CUDA and some other things on your instance virtual machine. The price is reasonable, and compared to vast.ai, as you mentioned in your post, your machine is always available.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2806143,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "05/10/2024 23:03:12",
      "content": "<p>You might try placing the data in a cloud bucket; I think this may be significantly faster for colab to access than Drive.</p>\n<p>You can test this by iterating over xeno-canto mp3's in our public GCS bucket; I'll be interested to know what you observe for read speeds compared to Drive:<br>\n<code>gs://chirp-public-bucket/xeno-canto/audio-data/*/*.mp3</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 2806153,
          "author_name": "tjamali",
          "author_url": "",
          "post_date": "05/10/2024 23:28:11",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2806183,
              "author_name": "tomdenton",
              "author_url": "",
              "post_date": "05/11/2024 00:09:33",
              "content": "<p>(Just a smol caution: the crawl in that bucket is rather out of date, and won't match the Kaggle datasets. Meaning to get around to updating it eventually.)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2806438,
                  "author_name": "tjamali",
                  "author_url": "",
                  "post_date": "05/11/2024 05:12:41",
                  "content": "<p>Thanks for letting me know.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2806303,
      "author_name": "neilus",
      "author_url": "",
      "post_date": "05/11/2024 03:01:38",
      "content": "<p>I have also been training using Colab.<br>\nLike you, I experienced delays when loading datasets from Google Drive.<br>\nTo address this, I moved the dataset (about 23.43 GB) from Google Drive to Colab's Disk, which allowed for faster processing.</p>\n<pre><code>!tar -xf drivebirdclef2024/dataset.tar\n# using dataset as dataset directory\n</code></pre>\n<p>GPU instance has 201.2 GB disk space. so about 100GB is available for dataset.<br>\nI think you can place it on the disk if it's at most about 100GB</p>\n<p>I hope this can be helpful to you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2806403,
          "author_name": "tjamali",
          "author_url": "",
          "post_date": "05/11/2024 04:41:19",
          "content": "<p>Thank you for your response. I've just begun the competition, and for the first stage, I need to pretrain on previous years' datasets, which are approximately 100 GB in size. Your strategy works well for handling this year's dataset size. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2807893,
      "author_name": "ollypowell",
      "author_url": "",
      "post_date": "05/11/2024 23:57:00",
      "content": "<p>I've grappled with this idea already, and ended up concluding that none of the current cloud options were worth the setup faf at the moment.  Sorry that's not the answer you're hoping for.</p>\n<p>But one possibly more useful observation is that maybe you don't need a 'high performance' GPU as much as you realise.  My old Dell G7 laptop on Ubuntu, with an NVIDIA 1060 is easily out-performing Kaggle's own platform in terms of speed (but not quite as much memory as I'd like).   I'm taking under 10 minutes per epoch, it was much longer on Kaggle.  I think Kaggle is mostly bottlenecked by data retrieval not GPU speed, probably just like Colab.</p>\n<p>Also my 1060 is on Ubuntu is easily beating my 3060 on Windows just due to Windows issues.   So I think for this comp, fancy GPU's aren't essential.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2808071,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "05/12/2024 04:28:35",
          "content": "<p>what are the problems with windows?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2809793,
              "author_name": "ollypowell",
              "author_url": "",
              "post_date": "05/13/2024 02:05:00",
              "content": "<p>The biggest problem I have with Windows comes when using multi-CPU processing for your data loading prior to sending off to the GPU.  For example if you set num_workers to anything other than 0 with Pytorch Lightning your IDE will crash.  It is because of the way windows spawns a new instance of your whole script, including the bit that does the spawning, so it becomes 2,4,8…, whereas Linux will just take the variables and states it needs for the forked process.   In a .py script there is an easy enough solution, protecting the critical line with an if <strong>name</strong> == 'main': so it doesn't get re-run when spawned,   but it still slows down the overall process because of the spawning.</p>\n<p>It is an even bigger pain in the bum to work with multi-cpu processing in a Jupyter Notebook environment on Windows.  I've wasted a fair bit of time over this, because my work computer needs to be in Windows.   There are work-arounds, but none of them are very appealing.</p>\n<p>Since this particular comp requires quite a lot of pre-processing, especially if you're working directly with sound files rather than spectrograms, I think the CPU part is quite critical.  If you're stuck with a windows machine, then maybe consider debugging in a notebook, or on Kaggle, but then converting to a .py script before long training runs?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2812012,
      "author_name": "faithk7u",
      "author_url": "",
      "post_date": "05/14/2024 04:10:05",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tjamali\" target=\"_blank\">@tjamali</a>, I had the similar question as you several days ago. Did some research on some popular cloud gpu platforms on the market and tried them out myself. </p>\n<p>Currently <strong><em>vast.ai</em></strong> works for me the best, but also see <strong><em>paperspace</em></strong> and <strong><em>runpod</em></strong> can be good choices depending on your use cases.</p>\n<p>Had a post in General (<a href=\"https://www.kaggle.com/discussions/general/494678\" target=\"_blank\">https://www.kaggle.com/discussions/general/494678</a>) where I shared my experience on using those cloud gpu platforms. If you find something better I would also be happy to know. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2812940,
          "author_name": "tjamali",
          "author_url": "",
          "post_date": "05/14/2024 13:39:13",
          "content": "<p>Hey! Thanks for your reply. I haven't tried these services yet. Currently, I'm working with Hyperstack, and it has been good so far. I went with it because of the price. The only issue I faced was that it wasn't clear how to connect using SSH the first time, and you need to install CUDA and some other things on your instance virtual machine. The price is reasonable, and compared to vast.ai, as you mentioned in your post, your machine is always available.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2806141": "This year, unfortunately, I don't have access to a high-performance GPU PC. I attempted to utilize Google Colab Pro+ for pretraining, but as the dataset size is too large to load directly into RAM, I linked Google Drive to Colab. However, this approach led to significant delays when loading the dataset. Can you recommend a GPU cloud service that would be suitable for my needs and also affordable?",
    "2806143": "You might try placing the data in a cloud bucket; I think this may be significantly faster for colab to access than Drive.\n\nYou can test this by iterating over xeno-canto mp3's in our public GCS bucket; I'll be interested to know what you observe for read speeds compared to Drive:\n```gs://chirp-public-bucket/xeno-canto/audio-data/*/*.mp3```",
    "2806153": "Thank you!",
    "2806183": "(Just a smol caution: the crawl in that bucket is rather out of date, and won't match the Kaggle datasets. Meaning to get around to updating it eventually.)",
    "2806303": "I have also been training using Colab.\nLike you, I experienced delays when loading datasets from Google Drive.\nTo address this, I moved the dataset (about 23.43 GB) from Google Drive to Colab's Disk, which allowed for faster processing.\n\n\n```\n!tar -xf /content/drive/MyDrive/birdclef2024/dataset.tar\n# using /content/dataset as dataset directory\n```\n\nGPU instance has 201.2 GB disk space. so about 100GB is available for dataset.\nI think you can place it on the disk if it's at most about 100GB\n\nI hope this can be helpful to you.",
    "2806403": "Thank you for your response. I've just begun the competition, and for the first stage, I need to pretrain on previous years' datasets, which are approximately 100 GB in size. Your strategy works well for handling this year's dataset size.",
    "2806438": "Thanks for letting me know.",
    "2807893": "I've grappled with this idea already, and ended up concluding that none of the current cloud options were worth the setup faf at the moment.  Sorry that's not the answer you're hoping for.\n\nBut one possibly more useful observation is that maybe you don't need a 'high performance' GPU as much as you realise.  My old Dell G7 laptop on Ubuntu, with an NVIDIA 1060 is easily out-performing Kaggle's own platform in terms of speed (but not quite as much memory as I'd like).   I'm taking under 10 minutes per epoch, it was much longer on Kaggle.  I think Kaggle is mostly bottlenecked by data retrieval not GPU speed, probably just like Colab.\n\nAlso my 1060 is on Ubuntu is easily beating my 3060 on Windows just due to Windows issues.   So I think for this comp, fancy GPU's aren't essential.",
    "2808071": "what are the problems with windows?",
    "2809793": "The biggest problem I have with Windows comes when using multi-CPU processing for your data loading prior to sending off to the GPU.  For example if you set num_workers to anything other than 0 with Pytorch Lightning your IDE will crash.  It is because of the way windows spawns a new instance of your whole script, including the bit that does the spawning, so it becomes 2,4,8..., whereas Linux will just take the variables and states it needs for the forked process.   In a .py script there is an easy enough solution, protecting the critical line with an if __name__ == 'main': so it doesn't get re-run when spawned,   but it still slows down the overall process because of the spawning.\n\nIt is an even bigger pain in the bum to work with multi-cpu processing in a Jupyter Notebook environment on Windows.  I've wasted a fair bit of time over this, because my work computer needs to be in Windows.   There are work-arounds, but none of them are very appealing.\n\nSince this particular comp requires quite a lot of pre-processing, especially if you're working directly with sound files rather than spectrograms, I think the CPU part is quite critical.  If you're stuck with a windows machine, then maybe consider debugging in a notebook, or on Kaggle, but then converting to a .py script before long training runs?",
    "2812012": "Hey @tjamali, I had the similar question as you several days ago. Did some research on some popular cloud gpu platforms on the market and tried them out myself. \n\nCurrently ***vast.ai*** works for me the best, but also see ***paperspace*** and ***runpod*** can be good choices depending on your use cases.\n\nHad a post in General (https://www.kaggle.com/discussions/general/494678) where I shared my experience on using those cloud gpu platforms. If you find something better I would also be happy to know. Thanks!",
    "2812940": "Hey! Thanks for your reply. I haven't tried these services yet. Currently, I'm working with Hyperstack, and it has been good so far. I went with it because of the price. The only issue I faced was that it wasn't clear how to connect using SSH the first time, and you need to install CUDA and some other things on your instance virtual machine. The price is reasonable, and compared to vast.ai, as you mentioned in your post, your machine is always available."
  },
  "source": "meta"
}