{
  "id": 226894,
  "title": "A word of advice to Colab + GCS users",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/226894",
  "author_name": "Chan Kha Vu",
  "post_date": "2021-03-18T06:03:14.211000",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi fellow kagglers with no home-baked deep learning rig because all GPUs in 2021 are out of stock thanks to miners 😄</p>\n<p>I'm using Colab Pro (they offer nice P100 and V100 GPUs, and even TPUs but it's harder to use if you're on PyTorch) and store data in Google Cloud Storage buckets. To improve GPU utilization, I'm using <a href=\"https://github.com/tmbdev/webdataset\" target=\"_blank\">WebDataset</a>. I believe that Colab is enough to get some decent results (for example, <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037\" target=\"_blank\">this champ</a> got 1st place on <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/overview\" target=\"_blank\">Google Landmarks Retrieval 2020</a> with just Colab Pro).</p>\n<p>Usually it's a very cheap option (for example, in <a href=\"https://www.kaggle.com/c/landmark-recognition-2020\" target=\"_blank\">Google Landmark Recognition 2020</a> challenge, my data flow for 5 parallel instances was 120GB per hour over 20 days and it costed me less than <strong>$20</strong>).</p>\n<p>But this time, with just 2 days of trying different experiment settings with 3 parallel instances, 11GB per 10 mins each instance, I just got a <strong>$275</strong> bill (which is like… 15 burgers).</p>\n<p>I forgot that Colab can give you instances from Asia and Europe (my chance of getting an instance from Taiwan was ~1/3), while my GCS buckets are located in the US. And the data transfer rate is… a whopping <strong>$0.12 per GB!</strong> Yikes!</p>\n<p>To get the information about the location of your Colab instance, put</p>\n<pre><code>!curl ipinfo.io\n</code></pre>\n<p>in your notebook. Don't overpay for services like me 😄😄😄</p>",
  "messages": [
    {
      "id": 1243293,
      "postDate": "2021-03-18T06:03:14.213Z",
      "content": "<p>Hi fellow kagglers with no home-baked deep learning rig because all GPUs in 2021 are out of stock thanks to miners 😄</p>\n<p>I'm using Colab Pro (they offer nice P100 and V100 GPUs, and even TPUs but it's harder to use if you're on PyTorch) and store data in Google Cloud Storage buckets. To improve GPU utilization, I'm using <a href=\"https://github.com/tmbdev/webdataset\" target=\"_blank\">WebDataset</a>. I believe that Colab is enough to get some decent results (for example, <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037\" target=\"_blank\">this champ</a> got 1st place on <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/overview\" target=\"_blank\">Google Landmarks Retrieval 2020</a> with just Colab Pro).</p>\n<p>Usually it's a very cheap option (for example, in <a href=\"https://www.kaggle.com/c/landmark-recognition-2020\" target=\"_blank\">Google Landmark Recognition 2020</a> challenge, my data flow for 5 parallel instances was 120GB per hour over 20 days and it costed me less than <strong>$20</strong>).</p>\n<p>But this time, with just 2 days of trying different experiment settings with 3 parallel instances, 11GB per 10 mins each instance, I just got a <strong>$275</strong> bill (which is like… 15 burgers).</p>\n<p>I forgot that Colab can give you instances from Asia and Europe (my chance of getting an instance from Taiwan was ~1/3), while my GCS buckets are located in the US. And the data transfer rate is… a whopping <strong>$0.12 per GB!</strong> Yikes!</p>\n<p>To get the information about the location of your Colab instance, put</p>\n<pre><code>!curl ipinfo.io\n</code></pre>\n<p>in your notebook. Don't overpay for services like me 😄😄😄</p>",
      "rawMarkdown": "Hi fellow kagglers with no home-baked deep learning rig because all GPUs in 2021 are out of stock thanks to miners 😄\n\nI'm using Colab Pro (they offer nice P100 and V100 GPUs, and even TPUs but it's harder to use if you're on PyTorch) and store data in Google Cloud Storage buckets. To improve GPU utilization, I'm using [WebDataset](https://github.com/tmbdev/webdataset). I believe that Colab is enough to get some decent results (for example, [this champ](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037) got 1st place on [Google Landmarks Retrieval 2020](https://www.kaggle.com/c/landmark-retrieval-2020/overview) with just Colab Pro).\n\nUsually it's a very cheap option (for example, in [Google Landmark Recognition 2020](https://www.kaggle.com/c/landmark-recognition-2020) challenge, my data flow for 5 parallel instances was 120GB per hour over 20 days and it costed me less than **$20**).\n\nBut this time, with just 2 days of trying different experiment settings with 3 parallel instances, 11GB per 10 mins each instance, I just got a **$275** bill (which is like... 15 burgers).\n\nI forgot that Colab can give you instances from Asia and Europe (my chance of getting an instance from Taiwan was ~1/3), while my GCS buckets are located in the US. And the data transfer rate is... a whopping **$0.12 per GB!** Yikes!\n\nTo get the information about the location of your Colab instance, put\n```\n!curl ipinfo.io\n```\nin your notebook. Don't overpay for services like me 😄😄😄",
      "votes": 13
    },
    {
      "id": 1251633,
      "postDate": "2021-03-25T00:59:44.210Z",
      "content": "<p>GCS egress cost between continents is painful! I've had same experience.</p>\n<p>So nowadays and I'm trying to management data transfer by making buckets per continental.</p>\n<p>And here are my small tips using colab pro.</p>\n<ul>\n<li>For dataset, download from kaggle's private or public dataset using kaggle api(from kaggle to Colab Pro's disk).</li>\n<li>Be careful of saving and loading checkpoints to GCS. Sometimes checkpoints are also large. If you use GCS and Colab Pro on different locations or on different continents, saving checkpoints frequently can cause VERY BIG COST. </li>\n</ul>",
      "rawMarkdown": "GCS egress cost between continents is painful! I've had same experience.\n\nSo nowadays and I'm trying to management data transfer by making buckets per continental.\n\nAnd here are my small tips using colab pro.\n- For dataset, download from kaggle's private or public dataset using kaggle api(from kaggle to Colab Pro's disk).\n- Be careful of saving and loading checkpoints to GCS. Sometimes checkpoints are also large. If you use GCS and Colab Pro on different locations or on different continents, saving checkpoints frequently can cause VERY BIG COST. ",
      "votes": 1
    },
    {
      "id": 1243830,
      "postDate": "2021-03-18T14:15:59.983Z",
      "content": "<p>I'm a noob w.r.t. Colab Pro (I've only used the free Colab and Kaggle kernels). So I had a question related to this.</p>\n<p>Would it make more sense to store the data as a reduced (remove tfrecords if not using them and separate train and test)(or use only tfrecord files train and test separated) .zip file(s) on Google Drive. Then when you launch Colab Pro you can mount your drive and unzip the files directly into your Colab environment?</p>\n<p>I think this would take approximately 5-10 minutes and there would be no additional transfer costs? You might need to upgrade Google Drive storage… but it should be a negligible cost compared to GCS transfer costs.</p>\n<p>Is there a reason this wouldn't work (or why your current setup is better)? I'm still learning so hopefully you don't take this as a challenge. I'm just trying to identify good approaches for future workflows. TBH I'm not at all familiar with WebDataset so some of my ignorance and/or confusion may be related to using that tool.</p>",
      "rawMarkdown": "I'm a noob w.r.t. Colab Pro (I've only used the free Colab and Kaggle kernels). So I had a question related to this.\n\nWould it make more sense to store the data as a reduced (remove tfrecords if not using them and separate train and test)(or use only tfrecord files train and test separated) .zip file(s) on Google Drive. Then when you launch Colab Pro you can mount your drive and unzip the files directly into your Colab environment?\n\nI think this would take approximately 5-10 minutes and there would be no additional transfer costs? You might need to upgrade Google Drive storage... but it should be a negligible cost compared to GCS transfer costs.\n\nIs there a reason this wouldn't work (or why your current setup is better)? I'm still learning so hopefully you don't take this as a challenge. I'm just trying to identify good approaches for future workflows. TBH I'm not at all familiar with WebDataset so some of my ignorance and/or confusion may be related to using that tool.",
      "votes": 1,
      "replies": [
        {
          "id": 1243851,
          "postDate": "2021-03-18T14:29:19.150Z",
          "content": "<p>I'm a noob as well, so please take my reasoning with a grain of salt:</p>\n<ul>\n<li>The main reason is familiarity -- I used the same pipeline for the Google Landmarks challenge. I know that it's dirt cheap and easy to use and pretty fast, and it was easy for me to code a pipeline.</li>\n<li>Colab Pro has limited disk space, so if I want to use more data (i.e. external data, larger image size, maybe even predictions from other models) -- it would be problematic.</li>\n<li>TFRecordDataset and WebDataset are designed for exactly this type of cloud training, and it's pretty fast. As for WebDataset -- it's so good that they're planning to merge it to PyTorch repository 😄 Here is a post about it <a href=\"https://pytorch.org/blog/efficient-pytorch-io-library-for-large-datasets-many-files-many-gpus/\" target=\"_blank\">https://pytorch.org/blog/efficient-pytorch-io-library-for-large-datasets-many-files-many-gpus/</a></li>\n<li>If I ever find a good solution and want to scale up to larger models, I don't need to change any code when I move to google cloud instances for training.</li>\n<li>With the right pipeline, GCS has no transfer cost as well if you're doing so from within Google's services (the cost I mentioned above for Google Landmarks is mostly because I transferred the data from outside of google services).</li>\n</ul>",
          "rawMarkdown": "I'm a noob as well, so please take my reasoning with a grain of salt:\n- The main reason is familiarity -- I used the same pipeline for the Google Landmarks challenge. I know that it's dirt cheap and easy to use and pretty fast, and it was easy for me to code a pipeline.\n- Colab Pro has limited disk space, so if I want to use more data (i.e. external data, larger image size, maybe even predictions from other models) -- it would be problematic.\n- TFRecordDataset and WebDataset are designed for exactly this type of cloud training, and it's pretty fast. As for WebDataset -- it's so good that they're planning to merge it to PyTorch repository 😄 Here is a post about it https://pytorch.org/blog/efficient-pytorch-io-library-for-large-datasets-many-files-many-gpus/\n- If I ever find a good solution and want to scale up to larger models, I don't need to change any code when I move to google cloud instances for training.\n- With the right pipeline, GCS has no transfer cost as well if you're doing so from within Google's services (the cost I mentioned above for Google Landmarks is mostly because I transferred the data from outside of google services).\n\n",
          "votes": 1
        },
        {
          "id": 1243868,
          "postDate": "2021-03-18T14:41:09.227Z",
          "content": "<p>That makes a lot of sense. Sounds like a good setup! Thanks for the response and the advice about Colab Pro Instance localization.</p>",
          "rawMarkdown": "That makes a lot of sense. Sounds like a good setup! Thanks for the response and the advice about Colab Pro Instance localization.",
          "votes": 1
        },
        {
          "id": 1243879,
          "postDate": "2021-03-18T14:52:59.010Z",
          "content": "<p>If you're using TPUs, you don't have to worry about localization, because all TPU v2-8 are located in <code>us-central</code> 😄😄😄</p>",
          "rawMarkdown": "If you're using TPUs, you don't have to worry about localization, because all TPU v2-8 are located in `us-central` 😄😄😄",
          "votes": 1
        },
        {
          "id": 1255487,
          "postDate": "2021-03-28T21:00:46.343Z",
          "content": "<p>Great tips. </p>\n<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> I used to use Google Drive with free colab earlier. But you see this can be hassle at times.</p>\n<p>Now I use Weights and Biases Artifacts. For free user account they provide 100 GB of artifact storage. Once you log your dataset as an artifacts it's just so much convenient to download it from anywhere be it local system, colab or kaggle. </p>\n<p>Hope it helps. </p>",
          "rawMarkdown": "Great tips. \n\n@dschettler8845 I used to use Google Drive with free colab earlier. But you see this can be hassle at times.\n\nNow I use Weights and Biases Artifacts. For free user account they provide 100 GB of artifact storage. Once you log your dataset as an artifacts it's just so much convenient to download it from anywhere be it local system, colab or kaggle. \n\nHope it helps. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1261993,
      "postDate": "2021-04-03T16:56:12.480Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a></p>\n<p>I am just trying to understand what is the reason for downloading the dataset in Google drive or using WebDataset.</p>\n<p>You can train the model from Google Cloud Storage directly.</p>\n<p>You just need to change the path of the images.</p>\n<p>Here <a href=\"https://www.kaggle.com/tt0721\" target=\"_blank\">@tt0721</a> and me trying to find the best way to do it.<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962</a></p>\n<p>here another discussion about it.<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/227062\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/227062</a></p>\n<p>All the best<br>\nHappy Learning 😀</p>",
      "rawMarkdown": "Hi @chankhavu\n\nI am just trying to understand what is the reason for downloading the dataset in Google drive or using WebDataset.\n\nYou can train the model from Google Cloud Storage directly.\n\nYou just need to change the path of the images.\n\nHere @tt0721 and me trying to find the best way to do it.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962\n\nhere another discussion about it.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/227062\n\nAll the best\nHappy Learning 😀"
    },
    {
      "id": 1255838,
      "postDate": "2021-03-29T09:01:59.560Z",
      "content": "<p>why not using Google Drive for dataset?</p>",
      "rawMarkdown": "why not using Google Drive for dataset?"
    }
  ],
  "comments": [
    {
      "id": 1251633,
      "author_name": "Sunghyun Jun",
      "author_url": "",
      "post_date": "2021-03-25T00:59:44.210000",
      "content": "<p>GCS egress cost between continents is painful! I've had same experience.</p>\n<p>So nowadays and I'm trying to management data transfer by making buckets per continental.</p>\n<p>And here are my small tips using colab pro.</p>\n<ul>\n<li>For dataset, download from kaggle's private or public dataset using kaggle api(from kaggle to Colab Pro's disk).</li>\n<li>Be careful of saving and loading checkpoints to GCS. Sometimes checkpoints are also large. If you use GCS and Colab Pro on different locations or on different continents, saving checkpoints frequently can cause VERY BIG COST. </li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1243830,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-03-18T14:15:59.983000",
      "content": "<p>I'm a noob w.r.t. Colab Pro (I've only used the free Colab and Kaggle kernels). So I had a question related to this.</p>\n<p>Would it make more sense to store the data as a reduced (remove tfrecords if not using them and separate train and test)(or use only tfrecord files train and test separated) .zip file(s) on Google Drive. Then when you launch Colab Pro you can mount your drive and unzip the files directly into your Colab environment?</p>\n<p>I think this would take approximately 5-10 minutes and there would be no additional transfer costs? You might need to upgrade Google Drive storage… but it should be a negligible cost compared to GCS transfer costs.</p>\n<p>Is there a reason this wouldn't work (or why your current setup is better)? I'm still learning so hopefully you don't take this as a challenge. I'm just trying to identify good approaches for future workflows. TBH I'm not at all familiar with WebDataset so some of my ignorance and/or confusion may be related to using that tool.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1243851,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2021-03-18T14:29:19.150000",
          "content": "<p>I'm a noob as well, so please take my reasoning with a grain of salt:</p>\n<ul>\n<li>The main reason is familiarity -- I used the same pipeline for the Google Landmarks challenge. I know that it's dirt cheap and easy to use and pretty fast, and it was easy for me to code a pipeline.</li>\n<li>Colab Pro has limited disk space, so if I want to use more data (i.e. external data, larger image size, maybe even predictions from other models) -- it would be problematic.</li>\n<li>TFRecordDataset and WebDataset are designed for exactly this type of cloud training, and it's pretty fast. As for WebDataset -- it's so good that they're planning to merge it to PyTorch repository 😄 Here is a post about it <a href=\"https://pytorch.org/blog/efficient-pytorch-io-library-for-large-datasets-many-files-many-gpus/\" target=\"_blank\">https://pytorch.org/blog/efficient-pytorch-io-library-for-large-datasets-many-files-many-gpus/</a></li>\n<li>If I ever find a good solution and want to scale up to larger models, I don't need to change any code when I move to google cloud instances for training.</li>\n<li>With the right pipeline, GCS has no transfer cost as well if you're doing so from within Google's services (the cost I mentioned above for Google Landmarks is mostly because I transferred the data from outside of google services).</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1243868,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-18T14:41:09.227000",
          "content": "<p>That makes a lot of sense. Sounds like a good setup! Thanks for the response and the advice about Colab Pro Instance localization.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1243879,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2021-03-18T14:52:59.010000",
          "content": "<p>If you're using TPUs, you don't have to worry about localization, because all TPU v2-8 are located in <code>us-central</code> 😄😄😄</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1255487,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-03-28T21:00:46.343000",
          "content": "<p>Great tips. </p>\n<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> I used to use Google Drive with free colab earlier. But you see this can be hassle at times.</p>\n<p>Now I use Weights and Biases Artifacts. For free user account they provide 100 GB of artifact storage. Once you log your dataset as an artifacts it's just so much convenient to download it from anywhere be it local system, colab or kaggle. </p>\n<p>Hope it helps. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1261993,
      "author_name": "Faisal Alsrheed",
      "author_url": "",
      "post_date": "2021-04-03T16:56:12.480000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a></p>\n<p>I am just trying to understand what is the reason for downloading the dataset in Google drive or using WebDataset.</p>\n<p>You can train the model from Google Cloud Storage directly.</p>\n<p>You just need to change the path of the images.</p>\n<p>Here <a href=\"https://www.kaggle.com/tt0721\" target=\"_blank\">@tt0721</a> and me trying to find the best way to do it.<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962</a></p>\n<p>here another discussion about it.<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/227062\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/227062</a></p>\n<p>All the best<br>\nHappy Learning 😀</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1255838,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2021-03-29T09:01:59.560000",
      "content": "<p>why not using Google Drive for dataset?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1243293": "Hi fellow kagglers with no home-baked deep learning rig because all GPUs in 2021 are out of stock thanks to miners 😄\n\nI'm using Colab Pro (they offer nice P100 and V100 GPUs, and even TPUs but it's harder to use if you're on PyTorch) and store data in Google Cloud Storage buckets. To improve GPU utilization, I'm using [WebDataset](https://github.com/tmbdev/webdataset). I believe that Colab is enough to get some decent results (for example, [this champ](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037) got 1st place on [Google Landmarks Retrieval 2020](https://www.kaggle.com/c/landmark-retrieval-2020/overview) with just Colab Pro).\n\nUsually it's a very cheap option (for example, in [Google Landmark Recognition 2020](https://www.kaggle.com/c/landmark-recognition-2020) challenge, my data flow for 5 parallel instances was 120GB per hour over 20 days and it costed me less than **$20**).\n\nBut this time, with just 2 days of trying different experiment settings with 3 parallel instances, 11GB per 10 mins each instance, I just got a **$275** bill (which is like... 15 burgers).\n\nI forgot that Colab can give you instances from Asia and Europe (my chance of getting an instance from Taiwan was ~1/3), while my GCS buckets are located in the US. And the data transfer rate is... a whopping **$0.12 per GB!** Yikes!\n\nTo get the information about the location of your Colab instance, put\n```\n!curl ipinfo.io\n```\nin your notebook. Don't overpay for services like me 😄😄😄",
    "1251633": "GCS egress cost between continents is painful! I've had same experience.\n\nSo nowadays and I'm trying to management data transfer by making buckets per continental.\n\nAnd here are my small tips using colab pro.\n- For dataset, download from kaggle's private or public dataset using kaggle api(from kaggle to Colab Pro's disk).\n- Be careful of saving and loading checkpoints to GCS. Sometimes checkpoints are also large. If you use GCS and Colab Pro on different locations or on different continents, saving checkpoints frequently can cause VERY BIG COST. ",
    "1243830": "I'm a noob w.r.t. Colab Pro (I've only used the free Colab and Kaggle kernels). So I had a question related to this.\n\nWould it make more sense to store the data as a reduced (remove tfrecords if not using them and separate train and test)(or use only tfrecord files train and test separated) .zip file(s) on Google Drive. Then when you launch Colab Pro you can mount your drive and unzip the files directly into your Colab environment?\n\nI think this would take approximately 5-10 minutes and there would be no additional transfer costs? You might need to upgrade Google Drive storage... but it should be a negligible cost compared to GCS transfer costs.\n\nIs there a reason this wouldn't work (or why your current setup is better)? I'm still learning so hopefully you don't take this as a challenge. I'm just trying to identify good approaches for future workflows. TBH I'm not at all familiar with WebDataset so some of my ignorance and/or confusion may be related to using that tool.",
    "1261993": "Hi @chankhavu\n\nI am just trying to understand what is the reason for downloading the dataset in Google drive or using WebDataset.\n\nYou can train the model from Google Cloud Storage directly.\n\nYou just need to change the path of the images.\n\nHere @tt0721 and me trying to find the best way to do it.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962\n\nhere another discussion about it.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/227062\n\nAll the best\nHappy Learning 😀",
    "1255838": "why not using Google Drive for dataset?"
  }
}