{
  "id": 227062,
  "title": "Dataset too large for Google Colab",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/227062",
  "author_name": "Muiz Alvi",
  "post_date": "2021-03-18T20:05:17.247000",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hey guys, I'm a new Kaggler and I use Google Colab to import datasets to my Google Drive, however Google Colab only allows a total of 107 GB of data to be downloaded and this dataset is about 150 GB. I also lack local disk space so downloading the data set locally is not an option. How should I go about downloading datasets?</p>",
  "messages": [
    {
      "id": 1244426,
      "postDate": "2021-03-19T01:32:44.837Z",
      "content": "<p>Well the dataset is about 150 GB but it contains both TF records and png files. You probably don't need both of them so the solution would be to download only files you need (either TF records or png files depending on your pipeline).</p>\n<p>I couldn't find a way how to download only a specific folder with Kaggle API or multiple files using wildcard but if you are downloading TF records you can probably just download the files one by one (see the <a href=\"https://github.com/Kaggle/kaggle-api#download-competition-files\" target=\"_blank\">Kaggle API github</a>)</p>\n<p>Or you can connect from your colab notebook to the competition Google Cloud Storage directly and download anything you need. To do so you have to first run this in the competition notebook here on kaggle:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('hpa-single-cell-image-classification')\nprint (GCS_DS_PATH)\n</code></pre>\n<p>which will print out something like this<br>\n<code>gs://kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab</code><br>\n(the path does change from time to time so if it stops working in your colab just rerun the code in your kaggle notebook)</p>\n<p>and then in your colab you can connect to the GCS and download what you want</p>\n<pre><code>!mkdir /home/data/train\n!mkdir /home/data/test\n\n!pip install gcsfs \nimport gcsfs\n\nfs = gcsfs.GCSFileSystem(project='my-project')\ndata_gcs = 'kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab'\nfs.get(f\"{data_gcs}/train/*.png\", '/home/data/train/')\nfs.get(f\"{data_gcs}/test/*.png\", '/home/data/test/')\n</code></pre>\n<p>I am not in this competition but i successfully used this in a different one so it should work just fine.<br>\nHope it helps. Good luck</p>",
      "rawMarkdown": "Well the dataset is about 150 GB but it contains both TF records and png files. You probably don't need both of them so the solution would be to download only files you need (either TF records or png files depending on your pipeline).\n\nI couldn't find a way how to download only a specific folder with Kaggle API or multiple files using wildcard but if you are downloading TF records you can probably just download the files one by one (see the [Kaggle API github](https://github.com/Kaggle/kaggle-api#download-competition-files))\n\nOr you can connect from your colab notebook to the competition Google Cloud Storage directly and download anything you need. To do so you have to first run this in the competition notebook here on kaggle:\n```\nfrom kaggle_datasets import KaggleDatasets\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('hpa-single-cell-image-classification')\nprint (GCS_DS_PATH)\n```\n\nwhich will print out something like this\n`gs://kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab`\n(the path does change from time to time so if it stops working in your colab just rerun the code in your kaggle notebook)\n\nand then in your colab you can connect to the GCS and download what you want\n```\n!mkdir /home/data/train\n!mkdir /home/data/test\n\n!pip install gcsfs \nimport gcsfs\n\nfs = gcsfs.GCSFileSystem(project='my-project')\ndata_gcs = 'kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab'\nfs.get(f\"{data_gcs}/train/*.png\", '/home/data/train/')\nfs.get(f\"{data_gcs}/test/*.png\", '/home/data/test/')\n```\n\nI am not in this competition but i successfully used this in a different one so it should work just fine.\nHope it helps. Good luck",
      "votes": 7,
      "replies": [
        {
          "id": 1261987,
          "postDate": "2021-04-03T16:49:39.630Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/michaln\" target=\"_blank\">@michaln</a> and <a href=\"https://www.kaggle.com/mgurevich\" target=\"_blank\">@mgurevich</a>, </p>\n<p>Thank you for sharing your steps </p>\n<p>I am just trying to understand what is the reason for downloading the dataset in Google drive.</p>\n<p>You can train the model from Google Cloud Storage directly. </p>\n<p>You just need to change the path of the images.  </p>\n<p>Here  <a href=\"https://www.kaggle.com/tt0721\" target=\"_blank\">@tt0721</a>  and me trying to find the best way to do it. <br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962</a></p>\n<p>All the best <br>\nHappy Learning 😀</p>",
          "rawMarkdown": "Hi @michaln and @mgurevich, \n\nThank you for sharing your steps \n\nI am just trying to understand what is the reason for downloading the dataset in Google drive.\n\nYou can train the model from Google Cloud Storage directly. \n\nYou just need to change the path of the images.  \n\nHere  @tt0721  and me trying to find the best way to do it. \nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962\n\nAll the best \nHappy Learning 😀",
          "votes": 1
        },
        {
          "id": 1262252,
          "postDate": "2021-04-04T02:46:51.763Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/faisalalsrheed\" target=\"_blank\">@faisalalsrheed</a>,</p>\n<p>I am no expert in TF/PyTorch, GCP/GCS so you can correct me if i am wrong :-)</p>\n<p>i definitely agree that if you are using Tensorflow you don't need to download the dataset at all and you can just access the data from GCS directly. I am not sure thought that the same works for PyTorch (i am not aware that there is a direct way and quick google search didn't give me a solution :-D ).</p>\n<p>Another reason that comes to my mind is that you might want to preprocess the data in some way and loading the files directly from GCS might not be supported by the library you want to use. Typically you would have to download the file from GCS anyway or read it as a bytes and then give it to the library. </p>\n<p>I believe recently i saw a topic here on Kaggle that there are some fees in case you are accessing GCS that is in different region then your server on GCP. If you are using colab then this does not concern you but if you use GCP then you should keep that in mind. (no idea what the fees are, when they are charged and if they would be charged in this case, i might be totally wrong again here).</p>\n<p>I would say in general the Tensorflow makes it nice and simple so if you want to use the original data and your pipeline is in TF then there might be no reason to download the dataset at all. If you are not using TF though then downloading it locally might be your best choice. Especially if you would read the data repeatedly as accessing the data from GCS might be much slower (as far as i know TF prefetches the data in parallel to avoid possible speed issues).</p>\n<p>In my case when using PyTorch i open images using cv2 or PIL and neither of them can read data from GCS directly. I would have to connect to the bucket, read the file as bytes and pass the bytes to the library. I am too lazy for that and i don't want to deal with GCS at all so i rather download the data. If there is a better and more efficient way I would love to know :-D</p>",
          "rawMarkdown": "Hi @faisalalsrheed,\n\nI am no expert in TF/PyTorch, GCP/GCS so you can correct me if i am wrong :-)\n\ni definitely agree that if you are using Tensorflow you don't need to download the dataset at all and you can just access the data from GCS directly. I am not sure thought that the same works for PyTorch (i am not aware that there is a direct way and quick google search didn't give me a solution :-D ).\n\nAnother reason that comes to my mind is that you might want to preprocess the data in some way and loading the files directly from GCS might not be supported by the library you want to use. Typically you would have to download the file from GCS anyway or read it as a bytes and then give it to the library. \n\nI believe recently i saw a topic here on Kaggle that there are some fees in case you are accessing GCS that is in different region then your server on GCP. If you are using colab then this does not concern you but if you use GCP then you should keep that in mind. (no idea what the fees are, when they are charged and if they would be charged in this case, i might be totally wrong again here).\n\nI would say in general the Tensorflow makes it nice and simple so if you want to use the original data and your pipeline is in TF then there might be no reason to download the dataset at all. If you are not using TF though then downloading it locally might be your best choice. Especially if you would read the data repeatedly as accessing the data from GCS might be much slower (as far as i know TF prefetches the data in parallel to avoid possible speed issues).\n\nIn my case when using PyTorch i open images using cv2 or PIL and neither of them can read data from GCS directly. I would have to connect to the bucket, read the file as bytes and pass the bytes to the library. I am too lazy for that and i don't want to deal with GCS at all so i rather download the data. If there is a better and more efficient way I would love to know :-D",
          "votes": 1
        },
        {
          "id": 1262682,
          "postDate": "2021-04-04T15:39:18.763Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/michaln\" target=\"_blank\">@michaln</a>, <br>\nI see the point now.<br>\nYou raised some great questions. I will keep looking for the answers.<br>\nall the best :)</p>",
          "rawMarkdown": "Thank you @michaln, \nI see the point now.\nYou raised some great questions. I will keep looking for the answers.\nall the best :)"
        }
      ]
    },
    {
      "id": 1244762,
      "postDate": "2021-03-19T08:01:15.887Z",
      "content": "<p>Hello! I ended up with following method:<br>\n1) Download dataset to a local machine<br>\n2) Remove all unnecessary data (in my case it was tfrecords)<br>\n3) Resize images to a size suitable for your needs (2048x2048 or 1024x1024 or even less)<br>\n4) Make a zip archive with all these data<br>\n5) Buy some space on google drive. Also buy colab pro (after that you will have 150Gb on a colab machine)<br>\n6) mount gdrive to your colab machine and unzip the archive:</p>\n<pre><code>from google.colab import drive\ndrive.mount('/content/drive')\n!unzip -q /content/drive/MyDrive/kaggle/train_1024_1024.zip -d data\n</code></pre>\n<p>in my case unzip took about 20 minutes. After that you can train your models on colab</p>",
      "rawMarkdown": "Hello! I ended up with following method:\n1) Download dataset to a local machine\n2) Remove all unnecessary data (in my case it was tfrecords)\n3) Resize images to a size suitable for your needs (2048x2048 or 1024x1024 or even less)\n4) Make a zip archive with all these data\n5) Buy some space on google drive. Also buy colab pro (after that you will have 150Gb on a colab machine)\n6) mount gdrive to your colab machine and unzip the archive:\n```\nfrom google.colab import drive\ndrive.mount('/content/drive')\n!unzip -q /content/drive/MyDrive/kaggle/train_1024_1024.zip -d data\n```\nin my case unzip took about 20 minutes. After that you can train your models on colab",
      "votes": 3,
      "replies": [
        {
          "id": 1261100,
          "postDate": "2021-04-02T17:34:42.690Z",
          "content": "<p>Hello! Are you training model without problem on google drive and colab?</p>\n<p>When I try to train model using cell-wise images, I get a drive error when loading the image.(I believe this is caused by having too many images.)<br>\nI would like to know how to solve it if you know. (Or you don’t use cell-wise images?)</p>\n<p>Thanks!</p>",
          "rawMarkdown": "Hello! Are you training model without problem on google drive and colab?\n\nWhen I try to train model using cell-wise images, I get a drive error when loading the image.(I believe this is caused by having too many images.)\nI would like to know how to solve it if you know. (Or you don’t use cell-wise images?)\n\nThanks!"
        },
        {
          "id": 1263157,
          "postDate": "2021-04-05T06:14:14.760Z",
          "content": "<p>Hello, I had a problem with google drive when I tried to download each image separately - I think in that way you exceed qouta pretty fast. </p>\n<p>Try following approach:<br>\n1) zip all you images<br>\n2) put archive on gdrive<br>\n3) download archive to your colab instance<br>\n4) extract images to the folder at colab instance (not gdrive)</p>",
          "rawMarkdown": "Hello, I had a problem with google drive when I tried to download each image separately - I think in that way you exceed qouta pretty fast. \n\nTry following approach:\n1) zip all you images\n2) put archive on gdrive\n3) download archive to your colab instance\n4) extract images to the folder at colab instance (not gdrive)\n\n"
        },
        {
          "id": 1263338,
          "postDate": "2021-04-05T10:28:10.500Z",
          "content": "<p>Thank you so much!</p>",
          "rawMarkdown": "Thank you so much!"
        },
        {
          "id": 1263341,
          "postDate": "2021-04-05T10:29:48.023Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mgurevich\" target=\"_blank\">@mgurevich</a> </p>\n<p>Thnak you for sharing.</p>\n<pre><code>You can access files in Drive in a number of ways, including:\nMounting your Google Drive in the runtime's virtual machine\nUsing a wrapper around the API such as PyDrive\nUsing the native REST API\n</code></pre>\n<p><a href=\"https://colab.research.google.com/notebooks/io.ipynb\" target=\"_blank\">https://colab.research.google.com/notebooks/io.ipynb</a></p>\n<p>Which one did you use?</p>",
          "rawMarkdown": "Hi @mgurevich \n\nThnak you for sharing.\n\n```\nYou can access files in Drive in a number of ways, including:\nMounting your Google Drive in the runtime's virtual machine\nUsing a wrapper around the API such as PyDrive\nUsing the native REST API\n```\nhttps://colab.research.google.com/notebooks/io.ipynb\n\nWhich one did you use?"
        },
        {
          "id": 1263433,
          "postDate": "2021-04-05T12:13:27.680Z",
          "content": "<p>Hello, I mounted Google Drive using following snippet:</p>\n<pre><code>from google.colab import drive\ndrive.mount('/content/drive', force_remount=True)\n</code></pre>",
          "rawMarkdown": "Hello, I mounted Google Drive using following snippet:\n```\nfrom google.colab import drive\ndrive.mount('/content/drive', force_remount=True)\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 1244240,
      "postDate": "2021-03-18T20:05:17.247Z",
      "content": "<p>Hey guys, I'm a new Kaggler and I use Google Colab to import datasets to my Google Drive, however Google Colab only allows a total of 107 GB of data to be downloaded and this dataset is about 150 GB. I also lack local disk space so downloading the data set locally is not an option. How should I go about downloading datasets?</p>",
      "rawMarkdown": "Hey guys, I'm a new Kaggler and I use Google Colab to import datasets to my Google Drive, however Google Colab only allows a total of 107 GB of data to be downloaded and this dataset is about 150 GB. I also lack local disk space so downloading the data set locally is not an option. How should I go about downloading datasets?",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1244426,
      "author_name": "Michal",
      "author_url": "",
      "post_date": "2021-03-19T01:32:44.837000",
      "content": "<p>Well the dataset is about 150 GB but it contains both TF records and png files. You probably don't need both of them so the solution would be to download only files you need (either TF records or png files depending on your pipeline).</p>\n<p>I couldn't find a way how to download only a specific folder with Kaggle API or multiple files using wildcard but if you are downloading TF records you can probably just download the files one by one (see the <a href=\"https://github.com/Kaggle/kaggle-api#download-competition-files\" target=\"_blank\">Kaggle API github</a>)</p>\n<p>Or you can connect from your colab notebook to the competition Google Cloud Storage directly and download anything you need. To do so you have to first run this in the competition notebook here on kaggle:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('hpa-single-cell-image-classification')\nprint (GCS_DS_PATH)\n</code></pre>\n<p>which will print out something like this<br>\n<code>gs://kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab</code><br>\n(the path does change from time to time so if it stops working in your colab just rerun the code in your kaggle notebook)</p>\n<p>and then in your colab you can connect to the GCS and download what you want</p>\n<pre><code>!mkdir /home/data/train\n!mkdir /home/data/test\n\n!pip install gcsfs \nimport gcsfs\n\nfs = gcsfs.GCSFileSystem(project='my-project')\ndata_gcs = 'kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab'\nfs.get(f\"{data_gcs}/train/*.png\", '/home/data/train/')\nfs.get(f\"{data_gcs}/test/*.png\", '/home/data/test/')\n</code></pre>\n<p>I am not in this competition but i successfully used this in a different one so it should work just fine.<br>\nHope it helps. Good luck</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1261987,
          "author_name": "Faisal Alsrheed",
          "author_url": "",
          "post_date": "2021-04-03T16:49:39.630000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/michaln\" target=\"_blank\">@michaln</a> and <a href=\"https://www.kaggle.com/mgurevich\" target=\"_blank\">@mgurevich</a>, </p>\n<p>Thank you for sharing your steps </p>\n<p>I am just trying to understand what is the reason for downloading the dataset in Google drive.</p>\n<p>You can train the model from Google Cloud Storage directly. </p>\n<p>You just need to change the path of the images.  </p>\n<p>Here  <a href=\"https://www.kaggle.com/tt0721\" target=\"_blank\">@tt0721</a>  and me trying to find the best way to do it. <br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229962</a></p>\n<p>All the best <br>\nHappy Learning 😀</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1262252,
          "author_name": "Michal",
          "author_url": "",
          "post_date": "2021-04-04T02:46:51.763000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/faisalalsrheed\" target=\"_blank\">@faisalalsrheed</a>,</p>\n<p>I am no expert in TF/PyTorch, GCP/GCS so you can correct me if i am wrong :-)</p>\n<p>i definitely agree that if you are using Tensorflow you don't need to download the dataset at all and you can just access the data from GCS directly. I am not sure thought that the same works for PyTorch (i am not aware that there is a direct way and quick google search didn't give me a solution :-D ).</p>\n<p>Another reason that comes to my mind is that you might want to preprocess the data in some way and loading the files directly from GCS might not be supported by the library you want to use. Typically you would have to download the file from GCS anyway or read it as a bytes and then give it to the library. </p>\n<p>I believe recently i saw a topic here on Kaggle that there are some fees in case you are accessing GCS that is in different region then your server on GCP. If you are using colab then this does not concern you but if you use GCP then you should keep that in mind. (no idea what the fees are, when they are charged and if they would be charged in this case, i might be totally wrong again here).</p>\n<p>I would say in general the Tensorflow makes it nice and simple so if you want to use the original data and your pipeline is in TF then there might be no reason to download the dataset at all. If you are not using TF though then downloading it locally might be your best choice. Especially if you would read the data repeatedly as accessing the data from GCS might be much slower (as far as i know TF prefetches the data in parallel to avoid possible speed issues).</p>\n<p>In my case when using PyTorch i open images using cv2 or PIL and neither of them can read data from GCS directly. I would have to connect to the bucket, read the file as bytes and pass the bytes to the library. I am too lazy for that and i don't want to deal with GCS at all so i rather download the data. If there is a better and more efficient way I would love to know :-D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1262682,
          "author_name": "Faisal Alsrheed",
          "author_url": "",
          "post_date": "2021-04-04T15:39:18.763000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/michaln\" target=\"_blank\">@michaln</a>, <br>\nI see the point now.<br>\nYou raised some great questions. I will keep looking for the answers.<br>\nall the best :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1244762,
      "author_name": "Mikhail Gurevich",
      "author_url": "",
      "post_date": "2021-03-19T08:01:15.887000",
      "content": "<p>Hello! I ended up with following method:<br>\n1) Download dataset to a local machine<br>\n2) Remove all unnecessary data (in my case it was tfrecords)<br>\n3) Resize images to a size suitable for your needs (2048x2048 or 1024x1024 or even less)<br>\n4) Make a zip archive with all these data<br>\n5) Buy some space on google drive. Also buy colab pro (after that you will have 150Gb on a colab machine)<br>\n6) mount gdrive to your colab machine and unzip the archive:</p>\n<pre><code>from google.colab import drive\ndrive.mount('/content/drive')\n!unzip -q /content/drive/MyDrive/kaggle/train_1024_1024.zip -d data\n</code></pre>\n<p>in my case unzip took about 20 minutes. After that you can train your models on colab</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1261100,
          "author_name": "tsujino",
          "author_url": "",
          "post_date": "2021-04-02T17:34:42.690000",
          "content": "<p>Hello! Are you training model without problem on google drive and colab?</p>\n<p>When I try to train model using cell-wise images, I get a drive error when loading the image.(I believe this is caused by having too many images.)<br>\nI would like to know how to solve it if you know. (Or you don’t use cell-wise images?)</p>\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1263157,
          "author_name": "Mikhail Gurevich",
          "author_url": "",
          "post_date": "2021-04-05T06:14:14.760000",
          "content": "<p>Hello, I had a problem with google drive when I tried to download each image separately - I think in that way you exceed qouta pretty fast. </p>\n<p>Try following approach:<br>\n1) zip all you images<br>\n2) put archive on gdrive<br>\n3) download archive to your colab instance<br>\n4) extract images to the folder at colab instance (not gdrive)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1263338,
          "author_name": "tsujino",
          "author_url": "",
          "post_date": "2021-04-05T10:28:10.500000",
          "content": "<p>Thank you so much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1263341,
          "author_name": "Faisal Alsrheed",
          "author_url": "",
          "post_date": "2021-04-05T10:29:48.023000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mgurevich\" target=\"_blank\">@mgurevich</a> </p>\n<p>Thnak you for sharing.</p>\n<pre><code>You can access files in Drive in a number of ways, including:\nMounting your Google Drive in the runtime's virtual machine\nUsing a wrapper around the API such as PyDrive\nUsing the native REST API\n</code></pre>\n<p><a href=\"https://colab.research.google.com/notebooks/io.ipynb\" target=\"_blank\">https://colab.research.google.com/notebooks/io.ipynb</a></p>\n<p>Which one did you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1263433,
          "author_name": "Mikhail Gurevich",
          "author_url": "",
          "post_date": "2021-04-05T12:13:27.680000",
          "content": "<p>Hello, I mounted Google Drive using following snippet:</p>\n<pre><code>from google.colab import drive\ndrive.mount('/content/drive', force_remount=True)\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1244426": "Well the dataset is about 150 GB but it contains both TF records and png files. You probably don't need both of them so the solution would be to download only files you need (either TF records or png files depending on your pipeline).\n\nI couldn't find a way how to download only a specific folder with Kaggle API or multiple files using wildcard but if you are downloading TF records you can probably just download the files one by one (see the [Kaggle API github](https://github.com/Kaggle/kaggle-api#download-competition-files))\n\nOr you can connect from your colab notebook to the competition Google Cloud Storage directly and download anything you need. To do so you have to first run this in the competition notebook here on kaggle:\n```\nfrom kaggle_datasets import KaggleDatasets\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('hpa-single-cell-image-classification')\nprint (GCS_DS_PATH)\n```\n\nwhich will print out something like this\n`gs://kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab`\n(the path does change from time to time so if it stops working in your colab just rerun the code in your kaggle notebook)\n\nand then in your colab you can connect to the GCS and download what you want\n```\n!mkdir /home/data/train\n!mkdir /home/data/test\n\n!pip install gcsfs \nimport gcsfs\n\nfs = gcsfs.GCSFileSystem(project='my-project')\ndata_gcs = 'kds-a7dd8b09d52950714bee1818d843af117accc125d4d289d9afcebfab'\nfs.get(f\"{data_gcs}/train/*.png\", '/home/data/train/')\nfs.get(f\"{data_gcs}/test/*.png\", '/home/data/test/')\n```\n\nI am not in this competition but i successfully used this in a different one so it should work just fine.\nHope it helps. Good luck",
    "1244762": "Hello! I ended up with following method:\n1) Download dataset to a local machine\n2) Remove all unnecessary data (in my case it was tfrecords)\n3) Resize images to a size suitable for your needs (2048x2048 or 1024x1024 or even less)\n4) Make a zip archive with all these data\n5) Buy some space on google drive. Also buy colab pro (after that you will have 150Gb on a colab machine)\n6) mount gdrive to your colab machine and unzip the archive:\n```\nfrom google.colab import drive\ndrive.mount('/content/drive')\n!unzip -q /content/drive/MyDrive/kaggle/train_1024_1024.zip -d data\n```\nin my case unzip took about 20 minutes. After that you can train your models on colab",
    "1244240": "Hey guys, I'm a new Kaggler and I use Google Colab to import datasets to my Google Drive, however Google Colab only allows a total of 107 GB of data to be downloaded and this dataset is about 150 GB. I also lack local disk space so downloading the data set locally is not an option. How should I go about downloading datasets?"
  }
}