{
  "id": 168020,
  "title": "how to work with this competition dataset on google colab?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168020",
  "author_name": "",
  "post_date": "2020-07-18T21:12:06.829668900Z",
  "votes": 2,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I was starting with small images like 32x32 and 64x64. With them I could put zip file with all the files on my google drive then unzip them in colab vm. That was working.</p>\n\n<p>But with 256x256 unziping is very long process. And I need it to perform it every time I start colab.</p>\n\n<p>I am searching for alternatives:\n- put zip file on google drive and use explode to mount zip as folder - works, but very slow (hours per epoch)\n- put uncompressed files into google drive - still trying, it will take hours to upload (single zip is uploading quickly)</p>\n\n<p>Do you know quick way to upload 30000 images into google drive? Maybe there is some way to use files from Kaggle dataset on colab without waiting for unzip?</p>",
  "messages": [
    {
      "id": "934886",
      "postDate": "07/18/2020 21:12:06",
      "content": "<p>I was starting with small images like 32x32 and 64x64. With them I could put zip file with all the files on my google drive then unzip them in colab vm. That was working.</p>\n\n<p>But with 256x256 unziping is very long process. And I need it to perform it every time I start colab.</p>\n\n<p>I am searching for alternatives:\n- put zip file on google drive and use explode to mount zip as folder - works, but very slow (hours per epoch)\n- put uncompressed files into google drive - still trying, it will take hours to upload (single zip is uploading quickly)</p>\n\n<p>Do you know quick way to upload 30000 images into google drive? Maybe there is some way to use files from Kaggle dataset on colab without waiting for unzip?</p>",
      "rawMarkdown": "I was starting with small images like 32x32 and 64x64. With them I could put zip file with all the files on my google drive then unzip them in colab vm. That was working.\n\nBut with 256x256 unziping is very long process. And I need it to perform it every time I start colab.\n\nI am searching for alternatives:\n- put zip file on google drive and use explode to mount zip as folder - works, but very slow (hours per epoch)\n- put uncompressed files into google drive - still trying, it will take hours to upload (single zip is uploading quickly)\n\nDo you know quick way to upload 30000 images into google drive? Maybe there is some way to use files from Kaggle dataset on colab without waiting for unzip?",
      "votes": null
    },
    {
      "id": "934903",
      "postDate": "07/18/2020 21:57:56",
      "content": "<p>As long as the dataset is public, you can access it as describe in this notebook by retrieving the gcs path : \n<a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses</a></p>\n\n<p>i don't know the access time from gdrive to colab notebook but i don't think it's better than accessing a gcs link.\nThen, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526</a> provides many datasets links in tfrecords and jpeg with many resolution.</p>",
      "rawMarkdown": "As long as the dataset is public, you can access it as describe in this notebook by retrieving the gcs path : \nhttps://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\n\ni don't know the access time from gdrive to colab notebook but i don't think it's better than accessing a gcs link.\nThen, https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526 provides many datasets links in tfrecords and jpeg with many resolution.",
      "votes": null
    },
    {
      "id": "934912",
      "postDate": "07/18/2020 22:28:05",
      "content": "<p>Yes, this is the right way to do it. And here is another notebook for looking up GCS paths for various public data sets:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">SIIM Show GCS bucket addresses</a></p>",
      "rawMarkdown": "Yes, this is the right way to do it. And here is another notebook for looking up GCS paths for various public data sets:\n\n[SIIM Show GCS bucket addresses](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses)",
      "votes": null
    },
    {
      "id": "934915",
      "postDate": "07/18/2020 22:34:16",
      "content": "<p>what if dataset is private?</p>",
      "rawMarkdown": "what if dataset is private?",
      "votes": null
    },
    {
      "id": "934920",
      "postDate": "07/18/2020 22:45:48",
      "content": "<p>I have never tried it with a private data set -- all my data sets are public. </p>",
      "rawMarkdown": "I have never tried it with a private data set -- all my data sets are public.",
      "votes": null
    },
    {
      "id": "934924",
      "postDate": "07/18/2020 22:52:23",
      "content": "<p>you can't get the gcs link. Another alternative is to create a google cloud account but there's some limitation as long as you don't pay.</p>\n\n<p>Btw, when you get this working, would you mind to tell me how long an epoch takes on pytorch (number of images / resolution / batch size) ? When i tried 256x256 with augmentation on the fly, a epoch of 60000 images with bs of 64 takes like 20 minutes...</p>\n\n<p>I stay with the fact that, with Pytorch, the data is loaded from notebook to TPU and so bandwidth should be the issue. This is not the case with tensorflow as we use tfrecords on gcs (but augmentations are limited to tensor operations).</p>",
      "rawMarkdown": "you can't get the gcs link. Another alternative is to create a google cloud account but there's some limitation as long as you don't pay.\n\nBtw, when you get this working, would you mind to tell me how long an epoch takes on pytorch (number of images / resolution / batch size) ? When i tried 256x256 with augmentation on the fly, a epoch of 60000 images with bs of 64 takes like 20 minutes...\n\nI stay with the fact that, with Pytorch, the data is loaded from notebook to TPU and so bandwidth should be the issue. This is not the case with tensorflow as we use tfrecords on gcs (but augmentations are limited to tensor operations).",
      "votes": null
    },
    {
      "id": "934930",
      "postDate": "07/18/2020 23:05:07",
      "content": "<p>256x256 PyTorch times:</p>\n\n<p>(please note this highly depends on augmentations and disk/CPU performance)</p>\n\n<ul>\n<li>on my local GPU with batch_size 30 it is 473.547519 (efficientnet-b1)</li>\n<li>on colab on GPU with batch_size 16 it is 593.404905 (efficientnet-b0)</li>\n<li>on colab on TPU with batch_size 16 it is 383.63376 (efficientnet-b0)</li>\n</ul>\n\n<p>I think I may have problem with this competition with finishing the training, as I can't do much with my local GPU (tried training b6 it is like half hour per epoch)</p>",
      "rawMarkdown": "256x256 PyTorch times:\n\n(please note this highly depends on augmentations and disk/CPU performance)\n\n- on my local GPU with batch_size 30 it is 473.547519 (efficientnet-b1)\n- on colab on GPU with batch_size 16 it is 593.404905 (efficientnet-b0)\n- on colab on TPU with batch_size 16 it is 383.63376 (efficientnet-b0)\n\nI think I may have problem with this competition with finishing the training, as I can't do much with my local GPU (tried training b6 it is like half hour per epoch)",
      "votes": null
    },
    {
      "id": "934941",
      "postDate": "07/18/2020 23:57:30",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>, for GPU you can download your dataset via the Kaggle API into colab and use it. </p>\n\n<p>However, TPU is faster for training and you can get the TFRecord data's bucket address as described <a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">here</a> at the time of running to use directly in your colab notebook. It is really fast.</p>",
      "rawMarkdown": "jacekpoplawski, for GPU you can download your dataset via the Kaggle API into colab and use it. \n\nHowever, TPU is faster for training and you can get the TFRecord data's bucket address as described [here](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses) at the time of running to use directly in your colab notebook. It is really fast.",
      "votes": null
    },
    {
      "id": "934942",
      "postDate": "07/19/2020 00:01:38",
      "content": "<p>I don't use tfrecords, I don't use tensorflow, I am quite familar with data now, what I am not familar with is google drive and colab for so many files, my current workaround is to decompress zip every time I start colab notebook</p>",
      "rawMarkdown": "I don't use tfrecords, I don't use tensorflow, I am quite familar with data now, what I am not familar with is google drive and colab for so many files, my current workaround is to decompress zip every time I start colab notebook",
      "votes": null
    },
    {
      "id": "934953",
      "postDate": "07/19/2020 00:17:48",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>, reading from google drive during training is definatley slow. The best way I believe is to have your data as a kaggle dataset and download it to the colab environment in your notebook before you start training. Downloading via the Kaggle Api is relatively fast.</p>",
      "rawMarkdown": "jacekpoplawski, reading from google drive during training is definatley slow. The best way I believe is to have your data as a kaggle dataset and download it to the colab environment in your notebook before you start training. Downloading via the Kaggle Api is relatively fast.",
      "votes": null
    },
    {
      "id": "935018",
      "postDate": "07/19/2020 02:49:59",
      "content": "<p>can I access private dataset with Kaggle Api?</p>",
      "rawMarkdown": "can I access private dataset with Kaggle Api?",
      "votes": null
    },
    {
      "id": "935133",
      "postDate": "07/19/2020 05:19:21",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  You can now use <a href=\"https://www.kaggle.com/product-feedback/163416\">TPUs with Private Datasets</a></p>",
      "rawMarkdown": "jacekpoplawski  You can now use [TPUs with Private Datasets](https://www.kaggle.com/product-feedback/163416)",
      "votes": null
    },
    {
      "id": "935304",
      "postDate": "07/19/2020 08:49:25",
      "content": "<p>With me,\n- Create a notebook \n- Add your dataset which you want read from the colab. \n- Use print ex:\n<code>print(KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images'))</code>\n to get address gs from that notebook.\n- Copy this address to your colab. </p>\n\n<p>Sometime the address will expried, you print to get new address.</p>",
      "rawMarkdown": "With me,\n- Create a notebook \n- Add your dataset which you want read from the colab. \n- Use print ex:\n`print(KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images'))`\n to get address gs from that notebook.\n- Copy this address to your colab. \n\nSometime the address will expried, you print to get new address.",
      "votes": null
    },
    {
      "id": "936753",
      "postDate": "07/20/2020 13:51:04",
      "content": "<p>This might be of help to you</p>\n\n<p><a href=\"https://towardsdatascience.com/setting-up-kaggle-in-google-colab-ebb281b61463\">https://towardsdatascience.com/setting-up-kaggle-in-google-colab-ebb281b61463</a> </p>",
      "rawMarkdown": "This might be of help to you\n\nhttps://towardsdatascience.com/setting-up-kaggle-in-google-colab-ebb281b61463",
      "votes": null
    },
    {
      "id": "939842",
      "postDate": "07/22/2020 14:07:36",
      "content": "<p>The safest way is to save the zip into your google drive and unzip.</p>\n\n<p>import zipfile<br>\nfrom google.colab import drive<br>\ndrive.mount('/content/drive/')<br>\nzip_ref = zipfile.ZipFile(\"/content/drive/My Drive/ML/DataSet.zip\", 'r')<br>\nzip_ref.extractall(\"/tmp\")<br>\nzip_ref.close()</p>",
      "rawMarkdown": "The safest way is to save the zip into your google drive and unzip.\n\nimport zipfile<br>\nfrom google.colab import drive<br>\ndrive.mount('/content/drive/')<br>\nzip_ref = zipfile.ZipFile(\"/content/drive/My Drive/ML/DataSet.zip\", 'r')<br>\nzip_ref.extractall(\"/tmp\")<br>\nzip_ref.close()",
      "votes": null
    },
    {
      "id": "949468",
      "postDate": "07/28/2020 16:48:10",
      "content": "<p><a href=\"/alincijov\">@alincijov</a>  Your approach proved to be not only the safest but the fastest way for me. \nI didn't manage to befriend Kaggle gcs paths and my custom Pytorch data loader. My data loader, namely PIL Image.open() command doesn't see files in subfolder \\train and returns 'file not found' error. I understand this error has something to do with permissions and admit it could be resolved easily. I just don't know how.</p>\n\n<p>import zipfile\nfrom google.colab import drive\ndrive.mount('/content/drive')</p>\n\n<p>zipref = zipfile.ZipFile(\"/content/drive/My Drive/data/jpeg-melanoma-256x256.zip\", 'r')\nzipref.extractall(\"/content/jpeg-melanoma-256x256\")\nzipref.close()</p>\n\n<p>What's good about this approach:\n- you don't spend hours to upload unzipped data to your google drive\n- unzip operation works pretty fast\n- we unzip dataset not to a google drive but to a local storage of VM. I am not aware of VM storage topology but assume this step makes our data \"closer\" than keeping it in a google drive.</p>\n\n<p>one epoch, batch=32, simple B1 model ~240 sec</p>\n\n<p>Many thanks!</p>",
      "rawMarkdown": "alincijov  Your approach proved to be not only the safest but the fastest way for me. \nI didn't manage to befriend Kaggle gcs paths and my custom Pytorch data loader. My data loader, namely PIL Image.open() command doesn't see files in subfolder \\train and returns 'file not found' error. I understand this error has something to do with permissions and admit it could be resolved easily. I just don't know how.\n\nimport zipfile\nfrom google.colab import drive\ndrive.mount('/content/drive')\n\nzipref = zipfile.ZipFile(\"/content/drive/My Drive/data/jpeg-melanoma-256x256.zip\", 'r')\nzipref.extractall(\"/content/jpeg-melanoma-256x256\")\nzipref.close()\n\nWhat's good about this approach:\n- you don't spend hours to upload unzipped data to your google drive\n- unzip operation works pretty fast\n- we unzip dataset not to a google drive but to a local storage of VM. I am not aware of VM storage topology but assume this step makes our data \"closer\" than keeping it in a google drive.\n\none epoch, batch=32, simple B1 model ~240 sec\n\nMany thanks!",
      "votes": null
    },
    {
      "id": "949573",
      "postDate": "07/28/2020 18:07:03",
      "content": "<p>Let me share a snippet, as I just tried to make it work myself recently\n<code>from google.colab import drive</code>\n<code>drive.mount(\"/content/drive\")</code>\n<code>os.environ['KAGGLE_CONFIG_DIR'] = \"/content/drive/My Drive/Kaggle\"</code>\nMake sure kaggle.json (API-key) is in this folder\n<code># Cris dataset for example</code>\n<code>!kaggle datasets download -d cdeotte/jpeg-melanoma-384x384</code>\n<code>!unzip jpeg-melanoma-384x384.zip -d /content/data/jpeg384</code></p>\n\n<p>Worked for me :) Hope, it helps </p>",
      "rawMarkdown": "Let me share a snippet, as I just tried to make it work myself recently\n`from google.colab import drive`\n`drive.mount(\"/content/drive\")`\n`os.environ['KAGGLE_CONFIG_DIR'] = \"/content/drive/My Drive/Kaggle\"`\nMake sure kaggle.json (API-key) is in this folder\n`# Cris dataset for example`\n`!kaggle datasets download -d cdeotte/jpeg-melanoma-384x384`\n`!unzip jpeg-melanoma-384x384.zip -d /content/data/jpeg384`\n\nWorked for me :) Hope, it helps",
      "votes": null
    },
    {
      "id": "949619",
      "postDate": "07/28/2020 18:44:45",
      "content": "<p>I'm glad to help you out :)</p>",
      "rawMarkdown": "I'm glad to help you out :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 934903,
      "author_name": "seb6084",
      "author_url": "",
      "post_date": "07/18/2020 21:57:56",
      "content": "<p>As long as the dataset is public, you can access it as describe in this notebook by retrieving the gcs path : \n<a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses</a></p>\n\n<p>i don't know the access time from gdrive to colab notebook but i don't think it's better than accessing a gcs link.\nThen, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526</a> provides many datasets links in tfrecords and jpeg with many resolution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 934912,
          "author_name": "graf10a",
          "author_url": "",
          "post_date": "07/18/2020 22:28:05",
          "content": "<p>Yes, this is the right way to do it. And here is another notebook for looking up GCS paths for various public data sets:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">SIIM Show GCS bucket addresses</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934915,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/18/2020 22:34:16",
          "content": "<p>what if dataset is private?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934920,
          "author_name": "graf10a",
          "author_url": "",
          "post_date": "07/18/2020 22:45:48",
          "content": "<p>I have never tried it with a private data set -- all my data sets are public. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934924,
          "author_name": "seb6084",
          "author_url": "",
          "post_date": "07/18/2020 22:52:23",
          "content": "<p>you can't get the gcs link. Another alternative is to create a google cloud account but there's some limitation as long as you don't pay.</p>\n\n<p>Btw, when you get this working, would you mind to tell me how long an epoch takes on pytorch (number of images / resolution / batch size) ? When i tried 256x256 with augmentation on the fly, a epoch of 60000 images with bs of 64 takes like 20 minutes...</p>\n\n<p>I stay with the fact that, with Pytorch, the data is loaded from notebook to TPU and so bandwidth should be the issue. This is not the case with tensorflow as we use tfrecords on gcs (but augmentations are limited to tensor operations).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934930,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/18/2020 23:05:07",
          "content": "<p>256x256 PyTorch times:</p>\n\n<p>(please note this highly depends on augmentations and disk/CPU performance)</p>\n\n<ul>\n<li>on my local GPU with batch_size 30 it is 473.547519 (efficientnet-b1)</li>\n<li>on colab on GPU with batch_size 16 it is 593.404905 (efficientnet-b0)</li>\n<li>on colab on TPU with batch_size 16 it is 383.63376 (efficientnet-b0)</li>\n</ul>\n\n<p>I think I may have problem with this competition with finishing the training, as I can't do much with my local GPU (tried training b6 it is like half hour per epoch)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 934941,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "07/18/2020 23:57:30",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>, for GPU you can download your dataset via the Kaggle API into colab and use it. </p>\n\n<p>However, TPU is faster for training and you can get the TFRecord data's bucket address as described <a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">here</a> at the time of running to use directly in your colab notebook. It is really fast.</p>",
      "votes": null,
      "replies": [
        {
          "id": 934942,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/19/2020 00:01:38",
          "content": "<p>I don't use tfrecords, I don't use tensorflow, I am quite familar with data now, what I am not familar with is google drive and colab for so many files, my current workaround is to decompress zip every time I start colab notebook</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934953,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "07/19/2020 00:17:48",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>, reading from google drive during training is definatley slow. The best way I believe is to have your data as a kaggle dataset and download it to the colab environment in your notebook before you start training. Downloading via the Kaggle Api is relatively fast.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 935018,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/19/2020 02:49:59",
          "content": "<p>can I access private dataset with Kaggle Api?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 935133,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/19/2020 05:19:21",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  You can now use <a href=\"https://www.kaggle.com/product-feedback/163416\">TPUs with Private Datasets</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 935304,
      "author_name": "truonghoang",
      "author_url": "",
      "post_date": "07/19/2020 08:49:25",
      "content": "<p>With me,\n- Create a notebook \n- Add your dataset which you want read from the colab. \n- Use print ex:\n<code>print(KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images'))</code>\n to get address gs from that notebook.\n- Copy this address to your colab. </p>\n\n<p>Sometime the address will expried, you print to get new address.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 936753,
      "author_name": "fireheart7",
      "author_url": "",
      "post_date": "07/20/2020 13:51:04",
      "content": "<p>This might be of help to you</p>\n\n<p><a href=\"https://towardsdatascience.com/setting-up-kaggle-in-google-colab-ebb281b61463\">https://towardsdatascience.com/setting-up-kaggle-in-google-colab-ebb281b61463</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 939842,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "07/22/2020 14:07:36",
      "content": "<p>The safest way is to save the zip into your google drive and unzip.</p>\n\n<p>import zipfile<br>\nfrom google.colab import drive<br>\ndrive.mount('/content/drive/')<br>\nzip_ref = zipfile.ZipFile(\"/content/drive/My Drive/ML/DataSet.zip\", 'r')<br>\nzip_ref.extractall(\"/tmp\")<br>\nzip_ref.close()</p>",
      "votes": null,
      "replies": [
        {
          "id": 949468,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "07/28/2020 16:48:10",
          "content": "<p><a href=\"/alincijov\">@alincijov</a>  Your approach proved to be not only the safest but the fastest way for me. \nI didn't manage to befriend Kaggle gcs paths and my custom Pytorch data loader. My data loader, namely PIL Image.open() command doesn't see files in subfolder \\train and returns 'file not found' error. I understand this error has something to do with permissions and admit it could be resolved easily. I just don't know how.</p>\n\n<p>import zipfile\nfrom google.colab import drive\ndrive.mount('/content/drive')</p>\n\n<p>zipref = zipfile.ZipFile(\"/content/drive/My Drive/data/jpeg-melanoma-256x256.zip\", 'r')\nzipref.extractall(\"/content/jpeg-melanoma-256x256\")\nzipref.close()</p>\n\n<p>What's good about this approach:\n- you don't spend hours to upload unzipped data to your google drive\n- unzip operation works pretty fast\n- we unzip dataset not to a google drive but to a local storage of VM. I am not aware of VM storage topology but assume this step makes our data \"closer\" than keeping it in a google drive.</p>\n\n<p>one epoch, batch=32, simple B1 model ~240 sec</p>\n\n<p>Many thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949619,
          "author_name": "alincijov",
          "author_url": "",
          "post_date": "07/28/2020 18:44:45",
          "content": "<p>I'm glad to help you out :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 949573,
      "author_name": "ademyanchuk",
      "author_url": "",
      "post_date": "07/28/2020 18:07:03",
      "content": "<p>Let me share a snippet, as I just tried to make it work myself recently\n<code>from google.colab import drive</code>\n<code>drive.mount(\"/content/drive\")</code>\n<code>os.environ['KAGGLE_CONFIG_DIR'] = \"/content/drive/My Drive/Kaggle\"</code>\nMake sure kaggle.json (API-key) is in this folder\n<code># Cris dataset for example</code>\n<code>!kaggle datasets download -d cdeotte/jpeg-melanoma-384x384</code>\n<code>!unzip jpeg-melanoma-384x384.zip -d /content/data/jpeg384</code></p>\n\n<p>Worked for me :) Hope, it helps </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "934886": "I was starting with small images like 32x32 and 64x64. With them I could put zip file with all the files on my google drive then unzip them in colab vm. That was working.\n\nBut with 256x256 unziping is very long process. And I need it to perform it every time I start colab.\n\nI am searching for alternatives:\n- put zip file on google drive and use explode to mount zip as folder - works, but very slow (hours per epoch)\n- put uncompressed files into google drive - still trying, it will take hours to upload (single zip is uploading quickly)\n\nDo you know quick way to upload 30000 images into google drive? Maybe there is some way to use files from Kaggle dataset on colab without waiting for unzip?",
    "934903": "As long as the dataset is public, you can access it as describe in this notebook by retrieving the gcs path : \nhttps://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\n\ni don't know the access time from gdrive to colab notebook but i don't think it's better than accessing a gcs link.\nThen, https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526 provides many datasets links in tfrecords and jpeg with many resolution.",
    "934912": "Yes, this is the right way to do it. And here is another notebook for looking up GCS paths for various public data sets:\n\n[SIIM Show GCS bucket addresses](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses)",
    "934915": "what if dataset is private?",
    "934920": "I have never tried it with a private data set -- all my data sets are public.",
    "934924": "you can't get the gcs link. Another alternative is to create a google cloud account but there's some limitation as long as you don't pay.\n\nBtw, when you get this working, would you mind to tell me how long an epoch takes on pytorch (number of images / resolution / batch size) ? When i tried 256x256 with augmentation on the fly, a epoch of 60000 images with bs of 64 takes like 20 minutes...\n\nI stay with the fact that, with Pytorch, the data is loaded from notebook to TPU and so bandwidth should be the issue. This is not the case with tensorflow as we use tfrecords on gcs (but augmentations are limited to tensor operations).",
    "934930": "256x256 PyTorch times:\n\n(please note this highly depends on augmentations and disk/CPU performance)\n\n- on my local GPU with batch_size 30 it is 473.547519 (efficientnet-b1)\n- on colab on GPU with batch_size 16 it is 593.404905 (efficientnet-b0)\n- on colab on TPU with batch_size 16 it is 383.63376 (efficientnet-b0)\n\nI think I may have problem with this competition with finishing the training, as I can't do much with my local GPU (tried training b6 it is like half hour per epoch)",
    "934941": "jacekpoplawski, for GPU you can download your dataset via the Kaggle API into colab and use it. \n\nHowever, TPU is faster for training and you can get the TFRecord data's bucket address as described [here](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses) at the time of running to use directly in your colab notebook. It is really fast.",
    "934942": "I don't use tfrecords, I don't use tensorflow, I am quite familar with data now, what I am not familar with is google drive and colab for so many files, my current workaround is to decompress zip every time I start colab notebook",
    "934953": "jacekpoplawski, reading from google drive during training is definatley slow. The best way I believe is to have your data as a kaggle dataset and download it to the colab environment in your notebook before you start training. Downloading via the Kaggle Api is relatively fast.",
    "935018": "can I access private dataset with Kaggle Api?",
    "935133": "jacekpoplawski  You can now use [TPUs with Private Datasets](https://www.kaggle.com/product-feedback/163416)",
    "935304": "With me,\n- Create a notebook \n- Add your dataset which you want read from the colab. \n- Use print ex:\n`print(KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images'))`\n to get address gs from that notebook.\n- Copy this address to your colab. \n\nSometime the address will expried, you print to get new address.",
    "936753": "This might be of help to you\n\nhttps://towardsdatascience.com/setting-up-kaggle-in-google-colab-ebb281b61463",
    "939842": "The safest way is to save the zip into your google drive and unzip.\n\nimport zipfile<br>\nfrom google.colab import drive<br>\ndrive.mount('/content/drive/')<br>\nzip_ref = zipfile.ZipFile(\"/content/drive/My Drive/ML/DataSet.zip\", 'r')<br>\nzip_ref.extractall(\"/tmp\")<br>\nzip_ref.close()",
    "949468": "alincijov  Your approach proved to be not only the safest but the fastest way for me. \nI didn't manage to befriend Kaggle gcs paths and my custom Pytorch data loader. My data loader, namely PIL Image.open() command doesn't see files in subfolder \\train and returns 'file not found' error. I understand this error has something to do with permissions and admit it could be resolved easily. I just don't know how.\n\nimport zipfile\nfrom google.colab import drive\ndrive.mount('/content/drive')\n\nzipref = zipfile.ZipFile(\"/content/drive/My Drive/data/jpeg-melanoma-256x256.zip\", 'r')\nzipref.extractall(\"/content/jpeg-melanoma-256x256\")\nzipref.close()\n\nWhat's good about this approach:\n- you don't spend hours to upload unzipped data to your google drive\n- unzip operation works pretty fast\n- we unzip dataset not to a google drive but to a local storage of VM. I am not aware of VM storage topology but assume this step makes our data \"closer\" than keeping it in a google drive.\n\none epoch, batch=32, simple B1 model ~240 sec\n\nMany thanks!",
    "949573": "Let me share a snippet, as I just tried to make it work myself recently\n`from google.colab import drive`\n`drive.mount(\"/content/drive\")`\n`os.environ['KAGGLE_CONFIG_DIR'] = \"/content/drive/My Drive/Kaggle\"`\nMake sure kaggle.json (API-key) is in this folder\n`# Cris dataset for example`\n`!kaggle datasets download -d cdeotte/jpeg-melanoma-384x384`\n`!unzip jpeg-melanoma-384x384.zip -d /content/data/jpeg384`\n\nWorked for me :) Hope, it helps",
    "949619": "I'm glad to help you out :)"
  },
  "source": "meta"
}