{
  "id": 241414,
  "title": "The pros and cons of the 3 options to deal with Kaggle Datasets in Colab .",
  "url": "/competitions/birdclef-2021/discussion/241414",
  "author_name": "",
  "post_date": "2021-05-24T13:43:59.366629100Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all 😃</p>\n<p>I am interested to know the pros and cons of dealing with the Kaggle Dataset in Colab.</p>\n<p>There are 3 options that I know:</p>\n<p>1- to download and keep the datasets to Colab VM (\"/tmp\")<br>\n2- to download and keep the datasets in Google Drive.<br>\n3- to Not download the dataset and use KaggleDatasets().get_gcs_path.</p>\n<p>Code examples:</p>\n<p>1- <a href=\"https://www.kaggle.com/jamshaidsohail5/colab-version-of-pytorch-training-birdclef2021/data\" target=\"_blank\">https://www.kaggle.com/jamshaidsohail5/colab-version-of-pytorch-training-birdclef2021/data</a><br>\n3-<a href=\"https://www.kaggle.com/dragonzhang/colab-version-bms-efficientnetv2-tpu-lb-3-18\" target=\"_blank\">https://www.kaggle.com/dragonzhang/colab-version-bms-efficientnetv2-tpu-lb-3-18</a> </p>\n<p>if you have any experience, please share it with us. </p>\n<p>thank you I appreciate it</p>\n<p>all the best <br>\nFaisal </p>",
  "messages": [
    {
      "id": "1321035",
      "postDate": "05/24/2021 13:43:59",
      "content": "<p>Hi all 😃</p>\n<p>I am interested to know the pros and cons of dealing with the Kaggle Dataset in Colab.</p>\n<p>There are 3 options that I know:</p>\n<p>1- to download and keep the datasets to Colab VM (\"/tmp\")<br>\n2- to download and keep the datasets in Google Drive.<br>\n3- to Not download the dataset and use KaggleDatasets().get_gcs_path.</p>\n<p>Code examples:</p>\n<p>1- <a href=\"https://www.kaggle.com/jamshaidsohail5/colab-version-of-pytorch-training-birdclef2021/data\" target=\"_blank\">https://www.kaggle.com/jamshaidsohail5/colab-version-of-pytorch-training-birdclef2021/data</a><br>\n3-<a href=\"https://www.kaggle.com/dragonzhang/colab-version-bms-efficientnetv2-tpu-lb-3-18\" target=\"_blank\">https://www.kaggle.com/dragonzhang/colab-version-bms-efficientnetv2-tpu-lb-3-18</a> </p>\n<p>if you have any experience, please share it with us. </p>\n<p>thank you I appreciate it</p>\n<p>all the best <br>\nFaisal </p>",
      "rawMarkdown": "Hi all 😃\n\nI am interested to know the pros and cons of dealing with the Kaggle Dataset in Colab.\n\nThere are 3 options that I know:\n\n1- to download and keep the datasets to Colab VM (\"/tmp\")\n2- to download and keep the datasets in Google Drive.\n3- to Not download the dataset and use KaggleDatasets().get_gcs_path.\n\n\nCode examples:\n\n1- https://www.kaggle.com/jamshaidsohail5/colab-version-of-pytorch-training-birdclef2021/data\n3-https://www.kaggle.com/dragonzhang/colab-version-bms-efficientnetv2-tpu-lb-3-18 \n\nif you have any experience, please share it with us. \n\nthank you I appreciate it\n\nall the best \nFaisal",
      "votes": null
    },
    {
      "id": "1321327",
      "postDate": "05/24/2021 16:09:37",
      "content": "<p>Hi,</p>\n<p>I have done I think 3 of them and here is my observation </p>\n<ol>\n<li>It takes up the hard disk, so if the data is too big you might not even able to unzip it in colab. If data is not so big then this is the best option to read from disk. Basically, if you use <code>pytorch</code> or <code>tf.dataset</code>, this will be faster while you train the models. Using this, you can speed up the training a bit rather than using google drive to store the data. Probably using the GCS path to get the data is also slower than this method.</li>\n<li>This is to store the data when data is too large. But using the data from google drive has the drawback of slow speed while training. This will sometimes cause the drive to collapse and the session to restart. So use it if the data is too large and you are fine with slow speed. </li>\n<li>I think using GCS paths for data is best if you are using TPUs. TPUs gets data from GCS anyways so having data already there and training with TPU makes things fast. But not sure how much this method is useful when you are dealing with GPUs because now one way or another you are downloading data from GCS and it will make things a bit slow if you can not access them fast. </li>\n</ol>\n<p>I hope it helps.</p>",
      "rawMarkdown": "Hi,\n\nI have done I think 3 of them and here is my observation \n1. It takes up the hard disk, so if the data is too big you might not even able to unzip it in colab. If data is not so big then this is the best option to read from disk. Basically, if you use `pytorch` or `tf.dataset`, this will be faster while you train the models. Using this, you can speed up the training a bit rather than using google drive to store the data. Probably using the GCS path to get the data is also slower than this method.\n2. This is to store the data when data is too large. But using the data from google drive has the drawback of slow speed while training. This will sometimes cause the drive to collapse and the session to restart. So use it if the data is too large and you are fine with slow speed. \n3. I think using GCS paths for data is best if you are using TPUs. TPUs gets data from GCS anyways so having data already there and training with TPU makes things fast. But not sure how much this method is useful when you are dealing with GPUs because now one way or another you are downloading data from GCS and it will make things a bit slow if you can not access them fast. \n\nI hope it helps.",
      "votes": null
    },
    {
      "id": "1321477",
      "postDate": "05/24/2021 17:48:19",
      "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a>  <br>\nThis is very helpful.</p>",
      "rawMarkdown": "Thank you very much, @urvishp80  \nThis is very helpful.",
      "votes": null
    },
    {
      "id": "1321967",
      "postDate": "05/25/2021 05:54:51",
      "content": "<p>You're welcome!</p>",
      "rawMarkdown": "You're welcome!",
      "votes": null
    },
    {
      "id": "1323415",
      "postDate": "05/26/2021 08:05:04",
      "content": "<p>Another way if there is space issues, not the fastest but stil, use !fuse-zip /content/…. /content/… to mount a drive and read directly from the zip file.</p>",
      "rawMarkdown": "Another way if there is space issues, not the fastest but stil, use !fuse-zip /content/.... /content/... to mount a drive and read directly from the zip file.",
      "votes": null
    },
    {
      "id": "1323586",
      "postDate": "05/26/2021 10:28:14",
      "content": "<p>Thnak you <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> <br>\ninteresting..<br>\nCan you please give us a code example? I want to try it.</p>",
      "rawMarkdown": "Thnak you @kirderf \ninteresting..\nCan you please give us a code example? I want to try it.",
      "votes": null
    },
    {
      "id": "1323599",
      "postDate": "05/26/2021 10:40:36",
      "content": "<p>I did <br>\n<code>!fuse-zip /content/…. /content/…</code></p>",
      "rawMarkdown": "I did \n`!fuse-zip /content/…. /content/…`",
      "votes": null
    },
    {
      "id": "1323647",
      "postDate": "05/26/2021 11:22:46",
      "content": "<p>Very nice trick here, thank you!</p>",
      "rawMarkdown": "Very nice trick here, thank you!",
      "votes": null
    },
    {
      "id": "1323676",
      "postDate": "05/26/2021 11:49:50",
      "content": "<p>oh I see now</p>\n<blockquote>\n  <p>fuse-zip is a FUSE file system to navigate, extract, create and modify ZIP.  With fuse-zip you really can work with ZIP archives as real directories.</p>\n</blockquote>\n<p>Resources<br>\n<a href=\"https://bitbucket.org/agalanin/fuse-zip/wiki/Home\" target=\"_blank\">https://bitbucket.org/agalanin/fuse-zip/wiki/Home</a><br>\n<a href=\"https://linux.die.net/man/1/fuse-zip\" target=\"_blank\">https://linux.die.net/man/1/fuse-zip</a><br>\n<a href=\"https://unix.stackexchange.com/questions/168807/mount-zip-file-as-a-read-only-filesystem\" target=\"_blank\">https://unix.stackexchange.com/questions/168807/mount-zip-file-as-a-read-only-filesystem</a></p>\n<p>Thank you </p>",
      "rawMarkdown": "oh I see now\n\n> fuse-zip is a FUSE file system to navigate, extract, create and modify ZIP.  With fuse-zip you really can work with ZIP archives as real directories.\n\nResources\nhttps://bitbucket.org/agalanin/fuse-zip/wiki/Home\nhttps://linux.die.net/man/1/fuse-zip\nhttps://unix.stackexchange.com/questions/168807/mount-zip-file-as-a-read-only-filesystem\n\nThank you",
      "votes": null
    },
    {
      "id": "1323718",
      "postDate": "05/26/2021 12:09:15",
      "content": "<p>Just install it with <code>!apt-get install -y fuse-zip</code></p>",
      "rawMarkdown": "Just install it with `!apt-get install -y fuse-zip`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1321327,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "05/24/2021 16:09:37",
      "content": "<p>Hi,</p>\n<p>I have done I think 3 of them and here is my observation </p>\n<ol>\n<li>It takes up the hard disk, so if the data is too big you might not even able to unzip it in colab. If data is not so big then this is the best option to read from disk. Basically, if you use <code>pytorch</code> or <code>tf.dataset</code>, this will be faster while you train the models. Using this, you can speed up the training a bit rather than using google drive to store the data. Probably using the GCS path to get the data is also slower than this method.</li>\n<li>This is to store the data when data is too large. But using the data from google drive has the drawback of slow speed while training. This will sometimes cause the drive to collapse and the session to restart. So use it if the data is too large and you are fine with slow speed. </li>\n<li>I think using GCS paths for data is best if you are using TPUs. TPUs gets data from GCS anyways so having data already there and training with TPU makes things fast. But not sure how much this method is useful when you are dealing with GPUs because now one way or another you are downloading data from GCS and it will make things a bit slow if you can not access them fast. </li>\n</ol>\n<p>I hope it helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1321477,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "05/24/2021 17:48:19",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a>  <br>\nThis is very helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1321967,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "05/25/2021 05:54:51",
          "content": "<p>You're welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1323415,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "05/26/2021 08:05:04",
      "content": "<p>Another way if there is space issues, not the fastest but stil, use !fuse-zip /content/…. /content/… to mount a drive and read directly from the zip file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1323586,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "05/26/2021 10:28:14",
          "content": "<p>Thnak you <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> <br>\ninteresting..<br>\nCan you please give us a code example? I want to try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1323599,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/26/2021 10:40:36",
          "content": "<p>I did <br>\n<code>!fuse-zip /content/…. /content/…</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1323647,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "05/26/2021 11:22:46",
          "content": "<p>Very nice trick here, thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1323676,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "05/26/2021 11:49:50",
          "content": "<p>oh I see now</p>\n<blockquote>\n  <p>fuse-zip is a FUSE file system to navigate, extract, create and modify ZIP.  With fuse-zip you really can work with ZIP archives as real directories.</p>\n</blockquote>\n<p>Resources<br>\n<a href=\"https://bitbucket.org/agalanin/fuse-zip/wiki/Home\" target=\"_blank\">https://bitbucket.org/agalanin/fuse-zip/wiki/Home</a><br>\n<a href=\"https://linux.die.net/man/1/fuse-zip\" target=\"_blank\">https://linux.die.net/man/1/fuse-zip</a><br>\n<a href=\"https://unix.stackexchange.com/questions/168807/mount-zip-file-as-a-read-only-filesystem\" target=\"_blank\">https://unix.stackexchange.com/questions/168807/mount-zip-file-as-a-read-only-filesystem</a></p>\n<p>Thank you </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1323718,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/26/2021 12:09:15",
          "content": "<p>Just install it with <code>!apt-get install -y fuse-zip</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1321035": "Hi all 😃\n\nI am interested to know the pros and cons of dealing with the Kaggle Dataset in Colab.\n\nThere are 3 options that I know:\n\n1- to download and keep the datasets to Colab VM (\"/tmp\")\n2- to download and keep the datasets in Google Drive.\n3- to Not download the dataset and use KaggleDatasets().get_gcs_path.\n\n\nCode examples:\n\n1- https://www.kaggle.com/jamshaidsohail5/colab-version-of-pytorch-training-birdclef2021/data\n3-https://www.kaggle.com/dragonzhang/colab-version-bms-efficientnetv2-tpu-lb-3-18 \n\nif you have any experience, please share it with us. \n\nthank you I appreciate it\n\nall the best \nFaisal",
    "1321327": "Hi,\n\nI have done I think 3 of them and here is my observation \n1. It takes up the hard disk, so if the data is too big you might not even able to unzip it in colab. If data is not so big then this is the best option to read from disk. Basically, if you use `pytorch` or `tf.dataset`, this will be faster while you train the models. Using this, you can speed up the training a bit rather than using google drive to store the data. Probably using the GCS path to get the data is also slower than this method.\n2. This is to store the data when data is too large. But using the data from google drive has the drawback of slow speed while training. This will sometimes cause the drive to collapse and the session to restart. So use it if the data is too large and you are fine with slow speed. \n3. I think using GCS paths for data is best if you are using TPUs. TPUs gets data from GCS anyways so having data already there and training with TPU makes things fast. But not sure how much this method is useful when you are dealing with GPUs because now one way or another you are downloading data from GCS and it will make things a bit slow if you can not access them fast. \n\nI hope it helps.",
    "1321477": "Thank you very much, @urvishp80  \nThis is very helpful.",
    "1321967": "You're welcome!",
    "1323415": "Another way if there is space issues, not the fastest but stil, use !fuse-zip /content/.... /content/... to mount a drive and read directly from the zip file.",
    "1323586": "Thnak you @kirderf \ninteresting..\nCan you please give us a code example? I want to try it.",
    "1323599": "I did \n`!fuse-zip /content/…. /content/…`",
    "1323647": "Very nice trick here, thank you!",
    "1323676": "oh I see now\n\n> fuse-zip is a FUSE file system to navigate, extract, create and modify ZIP.  With fuse-zip you really can work with ZIP archives as real directories.\n\nResources\nhttps://bitbucket.org/agalanin/fuse-zip/wiki/Home\nhttps://linux.die.net/man/1/fuse-zip\nhttps://unix.stackexchange.com/questions/168807/mount-zip-file-as-a-read-only-filesystem\n\nThank you",
    "1323718": "Just install it with `!apt-get install -y fuse-zip`"
  },
  "source": "meta"
}