{
  "id": 307995,
  "title": "How do I efficiently access competition dataset for Colab/External training?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307995",
  "author_name": "",
  "post_date": "2022-02-16T15:56:04.743123700Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I usually download competition data locally, and then upload it to wherever I would like to train. I've usually done this with 2-3 GB datasets, but this is the first time I have to deal with a 65 GB dataset. I don't have enough time or storage to move this data like I usually do, and I realize there must be a better way to load in huge datasets for external training. Otherwise, how else can researchers train on huge datasets like ImageNet? How should I go about loading and moving this huge dataset for training outside of Kaggle, like in Colab, that is efficient?</p>",
  "messages": [
    {
      "id": "1693340",
      "postDate": "02/16/2022 15:56:04",
      "content": "<p>I usually download competition data locally, and then upload it to wherever I would like to train. I've usually done this with 2-3 GB datasets, but this is the first time I have to deal with a 65 GB dataset. I don't have enough time or storage to move this data like I usually do, and I realize there must be a better way to load in huge datasets for external training. Otherwise, how else can researchers train on huge datasets like ImageNet? How should I go about loading and moving this huge dataset for training outside of Kaggle, like in Colab, that is efficient?</p>",
      "rawMarkdown": "I usually download competition data locally, and then upload it to wherever I would like to train. I've usually done this with 2-3 GB datasets, but this is the first time I have to deal with a 65 GB dataset. I don't have enough time or storage to move this data like I usually do, and I realize there must be a better way to load in huge datasets for external training. Otherwise, how else can researchers train on huge datasets like ImageNet? How should I go about loading and moving this huge dataset for training outside of Kaggle, like in Colab, that is efficient?",
      "votes": null
    },
    {
      "id": "1693476",
      "postDate": "02/16/2022 17:37:43",
      "content": "<p>You can download the data using <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a>. You can get your API token from the <code>Account</code> section listed under <code>Your Profile</code>. <a href=\"https://github.com/Kaggle/kaggle-api#competitions\" target=\"_blank\">This</a> part explains the commands to download the competition data. </p>",
      "rawMarkdown": "You can download the data using [Kaggle API](https://github.com/Kaggle/kaggle-api). You can get your API token from the `Account` section listed under `Your Profile`. [This](https://github.com/Kaggle/kaggle-api#competitions) part explains the commands to download the competition data.",
      "votes": null
    },
    {
      "id": "1693535",
      "postDate": "02/16/2022 18:12:10",
      "content": "<p><a href=\"https://www.kaggle.com/anandparthiban\" target=\"_blank\">@anandparthiban</a>  checkout this <a href=\"https://youtu.be/57N1g8k2Hwc\" target=\"_blank\">video</a></p>",
      "rawMarkdown": "anandparthiban  checkout this [video](https://youtu.be/57N1g8k2Hwc)",
      "votes": null
    },
    {
      "id": "1693847",
      "postDate": "02/17/2022 02:16:19",
      "content": "<p>If you are going to use a TPU on Colab it is not really necessary to move any data at all. It is possible for Colab to \"see\" the TFRecords by using the path to the records with a few lines from a Kaggle notebook.</p>\n<p>In the Kaggle notebook you'll need…</p>\n<pre><code>#TFRecords\nGCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')\nprint('GCS_DS_PATH', GCS_DS_PATH)\n\n#main data\nGCS_DS_PATH2=KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')\nprint('GCS_DS_PATH2', GCS_DS_PATH2)\n</code></pre>\n<p>In Colab you can access the info directly by using the path info you have found out.</p>\n<pre><code>GCS_DS_PATH='gs://kds-9b264a598efde6f592f9ef3ff55274d72e53ad51bec300328115530c'\ntrain_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-train*.tfrec')))\ntest_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-test*.tfrec')))\n\nGCS_DS_PATH2='gs://kds-ec2a2e61dd5345d807d47adc26eb93dcc160dd3b4da4084f45a859f0'\nsample_submission = pd.read_csv(GCS_DS_PATH2 + '/sample_submission.csv',index_col='image')\n</code></pre>\n<p>In Colab you'll have to install the Google File system</p>\n<p><code>!pip install gcsfs</code></p>\n<p>and as you won't need it anymore, comment out references to the Kaggle Dataset module</p>\n<p><code>#from kaggle_datasets import KaggleDatasets</code></p>\n<p>There are some public notebooks already set up for Colab use such as…<br>\n<a href=\"https://www.kaggle.com/librauee/train-arcfaceeffb6\" target=\"_blank\">【Train】ArcFaceEffB6</a> from <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a> </p>\n<p>The TFRecords example is from a <a href=\"https://www.kaggle.com/manojprabhaakr/whale-tfrecords-512\" target=\"_blank\">dataset</a> made available by <a href=\"https://www.kaggle.com/manojprabhaakr\" target=\"_blank\">@manojprabhaakr</a> </p>\n<p>The only real downside to this is that the link will expire after about a week, however all you have to do is run the notebook again in Kaggle … get the new GCS_DS_PATH, put it in Colab and resume using Colab. I have Google drive and it is possible to link your Colab to your Google drive. </p>\n<pre><code>from google.colab import drive\ndrive.mount('/content/drive')\n</code></pre>\n<p>I would not do that as a way of accessing large amounts of data though. It's way too slow. I use it for smaller files like the sample_submission.csv file so I don't have to worry about that path link expiring. </p>\n<p>Linking your Google drive is also very important for one other reason. When Colab finishes it's run it may expire the session if you don't get back to it in time. To avoid losing all your results and files, I simply put in a line at the end that copies all the files I care about over to my Google drive, with something like…</p>\n<p><code>!cp -R /content/working /content/drive/MyDrive/colab_models/WHALE_Model</code></p>\n<p>It will depend on how you set up the Colab and Google drives of course.</p>\n<p>If you have a look at some of the TPU based note books you will see some of the things I have mentioned.</p>\n<p>Hope that helps</p>",
      "rawMarkdown": "If you are going to use a TPU on Colab it is not really necessary to move any data at all. It is possible for Colab to \"see\" the TFRecords by using the path to the records with a few lines from a Kaggle notebook.\n\nIn the Kaggle notebook you'll need...\n\n```\n#TFRecords\nGCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')\nprint('GCS_DS_PATH', GCS_DS_PATH)\n\n#main data\nGCS_DS_PATH2=KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')\nprint('GCS_DS_PATH2', GCS_DS_PATH2)\n```\n\nIn Colab you can access the info directly by using the path info you have found out.\n\n```\nGCS_DS_PATH='gs://kds-9b264a598efde6f592f9ef3ff55274d72e53ad51bec300328115530c'\ntrain_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-train*.tfrec')))\ntest_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-test*.tfrec')))\n\nGCS_DS_PATH2='gs://kds-ec2a2e61dd5345d807d47adc26eb93dcc160dd3b4da4084f45a859f0'\nsample_submission = pd.read_csv(GCS_DS_PATH2 + '/sample_submission.csv',index_col='image')\n```\n\nIn Colab you'll have to install the Google File system\n\n`!pip install gcsfs`\n\nand as you won't need it anymore, comment out references to the Kaggle Dataset module\n\n`#from kaggle_datasets import KaggleDatasets`\n\nThere are some public notebooks already set up for Colab use such as...\n[【Train】ArcFaceEffB6](https://www.kaggle.com/librauee/train-arcfaceeffb6) from @librauee \n\nThe TFRecords example is from a [dataset](https://www.kaggle.com/manojprabhaakr/whale-tfrecords-512) made available by @manojprabhaakr \n\nThe only real downside to this is that the link will expire after about a week, however all you have to do is run the notebook again in Kaggle ... get the new GCS_DS_PATH, put it in Colab and resume using Colab. I have Google drive and it is possible to link your Colab to your Google drive. \n\n```\nfrom google.colab import drive\ndrive.mount('/content/drive')\n```\n\nI would not do that as a way of accessing large amounts of data though. It's way too slow. I use it for smaller files like the sample_submission.csv file so I don't have to worry about that path link expiring. \n\nLinking your Google drive is also very important for one other reason. When Colab finishes it's run it may expire the session if you don't get back to it in time. To avoid losing all your results and files, I simply put in a line at the end that copies all the files I care about over to my Google drive, with something like...\n\n`!cp -R /content/working /content/drive/MyDrive/colab_models/WHALE_Model`\n\nIt will depend on how you set up the Colab and Google drives of course.\n\nIf you have a look at some of the TPU based note books you will see some of the things I have mentioned.\n\nHope that helps",
      "votes": null
    },
    {
      "id": "1698834",
      "postDate": "02/20/2022 17:51:32",
      "content": "<p><a href=\"https://www.kaggle.com/BruceYoung\" target=\"_blank\">@BruceYoung</a> I really appreciate the answer you have given. I downloaded the dataset for the competition into Colab using Kaggle's API, but instead of downloading a single train.zip, only about 30 individually zipped images were downloaded. Has anyone else experienced a similar experience using the Kaggle API?</p>",
      "rawMarkdown": "BruceYoung I really appreciate the answer you have given. I downloaded the dataset for the competition into Colab using Kaggle's API, but instead of downloading a single train.zip, only about 30 individually zipped images were downloaded. Has anyone else experienced a similar experience using the Kaggle API?",
      "votes": null
    },
    {
      "id": "1700603",
      "postDate": "02/22/2022 05:34:08",
      "content": "<p>hi, i am getting this error when i run the code in the kaggle notebook :</p>\n<pre><code>BackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'whale-tfrecords-512', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5, 'details': []}, 'wasSuccessful': False}.\n</code></pre>",
      "rawMarkdown": "hi, i am getting this error when i run the code in the kaggle notebook :\n```\nBackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'whale-tfrecords-512', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5, 'details': []}, 'wasSuccessful': False}.\n```",
      "votes": null
    },
    {
      "id": "1700661",
      "postDate": "02/22/2022 06:17:08",
      "content": "<p>If you are using the notebook for training that I linked to you'll have to either add in the dataset for <code>whale-tfrecords-512</code> or just use the dataset he has <code>happywhale-tfrecords-v1</code></p>\n<p>Then just make sure this…</p>\n<p><code>GCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')</code></p>\n<p>…has the right info for the dataset you are using.</p>\n<p>If that is not what you are having trouble with,  just let me know.</p>",
      "rawMarkdown": "If you are using the notebook for training that I linked to you'll have to either add in the dataset for `whale-tfrecords-512` or just use the dataset he has `happywhale-tfrecords-v1`\n\nThen just make sure this...\n\n`GCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')`\n\n...has the right info for the dataset you are using.\n\nIf that is not what you are having trouble with,  just let me know.",
      "votes": null
    },
    {
      "id": "1700672",
      "postDate": "02/22/2022 06:24:39",
      "content": "<p>I don't think I've used the API more than about once and that was years ago. I download the lot as a zip file and upload it to my Googe drive. This is a very slow. It is also a slow way to use large amounts of data. You can, of course, copy the data you need from a Google drive to a colab drive but again, that is slow and data on your colab drive is not persistent. I got away with this in the COTS comp as only part of the data was needed so not all needed to be copied from my Google drive for each session.</p>\n<p>All this is why I love using a TPU with a link to the original data.</p>",
      "rawMarkdown": "I don't think I've used the API more than about once and that was years ago. I download the lot as a zip file and upload it to my Googe drive. This is a very slow. It is also a slow way to use large amounts of data. You can, of course, copy the data you need from a Google drive to a colab drive but again, that is slow and data on your colab drive is not persistent. I got away with this in the COTS comp as only part of the data was needed so not all needed to be copied from my Google drive for each session.\n\nAll this is why I love using a TPU with a link to the original data.",
      "votes": null
    },
    {
      "id": "1700684",
      "postDate": "02/22/2022 06:44:09",
      "content": "<p>Thank you for your explanation. It is very helpful. I've understood why tfrecords is important for TPU as well as how it could be imported in google colab.</p>\n<p>I am wondering is that approach (<code>KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')</code>) is suitable for non-TPU training?</p>",
      "rawMarkdown": "Thank you for your explanation. It is very helpful. I've understood why tfrecords is important for TPU as well as how it could be imported in google colab.\n\nI am wondering is that approach (`KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')`) is suitable for non-TPU training?",
      "votes": null
    },
    {
      "id": "1700831",
      "postDate": "02/22/2022 09:33:38",
      "content": "<p>thank you so much! this worked!</p>",
      "rawMarkdown": "thank you so much! this worked!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1693476,
      "author_name": "atharvaingle",
      "author_url": "",
      "post_date": "02/16/2022 17:37:43",
      "content": "<p>You can download the data using <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a>. You can get your API token from the <code>Account</code> section listed under <code>Your Profile</code>. <a href=\"https://github.com/Kaggle/kaggle-api#competitions\" target=\"_blank\">This</a> part explains the commands to download the competition data. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1693535,
      "author_name": "sarabhian",
      "author_url": "",
      "post_date": "02/16/2022 18:12:10",
      "content": "<p><a href=\"https://www.kaggle.com/anandparthiban\" target=\"_blank\">@anandparthiban</a>  checkout this <a href=\"https://youtu.be/57N1g8k2Hwc\" target=\"_blank\">video</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1693847,
      "author_name": "mutantspore",
      "author_url": "",
      "post_date": "02/17/2022 02:16:19",
      "content": "<p>If you are going to use a TPU on Colab it is not really necessary to move any data at all. It is possible for Colab to \"see\" the TFRecords by using the path to the records with a few lines from a Kaggle notebook.</p>\n<p>In the Kaggle notebook you'll need…</p>\n<pre><code>#TFRecords\nGCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')\nprint('GCS_DS_PATH', GCS_DS_PATH)\n\n#main data\nGCS_DS_PATH2=KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')\nprint('GCS_DS_PATH2', GCS_DS_PATH2)\n</code></pre>\n<p>In Colab you can access the info directly by using the path info you have found out.</p>\n<pre><code>GCS_DS_PATH='gs://kds-9b264a598efde6f592f9ef3ff55274d72e53ad51bec300328115530c'\ntrain_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-train*.tfrec')))\ntest_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-test*.tfrec')))\n\nGCS_DS_PATH2='gs://kds-ec2a2e61dd5345d807d47adc26eb93dcc160dd3b4da4084f45a859f0'\nsample_submission = pd.read_csv(GCS_DS_PATH2 + '/sample_submission.csv',index_col='image')\n</code></pre>\n<p>In Colab you'll have to install the Google File system</p>\n<p><code>!pip install gcsfs</code></p>\n<p>and as you won't need it anymore, comment out references to the Kaggle Dataset module</p>\n<p><code>#from kaggle_datasets import KaggleDatasets</code></p>\n<p>There are some public notebooks already set up for Colab use such as…<br>\n<a href=\"https://www.kaggle.com/librauee/train-arcfaceeffb6\" target=\"_blank\">【Train】ArcFaceEffB6</a> from <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a> </p>\n<p>The TFRecords example is from a <a href=\"https://www.kaggle.com/manojprabhaakr/whale-tfrecords-512\" target=\"_blank\">dataset</a> made available by <a href=\"https://www.kaggle.com/manojprabhaakr\" target=\"_blank\">@manojprabhaakr</a> </p>\n<p>The only real downside to this is that the link will expire after about a week, however all you have to do is run the notebook again in Kaggle … get the new GCS_DS_PATH, put it in Colab and resume using Colab. I have Google drive and it is possible to link your Colab to your Google drive. </p>\n<pre><code>from google.colab import drive\ndrive.mount('/content/drive')\n</code></pre>\n<p>I would not do that as a way of accessing large amounts of data though. It's way too slow. I use it for smaller files like the sample_submission.csv file so I don't have to worry about that path link expiring. </p>\n<p>Linking your Google drive is also very important for one other reason. When Colab finishes it's run it may expire the session if you don't get back to it in time. To avoid losing all your results and files, I simply put in a line at the end that copies all the files I care about over to my Google drive, with something like…</p>\n<p><code>!cp -R /content/working /content/drive/MyDrive/colab_models/WHALE_Model</code></p>\n<p>It will depend on how you set up the Colab and Google drives of course.</p>\n<p>If you have a look at some of the TPU based note books you will see some of the things I have mentioned.</p>\n<p>Hope that helps</p>",
      "votes": null,
      "replies": [
        {
          "id": 1700603,
          "author_name": "pjmathematician",
          "author_url": "",
          "post_date": "02/22/2022 05:34:08",
          "content": "<p>hi, i am getting this error when i run the code in the kaggle notebook :</p>\n<pre><code>BackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'whale-tfrecords-512', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5, 'details': []}, 'wasSuccessful': False}.\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1700661,
          "author_name": "mutantspore",
          "author_url": "",
          "post_date": "02/22/2022 06:17:08",
          "content": "<p>If you are using the notebook for training that I linked to you'll have to either add in the dataset for <code>whale-tfrecords-512</code> or just use the dataset he has <code>happywhale-tfrecords-v1</code></p>\n<p>Then just make sure this…</p>\n<p><code>GCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')</code></p>\n<p>…has the right info for the dataset you are using.</p>\n<p>If that is not what you are having trouble with,  just let me know.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1700684,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "02/22/2022 06:44:09",
          "content": "<p>Thank you for your explanation. It is very helpful. I've understood why tfrecords is important for TPU as well as how it could be imported in google colab.</p>\n<p>I am wondering is that approach (<code>KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')</code>) is suitable for non-TPU training?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1700831,
          "author_name": "pjmathematician",
          "author_url": "",
          "post_date": "02/22/2022 09:33:38",
          "content": "<p>thank you so much! this worked!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1698834,
      "author_name": "anandparthiban",
      "author_url": "",
      "post_date": "02/20/2022 17:51:32",
      "content": "<p><a href=\"https://www.kaggle.com/BruceYoung\" target=\"_blank\">@BruceYoung</a> I really appreciate the answer you have given. I downloaded the dataset for the competition into Colab using Kaggle's API, but instead of downloading a single train.zip, only about 30 individually zipped images were downloaded. Has anyone else experienced a similar experience using the Kaggle API?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1700672,
          "author_name": "mutantspore",
          "author_url": "",
          "post_date": "02/22/2022 06:24:39",
          "content": "<p>I don't think I've used the API more than about once and that was years ago. I download the lot as a zip file and upload it to my Googe drive. This is a very slow. It is also a slow way to use large amounts of data. You can, of course, copy the data you need from a Google drive to a colab drive but again, that is slow and data on your colab drive is not persistent. I got away with this in the COTS comp as only part of the data was needed so not all needed to be copied from my Google drive for each session.</p>\n<p>All this is why I love using a TPU with a link to the original data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1693340": "I usually download competition data locally, and then upload it to wherever I would like to train. I've usually done this with 2-3 GB datasets, but this is the first time I have to deal with a 65 GB dataset. I don't have enough time or storage to move this data like I usually do, and I realize there must be a better way to load in huge datasets for external training. Otherwise, how else can researchers train on huge datasets like ImageNet? How should I go about loading and moving this huge dataset for training outside of Kaggle, like in Colab, that is efficient?",
    "1693476": "You can download the data using [Kaggle API](https://github.com/Kaggle/kaggle-api). You can get your API token from the `Account` section listed under `Your Profile`. [This](https://github.com/Kaggle/kaggle-api#competitions) part explains the commands to download the competition data.",
    "1693535": "anandparthiban  checkout this [video](https://youtu.be/57N1g8k2Hwc)",
    "1693847": "If you are going to use a TPU on Colab it is not really necessary to move any data at all. It is possible for Colab to \"see\" the TFRecords by using the path to the records with a few lines from a Kaggle notebook.\n\nIn the Kaggle notebook you'll need...\n\n```\n#TFRecords\nGCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')\nprint('GCS_DS_PATH', GCS_DS_PATH)\n\n#main data\nGCS_DS_PATH2=KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')\nprint('GCS_DS_PATH2', GCS_DS_PATH2)\n```\n\nIn Colab you can access the info directly by using the path info you have found out.\n\n```\nGCS_DS_PATH='gs://kds-9b264a598efde6f592f9ef3ff55274d72e53ad51bec300328115530c'\ntrain_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-train*.tfrec')))\ntest_files = np.sort(np.array(tf.io.gfile.glob(GCS_DS_PATH + '/happywhale-2022-test*.tfrec')))\n\nGCS_DS_PATH2='gs://kds-ec2a2e61dd5345d807d47adc26eb93dcc160dd3b4da4084f45a859f0'\nsample_submission = pd.read_csv(GCS_DS_PATH2 + '/sample_submission.csv',index_col='image')\n```\n\nIn Colab you'll have to install the Google File system\n\n`!pip install gcsfs`\n\nand as you won't need it anymore, comment out references to the Kaggle Dataset module\n\n`#from kaggle_datasets import KaggleDatasets`\n\nThere are some public notebooks already set up for Colab use such as...\n[【Train】ArcFaceEffB6](https://www.kaggle.com/librauee/train-arcfaceeffb6) from @librauee \n\nThe TFRecords example is from a [dataset](https://www.kaggle.com/manojprabhaakr/whale-tfrecords-512) made available by @manojprabhaakr \n\nThe only real downside to this is that the link will expire after about a week, however all you have to do is run the notebook again in Kaggle ... get the new GCS_DS_PATH, put it in Colab and resume using Colab. I have Google drive and it is possible to link your Colab to your Google drive. \n\n```\nfrom google.colab import drive\ndrive.mount('/content/drive')\n```\n\nI would not do that as a way of accessing large amounts of data though. It's way too slow. I use it for smaller files like the sample_submission.csv file so I don't have to worry about that path link expiring. \n\nLinking your Google drive is also very important for one other reason. When Colab finishes it's run it may expire the session if you don't get back to it in time. To avoid losing all your results and files, I simply put in a line at the end that copies all the files I care about over to my Google drive, with something like...\n\n`!cp -R /content/working /content/drive/MyDrive/colab_models/WHALE_Model`\n\nIt will depend on how you set up the Colab and Google drives of course.\n\nIf you have a look at some of the TPU based note books you will see some of the things I have mentioned.\n\nHope that helps",
    "1698834": "BruceYoung I really appreciate the answer you have given. I downloaded the dataset for the competition into Colab using Kaggle's API, but instead of downloading a single train.zip, only about 30 individually zipped images were downloaded. Has anyone else experienced a similar experience using the Kaggle API?",
    "1700603": "hi, i am getting this error when i run the code in the kaggle notebook :\n```\nBackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'whale-tfrecords-512', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5, 'details': []}, 'wasSuccessful': False}.\n```",
    "1700661": "If you are using the notebook for training that I linked to you'll have to either add in the dataset for `whale-tfrecords-512` or just use the dataset he has `happywhale-tfrecords-v1`\n\nThen just make sure this...\n\n`GCS_DS_PATH=KaggleDatasets().get_gcs_path('whale-tfrecords-512')`\n\n...has the right info for the dataset you are using.\n\nIf that is not what you are having trouble with,  just let me know.",
    "1700672": "I don't think I've used the API more than about once and that was years ago. I download the lot as a zip file and upload it to my Googe drive. This is a very slow. It is also a slow way to use large amounts of data. You can, of course, copy the data you need from a Google drive to a colab drive but again, that is slow and data on your colab drive is not persistent. I got away with this in the COTS comp as only part of the data was needed so not all needed to be copied from my Google drive for each session.\n\nAll this is why I love using a TPU with a link to the original data.",
    "1700684": "Thank you for your explanation. It is very helpful. I've understood why tfrecords is important for TPU as well as how it could be imported in google colab.\n\nI am wondering is that approach (`KaggleDatasets().get_gcs_path('happy-whale-and-dolphin')`) is suitable for non-TPU training?",
    "1700831": "thank you so much! this worked!"
  },
  "source": "meta"
}