{
  "id": 201549,
  "title": "Create Kaggle dataset from files in Google Colab",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201549",
  "author_name": "",
  "post_date": "2020-12-05T14:47:30.250953700Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was trying to do some data processing and save it in TFRecord but couldn't due to the limited disk space of Kaggle's notebooks. Therefore, I'm bringing the notebook to Google Colab which works like a champ with much greater capacity.</p>\n<p>However, now I'm having an issue with bring the notebook's output from Colab back to Kaggle, preferably as a dataset in Kaggle. One obvious solution is to download the TFRecord files to local, then update them to Kaggle, but I'm hoping for alternative solutions (since the files take up ~50-60 GBs, which isn't easy to download/upload).</p>\n<p>So, my questions are:</p>\n<p>Is there a way to export files/outputs from Google Colab to Kaggle/Kaggle Datasets?<br>\nor<br>\nIs there a platform (with large disk space) where I can preprocess data, save it and then import back to Kaggle?<br>\nThanks everyone!</p>",
  "messages": [
    {
      "id": "1102995",
      "postDate": "12/05/2020 14:47:30",
      "content": "<p>I was trying to do some data processing and save it in TFRecord but couldn't due to the limited disk space of Kaggle's notebooks. Therefore, I'm bringing the notebook to Google Colab which works like a champ with much greater capacity.</p>\n<p>However, now I'm having an issue with bring the notebook's output from Colab back to Kaggle, preferably as a dataset in Kaggle. One obvious solution is to download the TFRecord files to local, then update them to Kaggle, but I'm hoping for alternative solutions (since the files take up ~50-60 GBs, which isn't easy to download/upload).</p>\n<p>So, my questions are:</p>\n<p>Is there a way to export files/outputs from Google Colab to Kaggle/Kaggle Datasets?<br>\nor<br>\nIs there a platform (with large disk space) where I can preprocess data, save it and then import back to Kaggle?<br>\nThanks everyone!</p>",
      "rawMarkdown": "I was trying to do some data processing and save it in TFRecord but couldn't due to the limited disk space of Kaggle's notebooks. Therefore, I'm bringing the notebook to Google Colab which works like a champ with much greater capacity.\n\nHowever, now I'm having an issue with bring the notebook's output from Colab back to Kaggle, preferably as a dataset in Kaggle. One obvious solution is to download the TFRecord files to local, then update them to Kaggle, but I'm hoping for alternative solutions (since the files take up ~50-60 GBs, which isn't easy to download/upload).\n\nSo, my questions are:\n\nIs there a way to export files/outputs from Google Colab to Kaggle/Kaggle Datasets?\nor\nIs there a platform (with large disk space) where I can preprocess data, save it and then import back to Kaggle?\nThanks everyone!",
      "votes": null
    },
    {
      "id": "1103028",
      "postDate": "12/05/2020 15:23:12",
      "content": "<p>EDIT: Now that i understood your requirement better, Just create a dataset using </p>\n<blockquote>\n  <p>!kaggle datasets create -p /your_output_files</p>\n</blockquote>\n<p>Make sure to store your api key in colab.</p>",
      "rawMarkdown": "EDIT: Now that i understood your requirement better, Just create a dataset using \n>!kaggle datasets create -p /your_output_files\n\nMake sure to store your api key in colab.",
      "votes": null
    },
    {
      "id": "1103040",
      "postDate": "12/05/2020 15:32:20",
      "content": "<p>You can download files from google drive to kaggle kernels (much faster than uploading, however you are limited to 20GB - I think) .  You can use output from kaggle kernels in your kaggle notebooks rather than kaggle dataset. </p>\n<p>Here's one script <a href=\"https://www.kaggle.com/rashmibanthia/google-drive-to-kernel\" target=\"_blank\">https://www.kaggle.com/rashmibanthia/google-drive-to-kernel</a> or try package <code>gdown</code>.  </p>\n<p>I wonder why do you have 50-60GB of data for this competition  🤔? </p>",
      "rawMarkdown": "You can download files from google drive to kaggle kernels (much faster than uploading, however you are limited to 20GB - I think) .  You can use output from kaggle kernels in your kaggle notebooks rather than kaggle dataset. \n\nHere's one script https://www.kaggle.com/rashmibanthia/google-drive-to-kernel or try package `gdown`.  \n\nI wonder why do you have 50-60GB of data for this competition  🤔?",
      "votes": null
    },
    {
      "id": "1103071",
      "postDate": "12/05/2020 15:56:34",
      "content": "<p>I appreciate your suggestion, though 20GB doesn't seem to work very efficiently for my case. </p>\n<p>My total dataset takes up about 150GB. I follow the preprocessing steps shown in <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">this notebook</a>. Long story short is that for each user, I create a sliding window of size 64 with a question as the last time step, so it's like the whole train set is repeated roughly 64 times, resulting in the exploding disk usage.</p>\n<p>I hope to have a better way to handle this, so feel free to comment any suggestions you have :) </p>\n<p>Thanks again!</p>",
      "rawMarkdown": "I appreciate your suggestion, though 20GB doesn't seem to work very efficiently for my case. \n\nMy total dataset takes up about 150GB. I follow the preprocessing steps shown in [this notebook](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public). Long story short is that for each user, I create a sliding window of size 64 with a question as the last time step, so it's like the whole train set is repeated roughly 64 times, resulting in the exploding disk usage.\n\nI hope to have a better way to handle this, so feel free to comment any suggestions you have :) \n\nThanks again!",
      "votes": null
    },
    {
      "id": "1103076",
      "postDate": "12/05/2020 16:01:33",
      "content": "<p>I wonder if this can be done directly in Kaggle notebook. Particularly I'm thinking that step 1 can be done in Kaggle instead of Colab and output will be save to Kaggle's /tmp scratch space, then I'll skip to step 3. This seems to be a nice thing to try out - thanks for the suggestion!</p>\n<p>Do you happen to know the size of the scratch space? </p>",
      "rawMarkdown": "I wonder if this can be done directly in Kaggle notebook. Particularly I'm thinking that step 1 can be done in Kaggle instead of Colab and output will be save to Kaggle's /tmp scratch space, then I'll skip to step 3. This seems to be a nice thing to try out - thanks for the suggestion!\n\nDo you happen to know the size of the scratch space?",
      "votes": null
    },
    {
      "id": "1103107",
      "postDate": "12/05/2020 16:34:05",
      "content": "<p>Please see my edit above.</p>",
      "rawMarkdown": "Please see my edit above.",
      "votes": null
    },
    {
      "id": "1103227",
      "postDate": "12/05/2020 18:35:33",
      "content": "<p><a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>, you can create a kaggle dataset directly from your colab notebook by using the kaggle API (see link below).</p>\n<p><a href=\"https://github.com/Kaggle/kaggle-api#initialize-metadata-file-for-dataset-creation\" target=\"_blank\">https://github.com/Kaggle/kaggle-api#initialize-metadata-file-for-dataset-creation</a></p>\n<p>Example:-</p>\n<p>!kaggle datasets init -p /path/to/dataset</p>",
      "rawMarkdown": "hoangnguyen719, you can create a kaggle dataset directly from your colab notebook by using the kaggle API (see link below).\n\nhttps://github.com/Kaggle/kaggle-api#initialize-metadata-file-for-dataset-creation\n\nExample:-\n\n!kaggle datasets init -p /path/to/dataset",
      "votes": null
    },
    {
      "id": "1103248",
      "postDate": "12/05/2020 19:00:26",
      "content": "<p>I believe, you should increase the seq_len, as you cannot create a dataset for 150GB here. Max is 100GB unless you make it public. (Go in Account Section to check the free capacity)</p>\n<p>Plus, starting with something simpler like splitting the sequences simply into len's of 128 might be a better option;</p>",
      "rawMarkdown": "I believe, you should increase the seq_len, as you cannot create a dataset for 150GB here. Max is 100GB unless you make it public. (Go in Account Section to check the free capacity)\n\nPlus, starting with something simpler like splitting the sequences simply into len's of 128 might be a better option;",
      "votes": null
    },
    {
      "id": "1103256",
      "postDate": "12/05/2020 19:06:18",
      "content": "<p><a href=\"https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end#Update-Submission-Dataset\" target=\"_blank\">https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end#Update-Submission-Dataset</a></p>\n<pre><code>    Path(KAGGLE_SUBMIT_DATASET).mkdir(exist_ok=True)\n\n    model.save_model(f'{KAGGLE_SUBMIT_DATASET}/model.xgb')\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/columns.json', 'w') as cj:\n            json.dump(columns_train, cj)\n\n    df_files = {\n        # 'df_users.pkl': df_users,\n        'df_users_content.pkl': df_users_content,\n        'df_questions.pkl': df_questions,\n    }\n\n    for file_path, df in df_files.items():\n        df.to_pickle(f'{KAGGLE_SUBMIT_DATASET}/{file_path}')\n\n    kaggle_id = f\"{os.getenv('KAGGLE_USERNAME')}/{KAGGLE_SUBMIT_DATASET}\"\n\n    metadata = {\n        \"licenses\": [{\"name\": \"CC0-1.0\"}],\n        \"id\": kaggle_id,\n        \"title\": KAGGLE_SUBMIT_DATASET\n           }\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/dataset-metadata.json', 'w') as f:\n        json.dump(metadata, f)\n\n    if kaggle_api.dataset_status(kaggle_id):\n        kaggle_api.dataset_create_version(KAGGLE_SUBMIT_DATASET,\n                                          version_notes='update dataset',\n                                          delete_old_versions=True,\n                                          dir_mode='tar',\n                                          quiet=True\n                                         )\n    else:\n        kaggle_api.dataset_create_new(KAGGLE_SUBMIT_DATASET,\n                                      dir_mode='tar', quiet=True)\n</code></pre>",
      "rawMarkdown": "https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end#Update-Submission-Dataset\n\n```\n    Path(KAGGLE_SUBMIT_DATASET).mkdir(exist_ok=True)\n\n    model.save_model(f'{KAGGLE_SUBMIT_DATASET}/model.xgb')\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/columns.json', 'w') as cj:\n            json.dump(columns_train, cj)\n    \n    df_files = {\n        # 'df_users.pkl': df_users,\n        'df_users_content.pkl': df_users_content,\n        'df_questions.pkl': df_questions,\n    }\n\n    for file_path, df in df_files.items():\n        df.to_pickle(f'{KAGGLE_SUBMIT_DATASET}/{file_path}')\n            \n    kaggle_id = f\"{os.getenv('KAGGLE_USERNAME')}/{KAGGLE_SUBMIT_DATASET}\"\n    \n    metadata = {\n        \"licenses\": [{\"name\": \"CC0-1.0\"}],\n        \"id\": kaggle_id,\n        \"title\": KAGGLE_SUBMIT_DATASET\n           }\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/dataset-metadata.json', 'w') as f:\n        json.dump(metadata, f)\n            \n    if kaggle_api.dataset_status(kaggle_id):\n        kaggle_api.dataset_create_version(KAGGLE_SUBMIT_DATASET,\n                                          version_notes='update dataset',\n                                          delete_old_versions=True,\n                                          dir_mode='tar',\n                                          quiet=True\n                                         )\n    else:\n        kaggle_api.dataset_create_new(KAGGLE_SUBMIT_DATASET,\n                                      dir_mode='tar', quiet=True)\n```",
      "votes": null
    },
    {
      "id": "1104140",
      "postDate": "12/06/2020 16:48:36",
      "content": "<p>You're correct, I'm going to make it public as I plan to use TPU also. Btw, given the preprocessing method mentioned above, increasing windows will actually increase data size </p>",
      "rawMarkdown": "You're correct, I'm going to make it public as I plan to use TPU also. Btw, given the preprocessing method mentioned above, increasing windows will actually increase data size",
      "votes": null
    },
    {
      "id": "1104141",
      "postDate": "12/06/2020 16:50:18",
      "content": "<p>Yep you're right, that's exactly what I've done (though I also needed to split it into 3 Kaggle datasets since the whole dataset ~ 180GB while Colab (free version) allows ~75GB only). Thanks a lot for the help!</p>",
      "rawMarkdown": "Yep you're right, that's exactly what I've done (though I also needed to split it into 3 Kaggle datasets since the whole dataset ~ 180GB while Colab (free version) allows ~75GB only). Thanks a lot for the help!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1103028,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "12/05/2020 15:23:12",
      "content": "<p>EDIT: Now that i understood your requirement better, Just create a dataset using </p>\n<blockquote>\n  <p>!kaggle datasets create -p /your_output_files</p>\n</blockquote>\n<p>Make sure to store your api key in colab.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103076,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "12/05/2020 16:01:33",
          "content": "<p>I wonder if this can be done directly in Kaggle notebook. Particularly I'm thinking that step 1 can be done in Kaggle instead of Colab and output will be save to Kaggle's /tmp scratch space, then I'll skip to step 3. This seems to be a nice thing to try out - thanks for the suggestion!</p>\n<p>Do you happen to know the size of the scratch space? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103107,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "12/05/2020 16:34:05",
          "content": "<p>Please see my edit above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104141,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "12/06/2020 16:50:18",
          "content": "<p>Yep you're right, that's exactly what I've done (though I also needed to split it into 3 Kaggle datasets since the whole dataset ~ 180GB while Colab (free version) allows ~75GB only). Thanks a lot for the help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103040,
      "author_name": "rashmibanthia",
      "author_url": "",
      "post_date": "12/05/2020 15:32:20",
      "content": "<p>You can download files from google drive to kaggle kernels (much faster than uploading, however you are limited to 20GB - I think) .  You can use output from kaggle kernels in your kaggle notebooks rather than kaggle dataset. </p>\n<p>Here's one script <a href=\"https://www.kaggle.com/rashmibanthia/google-drive-to-kernel\" target=\"_blank\">https://www.kaggle.com/rashmibanthia/google-drive-to-kernel</a> or try package <code>gdown</code>.  </p>\n<p>I wonder why do you have 50-60GB of data for this competition  🤔? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1103071,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "12/05/2020 15:56:34",
          "content": "<p>I appreciate your suggestion, though 20GB doesn't seem to work very efficiently for my case. </p>\n<p>My total dataset takes up about 150GB. I follow the preprocessing steps shown in <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">this notebook</a>. Long story short is that for each user, I create a sliding window of size 64 with a question as the last time step, so it's like the whole train set is repeated roughly 64 times, resulting in the exploding disk usage.</p>\n<p>I hope to have a better way to handle this, so feel free to comment any suggestions you have :) </p>\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103248,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/05/2020 19:00:26",
          "content": "<p>I believe, you should increase the seq_len, as you cannot create a dataset for 150GB here. Max is 100GB unless you make it public. (Go in Account Section to check the free capacity)</p>\n<p>Plus, starting with something simpler like splitting the sequences simply into len's of 128 might be a better option;</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104140,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "12/06/2020 16:48:36",
          "content": "<p>You're correct, I'm going to make it public as I plan to use TPU also. Btw, given the preprocessing method mentioned above, increasing windows will actually increase data size </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103227,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "12/05/2020 18:35:33",
      "content": "<p><a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>, you can create a kaggle dataset directly from your colab notebook by using the kaggle API (see link below).</p>\n<p><a href=\"https://github.com/Kaggle/kaggle-api#initialize-metadata-file-for-dataset-creation\" target=\"_blank\">https://github.com/Kaggle/kaggle-api#initialize-metadata-file-for-dataset-creation</a></p>\n<p>Example:-</p>\n<p>!kaggle datasets init -p /path/to/dataset</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103256,
      "author_name": "calebeverett",
      "author_url": "",
      "post_date": "12/05/2020 19:06:18",
      "content": "<p><a href=\"https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end#Update-Submission-Dataset\" target=\"_blank\">https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end#Update-Submission-Dataset</a></p>\n<pre><code>    Path(KAGGLE_SUBMIT_DATASET).mkdir(exist_ok=True)\n\n    model.save_model(f'{KAGGLE_SUBMIT_DATASET}/model.xgb')\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/columns.json', 'w') as cj:\n            json.dump(columns_train, cj)\n\n    df_files = {\n        # 'df_users.pkl': df_users,\n        'df_users_content.pkl': df_users_content,\n        'df_questions.pkl': df_questions,\n    }\n\n    for file_path, df in df_files.items():\n        df.to_pickle(f'{KAGGLE_SUBMIT_DATASET}/{file_path}')\n\n    kaggle_id = f\"{os.getenv('KAGGLE_USERNAME')}/{KAGGLE_SUBMIT_DATASET}\"\n\n    metadata = {\n        \"licenses\": [{\"name\": \"CC0-1.0\"}],\n        \"id\": kaggle_id,\n        \"title\": KAGGLE_SUBMIT_DATASET\n           }\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/dataset-metadata.json', 'w') as f:\n        json.dump(metadata, f)\n\n    if kaggle_api.dataset_status(kaggle_id):\n        kaggle_api.dataset_create_version(KAGGLE_SUBMIT_DATASET,\n                                          version_notes='update dataset',\n                                          delete_old_versions=True,\n                                          dir_mode='tar',\n                                          quiet=True\n                                         )\n    else:\n        kaggle_api.dataset_create_new(KAGGLE_SUBMIT_DATASET,\n                                      dir_mode='tar', quiet=True)\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1102995": "I was trying to do some data processing and save it in TFRecord but couldn't due to the limited disk space of Kaggle's notebooks. Therefore, I'm bringing the notebook to Google Colab which works like a champ with much greater capacity.\n\nHowever, now I'm having an issue with bring the notebook's output from Colab back to Kaggle, preferably as a dataset in Kaggle. One obvious solution is to download the TFRecord files to local, then update them to Kaggle, but I'm hoping for alternative solutions (since the files take up ~50-60 GBs, which isn't easy to download/upload).\n\nSo, my questions are:\n\nIs there a way to export files/outputs from Google Colab to Kaggle/Kaggle Datasets?\nor\nIs there a platform (with large disk space) where I can preprocess data, save it and then import back to Kaggle?\nThanks everyone!",
    "1103028": "EDIT: Now that i understood your requirement better, Just create a dataset using \n>!kaggle datasets create -p /your_output_files\n\nMake sure to store your api key in colab.",
    "1103040": "You can download files from google drive to kaggle kernels (much faster than uploading, however you are limited to 20GB - I think) .  You can use output from kaggle kernels in your kaggle notebooks rather than kaggle dataset. \n\nHere's one script https://www.kaggle.com/rashmibanthia/google-drive-to-kernel or try package `gdown`.  \n\nI wonder why do you have 50-60GB of data for this competition  🤔?",
    "1103071": "I appreciate your suggestion, though 20GB doesn't seem to work very efficiently for my case. \n\nMy total dataset takes up about 150GB. I follow the preprocessing steps shown in [this notebook](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public). Long story short is that for each user, I create a sliding window of size 64 with a question as the last time step, so it's like the whole train set is repeated roughly 64 times, resulting in the exploding disk usage.\n\nI hope to have a better way to handle this, so feel free to comment any suggestions you have :) \n\nThanks again!",
    "1103076": "I wonder if this can be done directly in Kaggle notebook. Particularly I'm thinking that step 1 can be done in Kaggle instead of Colab and output will be save to Kaggle's /tmp scratch space, then I'll skip to step 3. This seems to be a nice thing to try out - thanks for the suggestion!\n\nDo you happen to know the size of the scratch space?",
    "1103107": "Please see my edit above.",
    "1103227": "hoangnguyen719, you can create a kaggle dataset directly from your colab notebook by using the kaggle API (see link below).\n\nhttps://github.com/Kaggle/kaggle-api#initialize-metadata-file-for-dataset-creation\n\nExample:-\n\n!kaggle datasets init -p /path/to/dataset",
    "1103248": "I believe, you should increase the seq_len, as you cannot create a dataset for 150GB here. Max is 100GB unless you make it public. (Go in Account Section to check the free capacity)\n\nPlus, starting with something simpler like splitting the sequences simply into len's of 128 might be a better option;",
    "1103256": "https://www.kaggle.com/calebeverett/riiid-bigquery-xgboost-end-to-end#Update-Submission-Dataset\n\n```\n    Path(KAGGLE_SUBMIT_DATASET).mkdir(exist_ok=True)\n\n    model.save_model(f'{KAGGLE_SUBMIT_DATASET}/model.xgb')\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/columns.json', 'w') as cj:\n            json.dump(columns_train, cj)\n    \n    df_files = {\n        # 'df_users.pkl': df_users,\n        'df_users_content.pkl': df_users_content,\n        'df_questions.pkl': df_questions,\n    }\n\n    for file_path, df in df_files.items():\n        df.to_pickle(f'{KAGGLE_SUBMIT_DATASET}/{file_path}')\n            \n    kaggle_id = f\"{os.getenv('KAGGLE_USERNAME')}/{KAGGLE_SUBMIT_DATASET}\"\n    \n    metadata = {\n        \"licenses\": [{\"name\": \"CC0-1.0\"}],\n        \"id\": kaggle_id,\n        \"title\": KAGGLE_SUBMIT_DATASET\n           }\n\n    with open(f'{KAGGLE_SUBMIT_DATASET}/dataset-metadata.json', 'w') as f:\n        json.dump(metadata, f)\n            \n    if kaggle_api.dataset_status(kaggle_id):\n        kaggle_api.dataset_create_version(KAGGLE_SUBMIT_DATASET,\n                                          version_notes='update dataset',\n                                          delete_old_versions=True,\n                                          dir_mode='tar',\n                                          quiet=True\n                                         )\n    else:\n        kaggle_api.dataset_create_new(KAGGLE_SUBMIT_DATASET,\n                                      dir_mode='tar', quiet=True)\n```",
    "1104140": "You're correct, I'm going to make it public as I plan to use TPU also. Btw, given the preprocessing method mentioned above, increasing windows will actually increase data size",
    "1104141": "Yep you're right, that's exactly what I've done (though I also needed to split it into 3 Kaggle datasets since the whole dataset ~ 180GB while Colab (free version) allows ~75GB only). Thanks a lot for the help!"
  },
  "source": "meta"
}