{
  "id": 253912,
  "title": "Train Only Downloading (Colab)",
  "url": "/competitions/seti-breakthrough-listen/discussion/253912",
  "author_name": "kkiller",
  "post_date": "2021-07-19T10:15:35.669000",
  "votes": 38,
  "comment_count": 10,
  "views": 0,
  "content": "<p>As the whole dataset (new + old) is just enormous, we can hardly download it on small sized VMs, especially on Google Colab. I manage to find a workaround that let me download the train only. Let recall that, it's very difficult to do it within the official Kaggle API as one would have to make an API call  for each single file inside the train folder.</p>\n<pre><code>GCS_DS_PATH = \"gs://kds-c74c1b10274be658ad93d6c985165f826d0521eab9365a153a709320\"\n!gsutil -m cp -r {GCS_DS_PATH}/\"train\" .\n</code></pre>\n<p>The <strong>GCS_DS_PATH</strong> changes roughly each week. You can get the new one by yourself using the <strong>kaggle_datasets</strong> API inside a Kaggle notebook.</p>\n<p>For updated <strong>GCS_DS_PATH</strong>, please see the comment section.</p>",
  "messages": [
    {
      "id": 1393026,
      "postDate": "2021-07-19T10:15:35.670Z",
      "content": "<p>As the whole dataset (new + old) is just enormous, we can hardly download it on small sized VMs, especially on Google Colab. I manage to find a workaround that let me download the train only. Let recall that, it's very difficult to do it within the official Kaggle API as one would have to make an API call  for each single file inside the train folder.</p>\n<pre><code>GCS_DS_PATH = \"gs://kds-c74c1b10274be658ad93d6c985165f826d0521eab9365a153a709320\"\n!gsutil -m cp -r {GCS_DS_PATH}/\"train\" .\n</code></pre>\n<p>The <strong>GCS_DS_PATH</strong> changes roughly each week. You can get the new one by yourself using the <strong>kaggle_datasets</strong> API inside a Kaggle notebook.</p>\n<p>For updated <strong>GCS_DS_PATH</strong>, please see the comment section.</p>",
      "rawMarkdown": "As the whole dataset (new + old) is just enormous, we can hardly download it on small sized VMs, especially on Google Colab. I manage to find a workaround that let me download the train only. Let recall that, it's very difficult to do it within the official Kaggle API as one would have to make an API call  for each single file inside the train folder.\n\n```\nGCS_DS_PATH = \"gs://kds-c74c1b10274be658ad93d6c985165f826d0521eab9365a153a709320\"\n!gsutil -m cp -r {GCS_DS_PATH}/\"train\" .\n```\n\nThe **GCS_DS_PATH** changes roughly each week. You can get the new one by yourself using the **kaggle_datasets** API inside a Kaggle notebook.\n\nFor updated **GCS_DS_PATH**, please see the comment section.",
      "votes": 37
    },
    {
      "id": 1447489,
      "postDate": "2021-08-04T15:04:39.007Z",
      "content": "<p>Updated: <code>GCS_DS_PATH = \"gs://kds-06ebe33f115002b87e00190a4afa61624bf9aaf3b7f192044f43ffef\"</code></p>",
      "rawMarkdown": "Updated: `GCS_DS_PATH = \"gs://kds-06ebe33f115002b87e00190a4afa61624bf9aaf3b7f192044f43ffef\"`",
      "votes": 4
    },
    {
      "id": 1470003,
      "postDate": "2021-08-13T08:18:06.910Z",
      "content": "<p>Updated <code>GCS_DS_PATH = gs://kds-03e1a10468eed52cf1f7962028d75ca9e9ca67e4b3c30a4f640075f4</code></p>\n<p>You can run the following to get the <code>GCS_DS_PATH</code>:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nGCS_DS_PATH = KaggleDatasets().get_gcs_path(\"seti-breakthrough-listen\")\nprint(GCS_DS_PATH)\n</code></pre>",
      "rawMarkdown": "Updated `GCS_DS_PATH = gs://kds-03e1a10468eed52cf1f7962028d75ca9e9ca67e4b3c30a4f640075f4`\n\nYou can run the following to get the `GCS_DS_PATH`:\n\n    from kaggle_datasets import KaggleDatasets\n    GCS_DS_PATH = KaggleDatasets().get_gcs_path(\"seti-breakthrough-listen\")\n    print(GCS_DS_PATH)",
      "votes": 2
    },
    {
      "id": 1397409,
      "postDate": "2021-07-23T07:13:20.950Z",
      "content": "<p>Updated: <code>GCS_DS_PATH = \"gs://kds-796f503b3e2340cd6be862cdcf9c2040801cbc9be18c87ded69657a0\"</code></p>",
      "rawMarkdown": "Updated: `GCS_DS_PATH = \"gs://kds-796f503b3e2340cd6be862cdcf9c2040801cbc9be18c87ded69657a0\"`",
      "votes": 2,
      "replies": [
        {
          "id": 1397644,
          "postDate": "2021-07-23T11:27:57.863Z",
          "content": "<p>Good job <a href=\"https://www.kaggle.com/valleyzw\" target=\"_blank\">@valleyzw</a> </p>",
          "rawMarkdown": "Good job @valleyzw ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1393139,
      "postDate": "2021-07-19T12:12:11.463Z",
      "content": "<p>Great sharing <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> </p>",
      "rawMarkdown": "Great sharing @kneroma ",
      "votes": 2
    },
    {
      "id": 1477633,
      "postDate": "2021-08-17T14:59:50.637Z",
      "content": "<p>UPDATED: <code>gs://kds-bd53c2c14a8944835e167bf11ad1a3e3d9791601e8395c52e025c019</code></p>",
      "rawMarkdown": "UPDATED: `gs://kds-bd53c2c14a8944835e167bf11ad1a3e3d9791601e8395c52e025c019`"
    },
    {
      "id": 1471223,
      "postDate": "2021-08-14T04:05:20.133Z",
      "content": "<p>Thanks for sharing!<br>\nFYI, this took about 30 minutes when I downloaded the new train set in Colab (it should depend on the region of the instance/dataset).</p>",
      "rawMarkdown": "Thanks for sharing!\nFYI, this took about 30 minutes when I downloaded the new train set in Colab (it should depend on the region of the instance/dataset)."
    },
    {
      "id": 1447995,
      "postDate": "2021-08-04T16:28:54.680Z",
      "content": "<p>If you use TPU with TF on Colab, you don't even need to download the data. Just specify GCS_DS_PATH directly, without kaggle_datasets.</p>",
      "rawMarkdown": "If you use TPU with TF on Colab, you don't even need to download the data. Just specify GCS_DS_PATH directly, without kaggle_datasets.",
      "replies": [
        {
          "id": 1471211,
          "postDate": "2021-08-14T03:48:45.207Z",
          "content": "<p>Can you talk clearly about it? I don't know how to load directly</p>",
          "rawMarkdown": "Can you talk clearly about it? I don't know how to load directly"
        }
      ]
    },
    {
      "id": 1393584,
      "postDate": "2021-07-19T18:31:26.607Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1447489,
      "author_name": "Daniel Furman",
      "author_url": "",
      "post_date": "2021-08-04T15:04:39.007000",
      "content": "<p>Updated: <code>GCS_DS_PATH = \"gs://kds-06ebe33f115002b87e00190a4afa61624bf9aaf3b7f192044f43ffef\"</code></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1470003,
      "author_name": "Sanand M V",
      "author_url": "",
      "post_date": "2021-08-13T08:18:06.910000",
      "content": "<p>Updated <code>GCS_DS_PATH = gs://kds-03e1a10468eed52cf1f7962028d75ca9e9ca67e4b3c30a4f640075f4</code></p>\n<p>You can run the following to get the <code>GCS_DS_PATH</code>:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nGCS_DS_PATH = KaggleDatasets().get_gcs_path(\"seti-breakthrough-listen\")\nprint(GCS_DS_PATH)\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1397409,
      "author_name": "valley",
      "author_url": "",
      "post_date": "2021-07-23T07:13:20.950000",
      "content": "<p>Updated: <code>GCS_DS_PATH = \"gs://kds-796f503b3e2340cd6be862cdcf9c2040801cbc9be18c87ded69657a0\"</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1397644,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-07-23T11:27:57.863000",
          "content": "<p>Good job <a href=\"https://www.kaggle.com/valleyzw\" target=\"_blank\">@valleyzw</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1393139,
      "author_name": "Ulrich G.",
      "author_url": "",
      "post_date": "2021-07-19T12:12:11.463000",
      "content": "<p>Great sharing <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1477633,
      "author_name": "Shion Honda",
      "author_url": "",
      "post_date": "2021-08-17T14:59:50.637000",
      "content": "<p>UPDATED: <code>gs://kds-bd53c2c14a8944835e167bf11ad1a3e3d9791601e8395c52e025c019</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1471223,
      "author_name": "Shion Honda",
      "author_url": "",
      "post_date": "2021-08-14T04:05:20.133000",
      "content": "<p>Thanks for sharing!<br>\nFYI, this took about 30 minutes when I downloaded the new train set in Colab (it should depend on the region of the instance/dataset).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1447995,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2021-08-04T16:28:54.680000",
      "content": "<p>If you use TPU with TF on Colab, you don't even need to download the data. Just specify GCS_DS_PATH directly, without kaggle_datasets.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1471211,
          "author_name": "Cloudyy",
          "author_url": "",
          "post_date": "2021-08-14T03:48:45.207000",
          "content": "<p>Can you talk clearly about it? I don't know how to load directly</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1393584,
      "author_name": "Mohamed Hany",
      "author_url": "",
      "post_date": "2021-07-19T18:31:26.607000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1393026": "As the whole dataset (new + old) is just enormous, we can hardly download it on small sized VMs, especially on Google Colab. I manage to find a workaround that let me download the train only. Let recall that, it's very difficult to do it within the official Kaggle API as one would have to make an API call  for each single file inside the train folder.\n\n```\nGCS_DS_PATH = \"gs://kds-c74c1b10274be658ad93d6c985165f826d0521eab9365a153a709320\"\n!gsutil -m cp -r {GCS_DS_PATH}/\"train\" .\n```\n\nThe **GCS_DS_PATH** changes roughly each week. You can get the new one by yourself using the **kaggle_datasets** API inside a Kaggle notebook.\n\nFor updated **GCS_DS_PATH**, please see the comment section.",
    "1447489": "Updated: `GCS_DS_PATH = \"gs://kds-06ebe33f115002b87e00190a4afa61624bf9aaf3b7f192044f43ffef\"`",
    "1470003": "Updated `GCS_DS_PATH = gs://kds-03e1a10468eed52cf1f7962028d75ca9e9ca67e4b3c30a4f640075f4`\n\nYou can run the following to get the `GCS_DS_PATH`:\n\n    from kaggle_datasets import KaggleDatasets\n    GCS_DS_PATH = KaggleDatasets().get_gcs_path(\"seti-breakthrough-listen\")\n    print(GCS_DS_PATH)",
    "1397409": "Updated: `GCS_DS_PATH = \"gs://kds-796f503b3e2340cd6be862cdcf9c2040801cbc9be18c87ded69657a0\"`",
    "1393139": "Great sharing @kneroma ",
    "1477633": "UPDATED: `gs://kds-bd53c2c14a8944835e167bf11ad1a3e3d9791601e8395c52e025c019`",
    "1471223": "Thanks for sharing!\nFYI, this took about 30 minutes when I downloaded the new train set in Colab (it should depend on the region of the instance/dataset).",
    "1447995": "If you use TPU with TF on Colab, you don't even need to download the data. Just specify GCS_DS_PATH directly, without kaggle_datasets.",
    "1393584": "Thanks for sharing !"
  }
}