{
  "id": 271772,
  "title": "Best way to connect Kaggle data to Colab?",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/271772",
  "author_name": "",
  "post_date": "2021-09-12T14:24:38.562191600Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>So far, I have identified these two strategies: </p>\n<ol>\n<li>Copy the data from the <a href=\"https://cloud.google.com/storage/docs/creating-buckets\" target=\"_blank\">GCS bucket </a> to the attached hard drive. For example using: </li>\n</ol>\n<p><code>gsutil cp gs://BUCKET_NAME/OBJECT_NAME SAVE_TO_LOCATION</code></p>\n<ol>\n<li>Directly using the data from the GCS bucket. For example using the following package: <a href=\"https://fs-gcsfs.readthedocs.io/en/latest/\" target=\"_blank\">https://fs-gcsfs.readthedocs.io/en/latest/</a> or maybe <a href=\"https://gcsfs.readthedocs.io/en/latest/\" target=\"_blank\">https://gcsfs.readthedocs.io/en/latest/</a></li>\n</ol>\n<p>The first method copies data so that they will be faster to read later but will be removed once the session is recycled. The second is slower but doesn't need the inital copy.</p>\n<p>Which other methods have you tried and which ones are the best in terms of transfer speed and convenience. Thanks for the help!</p>",
  "messages": [
    {
      "id": "1510556",
      "postDate": "09/12/2021 14:24:38",
      "content": "<p>So far, I have identified these two strategies: </p>\n<ol>\n<li>Copy the data from the <a href=\"https://cloud.google.com/storage/docs/creating-buckets\" target=\"_blank\">GCS bucket </a> to the attached hard drive. For example using: </li>\n</ol>\n<p><code>gsutil cp gs://BUCKET_NAME/OBJECT_NAME SAVE_TO_LOCATION</code></p>\n<ol>\n<li>Directly using the data from the GCS bucket. For example using the following package: <a href=\"https://fs-gcsfs.readthedocs.io/en/latest/\" target=\"_blank\">https://fs-gcsfs.readthedocs.io/en/latest/</a> or maybe <a href=\"https://gcsfs.readthedocs.io/en/latest/\" target=\"_blank\">https://gcsfs.readthedocs.io/en/latest/</a></li>\n</ol>\n<p>The first method copies data so that they will be faster to read later but will be removed once the session is recycled. The second is slower but doesn't need the inital copy.</p>\n<p>Which other methods have you tried and which ones are the best in terms of transfer speed and convenience. Thanks for the help!</p>",
      "rawMarkdown": "So far, I have identified these two strategies: \n\n1. Copy the data from the [GCS bucket ](https://cloud.google.com/storage/docs/creating-buckets) to the attached hard drive. For example using: \n\n`gsutil cp gs://BUCKET_NAME/OBJECT_NAME SAVE_TO_LOCATION`\n\n\n2. Directly using the data from the GCS bucket. For example using the following package: https://fs-gcsfs.readthedocs.io/en/latest/ or maybe https://gcsfs.readthedocs.io/en/latest/\n\nThe first method copies data so that they will be faster to read later but will be removed once the session is recycled. The second is slower but doesn't need the inital copy.\n\nWhich other methods have you tried and which ones are the best in terms of transfer speed and convenience. Thanks for the help!",
      "votes": null
    },
    {
      "id": "1510635",
      "postDate": "09/12/2021 15:39:27",
      "content": "<p>You can use the TFRecord pipeline … It will be much faster I guess </p>",
      "rawMarkdown": "You can use the TFRecord pipeline ... It will be much faster I guess",
      "votes": null
    },
    {
      "id": "1510642",
      "postDate": "09/12/2021 15:47:16",
      "content": "<p>That's one option indeed, does this exist for PyTorch as well? 👀</p>",
      "rawMarkdown": "That's one option indeed, does this exist for PyTorch as well? 👀",
      "votes": null
    },
    {
      "id": "1510661",
      "postDate": "09/12/2021 16:11:32",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch</a></p>",
      "rawMarkdown": "https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch",
      "votes": null
    },
    {
      "id": "1510729",
      "postDate": "09/12/2021 17:22:42",
      "content": "<p>Thanks for the link! 👌</p>",
      "rawMarkdown": "Thanks for the link! 👌",
      "votes": null
    },
    {
      "id": "1559703",
      "postDate": "10/27/2021 07:07:50",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1510635,
      "author_name": "benihime91",
      "author_url": "",
      "post_date": "09/12/2021 15:39:27",
      "content": "<p>You can use the TFRecord pipeline … It will be much faster I guess </p>",
      "votes": null,
      "replies": [
        {
          "id": 1510642,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/12/2021 15:47:16",
          "content": "<p>That's one option indeed, does this exist for PyTorch as well? 👀</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1510661,
          "author_name": "benihime91",
          "author_url": "",
          "post_date": "09/12/2021 16:11:32",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1510729,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/12/2021 17:22:42",
          "content": "<p>Thanks for the link! 👌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559703,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:07:50",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1510556": "So far, I have identified these two strategies: \n\n1. Copy the data from the [GCS bucket ](https://cloud.google.com/storage/docs/creating-buckets) to the attached hard drive. For example using: \n\n`gsutil cp gs://BUCKET_NAME/OBJECT_NAME SAVE_TO_LOCATION`\n\n\n2. Directly using the data from the GCS bucket. For example using the following package: https://fs-gcsfs.readthedocs.io/en/latest/ or maybe https://gcsfs.readthedocs.io/en/latest/\n\nThe first method copies data so that they will be faster to read later but will be removed once the session is recycled. The second is slower but doesn't need the inital copy.\n\nWhich other methods have you tried and which ones are the best in terms of transfer speed and convenience. Thanks for the help!",
    "1510635": "You can use the TFRecord pipeline ... It will be much faster I guess",
    "1510642": "That's one option indeed, does this exist for PyTorch as well? 👀",
    "1510661": "https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch",
    "1510729": "Thanks for the link! 👌",
    "1559703": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}