{
  "id": 171207,
  "title": "Any way to avoid downloading 106 GB data? ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171207",
  "author_name": "",
  "post_date": "2020-07-30T20:25:16.678316200Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am using local computer and google colab for my code. I used the Kaggle API but it throws error as 429-Too many requests.</p>",
  "messages": [
    {
      "id": "952330",
      "postDate": "07/30/2020 20:25:16",
      "content": "<p>I am using local computer and google colab for my code. I used the Kaggle API but it throws error as 429-Too many requests.</p>",
      "rawMarkdown": "I am using local computer and google colab for my code. I used the Kaggle API but it throws error as 429-Too many requests.",
      "votes": null
    },
    {
      "id": "952432",
      "postDate": "07/30/2020 22:59:32",
      "content": "<p>There is no need to download the entire data set. If you know Google Cloud Storage bucket addresses for tfrecord files you can be transferring  data to your TPU/GPU directly from inside of your Colab notebook. Take a look at my public notebook illustrating the details of this process (BTW, the notebook, with some modifications, can be run on Colab): <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">EfficientNet BN+Tabular Features TF CV5 512x512</a>.</p>\n\n<p>And here are the discussion topic and another public notebook helping you to determine the most recent GCS addresses:</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165739\">An easy and convenient way to look up GCS bucket addresses!</a>\n<a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">SIIM Show GCS bucket addresses</a></p>\n\n<p>Hope it helps!</p>",
      "rawMarkdown": "There is no need to download the entire data set. If you know Google Cloud Storage bucket addresses for tfrecord files you can be transferring  data to your TPU/GPU directly from inside of your Colab notebook. Take a look at my public notebook illustrating the details of this process (BTW, the notebook, with some modifications, can be run on Colab): [EfficientNet BN+Tabular Features TF CV5 512x512](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512).\n\nAnd here are the discussion topic and another public notebook helping you to determine the most recent GCS addresses:\n\n[An easy and convenient way to look up GCS bucket addresses!](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165739)\n[SIIM Show GCS bucket addresses](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses)\n\nHope it helps!",
      "votes": null
    },
    {
      "id": "952446",
      "postDate": "07/30/2020 23:32:15",
      "content": "<p>In addition to Alexey's answer, it would also help to look at the top notebooks for the competition -- typically they forego downloading the dataset and use it straight from Kaggle (or use their own augmented/preprocessed dataset based on the competition set). The most useful one I've found to be is <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">Triple Stratified KFold with TFRecords</a>, but that may just be me. It is TPU-accelerated notebook with the competition data loaded in from a preprocessed user dataset. See <a href=\"/cdeotte\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">Triple Stratified Leak-Free KFold CV</a> for more info on the data preprocessing. When viewing the notebooks, you can fork the notebook and edit it as your own. I've found this really useful in getting started.</p>\n\n<p>Though I'm new to Kaggle it seems that it's a pretty collaborative place, so sorting the notebooks by the highest voted in a somewhat-aged competition will get you up to speed on how people are approaching the competition.</p>",
      "rawMarkdown": "In addition to Alexey's answer, it would also help to look at the top notebooks for the competition -- typically they forego downloading the dataset and use it straight from Kaggle (or use their own augmented/preprocessed dataset based on the competition set). The most useful one I've found to be is [Triple Stratified KFold with TFRecords](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords), but that may just be me. It is TPU-accelerated notebook with the competition data loaded in from a preprocessed user dataset. See @cdeotte's [Triple Stratified Leak-Free KFold CV](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) for more info on the data preprocessing. When viewing the notebooks, you can fork the notebook and edit it as your own. I've found this really useful in getting started.\n\nThough I'm new to Kaggle it seems that it's a pretty collaborative place, so sorting the notebooks by the highest voted in a somewhat-aged competition will get you up to speed on how people are approaching the competition.",
      "votes": null
    },
    {
      "id": "952639",
      "postDate": "07/31/2020 05:26:09",
      "content": "<p>Sure there is a way. You can download any link from this page <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. Each link contains everything you need to train, predict test, and submit to Kaggle. For example you can download all images resized to 128x128 which is only 240MB and still get a great CV LB score! Or download datasets of sizes 1, 2, or 5 GB and get a score in the top 50!</p>",
      "rawMarkdown": "Sure there is a way. You can download any link from this page [here][1]. Each link contains everything you need to train, predict test, and submit to Kaggle. For example you can download all images resized to 128x128 which is only 240MB and still get a great CV LB score! Or download datasets of sizes 1, 2, or 5 GB and get a score in the top 50!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
      "votes": null
    },
    {
      "id": "954681",
      "postDate": "08/02/2020 02:01:37",
      "content": "<p>Okay. Thanks a lot. I ll try that. </p>",
      "rawMarkdown": "Okay. Thanks a lot. I ll try that.",
      "votes": null
    },
    {
      "id": "954685",
      "postDate": "08/02/2020 02:02:15",
      "content": "<p>Alright. Thank you very much.</p>",
      "rawMarkdown": "Alright. Thank you very much.",
      "votes": null
    },
    {
      "id": "954686",
      "postDate": "08/02/2020 02:02:41",
      "content": "<p>Ah!!. Yes, I am new as well. This helps. Thanks</p>",
      "rawMarkdown": "Ah!!. Yes, I am new as well. This helps. Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 952432,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "07/30/2020 22:59:32",
      "content": "<p>There is no need to download the entire data set. If you know Google Cloud Storage bucket addresses for tfrecord files you can be transferring  data to your TPU/GPU directly from inside of your Colab notebook. Take a look at my public notebook illustrating the details of this process (BTW, the notebook, with some modifications, can be run on Colab): <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">EfficientNet BN+Tabular Features TF CV5 512x512</a>.</p>\n\n<p>And here are the discussion topic and another public notebook helping you to determine the most recent GCS addresses:</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165739\">An easy and convenient way to look up GCS bucket addresses!</a>\n<a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">SIIM Show GCS bucket addresses</a></p>\n\n<p>Hope it helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 954685,
          "author_name": "vishvavipinshah",
          "author_url": "",
          "post_date": "08/02/2020 02:02:15",
          "content": "<p>Alright. Thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 952446,
      "author_name": "jtan2231",
      "author_url": "",
      "post_date": "07/30/2020 23:32:15",
      "content": "<p>In addition to Alexey's answer, it would also help to look at the top notebooks for the competition -- typically they forego downloading the dataset and use it straight from Kaggle (or use their own augmented/preprocessed dataset based on the competition set). The most useful one I've found to be is <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">Triple Stratified KFold with TFRecords</a>, but that may just be me. It is TPU-accelerated notebook with the competition data loaded in from a preprocessed user dataset. See <a href=\"/cdeotte\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">Triple Stratified Leak-Free KFold CV</a> for more info on the data preprocessing. When viewing the notebooks, you can fork the notebook and edit it as your own. I've found this really useful in getting started.</p>\n\n<p>Though I'm new to Kaggle it seems that it's a pretty collaborative place, so sorting the notebooks by the highest voted in a somewhat-aged competition will get you up to speed on how people are approaching the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 954686,
          "author_name": "vishvavipinshah",
          "author_url": "",
          "post_date": "08/02/2020 02:02:41",
          "content": "<p>Ah!!. Yes, I am new as well. This helps. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 952639,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/31/2020 05:26:09",
      "content": "<p>Sure there is a way. You can download any link from this page <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. Each link contains everything you need to train, predict test, and submit to Kaggle. For example you can download all images resized to 128x128 which is only 240MB and still get a great CV LB score! Or download datasets of sizes 1, 2, or 5 GB and get a score in the top 50!</p>",
      "votes": null,
      "replies": [
        {
          "id": 954681,
          "author_name": "vishvavipinshah",
          "author_url": "",
          "post_date": "08/02/2020 02:01:37",
          "content": "<p>Okay. Thanks a lot. I ll try that. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "952330": "I am using local computer and google colab for my code. I used the Kaggle API but it throws error as 429-Too many requests.",
    "952432": "There is no need to download the entire data set. If you know Google Cloud Storage bucket addresses for tfrecord files you can be transferring  data to your TPU/GPU directly from inside of your Colab notebook. Take a look at my public notebook illustrating the details of this process (BTW, the notebook, with some modifications, can be run on Colab): [EfficientNet BN+Tabular Features TF CV5 512x512](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512).\n\nAnd here are the discussion topic and another public notebook helping you to determine the most recent GCS addresses:\n\n[An easy and convenient way to look up GCS bucket addresses!](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165739)\n[SIIM Show GCS bucket addresses](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses)\n\nHope it helps!",
    "952446": "In addition to Alexey's answer, it would also help to look at the top notebooks for the competition -- typically they forego downloading the dataset and use it straight from Kaggle (or use their own augmented/preprocessed dataset based on the competition set). The most useful one I've found to be is [Triple Stratified KFold with TFRecords](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords), but that may just be me. It is TPU-accelerated notebook with the competition data loaded in from a preprocessed user dataset. See @cdeotte's [Triple Stratified Leak-Free KFold CV](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) for more info on the data preprocessing. When viewing the notebooks, you can fork the notebook and edit it as your own. I've found this really useful in getting started.\n\nThough I'm new to Kaggle it seems that it's a pretty collaborative place, so sorting the notebooks by the highest voted in a somewhat-aged competition will get you up to speed on how people are approaching the competition.",
    "952639": "Sure there is a way. You can download any link from this page [here][1]. Each link contains everything you need to train, predict test, and submit to Kaggle. For example you can download all images resized to 128x128 which is only 240MB and still get a great CV LB score! Or download datasets of sizes 1, 2, or 5 GB and get a score in the top 50!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
    "954681": "Okay. Thanks a lot. I ll try that.",
    "954685": "Alright. Thank you very much.",
    "954686": "Ah!!. Yes, I am new as well. This helps. Thanks"
  },
  "source": "meta"
}