{
  "id": 186653,
  "title": "Google Colab",
  "url": "/competitions/landmark-recognition-2020/discussion/186653",
  "author_name": "",
  "post_date": "2020-09-25T11:55:34.575114500Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was wondering if anyone was able to perform this competition on Colab, considering:</p>\n<p>1) 429 - Too Many Requests received after running <br>\n<code>! kaggle competitions download -c landmark-recognition-2020</code><br>\n2) 70GB drive storage space per instance</p>",
  "messages": [
    {
      "id": "1026541",
      "postDate": "09/25/2020 11:55:34",
      "content": "<p>I was wondering if anyone was able to perform this competition on Colab, considering:</p>\n<p>1) 429 - Too Many Requests received after running <br>\n<code>! kaggle competitions download -c landmark-recognition-2020</code><br>\n2) 70GB drive storage space per instance</p>",
      "rawMarkdown": "I was wondering if anyone was able to perform this competition on Colab, considering:\n\n1) 429 - Too Many Requests received after running \n`! kaggle competitions download -c landmark-recognition-2020`\n2) 70GB drive storage space per instance",
      "votes": null
    },
    {
      "id": "1026677",
      "postDate": "09/25/2020 13:39:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shymammoth\" target=\"_blank\">@shymammoth</a> </p>\n<ol>\n<li><p>Try for error on request<br>\n!pip uninstall -y kaggle<br>\n!pip install --upgrade pip<br>\n!pip install kaggle==1.5.6</p></li>\n<li><p>Storage is definitely an issue because the competition data is somewhere around 100GB.<br>\nInstead of downloading the data use GCS path and access data files. The kernel shows an example of training with GCS path:<br>\n<a href=\"https://www.kaggle.com/ragnar123/efficientnetb3-data-pipeline-and-model\" target=\"_blank\">https://www.kaggle.com/ragnar123/efficientnetb3-data-pipeline-and-model</a></p></li>\n</ol>\n<p><strong>Hope this helps!</strong></p>",
      "rawMarkdown": "Hi @shymammoth \n1. Try for error on request\n!pip uninstall -y kaggle\n!pip install --upgrade pip\n!pip install kaggle==1.5.6\n\n2. Storage is definitely an issue because the competition data is somewhere around 100GB.\nInstead of downloading the data use GCS path and access data files. The kernel shows an example of training with GCS path:\nhttps://www.kaggle.com/ragnar123/efficientnetb3-data-pipeline-and-model\n\n**Hope this helps!**",
      "votes": null
    },
    {
      "id": "1027165",
      "postDate": "09/25/2020 22:49:17",
      "content": "<p>Yeah, I'm doing this competition exclusively on Colab Pro's TPU, and use GCS to store data. You can use <a href=\"https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting\" target=\"_blank\">this dirty JS trick</a> to prevent the Colab session from disconnecting.</p>\n<p>Also, <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">Keetar</a>, the guy in top 3 of this competition, is also using only Colab Pro + GCS (at least <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037\" target=\"_blank\">he said so in the 1st place writeup of the Retrieval Track of this competition</a>).</p>\n<p>Unfortunately, since there's only 4 days to go, you'll certainly won't be able to train a proper model for this competition. It take DAYS to train models for such a large dataset.</p>",
      "rawMarkdown": "Yeah, I'm doing this competition exclusively on Colab Pro's TPU, and use GCS to store data. You can use [this dirty JS trick](https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting) to prevent the Colab session from disconnecting.\n\nAlso, [Keetar](https://www.kaggle.com/keetar), the guy in top 3 of this competition, is also using only Colab Pro + GCS (at least [he said so in the 1st place writeup of the Retrieval Track of this competition](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037)).\n\nUnfortunately, since there's only 4 days to go, you'll certainly won't be able to train a proper model for this competition. It take DAYS to train models for such a large dataset.",
      "votes": null
    },
    {
      "id": "1027780",
      "postDate": "09/26/2020 11:07:01",
      "content": "<p>Thanks all for your valuable comments. Doing this competition as part as a school assignment so I guess position does not really matter for me. I do understand that GCS buckets are free for a certain threshold of request, but am quite afraid I will receive a bill shock uploading 90GB of data and querying from it.</p>\n<p>Hope I can do something decent within these next few days</p>",
      "rawMarkdown": "Thanks all for your valuable comments. Doing this competition as part as a school assignment so I guess position does not really matter for me. I do understand that GCS buckets are free for a certain threshold of request, but am quite afraid I will receive a bill shock uploading 90GB of data and querying from it.\n\nHope I can do something decent within these next few days",
      "votes": null
    },
    {
      "id": "1027791",
      "postDate": "09/26/2020 11:16:48",
      "content": "<p>Wow, that sounds like a really cool school assignment 😲😲 You can use the datasets that others has converted for this competition:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/180056\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/180056</a><br>\n(this one is even with an example training notebook - very cool)</li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/186487\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/186487</a></li>\n</ul>\n<p>They should be enough to beat the baseline :) </p>",
      "rawMarkdown": "Wow, that sounds like a really cool school assignment 😲😲 You can use the datasets that others has converted for this competition:\n\n- https://www.kaggle.com/c/landmark-recognition-2020/discussion/180056\n  (this one is even with an example training notebook - very cool)\n- https://www.kaggle.com/c/landmark-recognition-2020/discussion/186487\n\nThey should be enough to beat the baseline :)",
      "votes": null
    },
    {
      "id": "1028609",
      "postDate": "09/27/2020 03:17:41",
      "content": "<p>Thanks a lot there! :)</p>",
      "rawMarkdown": "Thanks a lot there! :)",
      "votes": null
    },
    {
      "id": "1029049",
      "postDate": "09/27/2020 12:29:27",
      "content": "<p><a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a> Thanks for sharing my dataset link. 👍</p>",
      "rawMarkdown": "chankhavu Thanks for sharing my dataset link. 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1026677,
      "author_name": "jagadish13",
      "author_url": "",
      "post_date": "09/25/2020 13:39:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shymammoth\" target=\"_blank\">@shymammoth</a> </p>\n<ol>\n<li><p>Try for error on request<br>\n!pip uninstall -y kaggle<br>\n!pip install --upgrade pip<br>\n!pip install kaggle==1.5.6</p></li>\n<li><p>Storage is definitely an issue because the competition data is somewhere around 100GB.<br>\nInstead of downloading the data use GCS path and access data files. The kernel shows an example of training with GCS path:<br>\n<a href=\"https://www.kaggle.com/ragnar123/efficientnetb3-data-pipeline-and-model\" target=\"_blank\">https://www.kaggle.com/ragnar123/efficientnetb3-data-pipeline-and-model</a></p></li>\n</ol>\n<p><strong>Hope this helps!</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1027165,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "09/25/2020 22:49:17",
      "content": "<p>Yeah, I'm doing this competition exclusively on Colab Pro's TPU, and use GCS to store data. You can use <a href=\"https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting\" target=\"_blank\">this dirty JS trick</a> to prevent the Colab session from disconnecting.</p>\n<p>Also, <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">Keetar</a>, the guy in top 3 of this competition, is also using only Colab Pro + GCS (at least <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037\" target=\"_blank\">he said so in the 1st place writeup of the Retrieval Track of this competition</a>).</p>\n<p>Unfortunately, since there's only 4 days to go, you'll certainly won't be able to train a proper model for this competition. It take DAYS to train models for such a large dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1027780,
          "author_name": "shymammoth",
          "author_url": "",
          "post_date": "09/26/2020 11:07:01",
          "content": "<p>Thanks all for your valuable comments. Doing this competition as part as a school assignment so I guess position does not really matter for me. I do understand that GCS buckets are free for a certain threshold of request, but am quite afraid I will receive a bill shock uploading 90GB of data and querying from it.</p>\n<p>Hope I can do something decent within these next few days</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1027791,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/26/2020 11:16:48",
          "content": "<p>Wow, that sounds like a really cool school assignment 😲😲 You can use the datasets that others has converted for this competition:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/180056\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/180056</a><br>\n(this one is even with an example training notebook - very cool)</li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/186487\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/186487</a></li>\n</ul>\n<p>They should be enough to beat the baseline :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1028609,
          "author_name": "shymammoth",
          "author_url": "",
          "post_date": "09/27/2020 03:17:41",
          "content": "<p>Thanks a lot there! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1029049,
          "author_name": "jkreddy",
          "author_url": "",
          "post_date": "09/27/2020 12:29:27",
          "content": "<p><a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a> Thanks for sharing my dataset link. 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1026541": "I was wondering if anyone was able to perform this competition on Colab, considering:\n\n1) 429 - Too Many Requests received after running \n`! kaggle competitions download -c landmark-recognition-2020`\n2) 70GB drive storage space per instance",
    "1026677": "Hi @shymammoth \n1. Try for error on request\n!pip uninstall -y kaggle\n!pip install --upgrade pip\n!pip install kaggle==1.5.6\n\n2. Storage is definitely an issue because the competition data is somewhere around 100GB.\nInstead of downloading the data use GCS path and access data files. The kernel shows an example of training with GCS path:\nhttps://www.kaggle.com/ragnar123/efficientnetb3-data-pipeline-and-model\n\n**Hope this helps!**",
    "1027165": "Yeah, I'm doing this competition exclusively on Colab Pro's TPU, and use GCS to store data. You can use [this dirty JS trick](https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting) to prevent the Colab session from disconnecting.\n\nAlso, [Keetar](https://www.kaggle.com/keetar), the guy in top 3 of this competition, is also using only Colab Pro + GCS (at least [he said so in the 1st place writeup of the Retrieval Track of this competition](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037)).\n\nUnfortunately, since there's only 4 days to go, you'll certainly won't be able to train a proper model for this competition. It take DAYS to train models for such a large dataset.",
    "1027780": "Thanks all for your valuable comments. Doing this competition as part as a school assignment so I guess position does not really matter for me. I do understand that GCS buckets are free for a certain threshold of request, but am quite afraid I will receive a bill shock uploading 90GB of data and querying from it.\n\nHope I can do something decent within these next few days",
    "1027791": "Wow, that sounds like a really cool school assignment 😲😲 You can use the datasets that others has converted for this competition:\n\n- https://www.kaggle.com/c/landmark-recognition-2020/discussion/180056\n  (this one is even with an example training notebook - very cool)\n- https://www.kaggle.com/c/landmark-recognition-2020/discussion/186487\n\nThey should be enough to beat the baseline :)",
    "1028609": "Thanks a lot there! :)",
    "1029049": "chankhavu Thanks for sharing my dataset link. 👍"
  },
  "source": "meta"
}