{
  "id": 172566,
  "title": "Make This Competition More Accessible",
  "url": "/competitions/landmark-retrieval-2020/discussion/172566",
  "author_name": "",
  "post_date": "2020-08-05T15:02:49.663154800Z",
  "votes": 32,
  "comment_count": 6,
  "views": 0,
  "content": "<p>So this competition is very brute force heavy. What I mean is that you need a lot of compute power and you need some way to store the competition data. </p>\n\n<p>However, from my experience in this competition so far, there are two things that the competition managers can do to make this more accessible. </p>\n\n<ol>\n<li>Create a TFRecord dataset from the competition data </li>\n<li>Provide a basic model (not .271 lb) that you can actually train in a kaggle notebook </li>\n</ol>\n\n<p>The problem is that new contestants have to jump through a huge number of hoops just to get started. For me I had do the following:</p>\n\n<ol>\n<li>Convert all the competition data to TFRecord format </li>\n<li>Host that on my own GCS bucket </li>\n<li>Hack apart the baseline submission model to understand how it works </li>\n<li>Train for minimum 12 hours to get a model at .20LB </li>\n</ol>\n\n<p>If we had a kernel that clearly describes how to train a basic model, then it would make the competition much more accessible. </p>\n\n<p>Also, if anyone knows how to upload a kaggle dataset from CGS, let me know and I'll upload my GCS bucket dataset. Its all the competition images stratified by class into 10 folds at 256x256. I can then make a simple TPU model and explain how to train it and get .10LB. </p>",
  "messages": [
    {
      "id": "959411",
      "postDate": "08/05/2020 15:02:49",
      "content": "<p>So this competition is very brute force heavy. What I mean is that you need a lot of compute power and you need some way to store the competition data. </p>\n\n<p>However, from my experience in this competition so far, there are two things that the competition managers can do to make this more accessible. </p>\n\n<ol>\n<li>Create a TFRecord dataset from the competition data </li>\n<li>Provide a basic model (not .271 lb) that you can actually train in a kaggle notebook </li>\n</ol>\n\n<p>The problem is that new contestants have to jump through a huge number of hoops just to get started. For me I had do the following:</p>\n\n<ol>\n<li>Convert all the competition data to TFRecord format </li>\n<li>Host that on my own GCS bucket </li>\n<li>Hack apart the baseline submission model to understand how it works </li>\n<li>Train for minimum 12 hours to get a model at .20LB </li>\n</ol>\n\n<p>If we had a kernel that clearly describes how to train a basic model, then it would make the competition much more accessible. </p>\n\n<p>Also, if anyone knows how to upload a kaggle dataset from CGS, let me know and I'll upload my GCS bucket dataset. Its all the competition images stratified by class into 10 folds at 256x256. I can then make a simple TPU model and explain how to train it and get .10LB. </p>",
      "rawMarkdown": "So this competition is very brute force heavy. What I mean is that you need a lot of compute power and you need some way to store the competition data. \n\nHowever, from my experience in this competition so far, there are two things that the competition managers can do to make this more accessible. \n\n1. Create a TFRecord dataset from the competition data \n2. Provide a basic model (not .271 lb) that you can actually train in a kaggle notebook \n\nThe problem is that new contestants have to jump through a huge number of hoops just to get started. For me I had do the following:\n\n1. Convert all the competition data to TFRecord format \n2. Host that on my own GCS bucket \n3. Hack apart the baseline submission model to understand how it works \n4. Train for minimum 12 hours to get a model at .20LB \n\nIf we had a kernel that clearly describes how to train a basic model, then it would make the competition much more accessible. \n\nAlso, if anyone knows how to upload a kaggle dataset from CGS, let me know and I'll upload my GCS bucket dataset. Its all the competition images stratified by class into 10 folds at 256x256. I can then make a simple TPU model and explain how to train it and get .10LB.",
      "votes": null
    },
    {
      "id": "960126",
      "postDate": "08/06/2020 06:53:49",
      "content": "<p>Fully agree with you, I assume that 12 hrs is for a model that uses TPUs, right? </p>\n\n<p>For an eligible submission we are also supposed to submit a max 9hr runtime GPU kernel without internet connection. I wonder if that is realistic at all? Are top submissions on the leaderboard fulfilling that criteria?</p>\n\n<p>One good thing about all of this is that I learned a lot about how to setup google cloud storage and service accounts and what not... This competition is going to reward the folks with the strongest ops background :D. </p>",
      "rawMarkdown": "Fully agree with you, I assume that 12 hrs is for a model that uses TPUs, right? \n\nFor an eligible submission we are also supposed to submit a max 9hr runtime GPU kernel without internet connection. I wonder if that is realistic at all? Are top submissions on the leaderboard fulfilling that criteria?\n\nOne good thing about all of this is that I learned a lot about how to setup google cloud storage and service accounts and what not... This competition is going to reward the folks with the strongest ops background :D.",
      "votes": null
    },
    {
      "id": "960437",
      "postDate": "08/06/2020 11:59:11",
      "content": "<p>Perhaps you can download the data to colab vm (about 100G storage), and then upload it to the kaggle dataset through kaggle cli api.</p>",
      "rawMarkdown": "Perhaps you can download the data to colab vm (about 100G storage), and then upload it to the kaggle dataset through kaggle cli api.",
      "votes": null
    },
    {
      "id": "960575",
      "postDate": "08/06/2020 14:11:54",
      "content": "<p>you need 101 GB memory but Colab has 95 GB memory inside each session</p>",
      "rawMarkdown": "you need 101 GB memory but Colab has 95 GB memory inside each session",
      "votes": null
    },
    {
      "id": "961859",
      "postDate": "08/07/2020 15:06:04",
      "content": "<p>I truly agree with you. </p>\n<p>I think it's really important to make competitions accessible and the less brute force heavy possible. It favors elegant solutions and it helps new comers to engage more easily in competitions. </p>",
      "rawMarkdown": "I truly agree with you. \n\nI think it's really important to make competitions accessible and the less brute force heavy possible. It favors elegant solutions and it helps new comers to engage more easily in competitions.",
      "votes": null
    },
    {
      "id": "964921",
      "postDate": "08/10/2020 09:11:08",
      "content": "<p><a href=\"https://www.kaggle.com/hooong\" target=\"_blank\">@hooong</a> is it possible for you to upload the said dataset?</p>",
      "rawMarkdown": "hooong is it possible for you to upload the said dataset?",
      "votes": null
    },
    {
      "id": "966720",
      "postDate": "08/11/2020 16:17:14",
      "content": "<p>Hope <a href=\"https://www.kaggle.com/philculliton/landmark-retrieval-2020-tfrecords\" target=\"_blank\">this</a>  dataset from kaggle team  is useful.<br>\n<a href=\"https://www.kaggle.com/hooong\" target=\"_blank\">@hooong</a> </p>",
      "rawMarkdown": "Hope [this](https://www.kaggle.com/philculliton/landmark-retrieval-2020-tfrecords)  dataset from kaggle team  is useful.\n@hooong",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961859,
      "author_name": "hugovergnesx",
      "author_url": "",
      "post_date": "08/07/2020 15:06:04",
      "content": "<p>I truly agree with you. </p>\n<p>I think it's really important to make competitions accessible and the less brute force heavy possible. It favors elegant solutions and it helps new comers to engage more easily in competitions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964921,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "08/10/2020 09:11:08",
      "content": "<p><a href=\"https://www.kaggle.com/hooong\" target=\"_blank\">@hooong</a> is it possible for you to upload the said dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 966720,
          "author_name": "feiwofeifeixiaowo",
          "author_url": "",
          "post_date": "08/11/2020 16:17:14",
          "content": "<p>Hope <a href=\"https://www.kaggle.com/philculliton/landmark-retrieval-2020-tfrecords\" target=\"_blank\">this</a>  dataset from kaggle team  is useful.<br>\n<a href=\"https://www.kaggle.com/hooong\" target=\"_blank\">@hooong</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 960126,
      "author_name": "nawidsayed",
      "author_url": "",
      "post_date": "08/06/2020 06:53:49",
      "content": "<p>Fully agree with you, I assume that 12 hrs is for a model that uses TPUs, right? </p>\n\n<p>For an eligible submission we are also supposed to submit a max 9hr runtime GPU kernel without internet connection. I wonder if that is realistic at all? Are top submissions on the leaderboard fulfilling that criteria?</p>\n\n<p>One good thing about all of this is that I learned a lot about how to setup google cloud storage and service accounts and what not... This competition is going to reward the folks with the strongest ops background :D. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 960437,
      "author_name": "feiwofeifeixiaowo",
      "author_url": "",
      "post_date": "08/06/2020 11:59:11",
      "content": "<p>Perhaps you can download the data to colab vm (about 100G storage), and then upload it to the kaggle dataset through kaggle cli api.</p>",
      "votes": null,
      "replies": [
        {
          "id": 960575,
          "author_name": "stenford23",
          "author_url": "",
          "post_date": "08/06/2020 14:11:54",
          "content": "<p>you need 101 GB memory but Colab has 95 GB memory inside each session</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "959411": "So this competition is very brute force heavy. What I mean is that you need a lot of compute power and you need some way to store the competition data. \n\nHowever, from my experience in this competition so far, there are two things that the competition managers can do to make this more accessible. \n\n1. Create a TFRecord dataset from the competition data \n2. Provide a basic model (not .271 lb) that you can actually train in a kaggle notebook \n\nThe problem is that new contestants have to jump through a huge number of hoops just to get started. For me I had do the following:\n\n1. Convert all the competition data to TFRecord format \n2. Host that on my own GCS bucket \n3. Hack apart the baseline submission model to understand how it works \n4. Train for minimum 12 hours to get a model at .20LB \n\nIf we had a kernel that clearly describes how to train a basic model, then it would make the competition much more accessible. \n\nAlso, if anyone knows how to upload a kaggle dataset from CGS, let me know and I'll upload my GCS bucket dataset. Its all the competition images stratified by class into 10 folds at 256x256. I can then make a simple TPU model and explain how to train it and get .10LB.",
    "960126": "Fully agree with you, I assume that 12 hrs is for a model that uses TPUs, right? \n\nFor an eligible submission we are also supposed to submit a max 9hr runtime GPU kernel without internet connection. I wonder if that is realistic at all? Are top submissions on the leaderboard fulfilling that criteria?\n\nOne good thing about all of this is that I learned a lot about how to setup google cloud storage and service accounts and what not... This competition is going to reward the folks with the strongest ops background :D.",
    "960437": "Perhaps you can download the data to colab vm (about 100G storage), and then upload it to the kaggle dataset through kaggle cli api.",
    "960575": "you need 101 GB memory but Colab has 95 GB memory inside each session",
    "961859": "I truly agree with you. \n\nI think it's really important to make competitions accessible and the less brute force heavy possible. It favors elegant solutions and it helps new comers to engage more easily in competitions.",
    "964921": "hooong is it possible for you to upload the said dataset?",
    "966720": "Hope [this](https://www.kaggle.com/philculliton/landmark-retrieval-2020-tfrecords)  dataset from kaggle team  is useful.\n@hooong"
  },
  "source": "meta"
}