{
  "id": 176487,
  "title": "Reflecting on last year edition",
  "url": "/competitions/landmark-recognition-2020/discussion/176487",
  "author_name": "",
  "post_date": "2020-08-22T00:00:15.467830200Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>People might be interested in using previous years successfull methods. Let's look at previous year's settings.</p>\n<table>\n<thead>\n<tr>\n<th>Landmark Recognition 2019</th>\n<th>Landmark Recognition 2020</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train: all 4M+ images</td>\n<td>Train: cleaned 1.5M images*</td>\n</tr>\n<tr>\n<td>Test: all 117k images</td>\n<td>Test: selected 10k images</td>\n</tr>\n<tr>\n<td>Inference: Upload csv with no restriction</td>\n<td>Inference: Under 12hrs as notebook</td>\n</tr>\n</tbody>\n</table>\n<p>* Of course you are free to use external data (eg: GLDv2 with 5M+ images)</p>\n<p>The baseline shared by the host which does recognition by retrieval on only 100k training images, which already takes 9-10hrs for inference on the test set. Might be useful to keep these limitations in mind before deciding methods to try out. (I will update this post in case something else pops in my head later)</p>",
  "messages": [
    {
      "id": "980862",
      "postDate": "08/22/2020 00:00:15",
      "content": "<p>People might be interested in using previous years successfull methods. Let's look at previous year's settings.</p>\n<table>\n<thead>\n<tr>\n<th>Landmark Recognition 2019</th>\n<th>Landmark Recognition 2020</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train: all 4M+ images</td>\n<td>Train: cleaned 1.5M images*</td>\n</tr>\n<tr>\n<td>Test: all 117k images</td>\n<td>Test: selected 10k images</td>\n</tr>\n<tr>\n<td>Inference: Upload csv with no restriction</td>\n<td>Inference: Under 12hrs as notebook</td>\n</tr>\n</tbody>\n</table>\n<p>* Of course you are free to use external data (eg: GLDv2 with 5M+ images)</p>\n<p>The baseline shared by the host which does recognition by retrieval on only 100k training images, which already takes 9-10hrs for inference on the test set. Might be useful to keep these limitations in mind before deciding methods to try out. (I will update this post in case something else pops in my head later)</p>",
      "rawMarkdown": "People might be interested in using previous years successfull methods. Let's look at previous year's settings.\n\n\n| Landmark Recognition 2019  | Landmark Recognition 2020 |\n| --- | --- |\n| Train: all 4M+ images  | Train: cleaned 1.5M images* |\n| Test: all 117k images  | Test: selected 10k images |\n| Inference: Upload csv with no restriction | Inference: Under 12hrs as notebook |\n\n\\* Of course you are free to use external data (eg: GLDv2 with 5M+ images)\n\nThe baseline shared by the host which does recognition by retrieval on only 100k training images, which already takes 9-10hrs for inference on the test set. Might be useful to keep these limitations in mind before deciding methods to try out. (I will update this post in case something else pops in my head later)",
      "votes": null
    },
    {
      "id": "981804",
      "postDate": "08/22/2020 17:43:59",
      "content": "<p>The main problem of training model for this competition is computational limits after all (especially if you don't have access to external GPU cluster).</p>",
      "rawMarkdown": "The main problem of training model for this competition is computational limits after all (especially if you don't have access to external GPU cluster).",
      "votes": null
    },
    {
      "id": "981809",
      "postDate": "08/22/2020 17:49:38",
      "content": "<p><a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037#979382\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037#979382</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037#979382",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 981804,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "08/22/2020 17:43:59",
      "content": "<p>The main problem of training model for this competition is computational limits after all (especially if you don't have access to external GPU cluster).</p>",
      "votes": null,
      "replies": [
        {
          "id": 981809,
          "author_name": "skrish13",
          "author_url": "",
          "post_date": "08/22/2020 17:49:38",
          "content": "<p><a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037#979382\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037#979382</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "980862": "People might be interested in using previous years successfull methods. Let's look at previous year's settings.\n\n\n| Landmark Recognition 2019  | Landmark Recognition 2020 |\n| --- | --- |\n| Train: all 4M+ images  | Train: cleaned 1.5M images* |\n| Test: all 117k images  | Test: selected 10k images |\n| Inference: Upload csv with no restriction | Inference: Under 12hrs as notebook |\n\n\\* Of course you are free to use external data (eg: GLDv2 with 5M+ images)\n\nThe baseline shared by the host which does recognition by retrieval on only 100k training images, which already takes 9-10hrs for inference on the test set. Might be useful to keep these limitations in mind before deciding methods to try out. (I will update this post in case something else pops in my head later)",
    "981804": "The main problem of training model for this competition is computational limits after all (especially if you don't have access to external GPU cluster).",
    "981809": "https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037#979382"
  },
  "source": "meta"
}