{
  "id": 275870,
  "title": "3rd place solution",
  "url": "/competitions/landmark-recognition-2021/writeups/bytedance-vision-tech-3rd-place-solution",
  "author_name": "",
  "post_date": "2021-10-02T12:30:16.927Z",
  "votes": 42,
  "comment_count": 7,
  "views": 0,
  "content": "<h2>3rd  place solution</h2>\n<p>Congrats to the winners, and thanks a lot to Kaggle &amp; Google for organising this wonderful Landmark Recognition competition yearly. This is great a teamwork for us, and we will briefly discuss our solution here. </p>\n<p>We are probably the most submitted team (200+ submissions) on LB, reason being we were exploring on different approaches quite independently in the beginning, where we only merged in the final week right before the deadline. We have explored quite many different ideas from the start, but our final ensemble pipeline differed a lot from what tried in the beginning. </p>\n<p>Probably like most others, we wasted quite a few submissions before realising that 1) there are at least 20K private test images to be provided during inference, instead of 10K as in test folder 2) there are landmark ids other than the cleaned 81313 version existing in private train.  </p>\n<h3>Solution Overview</h3>\n<p>Our final models include two sets - <strong>retrieval models</strong> and <strong>classification models</strong>. And our solution can be summarised the following: <br>\n<code>Concat &amp; Retrieval -&gt; \nclassification logits adjustment -&gt; \ndistractor score adjustment -&gt; \ntop 1 classification score</code></p>\n<h3>Dive into details for each part</h3>\n<h4>1) Concat &amp; Retrieval</h4>\n<p>We have 7 retrieval models in our final pipeline: </p>\n<p>OrangeX: Swin-l (384), Swin-l (384 , different training), EfficientnetV2 (800), EfficientB7 (800), CvT (384)<br>\nWeimin: EfficientB5 768</p>\n<p>We found that training models only on 3M (81313 landmarks) dataset gives better or at least equal retrieval performance compared to finetuning on full 4M datasets. And our single best retrieval model - Swin-l - gives 0.4 for its pure retrieval performance on Public LB. All others like B7, V2, B5 are a lot weaker in terms of individual retrieval performance. (Around 0.29 ~ 0.33). </p>\n<p>All our models give 512-d output features, and in our final pipeline we simply concat them to form a final embedding vector (512 x 6) for retrieval of top K matched training images from private train. We used K = 7. </p>\n<p>We found that using dynamic margin (or adaptive margins) actually helps improve training, thanks To Qishen's great solution last year! </p>\n<h4>2) classification logits adjustment</h4>\n<p>We found that it is crucial to use classification logits to support predictions. From the beginning, I managed to split the full 4M dataset into training and validation folds, and when I realised the private train could contain un-cleaned landmark ids, the validation mean_gap can be reliably used as measurement for classification performance, which is also consistent to LB. </p>\n<p>It turned out that this is important. As using a representative validation set, it is relatively easy to finetune models back on full 200k landmark ids training set (4M images). In the end, we found the top 1 pure classification accuracy of my B5 is around 0.39+ on LB, where for swin it is around 0.34.  We decided to go with all 4 EfficientNet that I trained on my side (B5 512 &amp; 768, B6 512 &amp; 768) as our main classification models for stage 2) here as well as 4) later on. </p>\n<p>At this stage 2), we have got top 7 retrieved training images from stage 1). So for each of the 7 images, we look up its classification logits from all 4 models chosen (B5 512&amp;768, B6 512&amp;768), and simply add the averaged logit to its corresponding cosine score as adjustment. </p>\n<h4>3) distractor score adjustment</h4>\n<p>Similarly from what Dieter did last year, we use the 2019 test set's nonlandmark images as index, and for each training image (4M), we found its top 3 matched scores and simply take its average as its distractor score. We generated the mapping between each 4M training image id to its score in a dict, and uploaded to Kaggle to use in submission. </p>\n<p>We then subtract the distractor score from each adjusted cosine score above from stage 2). <br>\nIn short, stage 1) - 3) can be viewed as: </p>\n<pre><code>Cosine score + classification logit - distractor score \n</code></pre>\n<h4>4) top 1 classification score</h4>\n<p>We found that using the top 1 classification logits from our best classification models (i.e. EffNet B5 and B6) can have another boost. </p>\n<p>At stage 3) above, we should have all top 7 indexed images with their score adjusted ready,  so we simply aggregate them to each's corresponding landmark id. We will, however, add another pair of <code>(top 1 classification landmark id, top 1 classification logit)</code> into the aggregation step. The classification logit used is just the raw top 1 logit, after we averaged all classification models' 200k prediction logits. </p>\n<p>We can't penalise the classification pair as it is not from any image like the top 7 matched, therefore it does not have a distractor score. But this turns out not to be a problem, as we found that top 1 classification score can be naturally used as penalty of non-landmark. See the distribution of landmark &amp; non-landmark images' top1 scores below: </p>\n<p><a href=\"https://drive.google.com/file/d/127O__NgWGIW8E73XX2ZqrK4X9m9asZY_/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/127O__NgWGIW8E73XX2ZqrK4X9m9asZY_/view?usp=sharing</a></p>\n<p>Our final selection of landmark id and score for each test image will be just the aggregation result from all 1) - 4) stages. </p>\n<p>Using stage 1) - stage 4) with a single B5 only (not fully trained yet) achieves 0.445/476 on LB. We achieved our final rankings using all models mentioned above. </p>",
  "messages": [
    {
      "id": "1531393",
      "postDate": "10/02/2021 00:10:18",
      "content": "<h2>3rd  place solution</h2>\n<p>Congrats to the winners, and thanks a lot to Kaggle &amp; Google for organising this wonderful Landmark Recognition competition yearly. This is great a teamwork for us, and we will briefly discuss our solution here. </p>\n<p>We are probably the most submitted team (200+ submissions) on LB, reason being we were exploring on different approaches quite independently in the beginning, where we only merged in the final week right before the deadline. We have explored quite many different ideas from the start, but our final ensemble pipeline differed a lot from what tried in the beginning. </p>\n<p>Probably like most others, we wasted quite a few submissions before realising that 1) there are at least 20K private test images to be provided during inference, instead of 10K as in test folder 2) there are landmark ids other than the cleaned 81313 version existing in private train.  </p>\n<h3>Solution Overview</h3>\n<p>Our final models include two sets - <strong>retrieval models</strong> and <strong>classification models</strong>. And our solution can be summarised the following: <br>\n<code>Concat &amp; Retrieval -&gt; \nclassification logits adjustment -&gt; \ndistractor score adjustment -&gt; \ntop 1 classification score</code></p>\n<h3>Dive into details for each part</h3>\n<h4>1) Concat &amp; Retrieval</h4>\n<p>We have 7 retrieval models in our final pipeline: </p>\n<p>OrangeX: Swin-l (384), Swin-l (384 , different training), EfficientnetV2 (800), EfficientB7 (800), CvT (384)<br>\nWeimin: EfficientB5 768</p>\n<p>We found that training models only on 3M (81313 landmarks) dataset gives better or at least equal retrieval performance compared to finetuning on full 4M datasets. And our single best retrieval model - Swin-l - gives 0.4 for its pure retrieval performance on Public LB. All others like B7, V2, B5 are a lot weaker in terms of individual retrieval performance. (Around 0.29 ~ 0.33). </p>\n<p>All our models give 512-d output features, and in our final pipeline we simply concat them to form a final embedding vector (512 x 6) for retrieval of top K matched training images from private train. We used K = 7. </p>\n<p>We found that using dynamic margin (or adaptive margins) actually helps improve training, thanks To Qishen's great solution last year! </p>\n<h4>2) classification logits adjustment</h4>\n<p>We found that it is crucial to use classification logits to support predictions. From the beginning, I managed to split the full 4M dataset into training and validation folds, and when I realised the private train could contain un-cleaned landmark ids, the validation mean_gap can be reliably used as measurement for classification performance, which is also consistent to LB. </p>\n<p>It turned out that this is important. As using a representative validation set, it is relatively easy to finetune models back on full 200k landmark ids training set (4M images). In the end, we found the top 1 pure classification accuracy of my B5 is around 0.39+ on LB, where for swin it is around 0.34.  We decided to go with all 4 EfficientNet that I trained on my side (B5 512 &amp; 768, B6 512 &amp; 768) as our main classification models for stage 2) here as well as 4) later on. </p>\n<p>At this stage 2), we have got top 7 retrieved training images from stage 1). So for each of the 7 images, we look up its classification logits from all 4 models chosen (B5 512&amp;768, B6 512&amp;768), and simply add the averaged logit to its corresponding cosine score as adjustment. </p>\n<h4>3) distractor score adjustment</h4>\n<p>Similarly from what Dieter did last year, we use the 2019 test set's nonlandmark images as index, and for each training image (4M), we found its top 3 matched scores and simply take its average as its distractor score. We generated the mapping between each 4M training image id to its score in a dict, and uploaded to Kaggle to use in submission. </p>\n<p>We then subtract the distractor score from each adjusted cosine score above from stage 2). <br>\nIn short, stage 1) - 3) can be viewed as: </p>\n<pre><code>Cosine score + classification logit - distractor score \n</code></pre>\n<h4>4) top 1 classification score</h4>\n<p>We found that using the top 1 classification logits from our best classification models (i.e. EffNet B5 and B6) can have another boost. </p>\n<p>At stage 3) above, we should have all top 7 indexed images with their score adjusted ready,  so we simply aggregate them to each's corresponding landmark id. We will, however, add another pair of <code>(top 1 classification landmark id, top 1 classification logit)</code> into the aggregation step. The classification logit used is just the raw top 1 logit, after we averaged all classification models' 200k prediction logits. </p>\n<p>We can't penalise the classification pair as it is not from any image like the top 7 matched, therefore it does not have a distractor score. But this turns out not to be a problem, as we found that top 1 classification score can be naturally used as penalty of non-landmark. See the distribution of landmark &amp; non-landmark images' top1 scores below: </p>\n<p><a href=\"https://drive.google.com/file/d/127O__NgWGIW8E73XX2ZqrK4X9m9asZY_/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/127O__NgWGIW8E73XX2ZqrK4X9m9asZY_/view?usp=sharing</a></p>\n<p>Our final selection of landmark id and score for each test image will be just the aggregation result from all 1) - 4) stages. </p>\n<p>Using stage 1) - stage 4) with a single B5 only (not fully trained yet) achieves 0.445/476 on LB. We achieved our final rankings using all models mentioned above. </p>",
      "rawMarkdown": "## 3rd  place solution \nCongrats to the winners, and thanks a lot to Kaggle & Google for organising this wonderful Landmark Recognition competition yearly. This is great a teamwork for us, and we will briefly discuss our solution here. \n\nWe are probably the most submitted team (200+ submissions) on LB, reason being we were exploring on different approaches quite independently in the beginning, where we only merged in the final week right before the deadline. We have explored quite many different ideas from the start, but our final ensemble pipeline differed a lot from what tried in the beginning. \n\nProbably like most others, we wasted quite a few submissions before realising that 1) there are at least 20K private test images to be provided during inference, instead of 10K as in test folder 2) there are landmark ids other than the cleaned 81313 version existing in private train.  \n\n### Solution Overview\n\nOur final models include two sets - **retrieval models** and **classification models**. And our solution can be summarised the following: \n`Concat & Retrieval -> \nclassification logits adjustment -> \ndistractor score adjustment -> \ntop 1 classification score `\n\n### Dive into details for each part\n\n####1) Concat & Retrieval \nWe have 7 retrieval models in our final pipeline: \n\nOrangeX: Swin-l (384), Swin-l (384 , different training), EfficientnetV2 (800), EfficientB7 (800), CvT (384)\nWeimin: EfficientB5 768\n\nWe found that training models only on 3M (81313 landmarks) dataset gives better or at least equal retrieval performance compared to finetuning on full 4M datasets. And our single best retrieval model - Swin-l - gives 0.4 for its pure retrieval performance on Public LB. All others like B7, V2, B5 are a lot weaker in terms of individual retrieval performance. (Around 0.29 ~ 0.33). \n\nAll our models give 512-d output features, and in our final pipeline we simply concat them to form a final embedding vector (512 x 6) for retrieval of top K matched training images from private train. We used K = 7. \n\nWe found that using dynamic margin (or adaptive margins) actually helps improve training, thanks To Qishen's great solution last year! \n\n####2) classification logits adjustment\nWe found that it is crucial to use classification logits to support predictions. From the beginning, I managed to split the full 4M dataset into training and validation folds, and when I realised the private train could contain un-cleaned landmark ids, the validation mean_gap can be reliably used as measurement for classification performance, which is also consistent to LB. \n\nIt turned out that this is important. As using a representative validation set, it is relatively easy to finetune models back on full 200k landmark ids training set (4M images). In the end, we found the top 1 pure classification accuracy of my B5 is around 0.39+ on LB, where for swin it is around 0.34.  We decided to go with all 4 EfficientNet that I trained on my side (B5 512 & 768, B6 512 & 768) as our main classification models for stage 2) here as well as 4) later on. \n\nAt this stage 2), we have got top 7 retrieved training images from stage 1). So for each of the 7 images, we look up its classification logits from all 4 models chosen (B5 512&768, B6 512&768), and simply add the averaged logit to its corresponding cosine score as adjustment. \n\n####3) distractor score adjustment\nSimilarly from what Dieter did last year, we use the 2019 test set's nonlandmark images as index, and for each training image (4M), we found its top 3 matched scores and simply take its average as its distractor score. We generated the mapping between each 4M training image id to its score in a dict, and uploaded to Kaggle to use in submission. \n\nWe then subtract the distractor score from each adjusted cosine score above from stage 2). \nIn short, stage 1) - 3) can be viewed as: \n\n```\nCosine score + classification logit - distractor score \n\n```\n####4) top 1 classification score \nWe found that using the top 1 classification logits from our best classification models (i.e. EffNet B5 and B6) can have another boost. \n\nAt stage 3) above, we should have all top 7 indexed images with their score adjusted ready,  so we simply aggregate them to each's corresponding landmark id. We will, however, add another pair of `(top 1 classification landmark id, top 1 classification logit)` into the aggregation step. The classification logit used is just the raw top 1 logit, after we averaged all classification models' 200k prediction logits. \n\nWe can't penalise the classification pair as it is not from any image like the top 7 matched, therefore it does not have a distractor score. But this turns out not to be a problem, as we found that top 1 classification score can be naturally used as penalty of non-landmark. See the distribution of landmark & non-landmark images' top1 scores below: \n\nhttps://drive.google.com/file/d/127O__NgWGIW8E73XX2ZqrK4X9m9asZY_/view?usp=sharing\n\nOur final selection of landmark id and score for each test image will be just the aggregation result from all 1) - 4) stages. \n\nUsing stage 1) - stage 4) with a single B5 only (not fully trained yet) achieves 0.445/476 on LB. We achieved our final rankings using all models mentioned above.",
      "votes": null
    },
    {
      "id": "1531401",
      "postDate": "10/02/2021 00:26:06",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a> what kind of hardware do you use for this competition? </p>",
      "rawMarkdown": "Hey @weimin what kind of hardware do you use for this competition?",
      "votes": null
    },
    {
      "id": "1531402",
      "postDate": "10/02/2021 00:26:49",
      "content": "<p>Congrats and thanks for the write-up! Great idea with the logits adjustment</p>",
      "rawMarkdown": "Congrats and thanks for the write-up! Great idea with the logits adjustment",
      "votes": null
    },
    {
      "id": "1531404",
      "postDate": "10/02/2021 00:31:28",
      "content": "<p>Thanks! It is similar to what top 3 used last year, but instead of multiplying we used addition. And our top 1 classification adjustment is new. </p>",
      "rawMarkdown": "Thanks! It is similar to what top 3 used last year, but instead of multiplying we used addition. And our top 1 classification adjustment is new.",
      "votes": null
    },
    {
      "id": "1531405",
      "postDate": "10/02/2021 00:32:50",
      "content": "<p>I only used 2 8-V100 machine to train my models - B5 and B6. </p>",
      "rawMarkdown": "I only used 2 8-V100 machine to train my models - B5 and B6.",
      "votes": null
    },
    {
      "id": "1531706",
      "postDate": "10/02/2021 09:37:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a>, congrats on the win and thanks for sharing your solution here!!<br>\nI was trying to compute something similar to your 'distractor score'  by finding the cosine scores between private train images and the 2019 test set's nonlandmark images during submission, but unfortunately ran into notebook memory errors.</p>\n<p>May I ask how did you decide on which of the 4M distractor scores to use on each of the adjusted cosine score? Is it simply by matching the private train image id with that of the image id of the 4M train images in your scoring dict? </p>\n<p>Also, can I assume that you computed a matching score between each of the 4M train images and the 100K 2019 test set's nonlandmark images? Meaning to say that you have (4M x 100K) distractor scores at the end? Would be great to learn from you if you could share the methods you used to compute these scores efficiently without any memory errors.:)</p>\n<p>Thank you!:)</p>",
      "rawMarkdown": "Hi @weimin, congrats on the win and thanks for sharing your solution here!!\nI was trying to compute something similar to your 'distractor score'  by finding the cosine scores between private train images and the 2019 test set's nonlandmark images during submission, but unfortunately ran into notebook memory errors.\n\nMay I ask how did you decide on which of the 4M distractor scores to use on each of the adjusted cosine score? Is it simply by matching the private train image id with that of the image id of the 4M train images in your scoring dict? \n\nAlso, can I assume that you computed a matching score between each of the 4M train images and the 100K 2019 test set's nonlandmark images? Meaning to say that you have (4M x 100K) distractor scores at the end? Would be great to learn from you if you could share the methods you used to compute these scores efficiently without any memory errors.:)\n\nThank you!:)",
      "votes": null
    },
    {
      "id": "1531825",
      "postDate": "10/02/2021 12:17:57",
      "content": "<p>To answer your two questions: </p>\n<p>1) Yes we uploaded a dict mapping for all 4M full train image id -&gt; non-landmark score. During inference, we retrieved top K private train images first, and for each image, we found its score from the mapping, and subtract that from its cosine</p>\n<p>2) Yes, all 4M from against 115K 2019 test nonlandmark images. You can use some modern GPU library such as faiss to compute the distance, and return only the top 3 best match, instead of full 100K. Then take the average of the top 3 cosines. Should take around 3 mins on GPU without memory issue I guess. </p>",
      "rawMarkdown": "To answer your two questions: \n\n1) Yes we uploaded a dict mapping for all 4M full train image id -> non-landmark score. During inference, we retrieved top K private train images first, and for each image, we found its score from the mapping, and subtract that from its cosine\n\n2) Yes, all 4M from against 115K 2019 test nonlandmark images. You can use some modern GPU library such as faiss to compute the distance, and return only the top 3 best match, instead of full 100K. Then take the average of the top 3 cosines. Should take around 3 mins on GPU without memory issue I guess.",
      "votes": null
    },
    {
      "id": "1532272",
      "postDate": "10/02/2021 19:17:10",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a> and your team! Thank you for that detailed write-up. Learned a lot!</p>",
      "rawMarkdown": "Congrats @weimin and your team! Thank you for that detailed write-up. Learned a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1531401,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "10/02/2021 00:26:06",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a> what kind of hardware do you use for this competition? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1531405,
          "author_name": "weimin",
          "author_url": "",
          "post_date": "10/02/2021 00:32:50",
          "content": "<p>I only used 2 8-V100 machine to train my models - B5 and B6. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1531402,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "10/02/2021 00:26:49",
      "content": "<p>Congrats and thanks for the write-up! Great idea with the logits adjustment</p>",
      "votes": null,
      "replies": [
        {
          "id": 1531404,
          "author_name": "weimin",
          "author_url": "",
          "post_date": "10/02/2021 00:31:28",
          "content": "<p>Thanks! It is similar to what top 3 used last year, but instead of multiplying we used addition. And our top 1 classification adjustment is new. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1531706,
      "author_name": "tmxxuan",
      "author_url": "",
      "post_date": "10/02/2021 09:37:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a>, congrats on the win and thanks for sharing your solution here!!<br>\nI was trying to compute something similar to your 'distractor score'  by finding the cosine scores between private train images and the 2019 test set's nonlandmark images during submission, but unfortunately ran into notebook memory errors.</p>\n<p>May I ask how did you decide on which of the 4M distractor scores to use on each of the adjusted cosine score? Is it simply by matching the private train image id with that of the image id of the 4M train images in your scoring dict? </p>\n<p>Also, can I assume that you computed a matching score between each of the 4M train images and the 100K 2019 test set's nonlandmark images? Meaning to say that you have (4M x 100K) distractor scores at the end? Would be great to learn from you if you could share the methods you used to compute these scores efficiently without any memory errors.:)</p>\n<p>Thank you!:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1531825,
          "author_name": "weimin",
          "author_url": "",
          "post_date": "10/02/2021 12:17:57",
          "content": "<p>To answer your two questions: </p>\n<p>1) Yes we uploaded a dict mapping for all 4M full train image id -&gt; non-landmark score. During inference, we retrieved top K private train images first, and for each image, we found its score from the mapping, and subtract that from its cosine</p>\n<p>2) Yes, all 4M from against 115K 2019 test nonlandmark images. You can use some modern GPU library such as faiss to compute the distance, and return only the top 3 best match, instead of full 100K. Then take the average of the top 3 cosines. Should take around 3 mins on GPU without memory issue I guess. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1532272,
      "author_name": "maxschfer",
      "author_url": "",
      "post_date": "10/02/2021 19:17:10",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a> and your team! Thank you for that detailed write-up. Learned a lot!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1531393": "## 3rd  place solution \nCongrats to the winners, and thanks a lot to Kaggle & Google for organising this wonderful Landmark Recognition competition yearly. This is great a teamwork for us, and we will briefly discuss our solution here. \n\nWe are probably the most submitted team (200+ submissions) on LB, reason being we were exploring on different approaches quite independently in the beginning, where we only merged in the final week right before the deadline. We have explored quite many different ideas from the start, but our final ensemble pipeline differed a lot from what tried in the beginning. \n\nProbably like most others, we wasted quite a few submissions before realising that 1) there are at least 20K private test images to be provided during inference, instead of 10K as in test folder 2) there are landmark ids other than the cleaned 81313 version existing in private train.  \n\n### Solution Overview\n\nOur final models include two sets - **retrieval models** and **classification models**. And our solution can be summarised the following: \n`Concat & Retrieval -> \nclassification logits adjustment -> \ndistractor score adjustment -> \ntop 1 classification score `\n\n### Dive into details for each part\n\n####1) Concat & Retrieval \nWe have 7 retrieval models in our final pipeline: \n\nOrangeX: Swin-l (384), Swin-l (384 , different training), EfficientnetV2 (800), EfficientB7 (800), CvT (384)\nWeimin: EfficientB5 768\n\nWe found that training models only on 3M (81313 landmarks) dataset gives better or at least equal retrieval performance compared to finetuning on full 4M datasets. And our single best retrieval model - Swin-l - gives 0.4 for its pure retrieval performance on Public LB. All others like B7, V2, B5 are a lot weaker in terms of individual retrieval performance. (Around 0.29 ~ 0.33). \n\nAll our models give 512-d output features, and in our final pipeline we simply concat them to form a final embedding vector (512 x 6) for retrieval of top K matched training images from private train. We used K = 7. \n\nWe found that using dynamic margin (or adaptive margins) actually helps improve training, thanks To Qishen's great solution last year! \n\n####2) classification logits adjustment\nWe found that it is crucial to use classification logits to support predictions. From the beginning, I managed to split the full 4M dataset into training and validation folds, and when I realised the private train could contain un-cleaned landmark ids, the validation mean_gap can be reliably used as measurement for classification performance, which is also consistent to LB. \n\nIt turned out that this is important. As using a representative validation set, it is relatively easy to finetune models back on full 200k landmark ids training set (4M images). In the end, we found the top 1 pure classification accuracy of my B5 is around 0.39+ on LB, where for swin it is around 0.34.  We decided to go with all 4 EfficientNet that I trained on my side (B5 512 & 768, B6 512 & 768) as our main classification models for stage 2) here as well as 4) later on. \n\nAt this stage 2), we have got top 7 retrieved training images from stage 1). So for each of the 7 images, we look up its classification logits from all 4 models chosen (B5 512&768, B6 512&768), and simply add the averaged logit to its corresponding cosine score as adjustment. \n\n####3) distractor score adjustment\nSimilarly from what Dieter did last year, we use the 2019 test set's nonlandmark images as index, and for each training image (4M), we found its top 3 matched scores and simply take its average as its distractor score. We generated the mapping between each 4M training image id to its score in a dict, and uploaded to Kaggle to use in submission. \n\nWe then subtract the distractor score from each adjusted cosine score above from stage 2). \nIn short, stage 1) - 3) can be viewed as: \n\n```\nCosine score + classification logit - distractor score \n\n```\n####4) top 1 classification score \nWe found that using the top 1 classification logits from our best classification models (i.e. EffNet B5 and B6) can have another boost. \n\nAt stage 3) above, we should have all top 7 indexed images with their score adjusted ready,  so we simply aggregate them to each's corresponding landmark id. We will, however, add another pair of `(top 1 classification landmark id, top 1 classification logit)` into the aggregation step. The classification logit used is just the raw top 1 logit, after we averaged all classification models' 200k prediction logits. \n\nWe can't penalise the classification pair as it is not from any image like the top 7 matched, therefore it does not have a distractor score. But this turns out not to be a problem, as we found that top 1 classification score can be naturally used as penalty of non-landmark. See the distribution of landmark & non-landmark images' top1 scores below: \n\nhttps://drive.google.com/file/d/127O__NgWGIW8E73XX2ZqrK4X9m9asZY_/view?usp=sharing\n\nOur final selection of landmark id and score for each test image will be just the aggregation result from all 1) - 4) stages. \n\nUsing stage 1) - stage 4) with a single B5 only (not fully trained yet) achieves 0.445/476 on LB. We achieved our final rankings using all models mentioned above.",
    "1531401": "Hey @weimin what kind of hardware do you use for this competition?",
    "1531402": "Congrats and thanks for the write-up! Great idea with the logits adjustment",
    "1531404": "Thanks! It is similar to what top 3 used last year, but instead of multiplying we used addition. And our top 1 classification adjustment is new.",
    "1531405": "I only used 2 8-V100 machine to train my models - B5 and B6.",
    "1531706": "Hi @weimin, congrats on the win and thanks for sharing your solution here!!\nI was trying to compute something similar to your 'distractor score'  by finding the cosine scores between private train images and the 2019 test set's nonlandmark images during submission, but unfortunately ran into notebook memory errors.\n\nMay I ask how did you decide on which of the 4M distractor scores to use on each of the adjusted cosine score? Is it simply by matching the private train image id with that of the image id of the 4M train images in your scoring dict? \n\nAlso, can I assume that you computed a matching score between each of the 4M train images and the 100K 2019 test set's nonlandmark images? Meaning to say that you have (4M x 100K) distractor scores at the end? Would be great to learn from you if you could share the methods you used to compute these scores efficiently without any memory errors.:)\n\nThank you!:)",
    "1531825": "To answer your two questions: \n\n1) Yes we uploaded a dict mapping for all 4M full train image id -> non-landmark score. During inference, we retrieved top K private train images first, and for each image, we found its score from the mapping, and subtract that from its cosine\n\n2) Yes, all 4M from against 115K 2019 test nonlandmark images. You can use some modern GPU library such as faiss to compute the distance, and return only the top 3 best match, instead of full 100K. Then take the average of the top 3 cosines. Should take around 3 mins on GPU without memory issue I guess.",
    "1532272": "Congrats @weimin and your team! Thank you for that detailed write-up. Learned a lot!"
  },
  "source": "meta"
}