{
  "id": 277386,
  "title": "4th place solution",
  "url": "/competitions/landmark-recognition-2021/writeups/all-data-are-ext-4th-place-solution",
  "author_name": "",
  "post_date": "2021-10-09T08:28:29.698868100Z",
  "votes": 27,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to my longtime teammate <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and new teammate <a href=\"https://www.kaggle.com/hongweizhang\" target=\"_blank\">@hongweizhang</a> . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.</p>\n<h2>Approach</h2>\n<p>Qishen and I were part of the 3rd place team (together with <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a> ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. </p>\n<p>This year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.</p>\n<h2>Training recipe</h2>\n<p>We mostly follow our last year's solution (see <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">here</a> for details). The recipes are</p>\n<ul>\n<li>sub-center ArcFace with dynamic margins</li>\n<li>progressive training with increasing image sizes</li>\n<li>the indispensable <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library</li>\n<li>cosine learning schedule with Adam/AdamW optimizer</li>\n<li>multi-GPU training with DistributedDataParallel</li>\n</ul>\n<h2>Model choices</h2>\n<p>There are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -&gt; 512 -&gt; 640/768). </p>\n<p>We found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.</p>\n<p>Besides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.</p>\n<h2>The ensemble</h2>\n<p>Our final ensemble consists of 11 models: 3 transformers, 3 new CNNs and 5 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).</p>\n<p>Our best retrieval ensemble has 4 fewer CNNs and only 7 models in total (3 transformers + 4 CNNs), because the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Image size</th>\n<th>Total epochs</th>\n<th>Finetune 2020</th>\n<th>cv GAP (recognition)</th>\n<th>cv mAP@100 (retrieval)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Swin base</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.7049</td>\n<td>0.4442</td>\n</tr>\n<tr>\n<td>Swin large</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.6775</td>\n<td>0.5161</td>\n</tr>\n<tr>\n<td>ViT large</td>\n<td>384</td>\n<td>50</td>\n<td></td>\n<td>0.6589</td>\n<td>0.5633</td>\n</tr>\n<tr>\n<td>ECA NFNet L2</td>\n<td>640</td>\n<td>38</td>\n<td></td>\n<td>0.7053</td>\n<td>0.3600</td>\n</tr>\n<tr>\n<td>EfficientNet v2m</td>\n<td>640</td>\n<td>45</td>\n<td></td>\n<td>0.7136</td>\n<td>0.3811</td>\n</tr>\n<tr>\n<td>EfficientNet v2l</td>\n<td>512</td>\n<td>30</td>\n<td></td>\n<td>0.7146</td>\n<td>0.4117</td>\n</tr>\n<tr>\n<td>EfficientNet B4</td>\n<td>768</td>\n<td>10</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B5</td>\n<td>768</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>512</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B7</td>\n<td>672</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>ReXNet 2.0</td>\n<td>768</td>\n<td>10</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Post-processing</h2>\n<p>Post-processing was shown to be very effective in last year's top solutions. We incorporated ideas of top3 solutions</p>\n<ol>\n<li>Combine cosine similarity and classification probabilities (our <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">3rd place solution</a>)</li>\n<li>Penalize similarity with non-landmark images (<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821\" target=\"_blank\">1st place solution</a>)</li>\n<li>Set similarity=0 if top5 similarities with non-landmarks &gt; 0.5 (<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299\" target=\"_blank\">2nd place solution</a>)</li>\n<li>Penalize CosSim(query, index) for index class’s count in train set (ours, new)</li>\n</ol>",
  "messages": [
    {
      "id": "1539242",
      "postDate": "10/09/2021 08:28:29",
      "content": "<p>Thanks to my longtime teammate <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and new teammate <a href=\"https://www.kaggle.com/hongweizhang\" target=\"_blank\">@hongweizhang</a> . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.</p>\n<h2>Approach</h2>\n<p>Qishen and I were part of the 3rd place team (together with <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a> ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. </p>\n<p>This year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.</p>\n<h2>Training recipe</h2>\n<p>We mostly follow our last year's solution (see <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">here</a> for details). The recipes are</p>\n<ul>\n<li>sub-center ArcFace with dynamic margins</li>\n<li>progressive training with increasing image sizes</li>\n<li>the indispensable <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library</li>\n<li>cosine learning schedule with Adam/AdamW optimizer</li>\n<li>multi-GPU training with DistributedDataParallel</li>\n</ul>\n<h2>Model choices</h2>\n<p>There are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -&gt; 512 -&gt; 640/768). </p>\n<p>We found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.</p>\n<p>Besides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.</p>\n<h2>The ensemble</h2>\n<p>Our final ensemble consists of 11 models: 3 transformers, 3 new CNNs and 5 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).</p>\n<p>Our best retrieval ensemble has 4 fewer CNNs and only 7 models in total (3 transformers + 4 CNNs), because the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Image size</th>\n<th>Total epochs</th>\n<th>Finetune 2020</th>\n<th>cv GAP (recognition)</th>\n<th>cv mAP@100 (retrieval)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Swin base</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.7049</td>\n<td>0.4442</td>\n</tr>\n<tr>\n<td>Swin large</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.6775</td>\n<td>0.5161</td>\n</tr>\n<tr>\n<td>ViT large</td>\n<td>384</td>\n<td>50</td>\n<td></td>\n<td>0.6589</td>\n<td>0.5633</td>\n</tr>\n<tr>\n<td>ECA NFNet L2</td>\n<td>640</td>\n<td>38</td>\n<td></td>\n<td>0.7053</td>\n<td>0.3600</td>\n</tr>\n<tr>\n<td>EfficientNet v2m</td>\n<td>640</td>\n<td>45</td>\n<td></td>\n<td>0.7136</td>\n<td>0.3811</td>\n</tr>\n<tr>\n<td>EfficientNet v2l</td>\n<td>512</td>\n<td>30</td>\n<td></td>\n<td>0.7146</td>\n<td>0.4117</td>\n</tr>\n<tr>\n<td>EfficientNet B4</td>\n<td>768</td>\n<td>10</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B5</td>\n<td>768</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>512</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B7</td>\n<td>672</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>ReXNet 2.0</td>\n<td>768</td>\n<td>10</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Post-processing</h2>\n<p>Post-processing was shown to be very effective in last year's top solutions. We incorporated ideas of top3 solutions</p>\n<ol>\n<li>Combine cosine similarity and classification probabilities (our <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">3rd place solution</a>)</li>\n<li>Penalize similarity with non-landmark images (<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821\" target=\"_blank\">1st place solution</a>)</li>\n<li>Set similarity=0 if top5 similarities with non-landmarks &gt; 0.5 (<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299\" target=\"_blank\">2nd place solution</a>)</li>\n<li>Penalize CosSim(query, index) for index class’s count in train set (ours, new)</li>\n</ol>",
      "rawMarkdown": "Thanks to my longtime teammate @haqishen and new teammate @hongweizhang . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.\n\n## Approach\nQishen and I were part of the 3rd place team (together with @garybios @alexanderliao ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. \n\nThis year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.\n\n## Training recipe\nWe mostly follow our last year's solution (see [here](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757) for details). The recipes are\n- sub-center ArcFace with dynamic margins\n- progressive training with increasing image sizes\n- the indispensable [timm](https://github.com/rwightman/pytorch-image-models) library\n- cosine learning schedule with Adam/AdamW optimizer\n- multi-GPU training with DistributedDataParallel\n\n## Model choices\nThere are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -> 512 -> 640/768). \n\nWe found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.\n\nBesides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.\n\n## The ensemble\nOur final ensemble consists of 11 models: 3 transformers, 3 new CNNs and 5 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).\n\nOur best retrieval ensemble has 4 fewer CNNs and only 7 models in total (3 transformers + 4 CNNs), because the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.\n\n|       Model      | Image size | Total epochs | Finetune 2020 | cv GAP (recognition) | cv mAP@100 (retrieval) |\n|:----------------:|:----------:|:------------:|:-------------:|:--------------------:|:----------------------:|\n|     Swin base    |     384    |      60      |               |        0.7049        |         0.4442         |\n|    Swin large    |     384    |      60      |               |        0.6775        |         0.5161         |\n|     ViT large    |     384    |      50      |               |        0.6589        |         0.5633         |\n|   ECA NFNet L2   |     640    |      38      |               |        0.7053        |         0.3600         |\n| EfficientNet v2m |     640    |      45      |               |        0.7136        |         0.3811         |\n| EfficientNet v2l |     512    |      30      |               |        0.7146        |         0.4117         |\n|  EfficientNet B4 |     768    |      10      |       ✓       |                      |                        |\n|  EfficientNet B5 |     768    |      20      |       ✓       |                      |                        |\n|  EfficientNet B6 |     512    |      20      |       ✓       |                      |                        |\n|  EfficientNet B7 |     672    |      20      |       ✓       |                      |                        |\n|    ReXNet 2.0    |     768    |      10      |       ✓       |                      |                        |\n\n## Post-processing \nPost-processing was shown to be very effective in last year's top solutions. We incorporated ideas of top3 solutions\n1. Combine cosine similarity and classification probabilities (our [3rd place solution](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757))\n2. Penalize similarity with non-landmark images ([1st place solution](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821))\n3. Set similarity=0 if top5 similarities with non-landmarks > 0.5 ([2nd place solution](https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299))\n4. Penalize CosSim(query, index) for index class’s count in train set (ours, new)",
      "votes": null
    },
    {
      "id": "1543582",
      "postDate": "10/13/2021 16:28:19",
      "content": "<p>Wow! Congrats <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> for your 4th place!</p>\n<p>I'd like to train a Swin, 60 epochs on full 200K landmarks dataset.</p>\n<p>I think it's going to take one life to train it on my single RTX 3090.</p>\n<p>Could you share plz the scheduler and the optimizer that you used to train the Swin base as well as hyperparams?</p>\n<p>Any other tip on training Swim / visual transformers is very appreciated!</p>",
      "rawMarkdown": "Wow! Congrats @boliu0 for your 4th place!\n\nI'd like to train a Swin, 60 epochs on full 200K landmarks dataset.\n\nI think it's going to take one life to train it on my single RTX 3090.\n\nCould you share plz the scheduler and the optimizer that you used to train the Swin base as well as hyperparams?\n\nAny other tip on training Swim / visual transformers is very appreciated!",
      "votes": null
    },
    {
      "id": "1544577",
      "postDate": "10/14/2021 13:13:00",
      "content": "<p>Thanks. For Swin Large at image size = 384, I used 8 V100 GPUs (32G) with batch size = 9 per GPU. It takes about 8.1 hours/epoch, or about 20 days for 60 epochs. </p>\n<p>Training schedule is cosine schedule (see our last <a href=\"https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution\" target=\"_blank\">solution</a>) with initial learning rate =2.5e6</p>\n<p>optimizer is AdamW</p>\n<p>On a single RTX 3090, you probably want to use Swin base, and train it on cGLD2 only, for fewer epochs. Just for fun though, it would take you only 20<em>(32</em>8/24) = 213 days to train 60 epochs on all data. If you start now, you can definitely make it for next year's landmark competitions 😉</p>",
      "rawMarkdown": "Thanks. For Swin Large at image size = 384, I used 8 V100 GPUs (32G) with batch size = 9 per GPU. It takes about 8.1 hours/epoch, or about 20 days for 60 epochs. \n\nTraining schedule is cosine schedule (see our last [solution](https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution)) with initial learning rate =2.5e6\n\noptimizer is AdamW\n\nOn a single RTX 3090, you probably want to use Swin base, and train it on cGLD2 only, for fewer epochs. Just for fun though, it would take you only 20*(32*8/24) = 213 days to train 60 epochs on all data. If you start now, you can definitely make it for next year's landmark competitions 😉",
      "votes": null
    },
    {
      "id": "1546541",
      "postDate": "10/16/2021 09:20:03",
      "content": "<p>213 days!!! Sure, I have time for the landmark 2022 competition … LOL!!!</p>\n<p>Definitely 213 days of training not an option for me. I hope SOTA brings a computationally cheaper vision transformer architecture within the next 213 days!</p>\n<p>I've to think in something else if I want to use Swin for this dataset</p>\n<p>Thanks a lot for your detailed response <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> !</p>",
      "rawMarkdown": "213 days!!! Sure, I have time for the landmark 2022 competition ... LOL!!!\n\nDefinitely 213 days of training not an option for me. I hope SOTA brings a computationally cheaper vision transformer architecture within the next 213 days!\n\nI've to think in something else if I want to use Swin for this dataset\n\nThanks a lot for your detailed response @boliu0 !",
      "votes": null
    },
    {
      "id": "1547621",
      "postDate": "10/17/2021 12:20:37",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": null
    },
    {
      "id": "1594793",
      "postDate": "11/25/2021 06:04:23",
      "content": "<p>Wow! Congratulations！！！<br>\n<code>the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.</code><br>\nThis is a very interesting point. Generally speaking, our intuition will think that CNN will be more suitable for retrieval. But your result is completely different, can you explain it? It makes me very curious. <br>\nCongratulations for your 4th place again!</p>",
      "rawMarkdown": "Wow! Congratulations！！！\n`the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.`\nThis is a very interesting point. Generally speaking, our intuition will think that CNN will be more suitable for retrieval. But your result is completely different, can you explain it? It makes me very curious. \nCongratulations for your 4th place again!",
      "votes": null
    },
    {
      "id": "1656353",
      "postDate": "01/19/2022 09:38:00",
      "content": "<p>What loss function have u used?</p>",
      "rawMarkdown": "What loss function have u used?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1543582,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "10/13/2021 16:28:19",
      "content": "<p>Wow! Congrats <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> for your 4th place!</p>\n<p>I'd like to train a Swin, 60 epochs on full 200K landmarks dataset.</p>\n<p>I think it's going to take one life to train it on my single RTX 3090.</p>\n<p>Could you share plz the scheduler and the optimizer that you used to train the Swin base as well as hyperparams?</p>\n<p>Any other tip on training Swim / visual transformers is very appreciated!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1544577,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "10/14/2021 13:13:00",
          "content": "<p>Thanks. For Swin Large at image size = 384, I used 8 V100 GPUs (32G) with batch size = 9 per GPU. It takes about 8.1 hours/epoch, or about 20 days for 60 epochs. </p>\n<p>Training schedule is cosine schedule (see our last <a href=\"https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution\" target=\"_blank\">solution</a>) with initial learning rate =2.5e6</p>\n<p>optimizer is AdamW</p>\n<p>On a single RTX 3090, you probably want to use Swin base, and train it on cGLD2 only, for fewer epochs. Just for fun though, it would take you only 20<em>(32</em>8/24) = 213 days to train 60 epochs on all data. If you start now, you can definitely make it for next year's landmark competitions 😉</p>",
          "votes": null,
          "replies": [
            {
              "id": 1546541,
              "author_name": "virilo",
              "author_url": "",
              "post_date": "10/16/2021 09:20:03",
              "content": "<p>213 days!!! Sure, I have time for the landmark 2022 competition … LOL!!!</p>\n<p>Definitely 213 days of training not an option for me. I hope SOTA brings a computationally cheaper vision transformer architecture within the next 213 days!</p>\n<p>I've to think in something else if I want to use Swin for this dataset</p>\n<p>Thanks a lot for your detailed response <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> !</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1547621,
      "author_name": "shakshyathedetector",
      "author_url": "",
      "post_date": "10/17/2021 12:20:37",
      "content": "<p>congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1594793,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "11/25/2021 06:04:23",
      "content": "<p>Wow! Congratulations！！！<br>\n<code>the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.</code><br>\nThis is a very interesting point. Generally speaking, our intuition will think that CNN will be more suitable for retrieval. But your result is completely different, can you explain it? It makes me very curious. <br>\nCongratulations for your 4th place again!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1656353,
      "author_name": "zelenadinja",
      "author_url": "",
      "post_date": "01/19/2022 09:38:00",
      "content": "<p>What loss function have u used?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1539242": "Thanks to my longtime teammate @haqishen and new teammate @hongweizhang . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.\n\n## Approach\nQishen and I were part of the 3rd place team (together with @garybios @alexanderliao ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. \n\nThis year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.\n\n## Training recipe\nWe mostly follow our last year's solution (see [here](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757) for details). The recipes are\n- sub-center ArcFace with dynamic margins\n- progressive training with increasing image sizes\n- the indispensable [timm](https://github.com/rwightman/pytorch-image-models) library\n- cosine learning schedule with Adam/AdamW optimizer\n- multi-GPU training with DistributedDataParallel\n\n## Model choices\nThere are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -> 512 -> 640/768). \n\nWe found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.\n\nBesides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.\n\n## The ensemble\nOur final ensemble consists of 11 models: 3 transformers, 3 new CNNs and 5 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).\n\nOur best retrieval ensemble has 4 fewer CNNs and only 7 models in total (3 transformers + 4 CNNs), because the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.\n\n|       Model      | Image size | Total epochs | Finetune 2020 | cv GAP (recognition) | cv mAP@100 (retrieval) |\n|:----------------:|:----------:|:------------:|:-------------:|:--------------------:|:----------------------:|\n|     Swin base    |     384    |      60      |               |        0.7049        |         0.4442         |\n|    Swin large    |     384    |      60      |               |        0.6775        |         0.5161         |\n|     ViT large    |     384    |      50      |               |        0.6589        |         0.5633         |\n|   ECA NFNet L2   |     640    |      38      |               |        0.7053        |         0.3600         |\n| EfficientNet v2m |     640    |      45      |               |        0.7136        |         0.3811         |\n| EfficientNet v2l |     512    |      30      |               |        0.7146        |         0.4117         |\n|  EfficientNet B4 |     768    |      10      |       ✓       |                      |                        |\n|  EfficientNet B5 |     768    |      20      |       ✓       |                      |                        |\n|  EfficientNet B6 |     512    |      20      |       ✓       |                      |                        |\n|  EfficientNet B7 |     672    |      20      |       ✓       |                      |                        |\n|    ReXNet 2.0    |     768    |      10      |       ✓       |                      |                        |\n\n## Post-processing \nPost-processing was shown to be very effective in last year's top solutions. We incorporated ideas of top3 solutions\n1. Combine cosine similarity and classification probabilities (our [3rd place solution](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757))\n2. Penalize similarity with non-landmark images ([1st place solution](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821))\n3. Set similarity=0 if top5 similarities with non-landmarks > 0.5 ([2nd place solution](https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299))\n4. Penalize CosSim(query, index) for index class’s count in train set (ours, new)",
    "1543582": "Wow! Congrats @boliu0 for your 4th place!\n\nI'd like to train a Swin, 60 epochs on full 200K landmarks dataset.\n\nI think it's going to take one life to train it on my single RTX 3090.\n\nCould you share plz the scheduler and the optimizer that you used to train the Swin base as well as hyperparams?\n\nAny other tip on training Swim / visual transformers is very appreciated!",
    "1544577": "Thanks. For Swin Large at image size = 384, I used 8 V100 GPUs (32G) with batch size = 9 per GPU. It takes about 8.1 hours/epoch, or about 20 days for 60 epochs. \n\nTraining schedule is cosine schedule (see our last [solution](https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution)) with initial learning rate =2.5e6\n\noptimizer is AdamW\n\nOn a single RTX 3090, you probably want to use Swin base, and train it on cGLD2 only, for fewer epochs. Just for fun though, it would take you only 20*(32*8/24) = 213 days to train 60 epochs on all data. If you start now, you can definitely make it for next year's landmark competitions 😉",
    "1546541": "213 days!!! Sure, I have time for the landmark 2022 competition ... LOL!!!\n\nDefinitely 213 days of training not an option for me. I hope SOTA brings a computationally cheaper vision transformer architecture within the next 213 days!\n\nI've to think in something else if I want to use Swin for this dataset\n\nThanks a lot for your detailed response @boliu0 !",
    "1547621": "congratulations",
    "1594793": "Wow! Congratulations！！！\n`the addition CNNs hurt the retrieval score. This is because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.`\nThis is a very interesting point. Generally speaking, our intuition will think that CNN will be more suitable for retrieval. But your result is completely different, can you explain it? It makes me very curious. \nCongratulations for your 4th place again!",
    "1656353": "What loss function have u used?"
  },
  "source": "meta"
}