{
  "id": 277382,
  "title": "3rd place solution",
  "url": "/competitions/landmark-retrieval-2021/writeups/all-data-are-ext-3rd-place-solution",
  "author_name": "",
  "post_date": "2021-10-09T08:14:47.087329700Z",
  "votes": 17,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to my longtime teammate <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and new teammate <a href=\"https://www.kaggle.com/hongweizhang\" target=\"_blank\">@hongweizhang</a> . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.</p>\n<h2>Approach</h2>\n<p>Qishen and I were part of the 3rd place team (together with <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a> ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. </p>\n<p>This year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.</p>\n<h2>Training recipe</h2>\n<p>We mostly follow our last year's solution (see <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">here</a> for details). The recipes are</p>\n<ul>\n<li>sub-center ArcFace with dynamic margins</li>\n<li>progressive training with increasing image sizes</li>\n<li>the indispensable <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library</li>\n<li>cosine learning schedule with Adam/AdamW optimizer</li>\n<li>multi-GPU training with DistributedDataParallel</li>\n</ul>\n<h2>Model choices</h2>\n<p>There are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -&gt; 512 -&gt; 640/768). </p>\n<p>We found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.</p>\n<p>Besides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.</p>\n<h2>The ensemble</h2>\n<p>Our final ensemble consists of 7 models: 3 transformers, 2 new CNNs and 2 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).</p>\n<p>Adding more CNNs to the ensemble hurts retrieval score but helps recognition score -- our recognition's best ensemble has 4 more CNNs, because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Image size</th>\n<th>Total epochs</th>\n<th>Finetune 2020</th>\n<th>cv GAP (recognition)</th>\n<th>cv mAP@100 (retrieval)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Swin base</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.7049</td>\n<td>0.4442</td>\n</tr>\n<tr>\n<td>Swin large</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.6775</td>\n<td>0.5161</td>\n</tr>\n<tr>\n<td>ViT large</td>\n<td>384</td>\n<td>50</td>\n<td></td>\n<td>0.6589</td>\n<td>0.5633</td>\n</tr>\n<tr>\n<td>ECA NFNet L2</td>\n<td>512</td>\n<td>30</td>\n<td></td>\n<td>0.7021</td>\n<td>0.3565</td>\n</tr>\n<tr>\n<td>EfficientNet v2l</td>\n<td>640</td>\n<td>40</td>\n<td></td>\n<td>0.7129</td>\n<td>0.4158</td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>512</td>\n<td>10</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B7</td>\n<td>672</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Post-processing a.k.a. reranking</h2>\n<p>The reranking is inspired by and improved upon 2019 Retrieval's <a href=\"https://arxiv.org/abs/1906.04087\" target=\"_blank\">winning solution</a>. Their method was to rerank all the \"positive\" index images before all the \"negative\" ones. The downside of this approach is that, the \"positive\" and \"negative\" are predicted by the model, which may be incorrect.</p>\n<p>We use a softer approach: when reranking, give a boost to the similarity score for \"positives\"  and give a penalty to the similarity score for \"negative\". The boost and penalty depends on how confident we are that an index image is a positive or negative.</p>",
  "messages": [
    {
      "id": "1539224",
      "postDate": "10/09/2021 08:14:47",
      "content": "<p>Thanks to my longtime teammate <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and new teammate <a href=\"https://www.kaggle.com/hongweizhang\" target=\"_blank\">@hongweizhang</a> . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.</p>\n<h2>Approach</h2>\n<p>Qishen and I were part of the 3rd place team (together with <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a> ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. </p>\n<p>This year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.</p>\n<h2>Training recipe</h2>\n<p>We mostly follow our last year's solution (see <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">here</a> for details). The recipes are</p>\n<ul>\n<li>sub-center ArcFace with dynamic margins</li>\n<li>progressive training with increasing image sizes</li>\n<li>the indispensable <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library</li>\n<li>cosine learning schedule with Adam/AdamW optimizer</li>\n<li>multi-GPU training with DistributedDataParallel</li>\n</ul>\n<h2>Model choices</h2>\n<p>There are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -&gt; 512 -&gt; 640/768). </p>\n<p>We found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.</p>\n<p>Besides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.</p>\n<h2>The ensemble</h2>\n<p>Our final ensemble consists of 7 models: 3 transformers, 2 new CNNs and 2 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).</p>\n<p>Adding more CNNs to the ensemble hurts retrieval score but helps recognition score -- our recognition's best ensemble has 4 more CNNs, because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Image size</th>\n<th>Total epochs</th>\n<th>Finetune 2020</th>\n<th>cv GAP (recognition)</th>\n<th>cv mAP@100 (retrieval)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Swin base</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.7049</td>\n<td>0.4442</td>\n</tr>\n<tr>\n<td>Swin large</td>\n<td>384</td>\n<td>60</td>\n<td></td>\n<td>0.6775</td>\n<td>0.5161</td>\n</tr>\n<tr>\n<td>ViT large</td>\n<td>384</td>\n<td>50</td>\n<td></td>\n<td>0.6589</td>\n<td>0.5633</td>\n</tr>\n<tr>\n<td>ECA NFNet L2</td>\n<td>512</td>\n<td>30</td>\n<td></td>\n<td>0.7021</td>\n<td>0.3565</td>\n</tr>\n<tr>\n<td>EfficientNet v2l</td>\n<td>640</td>\n<td>40</td>\n<td></td>\n<td>0.7129</td>\n<td>0.4158</td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>512</td>\n<td>10</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNet B7</td>\n<td>672</td>\n<td>20</td>\n<td>✓</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Post-processing a.k.a. reranking</h2>\n<p>The reranking is inspired by and improved upon 2019 Retrieval's <a href=\"https://arxiv.org/abs/1906.04087\" target=\"_blank\">winning solution</a>. Their method was to rerank all the \"positive\" index images before all the \"negative\" ones. The downside of this approach is that, the \"positive\" and \"negative\" are predicted by the model, which may be incorrect.</p>\n<p>We use a softer approach: when reranking, give a boost to the similarity score for \"positives\"  and give a penalty to the similarity score for \"negative\". The boost and penalty depends on how confident we are that an index image is a positive or negative.</p>",
      "rawMarkdown": "Thanks to my longtime teammate @haqishen and new teammate @hongweizhang . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.\n\n## Approach\nQishen and I were part of the 3rd place team (together with @garybios @alexanderliao ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. \n\nThis year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.\n\n## Training recipe\nWe mostly follow our last year's solution (see [here](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757) for details). The recipes are\n- sub-center ArcFace with dynamic margins\n- progressive training with increasing image sizes\n- the indispensable [timm](https://github.com/rwightman/pytorch-image-models) library\n- cosine learning schedule with Adam/AdamW optimizer\n- multi-GPU training with DistributedDataParallel\n\n## Model choices\nThere are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -> 512 -> 640/768). \n\nWe found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.\n\nBesides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.\n\n## The ensemble\nOur final ensemble consists of 7 models: 3 transformers, 2 new CNNs and 2 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).\n\nAdding more CNNs to the ensemble hurts retrieval score but helps recognition score -- our recognition's best ensemble has 4 more CNNs, because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.\n\n|       Model      | Image size | Total epochs | Finetune 2020 | cv GAP (recognition) | cv mAP@100 (retrieval) |\n|:----------------:|:----------:|:------------:|:-------------:|:--------------------:|:----------------------:|\n|     Swin base    |     384    |      60      |               |        0.7049        |         0.4442         |\n|    Swin large    |     384    |      60      |               |        0.6775        |         0.5161         |\n|     ViT large    |     384    |      50      |               |        0.6589        |         0.5633         |\n|   ECA NFNet L2   |     512    |      30      |               |        0.7021        |         0.3565         |\n| EfficientNet v2l |     640    |      40      |               |        0.7129        |         0.4158         |\n|  EfficientNet B6 |     512    |      10      |       ✓       |                      |                        |\n|  EfficientNet B7 |     672    |      20      |       ✓       |                      |                        |\n\n## Post-processing a.k.a. reranking\nThe reranking is inspired by and improved upon 2019 Retrieval's [winning solution](https://arxiv.org/abs/1906.04087). Their method was to rerank all the \"positive\" index images before all the \"negative\" ones. The downside of this approach is that, the \"positive\" and \"negative\" are predicted by the model, which may be incorrect.\n\nWe use a softer approach: when reranking, give a boost to the similarity score for \"positives\"  and give a penalty to the similarity score for \"negative\". The boost and penalty depends on how confident we are that an index image is a positive or negative.",
      "votes": null
    },
    {
      "id": "1540000",
      "postDate": "10/10/2021 04:35:37",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1540000,
      "author_name": "saleesh",
      "author_url": "",
      "post_date": "10/10/2021 04:35:37",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1539224": "Thanks to my longtime teammate @haqishen and new teammate @hongweizhang . It was again great teamwork. We got 3rd place in retrieval and 4th place in recognition competition. I'm posting the solutions in both forums for easy future reference, even though there is large overlap.\n\n## Approach\nQishen and I were part of the 3rd place team (together with @garybios @alexanderliao ) in recognition 2020. Last year we developed a strong model, i.e., sub-center ArcFace with dynamic margins. \n\nThis year two competitions run simultaneously with tighter deadlines, so we largely re-used last year's pipeline, with newer architectures and better post-processing. We trained models the same way for both retrieval and recognition competitions, but selected different model combinations and different post-processing for final submissions.\n\n## Training recipe\nWe mostly follow our last year's solution (see [here](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757) for details). The recipes are\n- sub-center ArcFace with dynamic margins\n- progressive training with increasing image sizes\n- the indispensable [timm](https://github.com/rwightman/pytorch-image-models) library\n- cosine learning schedule with Adam/AdamW optimizer\n- multi-GPU training with DistributedDataParallel\n\n## Model choices\nThere are quite a few new SOTA image models published in the past year, notably EfficientNet v2, NFNet, ViT, Swin transformers. We trained all these new models using last year's recipe. For transformers, we used 384 image size. For CNN models, we used progressively larger image sizes (256 -> 512 -> 640/768). \n\nWe found that the CNNs and transformers perform equally well on local GAP scores (recognition metric), but transformers have significantly higher local mAP (retrieval metric) even with smaller 384 image size. We think this is because transformers are patch-based and can capture more local features and their interactions, which are more important for retrieval tasks.\n\nBesides the new models, we also reused some of our last year's models by finetuning, mostly EfficientNets.\n\n## The ensemble\nOur final ensemble consists of 7 models: 3 transformers, 2 new CNNs and 2 old CNNs. Their specifications and local scores are below. Old CNNs' CV scores are not shown because they are leaky (last year's fold splits were different).\n\nAdding more CNNs to the ensemble hurts retrieval score but helps recognition score -- our recognition's best ensemble has 4 more CNNs, because transformers are more important for retrieval, and adding more CNNs would reduce transformers' weights.\n\n|       Model      | Image size | Total epochs | Finetune 2020 | cv GAP (recognition) | cv mAP@100 (retrieval) |\n|:----------------:|:----------:|:------------:|:-------------:|:--------------------:|:----------------------:|\n|     Swin base    |     384    |      60      |               |        0.7049        |         0.4442         |\n|    Swin large    |     384    |      60      |               |        0.6775        |         0.5161         |\n|     ViT large    |     384    |      50      |               |        0.6589        |         0.5633         |\n|   ECA NFNet L2   |     512    |      30      |               |        0.7021        |         0.3565         |\n| EfficientNet v2l |     640    |      40      |               |        0.7129        |         0.4158         |\n|  EfficientNet B6 |     512    |      10      |       ✓       |                      |                        |\n|  EfficientNet B7 |     672    |      20      |       ✓       |                      |                        |\n\n## Post-processing a.k.a. reranking\nThe reranking is inspired by and improved upon 2019 Retrieval's [winning solution](https://arxiv.org/abs/1906.04087). Their method was to rerank all the \"positive\" index images before all the \"negative\" ones. The downside of this approach is that, the \"positive\" and \"negative\" are predicted by the model, which may be incorrect.\n\nWe use a softer approach: when reranking, give a boost to the similarity score for \"positives\"  and give a penalty to the similarity score for \"negative\". The boost and penalty depends on how confident we are that an index image is a positive or negative.",
    "1540000": "Thanks for sharing."
  },
  "source": "meta"
}