{
  "id": 187961,
  "title": "6th place solution",
  "url": "/competitions/landmark-recognition-2020/writeups/yjcv-6th-place-solution",
  "author_name": "",
  "post_date": "2020-10-01T04:39:11.984440900Z",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We thank all organizers for this very exciting competition.  <br>\nCongratulations to all who finished the competition and to the winners.</p>\n<h2>Summary</h2>\n<ul>\n<li>kNN based on global descriptor</li>\n<li>Non-landmark filtering with similarity between test set and GLDv2 test set</li>\n<li>We tried local descriptor methods but they were not used in our final submission</li>\n</ul>\n<h2>Model Details</h2>\n<p>We trained CosFace based global features models.  <br>\nThe settings are almost the same as our models in the retrieval2020 (See <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175472\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175472</a>)</p>\n<ul>\n<li>Backbones: Ensemble of ResNeSt101, ResNeSt101, ResNeSt200</li>\n<li>Pooling: GeM (p=3) (Replace GeM p=3 with p=4 in testing)</li>\n<li>Head: FC-&gt;BN-&gt;L2</li>\n<li>Loss: CosFace with Label Smoothing</li>\n<li>Data Augmentation: HorizontalFlip, RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask</li>\n<li>LR: Cosine Annealing LR with warmup, training for 30 epochs + refine 5epochs</li>\n<li>Input image size in training: 352 (in refine: 640)</li>\n<li>Adding 600 non-landmark images sampled from the test set of the last year's recognition competition in training (this made the total 81,314 classes).</li>\n</ul>\n<p>Additionally, we tried \"Inplace Knowledge Distillation\" by MutualNet paper (<a href=\"https://arxiv.org/pdf/1909.12978.pdf\" target=\"_blank\">https://arxiv.org/pdf/1909.12978.pdf</a>) to train multi-scale image features efficiently.  <br>\nIt has been adopted for all models.</p>\n<h2>HOW descriptor based ASMK similarity</h2>\n<p>We tried HOW descriptor (<a href=\"https://arxiv.org/pdf/2007.13172.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.13172.pdf</a>) as an alternative of the global feature model.  <br>\nHOW aggregates CNN based local descriptors into a single global descriptor with ASMK.  <br>\nWe trained a ResNet50-based HOW descriptor model (our pytorch implementation) and got public score = 0.5284 (private = 0.5044).  <br>\nOne of our final submission was based on blending between global cosine similarity and ASMK similarity, that achieved the best public score = 0.6298. However, the other global-feature-only approach achieved the better private score.</p>\n<table>\n<thead>\n<tr>\n<th>Final submissions</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Global features</td>\n<td>0.5983</td>\n<td>0.6271</td>\n</tr>\n<tr>\n<td>Global + HOW ASMK</td>\n<td>0.5960</td>\n<td>0.6298</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<ul>\n<li>Reranking top-k with HOW ASMK similarity (Blending similarity was better)</li>\n<li>Reranking top-k with DELG local descriptors</li>\n<li>DBA</li>\n<li>Replacing the private train set with GLDv2clean<ul>\n<li>Undersampling up to 350 images per class</li>\n<li>Removing classes that do not exist in the private train set</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1033538",
      "postDate": "10/01/2020 04:39:11",
      "content": "<p>We thank all organizers for this very exciting competition.  <br>\nCongratulations to all who finished the competition and to the winners.</p>\n<h2>Summary</h2>\n<ul>\n<li>kNN based on global descriptor</li>\n<li>Non-landmark filtering with similarity between test set and GLDv2 test set</li>\n<li>We tried local descriptor methods but they were not used in our final submission</li>\n</ul>\n<h2>Model Details</h2>\n<p>We trained CosFace based global features models.  <br>\nThe settings are almost the same as our models in the retrieval2020 (See <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175472\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175472</a>)</p>\n<ul>\n<li>Backbones: Ensemble of ResNeSt101, ResNeSt101, ResNeSt200</li>\n<li>Pooling: GeM (p=3) (Replace GeM p=3 with p=4 in testing)</li>\n<li>Head: FC-&gt;BN-&gt;L2</li>\n<li>Loss: CosFace with Label Smoothing</li>\n<li>Data Augmentation: HorizontalFlip, RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask</li>\n<li>LR: Cosine Annealing LR with warmup, training for 30 epochs + refine 5epochs</li>\n<li>Input image size in training: 352 (in refine: 640)</li>\n<li>Adding 600 non-landmark images sampled from the test set of the last year's recognition competition in training (this made the total 81,314 classes).</li>\n</ul>\n<p>Additionally, we tried \"Inplace Knowledge Distillation\" by MutualNet paper (<a href=\"https://arxiv.org/pdf/1909.12978.pdf\" target=\"_blank\">https://arxiv.org/pdf/1909.12978.pdf</a>) to train multi-scale image features efficiently.  <br>\nIt has been adopted for all models.</p>\n<h2>HOW descriptor based ASMK similarity</h2>\n<p>We tried HOW descriptor (<a href=\"https://arxiv.org/pdf/2007.13172.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.13172.pdf</a>) as an alternative of the global feature model.  <br>\nHOW aggregates CNN based local descriptors into a single global descriptor with ASMK.  <br>\nWe trained a ResNet50-based HOW descriptor model (our pytorch implementation) and got public score = 0.5284 (private = 0.5044).  <br>\nOne of our final submission was based on blending between global cosine similarity and ASMK similarity, that achieved the best public score = 0.6298. However, the other global-feature-only approach achieved the better private score.</p>\n<table>\n<thead>\n<tr>\n<th>Final submissions</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Global features</td>\n<td>0.5983</td>\n<td>0.6271</td>\n</tr>\n<tr>\n<td>Global + HOW ASMK</td>\n<td>0.5960</td>\n<td>0.6298</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<ul>\n<li>Reranking top-k with HOW ASMK similarity (Blending similarity was better)</li>\n<li>Reranking top-k with DELG local descriptors</li>\n<li>DBA</li>\n<li>Replacing the private train set with GLDv2clean<ul>\n<li>Undersampling up to 350 images per class</li>\n<li>Removing classes that do not exist in the private train set</li></ul></li>\n</ul>",
      "rawMarkdown": "We thank all organizers for this very exciting competition.  \nCongratulations to all who finished the competition and to the winners.\n\n## Summary\n\n- kNN based on global descriptor\n- Non-landmark filtering with similarity between test set and GLDv2 test set\n- We tried local descriptor methods but they were not used in our final submission\n\n\n## Model Details\n\nWe trained CosFace based global features models.  \nThe settings are almost the same as our models in the retrieval2020 (See https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175472)\n\n- Backbones: Ensemble of ResNeSt101, ResNeSt101, ResNeSt200\n- Pooling: GeM (p=3) (Replace GeM p=3 with p=4 in testing)\n- Head: FC->BN->L2\n- Loss: CosFace with Label Smoothing\n- Data Augmentation: HorizontalFlip, RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask\n- LR: Cosine Annealing LR with warmup, training for 30 epochs + refine 5epochs\n- Input image size in training: 352 (in refine: 640)\n- Adding 600 non-landmark images sampled from the test set of the last year's recognition competition in training (this made the total 81,314 classes).\n\nAdditionally, we tried \"Inplace Knowledge Distillation\" by MutualNet paper (https://arxiv.org/pdf/1909.12978.pdf) to train multi-scale image features efficiently.  \nIt has been adopted for all models.\n\n\n## HOW descriptor based ASMK similarity\n\nWe tried HOW descriptor (https://arxiv.org/pdf/2007.13172.pdf) as an alternative of the global feature model.  \nHOW aggregates CNN based local descriptors into a single global descriptor with ASMK.  \nWe trained a ResNet50-based HOW descriptor model (our pytorch implementation) and got public score = 0.5284 (private = 0.5044).  \nOne of our final submission was based on blending between global cosine similarity and ASMK similarity, that achieved the best public score = 0.6298. However, the other global-feature-only approach achieved the better private score.\n\n| Final submissions | Private | Public |\n|-------------------|---------|--------|\n| Global features   | 0.5983  | 0.6271 |\n| Global + HOW ASMK | 0.5960  | 0.6298 |\n\n\n## What did not work\n\n- Reranking top-k with HOW ASMK similarity (Blending similarity was better)\n- Reranking top-k with DELG local descriptors\n- DBA\n- Replacing the private train set with GLDv2clean\n  - Undersampling up to 350 images per class\n  - Removing classes that do not exist in the private train set",
      "votes": null
    },
    {
      "id": "1033545",
      "postDate": "10/01/2020 04:49:55",
      "content": "<p>Congratulations with the strong finish! Seems like you guys were also using GLD v2 test set for non-landmark filtering as the 1st place team 😀 I have a question: what is your single model's performance, without filtering (using the organizer's pipeline)?</p>",
      "rawMarkdown": "Congratulations with the strong finish! Seems like you guys were also using GLD v2 test set for non-landmark filtering as the 1st place team 😀 I have a question: what is your single model's performance, without filtering (using the organizer's pipeline)?",
      "votes": null
    },
    {
      "id": "1034613",
      "postDate": "10/02/2020 01:02:05",
      "content": "<p>Thanks.</p>\n<p>The differences in performance with and without non-landmark filtering are as follows.</p>\n<table>\n<thead>\n<tr>\n<th>ResNeSt101 single model</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>with non-landmark filtering</td>\n<td>0.5950</td>\n<td>0.6148</td>\n</tr>\n<tr>\n<td>without non-landmark filtering</td>\n<td>0.5542</td>\n<td>0.5777</td>\n</tr>\n</tbody>\n</table>\n<p>(without non-landmark flltering = kNN with private train set only)</p>\n<p>But, since our model trained with the addition of a non-landmark class, so this model is likely to have some ability to identify non-landmarks.</p>",
      "rawMarkdown": "Thanks.\n\nThe differences in performance with and without non-landmark filtering are as follows.\n\n|ResNeSt101 single model        | Private | Public             |\n|:------------------------------|:-------|:----------------|\n|with non-landmark filtering     |0.5950 |0.6148|\n|without non-landmark filtering|0.5542 |0.5777|\n\n(without non-landmark flltering = kNN with private train set only)\n\nBut, since our model trained with the addition of a non-landmark class, so this model is likely to have some ability to identify non-landmarks.",
      "votes": null
    },
    {
      "id": "1034775",
      "postDate": "10/02/2020 06:58:20",
      "content": "<p>Thanks a lot for the details!</p>",
      "rawMarkdown": "Thanks a lot for the details!",
      "votes": null
    },
    {
      "id": "1035006",
      "postDate": "10/02/2020 11:30:16",
      "content": "<p>You have done it Great, Would love to know more on it 👍</p>",
      "rawMarkdown": "You have done it Great, Would love to know more on it 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1033545,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "10/01/2020 04:49:55",
      "content": "<p>Congratulations with the strong finish! Seems like you guys were also using GLD v2 test set for non-landmark filtering as the 1st place team 😀 I have a question: what is your single model's performance, without filtering (using the organizer's pipeline)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1034613,
          "author_name": "knjcode",
          "author_url": "",
          "post_date": "10/02/2020 01:02:05",
          "content": "<p>Thanks.</p>\n<p>The differences in performance with and without non-landmark filtering are as follows.</p>\n<table>\n<thead>\n<tr>\n<th>ResNeSt101 single model</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>with non-landmark filtering</td>\n<td>0.5950</td>\n<td>0.6148</td>\n</tr>\n<tr>\n<td>without non-landmark filtering</td>\n<td>0.5542</td>\n<td>0.5777</td>\n</tr>\n</tbody>\n</table>\n<p>(without non-landmark flltering = kNN with private train set only)</p>\n<p>But, since our model trained with the addition of a non-landmark class, so this model is likely to have some ability to identify non-landmarks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1034775,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "10/02/2020 06:58:20",
          "content": "<p>Thanks a lot for the details!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1035006,
      "author_name": "ankur95",
      "author_url": "",
      "post_date": "10/02/2020 11:30:16",
      "content": "<p>You have done it Great, Would love to know more on it 👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1033538": "We thank all organizers for this very exciting competition.  \nCongratulations to all who finished the competition and to the winners.\n\n## Summary\n\n- kNN based on global descriptor\n- Non-landmark filtering with similarity between test set and GLDv2 test set\n- We tried local descriptor methods but they were not used in our final submission\n\n\n## Model Details\n\nWe trained CosFace based global features models.  \nThe settings are almost the same as our models in the retrieval2020 (See https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175472)\n\n- Backbones: Ensemble of ResNeSt101, ResNeSt101, ResNeSt200\n- Pooling: GeM (p=3) (Replace GeM p=3 with p=4 in testing)\n- Head: FC->BN->L2\n- Loss: CosFace with Label Smoothing\n- Data Augmentation: HorizontalFlip, RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask\n- LR: Cosine Annealing LR with warmup, training for 30 epochs + refine 5epochs\n- Input image size in training: 352 (in refine: 640)\n- Adding 600 non-landmark images sampled from the test set of the last year's recognition competition in training (this made the total 81,314 classes).\n\nAdditionally, we tried \"Inplace Knowledge Distillation\" by MutualNet paper (https://arxiv.org/pdf/1909.12978.pdf) to train multi-scale image features efficiently.  \nIt has been adopted for all models.\n\n\n## HOW descriptor based ASMK similarity\n\nWe tried HOW descriptor (https://arxiv.org/pdf/2007.13172.pdf) as an alternative of the global feature model.  \nHOW aggregates CNN based local descriptors into a single global descriptor with ASMK.  \nWe trained a ResNet50-based HOW descriptor model (our pytorch implementation) and got public score = 0.5284 (private = 0.5044).  \nOne of our final submission was based on blending between global cosine similarity and ASMK similarity, that achieved the best public score = 0.6298. However, the other global-feature-only approach achieved the better private score.\n\n| Final submissions | Private | Public |\n|-------------------|---------|--------|\n| Global features   | 0.5983  | 0.6271 |\n| Global + HOW ASMK | 0.5960  | 0.6298 |\n\n\n## What did not work\n\n- Reranking top-k with HOW ASMK similarity (Blending similarity was better)\n- Reranking top-k with DELG local descriptors\n- DBA\n- Replacing the private train set with GLDv2clean\n  - Undersampling up to 350 images per class\n  - Removing classes that do not exist in the private train set",
    "1033545": "Congratulations with the strong finish! Seems like you guys were also using GLD v2 test set for non-landmark filtering as the 1st place team 😀 I have a question: what is your single model's performance, without filtering (using the organizer's pipeline)?",
    "1034613": "Thanks.\n\nThe differences in performance with and without non-landmark filtering are as follows.\n\n|ResNeSt101 single model        | Private | Public             |\n|:------------------------------|:-------|:----------------|\n|with non-landmark filtering     |0.5950 |0.6148|\n|without non-landmark filtering|0.5542 |0.5777|\n\n(without non-landmark flltering = kNN with private train set only)\n\nBut, since our model trained with the addition of a non-landmark class, so this model is likely to have some ability to identify non-landmarks.",
    "1034775": "Thanks a lot for the details!",
    "1035006": "You have done it Great, Would love to know more on it 👍"
  },
  "source": "meta"
}