{
  "id": 175472,
  "title": "9th place solution overview",
  "url": "/competitions/landmark-retrieval-2020/discussion/175472",
  "author_name": "kenji",
  "post_date": "2020-08-18T09:22:37.546000",
  "votes": 39,
  "comment_count": 17,
  "views": 0,
  "content": "<p>We thank all organizers for this very exciting competition.<br>\nCongratulations to all who finished the competition and to the winners.</p>\n<p>We trained models using only GLDv2clean dataset on PyTorch and converted them to TensorFlow’s saved_model.<br>\nThen, we validated them using GLR2019 public and private datasets.</p>\n<h2>Final submission</h2>\n<p>ResNeSt50 + ResNet101 + ResNet152 + SEResNeXt101</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>output dim</th>\n<th>input size</th>\n<th>GLR2019 mAP@100</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNeSt50</td>\n<td>512</td>\n<td>416</td>\n<td>0.3114</td>\n<td>0.335</td>\n<td>0.292</td>\n</tr>\n<tr>\n<td>ResNet101</td>\n<td>512</td>\n<td>608</td>\n<td>0.3243</td>\n<td>0.346</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>ResNet152</td>\n<td>512</td>\n<td>480</td>\n<td>0.3180</td>\n<td>0.336</td>\n<td>0.288</td>\n</tr>\n<tr>\n<td>SEResNeXt101</td>\n<td>512</td>\n<td>480</td>\n<td>0.3209</td>\n<td>0.337</td>\n<td>0.296</td>\n</tr>\n<tr>\n<td>Ensemble of 4 models</td>\n<td>2048 (concat)</td>\n<td>-</td>\n<td>0.3396</td>\n<td>0.361</td>\n<td>0.317</td>\n</tr>\n</tbody>\n</table>\n<h2>Model details</h2>\n<ul>\n<li>Backbones: Ensemble of ResNeSt50, ResNet101, ResNet152 and SEResNeXt101</li>\n<li>Pooling: GeM (p=3)</li>\n<li>Head: FC-&gt;BN-&gt;L2 (the same as the last year’s first place team)</li>\n<li>Loss: CosFace with Label Smoothing (ArcFace was also good, but CosFace was better)</li>\n<li>Data Augmentation: RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask</li>\n<li>LR: Cosine Annealing LR with warmup, training for 30 epochs</li>\n<li>Input image size in training: 352</li>\n</ul>\n<h2>What we tried and worked</h2>\n<ul>\n<li>Automatic mixed precision training</li>\n<li>Replace GeM p=3 with p=4 in testing</li>\n<li>Increase input image size last few epochs of training with freezed BN</li>\n<li>Large and multiple input image sizes in testing: 416, 480 and 608</li>\n<li>MVArcFace</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>PCA whitening</li>\n<li>Maintaining the aspect ratio of input images in testing (perhaps because our models were trained on square images)</li>\n<li>Combination of arcface loss and pairwise (e.g., triplet) loss<br>\nWe have tried arcface loss with pairwise loss (specially multi-similarity loss). However, the single arcface loss was better than multiple losses.</li>\n<li>EfficientNet</li>\n<li>Circle loss</li>\n<li>Removing noisy classes<br>\nWe have tried to remove the worst 3 noisy classes which have high variance based on arcface class weight, but it did not work.</li>\n</ul>\n<h2>What we have not tried</h2>\n<ul>\n<li>Training with GLDv1 dataset</li>\n<li>Training without changing the aspect ratio of images </li>\n</ul>",
  "messages": [
    {
      "id": 975396,
      "postDate": "2020-08-18T09:22:37.547Z",
      "content": "<p>We thank all organizers for this very exciting competition.<br>\nCongratulations to all who finished the competition and to the winners.</p>\n<p>We trained models using only GLDv2clean dataset on PyTorch and converted them to TensorFlow’s saved_model.<br>\nThen, we validated them using GLR2019 public and private datasets.</p>\n<h2>Final submission</h2>\n<p>ResNeSt50 + ResNet101 + ResNet152 + SEResNeXt101</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>output dim</th>\n<th>input size</th>\n<th>GLR2019 mAP@100</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNeSt50</td>\n<td>512</td>\n<td>416</td>\n<td>0.3114</td>\n<td>0.335</td>\n<td>0.292</td>\n</tr>\n<tr>\n<td>ResNet101</td>\n<td>512</td>\n<td>608</td>\n<td>0.3243</td>\n<td>0.346</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>ResNet152</td>\n<td>512</td>\n<td>480</td>\n<td>0.3180</td>\n<td>0.336</td>\n<td>0.288</td>\n</tr>\n<tr>\n<td>SEResNeXt101</td>\n<td>512</td>\n<td>480</td>\n<td>0.3209</td>\n<td>0.337</td>\n<td>0.296</td>\n</tr>\n<tr>\n<td>Ensemble of 4 models</td>\n<td>2048 (concat)</td>\n<td>-</td>\n<td>0.3396</td>\n<td>0.361</td>\n<td>0.317</td>\n</tr>\n</tbody>\n</table>\n<h2>Model details</h2>\n<ul>\n<li>Backbones: Ensemble of ResNeSt50, ResNet101, ResNet152 and SEResNeXt101</li>\n<li>Pooling: GeM (p=3)</li>\n<li>Head: FC-&gt;BN-&gt;L2 (the same as the last year’s first place team)</li>\n<li>Loss: CosFace with Label Smoothing (ArcFace was also good, but CosFace was better)</li>\n<li>Data Augmentation: RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask</li>\n<li>LR: Cosine Annealing LR with warmup, training for 30 epochs</li>\n<li>Input image size in training: 352</li>\n</ul>\n<h2>What we tried and worked</h2>\n<ul>\n<li>Automatic mixed precision training</li>\n<li>Replace GeM p=3 with p=4 in testing</li>\n<li>Increase input image size last few epochs of training with freezed BN</li>\n<li>Large and multiple input image sizes in testing: 416, 480 and 608</li>\n<li>MVArcFace</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>PCA whitening</li>\n<li>Maintaining the aspect ratio of input images in testing (perhaps because our models were trained on square images)</li>\n<li>Combination of arcface loss and pairwise (e.g., triplet) loss<br>\nWe have tried arcface loss with pairwise loss (specially multi-similarity loss). However, the single arcface loss was better than multiple losses.</li>\n<li>EfficientNet</li>\n<li>Circle loss</li>\n<li>Removing noisy classes<br>\nWe have tried to remove the worst 3 noisy classes which have high variance based on arcface class weight, but it did not work.</li>\n</ul>\n<h2>What we have not tried</h2>\n<ul>\n<li>Training with GLDv1 dataset</li>\n<li>Training without changing the aspect ratio of images </li>\n</ul>",
      "rawMarkdown": "We thank all organizers for this very exciting competition.\nCongratulations to all who finished the competition and to the winners.\n\nWe trained models using only GLDv2clean dataset on PyTorch and converted them to TensorFlow’s saved_model.\nThen, we validated them using GLR2019 public and private datasets.\n\n## Final submission\n\nResNeSt50 + ResNet101 + ResNet152 + SEResNeXt101\n\n|model                    | output dim  | input size | GLR2019 mAP@100 | Public LB | Private LB |\n|:------------------------|:------------|:-----------|:--------|:----------|:-----------|\n|ResNeSt50                |512          |416         | 0.3114  | 0.335     | 0.292      |\n|ResNet101                |512          |608         | 0.3243  | 0.346     | 0.303      |\n|ResNet152                |512          |480         | 0.3180  | 0.336     | 0.288      |\n|SEResNeXt101             |512          |480         | 0.3209  | 0.337     | 0.296      |\n|Ensemble of 4 models     |2048 (concat)|-           | 0.3396  | 0.361     | 0.317      |\n\n\n## Model details\n- Backbones: Ensemble of ResNeSt50, ResNet101, ResNet152 and SEResNeXt101\n- Pooling: GeM (p=3)\n- Head: FC->BN->L2 (the same as the last year’s first place team)\n- Loss: CosFace with Label Smoothing (ArcFace was also good, but CosFace was better)\n- Data Augmentation: RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask\n- LR: Cosine Annealing LR with warmup, training for 30 epochs\n- Input image size in training: 352\n\n## What we tried and worked\n- Automatic mixed precision training\n- Replace GeM p=3 with p=4 in testing\n- Increase input image size last few epochs of training with freezed BN\n- Large and multiple input image sizes in testing: 416, 480 and 608\n- MVArcFace\n\n## What did not work\n- PCA whitening\n- Maintaining the aspect ratio of input images in testing (perhaps because our models were trained on square images)\n- Combination of arcface loss and pairwise (e.g., triplet) loss\n  We have tried arcface loss with pairwise loss (specially multi-similarity loss). However, the single arcface loss was better than multiple losses.\n- EfficientNet\n- Circle loss\n- Removing noisy classes\n  We have tried to remove the worst 3 noisy classes which have high variance based on arcface class weight, but it did not work.\n\n## What we have not tried\n- Training with GLDv1 dataset\n- Training without changing the aspect ratio of images ",
      "votes": 39
    },
    {
      "id": 975814,
      "postDate": "2020-08-18T13:30:47.160Z",
      "content": "<p>Great solution! Have you tried learning the p paramter of GeM on-the-fly?</p>",
      "rawMarkdown": "Great solution! Have you tried learning the p paramter of GeM on-the-fly?",
      "votes": 4,
      "replies": [
        {
          "id": 976561,
          "postDate": "2020-08-19T00:28:12.027Z",
          "content": "<p>I also tried learning the p parameter of GeM, but it didn't work.</p>",
          "rawMarkdown": "I also tried learning the p parameter of GeM, but it didn't work.",
          "votes": 4
        },
        {
          "id": 976585,
          "postDate": "2020-08-19T00:58:42.770Z",
          "content": "<p>Interesting. I tried that and p converged to around 2.93, which is how I got my current score.</p>",
          "rawMarkdown": "Interesting. I tried that and p converged to around 2.93, which is how I got my current score.",
          "votes": 2
        },
        {
          "id": 977083,
          "postDate": "2020-08-19T09:42:59.523Z",
          "content": "<p>An interesting experiment I can think of is to init. p as a parameter vector with values=3 and dimensions equal to num. channels from last feature map, then train the network. In this challenge, I didn't have enough time to run it 😄</p>",
          "rawMarkdown": "An interesting experiment I can think of is to init. p as a parameter vector with values=3 and dimensions equal to num. channels from last feature map, then train the network. In this challenge, I didn't have enough time to run it 😄",
          "votes": 2
        }
      ]
    },
    {
      "id": 982968,
      "postDate": "2020-08-23T22:32:14.747Z",
      "content": "<p>Congratulations 👏 for your gold.  Can you please explain or maybe list some resources for implementing pca whitening. Preferably torch implementation but tensorflow will work as well. Thanks.</p>",
      "rawMarkdown": "Congratulations 👏 for your gold.  Can you please explain or maybe list some resources for implementing pca whitening. Preferably torch implementation but tensorflow will work as well. Thanks.",
      "votes": 1
    },
    {
      "id": 976256,
      "postDate": "2020-08-18T18:42:50.673Z",
      "content": "<p>Thanks for sharing, and congrats for 9th place!</p>\n<p>May I ask you about your training hardware, and batch sizes you used? Also, you mentioned that \"Maintaining the aspect ratio of input images in testing\" did not worked, opposite to the baseline example. What is your intuition on that? Will training without aspect ratio changed improve your end result?</p>\n<p>Your findings on <code>p=4</code> is really interesting - when I read the DELG paper, my biggest question was \"why p=3 and not larger?\", since, in my intuition, using Lp-norm is a way to combat the weird effects of L2 in higher dimensions (see \"hypersphere packing problem\"). What is your intuition on that?</p>",
      "rawMarkdown": "Thanks for sharing, and congrats for 9th place!\n\nMay I ask you about your training hardware, and batch sizes you used? Also, you mentioned that \"Maintaining the aspect ratio of input images in testing\" did not worked, opposite to the baseline example. What is your intuition on that? Will training without aspect ratio changed improve your end result?\n\nYour findings on `p=4` is really interesting - when I read the DELG paper, my biggest question was \"why p=3 and not larger?\", since, in my intuition, using Lp-norm is a way to combat the weird effects of L2 in higher dimensions (see \"hypersphere packing problem\"). What is your intuition on that?",
      "votes": 1
    },
    {
      "id": 975584,
      "postDate": "2020-08-18T11:16:26.087Z",
      "content": "<p>Great..can you give more details of data strategy? how you sampled anchor positive and negative samples? whats your hardware to try those many models? that too for 30 epochs each..</p>",
      "rawMarkdown": "Great..can you give more details of data strategy? how you sampled anchor positive and negative samples? whats your hardware to try those many models? that too for 30 epochs each..",
      "votes": 2
    },
    {
      "id": 975418,
      "postDate": "2020-08-18T09:38:43.773Z",
      "content": "<p>Congrats!<br>\nI am interest with 'Replace GeM p=3 with p=4 in testing' . How much is the improvement.<br>\nIs it leading to overfit the public? Thanks!</p>",
      "rawMarkdown": "Congrats!\nI am interest with 'Replace GeM p=3 with p=4 in testing' . How much is the improvement.\nIs it leading to overfit the public? Thanks!",
      "votes": 2,
      "replies": [
        {
          "id": 975588,
          "postDate": "2020-08-18T11:17:32.077Z",
          "content": "<p>Thanks.<br>\nChanging the GeM p in testing doesn't seem to work in all cases, but I changed the GeM p to account for the GLR2019 validation score as well.<br>\nAs long as I checked some results, it seems to be effective including Private LB, as shown in the example below.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>input_size</th>\n<th>GLR2019  mAP@100</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet101 (GeMp=3)</td>\n<td>576</td>\n<td>0.3201</td>\n<td>0.344</td>\n<td>0.298</td>\n</tr>\n<tr>\n<td>ResNet101 (GeMp=4)</td>\n<td>576</td>\n<td>0.3209</td>\n<td>0.347</td>\n<td>0.300</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Thanks.\nChanging the GeM p in testing doesn't seem to work in all cases, but I changed the GeM p to account for the GLR2019 validation score as well.\nAs long as I checked some results, it seems to be effective including Private LB, as shown in the example below.\n\n|model                    |input_size  | GLR2019  mAP@100 | Public LB | Private LB |\n|:------------------------|:-----------|:--------|:----------|:-----------|\n|ResNet101 (GeMp=3)       | 576        | 0.3201  | 0.344     | 0.298      |\n|ResNet101 (GeMp=4)       | 576        | 0.3209  | 0.347     | 0.300      |\n",
          "votes": 4
        },
        {
          "id": 976566,
          "postDate": "2020-08-19T00:29:46.867Z",
          "content": "<p>Thanks! well done!</p>",
          "rawMarkdown": "Thanks! well done!"
        }
      ]
    },
    {
      "id": 977901,
      "postDate": "2020-08-19T19:36:12.377Z",
      "content": "<p>Good job! Do you know roughly how much improvement you get from automatic mixed-precision training?</p>",
      "rawMarkdown": "Good job! Do you know roughly how much improvement you get from automatic mixed-precision training?",
      "replies": [
        {
          "id": 978436,
          "postDate": "2020-08-20T07:14:42.137Z",
          "content": "<p>Amp (Automatic Mixed Precision) was effective in that it allowed for faster training and more efficient experimentation in a limited amount of time.<br>\nAlthough unconfirmed, we believe that a model without amp may have better accuracy than a model with amp.<br>\nThe final model we submitted was trained without amp.</p>",
          "rawMarkdown": "Amp (Automatic Mixed Precision) was effective in that it allowed for faster training and more efficient experimentation in a limited amount of time.\nAlthough unconfirmed, we believe that a model without amp may have better accuracy than a model with amp.\nThe final model we submitted was trained without amp.",
          "votes": 2
        }
      ]
    },
    {
      "id": 976159,
      "postDate": "2020-08-18T17:15:47.383Z",
      "content": "<p>Congrats, good job</p>",
      "rawMarkdown": "Congrats, good job"
    },
    {
      "id": 975736,
      "postDate": "2020-08-18T12:52:05.100Z",
      "content": "<p>Great job! can you share the code?</p>",
      "rawMarkdown": "Great job! can you share the code?"
    },
    {
      "id": 3253625,
      "postDate": "2025-07-25T02:30:03.150Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 977019,
      "postDate": "2020-08-19T08:29:51.623Z",
      "content": "<p>Thanks for sharing and congrats! </p>",
      "rawMarkdown": "Thanks for sharing and congrats! "
    },
    {
      "id": 975862,
      "postDate": "2020-08-18T13:54:27.270Z",
      "content": "<p>thanks for your sharing!</p>",
      "rawMarkdown": "thanks for your sharing!"
    }
  ],
  "comments": [
    {
      "id": 975814,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2020-08-18T13:30:47.160000",
      "content": "<p>Great solution! Have you tried learning the p paramter of GeM on-the-fly?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 976561,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2020-08-19T00:28:12.027000",
          "content": "<p>I also tried learning the p parameter of GeM, but it didn't work.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 976585,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2020-08-19T00:58:42.770000",
          "content": "<p>Interesting. I tried that and p converged to around 2.93, which is how I got my current score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 977083,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2020-08-19T09:42:59.523000",
          "content": "<p>An interesting experiment I can think of is to init. p as a parameter vector with values=3 and dimensions equal to num. channels from last feature map, then train the network. In this challenge, I didn't have enough time to run it 😄</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 982968,
      "author_name": "torch",
      "author_url": "",
      "post_date": "2020-08-23T22:32:14.747000",
      "content": "<p>Congratulations 👏 for your gold.  Can you please explain or maybe list some resources for implementing pca whitening. Preferably torch implementation but tensorflow will work as well. Thanks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976256,
      "author_name": "Chan Kha Vu",
      "author_url": "",
      "post_date": "2020-08-18T18:42:50.673000",
      "content": "<p>Thanks for sharing, and congrats for 9th place!</p>\n<p>May I ask you about your training hardware, and batch sizes you used? Also, you mentioned that \"Maintaining the aspect ratio of input images in testing\" did not worked, opposite to the baseline example. What is your intuition on that? Will training without aspect ratio changed improve your end result?</p>\n<p>Your findings on <code>p=4</code> is really interesting - when I read the DELG paper, my biggest question was \"why p=3 and not larger?\", since, in my intuition, using Lp-norm is a way to combat the weird effects of L2 in higher dimensions (see \"hypersphere packing problem\"). What is your intuition on that?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975584,
      "author_name": "Uday Kumar Gurugubelli",
      "author_url": "",
      "post_date": "2020-08-18T11:16:26.087000",
      "content": "<p>Great..can you give more details of data strategy? how you sampled anchor positive and negative samples? whats your hardware to try those many models? that too for 30 epochs each..</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 975418,
      "author_name": "Octo",
      "author_url": "",
      "post_date": "2020-08-18T09:38:43.773000",
      "content": "<p>Congrats!<br>\nI am interest with 'Replace GeM p=3 with p=4 in testing' . How much is the improvement.<br>\nIs it leading to overfit the public? Thanks!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 975588,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2020-08-18T11:17:32.077000",
          "content": "<p>Thanks.<br>\nChanging the GeM p in testing doesn't seem to work in all cases, but I changed the GeM p to account for the GLR2019 validation score as well.<br>\nAs long as I checked some results, it seems to be effective including Private LB, as shown in the example below.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>input_size</th>\n<th>GLR2019  mAP@100</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet101 (GeMp=3)</td>\n<td>576</td>\n<td>0.3201</td>\n<td>0.344</td>\n<td>0.298</td>\n</tr>\n<tr>\n<td>ResNet101 (GeMp=4)</td>\n<td>576</td>\n<td>0.3209</td>\n<td>0.347</td>\n<td>0.300</td>\n</tr>\n</tbody>\n</table>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 976566,
          "author_name": "Octo",
          "author_url": "",
          "post_date": "2020-08-19T00:29:46.867000",
          "content": "<p>Thanks! well done!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 977901,
      "author_name": "Jeremy Ma",
      "author_url": "",
      "post_date": "2020-08-19T19:36:12.377000",
      "content": "<p>Good job! Do you know roughly how much improvement you get from automatic mixed-precision training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 978436,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2020-08-20T07:14:42.137000",
          "content": "<p>Amp (Automatic Mixed Precision) was effective in that it allowed for faster training and more efficient experimentation in a limited amount of time.<br>\nAlthough unconfirmed, we believe that a model without amp may have better accuracy than a model with amp.<br>\nThe final model we submitted was trained without amp.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 976159,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2020-08-18T17:15:47.383000",
      "content": "<p>Congrats, good job</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975736,
      "author_name": "Shivam",
      "author_url": "",
      "post_date": "2020-08-18T12:52:05.100000",
      "content": "<p>Great job! can you share the code?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3253625,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-25T02:30:03.150000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 977019,
      "author_name": "Abishek Sudarshan",
      "author_url": "",
      "post_date": "2020-08-19T08:29:51.623000",
      "content": "<p>Thanks for sharing and congrats! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975862,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2020-08-18T13:54:27.270000",
      "content": "<p>thanks for your sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "975396": "We thank all organizers for this very exciting competition.\nCongratulations to all who finished the competition and to the winners.\n\nWe trained models using only GLDv2clean dataset on PyTorch and converted them to TensorFlow’s saved_model.\nThen, we validated them using GLR2019 public and private datasets.\n\n## Final submission\n\nResNeSt50 + ResNet101 + ResNet152 + SEResNeXt101\n\n|model                    | output dim  | input size | GLR2019 mAP@100 | Public LB | Private LB |\n|:------------------------|:------------|:-----------|:--------|:----------|:-----------|\n|ResNeSt50                |512          |416         | 0.3114  | 0.335     | 0.292      |\n|ResNet101                |512          |608         | 0.3243  | 0.346     | 0.303      |\n|ResNet152                |512          |480         | 0.3180  | 0.336     | 0.288      |\n|SEResNeXt101             |512          |480         | 0.3209  | 0.337     | 0.296      |\n|Ensemble of 4 models     |2048 (concat)|-           | 0.3396  | 0.361     | 0.317      |\n\n\n## Model details\n- Backbones: Ensemble of ResNeSt50, ResNet101, ResNet152 and SEResNeXt101\n- Pooling: GeM (p=3)\n- Head: FC->BN->L2 (the same as the last year’s first place team)\n- Loss: CosFace with Label Smoothing (ArcFace was also good, but CosFace was better)\n- Data Augmentation: RandomResizedCrop, Rotation, RandomGrayScale, ColorJitter, GaussianNoise, Normalize, and GridMask\n- LR: Cosine Annealing LR with warmup, training for 30 epochs\n- Input image size in training: 352\n\n## What we tried and worked\n- Automatic mixed precision training\n- Replace GeM p=3 with p=4 in testing\n- Increase input image size last few epochs of training with freezed BN\n- Large and multiple input image sizes in testing: 416, 480 and 608\n- MVArcFace\n\n## What did not work\n- PCA whitening\n- Maintaining the aspect ratio of input images in testing (perhaps because our models were trained on square images)\n- Combination of arcface loss and pairwise (e.g., triplet) loss\n  We have tried arcface loss with pairwise loss (specially multi-similarity loss). However, the single arcface loss was better than multiple losses.\n- EfficientNet\n- Circle loss\n- Removing noisy classes\n  We have tried to remove the worst 3 noisy classes which have high variance based on arcface class weight, but it did not work.\n\n## What we have not tried\n- Training with GLDv1 dataset\n- Training without changing the aspect ratio of images ",
    "975814": "Great solution! Have you tried learning the p paramter of GeM on-the-fly?",
    "982968": "Congratulations 👏 for your gold.  Can you please explain or maybe list some resources for implementing pca whitening. Preferably torch implementation but tensorflow will work as well. Thanks.",
    "976256": "Thanks for sharing, and congrats for 9th place!\n\nMay I ask you about your training hardware, and batch sizes you used? Also, you mentioned that \"Maintaining the aspect ratio of input images in testing\" did not worked, opposite to the baseline example. What is your intuition on that? Will training without aspect ratio changed improve your end result?\n\nYour findings on `p=4` is really interesting - when I read the DELG paper, my biggest question was \"why p=3 and not larger?\", since, in my intuition, using Lp-norm is a way to combat the weird effects of L2 in higher dimensions (see \"hypersphere packing problem\"). What is your intuition on that?",
    "975584": "Great..can you give more details of data strategy? how you sampled anchor positive and negative samples? whats your hardware to try those many models? that too for 30 epochs each..",
    "975418": "Congrats!\nI am interest with 'Replace GeM p=3 with p=4 in testing' . How much is the improvement.\nIs it leading to overfit the public? Thanks!",
    "977901": "Good job! Do you know roughly how much improvement you get from automatic mixed-precision training?",
    "976159": "Congrats, good job",
    "975736": "Great job! can you share the code?",
    "3253625": "",
    "977019": "Thanks for sharing and congrats! ",
    "975862": "thanks for your sharing!"
  }
}