{
  "id": 187757,
  "title": "3rd Place Solution — A Pure Global Feature Approach",
  "url": "/competitions/landmark-recognition-2020/writeups/all-data-are-ext-3rd-place-solution-a-pure-global-",
  "author_name": "",
  "post_date": "2020-10-14T04:00:14.940Z",
  "votes": 91,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Thanks to the organizers and congrats to all the winners! Our ( <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a>) solution is a pure global feature metric learning approach with some tricks.</p>\n<h2>architecture: sub-center ArcFace with dynamic margins</h2>\n<p>ArcFace (<a href=\"https://arxiv.org/abs/1801.07698\" target=\"_blank\">paper</a>) has become a standard metric learning method on Kaggle over the past two years or so. In this competition, we used Sub-center ArcFace (<a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\" target=\"_blank\">paper</a>), a recent improvement over ArcFace by the same authors. The idea is that each class may have more than one class center. For example, a certain landmark's photos may have a few clusters (e.g. from different angles). Sub-center ArcFace's weights store multiple class centers' representations, which can increase classification accuracy and improve global feature's quality.</p>\n<p>The classes in GLD dataset are extremely imbalanced with longs tails. For models to converge better in the presence of heavy imbalance, smaller classes need to have bigger margins as they are harder to learn. Instead of manually setting different margin levels based on class size, we introduce <em>dynamic margins</em>, a family of continuous functions mapping class size to margin level. This give us some major boost. Details to be described in paper.</p>\n<p>We train the model using ArcFace loss only.</p>\n<h2>validation scheme</h2>\n<p>Stratified 5-fold. Training using 4 folds, validate on only 1/15 of 1 fold to save time. Use different folds for different single models, for maximal diversity in ensemble.</p>\n<p>We use GAP metric to validate. CV GAP is very high compared to LB. All model's CV GAP are over 0.967. Good news is that CV GAP and LB have high correlation.</p>\n<h2>test set predicting strategy</h2>\n<ul>\n<li>If we use model's ArcFace head to predict test set, our best single fold model's public LB is only <strong>0.564</strong></li>\n<li>A better strategy is to calculate the global feature cosine similarity of each [private train image, private test image] pair, and use the top1 neighbor and corresponding cosine similarity of each test image as prediction. This gives us <strong>0.604</strong> public LB for the same single fold model</li>\n<li>Instead of using top1 neighbor only, we can improve it by combining top5 neighbors and their similarities. Best combining function is 8th power. <br>\nE.g. let's say test image A's top5 neighbors and their cosine similarities are: class 1 (0.9), class 2 (0.8), class 2 (0.7), class 1 (0.5), class 3 (0.45). Then class 1's total score is <code>0.9**8 + 0.5**8 = 0.434</code>; class 2's total score is <code>0.8**8 + 0.7**8 = 0.225</code>. So we predict class 1 with p=0.434<br>\nLB increases to <strong>0.610</strong></li>\n<li>We can further improve it by incorporating ArcFace head's predictions with 12th power. In above example, assuming image A's ArcFace head give class 1 score = 0.75, class 2 score = 0.88. Then class 1's total score becomes <code>(0.9**8 + 0.5**8) * 0.75**12 = 0.0138</code>; class 2's total score becomes <code>(0.8**8 + 0.7**8) * 0.88**12 = 0.0486</code>. Now we predict class 2 with p=0.0486<br>\nLB increases to <strong>0.618</strong></li>\n</ul>\n<h2>augmentations</h2>\n<p>We resize all images to square shape without cropping.</p>\n<pre><code>import albumentations as A\nA.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),    \n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n        A.Normalize()\n    ])\n</code></pre>\n<h2>pretraining and finetuning</h2>\n<ul>\n<li>In cGLDv2 (cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.</li>\n<li>In GLDv2, there are 4.1m training images and 200k classes, among which 3.2m images belong to the 81k classes in cGLDv2.</li>\n</ul>\n<p>We noticed that (1) training with the 3.2m data gives better results than only the 1.6m competition cGLDv2 data, (2) pretraining on all 4.1m data then finetuning on 3.2m data gives even better results.</p>\n<h2>3-stage training schedule</h2>\n<ul>\n<li>Stage 1 (pretrain): 10 epochs with small image size (256) on 4.1m data</li>\n<li>Stage 2 (finetune): about 16 epochs with medium image size (512 to 768 depending on model) on 3.2m data. Number of epochs varies by model and ranges from 13 to 21. </li>\n<li>Stage 3 (finetune): 1 epoch with large image size (672 to 1024) on 3.2m data</li>\n</ul>\n<p>Note: above schedules are for Sub1 (private 0.6289). Sub2 has higher score (private 0.6344) but more complex schedules, i.e. longer with more rounds of finetuning. Both submissions are 3rd place.</p>\n<h2>ensemble</h2>\n<p>7 models: EfficientNet B7, B6, B5, B4, B3, <a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">ResNeSt101</a>, <a href=\"https://github.com/clovaai/rexnet\" target=\"_blank\">ReXNet2.0</a></p>\n<p>For global feature neighbor search, we concatenate each model's 512-dimension feature; for ArcFace head, we take simple average of each model's logits.</p>\n<p>Best single model is EfficientB6. <strong>private = 0.6005, public = 0.6179</strong> (We didn't submit all the epochs, there may be higher ones)<br>\nSub1 is 7 model ensemble. <strong>private = 0.6289, public = 0.6604</strong><br>\nSub2 is 9 model ensemble (B5 and B6 twice). <strong>private = 0.6344, public = 0.6581</strong></p>\n<h3>[update 10/12/2020]</h3>\n<p>paper: <a href=\"https://arxiv.org/abs/2010.05350\" target=\"_blank\">https://arxiv.org/abs/2010.05350</a><br>\nrepo: <a href=\"https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution\" target=\"_blank\">https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution</a></p>",
  "messages": [
    {
      "id": "1032394",
      "postDate": "09/30/2020 07:13:45",
      "content": "<p>Thanks to the organizers and congrats to all the winners! Our ( <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a>) solution is a pure global feature metric learning approach with some tricks.</p>\n<h2>architecture: sub-center ArcFace with dynamic margins</h2>\n<p>ArcFace (<a href=\"https://arxiv.org/abs/1801.07698\" target=\"_blank\">paper</a>) has become a standard metric learning method on Kaggle over the past two years or so. In this competition, we used Sub-center ArcFace (<a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\" target=\"_blank\">paper</a>), a recent improvement over ArcFace by the same authors. The idea is that each class may have more than one class center. For example, a certain landmark's photos may have a few clusters (e.g. from different angles). Sub-center ArcFace's weights store multiple class centers' representations, which can increase classification accuracy and improve global feature's quality.</p>\n<p>The classes in GLD dataset are extremely imbalanced with longs tails. For models to converge better in the presence of heavy imbalance, smaller classes need to have bigger margins as they are harder to learn. Instead of manually setting different margin levels based on class size, we introduce <em>dynamic margins</em>, a family of continuous functions mapping class size to margin level. This give us some major boost. Details to be described in paper.</p>\n<p>We train the model using ArcFace loss only.</p>\n<h2>validation scheme</h2>\n<p>Stratified 5-fold. Training using 4 folds, validate on only 1/15 of 1 fold to save time. Use different folds for different single models, for maximal diversity in ensemble.</p>\n<p>We use GAP metric to validate. CV GAP is very high compared to LB. All model's CV GAP are over 0.967. Good news is that CV GAP and LB have high correlation.</p>\n<h2>test set predicting strategy</h2>\n<ul>\n<li>If we use model's ArcFace head to predict test set, our best single fold model's public LB is only <strong>0.564</strong></li>\n<li>A better strategy is to calculate the global feature cosine similarity of each [private train image, private test image] pair, and use the top1 neighbor and corresponding cosine similarity of each test image as prediction. This gives us <strong>0.604</strong> public LB for the same single fold model</li>\n<li>Instead of using top1 neighbor only, we can improve it by combining top5 neighbors and their similarities. Best combining function is 8th power. <br>\nE.g. let's say test image A's top5 neighbors and their cosine similarities are: class 1 (0.9), class 2 (0.8), class 2 (0.7), class 1 (0.5), class 3 (0.45). Then class 1's total score is <code>0.9**8 + 0.5**8 = 0.434</code>; class 2's total score is <code>0.8**8 + 0.7**8 = 0.225</code>. So we predict class 1 with p=0.434<br>\nLB increases to <strong>0.610</strong></li>\n<li>We can further improve it by incorporating ArcFace head's predictions with 12th power. In above example, assuming image A's ArcFace head give class 1 score = 0.75, class 2 score = 0.88. Then class 1's total score becomes <code>(0.9**8 + 0.5**8) * 0.75**12 = 0.0138</code>; class 2's total score becomes <code>(0.8**8 + 0.7**8) * 0.88**12 = 0.0486</code>. Now we predict class 2 with p=0.0486<br>\nLB increases to <strong>0.618</strong></li>\n</ul>\n<h2>augmentations</h2>\n<p>We resize all images to square shape without cropping.</p>\n<pre><code>import albumentations as A\nA.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),    \n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n        A.Normalize()\n    ])\n</code></pre>\n<h2>pretraining and finetuning</h2>\n<ul>\n<li>In cGLDv2 (cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.</li>\n<li>In GLDv2, there are 4.1m training images and 200k classes, among which 3.2m images belong to the 81k classes in cGLDv2.</li>\n</ul>\n<p>We noticed that (1) training with the 3.2m data gives better results than only the 1.6m competition cGLDv2 data, (2) pretraining on all 4.1m data then finetuning on 3.2m data gives even better results.</p>\n<h2>3-stage training schedule</h2>\n<ul>\n<li>Stage 1 (pretrain): 10 epochs with small image size (256) on 4.1m data</li>\n<li>Stage 2 (finetune): about 16 epochs with medium image size (512 to 768 depending on model) on 3.2m data. Number of epochs varies by model and ranges from 13 to 21. </li>\n<li>Stage 3 (finetune): 1 epoch with large image size (672 to 1024) on 3.2m data</li>\n</ul>\n<p>Note: above schedules are for Sub1 (private 0.6289). Sub2 has higher score (private 0.6344) but more complex schedules, i.e. longer with more rounds of finetuning. Both submissions are 3rd place.</p>\n<h2>ensemble</h2>\n<p>7 models: EfficientNet B7, B6, B5, B4, B3, <a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">ResNeSt101</a>, <a href=\"https://github.com/clovaai/rexnet\" target=\"_blank\">ReXNet2.0</a></p>\n<p>For global feature neighbor search, we concatenate each model's 512-dimension feature; for ArcFace head, we take simple average of each model's logits.</p>\n<p>Best single model is EfficientB6. <strong>private = 0.6005, public = 0.6179</strong> (We didn't submit all the epochs, there may be higher ones)<br>\nSub1 is 7 model ensemble. <strong>private = 0.6289, public = 0.6604</strong><br>\nSub2 is 9 model ensemble (B5 and B6 twice). <strong>private = 0.6344, public = 0.6581</strong></p>\n<h3>[update 10/12/2020]</h3>\n<p>paper: <a href=\"https://arxiv.org/abs/2010.05350\" target=\"_blank\">https://arxiv.org/abs/2010.05350</a><br>\nrepo: <a href=\"https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution\" target=\"_blank\">https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution</a></p>",
      "rawMarkdown": "Thanks to the organizers and congrats to all the winners! Our ( @haqishen @boliu0 @garybios @alexanderliao) solution is a pure global feature metric learning approach with some tricks.\n\n## architecture: sub-center ArcFace with dynamic margins\nArcFace ([paper](https://arxiv.org/abs/1801.07698)) has become a standard metric learning method on Kaggle over the past two years or so. In this competition, we used Sub-center ArcFace ([paper](https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf)), a recent improvement over ArcFace by the same authors. The idea is that each class may have more than one class center. For example, a certain landmark's photos may have a few clusters (e.g. from different angles). Sub-center ArcFace's weights store multiple class centers' representations, which can increase classification accuracy and improve global feature's quality.\n\nThe classes in GLD dataset are extremely imbalanced with longs tails. For models to converge better in the presence of heavy imbalance, smaller classes need to have bigger margins as they are harder to learn. Instead of manually setting different margin levels based on class size, we introduce *dynamic margins*, a family of continuous functions mapping class size to margin level. This give us some major boost. Details to be described in paper.\n\nWe train the model using ArcFace loss only.\n\n## validation scheme\nStratified 5-fold. Training using 4 folds, validate on only 1/15 of 1 fold to save time. Use different folds for different single models, for maximal diversity in ensemble.\n\nWe use GAP metric to validate. CV GAP is very high compared to LB. All model's CV GAP are over 0.967. Good news is that CV GAP and LB have high correlation.\n\n## test set predicting strategy\n- If we use model's ArcFace head to predict test set, our best single fold model's public LB is only **0.564**\n- A better strategy is to calculate the global feature cosine similarity of each [private train image, private test image] pair, and use the top1 neighbor and corresponding cosine similarity of each test image as prediction. This gives us **0.604** public LB for the same single fold model\n- Instead of using top1 neighbor only, we can improve it by combining top5 neighbors and their similarities. Best combining function is 8th power. \nE.g. let's say test image A's top5 neighbors and their cosine similarities are: class 1 (0.9), class 2 (0.8), class 2 (0.7), class 1 (0.5), class 3 (0.45). Then class 1's total score is `0.9**8 + 0.5**8 = 0.434`; class 2's total score is `0.8**8 + 0.7**8 = 0.225`. So we predict class 1 with p=0.434\nLB increases to **0.610**\n- We can further improve it by incorporating ArcFace head's predictions with 12th power. In above example, assuming image A's ArcFace head give class 1 score = 0.75, class 2 score = 0.88. Then class 1's total score becomes `(0.9**8 + 0.5**8) * 0.75**12 = 0.0138`; class 2's total score becomes `(0.8**8 + 0.7**8) * 0.88**12 = 0.0486`. Now we predict class 2 with p=0.0486\nLB increases to **0.618**\n\n\n## augmentations\nWe resize all images to square shape without cropping.\n```\nimport albumentations as A\nA.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),    \n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n        A.Normalize()\n    ])\n```\n\n## pretraining and finetuning\n- In cGLDv2 (cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.\n- In GLDv2, there are 4.1m training images and 200k classes, among which 3.2m images belong to the 81k classes in cGLDv2.\n\nWe noticed that (1) training with the 3.2m data gives better results than only the 1.6m competition cGLDv2 data, (2) pretraining on all 4.1m data then finetuning on 3.2m data gives even better results.\n\n## 3-stage training schedule\n- Stage 1 (pretrain): 10 epochs with small image size (256) on 4.1m data\n- Stage 2 (finetune): about 16 epochs with medium image size (512 to 768 depending on model) on 3.2m data. Number of epochs varies by model and ranges from 13 to 21. \n- Stage 3 (finetune): 1 epoch with large image size (672 to 1024) on 3.2m data\n\nNote: above schedules are for Sub1 (private 0.6289). Sub2 has higher score (private 0.6344) but more complex schedules, i.e. longer with more rounds of finetuning. Both submissions are 3rd place.\n\n## ensemble\n7 models: EfficientNet B7, B6, B5, B4, B3, [ResNeSt101](https://github.com/zhanghang1989/ResNeSt), [ReXNet2.0](https://github.com/clovaai/rexnet)\n\nFor global feature neighbor search, we concatenate each model's 512-dimension feature; for ArcFace head, we take simple average of each model's logits.\n\nBest single model is EfficientB6. **private = 0.6005, public = 0.6179** (We didn't submit all the epochs, there may be higher ones)\nSub1 is 7 model ensemble. **private = 0.6289, public = 0.6604**\nSub2 is 9 model ensemble (B5 and B6 twice). **private = 0.6344, public = 0.6581**\n\n\n### [update 10/12/2020]\npaper: https://arxiv.org/abs/2010.05350\nrepo: https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution",
      "votes": null
    },
    {
      "id": "1032503",
      "postDate": "09/30/2020 08:51:13",
      "content": "<p>As an addition to <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> 's wonderful post, we will also share techniques that we tried but failed to work. This includes classical approaches in retrieval such as DBA, QE; distractor removal using object detection (OID-pretrained), and training DELF from scratch with our own PyTorch implementation.</p>",
      "rawMarkdown": "As an addition to @boliu0 's wonderful post, we will also share techniques that we tried but failed to work. This includes classical approaches in retrieval such as DBA, QE; distractor removal using object detection (OID-pretrained), and training DELF from scratch with our own PyTorch implementation.",
      "votes": null
    },
    {
      "id": "1032529",
      "postDate": "09/30/2020 09:11:09",
      "content": "<p>Congrats and thx <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and the rest of the team for the write-up. Sub-center ArcFace with dynamic margins definitely sounds interesting, will have a look :)</p>",
      "rawMarkdown": "Congrats and thx @boliu0 and the rest of the team for the write-up. Sub-center ArcFace with dynamic margins definitely sounds interesting, will have a look :)",
      "votes": null
    },
    {
      "id": "1032531",
      "postDate": "09/30/2020 09:12:58",
      "content": "<p>Wow, <strong>congratulations!</strong> 🎉🎉🎉 It's crazy that you managed to achieve such a high place with just global features.</p>\n<p>I think a teammate of yours mentioned about the dynamic margin idea (different margin for different classes) somewhere in the comments to the Retrieval competition. Glad that it played out well! Can't wait to read your paper!</p>\n<p>A few questions about it:</p>\n<ul>\n<li>Are the ArcFace's margin paramaters learnable or changing over time? If so, you only learn/update the parameters of the continuous function <code>size -&gt; margin</code>, right?</li>\n<li>Why did you make Dynamic Margins to be a function of class sizes, not of class indices?</li>\n</ul>",
      "rawMarkdown": "Wow, **congratulations!** 🎉🎉🎉 It's crazy that you managed to achieve such a high place with just global features.\n\nI think a teammate of yours mentioned about the dynamic margin idea (different margin for different classes) somewhere in the comments to the Retrieval competition. Glad that it played out well! Can't wait to read your paper!\n\nA few questions about it:\n- Are the ArcFace's margin paramaters learnable or changing over time? If so, you only learn/update the parameters of the continuous function `size -> margin`, right?\n- Why did you make Dynamic Margins to be a function of class sizes, not of class indices?",
      "votes": null
    },
    {
      "id": "1032579",
      "postDate": "09/30/2020 09:53:36",
      "content": "<p>Wow, what an elegant solution. Congrats. May I ask what hardware you used for training models?</p>",
      "rawMarkdown": "Wow, what an elegant solution. Congrats. May I ask what hardware you used for training models?",
      "votes": null
    },
    {
      "id": "1032612",
      "postDate": "09/30/2020 10:12:22",
      "content": "<p>I believe you can even get good models with just Colab Pro + a lot of time in this competition.</p>",
      "rawMarkdown": "I believe you can even get good models with just Colab Pro + a lot of time in this competition.",
      "votes": null
    },
    {
      "id": "1032617",
      "postDate": "09/30/2020 10:15:00",
      "content": "<p>Congratzz on your 3rd place. Nice write up and really helpful. </p>\n<blockquote>\n  <p>If we use model's ArcFace head to predict test set, our best single fold model's public LB is only 0.564</p>\n</blockquote>\n<p>How are you predicting with Arcface head on the test set? I mean, arcface takes the label as input and so don't we need labels to predict? Are you just passing a 81313 dimensional 0-vector as label?</p>",
      "rawMarkdown": "Congratzz on your 3rd place. Nice write up and really helpful. \n\n> If we use model's ArcFace head to predict test set, our best single fold model's public LB is only 0.564\n\nHow are you predicting with Arcface head on the test set? I mean, arcface takes the label as input and so don't we need labels to predict? Are you just passing a 81313 dimensional 0-vector as label?",
      "votes": null
    },
    {
      "id": "1032632",
      "postDate": "09/30/2020 10:23:22",
      "content": "<p>Yes ArcFace's margin paramaters are learnable. <br>\nBut I don't really get your other questions.</p>",
      "rawMarkdown": "Yes ArcFace's margin paramaters are learnable. \nBut I don't really get your other questions.",
      "votes": null
    },
    {
      "id": "1032633",
      "postDate": "09/30/2020 10:25:58",
      "content": "<p>Oh, sorry for not stating it clear. I mean why you decided to have a function mapping class sizes to margin level, instead of somehow having individual dynamic margins for each class?</p>",
      "rawMarkdown": "Oh, sorry for not stating it clear. I mean why you decided to have a function mapping class sizes to margin level, instead of somehow having individual dynamic margins for each class?",
      "votes": null
    },
    {
      "id": "1032635",
      "postDate": "09/30/2020 10:26:54",
      "content": "<p>We use arcface loss to update the arcface parameters, using the output of arcface head (arc margin product) and labels. So we don't need label to do the prediction.</p>",
      "rawMarkdown": "We use arcface loss to update the arcface parameters, using the output of arcface head (arc margin product) and labels. So we don't need label to do the prediction.",
      "votes": null
    },
    {
      "id": "1032642",
      "postDate": "09/30/2020 10:34:17",
      "content": "<p>What's yor thought on having individual dynamic margins for each class? <br>\nAs you might definately not want to assign them manually.</p>",
      "rawMarkdown": "What's yor thought on having individual dynamic margins for each class? \nAs you might definately not want to assign them manually.",
      "votes": null
    },
    {
      "id": "1032657",
      "postDate": "09/30/2020 10:46:38",
      "content": "<p>Hmm, I don't have anything specific in mind. Maybe just as simple as increasing/decreasing class-specific margins based on the model's precision on given class over the last K mini-batches?</p>\n<p>Anyways, congratulations to your teams again! Your solution looks elegant - not having local features and non-landmark removal means your global descriptors are insanely good!</p>",
      "rawMarkdown": "Hmm, I don't have anything specific in mind. Maybe just as simple as increasing/decreasing class-specific margins based on the model's precision on given class over the last K mini-batches?\n\nAnyways, congratulations to your teams again! Your solution looks elegant - not having local features and non-landmark removal means your global descriptors are insanely good!",
      "votes": null
    },
    {
      "id": "1032674",
      "postDate": "09/30/2020 11:03:10",
      "content": "<p>I actually tried Sub-center but it did not help us much. But seeing your success with it it might have been useful to explore more.</p>",
      "rawMarkdown": "I actually tried Sub-center but it did not help us much. But seeing your success with it it might have been useful to explore more.",
      "votes": null
    },
    {
      "id": "1032700",
      "postDate": "09/30/2020 11:40:18",
      "content": "<p>Congratulations.</p>",
      "rawMarkdown": "Congratulations.",
      "votes": null
    },
    {
      "id": "1032712",
      "postDate": "09/30/2020 11:50:55",
      "content": "<p>Congratulations and many thanks for the write up! :)<br>\nDo you have any estimate on how much the Sub-center ArcFace and dynamic margin helped?</p>",
      "rawMarkdown": "Congratulations and many thanks for the write up! :)\nDo you have any estimate on how much the Sub-center ArcFace and dynamic margin helped?",
      "votes": null
    },
    {
      "id": "1032724",
      "postDate": "09/30/2020 12:07:11",
      "content": "<p>Wow, great work 😍<br>\nCongratulations 👍</p>",
      "rawMarkdown": "Wow, great work 😍\nCongratulations 👍",
      "votes": null
    },
    {
      "id": "1032739",
      "postDate": "09/30/2020 12:20:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> </p>",
      "rawMarkdown": "Congratulations @boliu0",
      "votes": null
    },
    {
      "id": "1032746",
      "postDate": "09/30/2020 12:27:48",
      "content": "<p>Awesome prize win after a victory in previous comp.  Dream team!</p>",
      "rawMarkdown": "Awesome prize win after a victory in previous comp.  Dream team!",
      "votes": null
    },
    {
      "id": "1032778",
      "postDate": "09/30/2020 12:47:55",
      "content": "<p>You have an idea, you try it, you find it useful/not useful.<br>\nThat's how our final solution comes out ;)</p>\n<p>Be brave to try your own idea next time!</p>",
      "rawMarkdown": "You have an idea, you try it, you find it useful/not useful.\nThat's how our final solution comes out ;)\n\nBe brave to try your own idea next time!",
      "votes": null
    },
    {
      "id": "1032786",
      "postDate": "09/30/2020 12:56:48",
      "content": "<p>dynamic margin: majoy improvement<br>\nsubcenter: minor improvement</p>\n<p>we plan to estimate our dynamic margin better on several datasets then publish the experiment results in our paper.</p>",
      "rawMarkdown": "dynamic margin: majoy improvement\nsubcenter: minor improvement\n\nwe plan to estimate our dynamic margin better on several datasets then publish the experiment results in our paper.",
      "votes": null
    },
    {
      "id": "1032799",
      "postDate": "09/30/2020 13:03:12",
      "content": "<p>Great solution! Very impressive global feature quality. Did you try to also identify/re-rank nonlandmarks in some way like some of the past solutions did? </p>",
      "rawMarkdown": "Great solution! Very impressive global feature quality. Did you try to also identify/re-rank nonlandmarks in some way like some of the past solutions did?",
      "votes": null
    },
    {
      "id": "1032801",
      "postDate": "09/30/2020 13:03:38",
      "content": "<p>That's great, I'm looking forward to it :)<br>\nThanks</p>",
      "rawMarkdown": "That's great, I'm looking forward to it :)\nThanks",
      "votes": null
    },
    {
      "id": "1032814",
      "postDate": "09/30/2020 13:10:46",
      "content": "<p>thanks, psi. It's pity that we did not try to identify the non-landmarks. About this, I learn a lot from your team solution. </p>",
      "rawMarkdown": "thanks, psi. It's pity that we did not try to identify the non-landmarks. About this, I learn a lot from your team solution.",
      "votes": null
    },
    {
      "id": "1032820",
      "postDate": "09/30/2020 13:14:06",
      "content": "<p>Yeah, just from what you describe it looks like you could get quite good boosts from it, even using something similar like last year winners (<a href=\"https://www.kaggle.com/c/landmark-recognition-2019/discussion/94523\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2019/discussion/94523</a>) could net you good boosts. That said, again, very very impressive score with mostly using just the global embedding model.</p>",
      "rawMarkdown": "Yeah, just from what you describe it looks like you could get quite good boosts from it, even using something similar like last year winners (https://www.kaggle.com/c/landmark-recognition-2019/discussion/94523) could net you good boosts. That said, again, very very impressive score with mostly using just the global embedding model.",
      "votes": null
    },
    {
      "id": "1032821",
      "postDate": "09/30/2020 13:14:07",
      "content": "<p>I'm now very curious to know what if you guys do a crossover (awesome global features of 3rd place team + re-ranking technique of the 1st place team) 😲😲😲</p>",
      "rawMarkdown": "I'm now very curious to know what if you guys do a crossover (awesome global features of 3rd place team + re-ranking technique of the 1st place team) 😲😲😲",
      "votes": null
    },
    {
      "id": "1032822",
      "postDate": "09/30/2020 13:15:06",
      "content": "<p>Congratulations and many thanks for the write-up! :) </p>",
      "rawMarkdown": "Congratulations and many thanks for the write-up! :)",
      "votes": null
    },
    {
      "id": "1032824",
      "postDate": "09/30/2020 13:16:30",
      "content": "<p>Yeah - this competition has a lot of potential to combine different solutions.</p>",
      "rawMarkdown": "Yeah - this competition has a lot of potential to combine different solutions.",
      "votes": null
    },
    {
      "id": "1032847",
      "postDate": "09/30/2020 13:28:29",
      "content": "<p>Congrats and well done.</p>",
      "rawMarkdown": "Congrats and well done.",
      "votes": null
    },
    {
      "id": "1032932",
      "postDate": "09/30/2020 14:18:53",
      "content": "<p>One more question, if I may; what optimizer did you guys used?</p>",
      "rawMarkdown": "One more question, if I may; what optimizer did you guys used?",
      "votes": null
    },
    {
      "id": "1032933",
      "postDate": "09/30/2020 14:19:32",
      "content": "<p>Adam, no weight decay</p>",
      "rawMarkdown": "Adam, no weight decay",
      "votes": null
    },
    {
      "id": "1032941",
      "postDate": "09/30/2020 14:26:54",
      "content": "<p>Impressive, thanks for sharing!</p>",
      "rawMarkdown": "Impressive, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1032959",
      "postDate": "09/30/2020 14:54:29",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": null
    },
    {
      "id": "1033166",
      "postDate": "09/30/2020 18:01:24",
      "content": "<p>Very nice solution and impressive to just have used the global features. Congrats!</p>",
      "rawMarkdown": "Very nice solution and impressive to just have used the global features. Congrats!",
      "votes": null
    },
    {
      "id": "1033327",
      "postDate": "09/30/2020 21:09:26",
      "content": "<p>The remarkable thing is, the solutions of both of your teams are <strong>extremely scalable</strong> and <strong>cheap</strong>. Basically, the 3rd place team created a compact yet effective representation, and your team took care of the rest of the pipeline.</p>\n<p>I can only imagine how it could be easily integrated into Google's large-scale image search engine. Local features are slow because you can't pre-compute RANSAC matches (well, at least it would be very un-economical and you would need to cache them somehow). Storing hundreds of local descriptors for each image are un-economical as well.</p>",
      "rawMarkdown": "The remarkable thing is, the solutions of both of your teams are **extremely scalable** and **cheap**. Basically, the 3rd place team created a compact yet effective representation, and your team took care of the rest of the pipeline.\n\nI can only imagine how it could be easily integrated into Google's large-scale image search engine. Local features are slow because you can't pre-compute RANSAC matches (well, at least it would be very un-economical and you would need to cache them somehow). Storing hundreds of local descriptors for each image are un-economical as well.",
      "votes": null
    },
    {
      "id": "1034159",
      "postDate": "10/01/2020 14:50:24",
      "content": "<p>hearty congratulations</p>",
      "rawMarkdown": "hearty congratulations",
      "votes": null
    },
    {
      "id": "1034575",
      "postDate": "10/01/2020 23:15:45",
      "content": "<p>Thank you for sharing the solution and congratulations on 3rd position finish :) </p>\n<p>Any chance you're also able to share with us the training code please?</p>",
      "rawMarkdown": "Thank you for sharing the solution and congratulations on 3rd position finish :) \n\nAny chance you're also able to share with us the training code please?",
      "votes": null
    },
    {
      "id": "1034583",
      "postDate": "10/01/2020 23:45:39",
      "content": "<p>Nice, congrants !!!</p>",
      "rawMarkdown": "Nice, congrants !!!",
      "votes": null
    },
    {
      "id": "1034928",
      "postDate": "10/02/2020 09:58:28",
      "content": "<p>While waiting to see more detailed explanation of dynamic margins, found one article which proposed very similar idea to classification task:<br>\nLearning Imbalanced Datasets with Label-Distribution-Aware Margin Loss<br>\n<a href=\"https://openreview.net/pdf/29888c3e4c2d0c995bc9e0bf7eca07d1e88b2721.pdf\" target=\"_blank\">https://openreview.net/pdf/29888c3e4c2d0c995bc9e0bf7eca07d1e88b2721.pdf</a></p>\n<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> have you seen it?</p>",
      "rawMarkdown": "While waiting to see more detailed explanation of dynamic margins, found one article which proposed very similar idea to classification task:\nLearning Imbalanced Datasets with Label-Distribution-Aware Margin Loss\nhttps://openreview.net/pdf/29888c3e4c2d0c995bc9e0bf7eca07d1e88b2721.pdf\n\n@boliu0 have you seen it?",
      "votes": null
    },
    {
      "id": "1035148",
      "postDate": "10/02/2020 13:45:02",
      "content": "<p>Hi Zakirov, I was not aware of this paper. Thanks for letting us know. I will take a look.</p>",
      "rawMarkdown": "Hi Zakirov, I was not aware of this paper. Thanks for letting us know. I will take a look.",
      "votes": null
    },
    {
      "id": "1035538",
      "postDate": "10/02/2020 19:17:07",
      "content": "<p>Simply great explanation ,i just learned a lot from your explanation and <br>\nCongratulations </p>",
      "rawMarkdown": "Simply great explanation ,i just learned a lot from your explanation and \nCongratulations",
      "votes": null
    },
    {
      "id": "1036875",
      "postDate": "10/04/2020 11:16:20",
      "content": "<p>Congrats and thanks for sharing :)</p>",
      "rawMarkdown": "Congrats and thanks for sharing :)",
      "votes": null
    },
    {
      "id": "1043743",
      "postDate": "10/09/2020 08:01:09",
      "content": "<p>Congrats and thanks for your sharing! It is amazing that the pure global model only can achieve such great success! Can you give a detailed description about the structure of the neck and the head? (backbone-&gt;neck(pooling)-&gt;head)</p>",
      "rawMarkdown": "Congrats and thanks for your sharing! It is amazing that the pure global model only can achieve such great success! Can you give a detailed description about the structure of the neck and the head? (backbone->neck(pooling)->head)",
      "votes": null
    },
    {
      "id": "1049046",
      "postDate": "10/14/2020 03:59:45",
      "content": "<p>Hi Thanks HangWang. We just open sourced the solution: <a href=\"https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution\" target=\"_blank\">https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution</a></p>",
      "rawMarkdown": "Hi Thanks HangWang. We just open sourced the solution: https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution",
      "votes": null
    },
    {
      "id": "1320869",
      "postDate": "05/24/2021 11:41:03",
      "content": "<p>Hello thanks for sharing your solution. May I know if there is a Kaggle notebook implementation for this by any chance? </p>",
      "rawMarkdown": "Hello thanks for sharing your solution. May I know if there is a Kaggle notebook implementation for this by any chance?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1032503,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "09/30/2020 08:51:13",
      "content": "<p>As an addition to <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> 's wonderful post, we will also share techniques that we tried but failed to work. This includes classical approaches in retrieval such as DBA, QE; distractor removal using object detection (OID-pretrained), and training DELF from scratch with our own PyTorch implementation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032529,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "09/30/2020 09:11:09",
      "content": "<p>Congrats and thx <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and the rest of the team for the write-up. Sub-center ArcFace with dynamic margins definitely sounds interesting, will have a look :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1032674,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/30/2020 11:03:10",
          "content": "<p>I actually tried Sub-center but it did not help us much. But seeing your success with it it might have been useful to explore more.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1032531,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "09/30/2020 09:12:58",
      "content": "<p>Wow, <strong>congratulations!</strong> 🎉🎉🎉 It's crazy that you managed to achieve such a high place with just global features.</p>\n<p>I think a teammate of yours mentioned about the dynamic margin idea (different margin for different classes) somewhere in the comments to the Retrieval competition. Glad that it played out well! Can't wait to read your paper!</p>\n<p>A few questions about it:</p>\n<ul>\n<li>Are the ArcFace's margin paramaters learnable or changing over time? If so, you only learn/update the parameters of the continuous function <code>size -&gt; margin</code>, right?</li>\n<li>Why did you make Dynamic Margins to be a function of class sizes, not of class indices?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1032632,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/30/2020 10:23:22",
          "content": "<p>Yes ArcFace's margin paramaters are learnable. <br>\nBut I don't really get your other questions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032633,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/30/2020 10:25:58",
          "content": "<p>Oh, sorry for not stating it clear. I mean why you decided to have a function mapping class sizes to margin level, instead of somehow having individual dynamic margins for each class?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032642,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/30/2020 10:34:17",
          "content": "<p>What's yor thought on having individual dynamic margins for each class? <br>\nAs you might definately not want to assign them manually.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032657,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/30/2020 10:46:38",
          "content": "<p>Hmm, I don't have anything specific in mind. Maybe just as simple as increasing/decreasing class-specific margins based on the model's precision on given class over the last K mini-batches?</p>\n<p>Anyways, congratulations to your teams again! Your solution looks elegant - not having local features and non-landmark removal means your global descriptors are insanely good!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032778,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/30/2020 12:47:55",
          "content": "<p>You have an idea, you try it, you find it useful/not useful.<br>\nThat's how our final solution comes out ;)</p>\n<p>Be brave to try your own idea next time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1032579,
      "author_name": "hav4ik",
      "author_url": "",
      "post_date": "09/30/2020 09:53:36",
      "content": "<p>Wow, what an elegant solution. Congrats. May I ask what hardware you used for training models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1032612,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/30/2020 10:12:22",
          "content": "<p>I believe you can even get good models with just Colab Pro + a lot of time in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1032617,
      "author_name": "josealways123",
      "author_url": "",
      "post_date": "09/30/2020 10:15:00",
      "content": "<p>Congratzz on your 3rd place. Nice write up and really helpful. </p>\n<blockquote>\n  <p>If we use model's ArcFace head to predict test set, our best single fold model's public LB is only 0.564</p>\n</blockquote>\n<p>How are you predicting with Arcface head on the test set? I mean, arcface takes the label as input and so don't we need labels to predict? Are you just passing a 81313 dimensional 0-vector as label?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1032635,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/30/2020 10:26:54",
          "content": "<p>We use arcface loss to update the arcface parameters, using the output of arcface head (arc margin product) and labels. So we don't need label to do the prediction.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1032700,
      "author_name": "priyasinha1225",
      "author_url": "",
      "post_date": "09/30/2020 11:40:18",
      "content": "<p>Congratulations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032712,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "09/30/2020 11:50:55",
      "content": "<p>Congratulations and many thanks for the write up! :)<br>\nDo you have any estimate on how much the Sub-center ArcFace and dynamic margin helped?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1032786,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/30/2020 12:56:48",
          "content": "<p>dynamic margin: majoy improvement<br>\nsubcenter: minor improvement</p>\n<p>we plan to estimate our dynamic margin better on several datasets then publish the experiment results in our paper.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032801,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "09/30/2020 13:03:38",
          "content": "<p>That's great, I'm looking forward to it :)<br>\nThanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032932,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "09/30/2020 14:18:53",
          "content": "<p>One more question, if I may; what optimizer did you guys used?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032933,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/30/2020 14:19:32",
          "content": "<p>Adam, no weight decay</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1032724,
      "author_name": "rahim3",
      "author_url": "",
      "post_date": "09/30/2020 12:07:11",
      "content": "<p>Wow, great work 😍<br>\nCongratulations 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032739,
      "author_name": "data855",
      "author_url": "",
      "post_date": "09/30/2020 12:20:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032746,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/30/2020 12:27:48",
      "content": "<p>Awesome prize win after a victory in previous comp.  Dream team!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032799,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "09/30/2020 13:03:12",
      "content": "<p>Great solution! Very impressive global feature quality. Did you try to also identify/re-rank nonlandmarks in some way like some of the past solutions did? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1032814,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "09/30/2020 13:10:46",
          "content": "<p>thanks, psi. It's pity that we did not try to identify the non-landmarks. About this, I learn a lot from your team solution. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032820,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/30/2020 13:14:06",
          "content": "<p>Yeah, just from what you describe it looks like you could get quite good boosts from it, even using something similar like last year winners (<a href=\"https://www.kaggle.com/c/landmark-recognition-2019/discussion/94523\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2019/discussion/94523</a>) could net you good boosts. That said, again, very very impressive score with mostly using just the global embedding model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032821,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/30/2020 13:14:07",
          "content": "<p>I'm now very curious to know what if you guys do a crossover (awesome global features of 3rd place team + re-ranking technique of the 1st place team) 😲😲😲</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032824,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/30/2020 13:16:30",
          "content": "<p>Yeah - this competition has a lot of potential to combine different solutions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1033327,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/30/2020 21:09:26",
          "content": "<p>The remarkable thing is, the solutions of both of your teams are <strong>extremely scalable</strong> and <strong>cheap</strong>. Basically, the 3rd place team created a compact yet effective representation, and your team took care of the rest of the pipeline.</p>\n<p>I can only imagine how it could be easily integrated into Google's large-scale image search engine. Local features are slow because you can't pre-compute RANSAC matches (well, at least it would be very un-economical and you would need to cache them somehow). Storing hundreds of local descriptors for each image are un-economical as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1032822,
      "author_name": "piyushagni5",
      "author_url": "",
      "post_date": "09/30/2020 13:15:06",
      "content": "<p>Congratulations and many thanks for the write-up! :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032847,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "09/30/2020 13:28:29",
      "content": "<p>Congrats and well done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032941,
      "author_name": "delllectron",
      "author_url": "",
      "post_date": "09/30/2020 14:26:54",
      "content": "<p>Impressive, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1032959,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "09/30/2020 14:54:29",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1033166,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "09/30/2020 18:01:24",
      "content": "<p>Very nice solution and impressive to just have used the global features. Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1034575,
      "author_name": "aroraaman",
      "author_url": "",
      "post_date": "10/01/2020 23:15:45",
      "content": "<p>Thank you for sharing the solution and congratulations on 3rd position finish :) </p>\n<p>Any chance you're also able to share with us the training code please?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1034583,
      "author_name": "tiagorosa",
      "author_url": "",
      "post_date": "10/01/2020 23:45:39",
      "content": "<p>Nice, congrants !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1034928,
      "author_name": "zakajd",
      "author_url": "",
      "post_date": "10/02/2020 09:58:28",
      "content": "<p>While waiting to see more detailed explanation of dynamic margins, found one article which proposed very similar idea to classification task:<br>\nLearning Imbalanced Datasets with Label-Distribution-Aware Margin Loss<br>\n<a href=\"https://openreview.net/pdf/29888c3e4c2d0c995bc9e0bf7eca07d1e88b2721.pdf\" target=\"_blank\">https://openreview.net/pdf/29888c3e4c2d0c995bc9e0bf7eca07d1e88b2721.pdf</a></p>\n<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> have you seen it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1035148,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "10/02/2020 13:45:02",
          "content": "<p>Hi Zakirov, I was not aware of this paper. Thanks for letting us know. I will take a look.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1035538,
      "author_name": "surekharamireddy",
      "author_url": "",
      "post_date": "10/02/2020 19:17:07",
      "content": "<p>Simply great explanation ,i just learned a lot from your explanation and <br>\nCongratulations </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1036875,
      "author_name": "khatkeashish",
      "author_url": "",
      "post_date": "10/04/2020 11:16:20",
      "content": "<p>Congrats and thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1043743,
      "author_name": "sjtuwh",
      "author_url": "",
      "post_date": "10/09/2020 08:01:09",
      "content": "<p>Congrats and thanks for your sharing! It is amazing that the pure global model only can achieve such great success! Can you give a detailed description about the structure of the neck and the head? (backbone-&gt;neck(pooling)-&gt;head)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1049046,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "10/14/2020 03:59:45",
          "content": "<p>Hi Thanks HangWang. We just open sourced the solution: <a href=\"https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution\" target=\"_blank\">https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1320869,
      "author_name": "yixiang",
      "author_url": "",
      "post_date": "05/24/2021 11:41:03",
      "content": "<p>Hello thanks for sharing your solution. May I know if there is a Kaggle notebook implementation for this by any chance? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1034159,
      "author_name": "meesalasaidhanush",
      "author_url": "",
      "post_date": "10/01/2020 14:50:24",
      "content": "<p>hearty congratulations</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1032394": "Thanks to the organizers and congrats to all the winners! Our ( @haqishen @boliu0 @garybios @alexanderliao) solution is a pure global feature metric learning approach with some tricks.\n\n## architecture: sub-center ArcFace with dynamic margins\nArcFace ([paper](https://arxiv.org/abs/1801.07698)) has become a standard metric learning method on Kaggle over the past two years or so. In this competition, we used Sub-center ArcFace ([paper](https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf)), a recent improvement over ArcFace by the same authors. The idea is that each class may have more than one class center. For example, a certain landmark's photos may have a few clusters (e.g. from different angles). Sub-center ArcFace's weights store multiple class centers' representations, which can increase classification accuracy and improve global feature's quality.\n\nThe classes in GLD dataset are extremely imbalanced with longs tails. For models to converge better in the presence of heavy imbalance, smaller classes need to have bigger margins as they are harder to learn. Instead of manually setting different margin levels based on class size, we introduce *dynamic margins*, a family of continuous functions mapping class size to margin level. This give us some major boost. Details to be described in paper.\n\nWe train the model using ArcFace loss only.\n\n## validation scheme\nStratified 5-fold. Training using 4 folds, validate on only 1/15 of 1 fold to save time. Use different folds for different single models, for maximal diversity in ensemble.\n\nWe use GAP metric to validate. CV GAP is very high compared to LB. All model's CV GAP are over 0.967. Good news is that CV GAP and LB have high correlation.\n\n## test set predicting strategy\n- If we use model's ArcFace head to predict test set, our best single fold model's public LB is only **0.564**\n- A better strategy is to calculate the global feature cosine similarity of each [private train image, private test image] pair, and use the top1 neighbor and corresponding cosine similarity of each test image as prediction. This gives us **0.604** public LB for the same single fold model\n- Instead of using top1 neighbor only, we can improve it by combining top5 neighbors and their similarities. Best combining function is 8th power. \nE.g. let's say test image A's top5 neighbors and their cosine similarities are: class 1 (0.9), class 2 (0.8), class 2 (0.7), class 1 (0.5), class 3 (0.45). Then class 1's total score is `0.9**8 + 0.5**8 = 0.434`; class 2's total score is `0.8**8 + 0.7**8 = 0.225`. So we predict class 1 with p=0.434\nLB increases to **0.610**\n- We can further improve it by incorporating ArcFace head's predictions with 12th power. In above example, assuming image A's ArcFace head give class 1 score = 0.75, class 2 score = 0.88. Then class 1's total score becomes `(0.9**8 + 0.5**8) * 0.75**12 = 0.0138`; class 2's total score becomes `(0.8**8 + 0.7**8) * 0.88**12 = 0.0486`. Now we predict class 2 with p=0.0486\nLB increases to **0.618**\n\n\n## augmentations\nWe resize all images to square shape without cropping.\n```\nimport albumentations as A\nA.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),    \n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n        A.Normalize()\n    ])\n```\n\n## pretraining and finetuning\n- In cGLDv2 (cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.\n- In GLDv2, there are 4.1m training images and 200k classes, among which 3.2m images belong to the 81k classes in cGLDv2.\n\nWe noticed that (1) training with the 3.2m data gives better results than only the 1.6m competition cGLDv2 data, (2) pretraining on all 4.1m data then finetuning on 3.2m data gives even better results.\n\n## 3-stage training schedule\n- Stage 1 (pretrain): 10 epochs with small image size (256) on 4.1m data\n- Stage 2 (finetune): about 16 epochs with medium image size (512 to 768 depending on model) on 3.2m data. Number of epochs varies by model and ranges from 13 to 21. \n- Stage 3 (finetune): 1 epoch with large image size (672 to 1024) on 3.2m data\n\nNote: above schedules are for Sub1 (private 0.6289). Sub2 has higher score (private 0.6344) but more complex schedules, i.e. longer with more rounds of finetuning. Both submissions are 3rd place.\n\n## ensemble\n7 models: EfficientNet B7, B6, B5, B4, B3, [ResNeSt101](https://github.com/zhanghang1989/ResNeSt), [ReXNet2.0](https://github.com/clovaai/rexnet)\n\nFor global feature neighbor search, we concatenate each model's 512-dimension feature; for ArcFace head, we take simple average of each model's logits.\n\nBest single model is EfficientB6. **private = 0.6005, public = 0.6179** (We didn't submit all the epochs, there may be higher ones)\nSub1 is 7 model ensemble. **private = 0.6289, public = 0.6604**\nSub2 is 9 model ensemble (B5 and B6 twice). **private = 0.6344, public = 0.6581**\n\n\n### [update 10/12/2020]\npaper: https://arxiv.org/abs/2010.05350\nrepo: https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution",
    "1032503": "As an addition to @boliu0 's wonderful post, we will also share techniques that we tried but failed to work. This includes classical approaches in retrieval such as DBA, QE; distractor removal using object detection (OID-pretrained), and training DELF from scratch with our own PyTorch implementation.",
    "1032529": "Congrats and thx @boliu0 and the rest of the team for the write-up. Sub-center ArcFace with dynamic margins definitely sounds interesting, will have a look :)",
    "1032531": "Wow, **congratulations!** 🎉🎉🎉 It's crazy that you managed to achieve such a high place with just global features.\n\nI think a teammate of yours mentioned about the dynamic margin idea (different margin for different classes) somewhere in the comments to the Retrieval competition. Glad that it played out well! Can't wait to read your paper!\n\nA few questions about it:\n- Are the ArcFace's margin paramaters learnable or changing over time? If so, you only learn/update the parameters of the continuous function `size -> margin`, right?\n- Why did you make Dynamic Margins to be a function of class sizes, not of class indices?",
    "1032579": "Wow, what an elegant solution. Congrats. May I ask what hardware you used for training models?",
    "1032612": "I believe you can even get good models with just Colab Pro + a lot of time in this competition.",
    "1032617": "Congratzz on your 3rd place. Nice write up and really helpful. \n\n> If we use model's ArcFace head to predict test set, our best single fold model's public LB is only 0.564\n\nHow are you predicting with Arcface head on the test set? I mean, arcface takes the label as input and so don't we need labels to predict? Are you just passing a 81313 dimensional 0-vector as label?",
    "1032632": "Yes ArcFace's margin paramaters are learnable. \nBut I don't really get your other questions.",
    "1032633": "Oh, sorry for not stating it clear. I mean why you decided to have a function mapping class sizes to margin level, instead of somehow having individual dynamic margins for each class?",
    "1032635": "We use arcface loss to update the arcface parameters, using the output of arcface head (arc margin product) and labels. So we don't need label to do the prediction.",
    "1032642": "What's yor thought on having individual dynamic margins for each class? \nAs you might definately not want to assign them manually.",
    "1032657": "Hmm, I don't have anything specific in mind. Maybe just as simple as increasing/decreasing class-specific margins based on the model's precision on given class over the last K mini-batches?\n\nAnyways, congratulations to your teams again! Your solution looks elegant - not having local features and non-landmark removal means your global descriptors are insanely good!",
    "1032674": "I actually tried Sub-center but it did not help us much. But seeing your success with it it might have been useful to explore more.",
    "1032700": "Congratulations.",
    "1032712": "Congratulations and many thanks for the write up! :)\nDo you have any estimate on how much the Sub-center ArcFace and dynamic margin helped?",
    "1032724": "Wow, great work 😍\nCongratulations 👍",
    "1032739": "Congratulations @boliu0",
    "1032746": "Awesome prize win after a victory in previous comp.  Dream team!",
    "1032778": "You have an idea, you try it, you find it useful/not useful.\nThat's how our final solution comes out ;)\n\nBe brave to try your own idea next time!",
    "1032786": "dynamic margin: majoy improvement\nsubcenter: minor improvement\n\nwe plan to estimate our dynamic margin better on several datasets then publish the experiment results in our paper.",
    "1032799": "Great solution! Very impressive global feature quality. Did you try to also identify/re-rank nonlandmarks in some way like some of the past solutions did?",
    "1032801": "That's great, I'm looking forward to it :)\nThanks",
    "1032814": "thanks, psi. It's pity that we did not try to identify the non-landmarks. About this, I learn a lot from your team solution.",
    "1032820": "Yeah, just from what you describe it looks like you could get quite good boosts from it, even using something similar like last year winners (https://www.kaggle.com/c/landmark-recognition-2019/discussion/94523) could net you good boosts. That said, again, very very impressive score with mostly using just the global embedding model.",
    "1032821": "I'm now very curious to know what if you guys do a crossover (awesome global features of 3rd place team + re-ranking technique of the 1st place team) 😲😲😲",
    "1032822": "Congratulations and many thanks for the write-up! :)",
    "1032824": "Yeah - this competition has a lot of potential to combine different solutions.",
    "1032847": "Congrats and well done.",
    "1032932": "One more question, if I may; what optimizer did you guys used?",
    "1032933": "Adam, no weight decay",
    "1032941": "Impressive, thanks for sharing!",
    "1032959": "Congrats and thanks for sharing!",
    "1033166": "Very nice solution and impressive to just have used the global features. Congrats!",
    "1033327": "The remarkable thing is, the solutions of both of your teams are **extremely scalable** and **cheap**. Basically, the 3rd place team created a compact yet effective representation, and your team took care of the rest of the pipeline.\n\nI can only imagine how it could be easily integrated into Google's large-scale image search engine. Local features are slow because you can't pre-compute RANSAC matches (well, at least it would be very un-economical and you would need to cache them somehow). Storing hundreds of local descriptors for each image are un-economical as well.",
    "1034159": "hearty congratulations",
    "1034575": "Thank you for sharing the solution and congratulations on 3rd position finish :) \n\nAny chance you're also able to share with us the training code please?",
    "1034583": "Nice, congrants !!!",
    "1034928": "While waiting to see more detailed explanation of dynamic margins, found one article which proposed very similar idea to classification task:\nLearning Imbalanced Datasets with Label-Distribution-Aware Margin Loss\nhttps://openreview.net/pdf/29888c3e4c2d0c995bc9e0bf7eca07d1e88b2721.pdf\n\n@boliu0 have you seen it?",
    "1035148": "Hi Zakirov, I was not aware of this paper. Thanks for letting us know. I will take a look.",
    "1035538": "Simply great explanation ,i just learned a lot from your explanation and \nCongratulations",
    "1036875": "Congrats and thanks for sharing :)",
    "1043743": "Congrats and thanks for your sharing! It is amazing that the pure global model only can achieve such great success! Can you give a detailed description about the structure of the neck and the head? (backbone->neck(pooling)->head)",
    "1049046": "Hi Thanks HangWang. We just open sourced the solution: https://github.com/haqishen/Google-Landmark-Recognition-2020-3rd-Place-Solution",
    "1320869": "Hello thanks for sharing your solution. May I know if there is a Kaggle notebook implementation for this by any chance?"
  },
  "source": "meta"
}