{
  "id": 187894,
  "title": "7th place solution (with inference kernel + training code)",
  "url": "/competitions/landmark-recognition-2020/writeups/dna-7th-place-solution-with-inference-kernel-train",
  "author_name": "",
  "post_date": "2020-09-30T18:28:23.937Z",
  "votes": 34,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Firstly, I want to thank my teammates <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> and <a href=\"https://www.kaggle.com/dattran2346\" target=\"_blank\">@dattran2346</a> for their hard work. In this comp., they focused on improving global models so that I could spend time on local part.</p>\n<p>Our best sub. was geometric mean of similarities from 4 models (2 backbones SEResNext101+ResNext101-32x4d at 512x512 and 736x736 image sizes) + re-ranking by SuperPoint + SuperGlue. We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.</p>\n<p><strong>1. Global models</strong><br>\nWe continued training from our best checkpoints from retrieval challenge (<a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176151)\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176151)</a>. Our models were originally trained with focal smoothing loss with a tweak to work nicely with CosFace (thanks to Ahmet) on GLv2 clean dataset. For this comp., we used all the remaining train images from 81k classes (3.2m images) for re-training. We also came across the seesaw loss paper (<a href=\"https://arxiv.org/abs/2008.10032\" target=\"_blank\">https://arxiv.org/abs/2008.10032</a>) and decided to give it a try (thanks to Dat).<br>\nModel architecture was just simply CNN + GEM + Linear + BN + CosFace, like many other teams. <br>\nOur 2 models were then re-trained with these two losses as following:</p>\n<ul>\n<li>Stage 1: 20 epochs at size 512x512</li>\n<li>Stage 2: froze batchnorm layers, then fine-tuned at size 736x736 for 2 epochs.<br>\nCompared to the retrieval comp., our local validation scores were boosted by 5% and correlated well with the LB. From our estimates, 3% was due to adding extra data, 1% from the seesaw loss and 1% from larger image size.</li>\n</ul>\n<p><strong>2. Local model</strong><br>\nAfter obtaining the top-k nearest train ids for each test image, we used SuperPoint + HRNetv2 pre-trained on ADE20k dataset to filter predicted keypoints on sky/person/flowers/tree classes + SuperGlue (copied from the winning sol. at CVPR this year) to calculate the number of inliers between each image pair. Local score was then obtained using the formula provided by the host team (with max_num_inliers=200) and then multiplied with global score to get final score. This post-processing step gave a very strong boost (4-6% in public LB to our above global models). However, we noticed that the better the global models were, the less improvement this local re-ranking step gave us. In addition, we couldn't come up with a reliable strategy to evaluate SuperPoint+SuperGlue when integrating with global models locally. Relying solely on LB was pretty dangerous and we were lucky to stay at 7th 😅.<br>\nI also tried to re-train SuperPoint and SuperGlue on a subset of 200k clean train images but couldn't obtain any positive improvement.</p>\n<p>Our kernel + training code + model checkpoints were here: <a href=\"https://www.kaggle.com/andy2709/fork-of-recognition-notebook-16367c-511fa4-95d3bf?scriptVersionId=43685736\" target=\"_blank\">https://www.kaggle.com/andy2709/fork-of-recognition-notebook-16367c-511fa4-95d3bf?scriptVersionId=43685736</a>. </p>\n<p>P/s: The DELG weights + training config. files for 3 different backbones (res50, res101 and seres101) will be uploaded soon 😃. They were all trained with 512x512  images for 10 epochs with AdamW and cosine scheduler. In retrieval challenge, res50's perf was 0.299/ 0.268 (public/ private); others were higher so I'm quite confident in the implementation correctness. </p>",
  "messages": [
    {
      "id": "1033183",
      "postDate": "09/30/2020 18:15:45",
      "content": "<p>Firstly, I want to thank my teammates <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> and <a href=\"https://www.kaggle.com/dattran2346\" target=\"_blank\">@dattran2346</a> for their hard work. In this comp., they focused on improving global models so that I could spend time on local part.</p>\n<p>Our best sub. was geometric mean of similarities from 4 models (2 backbones SEResNext101+ResNext101-32x4d at 512x512 and 736x736 image sizes) + re-ranking by SuperPoint + SuperGlue. We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.</p>\n<p><strong>1. Global models</strong><br>\nWe continued training from our best checkpoints from retrieval challenge (<a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176151)\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176151)</a>. Our models were originally trained with focal smoothing loss with a tweak to work nicely with CosFace (thanks to Ahmet) on GLv2 clean dataset. For this comp., we used all the remaining train images from 81k classes (3.2m images) for re-training. We also came across the seesaw loss paper (<a href=\"https://arxiv.org/abs/2008.10032\" target=\"_blank\">https://arxiv.org/abs/2008.10032</a>) and decided to give it a try (thanks to Dat).<br>\nModel architecture was just simply CNN + GEM + Linear + BN + CosFace, like many other teams. <br>\nOur 2 models were then re-trained with these two losses as following:</p>\n<ul>\n<li>Stage 1: 20 epochs at size 512x512</li>\n<li>Stage 2: froze batchnorm layers, then fine-tuned at size 736x736 for 2 epochs.<br>\nCompared to the retrieval comp., our local validation scores were boosted by 5% and correlated well with the LB. From our estimates, 3% was due to adding extra data, 1% from the seesaw loss and 1% from larger image size.</li>\n</ul>\n<p><strong>2. Local model</strong><br>\nAfter obtaining the top-k nearest train ids for each test image, we used SuperPoint + HRNetv2 pre-trained on ADE20k dataset to filter predicted keypoints on sky/person/flowers/tree classes + SuperGlue (copied from the winning sol. at CVPR this year) to calculate the number of inliers between each image pair. Local score was then obtained using the formula provided by the host team (with max_num_inliers=200) and then multiplied with global score to get final score. This post-processing step gave a very strong boost (4-6% in public LB to our above global models). However, we noticed that the better the global models were, the less improvement this local re-ranking step gave us. In addition, we couldn't come up with a reliable strategy to evaluate SuperPoint+SuperGlue when integrating with global models locally. Relying solely on LB was pretty dangerous and we were lucky to stay at 7th 😅.<br>\nI also tried to re-train SuperPoint and SuperGlue on a subset of 200k clean train images but couldn't obtain any positive improvement.</p>\n<p>Our kernel + training code + model checkpoints were here: <a href=\"https://www.kaggle.com/andy2709/fork-of-recognition-notebook-16367c-511fa4-95d3bf?scriptVersionId=43685736\" target=\"_blank\">https://www.kaggle.com/andy2709/fork-of-recognition-notebook-16367c-511fa4-95d3bf?scriptVersionId=43685736</a>. </p>\n<p>P/s: The DELG weights + training config. files for 3 different backbones (res50, res101 and seres101) will be uploaded soon 😃. They were all trained with 512x512  images for 10 epochs with AdamW and cosine scheduler. In retrieval challenge, res50's perf was 0.299/ 0.268 (public/ private); others were higher so I'm quite confident in the implementation correctness. </p>",
      "rawMarkdown": "Firstly, I want to thank my teammates @aerdem4 and @dattran2346 for their hard work. In this comp., they focused on improving global models so that I could spend time on local part.\n\nOur best sub. was geometric mean of similarities from 4 models (2 backbones SEResNext101+ResNext101-32x4d at 512x512 and 736x736 image sizes) + re-ranking by SuperPoint + SuperGlue. We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.\n\n**1. Global models**\nWe continued training from our best checkpoints from retrieval challenge (https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176151). Our models were originally trained with focal smoothing loss with a tweak to work nicely with CosFace (thanks to Ahmet) on GLv2 clean dataset. For this comp., we used all the remaining train images from 81k classes (3.2m images) for re-training. We also came across the seesaw loss paper (https://arxiv.org/abs/2008.10032) and decided to give it a try (thanks to Dat).\nModel architecture was just simply CNN + GEM + Linear + BN + CosFace, like many other teams. \nOur 2 models were then re-trained with these two losses as following:\n+ Stage 1: 20 epochs at size 512x512\n+ Stage 2: froze batchnorm layers, then fine-tuned at size 736x736 for 2 epochs.\nCompared to the retrieval comp., our local validation scores were boosted by 5% and correlated well with the LB. From our estimates, 3% was due to adding extra data, 1% from the seesaw loss and 1% from larger image size.\n\n**2. Local model**\nAfter obtaining the top-k nearest train ids for each test image, we used SuperPoint + HRNetv2 pre-trained on ADE20k dataset to filter predicted keypoints on sky/person/flowers/tree classes + SuperGlue (copied from the winning sol. at CVPR this year) to calculate the number of inliers between each image pair. Local score was then obtained using the formula provided by the host team (with max_num_inliers=200) and then multiplied with global score to get final score. This post-processing step gave a very strong boost (4-6% in public LB to our above global models). However, we noticed that the better the global models were, the less improvement this local re-ranking step gave us. In addition, we couldn't come up with a reliable strategy to evaluate SuperPoint+SuperGlue when integrating with global models locally. Relying solely on LB was pretty dangerous and we were lucky to stay at 7th 😅.\nI also tried to re-train SuperPoint and SuperGlue on a subset of 200k clean train images but couldn't obtain any positive improvement.\n\nOur kernel + training code + model checkpoints were here: https://www.kaggle.com/andy2709/fork-of-recognition-notebook-16367c-511fa4-95d3bf?scriptVersionId=43685736. \n\nP/s: The DELG weights + training config. files for 3 different backbones (res50, res101 and seres101) will be uploaded soon 😃. They were all trained with 512x512  images for 10 epochs with AdamW and cosine scheduler. In retrieval challenge, res50's perf was 0.299/ 0.268 (public/ private); others were higher so I'm quite confident in the implementation correctness.",
      "votes": null
    },
    {
      "id": "1033227",
      "postDate": "09/30/2020 18:49:45",
      "content": "<p>Amazing approach 😍<br>\nCongratulations to your team, great work 👍</p>",
      "rawMarkdown": "Amazing approach 😍\nCongratulations to your team, great work 👍",
      "votes": null
    },
    {
      "id": "1033402",
      "postDate": "10/01/2020 00:38:20",
      "content": "<p>Thanks for the write-up, and congratulations on achieving such a remarkable results! It is interesting to see that SuperPoints and SuperGlues works so well outside the data distribution that they've trained on.</p>",
      "rawMarkdown": "Thanks for the write-up, and congratulations on achieving such a remarkable results! It is interesting to see that SuperPoints and SuperGlues works so well outside the data distribution that they've trained on.",
      "votes": null
    },
    {
      "id": "1033479",
      "postDate": "10/01/2020 03:11:34",
      "content": "<p>Congratulations to you for your second consecutive gold medal bro Nhan <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
      "rawMarkdown": "Congratulations to you for your second consecutive gold medal bro Nhan @andy2709",
      "votes": null
    },
    {
      "id": "1043838",
      "postDate": "10/09/2020 09:30:55",
      "content": "<p>Congrats! Thanks for your sharing!  <br>\nI wonder if it works?</p>\n<blockquote>\n  <p>We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.</p>\n</blockquote>\n<p>And do you use any non-landmark filtering tricks?</p>",
      "rawMarkdown": "Congrats! Thanks for your sharing!  \nI wonder if it works?\n> We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.\n\nAnd do you use any non-landmark filtering tricks?",
      "votes": null
    },
    {
      "id": "1047354",
      "postDate": "10/12/2020 14:04:05",
      "content": "<p>thanks for share, what is HRNetv2?</p>",
      "rawMarkdown": "thanks for share, what is HRNetv2?",
      "votes": null
    },
    {
      "id": "1047416",
      "postDate": "10/12/2020 15:23:19",
      "content": "<p>It's a high-performant segmentation model. You can find it here: <a href=\"https://github.com/CSAILVision/semantic-segmentation-pytorch\" target=\"_blank\">https://github.com/CSAILVision/semantic-segmentation-pytorch</a></p>",
      "rawMarkdown": "It's a high-performant segmentation model. You can find it here: https://github.com/CSAILVision/semantic-segmentation-pytorch",
      "votes": null
    },
    {
      "id": "1048319",
      "postDate": "10/13/2020 11:37:30",
      "content": "<p>ok, i get, use HRNetv2 to filter the keypoint on sky/person/flowers/tree classes</p>",
      "rawMarkdown": "ok, i get, use HRNetv2 to filter the keypoint on sky/person/flowers/tree classes",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1033227,
      "author_name": "rahim3",
      "author_url": "",
      "post_date": "09/30/2020 18:49:45",
      "content": "<p>Amazing approach 😍<br>\nCongratulations to your team, great work 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1033402,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "10/01/2020 00:38:20",
      "content": "<p>Thanks for the write-up, and congratulations on achieving such a remarkable results! It is interesting to see that SuperPoints and SuperGlues works so well outside the data distribution that they've trained on.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1033479,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "10/01/2020 03:11:34",
      "content": "<p>Congratulations to you for your second consecutive gold medal bro Nhan <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1043838,
      "author_name": "sjtuwh",
      "author_url": "",
      "post_date": "10/09/2020 09:30:55",
      "content": "<p>Congrats! Thanks for your sharing!  <br>\nI wonder if it works?</p>\n<blockquote>\n  <p>We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.</p>\n</blockquote>\n<p>And do you use any non-landmark filtering tricks?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047354,
      "author_name": "finlay",
      "author_url": "",
      "post_date": "10/12/2020 14:04:05",
      "content": "<p>thanks for share, what is HRNetv2?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1047416,
          "author_name": "andy2709",
          "author_url": "",
          "post_date": "10/12/2020 15:23:19",
          "content": "<p>It's a high-performant segmentation model. You can find it here: <a href=\"https://github.com/CSAILVision/semantic-segmentation-pytorch\" target=\"_blank\">https://github.com/CSAILVision/semantic-segmentation-pytorch</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048319,
          "author_name": "finlay",
          "author_url": "",
          "post_date": "10/13/2020 11:37:30",
          "content": "<p>ok, i get, use HRNetv2 to filter the keypoint on sky/person/flowers/tree classes</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1033183": "Firstly, I want to thank my teammates @aerdem4 and @dattran2346 for their hard work. In this comp., they focused on improving global models so that I could spend time on local part.\n\nOur best sub. was geometric mean of similarities from 4 models (2 backbones SEResNext101+ResNext101-32x4d at 512x512 and 736x736 image sizes) + re-ranking by SuperPoint + SuperGlue. We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.\n\n**1. Global models**\nWe continued training from our best checkpoints from retrieval challenge (https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176151). Our models were originally trained with focal smoothing loss with a tweak to work nicely with CosFace (thanks to Ahmet) on GLv2 clean dataset. For this comp., we used all the remaining train images from 81k classes (3.2m images) for re-training. We also came across the seesaw loss paper (https://arxiv.org/abs/2008.10032) and decided to give it a try (thanks to Dat).\nModel architecture was just simply CNN + GEM + Linear + BN + CosFace, like many other teams. \nOur 2 models were then re-trained with these two losses as following:\n+ Stage 1: 20 epochs at size 512x512\n+ Stage 2: froze batchnorm layers, then fine-tuned at size 736x736 for 2 epochs.\nCompared to the retrieval comp., our local validation scores were boosted by 5% and correlated well with the LB. From our estimates, 3% was due to adding extra data, 1% from the seesaw loss and 1% from larger image size.\n\n**2. Local model**\nAfter obtaining the top-k nearest train ids for each test image, we used SuperPoint + HRNetv2 pre-trained on ADE20k dataset to filter predicted keypoints on sky/person/flowers/tree classes + SuperGlue (copied from the winning sol. at CVPR this year) to calculate the number of inliers between each image pair. Local score was then obtained using the formula provided by the host team (with max_num_inliers=200) and then multiplied with global score to get final score. This post-processing step gave a very strong boost (4-6% in public LB to our above global models). However, we noticed that the better the global models were, the less improvement this local re-ranking step gave us. In addition, we couldn't come up with a reliable strategy to evaluate SuperPoint+SuperGlue when integrating with global models locally. Relying solely on LB was pretty dangerous and we were lucky to stay at 7th 😅.\nI also tried to re-train SuperPoint and SuperGlue on a subset of 200k clean train images but couldn't obtain any positive improvement.\n\nOur kernel + training code + model checkpoints were here: https://www.kaggle.com/andy2709/fork-of-recognition-notebook-16367c-511fa4-95d3bf?scriptVersionId=43685736. \n\nP/s: The DELG weights + training config. files for 3 different backbones (res50, res101 and seres101) will be uploaded soon 😃. They were all trained with 512x512  images for 10 epochs with AdamW and cosine scheduler. In retrieval challenge, res50's perf was 0.299/ 0.268 (public/ private); others were higher so I'm quite confident in the implementation correctness.",
    "1033227": "Amazing approach 😍\nCongratulations to your team, great work 👍",
    "1033402": "Thanks for the write-up, and congratulations on achieving such a remarkable results! It is interesting to see that SuperPoints and SuperGlues works so well outside the data distribution that they've trained on.",
    "1033479": "Congratulations to you for your second consecutive gold medal bro Nhan @andy2709",
    "1043838": "Congrats! Thanks for your sharing!  \nI wonder if it works?\n> We pre-computed 1.6m train embeddings for 4 models, added them as ext. data and simply filtered for the relevant private image ids (100k) for inference.\n\nAnd do you use any non-landmark filtering tricks?",
    "1047354": "thanks for share, what is HRNetv2?",
    "1047416": "It's a high-performant segmentation model. You can find it here: https://github.com/CSAILVision/semantic-segmentation-pytorch",
    "1048319": "ok, i get, use HRNetv2 to filter the keypoint on sky/person/flowers/tree classes"
  },
  "source": "meta"
}