{
  "id": 188299,
  "title": "2nd place solution",
  "url": "/competitions/landmark-recognition-2020/discussion/188299",
  "author_name": "bestfitting",
  "post_date": "2020-10-02T16:46:12.867000",
  "votes": 97,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Congratulations to all the winners, thanks to kaggle and Google for hosting Landmark Retrieval/Recognition competition. </p>\n<p>Here is my brief solution to this competition.</p>\n<p><strong>Definition</strong><br>\n<strong>GLD_v2c</strong>(cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.<br>\n<strong>GLD_v2x</strong>, in GLDv2, there are 3.2 million images belong to the 81k classes in GLD_v2c. I define these 3.2m images as GLD_v2x.<br>\n<strong>query image</strong>, the test image when submitting to kaggle or the image in validation set when validate locally.<br>\n<strong>Index image set</strong>, the images from train set, as described in Data page of this competition, “subset contains all of the training set images associated with the landmarks in the private test set”<br>\n<strong>non-landmark image set</strong>, sub set of 5000 non-landmark images form test set of GLDv2.</p>\n<p><strong>Models to retrieve images</strong></p>\n<p>It’s very important to search related images from index image set accurately. I trained the efficientnet B5, B6, B7 and Resnet152 models according to the first place solution of Landmark Retrieval 2020 competition from <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a>, <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037</a>. I will only describe the differences of my training strategy, please refer to his great solution for details.<br>\n1.Train 448x448 images from GLD_v2c for 5-6 epochs.<br>\n2.Finetune the model of step1 on 448x448 images from GLD_v2x<br>\n3.Finetune the model of step2 on 512x512 images from GLD_v2x for 5-6 epoch. I used GLD_v2x instead of GLD_v2 all.<br>\n4.640x640 for 3-5 epochs<br>\n5.736x736 for 2-3 epochs</p>\n<p><strong>Loss function</strong>: arcface instead of adacos loss.<br>\n<strong>Optimizer</strong>: SGD(0.01, momentum=0.9, decay=1e-5) <br>\n<strong>Inference</strong>: extract embeddings by feeding 800x800 images.<br>\n<strong>Validate set</strong>: sample 200 images from GLDv2 test set as val set and all the ground truth images of Google Landmark Retrieval Competition 2019  as index dataset</p>\n<p>After replacing the model in baseline kernel <a href=\"https://www.kaggle.com/camaskew/host-baseline-example\" target=\"_blank\">https://www.kaggle.com/camaskew/host-baseline-example</a> from the host with trained efficientnet B7 model, <br>\nthe public and private score of B7 model are 0.5927/0.5582.</p>\n<p><strong>Validation strategy for Landmark Recognition Task</strong><br>\nThe val set part 1: the 1.3k landmark images from GLDv2 test set( exclude those not in 81k classes) .   <br>\nThe val set part 2: sample 2.7k images from GLD_v2x-GLD_v2c.</p>\n<p>The index image set for val set: all the images of related landmarks from GLD_v2c train set and sample some other images to get 200k images.</p>\n<p>This strategy is quite stable during the whole competition, but unfortunately, the private test set distribution is a little different from my local CV and public test set. I should have used all the GLD-v2x images to generate index image set as many landmark images are not included in GLD-v2c. </p>\n<p><strong>SuperPoint+SuperGlue+pydegensac</strong><br>\nThis combination is better than delf+kdtree+pydegensac.<br>\nThe scored improved from 0.5927/0.5582 to 0.6146/0.5756, which can be top 10 on leaderboard.</p>\n<p><strong>Post-Processing</strong></p>\n<p>Although I force myself not pay to much efforts on post-processing, finding magic or reverse engineering, post-processing is so important in this competition that I had to spend a lot of time to analyse the model results and design rules on the validation set, I tried rules from winner solutions of last year, the following are the effective rules:<br>\n1.Search top 3 non-landmark images from no-landmark image set for query a image, if the similarity of top3 &gt;0.3, then decrease the score of the query image.<br>\n2.If a landmark is predicted &gt;20 times in the test set, then treat all the images of that landmark as non-landmarks.</p>\n<p>As many features(ransac inliers, similarity to index images, similarity to non-landmark images…) can be used for determining whether an image is non-landmark or not, I developed a model which can be called re-rank model. As to re-rank, we can refer to this post:<br>\n<a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/159315\" target=\"_blank\">https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/159315</a></p>\n<p>After post-processing, the score of efficientnet B7 model improved from 0.6146/0.5756 to 0.6797/0.6301 which can be top-3 on leaderboard.<br>\n<strong>Ensemble</strong><br>\nMy final scored was achieve by ensemble efficientnet B7, B6, B5, Resnet152 models, which is not a big improvement considering the complexity and computation resources.</p>\n<p>repo:<a href=\"https://github.com/bestfitting/instance_level_recognition\" target=\"_blank\">https://github.com/bestfitting/instance_level_recognition</a><br>\n[Update] with paper attached</p>",
  "messages": [
    {
      "id": 1035342,
      "postDate": "2020-10-02T16:46:12.867Z",
      "content": "<p>Congratulations to all the winners, thanks to kaggle and Google for hosting Landmark Retrieval/Recognition competition. </p>\n<p>Here is my brief solution to this competition.</p>\n<p><strong>Definition</strong><br>\n<strong>GLD_v2c</strong>(cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.<br>\n<strong>GLD_v2x</strong>, in GLDv2, there are 3.2 million images belong to the 81k classes in GLD_v2c. I define these 3.2m images as GLD_v2x.<br>\n<strong>query image</strong>, the test image when submitting to kaggle or the image in validation set when validate locally.<br>\n<strong>Index image set</strong>, the images from train set, as described in Data page of this competition, “subset contains all of the training set images associated with the landmarks in the private test set”<br>\n<strong>non-landmark image set</strong>, sub set of 5000 non-landmark images form test set of GLDv2.</p>\n<p><strong>Models to retrieve images</strong></p>\n<p>It’s very important to search related images from index image set accurately. I trained the efficientnet B5, B6, B7 and Resnet152 models according to the first place solution of Landmark Retrieval 2020 competition from <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a>, <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037</a>. I will only describe the differences of my training strategy, please refer to his great solution for details.<br>\n1.Train 448x448 images from GLD_v2c for 5-6 epochs.<br>\n2.Finetune the model of step1 on 448x448 images from GLD_v2x<br>\n3.Finetune the model of step2 on 512x512 images from GLD_v2x for 5-6 epoch. I used GLD_v2x instead of GLD_v2 all.<br>\n4.640x640 for 3-5 epochs<br>\n5.736x736 for 2-3 epochs</p>\n<p><strong>Loss function</strong>: arcface instead of adacos loss.<br>\n<strong>Optimizer</strong>: SGD(0.01, momentum=0.9, decay=1e-5) <br>\n<strong>Inference</strong>: extract embeddings by feeding 800x800 images.<br>\n<strong>Validate set</strong>: sample 200 images from GLDv2 test set as val set and all the ground truth images of Google Landmark Retrieval Competition 2019  as index dataset</p>\n<p>After replacing the model in baseline kernel <a href=\"https://www.kaggle.com/camaskew/host-baseline-example\" target=\"_blank\">https://www.kaggle.com/camaskew/host-baseline-example</a> from the host with trained efficientnet B7 model, <br>\nthe public and private score of B7 model are 0.5927/0.5582.</p>\n<p><strong>Validation strategy for Landmark Recognition Task</strong><br>\nThe val set part 1: the 1.3k landmark images from GLDv2 test set( exclude those not in 81k classes) .   <br>\nThe val set part 2: sample 2.7k images from GLD_v2x-GLD_v2c.</p>\n<p>The index image set for val set: all the images of related landmarks from GLD_v2c train set and sample some other images to get 200k images.</p>\n<p>This strategy is quite stable during the whole competition, but unfortunately, the private test set distribution is a little different from my local CV and public test set. I should have used all the GLD-v2x images to generate index image set as many landmark images are not included in GLD-v2c. </p>\n<p><strong>SuperPoint+SuperGlue+pydegensac</strong><br>\nThis combination is better than delf+kdtree+pydegensac.<br>\nThe scored improved from 0.5927/0.5582 to 0.6146/0.5756, which can be top 10 on leaderboard.</p>\n<p><strong>Post-Processing</strong></p>\n<p>Although I force myself not pay to much efforts on post-processing, finding magic or reverse engineering, post-processing is so important in this competition that I had to spend a lot of time to analyse the model results and design rules on the validation set, I tried rules from winner solutions of last year, the following are the effective rules:<br>\n1.Search top 3 non-landmark images from no-landmark image set for query a image, if the similarity of top3 &gt;0.3, then decrease the score of the query image.<br>\n2.If a landmark is predicted &gt;20 times in the test set, then treat all the images of that landmark as non-landmarks.</p>\n<p>As many features(ransac inliers, similarity to index images, similarity to non-landmark images…) can be used for determining whether an image is non-landmark or not, I developed a model which can be called re-rank model. As to re-rank, we can refer to this post:<br>\n<a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/159315\" target=\"_blank\">https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/159315</a></p>\n<p>After post-processing, the score of efficientnet B7 model improved from 0.6146/0.5756 to 0.6797/0.6301 which can be top-3 on leaderboard.<br>\n<strong>Ensemble</strong><br>\nMy final scored was achieve by ensemble efficientnet B7, B6, B5, Resnet152 models, which is not a big improvement considering the complexity and computation resources.</p>\n<p>repo:<a href=\"https://github.com/bestfitting/instance_level_recognition\" target=\"_blank\">https://github.com/bestfitting/instance_level_recognition</a><br>\n[Update] with paper attached</p>",
      "rawMarkdown": "Congratulations to all the winners, thanks to kaggle and Google for hosting Landmark Retrieval/Recognition competition. \n\nHere is my brief solution to this competition.\n\n**Definition**\n**GLD_v2c**(cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.\n**GLD_v2x**, in GLDv2, there are 3.2 million images belong to the 81k classes in GLD_v2c. I define these 3.2m images as GLD_v2x.\n**query image**, the test image when submitting to kaggle or the image in validation set when validate locally.\n**Index image set**, the images from train set, as described in Data page of this competition, “subset contains all of the training set images associated with the landmarks in the private test set”\n**non-landmark image set**, sub set of 5000 non-landmark images form test set of GLDv2.\n\n**Models to retrieve images**\n\nIt’s very important to search related images from index image set accurately. I trained the efficientnet B5, B6, B7 and Resnet152 models according to the first place solution of Landmark Retrieval 2020 competition from @keetar, https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037. I will only describe the differences of my training strategy, please refer to his great solution for details.\n1.Train 448x448 images from GLD_v2c for 5-6 epochs.\n2.Finetune the model of step1 on 448x448 images from GLD_v2x\n3.Finetune the model of step2 on 512x512 images from GLD_v2x for 5-6 epoch. I used GLD_v2x instead of GLD_v2 all.\n4.640x640 for 3-5 epochs\n5.736x736 for 2-3 epochs\n \n**Loss function**: arcface instead of adacos loss.\n**Optimizer**: SGD(0.01, momentum=0.9, decay=1e-5) \n**Inference**: extract embeddings by feeding 800x800 images.\n**Validate set**: sample 200 images from GLDv2 test set as val set and all the ground truth images of Google Landmark Retrieval Competition 2019  as index dataset\n\nAfter replacing the model in baseline kernel https://www.kaggle.com/camaskew/host-baseline-example from the host with trained efficientnet B7 model, \nthe public and private score of B7 model are 0.5927/0.5582.\n\n**Validation strategy for Landmark Recognition Task**\nThe val set part 1: the 1.3k landmark images from GLDv2 test set( exclude those not in 81k classes) .   \nThe val set part 2: sample 2.7k images from GLD_v2x-GLD_v2c.\n\nThe index image set for val set: all the images of related landmarks from GLD_v2c train set and sample some other images to get 200k images.\n\nThis strategy is quite stable during the whole competition, but unfortunately, the private test set distribution is a little different from my local CV and public test set. I should have used all the GLD-v2x images to generate index image set as many landmark images are not included in GLD-v2c. \n\n\n**SuperPoint+SuperGlue+pydegensac**\nThis combination is better than delf+kdtree+pydegensac.\nThe scored improved from 0.5927/0.5582 to 0.6146/0.5756, which can be top 10 on leaderboard.\n\n**Post-Processing**\n\nAlthough I force myself not pay to much efforts on post-processing, finding magic or reverse engineering, post-processing is so important in this competition that I had to spend a lot of time to analyse the model results and design rules on the validation set, I tried rules from winner solutions of last year, the following are the effective rules:\n1.Search top 3 non-landmark images from no-landmark image set for query a image, if the similarity of top3 >0.3, then decrease the score of the query image.\n2.If a landmark is predicted >20 times in the test set, then treat all the images of that landmark as non-landmarks.\n\nAs many features(ransac inliers, similarity to index images, similarity to non-landmark images...) can be used for determining whether an image is non-landmark or not, I developed a model which can be called re-rank model. As to re-rank, we can refer to this post:\nhttps://www.kaggle.com/c/tweet-sentiment-extraction/discussion/159315\n\nAfter post-processing, the score of efficientnet B7 model improved from 0.6146/0.5756 to 0.6797/0.6301 which can be top-3 on leaderboard.\n**Ensemble**\nMy final scored was achieve by ensemble efficientnet B7, B6, B5, Resnet152 models, which is not a big improvement considering the complexity and computation resources.\n\nrepo:https://github.com/bestfitting/instance_level_recognition\n[Update] with paper attached",
      "votes": 97
    },
    {
      "id": 1035637,
      "postDate": "2020-10-02T22:37:03.380Z",
      "content": "<p>Congratulations on the strong finish, and also for getting that #1 spot back !</p>",
      "rawMarkdown": "Congratulations on the strong finish, and also for getting that #1 spot back !",
      "votes": 3
    },
    {
      "id": 1035584,
      "postDate": "2020-10-02T20:43:06.130Z",
      "content": "<p>great solution! I got almost same score as yours 0.6146/0.5756 with global model + superpoints+superglue (my training approach is similar to yours) The critical missing part is your Post-Processing idea in the end which gave the major boost on top of that. </p>",
      "rawMarkdown": "great solution! I got almost same score as yours 0.6146/0.5756 with global model + superpoints+superglue (my training approach is similar to yours) The critical missing part is your Post-Processing idea in the end which gave the major boost on top of that. ",
      "votes": 3,
      "replies": [
        {
          "id": 1035743,
          "postDate": "2020-10-03T03:28:31.593Z",
          "content": "<p>Same here ! We also got 0.613 or the combination of ResNext101 32x4d + superpoint + superglue. EffNet or ResNet doesn't matter that much as long as you train them right; proper post-processing was required for the prize spots 😄<br>\nThanks for sharing the solution and congrats to you <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> </p>",
          "rawMarkdown": "Same here ! We also got 0.613 or the combination of ResNext101 32x4d + superpoint + superglue. EffNet or ResNet doesn't matter that much as long as you train them right; proper post-processing was required for the prize spots 😄\nThanks for sharing the solution and congrats to you @bestfitting ",
          "votes": 2
        },
        {
          "id": 1037365,
          "postDate": "2020-10-04T23:04:36.403Z",
          "content": "<p><a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a>, <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> yes! Post-processing was even important in last year's competition as there are much more distractors ( with less than 2k landmark images and more than 100k distractors). </p>",
          "rawMarkdown": "@weimin, @andy2709 yes! Post-processing was even important in last year's competition as there are much more distractors ( with less than 2k landmark images and more than 100k distractors). ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1035360,
      "postDate": "2020-10-02T16:55:53.537Z",
      "content": "<p>Great solution! The battle on public leaderboard was so fierce and fun! Our solutions look in a grand-scheme very similar, except that we also used GLD_v2x as index set and this seems to really help as you suspect.</p>",
      "rawMarkdown": "Great solution! The battle on public leaderboard was so fierce and fun! Our solutions look in a grand-scheme very similar, except that we also used GLD_v2x as index set and this seems to really help as you suspect.",
      "votes": 4,
      "replies": [
        {
          "id": 1035367,
          "postDate": "2020-10-02T17:11:49.970Z",
          "content": "<p>Congrats to you and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, you've done a very good job. I quite enjoy the \"battle\" and satisfied with the result.</p>\n<p>Yes, the usage of the full-version of GLDv2 is the key to this competition indeed, as there are so many landmark images were mis-cleaned from GLDv2 dataset.</p>\n<p>It's a pity that I analysed the rules of generating the GLDv2 clean dataset, and I quite sure I should use all the images when training but missed the kNN search part.</p>",
          "rawMarkdown": "Congrats to you and @christofhenkel, you've done a very good job. I quite enjoy the \"battle\" and satisfied with the result.\n\nYes, the usage of the full-version of GLDv2 is the key to this competition indeed, as there are so many landmark images were mis-cleaned from GLDv2 dataset.\n\nIt's a pity that I analysed the rules of generating the GLDv2 clean dataset, and I quite sure I should use all the images when training but missed the kNN search part.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1052126,
      "postDate": "2020-10-17T10:56:24.713Z",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> Excellent solution.<br>\nCan i have question to you <br>\nWhat is hardware configuration? and what time it took for you to complete on model</p>",
      "rawMarkdown": "Congratulation @bestfitting Excellent solution.\nCan i have question to you \nWhat is hardware configuration? and what time it took for you to complete on model",
      "votes": 1,
      "replies": [
        {
          "id": 1052608,
          "postDate": "2020-10-18T02:09:44.183Z",
          "content": "<p>I mentioned here:<br>\n<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299#1037363\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299#1037363</a></p>",
          "rawMarkdown": "I mentioned here:\nhttps://www.kaggle.com/c/landmark-recognition-2020/discussion/188299#1037363"
        }
      ]
    },
    {
      "id": 1043704,
      "postDate": "2020-10-09T07:05:44.687Z",
      "content": "<p>Great work, congratulations!👍</p>",
      "rawMarkdown": "Great work, congratulations!👍",
      "votes": 1
    },
    {
      "id": 1037017,
      "postDate": "2020-10-04T14:07:01.583Z",
      "content": "<p>Lots of things to learn, thanks for sharing! </p>",
      "rawMarkdown": "Lots of things to learn, thanks for sharing! ",
      "votes": 1
    },
    {
      "id": 1035938,
      "postDate": "2020-10-03T09:01:47.220Z",
      "content": "<p>Great work, congratulations! :)</p>",
      "rawMarkdown": "Great work, congratulations! :)",
      "votes": 1
    },
    {
      "id": 1035376,
      "postDate": "2020-10-02T17:24:10.460Z",
      "content": "<p>Thank you sharing your approach. It's great to see you back in action :) Your posts are inspirations to many kagglers like me. I wonder why did you choose 800x800 as your inference image size, given that it was not among the image size in your multi-scale training of models? Did you try to do inference with 736x736 image size?</p>",
      "rawMarkdown": "Thank you sharing your approach. It's great to see you back in action :) Your posts are inspirations to many kagglers like me. I wonder why did you choose 800x800 as your inference image size, given that it was not among the image size in your multi-scale training of models? Did you try to do inference with 736x736 image size?",
      "votes": 1,
      "replies": [
        {
          "id": 1035397,
          "postDate": "2020-10-02T17:34:06.093Z",
          "content": "<p>Inference on larger images than training worked well here.</p>",
          "rawMarkdown": "Inference on larger images than training worked well here.",
          "votes": 3
        },
        {
          "id": 1035416,
          "postDate": "2020-10-02T17:42:06.153Z",
          "content": "<p>800x800 is a little better. <br>\nTrain on small images and inference on large images is a useful trick when we want to make use of the details of an image.</p>",
          "rawMarkdown": "800x800 is a little better. \nTrain on small images and inference on large images is a useful trick when we want to make use of the details of an image.\n\n",
          "votes": 7
        }
      ]
    },
    {
      "id": 1035401,
      "postDate": "2020-10-02T17:35:24.790Z",
      "content": "<p>Saw your name in the competition, felt like when Mr Bolt set foot on the 100m track, fantastic progress, speed to the top and the solution was top notch. Thanks for the show!</p>",
      "rawMarkdown": "Saw your name in the competition, felt like when Mr Bolt set foot on the 100m track, fantastic progress, speed to the top and the solution was top notch. Thanks for the show!",
      "votes": 2,
      "replies": [
        {
          "id": 1035433,
          "postDate": "2020-10-02T17:50:16.973Z",
          "content": "<p>Thanks for your kind word, but I learned so much from kaggle community, without  <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> 's sharing, I can not train a good retrieval model. So, let's compete and learn from each other.</p>",
          "rawMarkdown": "Thanks for your kind word, but I learned so much from kaggle community, without  @keetar 's sharing, I can not train a good retrieval model. So, let's compete and learn from each other.",
          "votes": 2
        },
        {
          "id": 1035580,
          "postDate": "2020-10-02T20:25:30.810Z",
          "content": "<p>Yeah, I was feeling the same :)) </p>\n<ul>\n<li>Beginning of September: <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> joined the game <br>\nMe: interesting, but ok</li>\n<li>One week later: <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> is in the top 10<br>\nMe: holy……………</li>\n</ul>\n<p>By the way, sems like after this competition you will be <strong>top 1</strong> on Kaggle's ranking again. So congratulations for that also 🎉🎉🎉</p>",
          "rawMarkdown": "Yeah, I was feeling the same :)) \n- Beginning of September: @bestfitting joined the game \n   Me: interesting, but ok\n- One week later: @bestfitting is in the top 10\n   Me: holy...............\n\nBy the way, sems like after this competition you will be **top 1** on Kaggle's ranking again. So congratulations for that also 🎉🎉🎉",
          "votes": 3
        }
      ]
    },
    {
      "id": 1036262,
      "postDate": "2020-10-03T15:30:49.880Z",
      "content": "<p>Thank you. Could I please ask:</p>\n<ul>\n<li>How was your hardware setup like?</li>\n<li>How long did it take to train your models?</li>\n<li>Any chance you’d be willing to share code?</li>\n</ul>",
      "rawMarkdown": "Thank you. Could I please ask:\n\n- How was your hardware setup like?\n- How long did it take to train your models?\n- Any chance you’d be willing to share code?",
      "replies": [
        {
          "id": 1037363,
          "postDate": "2020-10-04T22:57:42.733Z",
          "content": "<p>I struggled with the big dataset and tried hard to feed larger images to the models during the whole period of competition since I entered (~40 days) on my server with 8 GPUs, I also tried Colab Pro, as the winner of  Landmark Retrieval Competition 2020 trained similar models successfully on it. </p>",
          "rawMarkdown": "I struggled with the big dataset and tried hard to feed larger images to the models during the whole period of competition since I entered (~40 days) on my server with 8 GPUs, I also tried Colab Pro, as the winner of  Landmark Retrieval Competition 2020 trained similar models successfully on it. \n",
          "votes": 3
        },
        {
          "id": 1037649,
          "postDate": "2020-10-05T07:47:28.573Z",
          "content": "<p>Thank you :)</p>",
          "rawMarkdown": "Thank you :)"
        }
      ]
    },
    {
      "id": 1467341,
      "postDate": "2021-08-11T23:57:34.747Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1035751,
      "postDate": "2020-10-03T03:43:35.123Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.",
      "votes": 1
    },
    {
      "id": 1485308,
      "postDate": "2021-08-22T01:39:42.410Z",
      "content": "<p>Great work. Thanks for sharing</p>",
      "rawMarkdown": "Great work. Thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 1035637,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-10-02T22:37:03.380000",
      "content": "<p>Congratulations on the strong finish, and also for getting that #1 spot back !</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1035584,
      "author_name": "Weimin Wang",
      "author_url": "",
      "post_date": "2020-10-02T20:43:06.130000",
      "content": "<p>great solution! I got almost same score as yours 0.6146/0.5756 with global model + superpoints+superglue (my training approach is similar to yours) The critical missing part is your Post-Processing idea in the end which gave the major boost on top of that. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1035743,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2020-10-03T03:28:31.593000",
          "content": "<p>Same here ! We also got 0.613 or the combination of ResNext101 32x4d + superpoint + superglue. EffNet or ResNet doesn't matter that much as long as you train them right; proper post-processing was required for the prize spots 😄<br>\nThanks for sharing the solution and congrats to you <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1037365,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2020-10-04T23:04:36.403000",
          "content": "<p><a href=\"https://www.kaggle.com/weimin\" target=\"_blank\">@weimin</a>, <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> yes! Post-processing was even important in last year's competition as there are much more distractors ( with less than 2k landmark images and more than 100k distractors). </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1035360,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-10-02T16:55:53.537000",
      "content": "<p>Great solution! The battle on public leaderboard was so fierce and fun! Our solutions look in a grand-scheme very similar, except that we also used GLD_v2x as index set and this seems to really help as you suspect.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1035367,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2020-10-02T17:11:49.970000",
          "content": "<p>Congrats to you and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, you've done a very good job. I quite enjoy the \"battle\" and satisfied with the result.</p>\n<p>Yes, the usage of the full-version of GLDv2 is the key to this competition indeed, as there are so many landmark images were mis-cleaned from GLDv2 dataset.</p>\n<p>It's a pity that I analysed the rules of generating the GLDv2 clean dataset, and I quite sure I should use all the images when training but missed the kNN search part.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1052126,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-17T10:56:24.713000",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> Excellent solution.<br>\nCan i have question to you <br>\nWhat is hardware configuration? and what time it took for you to complete on model</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1052608,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2020-10-18T02:09:44.183000",
          "content": "<p>I mentioned here:<br>\n<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299#1037363\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299#1037363</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1043704,
      "author_name": "HangWang",
      "author_url": "",
      "post_date": "2020-10-09T07:05:44.687000",
      "content": "<p>Great work, congratulations!👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1037017,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-10-04T14:07:01.583000",
      "content": "<p>Lots of things to learn, thanks for sharing! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1035938,
      "author_name": "keetar",
      "author_url": "",
      "post_date": "2020-10-03T09:01:47.220000",
      "content": "<p>Great work, congratulations! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1035376,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-10-02T17:24:10.460000",
      "content": "<p>Thank you sharing your approach. It's great to see you back in action :) Your posts are inspirations to many kagglers like me. I wonder why did you choose 800x800 as your inference image size, given that it was not among the image size in your multi-scale training of models? Did you try to do inference with 736x736 image size?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1035397,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-10-02T17:34:06.093000",
          "content": "<p>Inference on larger images than training worked well here.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1035416,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2020-10-02T17:42:06.153000",
          "content": "<p>800x800 is a little better. <br>\nTrain on small images and inference on large images is a useful trick when we want to make use of the details of an image.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1035401,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-10-02T17:35:24.790000",
      "content": "<p>Saw your name in the competition, felt like when Mr Bolt set foot on the 100m track, fantastic progress, speed to the top and the solution was top notch. Thanks for the show!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1035433,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2020-10-02T17:50:16.973000",
          "content": "<p>Thanks for your kind word, but I learned so much from kaggle community, without  <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> 's sharing, I can not train a good retrieval model. So, let's compete and learn from each other.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1035580,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2020-10-02T20:25:30.810000",
          "content": "<p>Yeah, I was feeling the same :)) </p>\n<ul>\n<li>Beginning of September: <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> joined the game <br>\nMe: interesting, but ok</li>\n<li>One week later: <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> is in the top 10<br>\nMe: holy……………</li>\n</ul>\n<p>By the way, sems like after this competition you will be <strong>top 1</strong> on Kaggle's ranking again. So congratulations for that also 🎉🎉🎉</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1036262,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-10-03T15:30:49.880000",
      "content": "<p>Thank you. Could I please ask:</p>\n<ul>\n<li>How was your hardware setup like?</li>\n<li>How long did it take to train your models?</li>\n<li>Any chance you’d be willing to share code?</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 1037363,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2020-10-04T22:57:42.733000",
          "content": "<p>I struggled with the big dataset and tried hard to feed larger images to the models during the whole period of competition since I entered (~40 days) on my server with 8 GPUs, I also tried Colab Pro, as the winner of  Landmark Retrieval Competition 2020 trained similar models successfully on it. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1037649,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-10-05T07:47:28.573000",
          "content": "<p>Thank you :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1467341,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-11T23:57:34.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1035751,
      "author_name": "Vandit",
      "author_url": "",
      "post_date": "2020-10-03T03:43:35.123000",
      "content": "<p>Thank you for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1485308,
      "author_name": "Rakesh Jarupula",
      "author_url": "",
      "post_date": "2021-08-22T01:39:42.410000",
      "content": "<p>Great work. Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1035342": "Congratulations to all the winners, thanks to kaggle and Google for hosting Landmark Retrieval/Recognition competition. \n\nHere is my brief solution to this competition.\n\n**Definition**\n**GLD_v2c**(cleaned GLDv2), there are 1.6 million training images and 81k classes. All landmark test images belong to these classes.\n**GLD_v2x**, in GLDv2, there are 3.2 million images belong to the 81k classes in GLD_v2c. I define these 3.2m images as GLD_v2x.\n**query image**, the test image when submitting to kaggle or the image in validation set when validate locally.\n**Index image set**, the images from train set, as described in Data page of this competition, “subset contains all of the training set images associated with the landmarks in the private test set”\n**non-landmark image set**, sub set of 5000 non-landmark images form test set of GLDv2.\n\n**Models to retrieve images**\n\nIt’s very important to search related images from index image set accurately. I trained the efficientnet B5, B6, B7 and Resnet152 models according to the first place solution of Landmark Retrieval 2020 competition from @keetar, https://www.kaggle.com/c/landmark-retrieval-2020/discussion/176037. I will only describe the differences of my training strategy, please refer to his great solution for details.\n1.Train 448x448 images from GLD_v2c for 5-6 epochs.\n2.Finetune the model of step1 on 448x448 images from GLD_v2x\n3.Finetune the model of step2 on 512x512 images from GLD_v2x for 5-6 epoch. I used GLD_v2x instead of GLD_v2 all.\n4.640x640 for 3-5 epochs\n5.736x736 for 2-3 epochs\n \n**Loss function**: arcface instead of adacos loss.\n**Optimizer**: SGD(0.01, momentum=0.9, decay=1e-5) \n**Inference**: extract embeddings by feeding 800x800 images.\n**Validate set**: sample 200 images from GLDv2 test set as val set and all the ground truth images of Google Landmark Retrieval Competition 2019  as index dataset\n\nAfter replacing the model in baseline kernel https://www.kaggle.com/camaskew/host-baseline-example from the host with trained efficientnet B7 model, \nthe public and private score of B7 model are 0.5927/0.5582.\n\n**Validation strategy for Landmark Recognition Task**\nThe val set part 1: the 1.3k landmark images from GLDv2 test set( exclude those not in 81k classes) .   \nThe val set part 2: sample 2.7k images from GLD_v2x-GLD_v2c.\n\nThe index image set for val set: all the images of related landmarks from GLD_v2c train set and sample some other images to get 200k images.\n\nThis strategy is quite stable during the whole competition, but unfortunately, the private test set distribution is a little different from my local CV and public test set. I should have used all the GLD-v2x images to generate index image set as many landmark images are not included in GLD-v2c. \n\n\n**SuperPoint+SuperGlue+pydegensac**\nThis combination is better than delf+kdtree+pydegensac.\nThe scored improved from 0.5927/0.5582 to 0.6146/0.5756, which can be top 10 on leaderboard.\n\n**Post-Processing**\n\nAlthough I force myself not pay to much efforts on post-processing, finding magic or reverse engineering, post-processing is so important in this competition that I had to spend a lot of time to analyse the model results and design rules on the validation set, I tried rules from winner solutions of last year, the following are the effective rules:\n1.Search top 3 non-landmark images from no-landmark image set for query a image, if the similarity of top3 >0.3, then decrease the score of the query image.\n2.If a landmark is predicted >20 times in the test set, then treat all the images of that landmark as non-landmarks.\n\nAs many features(ransac inliers, similarity to index images, similarity to non-landmark images...) can be used for determining whether an image is non-landmark or not, I developed a model which can be called re-rank model. As to re-rank, we can refer to this post:\nhttps://www.kaggle.com/c/tweet-sentiment-extraction/discussion/159315\n\nAfter post-processing, the score of efficientnet B7 model improved from 0.6146/0.5756 to 0.6797/0.6301 which can be top-3 on leaderboard.\n**Ensemble**\nMy final scored was achieve by ensemble efficientnet B7, B6, B5, Resnet152 models, which is not a big improvement considering the complexity and computation resources.\n\nrepo:https://github.com/bestfitting/instance_level_recognition\n[Update] with paper attached",
    "1035637": "Congratulations on the strong finish, and also for getting that #1 spot back !",
    "1035584": "great solution! I got almost same score as yours 0.6146/0.5756 with global model + superpoints+superglue (my training approach is similar to yours) The critical missing part is your Post-Processing idea in the end which gave the major boost on top of that. ",
    "1035360": "Great solution! The battle on public leaderboard was so fierce and fun! Our solutions look in a grand-scheme very similar, except that we also used GLD_v2x as index set and this seems to really help as you suspect.",
    "1052126": "Congratulation @bestfitting Excellent solution.\nCan i have question to you \nWhat is hardware configuration? and what time it took for you to complete on model",
    "1043704": "Great work, congratulations!👍",
    "1037017": "Lots of things to learn, thanks for sharing! ",
    "1035938": "Great work, congratulations! :)",
    "1035376": "Thank you sharing your approach. It's great to see you back in action :) Your posts are inspirations to many kagglers like me. I wonder why did you choose 800x800 as your inference image size, given that it was not among the image size in your multi-scale training of models? Did you try to do inference with 736x736 image size?",
    "1035401": "Saw your name in the competition, felt like when Mr Bolt set foot on the 100m track, fantastic progress, speed to the top and the solution was top notch. Thanks for the show!",
    "1036262": "Thank you. Could I please ask:\n\n- How was your hardware setup like?\n- How long did it take to train your models?\n- Any chance you’d be willing to share code?",
    "1467341": "",
    "1035751": "Thank you for sharing.",
    "1485308": "Great work. Thanks for sharing"
  }
}