{
  "id": 277273,
  "title": "2nd Place Solution",
  "url": "/competitions/landmark-retrieval-2021/writeups/zhangwesley-2nd-place-solution",
  "author_name": "",
  "post_date": "2021-10-11T08:02:58.530Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Code: <a href=\"https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution\" target=\"_blank\">https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution</a><br>\nPaper:  <br>\n<a href=\"https://arxiv.org/abs/2110.04294\" target=\"_blank\">https://arxiv.org/abs/2110.04294</a>       OR<br>\n<a href=\"https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/blob/master/ILR2021_2nd_solution.pdf\" target=\"_blank\">https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/blob/master/ILR2021_2nd_solution.pdf</a>. <br>\nSubmission Notebook: <a href=\"https://www.kaggle.com/zhangwesley/ilr2021-retrieval-2nd-solution\" target=\"_blank\">https://www.kaggle.com/zhangwesley/ilr2021-retrieval-2nd-solution</a></p>\n<h1>Introduction</h1>\n<p>Image retrieval is a very important computer vision task which aims at finding images similar to the query image. It is different from instance-level retrieval. Image retrieval aims to retrieve objects holding the same appearance with the query, even they are not the same instance. This makes the task more easier compared to instance-level retrieval. <br>\nOn the other hand, it is different from the fine-grained level image retrieval. The fine-grained level image retrieval pays more attention on the local attentions to discover more details due to its small intra-category variance, such as person re-identification.<br>\nLandmark retrieval is an instance-level retrieval task, which aims to search the same landmark from a large candidate set. In this paper, we will introduce our techniques used in the fourth landmark retrieval competition, Google Landmark Retrieval 2021 held on Kaggle. Some of them are inspired from the state of the art algorithms in person re-identification.</p>\n<p>Besides, we also involve many techniques that are commonly used in previous competitions including model structures training strategies, loss functions. These techniques have been well explored and introduced in previous competitions, so we only introduce our methods, denoted as new contributions listed below.</p>\n<ul>\n<li>We involve bags-of-tricks from person re-identification and conduct careful experiments on these tricks.</li>\n<li>We propose a continent-aware sampler to balance the distribution of training images based on their continent tags.</li>\n<li>We design a Landmark-Country aware reranking algorithm and integrate it with the K-reciprocal reranking method.</li>\n</ul>\n<h1>Method and experiments</h1>\n<h2>Train and validation set</h2>\n<p>The official GLDv2 dataset has provided a clean and a full version. As pointed out by previous works, many noisy images with the same IDs as clean set have been filtered out during the cleaning stage. Expanding these noisy data into clean set results in ‘c2x‘. Although dis-similar, these noisy images may contain very valuable information, e.g., the in- door or outdoor of a building. What’s more, we also include the index set from GLDv2, which shares many common ids with trainfull. We list the dataset below for a better understanding.</p>\n<table>\n<thead>\n<tr>\n<th>Trainset</th>\n<th># Samples</th>\n<th># Labels</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Clean</td>\n<td>1,580,470</td>\n<td>81,313</td>\n</tr>\n<tr>\n<td>C2x</td>\n<td>3,223,078</td>\n<td>81,313</td>\n</tr>\n<tr>\n<td>Trainfull</td>\n<td>4,132,914</td>\n<td>203,094</td>\n</tr>\n<tr>\n<td>All</td>\n<td>4,825,830</td>\n<td>203,094</td>\n</tr>\n</tbody>\n</table>\n<p>We use the 1129 GLDv2 test set together with the 76,176 index set from the competition. We put all GT images of each query into the index set and expand it to 78,959 im- ages. In this case, any query could find all its GTs in the index set.</p>\n<h2>Baseline network</h2>\n<p>We select several large CNN networks including SE-ResNet-101, ResNeXt-101, ResNeSt101 and ResNeSt269 as backbones. IBN extension is used for SE-ResNet and ResNeXt-101. The input size is selected as 384 for pretraining and 512 for the last fine-tuning. The last stride of the CNN network is set to 1.<br>\nWe use generalized mean-pooling (GeM) for pooling method with p=3.0. Arcface loss with scale=30 and margin=0.3 is used. We use weight decay=0.0005. Training details with gradually enlarging input size and data scale can be found in implementation details.</p>\n<p>Some tricks from person re-identification has been explored.  Since the dataset is very large, we use R50 backbone with an in- put size of 256 × 256. We list the validation accuracy as well as public/private scores. Random Erasing randomly erases out image patches and has shown great success in many fields. Label smoothing by using soft targets that are a weighted average of the hard targets can often be useful in many computer vision tasks. From the table below, ransom erasing proves to be effective while label smoothing fails.</p>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Validation</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>32.60</td>\n<td>28.44</td>\n<td>30.59</td>\n</tr>\n<tr>\n<td>+RE</td>\n<td>32.78</td>\n<td>29.29</td>\n<td>30.60</td>\n</tr>\n<tr>\n<td>+label smooth</td>\n<td>32.55</td>\n<td>28.21</td>\n<td>29.79</td>\n</tr>\n</tbody>\n</table>\n<h2>Sampling Strategy</h2>\n<p>On one hand, in person re-identification or face recognition, id-uniform is widely used as a data sampling strategy. For a batch, we randomly select P Ids and then K images for each Id. Thus we have P*K images as a batch. Each id is treated fair for this setting. On the other hand, softmax sampling has been widely used in previous competitions. The softmax sampling just shuffle all dataset once at the beginning of the epoch and then samples iterative through the data. Head data which appears more will be put more attention with this setting.<br>\nWe have tried these sampling strategies and find neither of them are good enough for our task.</p>\n<p>As the provided landmark dataset pays more concentration on Asia landmarks, we manage to design a sampling strategy based on their continent labels. <br>\nFirst, we use the country-and-continent-codes-list to find how many countries each continent has. Then we search and list all the landmarks in each country. Based on this processing, we can get the country tag and continent tag for every landmark.</p>\n<p>We setup a continent sampling prob by {'Asia': 0.5, 'Europe': 0.2, 'Africa': 0.15, 'North America': 0.1, 'South America': 0.02, 'Antarctica': 0.01, 'Oceania': 0.01, 'OTHER': 0.01}. For an epoch of images, we sample continent images by the corresponding ratios. Also, we learn from previous paper which set 0.66 probability for clean data and 0.33 probability for noisy data.</p>\n<p>We list results for different sampling strategy below. The widely used id-uniform strategy fails for landmark retrieval. We think the reason is due to the large amount of tail data. These tail data may be noisy and thus degrades the model training. The softmax strategy works better than id- uniform. For continent-aware strategy, it achieves the best performance.</p>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Validation</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Id-uniform</td>\n<td>31.05</td>\n<td>24.78</td>\n<td>27.28</td>\n</tr>\n<tr>\n<td>Softmax</td>\n<td>32.60</td>\n<td>28.44</td>\n<td>30.59</td>\n</tr>\n<tr>\n<td>Continent-aware</td>\n<td>33.07</td>\n<td>31.37</td>\n<td>32.44</td>\n</tr>\n</tbody>\n</table>\n<h2>Reranking</h2>\n<p>Reranking is very essential to the final performance. Besides K-reciprocal reranking, a Landmark-Country aware reranking algorithm is specially designed for our task.</p>\n<p>We observe that there are many images which contain the same landmark and can't be easily retrieved by visual features due to great variations caused by views and illumination. Considering this, a Landmark-Country aware reranking is proposed by taking fully use of the training set. Specifically, as each image in training set has its landmark tag and its country tag, we first assign query and index images in the testing set with a training tag, and then retrieval the query from index images by the assigned tags. </p>\n<p>For query image tagging, We give a list of potential landmark tags and country tags to every query image according to its top K similar images in training set, and each landmark tag and country tag is scored by the accumulation of similarity in top k . With the Landmark-Country aware reranking, images have the same landmark tag or country tag with the potential tag of query are advanced in the retrieval sorting. An illustration can be found below and more details and equations are listed in paper.</p>\n<p><img src=\"https://raw.githubusercontent.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/master/figs/rerank_vis.jpg\" alt=\"rerank\"></p>\n<h2>Implementation details</h2>\n<p>Motivated by previous solutions, we first train the models on ‘clean‘ subset with an input size of 384(For ResNeSt269, the size is 448). The initial learning rate is 0.01 and we train for 10 epochs with the first epoch as warmup. Then we keep the same category number but with more data as ‘c2x‘. The initial learning rate is 0.001 and we train for 6 epochs. Then we expand dataset to ‘trainfull‘ and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512. At last, we use ‘all‘ data from GLDv2 and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512.</p>\n<h1>Conclusion</h1>\n<p>In this paper, we import techniques and bag-of-tricks from person re-identification. For the specific task of landmark retrieval, we propose continent-aware sampling strategy and Landmark-Country aware post processing, which has proven to be very effective on the private leaderboard. The relationship between landmark retrieval and landmark recognition would be studied in the future.</p>",
  "messages": [
    {
      "id": "1538725",
      "postDate": "10/08/2021 16:58:46",
      "content": "<p>Code: <a href=\"https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution\" target=\"_blank\">https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution</a><br>\nPaper:  <br>\n<a href=\"https://arxiv.org/abs/2110.04294\" target=\"_blank\">https://arxiv.org/abs/2110.04294</a>       OR<br>\n<a href=\"https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/blob/master/ILR2021_2nd_solution.pdf\" target=\"_blank\">https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/blob/master/ILR2021_2nd_solution.pdf</a>. <br>\nSubmission Notebook: <a href=\"https://www.kaggle.com/zhangwesley/ilr2021-retrieval-2nd-solution\" target=\"_blank\">https://www.kaggle.com/zhangwesley/ilr2021-retrieval-2nd-solution</a></p>\n<h1>Introduction</h1>\n<p>Image retrieval is a very important computer vision task which aims at finding images similar to the query image. It is different from instance-level retrieval. Image retrieval aims to retrieve objects holding the same appearance with the query, even they are not the same instance. This makes the task more easier compared to instance-level retrieval. <br>\nOn the other hand, it is different from the fine-grained level image retrieval. The fine-grained level image retrieval pays more attention on the local attentions to discover more details due to its small intra-category variance, such as person re-identification.<br>\nLandmark retrieval is an instance-level retrieval task, which aims to search the same landmark from a large candidate set. In this paper, we will introduce our techniques used in the fourth landmark retrieval competition, Google Landmark Retrieval 2021 held on Kaggle. Some of them are inspired from the state of the art algorithms in person re-identification.</p>\n<p>Besides, we also involve many techniques that are commonly used in previous competitions including model structures training strategies, loss functions. These techniques have been well explored and introduced in previous competitions, so we only introduce our methods, denoted as new contributions listed below.</p>\n<ul>\n<li>We involve bags-of-tricks from person re-identification and conduct careful experiments on these tricks.</li>\n<li>We propose a continent-aware sampler to balance the distribution of training images based on their continent tags.</li>\n<li>We design a Landmark-Country aware reranking algorithm and integrate it with the K-reciprocal reranking method.</li>\n</ul>\n<h1>Method and experiments</h1>\n<h2>Train and validation set</h2>\n<p>The official GLDv2 dataset has provided a clean and a full version. As pointed out by previous works, many noisy images with the same IDs as clean set have been filtered out during the cleaning stage. Expanding these noisy data into clean set results in ‘c2x‘. Although dis-similar, these noisy images may contain very valuable information, e.g., the in- door or outdoor of a building. What’s more, we also include the index set from GLDv2, which shares many common ids with trainfull. We list the dataset below for a better understanding.</p>\n<table>\n<thead>\n<tr>\n<th>Trainset</th>\n<th># Samples</th>\n<th># Labels</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Clean</td>\n<td>1,580,470</td>\n<td>81,313</td>\n</tr>\n<tr>\n<td>C2x</td>\n<td>3,223,078</td>\n<td>81,313</td>\n</tr>\n<tr>\n<td>Trainfull</td>\n<td>4,132,914</td>\n<td>203,094</td>\n</tr>\n<tr>\n<td>All</td>\n<td>4,825,830</td>\n<td>203,094</td>\n</tr>\n</tbody>\n</table>\n<p>We use the 1129 GLDv2 test set together with the 76,176 index set from the competition. We put all GT images of each query into the index set and expand it to 78,959 im- ages. In this case, any query could find all its GTs in the index set.</p>\n<h2>Baseline network</h2>\n<p>We select several large CNN networks including SE-ResNet-101, ResNeXt-101, ResNeSt101 and ResNeSt269 as backbones. IBN extension is used for SE-ResNet and ResNeXt-101. The input size is selected as 384 for pretraining and 512 for the last fine-tuning. The last stride of the CNN network is set to 1.<br>\nWe use generalized mean-pooling (GeM) for pooling method with p=3.0. Arcface loss with scale=30 and margin=0.3 is used. We use weight decay=0.0005. Training details with gradually enlarging input size and data scale can be found in implementation details.</p>\n<p>Some tricks from person re-identification has been explored.  Since the dataset is very large, we use R50 backbone with an in- put size of 256 × 256. We list the validation accuracy as well as public/private scores. Random Erasing randomly erases out image patches and has shown great success in many fields. Label smoothing by using soft targets that are a weighted average of the hard targets can often be useful in many computer vision tasks. From the table below, ransom erasing proves to be effective while label smoothing fails.</p>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Validation</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>32.60</td>\n<td>28.44</td>\n<td>30.59</td>\n</tr>\n<tr>\n<td>+RE</td>\n<td>32.78</td>\n<td>29.29</td>\n<td>30.60</td>\n</tr>\n<tr>\n<td>+label smooth</td>\n<td>32.55</td>\n<td>28.21</td>\n<td>29.79</td>\n</tr>\n</tbody>\n</table>\n<h2>Sampling Strategy</h2>\n<p>On one hand, in person re-identification or face recognition, id-uniform is widely used as a data sampling strategy. For a batch, we randomly select P Ids and then K images for each Id. Thus we have P*K images as a batch. Each id is treated fair for this setting. On the other hand, softmax sampling has been widely used in previous competitions. The softmax sampling just shuffle all dataset once at the beginning of the epoch and then samples iterative through the data. Head data which appears more will be put more attention with this setting.<br>\nWe have tried these sampling strategies and find neither of them are good enough for our task.</p>\n<p>As the provided landmark dataset pays more concentration on Asia landmarks, we manage to design a sampling strategy based on their continent labels. <br>\nFirst, we use the country-and-continent-codes-list to find how many countries each continent has. Then we search and list all the landmarks in each country. Based on this processing, we can get the country tag and continent tag for every landmark.</p>\n<p>We setup a continent sampling prob by {'Asia': 0.5, 'Europe': 0.2, 'Africa': 0.15, 'North America': 0.1, 'South America': 0.02, 'Antarctica': 0.01, 'Oceania': 0.01, 'OTHER': 0.01}. For an epoch of images, we sample continent images by the corresponding ratios. Also, we learn from previous paper which set 0.66 probability for clean data and 0.33 probability for noisy data.</p>\n<p>We list results for different sampling strategy below. The widely used id-uniform strategy fails for landmark retrieval. We think the reason is due to the large amount of tail data. These tail data may be noisy and thus degrades the model training. The softmax strategy works better than id- uniform. For continent-aware strategy, it achieves the best performance.</p>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Validation</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Id-uniform</td>\n<td>31.05</td>\n<td>24.78</td>\n<td>27.28</td>\n</tr>\n<tr>\n<td>Softmax</td>\n<td>32.60</td>\n<td>28.44</td>\n<td>30.59</td>\n</tr>\n<tr>\n<td>Continent-aware</td>\n<td>33.07</td>\n<td>31.37</td>\n<td>32.44</td>\n</tr>\n</tbody>\n</table>\n<h2>Reranking</h2>\n<p>Reranking is very essential to the final performance. Besides K-reciprocal reranking, a Landmark-Country aware reranking algorithm is specially designed for our task.</p>\n<p>We observe that there are many images which contain the same landmark and can't be easily retrieved by visual features due to great variations caused by views and illumination. Considering this, a Landmark-Country aware reranking is proposed by taking fully use of the training set. Specifically, as each image in training set has its landmark tag and its country tag, we first assign query and index images in the testing set with a training tag, and then retrieval the query from index images by the assigned tags. </p>\n<p>For query image tagging, We give a list of potential landmark tags and country tags to every query image according to its top K similar images in training set, and each landmark tag and country tag is scored by the accumulation of similarity in top k . With the Landmark-Country aware reranking, images have the same landmark tag or country tag with the potential tag of query are advanced in the retrieval sorting. An illustration can be found below and more details and equations are listed in paper.</p>\n<p><img src=\"https://raw.githubusercontent.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/master/figs/rerank_vis.jpg\" alt=\"rerank\"></p>\n<h2>Implementation details</h2>\n<p>Motivated by previous solutions, we first train the models on ‘clean‘ subset with an input size of 384(For ResNeSt269, the size is 448). The initial learning rate is 0.01 and we train for 10 epochs with the first epoch as warmup. Then we keep the same category number but with more data as ‘c2x‘. The initial learning rate is 0.001 and we train for 6 epochs. Then we expand dataset to ‘trainfull‘ and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512. At last, we use ‘all‘ data from GLDv2 and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512.</p>\n<h1>Conclusion</h1>\n<p>In this paper, we import techniques and bag-of-tricks from person re-identification. For the specific task of landmark retrieval, we propose continent-aware sampling strategy and Landmark-Country aware post processing, which has proven to be very effective on the private leaderboard. The relationship between landmark retrieval and landmark recognition would be studied in the future.</p>",
      "rawMarkdown": "Code: https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution\nPaper:  \nhttps://arxiv.org/abs/2110.04294       OR\nhttps://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/blob/master/ILR2021_2nd_solution.pdf. \nSubmission Notebook: https://www.kaggle.com/zhangwesley/ilr2021-retrieval-2nd-solution\n\n# Introduction\nImage retrieval is a very important computer vision task which aims at finding images similar to the query image. It is different from instance-level retrieval. Image retrieval aims to retrieve objects holding the same appearance with the query, even they are not the same instance. This makes the task more easier compared to instance-level retrieval. \nOn the other hand, it is different from the fine-grained level image retrieval. The fine-grained level image retrieval pays more attention on the local attentions to discover more details due to its small intra-category variance, such as person re-identification.\nLandmark retrieval is an instance-level retrieval task, which aims to search the same landmark from a large candidate set. In this paper, we will introduce our techniques used in the fourth landmark retrieval competition, Google Landmark Retrieval 2021 held on Kaggle. Some of them are inspired from the state of the art algorithms in person re-identification.\n\nBesides, we also involve many techniques that are commonly used in previous competitions including model structures training strategies, loss functions. These techniques have been well explored and introduced in previous competitions, so we only introduce our methods, denoted as new contributions listed below.\n- We involve bags-of-tricks from person re-identification and conduct careful experiments on these tricks.\n- We propose a continent-aware sampler to balance the distribution of training images based on their continent tags.\n- We design a Landmark-Country aware reranking algorithm and integrate it with the K-reciprocal reranking method.\n\n# Method and experiments\n## Train and validation set\nThe official GLDv2 dataset has provided a clean and a full version. As pointed out by previous works, many noisy images with the same IDs as clean set have been filtered out during the cleaning stage. Expanding these noisy data into clean set results in ‘c2x‘. Although dis-similar, these noisy images may contain very valuable information, e.g., the in- door or outdoor of a building. What’s more, we also include the index set from GLDv2, which shares many common ids with trainfull. We list the dataset below for a better understanding.\n\n| Trainset  | \\# Samples | \\# Labels |\n| :-------: | :--------: | :-------: |\n|   Clean   | 1,580,470  |  81,313   |\n|    C2x    | 3,223,078  |  81,313   |\n| Trainfull | 4,132,914  |  203,094  |\n|    All    | 4,825,830  |  203,094  |\n\nWe use the 1129 GLDv2 test set together with the 76,176 index set from the competition. We put all GT images of each query into the index set and expand it to 78,959 im- ages. In this case, any query could find all its GTs in the index set.\n\n\n## Baseline network\nWe select several large CNN networks including SE-ResNet-101, ResNeXt-101, ResNeSt101 and ResNeSt269 as backbones. IBN extension is used for SE-ResNet and ResNeXt-101. The input size is selected as 384 for pretraining and 512 for the last fine-tuning. The last stride of the CNN network is set to 1.\nWe use generalized mean-pooling (GeM) for pooling method with p=3.0. Arcface loss with scale=30 and margin=0.3 is used. We use weight decay=0.0005. Training details with gradually enlarging input size and data scale can be found in implementation details.\n\nSome tricks from person re-identification has been explored.  Since the dataset is very large, we use R50 backbone with an in- put size of 256 × 256. We list the validation accuracy as well as public/private scores. Random Erasing randomly erases out image patches and has shown great success in many fields. Label smoothing by using soft targets that are a weighted average of the hard targets can often be useful in many computer vision tasks. From the table below, ransom erasing proves to be effective while label smoothing fails.\n\n|    Setting    | Validation | Public | Private |\n| :-----------: | :--------: | :----: | :-----: |\n|   baseline    |   32.60    | 28.44  |  30.59  |\n|      +RE      |   32.78    | 29.29  |  30.60  |\n| +label smooth |   32.55    | 28.21  |  29.79  |\n\n\n##  Sampling Strategy\nOn one hand, in person re-identification or face recognition, id-uniform is widely used as a data sampling strategy. For a batch, we randomly select P Ids and then K images for each Id. Thus we have P*K images as a batch. Each id is treated fair for this setting. On the other hand, softmax sampling has been widely used in previous competitions. The softmax sampling just shuffle all dataset once at the beginning of the epoch and then samples iterative through the data. Head data which appears more will be put more attention with this setting.\nWe have tried these sampling strategies and find neither of them are good enough for our task.\n\nAs the provided landmark dataset pays more concentration on Asia landmarks, we manage to design a sampling strategy based on their continent labels. \nFirst, we use the country-and-continent-codes-list to find how many countries each continent has. Then we search and list all the landmarks in each country. Based on this processing, we can get the country tag and continent tag for every landmark.\n\nWe setup a continent sampling prob by {'Asia': 0.5, 'Europe': 0.2, 'Africa': 0.15, 'North America': 0.1, 'South America': 0.02, 'Antarctica': 0.01, 'Oceania': 0.01, 'OTHER': 0.01}. For an epoch of images, we sample continent images by the corresponding ratios. Also, we learn from previous paper which set 0.66 probability for clean data and 0.33 probability for noisy data.\n\n\nWe list results for different sampling strategy below. The widely used id-uniform strategy fails for landmark retrieval. We think the reason is due to the large amount of tail data. These tail data may be noisy and thus degrades the model training. The softmax strategy works better than id- uniform. For continent-aware strategy, it achieves the best performance.\n\n|     Setting     | Validation | Public | Private |\n| :-------------: | :--------: | :----: | :-----: |\n|   Id-uniform    |   31.05    | 24.78  |  27.28  |\n|     Softmax     |   32.60    | 28.44  |  30.59  |\n| Continent-aware |   33.07    | 31.37  |  32.44  |\n\n\n\n##  Reranking\nReranking is very essential to the final performance. Besides K-reciprocal reranking, a Landmark-Country aware reranking algorithm is specially designed for our task.\n\nWe observe that there are many images which contain the same landmark and can't be easily retrieved by visual features due to great variations caused by views and illumination. Considering this, a Landmark-Country aware reranking is proposed by taking fully use of the training set. Specifically, as each image in training set has its landmark tag and its country tag, we first assign query and index images in the testing set with a training tag, and then retrieval the query from index images by the assigned tags. \n\nFor query image tagging, We give a list of potential landmark tags and country tags to every query image according to its top K similar images in training set, and each landmark tag and country tag is scored by the accumulation of similarity in top k . With the Landmark-Country aware reranking, images have the same landmark tag or country tag with the potential tag of query are advanced in the retrieval sorting. An illustration can be found below and more details and equations are listed in paper.\n\n![rerank](https://raw.githubusercontent.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/master/figs/rerank_vis.jpg)\n\n##  Implementation details\nMotivated by previous solutions, we first train the models on ‘clean‘ subset with an input size of 384(For ResNeSt269, the size is 448). The initial learning rate is 0.01 and we train for 10 epochs with the first epoch as warmup. Then we keep the same category number but with more data as ‘c2x‘. The initial learning rate is 0.001 and we train for 6 epochs. Then we expand dataset to ‘trainfull‘ and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512. At last, we use ‘all‘ data from GLDv2 and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512.\n\n\n\n# Conclusion\nIn this paper, we import techniques and bag-of-tricks from person re-identification. For the specific task of landmark retrieval, we propose continent-aware sampling strategy and Landmark-Country aware post processing, which has proven to be very effective on the private leaderboard. The relationship between landmark retrieval and landmark recognition would be studied in the future.",
      "votes": null
    },
    {
      "id": "1539213",
      "postDate": "10/09/2021 07:43:14",
      "content": "<p>Congratulations.</p>",
      "rawMarkdown": "Congratulations.",
      "votes": null
    },
    {
      "id": "1540002",
      "postDate": "10/10/2021 04:36:05",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2315269",
      "postDate": "06/24/2023 01:41:18",
      "content": "<p>May I ask how to obtain this pkl file: 'GLDv2_search_label_competition_2021.pkl'</p>",
      "rawMarkdown": "May I ask how to obtain this pkl file: 'GLDv2_search_label_competition_2021.pkl'",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1539213,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "10/09/2021 07:43:14",
      "content": "<p>Congratulations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1540002,
      "author_name": "saleesh",
      "author_url": "",
      "post_date": "10/10/2021 04:36:05",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2315269,
      "author_name": "dafuni",
      "author_url": "",
      "post_date": "06/24/2023 01:41:18",
      "content": "<p>May I ask how to obtain this pkl file: 'GLDv2_search_label_competition_2021.pkl'</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1538725": "Code: https://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution\nPaper:  \nhttps://arxiv.org/abs/2110.04294       OR\nhttps://github.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/blob/master/ILR2021_2nd_solution.pdf. \nSubmission Notebook: https://www.kaggle.com/zhangwesley/ilr2021-retrieval-2nd-solution\n\n# Introduction\nImage retrieval is a very important computer vision task which aims at finding images similar to the query image. It is different from instance-level retrieval. Image retrieval aims to retrieve objects holding the same appearance with the query, even they are not the same instance. This makes the task more easier compared to instance-level retrieval. \nOn the other hand, it is different from the fine-grained level image retrieval. The fine-grained level image retrieval pays more attention on the local attentions to discover more details due to its small intra-category variance, such as person re-identification.\nLandmark retrieval is an instance-level retrieval task, which aims to search the same landmark from a large candidate set. In this paper, we will introduce our techniques used in the fourth landmark retrieval competition, Google Landmark Retrieval 2021 held on Kaggle. Some of them are inspired from the state of the art algorithms in person re-identification.\n\nBesides, we also involve many techniques that are commonly used in previous competitions including model structures training strategies, loss functions. These techniques have been well explored and introduced in previous competitions, so we only introduce our methods, denoted as new contributions listed below.\n- We involve bags-of-tricks from person re-identification and conduct careful experiments on these tricks.\n- We propose a continent-aware sampler to balance the distribution of training images based on their continent tags.\n- We design a Landmark-Country aware reranking algorithm and integrate it with the K-reciprocal reranking method.\n\n# Method and experiments\n## Train and validation set\nThe official GLDv2 dataset has provided a clean and a full version. As pointed out by previous works, many noisy images with the same IDs as clean set have been filtered out during the cleaning stage. Expanding these noisy data into clean set results in ‘c2x‘. Although dis-similar, these noisy images may contain very valuable information, e.g., the in- door or outdoor of a building. What’s more, we also include the index set from GLDv2, which shares many common ids with trainfull. We list the dataset below for a better understanding.\n\n| Trainset  | \\# Samples | \\# Labels |\n| :-------: | :--------: | :-------: |\n|   Clean   | 1,580,470  |  81,313   |\n|    C2x    | 3,223,078  |  81,313   |\n| Trainfull | 4,132,914  |  203,094  |\n|    All    | 4,825,830  |  203,094  |\n\nWe use the 1129 GLDv2 test set together with the 76,176 index set from the competition. We put all GT images of each query into the index set and expand it to 78,959 im- ages. In this case, any query could find all its GTs in the index set.\n\n\n## Baseline network\nWe select several large CNN networks including SE-ResNet-101, ResNeXt-101, ResNeSt101 and ResNeSt269 as backbones. IBN extension is used for SE-ResNet and ResNeXt-101. The input size is selected as 384 for pretraining and 512 for the last fine-tuning. The last stride of the CNN network is set to 1.\nWe use generalized mean-pooling (GeM) for pooling method with p=3.0. Arcface loss with scale=30 and margin=0.3 is used. We use weight decay=0.0005. Training details with gradually enlarging input size and data scale can be found in implementation details.\n\nSome tricks from person re-identification has been explored.  Since the dataset is very large, we use R50 backbone with an in- put size of 256 × 256. We list the validation accuracy as well as public/private scores. Random Erasing randomly erases out image patches and has shown great success in many fields. Label smoothing by using soft targets that are a weighted average of the hard targets can often be useful in many computer vision tasks. From the table below, ransom erasing proves to be effective while label smoothing fails.\n\n|    Setting    | Validation | Public | Private |\n| :-----------: | :--------: | :----: | :-----: |\n|   baseline    |   32.60    | 28.44  |  30.59  |\n|      +RE      |   32.78    | 29.29  |  30.60  |\n| +label smooth |   32.55    | 28.21  |  29.79  |\n\n\n##  Sampling Strategy\nOn one hand, in person re-identification or face recognition, id-uniform is widely used as a data sampling strategy. For a batch, we randomly select P Ids and then K images for each Id. Thus we have P*K images as a batch. Each id is treated fair for this setting. On the other hand, softmax sampling has been widely used in previous competitions. The softmax sampling just shuffle all dataset once at the beginning of the epoch and then samples iterative through the data. Head data which appears more will be put more attention with this setting.\nWe have tried these sampling strategies and find neither of them are good enough for our task.\n\nAs the provided landmark dataset pays more concentration on Asia landmarks, we manage to design a sampling strategy based on their continent labels. \nFirst, we use the country-and-continent-codes-list to find how many countries each continent has. Then we search and list all the landmarks in each country. Based on this processing, we can get the country tag and continent tag for every landmark.\n\nWe setup a continent sampling prob by {'Asia': 0.5, 'Europe': 0.2, 'Africa': 0.15, 'North America': 0.1, 'South America': 0.02, 'Antarctica': 0.01, 'Oceania': 0.01, 'OTHER': 0.01}. For an epoch of images, we sample continent images by the corresponding ratios. Also, we learn from previous paper which set 0.66 probability for clean data and 0.33 probability for noisy data.\n\n\nWe list results for different sampling strategy below. The widely used id-uniform strategy fails for landmark retrieval. We think the reason is due to the large amount of tail data. These tail data may be noisy and thus degrades the model training. The softmax strategy works better than id- uniform. For continent-aware strategy, it achieves the best performance.\n\n|     Setting     | Validation | Public | Private |\n| :-------------: | :--------: | :----: | :-----: |\n|   Id-uniform    |   31.05    | 24.78  |  27.28  |\n|     Softmax     |   32.60    | 28.44  |  30.59  |\n| Continent-aware |   33.07    | 31.37  |  32.44  |\n\n\n\n##  Reranking\nReranking is very essential to the final performance. Besides K-reciprocal reranking, a Landmark-Country aware reranking algorithm is specially designed for our task.\n\nWe observe that there are many images which contain the same landmark and can't be easily retrieved by visual features due to great variations caused by views and illumination. Considering this, a Landmark-Country aware reranking is proposed by taking fully use of the training set. Specifically, as each image in training set has its landmark tag and its country tag, we first assign query and index images in the testing set with a training tag, and then retrieval the query from index images by the assigned tags. \n\nFor query image tagging, We give a list of potential landmark tags and country tags to every query image according to its top K similar images in training set, and each landmark tag and country tag is scored by the accumulation of similarity in top k . With the Landmark-Country aware reranking, images have the same landmark tag or country tag with the potential tag of query are advanced in the retrieval sorting. An illustration can be found below and more details and equations are listed in paper.\n\n![rerank](https://raw.githubusercontent.com/WesleyZhang1991/Google_Landmark_Retrieval_2021_2nd_Place_Solution/master/figs/rerank_vis.jpg)\n\n##  Implementation details\nMotivated by previous solutions, we first train the models on ‘clean‘ subset with an input size of 384(For ResNeSt269, the size is 448). The initial learning rate is 0.01 and we train for 10 epochs with the first epoch as warmup. Then we keep the same category number but with more data as ‘c2x‘. The initial learning rate is 0.001 and we train for 6 epochs. Then we expand dataset to ‘trainfull‘ and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512. At last, we use ‘all‘ data from GLDv2 and train for another 2 epochs with the initial learning rate as 0.0001 and input size as 512.\n\n\n\n# Conclusion\nIn this paper, we import techniques and bag-of-tricks from person re-identification. For the specific task of landmark retrieval, we propose continent-aware sampling strategy and Landmark-Country aware post processing, which has proven to be very effective on the private leaderboard. The relationship between landmark retrieval and landmark recognition would be studied in the future.",
    "1539213": "Congratulations.",
    "1540002": "Thanks for sharing.",
    "2315269": "May I ask how to obtain this pkl file: 'GLDv2_search_label_competition_2021.pkl'"
  },
  "source": "meta"
}