{
  "id": 276632,
  "title": "6th place solution",
  "url": "/competitions/landmark-retrieval-2021/discussion/276632",
  "author_name": "Inoichan",
  "post_date": "2021-10-05T15:17:35.337000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Fisrst of all, many thanks to Kaggle and the hosts for hosting such an interesting competition, and congratulations to all the winners. And a special thanks to all my teammates <a href=\"https://www.kaggle.com/takuok\" target=\"_blank\">@takuok</a> and <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> .</p>\n<h1>Data definition</h1>\n<ul>\n<li>train_clean:  there are 1.6 million training images and 81k classes.</li>\n<li>train3.2m: there are 3.2 million images belonging to the 81k classes in train_clean.</li>\n<li>train4.1m: all train data in GLDv2, that is 4.1 million images.</li>\n<li>train4.9m: all train data in GLDv2 and index images of 2019 competition.</li>\n</ul>\n<h1>Tips in Retrieval</h1>\n<p>We used reranking techniques proposed by smlyaka team of the 2019 GLR champion. <br>\n<strong>Github</strong>: <a href=\"https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution\" target=\"_blank\">https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution</a><br>\n<strong>Paper</strong>: <a href=\"https://arxiv.org/abs/2003.11211\" target=\"_blank\">https://arxiv.org/abs/2003.11211</a><br>\nThere are two steps, sort-step and insert-step. First, landmark_id and score of index set and query set are determined using train4.9m. Then, in the sort-step, retrieved index results are sorted based on their label similarity to the query images. In the insert-step, we added samples that are not retrieved by their image similarity but have same landmark_id as the query image.<br>\nThis improved the base score from public/private 0.31324/0.32614 to 0.39905/0.43219.</p>\n<h2>Model</h2>\n<p>The models we used are almost same as Recognition ones.<br>\nWe concatenate embedding vectors of our models shown below, then use them to calculate similarities.</p>\n<h3>takuoko part</h3>\n<p>I used almost the same architecture that the last year 1st team used. <code>ViT backbone -&gt; Dense(1024) -&gt; ArcFace</code> with class weight. I chose ViT backbone because it works well on <a href=\"https://paperswithcode.com/task/fine-grained-image-classification\" target=\"_blank\">FGVC tasks</a>. I had thought about using FGVC tricks to get local and global features, but I didn't have enough time.<br>\nI also suffered from the phenomenon of loss becoming nan when learning with fp16. I got around this by the grad_clip and calculating Attention layer using fp32.</p>\n<p>Base settings</p>\n<ul>\n<li>image size: 384</li>\n<li>optimizer: AdamW</li>\n<li>scheduler: cosine with warmup</li>\n<li>batch size: 16*2GPU</li>\n<li>use fp16</li>\n</ul>\n<p><strong>BeiT-large: 2epochs@(train_clean) -&gt; 8epochs@(train4.1m)</strong><br>\n<strong>Swin-base: 30epochs@(train_clean) -&gt; 6epochs@(train4.1m)</strong><br>\nSadly, the 2nd part of my swin didn’t work well.</p>\n<h3>tascj part</h3>\n<p>I used augmentations, loss (sub-center ArcFace) and training strategies following last year’s top solutions.</p>\n<p><strong>b4: 10epochs@(train_clean, 512) -&gt; 6epochs@(train4.9m, 512)</strong><br>\n<strong>b6: 10epochs@(train_clean, 256) -&gt; 6epochs@(train4.9m, 512) -&gt; 2epochs@(train4.9m, 640)</strong><br>\n<strong>xcit_small_24_p16_384: 6epochs@(train_clean, 384) -&gt; 6epochs@(train4.9m, 384)</strong></p>\n<h3>inoichan part</h3>\n<p>Architecture is sub-center ArcFace with dynamic margins which was used in last year’s 3rd-place team. Backbone -&gt; Dense(512) -&gt; sub-center Arc. Training strategy is as following:</p>\n<h5>efficientnet v2m</h5>\n<p><strong>1st</strong><br>\ndata: train_clean<br>\nimage size: 256 crop from 300<br>\nepoch: 15<br>\nbatch size: 64<br>\nscheduler: warmup</p>\n<p><strong>2nd</strong><br>\ndata: train4.1m<br>\nimage size: 512 crop from 600<br>\nepoch: 15<br>\nbatch size: 8 x 16 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<p><strong>3rd</strong><br>\ndata: train4.1m<br>\nimage size: 640 crop from 720<br>\nepoch: 3<br>\nbatch size: 4 x 64 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<h5>efficientnet b5</h5>\n<p><strong>1st</strong><br>\ndata: train4.1m<br>\nimage size: 256<br>\nepoch: 10<br>\nbatch size: 32 x 8 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<p><strong>2nd</strong><br>\ndata: train3.2m<br>\nimage size: 512<br>\nepoch: 7<br>\nbatch size: 16 x 16 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<p><strong>3rd</strong><br>\ndata: train4.1m<br>\nimage size: 512<br>\nepoch: 9 (freeze backbone in first two epoch, then unfreeze)<br>\nbatch size: 16 x 16 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<h1>Acknowledge</h1>\n<p>takuoko is a member of Z by HP &amp; NVIDIA Data Science Global Ambassadors.<br>\nSpecial Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.<br>\nThis competition has the big dataset.<br>\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and it helped a lot.</p>",
  "messages": [
    {
      "id": 1535219,
      "postDate": "2021-10-05T15:17:35.337Z",
      "content": "<p>Fisrst of all, many thanks to Kaggle and the hosts for hosting such an interesting competition, and congratulations to all the winners. And a special thanks to all my teammates <a href=\"https://www.kaggle.com/takuok\" target=\"_blank\">@takuok</a> and <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> .</p>\n<h1>Data definition</h1>\n<ul>\n<li>train_clean:  there are 1.6 million training images and 81k classes.</li>\n<li>train3.2m: there are 3.2 million images belonging to the 81k classes in train_clean.</li>\n<li>train4.1m: all train data in GLDv2, that is 4.1 million images.</li>\n<li>train4.9m: all train data in GLDv2 and index images of 2019 competition.</li>\n</ul>\n<h1>Tips in Retrieval</h1>\n<p>We used reranking techniques proposed by smlyaka team of the 2019 GLR champion. <br>\n<strong>Github</strong>: <a href=\"https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution\" target=\"_blank\">https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution</a><br>\n<strong>Paper</strong>: <a href=\"https://arxiv.org/abs/2003.11211\" target=\"_blank\">https://arxiv.org/abs/2003.11211</a><br>\nThere are two steps, sort-step and insert-step. First, landmark_id and score of index set and query set are determined using train4.9m. Then, in the sort-step, retrieved index results are sorted based on their label similarity to the query images. In the insert-step, we added samples that are not retrieved by their image similarity but have same landmark_id as the query image.<br>\nThis improved the base score from public/private 0.31324/0.32614 to 0.39905/0.43219.</p>\n<h2>Model</h2>\n<p>The models we used are almost same as Recognition ones.<br>\nWe concatenate embedding vectors of our models shown below, then use them to calculate similarities.</p>\n<h3>takuoko part</h3>\n<p>I used almost the same architecture that the last year 1st team used. <code>ViT backbone -&gt; Dense(1024) -&gt; ArcFace</code> with class weight. I chose ViT backbone because it works well on <a href=\"https://paperswithcode.com/task/fine-grained-image-classification\" target=\"_blank\">FGVC tasks</a>. I had thought about using FGVC tricks to get local and global features, but I didn't have enough time.<br>\nI also suffered from the phenomenon of loss becoming nan when learning with fp16. I got around this by the grad_clip and calculating Attention layer using fp32.</p>\n<p>Base settings</p>\n<ul>\n<li>image size: 384</li>\n<li>optimizer: AdamW</li>\n<li>scheduler: cosine with warmup</li>\n<li>batch size: 16*2GPU</li>\n<li>use fp16</li>\n</ul>\n<p><strong>BeiT-large: 2epochs@(train_clean) -&gt; 8epochs@(train4.1m)</strong><br>\n<strong>Swin-base: 30epochs@(train_clean) -&gt; 6epochs@(train4.1m)</strong><br>\nSadly, the 2nd part of my swin didn’t work well.</p>\n<h3>tascj part</h3>\n<p>I used augmentations, loss (sub-center ArcFace) and training strategies following last year’s top solutions.</p>\n<p><strong>b4: 10epochs@(train_clean, 512) -&gt; 6epochs@(train4.9m, 512)</strong><br>\n<strong>b6: 10epochs@(train_clean, 256) -&gt; 6epochs@(train4.9m, 512) -&gt; 2epochs@(train4.9m, 640)</strong><br>\n<strong>xcit_small_24_p16_384: 6epochs@(train_clean, 384) -&gt; 6epochs@(train4.9m, 384)</strong></p>\n<h3>inoichan part</h3>\n<p>Architecture is sub-center ArcFace with dynamic margins which was used in last year’s 3rd-place team. Backbone -&gt; Dense(512) -&gt; sub-center Arc. Training strategy is as following:</p>\n<h5>efficientnet v2m</h5>\n<p><strong>1st</strong><br>\ndata: train_clean<br>\nimage size: 256 crop from 300<br>\nepoch: 15<br>\nbatch size: 64<br>\nscheduler: warmup</p>\n<p><strong>2nd</strong><br>\ndata: train4.1m<br>\nimage size: 512 crop from 600<br>\nepoch: 15<br>\nbatch size: 8 x 16 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<p><strong>3rd</strong><br>\ndata: train4.1m<br>\nimage size: 640 crop from 720<br>\nepoch: 3<br>\nbatch size: 4 x 64 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<h5>efficientnet b5</h5>\n<p><strong>1st</strong><br>\ndata: train4.1m<br>\nimage size: 256<br>\nepoch: 10<br>\nbatch size: 32 x 8 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<p><strong>2nd</strong><br>\ndata: train3.2m<br>\nimage size: 512<br>\nepoch: 7<br>\nbatch size: 16 x 16 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<p><strong>3rd</strong><br>\ndata: train4.1m<br>\nimage size: 512<br>\nepoch: 9 (freeze backbone in first two epoch, then unfreeze)<br>\nbatch size: 16 x 16 steps w/ freeze batch norm<br>\nscheduler: cosine annealing</p>\n<h1>Acknowledge</h1>\n<p>takuoko is a member of Z by HP &amp; NVIDIA Data Science Global Ambassadors.<br>\nSpecial Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.<br>\nThis competition has the big dataset.<br>\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and it helped a lot.</p>",
      "rawMarkdown": "Fisrst of all, many thanks to Kaggle and the hosts for hosting such an interesting competition, and congratulations to all the winners. And a special thanks to all my teammates @takuok and @tascj0 .\n\n# Data definition\n- train_clean:  there are 1.6 million training images and 81k classes.\n- train3.2m: there are 3.2 million images belonging to the 81k classes in train_clean.\n- train4.1m: all train data in GLDv2, that is 4.1 million images.\n- train4.9m: all train data in GLDv2 and index images of 2019 competition.\n\n# Tips in Retrieval\n\nWe used reranking techniques proposed by smlyaka team of the 2019 GLR champion. \n**Github**: https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution\n**Paper**: https://arxiv.org/abs/2003.11211\nThere are two steps, sort-step and insert-step. First, landmark_id and score of index set and query set are determined using train4.9m. Then, in the sort-step, retrieved index results are sorted based on their label similarity to the query images. In the insert-step, we added samples that are not retrieved by their image similarity but have same landmark_id as the query image.\nThis improved the base score from public/private 0.31324/0.32614 to 0.39905/0.43219.\n\n## Model\n\nThe models we used are almost same as Recognition ones.\nWe concatenate embedding vectors of our models shown below, then use them to calculate similarities.\n\n\n### takuoko part\n\nI used almost the same architecture that the last year 1st team used. `ViT backbone -> Dense(1024) -> ArcFace` with class weight. I chose ViT backbone because it works well on [FGVC tasks](https://paperswithcode.com/task/fine-grained-image-classification). I had thought about using FGVC tricks to get local and global features, but I didn't have enough time.\nI also suffered from the phenomenon of loss becoming nan when learning with fp16. I got around this by the grad_clip and calculating Attention layer using fp32.\n\nBase settings\n- image size: 384\n- optimizer: AdamW\n- scheduler: cosine with warmup\n- batch size: 16*2GPU\n- use fp16\n\n**BeiT-large: 2epochs@(train_clean) -> 8epochs@(train4.1m)**\n**Swin-base: 30epochs@(train_clean) -> 6epochs@(train4.1m)**\nSadly, the 2nd part of my swin didn’t work well.\n\n### tascj part\nI used augmentations, loss (sub-center ArcFace) and training strategies following last year’s top solutions.\n\n**b4: 10epochs@(train_clean, 512) -> 6epochs@(train4.9m, 512)**\n**b6: 10epochs@(train_clean, 256) -> 6epochs@(train4.9m, 512) -> 2epochs@(train4.9m, 640)**\n**xcit_small_24_p16_384: 6epochs@(train_clean, 384) -> 6epochs@(train4.9m, 384)**\n\n\n### inoichan part\nArchitecture is sub-center ArcFace with dynamic margins which was used in last year’s 3rd-place team. Backbone -> Dense(512) -> sub-center Arc. Training strategy is as following:\n\n##### efficientnet v2m\n**1st**\ndata: train_clean\nimage size: 256 crop from 300\nepoch: 15\nbatch size: 64\nscheduler: warmup\n\n**2nd**\ndata: train4.1m\nimage size: 512 crop from 600\nepoch: 15\nbatch size: 8 x 16 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n**3rd**\ndata: train4.1m\nimage size: 640 crop from 720\nepoch: 3\nbatch size: 4 x 64 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n\n##### efficientnet b5\n**1st**\ndata: train4.1m\nimage size: 256\nepoch: 10\nbatch size: 32 x 8 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n**2nd**\ndata: train3.2m\nimage size: 512\nepoch: 7\nbatch size: 16 x 16 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n**3rd**\ndata: train4.1m\nimage size: 512\nepoch: 9 (freeze backbone in first two epoch, then unfreeze)\nbatch size: 16 x 16 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n\n# Acknowledge\ntakuoko is a member of Z by HP & NVIDIA Data Science Global Ambassadors.\nSpecial Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.\nThis competition has the big dataset.\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and it helped a lot.\n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1535219": "Fisrst of all, many thanks to Kaggle and the hosts for hosting such an interesting competition, and congratulations to all the winners. And a special thanks to all my teammates @takuok and @tascj0 .\n\n# Data definition\n- train_clean:  there are 1.6 million training images and 81k classes.\n- train3.2m: there are 3.2 million images belonging to the 81k classes in train_clean.\n- train4.1m: all train data in GLDv2, that is 4.1 million images.\n- train4.9m: all train data in GLDv2 and index images of 2019 competition.\n\n# Tips in Retrieval\n\nWe used reranking techniques proposed by smlyaka team of the 2019 GLR champion. \n**Github**: https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution\n**Paper**: https://arxiv.org/abs/2003.11211\nThere are two steps, sort-step and insert-step. First, landmark_id and score of index set and query set are determined using train4.9m. Then, in the sort-step, retrieved index results are sorted based on their label similarity to the query images. In the insert-step, we added samples that are not retrieved by their image similarity but have same landmark_id as the query image.\nThis improved the base score from public/private 0.31324/0.32614 to 0.39905/0.43219.\n\n## Model\n\nThe models we used are almost same as Recognition ones.\nWe concatenate embedding vectors of our models shown below, then use them to calculate similarities.\n\n\n### takuoko part\n\nI used almost the same architecture that the last year 1st team used. `ViT backbone -> Dense(1024) -> ArcFace` with class weight. I chose ViT backbone because it works well on [FGVC tasks](https://paperswithcode.com/task/fine-grained-image-classification). I had thought about using FGVC tricks to get local and global features, but I didn't have enough time.\nI also suffered from the phenomenon of loss becoming nan when learning with fp16. I got around this by the grad_clip and calculating Attention layer using fp32.\n\nBase settings\n- image size: 384\n- optimizer: AdamW\n- scheduler: cosine with warmup\n- batch size: 16*2GPU\n- use fp16\n\n**BeiT-large: 2epochs@(train_clean) -> 8epochs@(train4.1m)**\n**Swin-base: 30epochs@(train_clean) -> 6epochs@(train4.1m)**\nSadly, the 2nd part of my swin didn’t work well.\n\n### tascj part\nI used augmentations, loss (sub-center ArcFace) and training strategies following last year’s top solutions.\n\n**b4: 10epochs@(train_clean, 512) -> 6epochs@(train4.9m, 512)**\n**b6: 10epochs@(train_clean, 256) -> 6epochs@(train4.9m, 512) -> 2epochs@(train4.9m, 640)**\n**xcit_small_24_p16_384: 6epochs@(train_clean, 384) -> 6epochs@(train4.9m, 384)**\n\n\n### inoichan part\nArchitecture is sub-center ArcFace with dynamic margins which was used in last year’s 3rd-place team. Backbone -> Dense(512) -> sub-center Arc. Training strategy is as following:\n\n##### efficientnet v2m\n**1st**\ndata: train_clean\nimage size: 256 crop from 300\nepoch: 15\nbatch size: 64\nscheduler: warmup\n\n**2nd**\ndata: train4.1m\nimage size: 512 crop from 600\nepoch: 15\nbatch size: 8 x 16 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n**3rd**\ndata: train4.1m\nimage size: 640 crop from 720\nepoch: 3\nbatch size: 4 x 64 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n\n##### efficientnet b5\n**1st**\ndata: train4.1m\nimage size: 256\nepoch: 10\nbatch size: 32 x 8 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n**2nd**\ndata: train3.2m\nimage size: 512\nepoch: 7\nbatch size: 16 x 16 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n**3rd**\ndata: train4.1m\nimage size: 512\nepoch: 9 (freeze backbone in first two epoch, then unfreeze)\nbatch size: 16 x 16 steps w/ freeze batch norm\nscheduler: cosine annealing\n\n\n# Acknowledge\ntakuoko is a member of Z by HP & NVIDIA Data Science Global Ambassadors.\nSpecial Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.\nThis competition has the big dataset.\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and it helped a lot.\n"
  }
}