{
  "id": 242207,
  "title": "8th place solution: arcface + cosface + classification",
  "url": "/competitions/hotel-id-2021-fgvc8/writeups/michaln-8th-place-solution-arcface-cosface-classif",
  "author_name": "",
  "post_date": "2021-06-06T23:38:08.363Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Many thanks to the organizers for hosting such an interesting challenge and I hope some of the solutions will help to improve the current methods used in trafficking investigations.</p>\n<p>Congrats to the winners, the final scores are pretty impressive and I can't wait to learn what was the secret sauce of your solutions.</p>\n<h2>Motivation</h2>\n<p>I've never worked on reverse image search or image similary problem before so i wanted to try out some methods and learn something new. On top of that this competition is very interesting and the solutions might have a real world impact which made it very tempting to join.</p>\n<h2>Overview</h2>\n<p>My solution is nothing special. I trained 3 types of models<br>\nArcMargin model: <a href=\"https://www.kaggle.com/michaln/hotel-id-arcmargin-training\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-arcmargin-training</a><br>\nCosFace model: <a href=\"https://www.kaggle.com/michaln/hotel-id-cosface-training\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-cosface-training</a><br>\nSimple classification model: <a href=\"https://www.kaggle.com/michaln/hotel-id-classification-training\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-classification-training</a></p>\n<p>With same parameters: Lookahead + AdamW optimizer, OneCycleLR scheduler, CrossEntropyLoss/CosFace loss</p>\n<p>Models used as backbones: eca_nfnet_l0, efficientnet_b1,  ecaresnet50d_pruned</p>\n<p>These models were then used to generate embeddings for the images which were then used to calculated cosine similarity of the test images to the train dataset. To ensemble I just calculated the product of similarities from different models and then found the top 5 most similar images from different hotels.</p>\n<p>I planned to train more models but it was painfully slow (2-3 hours for a single epoch on Colab using T4 GPU) and I ran out of time. So in the end I used only 4 models in my final ensemble.<br>\nInference notebook: <a href=\"https://www.kaggle.com/michaln/hotel-id-inference\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-inference</a><br>\nDataset with trained models: <a href=\"https://www.kaggle.com/michaln/hotelid-trained-models\" target=\"_blank\">https://www.kaggle.com/michaln/hotelid-trained-models</a></p>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Backbone</th>\n<th>Embed size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Epochs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ArcMargin</td>\n<td>eca_nfnet_l0</td>\n<td>1024</td>\n<td>0.6564</td>\n<td>0.6704</td>\n<td>6/6</td>\n</tr>\n<tr>\n<td>ArcMargin</td>\n<td>efficientnet_b1</td>\n<td>4096</td>\n<td>0.6780</td>\n<td>0.6962</td>\n<td>9/9</td>\n</tr>\n<tr>\n<td>Classification</td>\n<td>eca_nfnet_l0</td>\n<td>4096</td>\n<td>0.6691</td>\n<td>0.6875</td>\n<td>6/9</td>\n</tr>\n<tr>\n<td>CosFace</td>\n<td>ecaresnet50d_pruned</td>\n<td>4096</td>\n<td>0.6702</td>\n<td>0.6796</td>\n<td>9/9</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td></td>\n<td></td>\n<td>0.7273</td>\n<td>0.7446</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>git: <a href=\"https://github.com/michal-nahlik/kaggle-hotel-id-2021\" target=\"_blank\">https://github.com/michal-nahlik/kaggle-hotel-id-2021</a></p>\n<h2>Data</h2>\n<p>I used only competition data as it was never really confirmed that we can use external datasets like Hotels50k. I rescaled images to 512x512 and padded them when it was needed.</p>\n<p>Image preprocessing notebook: <a href=\"https://www.kaggle.com/michaln/hotel-id-preprocess-images\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-preprocess-images</a><br>\n512x512 dataset: <a href=\"https://www.kaggle.com/michaln/hotelid-images-512x512-padded\" target=\"_blank\">https://www.kaggle.com/michaln/hotelid-images-512x512-padded</a><br>\n256x256 dataset: <a href=\"https://www.kaggle.com/michaln/hotelid-images-256x256-padded\" target=\"_blank\">https://www.kaggle.com/michaln/hotelid-images-256x256-padded</a></p>\n<h2>What worked</h2>\n<ul>\n<li>heavy augmentations</li>\n<li>cosine similarity (I tried euclidean distance, SNR Distance and other buts cosine worked the best)</li>\n<li>lookahead + AdamW optimizer</li>\n<li>bigger embedding layer - in general I saw best results with embedding layer with at least half of the size of final targets (used 4096 for 7770 targets)</li>\n<li>bigger images (512x512 was better than 256x256, 1024 was even better but too slow to train)</li>\n<li>eca_nfnet_l0 - this model got the best score out of all, it was a hidden gem dicovered in Shopee competition for me</li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>TTA - I tried to use TTA to rotate the image and use the maximum similarity (or minumum distance) over the TTAs but in the end it did not improve the score</li>\n<li>clustering + voting - I tried clustering to get the predictions based on embeddings and then use voting to ensemble different models but simple search for most similar image and product of similarites worked better</li>\n<li>triplet loss - couldn't get a decent score, it got much better when I added classification head but then pure classification got better results by itself so I dropped triplet completely</li>\n<li>pretraining on smaller images + pretraining and freezing some layers - full training was just much better</li>\n</ul>\n<h2>What worked but I didn't use</h2>\n<ul>\n<li>More data -  I downloaded parts of the Hotels 50k dataset and added them to the competition data and it did improve the score. The training was pretty slow already and it was never confirmed that we can use it so I decided to stick with the competition data.</li>\n</ul>",
  "messages": [
    {
      "id": "1325712",
      "postDate": "05/28/2021 00:48:11",
      "content": "<p>Many thanks to the organizers for hosting such an interesting challenge and I hope some of the solutions will help to improve the current methods used in trafficking investigations.</p>\n<p>Congrats to the winners, the final scores are pretty impressive and I can't wait to learn what was the secret sauce of your solutions.</p>\n<h2>Motivation</h2>\n<p>I've never worked on reverse image search or image similary problem before so i wanted to try out some methods and learn something new. On top of that this competition is very interesting and the solutions might have a real world impact which made it very tempting to join.</p>\n<h2>Overview</h2>\n<p>My solution is nothing special. I trained 3 types of models<br>\nArcMargin model: <a href=\"https://www.kaggle.com/michaln/hotel-id-arcmargin-training\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-arcmargin-training</a><br>\nCosFace model: <a href=\"https://www.kaggle.com/michaln/hotel-id-cosface-training\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-cosface-training</a><br>\nSimple classification model: <a href=\"https://www.kaggle.com/michaln/hotel-id-classification-training\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-classification-training</a></p>\n<p>With same parameters: Lookahead + AdamW optimizer, OneCycleLR scheduler, CrossEntropyLoss/CosFace loss</p>\n<p>Models used as backbones: eca_nfnet_l0, efficientnet_b1,  ecaresnet50d_pruned</p>\n<p>These models were then used to generate embeddings for the images which were then used to calculated cosine similarity of the test images to the train dataset. To ensemble I just calculated the product of similarities from different models and then found the top 5 most similar images from different hotels.</p>\n<p>I planned to train more models but it was painfully slow (2-3 hours for a single epoch on Colab using T4 GPU) and I ran out of time. So in the end I used only 4 models in my final ensemble.<br>\nInference notebook: <a href=\"https://www.kaggle.com/michaln/hotel-id-inference\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-inference</a><br>\nDataset with trained models: <a href=\"https://www.kaggle.com/michaln/hotelid-trained-models\" target=\"_blank\">https://www.kaggle.com/michaln/hotelid-trained-models</a></p>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Backbone</th>\n<th>Embed size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Epochs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ArcMargin</td>\n<td>eca_nfnet_l0</td>\n<td>1024</td>\n<td>0.6564</td>\n<td>0.6704</td>\n<td>6/6</td>\n</tr>\n<tr>\n<td>ArcMargin</td>\n<td>efficientnet_b1</td>\n<td>4096</td>\n<td>0.6780</td>\n<td>0.6962</td>\n<td>9/9</td>\n</tr>\n<tr>\n<td>Classification</td>\n<td>eca_nfnet_l0</td>\n<td>4096</td>\n<td>0.6691</td>\n<td>0.6875</td>\n<td>6/9</td>\n</tr>\n<tr>\n<td>CosFace</td>\n<td>ecaresnet50d_pruned</td>\n<td>4096</td>\n<td>0.6702</td>\n<td>0.6796</td>\n<td>9/9</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td></td>\n<td></td>\n<td>0.7273</td>\n<td>0.7446</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>git: <a href=\"https://github.com/michal-nahlik/kaggle-hotel-id-2021\" target=\"_blank\">https://github.com/michal-nahlik/kaggle-hotel-id-2021</a></p>\n<h2>Data</h2>\n<p>I used only competition data as it was never really confirmed that we can use external datasets like Hotels50k. I rescaled images to 512x512 and padded them when it was needed.</p>\n<p>Image preprocessing notebook: <a href=\"https://www.kaggle.com/michaln/hotel-id-preprocess-images\" target=\"_blank\">https://www.kaggle.com/michaln/hotel-id-preprocess-images</a><br>\n512x512 dataset: <a href=\"https://www.kaggle.com/michaln/hotelid-images-512x512-padded\" target=\"_blank\">https://www.kaggle.com/michaln/hotelid-images-512x512-padded</a><br>\n256x256 dataset: <a href=\"https://www.kaggle.com/michaln/hotelid-images-256x256-padded\" target=\"_blank\">https://www.kaggle.com/michaln/hotelid-images-256x256-padded</a></p>\n<h2>What worked</h2>\n<ul>\n<li>heavy augmentations</li>\n<li>cosine similarity (I tried euclidean distance, SNR Distance and other buts cosine worked the best)</li>\n<li>lookahead + AdamW optimizer</li>\n<li>bigger embedding layer - in general I saw best results with embedding layer with at least half of the size of final targets (used 4096 for 7770 targets)</li>\n<li>bigger images (512x512 was better than 256x256, 1024 was even better but too slow to train)</li>\n<li>eca_nfnet_l0 - this model got the best score out of all, it was a hidden gem dicovered in Shopee competition for me</li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>TTA - I tried to use TTA to rotate the image and use the maximum similarity (or minumum distance) over the TTAs but in the end it did not improve the score</li>\n<li>clustering + voting - I tried clustering to get the predictions based on embeddings and then use voting to ensemble different models but simple search for most similar image and product of similarites worked better</li>\n<li>triplet loss - couldn't get a decent score, it got much better when I added classification head but then pure classification got better results by itself so I dropped triplet completely</li>\n<li>pretraining on smaller images + pretraining and freezing some layers - full training was just much better</li>\n</ul>\n<h2>What worked but I didn't use</h2>\n<ul>\n<li>More data -  I downloaded parts of the Hotels 50k dataset and added them to the competition data and it did improve the score. The training was pretty slow already and it was never confirmed that we can use it so I decided to stick with the competition data.</li>\n</ul>",
      "rawMarkdown": "Many thanks to the organizers for hosting such an interesting challenge and I hope some of the solutions will help to improve the current methods used in trafficking investigations.\n\nCongrats to the winners, the final scores are pretty impressive and I can't wait to learn what was the secret sauce of your solutions.\n\n## Motivation\nI've never worked on reverse image search or image similary problem before so i wanted to try out some methods and learn something new. On top of that this competition is very interesting and the solutions might have a real world impact which made it very tempting to join.\n\n## Overview\nMy solution is nothing special. I trained 3 types of models\nArcMargin model: https://www.kaggle.com/michaln/hotel-id-arcmargin-training\nCosFace model: https://www.kaggle.com/michaln/hotel-id-cosface-training\nSimple classification model: https://www.kaggle.com/michaln/hotel-id-classification-training\n\nWith same parameters: Lookahead + AdamW optimizer, OneCycleLR scheduler, CrossEntropyLoss/CosFace loss\n\nModels used as backbones: eca_nfnet_l0, efficientnet_b1,  ecaresnet50d_pruned\n\nThese models were then used to generate embeddings for the images which were then used to calculated cosine similarity of the test images to the train dataset. To ensemble I just calculated the product of similarities from different models and then found the top 5 most similar images from different hotels.\n\nI planned to train more models but it was painfully slow (2-3 hours for a single epoch on Colab using T4 GPU) and I ran out of time. So in the end I used only 4 models in my final ensemble.\nInference notebook: https://www.kaggle.com/michaln/hotel-id-inference\nDataset with trained models: https://www.kaggle.com/michaln/hotelid-trained-models\n\n| Type | Backbone | Embed size | Public LB| Private LB | Epochs | \n| --- | --- | --- | --- | --- | --- |\n| ArcMargin | eca_nfnet_l0 | 1024 | 0.6564 | 0.6704 | 6/6 |\n| ArcMargin | efficientnet_b1 | 4096 | 0.6780 | 0.6962 | 9/9 |\n| Classification | eca_nfnet_l0 | 4096 | 0.6691 | 0.6875 | 6/9|\n| CosFace | ecaresnet50d_pruned| 4096 | 0.6702 | 0.6796 | 9/9 |\n| Ensemble |  |  | 0.7273 | 0.7446 | |\n\ngit: https://github.com/michal-nahlik/kaggle-hotel-id-2021\n\n## Data\nI used only competition data as it was never really confirmed that we can use external datasets like Hotels50k. I rescaled images to 512x512 and padded them when it was needed.\n\nImage preprocessing notebook: https://www.kaggle.com/michaln/hotel-id-preprocess-images\n512x512 dataset: https://www.kaggle.com/michaln/hotelid-images-512x512-padded\n256x256 dataset: https://www.kaggle.com/michaln/hotelid-images-256x256-padded\n\n\n## What worked\n- heavy augmentations\n- cosine similarity (I tried euclidean distance, SNR Distance and other buts cosine worked the best)\n- lookahead + AdamW optimizer\n- bigger embedding layer - in general I saw best results with embedding layer with at least half of the size of final targets (used 4096 for 7770 targets)\n- bigger images (512x512 was better than 256x256, 1024 was even better but too slow to train)\n- eca_nfnet_l0 - this model got the best score out of all, it was a hidden gem dicovered in Shopee competition for me\n\n## What didn't work\n- TTA - I tried to use TTA to rotate the image and use the maximum similarity (or minumum distance) over the TTAs but in the end it did not improve the score\n- clustering + voting - I tried clustering to get the predictions based on embeddings and then use voting to ensemble different models but simple search for most similar image and product of similarites worked better\n- triplet loss - couldn't get a decent score, it got much better when I added classification head but then pure classification got better results by itself so I dropped triplet completely\n- pretraining on smaller images + pretraining and freezing some layers - full training was just much better\n\n## What worked but I didn't use\n- More data -  I downloaded parts of the Hotels 50k dataset and added them to the competition data and it did improve the score. The training was pretty slow already and it was never confirmed that we can use it so I decided to stick with the competition data.",
      "votes": null
    },
    {
      "id": "1326608",
      "postDate": "05/28/2021 15:10:46",
      "content": "<p>thanks for your sharing,good solution👍</p>",
      "rawMarkdown": "thanks for your sharing,good solution👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1326608,
      "author_name": "xingkuizhu",
      "author_url": "",
      "post_date": "05/28/2021 15:10:46",
      "content": "<p>thanks for your sharing,good solution👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1325712": "Many thanks to the organizers for hosting such an interesting challenge and I hope some of the solutions will help to improve the current methods used in trafficking investigations.\n\nCongrats to the winners, the final scores are pretty impressive and I can't wait to learn what was the secret sauce of your solutions.\n\n## Motivation\nI've never worked on reverse image search or image similary problem before so i wanted to try out some methods and learn something new. On top of that this competition is very interesting and the solutions might have a real world impact which made it very tempting to join.\n\n## Overview\nMy solution is nothing special. I trained 3 types of models\nArcMargin model: https://www.kaggle.com/michaln/hotel-id-arcmargin-training\nCosFace model: https://www.kaggle.com/michaln/hotel-id-cosface-training\nSimple classification model: https://www.kaggle.com/michaln/hotel-id-classification-training\n\nWith same parameters: Lookahead + AdamW optimizer, OneCycleLR scheduler, CrossEntropyLoss/CosFace loss\n\nModels used as backbones: eca_nfnet_l0, efficientnet_b1,  ecaresnet50d_pruned\n\nThese models were then used to generate embeddings for the images which were then used to calculated cosine similarity of the test images to the train dataset. To ensemble I just calculated the product of similarities from different models and then found the top 5 most similar images from different hotels.\n\nI planned to train more models but it was painfully slow (2-3 hours for a single epoch on Colab using T4 GPU) and I ran out of time. So in the end I used only 4 models in my final ensemble.\nInference notebook: https://www.kaggle.com/michaln/hotel-id-inference\nDataset with trained models: https://www.kaggle.com/michaln/hotelid-trained-models\n\n| Type | Backbone | Embed size | Public LB| Private LB | Epochs | \n| --- | --- | --- | --- | --- | --- |\n| ArcMargin | eca_nfnet_l0 | 1024 | 0.6564 | 0.6704 | 6/6 |\n| ArcMargin | efficientnet_b1 | 4096 | 0.6780 | 0.6962 | 9/9 |\n| Classification | eca_nfnet_l0 | 4096 | 0.6691 | 0.6875 | 6/9|\n| CosFace | ecaresnet50d_pruned| 4096 | 0.6702 | 0.6796 | 9/9 |\n| Ensemble |  |  | 0.7273 | 0.7446 | |\n\ngit: https://github.com/michal-nahlik/kaggle-hotel-id-2021\n\n## Data\nI used only competition data as it was never really confirmed that we can use external datasets like Hotels50k. I rescaled images to 512x512 and padded them when it was needed.\n\nImage preprocessing notebook: https://www.kaggle.com/michaln/hotel-id-preprocess-images\n512x512 dataset: https://www.kaggle.com/michaln/hotelid-images-512x512-padded\n256x256 dataset: https://www.kaggle.com/michaln/hotelid-images-256x256-padded\n\n\n## What worked\n- heavy augmentations\n- cosine similarity (I tried euclidean distance, SNR Distance and other buts cosine worked the best)\n- lookahead + AdamW optimizer\n- bigger embedding layer - in general I saw best results with embedding layer with at least half of the size of final targets (used 4096 for 7770 targets)\n- bigger images (512x512 was better than 256x256, 1024 was even better but too slow to train)\n- eca_nfnet_l0 - this model got the best score out of all, it was a hidden gem dicovered in Shopee competition for me\n\n## What didn't work\n- TTA - I tried to use TTA to rotate the image and use the maximum similarity (or minumum distance) over the TTAs but in the end it did not improve the score\n- clustering + voting - I tried clustering to get the predictions based on embeddings and then use voting to ensemble different models but simple search for most similar image and product of similarites worked better\n- triplet loss - couldn't get a decent score, it got much better when I added classification head but then pure classification got better results by itself so I dropped triplet completely\n- pretraining on smaller images + pretraining and freezing some layers - full training was just much better\n\n## What worked but I didn't use\n- More data -  I downloaded parts of the Hotels 50k dataset and added them to the competition data and it did improve the score. The training was pretty slow already and it was never confirmed that we can use it so I decided to stick with the competition data.",
    "1326608": "thanks for your sharing,good solution👍"
  },
  "source": "meta"
}