{
  "id": 276338,
  "title": "11th Place Solution: Supervised Contrastive Pretraining + Query and Database expansion",
  "url": "/competitions/landmark-retrieval-2021/writeups/fat-cat-biba-boba-11th-place-solution-supervised-c",
  "author_name": "",
  "post_date": "2021-10-04T10:40:08.041670300Z",
  "votes": 33,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to Google and Kaggle for organizing this interesting competition! <br>\nCongratulations to the winners!</p>\n<p>It was the first large-scale competition for both teams <a href=\"https://www.kaggle.com/ofitserovlad\" target=\"_blank\">@ofitserovlad</a>, <a href=\"https://www.kaggle.com/evgenysidorov\" target=\"_blank\">@evgenysidorov</a>, <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> and we spent a lot of time on the experiments and came up with a strong pipeline pretty close to the end of the competition. Even though we finished just short of the gold medal we are still more than satisfied with the results and enjoyed the competition immensely.</p>\n<h1>Solution Overview</h1>\n<p><img src=\"https://user-images.githubusercontent.com/14181915/135837030-4bdf97b7-aa89-42db-9b2c-f97c837cdb58.png\" alt=\"\"></p>\n<h1>1. Supervised Contrastive Pretraining</h1>\n<h3>Models</h3>\n<p>Efficientnetv2_m</p>\n<h3>Validation</h3>\n<p>mAP calculated with Index (100k sampled) and Test (1k) data from google’s repository. CV and LB had good correlation with LB being almost always 0.015-0.02 lower than CV (ex. CV 0.380 -&gt; LB 0.360)</p>\n<h3>Training Pipeline</h3>\n<p><img src=\"https://user-images.githubusercontent.com/14181915/135837248-2b9063ec-aaed-40a1-987b-2bc3f5181538.png\" alt=\"\"></p>\n<p>Team Fat Cat finalized a pretty strong 3 stage training process a bit too late (1.5 week before the end) into the competition:<br>\nArtificially Sampled Supervised Contrastive Pretraining with class balanced temperature on 448 resolution and clean dataset <br>\nFinetuning with AdaCos with class balanced Focal Loss with 512 resolution on clean dataset <br>\nSame as stage 2 but with 604 resolution and full dataset </p>\n<h3>Artificially Sampled Supervised Contrastive Pretraining with class balanced temperature</h3>\n<p>Pretraining part was inspired by this <a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">paper</a>. At this stage the model should learn easy image features and be more generalizable before finetuning stage. The premise of the paper is pretty similar to SimCLR, which takes 2 augmented (color jitter, random flip) random cuts of the image and makes them close in the cosine space. The extension of Supervised Contrastive paper is to make cuts of the image of the same class in the batch also close together. This paper also normalizes the image embeddings before feeding them into the nonlinear projection head. Outputs of the projection head are then used for supervised contrastive loss. </p>\n<p>The main problem with implementation of the original paper was that they used enormous batch size (8192) and only 1000 classes and by doing so are somewhat guaranteed to have multiple instances of the same class in the batch. When I tried that kind of approach with a batch of 256 cv was at ~0.180, clearly indicating that there was very little probability of the same class occurring in the batch. </p>\n<p>1.5 weeks before the end I realized that I can artificially sample cuts of the different images of the same class into 1 batch. I shuffled my dataset in such a way that each 4 consecutive images were augmented cuts of the same class. This kind of pretraining resulted in cv mAP of .273 (and was much faster than cutting 2 parts of the same image), after 10 epochs of pretraining. Graphs showed that given more time it could reach .300. For pretraining I used SGD with cosine warmup and annealing. </p>\n<p>Additionally, individual temperature for each class was used. It was derived using class-balanced strategy and omitting final normalization (i.e sum of the weights equals to the number of classes). </p>\n<h3>Finetuning</h3>\n<p>For finetuning Adaptive Cosine loss was used. It is an extension of AcrFace and was also used by the previous year’s winner. Supervised Contrastive pretraining was a crucial step for faster convergence during finetuning. Without pretraining after 1 epoch training on clean data and 512 resolution CV mAP was at .210 and with pretraining it started at 0.31 and reached 0.35 after 10 epochs with SGD and cosine scheduling. However, with better hyperparameters it potentially could have climbed even higher. Then, the resolution was increased to 604 and finetuned on the full dataset. </p>\n<p>At the very end of the competition we added additional weight to landmarks from Asia, Africa, and Oceania to try to match the distribution of the test data which was described by Google's paper. It helped to bridge the gap between LB and CV. </p>\n<h1>2. Encoder + GEM + ArcHead</h1>\n<h3>Networks:</h3>\n<p>EfficientnetB6, 704x704<br>\nEfficientnetB7, 664x664<br>\nSeresnext101_32x4d, 512x512</p>\n<h3>Training procedure:</h3>\n<p>ArcFace loss based on Focal Loss with various modifications was used to train the models. <br>\nWe quickly realized that big models with big resolution give improvements up to 0.03-0.05.<br>\nTraining big models was optimized by using FP16 precision, RandomCrop augmentations to reduce training image size (for example, from 704 to 664) and doing all the augmentations on GPU with Kornia.<br>\nOne model training can be divided into two steps:<br>\nTraining on GLDv2 cleaned<br>\nFine tuning on GLDv2 full dataset with smooth labeling and doubled weights for cleaned samples to reduce effect of the noisiness </p>\n<h3>3. Post Processing</h3>\n<p>Two stage post processing procedure was suggested by Fat Cat. It consisted of query expansion and database expansion:</p>\n<p>Query expansion - for each sample from test set we find top2 samples from index set by cosine similarity and generate new embedding for test sample by normalizing sum of weighted by cosine similarity score top2 index samples and test sample<br>\nDatabase expansion - for each test and index sample we find top16 samples from GLDv2 train set and based on their classes define the class of the test/index sample. Using this classes we re-rank top100 list of index samples for each test sample that we get after qe expansion. We put on top of the new list samples from the top100 list that have the same class as the test sample in the same order. Then we put samples from the index set that have the same class as the test sample but weren't in the top100 list before sorted by cosine similarity score, to the end of the new list we put all samples from the top100 list that have different class than the test sample. Finally, top100 samples from the new list are taken.<br>\nThis post processing strategy improved our LB score by around 0.07.</p>\n<h3>Ensemble</h3>\n<p>For all the post processing procedures we used faiss to speed it up using GPU instead of CPU. For testing the ensemble we used Biba &amp; Boba’s Seresnext101_32x4d model (0.418 LB) and Fat Cat’s model (0.425 LB). Our team explored two different ways of merging models.</p>\n<ul>\n<li>“Majority voting” - from each model in the ensemble top100 samples was found independently and then all the image ids were sorted based on the sum of the positions that this ids has in each model’s top100 list. Finally, the top 100 indices from the sorted list were taken. (0.433 on LB)</li>\n<li>“Concat” - concatenating normalized model’s embeddings. Before concat each embedding was multiplied by the weight given to its model. When all embeddings are concatenated we normalize the new big one. To new big embedding we applied two staged post processing described above. (0.455 on LB)</li>\n</ul>\n<p>The final ensemble were made with:<br>\nTwo EffientnetV2 M models trained with supervised contrastive pretraining + AdaCos<br>\nSeresnext101_32x4d and EfficientnetB6 “Encoder + GEM + ArcHead” models</p>\n<h3>Final Thoughts</h3>\n<p>It was both teams' first large-scale kaggle competition so a lot of different mistakes were made along the way. After all experiments here are the ingredients of what we believe could be very strong training pipeline</p>\n<ul>\n<li>Better cropping <a href=\"https://arxiv.org/abs/1906.06423\" target=\"_blank\">paper</a></li>\n<li>Artificially sampled Supervised contrastive pretraining with 8 cuts of different images (here we used 4) with 448 validation resolution and 20 epochs on clean data</li>\n<li>Finetuning with ArcFace 512 validation resolution and class balanced focal loss for 20 epoch with sgd without schedulers on clean data</li>\n<li>Same as previous previous step but with 736 and total data</li>\n<li>Ensemble by concatenation </li>\n<li>Post processing: qe + db expansion </li>\n</ul>",
  "messages": [
    {
      "id": "1533758",
      "postDate": "10/04/2021 10:40:08",
      "content": "<p>Thanks to Google and Kaggle for organizing this interesting competition! <br>\nCongratulations to the winners!</p>\n<p>It was the first large-scale competition for both teams <a href=\"https://www.kaggle.com/ofitserovlad\" target=\"_blank\">@ofitserovlad</a>, <a href=\"https://www.kaggle.com/evgenysidorov\" target=\"_blank\">@evgenysidorov</a>, <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> and we spent a lot of time on the experiments and came up with a strong pipeline pretty close to the end of the competition. Even though we finished just short of the gold medal we are still more than satisfied with the results and enjoyed the competition immensely.</p>\n<h1>Solution Overview</h1>\n<p><img src=\"https://user-images.githubusercontent.com/14181915/135837030-4bdf97b7-aa89-42db-9b2c-f97c837cdb58.png\" alt=\"\"></p>\n<h1>1. Supervised Contrastive Pretraining</h1>\n<h3>Models</h3>\n<p>Efficientnetv2_m</p>\n<h3>Validation</h3>\n<p>mAP calculated with Index (100k sampled) and Test (1k) data from google’s repository. CV and LB had good correlation with LB being almost always 0.015-0.02 lower than CV (ex. CV 0.380 -&gt; LB 0.360)</p>\n<h3>Training Pipeline</h3>\n<p><img src=\"https://user-images.githubusercontent.com/14181915/135837248-2b9063ec-aaed-40a1-987b-2bc3f5181538.png\" alt=\"\"></p>\n<p>Team Fat Cat finalized a pretty strong 3 stage training process a bit too late (1.5 week before the end) into the competition:<br>\nArtificially Sampled Supervised Contrastive Pretraining with class balanced temperature on 448 resolution and clean dataset <br>\nFinetuning with AdaCos with class balanced Focal Loss with 512 resolution on clean dataset <br>\nSame as stage 2 but with 604 resolution and full dataset </p>\n<h3>Artificially Sampled Supervised Contrastive Pretraining with class balanced temperature</h3>\n<p>Pretraining part was inspired by this <a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">paper</a>. At this stage the model should learn easy image features and be more generalizable before finetuning stage. The premise of the paper is pretty similar to SimCLR, which takes 2 augmented (color jitter, random flip) random cuts of the image and makes them close in the cosine space. The extension of Supervised Contrastive paper is to make cuts of the image of the same class in the batch also close together. This paper also normalizes the image embeddings before feeding them into the nonlinear projection head. Outputs of the projection head are then used for supervised contrastive loss. </p>\n<p>The main problem with implementation of the original paper was that they used enormous batch size (8192) and only 1000 classes and by doing so are somewhat guaranteed to have multiple instances of the same class in the batch. When I tried that kind of approach with a batch of 256 cv was at ~0.180, clearly indicating that there was very little probability of the same class occurring in the batch. </p>\n<p>1.5 weeks before the end I realized that I can artificially sample cuts of the different images of the same class into 1 batch. I shuffled my dataset in such a way that each 4 consecutive images were augmented cuts of the same class. This kind of pretraining resulted in cv mAP of .273 (and was much faster than cutting 2 parts of the same image), after 10 epochs of pretraining. Graphs showed that given more time it could reach .300. For pretraining I used SGD with cosine warmup and annealing. </p>\n<p>Additionally, individual temperature for each class was used. It was derived using class-balanced strategy and omitting final normalization (i.e sum of the weights equals to the number of classes). </p>\n<h3>Finetuning</h3>\n<p>For finetuning Adaptive Cosine loss was used. It is an extension of AcrFace and was also used by the previous year’s winner. Supervised Contrastive pretraining was a crucial step for faster convergence during finetuning. Without pretraining after 1 epoch training on clean data and 512 resolution CV mAP was at .210 and with pretraining it started at 0.31 and reached 0.35 after 10 epochs with SGD and cosine scheduling. However, with better hyperparameters it potentially could have climbed even higher. Then, the resolution was increased to 604 and finetuned on the full dataset. </p>\n<p>At the very end of the competition we added additional weight to landmarks from Asia, Africa, and Oceania to try to match the distribution of the test data which was described by Google's paper. It helped to bridge the gap between LB and CV. </p>\n<h1>2. Encoder + GEM + ArcHead</h1>\n<h3>Networks:</h3>\n<p>EfficientnetB6, 704x704<br>\nEfficientnetB7, 664x664<br>\nSeresnext101_32x4d, 512x512</p>\n<h3>Training procedure:</h3>\n<p>ArcFace loss based on Focal Loss with various modifications was used to train the models. <br>\nWe quickly realized that big models with big resolution give improvements up to 0.03-0.05.<br>\nTraining big models was optimized by using FP16 precision, RandomCrop augmentations to reduce training image size (for example, from 704 to 664) and doing all the augmentations on GPU with Kornia.<br>\nOne model training can be divided into two steps:<br>\nTraining on GLDv2 cleaned<br>\nFine tuning on GLDv2 full dataset with smooth labeling and doubled weights for cleaned samples to reduce effect of the noisiness </p>\n<h3>3. Post Processing</h3>\n<p>Two stage post processing procedure was suggested by Fat Cat. It consisted of query expansion and database expansion:</p>\n<p>Query expansion - for each sample from test set we find top2 samples from index set by cosine similarity and generate new embedding for test sample by normalizing sum of weighted by cosine similarity score top2 index samples and test sample<br>\nDatabase expansion - for each test and index sample we find top16 samples from GLDv2 train set and based on their classes define the class of the test/index sample. Using this classes we re-rank top100 list of index samples for each test sample that we get after qe expansion. We put on top of the new list samples from the top100 list that have the same class as the test sample in the same order. Then we put samples from the index set that have the same class as the test sample but weren't in the top100 list before sorted by cosine similarity score, to the end of the new list we put all samples from the top100 list that have different class than the test sample. Finally, top100 samples from the new list are taken.<br>\nThis post processing strategy improved our LB score by around 0.07.</p>\n<h3>Ensemble</h3>\n<p>For all the post processing procedures we used faiss to speed it up using GPU instead of CPU. For testing the ensemble we used Biba &amp; Boba’s Seresnext101_32x4d model (0.418 LB) and Fat Cat’s model (0.425 LB). Our team explored two different ways of merging models.</p>\n<ul>\n<li>“Majority voting” - from each model in the ensemble top100 samples was found independently and then all the image ids were sorted based on the sum of the positions that this ids has in each model’s top100 list. Finally, the top 100 indices from the sorted list were taken. (0.433 on LB)</li>\n<li>“Concat” - concatenating normalized model’s embeddings. Before concat each embedding was multiplied by the weight given to its model. When all embeddings are concatenated we normalize the new big one. To new big embedding we applied two staged post processing described above. (0.455 on LB)</li>\n</ul>\n<p>The final ensemble were made with:<br>\nTwo EffientnetV2 M models trained with supervised contrastive pretraining + AdaCos<br>\nSeresnext101_32x4d and EfficientnetB6 “Encoder + GEM + ArcHead” models</p>\n<h3>Final Thoughts</h3>\n<p>It was both teams' first large-scale kaggle competition so a lot of different mistakes were made along the way. After all experiments here are the ingredients of what we believe could be very strong training pipeline</p>\n<ul>\n<li>Better cropping <a href=\"https://arxiv.org/abs/1906.06423\" target=\"_blank\">paper</a></li>\n<li>Artificially sampled Supervised contrastive pretraining with 8 cuts of different images (here we used 4) with 448 validation resolution and 20 epochs on clean data</li>\n<li>Finetuning with ArcFace 512 validation resolution and class balanced focal loss for 20 epoch with sgd without schedulers on clean data</li>\n<li>Same as previous previous step but with 736 and total data</li>\n<li>Ensemble by concatenation </li>\n<li>Post processing: qe + db expansion </li>\n</ul>",
      "rawMarkdown": "Thanks to Google and Kaggle for organizing this interesting competition! \nCongratulations to the winners!\n\nIt was the first large-scale competition for both teams @ofitserovlad, @evgenysidorov, @joven1997 and we spent a lot of time on the experiments and came up with a strong pipeline pretty close to the end of the competition. Even though we finished just short of the gold medal we are still more than satisfied with the results and enjoyed the competition immensely.\n\n# Solution Overview\n![](https://user-images.githubusercontent.com/14181915/135837030-4bdf97b7-aa89-42db-9b2c-f97c837cdb58.png)\n\n# 1. Supervised Contrastive Pretraining\n### Models\nEfficientnetv2_m\n\n### Validation \nmAP calculated with Index (100k sampled) and Test (1k) data from google’s repository. CV and LB had good correlation with LB being almost always 0.015-0.02 lower than CV (ex. CV 0.380 -> LB 0.360)\n\n### Training Pipeline \n![](https://user-images.githubusercontent.com/14181915/135837248-2b9063ec-aaed-40a1-987b-2bc3f5181538.png)\n\nTeam Fat Cat finalized a pretty strong 3 stage training process a bit too late (1.5 week before the end) into the competition:\nArtificially Sampled Supervised Contrastive Pretraining with class balanced temperature on 448 resolution and clean dataset \nFinetuning with AdaCos with class balanced Focal Loss with 512 resolution on clean dataset \nSame as stage 2 but with 604 resolution and full dataset \n\n### Artificially Sampled Supervised Contrastive Pretraining with class balanced temperature \nPretraining part was inspired by this [paper](https://arxiv.org/abs/2004.11362). At this stage the model should learn easy image features and be more generalizable before finetuning stage. The premise of the paper is pretty similar to SimCLR, which takes 2 augmented (color jitter, random flip) random cuts of the image and makes them close in the cosine space. The extension of Supervised Contrastive paper is to make cuts of the image of the same class in the batch also close together. This paper also normalizes the image embeddings before feeding them into the nonlinear projection head. Outputs of the projection head are then used for supervised contrastive loss. \n\nThe main problem with implementation of the original paper was that they used enormous batch size (8192) and only 1000 classes and by doing so are somewhat guaranteed to have multiple instances of the same class in the batch. When I tried that kind of approach with a batch of 256 cv was at ~0.180, clearly indicating that there was very little probability of the same class occurring in the batch. \n\n1.5 weeks before the end I realized that I can artificially sample cuts of the different images of the same class into 1 batch. I shuffled my dataset in such a way that each 4 consecutive images were augmented cuts of the same class. This kind of pretraining resulted in cv mAP of .273 (and was much faster than cutting 2 parts of the same image), after 10 epochs of pretraining. Graphs showed that given more time it could reach .300. For pretraining I used SGD with cosine warmup and annealing. \n\nAdditionally, individual temperature for each class was used. It was derived using class-balanced strategy and omitting final normalization (i.e sum of the weights equals to the number of classes). \n\n### Finetuning \nFor finetuning Adaptive Cosine loss was used. It is an extension of AcrFace and was also used by the previous year’s winner. Supervised Contrastive pretraining was a crucial step for faster convergence during finetuning. Without pretraining after 1 epoch training on clean data and 512 resolution CV mAP was at .210 and with pretraining it started at 0.31 and reached 0.35 after 10 epochs with SGD and cosine scheduling. However, with better hyperparameters it potentially could have climbed even higher. Then, the resolution was increased to 604 and finetuned on the full dataset. \n\nAt the very end of the competition we added additional weight to landmarks from Asia, Africa, and Oceania to try to match the distribution of the test data which was described by Google's paper. It helped to bridge the gap between LB and CV. \n\n# 2. Encoder + GEM + ArcHead\n \n### Networks:\nEfficientnetB6, 704x704\nEfficientnetB7, 664x664\nSeresnext101_32x4d, 512x512\n\n### Training procedure:\nArcFace loss based on Focal Loss with various modifications was used to train the models. \nWe quickly realized that big models with big resolution give improvements up to 0.03-0.05.\nTraining big models was optimized by using FP16 precision, RandomCrop augmentations to reduce training image size (for example, from 704 to 664) and doing all the augmentations on GPU with Kornia.\nOne model training can be divided into two steps:\nTraining on GLDv2 cleaned\nFine tuning on GLDv2 full dataset with smooth labeling and doubled weights for cleaned samples to reduce effect of the noisiness \n \n### 3. Post Processing\nTwo stage post processing procedure was suggested by Fat Cat. It consisted of query expansion and database expansion:\n\n\nQuery expansion - for each sample from test set we find top2 samples from index set by cosine similarity and generate new embedding for test sample by normalizing sum of weighted by cosine similarity score top2 index samples and test sample\nDatabase expansion - for each test and index sample we find top16 samples from GLDv2 train set and based on their classes define the class of the test/index sample. Using this classes we re-rank top100 list of index samples for each test sample that we get after qe expansion. We put on top of the new list samples from the top100 list that have the same class as the test sample in the same order. Then we put samples from the index set that have the same class as the test sample but weren't in the top100 list before sorted by cosine similarity score, to the end of the new list we put all samples from the top100 list that have different class than the test sample. Finally, top100 samples from the new list are taken.\nThis post processing strategy improved our LB score by around 0.07.\n \n### Ensemble\nFor all the post processing procedures we used faiss to speed it up using GPU instead of CPU. For testing the ensemble we used Biba & Boba’s Seresnext101_32x4d model (0.418 LB) and Fat Cat’s model (0.425 LB). Our team explored two different ways of merging models.\n- “Majority voting” - from each model in the ensemble top100 samples was found independently and then all the image ids were sorted based on the sum of the positions that this ids has in each model’s top100 list. Finally, the top 100 indices from the sorted list were taken. (0.433 on LB)\n- “Concat” - concatenating normalized model’s embeddings. Before concat each embedding was multiplied by the weight given to its model. When all embeddings are concatenated we normalize the new big one. To new big embedding we applied two staged post processing described above. (0.455 on LB)\n \nThe final ensemble were made with:\nTwo EffientnetV2 M models trained with supervised contrastive pretraining + AdaCos\nSeresnext101_32x4d and EfficientnetB6 “Encoder + GEM + ArcHead” models\n\n \n### Final Thoughts \nIt was both teams' first large-scale kaggle competition so a lot of different mistakes were made along the way. After all experiments here are the ingredients of what we believe could be very strong training pipeline\n- Better cropping [paper](https://arxiv.org/abs/1906.06423)\n- Artificially sampled Supervised contrastive pretraining with 8 cuts of different images (here we used 4) with 448 validation resolution and 20 epochs on clean data\n- Finetuning with ArcFace 512 validation resolution and class balanced focal loss for 20 epoch with sgd without schedulers on clean data\n- Same as previous previous step but with 736 and total data\n- Ensemble by concatenation \n- Post processing: qe + db expansion",
      "votes": null
    },
    {
      "id": "1545621",
      "postDate": "10/15/2021 12:05:11",
      "content": "<p>Are you planning to release the code?<br>\nI was particularly interested in the Supervised Contrastive pretraining part</p>",
      "rawMarkdown": "Are you planning to release the code?\nI was particularly interested in the Supervised Contrastive pretraining part",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1545621,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "10/15/2021 12:05:11",
      "content": "<p>Are you planning to release the code?<br>\nI was particularly interested in the Supervised Contrastive pretraining part</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1533758": "Thanks to Google and Kaggle for organizing this interesting competition! \nCongratulations to the winners!\n\nIt was the first large-scale competition for both teams @ofitserovlad, @evgenysidorov, @joven1997 and we spent a lot of time on the experiments and came up with a strong pipeline pretty close to the end of the competition. Even though we finished just short of the gold medal we are still more than satisfied with the results and enjoyed the competition immensely.\n\n# Solution Overview\n![](https://user-images.githubusercontent.com/14181915/135837030-4bdf97b7-aa89-42db-9b2c-f97c837cdb58.png)\n\n# 1. Supervised Contrastive Pretraining\n### Models\nEfficientnetv2_m\n\n### Validation \nmAP calculated with Index (100k sampled) and Test (1k) data from google’s repository. CV and LB had good correlation with LB being almost always 0.015-0.02 lower than CV (ex. CV 0.380 -> LB 0.360)\n\n### Training Pipeline \n![](https://user-images.githubusercontent.com/14181915/135837248-2b9063ec-aaed-40a1-987b-2bc3f5181538.png)\n\nTeam Fat Cat finalized a pretty strong 3 stage training process a bit too late (1.5 week before the end) into the competition:\nArtificially Sampled Supervised Contrastive Pretraining with class balanced temperature on 448 resolution and clean dataset \nFinetuning with AdaCos with class balanced Focal Loss with 512 resolution on clean dataset \nSame as stage 2 but with 604 resolution and full dataset \n\n### Artificially Sampled Supervised Contrastive Pretraining with class balanced temperature \nPretraining part was inspired by this [paper](https://arxiv.org/abs/2004.11362). At this stage the model should learn easy image features and be more generalizable before finetuning stage. The premise of the paper is pretty similar to SimCLR, which takes 2 augmented (color jitter, random flip) random cuts of the image and makes them close in the cosine space. The extension of Supervised Contrastive paper is to make cuts of the image of the same class in the batch also close together. This paper also normalizes the image embeddings before feeding them into the nonlinear projection head. Outputs of the projection head are then used for supervised contrastive loss. \n\nThe main problem with implementation of the original paper was that they used enormous batch size (8192) and only 1000 classes and by doing so are somewhat guaranteed to have multiple instances of the same class in the batch. When I tried that kind of approach with a batch of 256 cv was at ~0.180, clearly indicating that there was very little probability of the same class occurring in the batch. \n\n1.5 weeks before the end I realized that I can artificially sample cuts of the different images of the same class into 1 batch. I shuffled my dataset in such a way that each 4 consecutive images were augmented cuts of the same class. This kind of pretraining resulted in cv mAP of .273 (and was much faster than cutting 2 parts of the same image), after 10 epochs of pretraining. Graphs showed that given more time it could reach .300. For pretraining I used SGD with cosine warmup and annealing. \n\nAdditionally, individual temperature for each class was used. It was derived using class-balanced strategy and omitting final normalization (i.e sum of the weights equals to the number of classes). \n\n### Finetuning \nFor finetuning Adaptive Cosine loss was used. It is an extension of AcrFace and was also used by the previous year’s winner. Supervised Contrastive pretraining was a crucial step for faster convergence during finetuning. Without pretraining after 1 epoch training on clean data and 512 resolution CV mAP was at .210 and with pretraining it started at 0.31 and reached 0.35 after 10 epochs with SGD and cosine scheduling. However, with better hyperparameters it potentially could have climbed even higher. Then, the resolution was increased to 604 and finetuned on the full dataset. \n\nAt the very end of the competition we added additional weight to landmarks from Asia, Africa, and Oceania to try to match the distribution of the test data which was described by Google's paper. It helped to bridge the gap between LB and CV. \n\n# 2. Encoder + GEM + ArcHead\n \n### Networks:\nEfficientnetB6, 704x704\nEfficientnetB7, 664x664\nSeresnext101_32x4d, 512x512\n\n### Training procedure:\nArcFace loss based on Focal Loss with various modifications was used to train the models. \nWe quickly realized that big models with big resolution give improvements up to 0.03-0.05.\nTraining big models was optimized by using FP16 precision, RandomCrop augmentations to reduce training image size (for example, from 704 to 664) and doing all the augmentations on GPU with Kornia.\nOne model training can be divided into two steps:\nTraining on GLDv2 cleaned\nFine tuning on GLDv2 full dataset with smooth labeling and doubled weights for cleaned samples to reduce effect of the noisiness \n \n### 3. Post Processing\nTwo stage post processing procedure was suggested by Fat Cat. It consisted of query expansion and database expansion:\n\n\nQuery expansion - for each sample from test set we find top2 samples from index set by cosine similarity and generate new embedding for test sample by normalizing sum of weighted by cosine similarity score top2 index samples and test sample\nDatabase expansion - for each test and index sample we find top16 samples from GLDv2 train set and based on their classes define the class of the test/index sample. Using this classes we re-rank top100 list of index samples for each test sample that we get after qe expansion. We put on top of the new list samples from the top100 list that have the same class as the test sample in the same order. Then we put samples from the index set that have the same class as the test sample but weren't in the top100 list before sorted by cosine similarity score, to the end of the new list we put all samples from the top100 list that have different class than the test sample. Finally, top100 samples from the new list are taken.\nThis post processing strategy improved our LB score by around 0.07.\n \n### Ensemble\nFor all the post processing procedures we used faiss to speed it up using GPU instead of CPU. For testing the ensemble we used Biba & Boba’s Seresnext101_32x4d model (0.418 LB) and Fat Cat’s model (0.425 LB). Our team explored two different ways of merging models.\n- “Majority voting” - from each model in the ensemble top100 samples was found independently and then all the image ids were sorted based on the sum of the positions that this ids has in each model’s top100 list. Finally, the top 100 indices from the sorted list were taken. (0.433 on LB)\n- “Concat” - concatenating normalized model’s embeddings. Before concat each embedding was multiplied by the weight given to its model. When all embeddings are concatenated we normalize the new big one. To new big embedding we applied two staged post processing described above. (0.455 on LB)\n \nThe final ensemble were made with:\nTwo EffientnetV2 M models trained with supervised contrastive pretraining + AdaCos\nSeresnext101_32x4d and EfficientnetB6 “Encoder + GEM + ArcHead” models\n\n \n### Final Thoughts \nIt was both teams' first large-scale kaggle competition so a lot of different mistakes were made along the way. After all experiments here are the ingredients of what we believe could be very strong training pipeline\n- Better cropping [paper](https://arxiv.org/abs/1906.06423)\n- Artificially sampled Supervised contrastive pretraining with 8 cuts of different images (here we used 4) with 448 validation resolution and 20 epochs on clean data\n- Finetuning with ArcFace 512 validation resolution and class balanced focal loss for 20 epoch with sgd without schedulers on clean data\n- Same as previous previous step but with 736 and total data\n- Ensemble by concatenation \n- Post processing: qe + db expansion",
    "1545621": "Are you planning to release the code?\nI was particularly interested in the Supervised Contrastive pretraining part"
  },
  "source": "meta"
}