{
  "id": 242087,
  "title": "1st place solution: Swin + Arcface + Label-constrained DBA",
  "url": "/competitions/hotel-id-2021-fgvc8/discussion/242087",
  "author_name": "Kohei",
  "post_date": "2021-05-27T13:01:58.479000",
  "votes": 50,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to the organizers for this interesting challenge! I am very proud to have participated in such a socially meaningful competition.</p>\n<p>Code: <a href=\"https://github.com/smly/hotelid-2021-first-place-solution\" target=\"_blank\">https://github.com/smly/hotelid-2021-first-place-solution</a></p>\n<h2>Preprocessing: Rotation Correction</h2>\n<p>Many Traffickcam images do not have the proper rotation angle. I created a classification model for correcting the rotation angle and applied it to all traffickcam images. This model was trained on the Hotels50k dataset. In the end, the rotation angles of 4859 out of 97554 training images were corrected.</p>\n<h2>Modeling</h2>\n<ul>\n<li>arch: Arcface (backbone: Swin, RegNetY, ResNeSt)</li>\n<li>optimizer: AdamW, scheduler=CosineAnnealing</li>\n<li>data augmentation: RandomResizedCrop, Resize, HorizontalFlip, RandomBrightness, RandomContrast, RandomGamma, ShiftScaleRotate and Cutout</li>\n</ul>\n<p>I solved this contest by metric learning and nearest neighbor search. The image representations were trained by Arcface. For the backbone, I used ResNeSt101e, RegNetY120, and Swin Transformer (Swin-L). All models I trained scores almost the same in public LB.</p>\n<p>For imbalances, in the nearest neighbor search, only the TOP1 nearest neighbor similarity score for each hotel was used as confidence. When aggregated with the sum of topk scores, the result is worse.</p>\n<h2>Label-constrained Database Augmentation</h2>\n<p>By updating the image representations with Database augmentation (k=5), the LB score can be improved. However, DBA is easily affected by different classes of image representations because there are few similar images. Thus, I added the following label constraint into DBA. (<a href=\"https://gist.github.com/smly/ce2457841941789cc4b9dcc670765dca\" target=\"_blank\">code</a>)</p>\n<p><img src=\"https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212650.png\" alt=\"\"></p>\n<h2>Database Expansion</h2>\n<p>I performed a search using Hotels50k as an index with Hotel-ID as the query and mapped the hotels. Adding H50k images to the training set to the index set in nearest neighbor search further improved the classification accuracy.</p>\n<p><img src=\"https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212607.png\" alt=\"\"></p>\n<h2>Ablation Study</h2>\n<h3>Single model performance</h3>\n<p>I used three backbone, ResNeSt101e, RegNetY120, and Swin Transformer, to create models with various combinations of training and index sets. The major improvement is from the difference in the training set and index set, not from the modeling part.</p>\n<p>HID+H50k(a) indicates the image set with added traffickcam images in H50k.  HID+H50k(b) indicates the image set with added traffickcam &amp; travel site images in H50k. HID+H50k(a) and HID+H50k(b) correspond to <code>train_hotelidv3.csv</code> and <code>train_hotelidv4.csv</code> in <a href=\"https://www.kaggle.com/confirm/hotelid-models\" target=\"_blank\">https://www.kaggle.com/confirm/hotelid-models</a> respectively.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>input size</th>\n<th>train set</th>\n<th>index set</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNeSt101e</td>\n<td>512</td>\n<td>H50k-&gt;HID</td>\n<td>HID only</td>\n<td>0.7852</td>\n<td>0.8027</td>\n</tr>\n<tr>\n<td>ResNeSt101e</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(a)</td>\n<td>0.8197</td>\n<td>0.8250</td>\n</tr>\n<tr>\n<td>ResNeSt101e</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(b)</td>\n<td>tbd</td>\n<td>tbd</td>\n</tr>\n<tr>\n<td>RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(a)</td>\n<td>tbd</td>\n<td>tbd</td>\n</tr>\n<tr>\n<td>RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(a)</td>\n<td>best</td>\n<td>best</td>\n</tr>\n<tr>\n<td>Swin</td>\n<td>384</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(a)</td>\n<td>0.8126</td>\n<td>0.8218</td>\n</tr>\n<tr>\n<td>Swin</td>\n<td>384</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(b)</td>\n<td>0.8181</td>\n<td>0.8259</td>\n</tr>\n</tbody>\n</table>\n<h3>Ensemble</h3>\n<p>I checked the performance contribution on the ensemble by removing one of the four models I created. The lower score indicates that the model is more important to the ensemble. It can be seen that Swin makes a significant contribution to the ensemble.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>input size</th>\n<th>train set</th>\n<th>index set</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>- ResNeSt101e</td>\n<td>512</td>\n<td>H50k-&gt;HID+H50k(a)</td>\n<td>HID+H50k(b)</td>\n<td>0.8502</td>\n<td>0.8551</td>\n</tr>\n<tr>\n<td>- RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(b)</td>\n<td>0.8525</td>\n<td>0.8594</td>\n</tr>\n<tr>\n<td>- RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(b)</td>\n<td>0.8523</td>\n<td>0.8589</td>\n</tr>\n<tr>\n<td>- Swin</td>\n<td>384</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(b)</td>\n<td><strong>0.8442</strong></td>\n<td><strong>0.8508</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Database Augmentation</h3>\n<p>The table below compares the effect of DBA after ensemble.</p>\n<p>HID+H50k(b) gives higher accuracy, but the number of images in the index set is large. This causes an out-of-memory error when trying to do DBA.</p>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Index set</th>\n<th>DBA method</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>#1</td>\n<td>HID+H50k(a)</td>\n<td>No QE&amp;DBA</td>\n<td>0.8514</td>\n<td>0.8557</td>\n</tr>\n<tr>\n<td>#2</td>\n<td>HID+H50k(a)</td>\n<td>QE&amp;DBA*2</td>\n<td>0.8527</td>\n<td>0.8570</td>\n</tr>\n<tr>\n<td>#3</td>\n<td>HID+H50k(a)</td>\n<td>QE&amp;LC-DBA*2</td>\n<td><strong>0.8576</strong></td>\n<td><strong>0.8596</strong></td>\n</tr>\n<tr>\n<td>#4</td>\n<td>HID+H50k(b)</td>\n<td>No QE&amp;DBA</td>\n<td>0.8541</td>\n<td>0.8595</td>\n</tr>\n<tr>\n<td>#5</td>\n<td>HID+H50k(b)</td>\n<td>QE&amp;DBA*2</td>\n<td>OOM Error</td>\n<td>OOM Error</td>\n</tr>\n<tr>\n<td>#6</td>\n<td>HID+H50k(b)</td>\n<td>QE&amp;LC-DBA*2</td>\n<td>OOM Error</td>\n<td>OOM Error</td>\n</tr>\n</tbody>\n</table>\n<p>Finally, my best submission is the simple average of (#3) and (#4). It scores 0.859 in public LB.</p>\n<h2>Things that did not work well</h2>\n<p>I tried to add a chain classifier head, but it didn't work well. Since the number of images for many hotels is very small, I assumed that the information about the same chain would be useful in training the model. I think the reason why it did not work was that the diversity of images in the same chain was too large.</p>\n<ul>\n<li>OSME block</li>\n<li>chain_id Head</li>\n<li>DELG descriptors</li>\n<li>Reranking Transformer (My code might be wrong.)</li>\n<li>Scoring similarities with using local feature matching</li>\n<li>TransReID</li>\n<li>Graph NN</li>\n<li>Graph-based SSL</li>\n</ul>",
  "messages": [
    {
      "id": 1325037,
      "postDate": "2021-05-27T13:01:58.480Z",
      "content": "<p>Thanks to the organizers for this interesting challenge! I am very proud to have participated in such a socially meaningful competition.</p>\n<p>Code: <a href=\"https://github.com/smly/hotelid-2021-first-place-solution\" target=\"_blank\">https://github.com/smly/hotelid-2021-first-place-solution</a></p>\n<h2>Preprocessing: Rotation Correction</h2>\n<p>Many Traffickcam images do not have the proper rotation angle. I created a classification model for correcting the rotation angle and applied it to all traffickcam images. This model was trained on the Hotels50k dataset. In the end, the rotation angles of 4859 out of 97554 training images were corrected.</p>\n<h2>Modeling</h2>\n<ul>\n<li>arch: Arcface (backbone: Swin, RegNetY, ResNeSt)</li>\n<li>optimizer: AdamW, scheduler=CosineAnnealing</li>\n<li>data augmentation: RandomResizedCrop, Resize, HorizontalFlip, RandomBrightness, RandomContrast, RandomGamma, ShiftScaleRotate and Cutout</li>\n</ul>\n<p>I solved this contest by metric learning and nearest neighbor search. The image representations were trained by Arcface. For the backbone, I used ResNeSt101e, RegNetY120, and Swin Transformer (Swin-L). All models I trained scores almost the same in public LB.</p>\n<p>For imbalances, in the nearest neighbor search, only the TOP1 nearest neighbor similarity score for each hotel was used as confidence. When aggregated with the sum of topk scores, the result is worse.</p>\n<h2>Label-constrained Database Augmentation</h2>\n<p>By updating the image representations with Database augmentation (k=5), the LB score can be improved. However, DBA is easily affected by different classes of image representations because there are few similar images. Thus, I added the following label constraint into DBA. (<a href=\"https://gist.github.com/smly/ce2457841941789cc4b9dcc670765dca\" target=\"_blank\">code</a>)</p>\n<p><img src=\"https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212650.png\" alt=\"\"></p>\n<h2>Database Expansion</h2>\n<p>I performed a search using Hotels50k as an index with Hotel-ID as the query and mapped the hotels. Adding H50k images to the training set to the index set in nearest neighbor search further improved the classification accuracy.</p>\n<p><img src=\"https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212607.png\" alt=\"\"></p>\n<h2>Ablation Study</h2>\n<h3>Single model performance</h3>\n<p>I used three backbone, ResNeSt101e, RegNetY120, and Swin Transformer, to create models with various combinations of training and index sets. The major improvement is from the difference in the training set and index set, not from the modeling part.</p>\n<p>HID+H50k(a) indicates the image set with added traffickcam images in H50k.  HID+H50k(b) indicates the image set with added traffickcam &amp; travel site images in H50k. HID+H50k(a) and HID+H50k(b) correspond to <code>train_hotelidv3.csv</code> and <code>train_hotelidv4.csv</code> in <a href=\"https://www.kaggle.com/confirm/hotelid-models\" target=\"_blank\">https://www.kaggle.com/confirm/hotelid-models</a> respectively.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>input size</th>\n<th>train set</th>\n<th>index set</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNeSt101e</td>\n<td>512</td>\n<td>H50k-&gt;HID</td>\n<td>HID only</td>\n<td>0.7852</td>\n<td>0.8027</td>\n</tr>\n<tr>\n<td>ResNeSt101e</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(a)</td>\n<td>0.8197</td>\n<td>0.8250</td>\n</tr>\n<tr>\n<td>ResNeSt101e</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(b)</td>\n<td>tbd</td>\n<td>tbd</td>\n</tr>\n<tr>\n<td>RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(a)</td>\n<td>tbd</td>\n<td>tbd</td>\n</tr>\n<tr>\n<td>RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(a)</td>\n<td>best</td>\n<td>best</td>\n</tr>\n<tr>\n<td>Swin</td>\n<td>384</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(a)</td>\n<td>0.8126</td>\n<td>0.8218</td>\n</tr>\n<tr>\n<td>Swin</td>\n<td>384</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(b)</td>\n<td>0.8181</td>\n<td>0.8259</td>\n</tr>\n</tbody>\n</table>\n<h3>Ensemble</h3>\n<p>I checked the performance contribution on the ensemble by removing one of the four models I created. The lower score indicates that the model is more important to the ensemble. It can be seen that Swin makes a significant contribution to the ensemble.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>input size</th>\n<th>train set</th>\n<th>index set</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>- ResNeSt101e</td>\n<td>512</td>\n<td>H50k-&gt;HID+H50k(a)</td>\n<td>HID+H50k(b)</td>\n<td>0.8502</td>\n<td>0.8551</td>\n</tr>\n<tr>\n<td>- RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(a)</td>\n<td>HID+H50k(b)</td>\n<td>0.8525</td>\n<td>0.8594</td>\n</tr>\n<tr>\n<td>- RegNetY120</td>\n<td>512</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(b)</td>\n<td>0.8523</td>\n<td>0.8589</td>\n</tr>\n<tr>\n<td>- Swin</td>\n<td>384</td>\n<td>H50k→HID+H50k(b)</td>\n<td>HID+H50k(b)</td>\n<td><strong>0.8442</strong></td>\n<td><strong>0.8508</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Database Augmentation</h3>\n<p>The table below compares the effect of DBA after ensemble.</p>\n<p>HID+H50k(b) gives higher accuracy, but the number of images in the index set is large. This causes an out-of-memory error when trying to do DBA.</p>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Index set</th>\n<th>DBA method</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>#1</td>\n<td>HID+H50k(a)</td>\n<td>No QE&amp;DBA</td>\n<td>0.8514</td>\n<td>0.8557</td>\n</tr>\n<tr>\n<td>#2</td>\n<td>HID+H50k(a)</td>\n<td>QE&amp;DBA*2</td>\n<td>0.8527</td>\n<td>0.8570</td>\n</tr>\n<tr>\n<td>#3</td>\n<td>HID+H50k(a)</td>\n<td>QE&amp;LC-DBA*2</td>\n<td><strong>0.8576</strong></td>\n<td><strong>0.8596</strong></td>\n</tr>\n<tr>\n<td>#4</td>\n<td>HID+H50k(b)</td>\n<td>No QE&amp;DBA</td>\n<td>0.8541</td>\n<td>0.8595</td>\n</tr>\n<tr>\n<td>#5</td>\n<td>HID+H50k(b)</td>\n<td>QE&amp;DBA*2</td>\n<td>OOM Error</td>\n<td>OOM Error</td>\n</tr>\n<tr>\n<td>#6</td>\n<td>HID+H50k(b)</td>\n<td>QE&amp;LC-DBA*2</td>\n<td>OOM Error</td>\n<td>OOM Error</td>\n</tr>\n</tbody>\n</table>\n<p>Finally, my best submission is the simple average of (#3) and (#4). It scores 0.859 in public LB.</p>\n<h2>Things that did not work well</h2>\n<p>I tried to add a chain classifier head, but it didn't work well. Since the number of images for many hotels is very small, I assumed that the information about the same chain would be useful in training the model. I think the reason why it did not work was that the diversity of images in the same chain was too large.</p>\n<ul>\n<li>OSME block</li>\n<li>chain_id Head</li>\n<li>DELG descriptors</li>\n<li>Reranking Transformer (My code might be wrong.)</li>\n<li>Scoring similarities with using local feature matching</li>\n<li>TransReID</li>\n<li>Graph NN</li>\n<li>Graph-based SSL</li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers for this interesting challenge! I am very proud to have participated in such a socially meaningful competition.\n\nCode: https://github.com/smly/hotelid-2021-first-place-solution\n\n## Preprocessing: Rotation Correction\n\nMany Traffickcam images do not have the proper rotation angle. I created a classification model for correcting the rotation angle and applied it to all traffickcam images. This model was trained on the Hotels50k dataset. In the end, the rotation angles of 4859 out of 97554 training images were corrected.\n\n## Modeling\n\n* arch: Arcface (backbone: Swin, RegNetY, ResNeSt)\n* optimizer: AdamW, scheduler=CosineAnnealing\n* data augmentation: RandomResizedCrop, Resize, HorizontalFlip, RandomBrightness, RandomContrast, RandomGamma, ShiftScaleRotate and Cutout\n\nI solved this contest by metric learning and nearest neighbor search. The image representations were trained by Arcface. For the backbone, I used ResNeSt101e, RegNetY120, and Swin Transformer (Swin-L). All models I trained scores almost the same in public LB.\n\nFor imbalances, in the nearest neighbor search, only the TOP1 nearest neighbor similarity score for each hotel was used as confidence. When aggregated with the sum of topk scores, the result is worse.\n\n## Label-constrained Database Augmentation\n\nBy updating the image representations with Database augmentation (k=5), the LB score can be improved. However, DBA is easily affected by different classes of image representations because there are few similar images. Thus, I added the following label constraint into DBA. ([code](https://gist.github.com/smly/ce2457841941789cc4b9dcc670765dca))\n\n![](https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212650.png)\n\n## Database Expansion\n\nI performed a search using Hotels50k as an index with Hotel-ID as the query and mapped the hotels. Adding H50k images to the training set to the index set in nearest neighbor search further improved the classification accuracy.\n\n![](https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212607.png)\n\n## Ablation Study\n\n### Single model performance\n\nI used three backbone, ResNeSt101e, RegNetY120, and Swin Transformer, to create models with various combinations of training and index sets. The major improvement is from the difference in the training set and index set, not from the modeling part.\n\nHID+H50k(a) indicates the image set with added traffickcam images in H50k.  HID+H50k(b) indicates the image set with added traffickcam & travel site images in H50k. HID+H50k(a) and HID+H50k(b) correspond to `train_hotelidv3.csv` and `train_hotelidv4.csv` in https://www.kaggle.com/confirm/hotelid-models respectively.\n\n| Backbone | input size | train set | index set | public LB | private LB |\n|------|---------|------|--------|------------|----------|\n| ResNeSt101e | 512 | H50k->HID | HID only | 0.7852 | 0.8027 |\n| ResNeSt101e | 512 | H50k→HID+H50k(a) | HID+H50k(a) | 0.8197 | 0.8250 |\n| ResNeSt101e | 512 | H50k→HID+H50k(a) | HID+H50k(b) | tbd | tbd |\n| RegNetY120 | 512 | H50k→HID+H50k(a) | HID+H50k(a) | tbd | tbd |\n| RegNetY120 | 512 | H50k→HID+H50k(b) | HID+H50k(a) | best | best |\n| Swin | 384 | H50k→HID+H50k(b) | HID+H50k(a) | 0.8126 | 0.8218 |\n| Swin | 384 | H50k→HID+H50k(b) | HID+H50k(b) | 0.8181 | 0.8259 |\n\n### Ensemble\n\nI checked the performance contribution on the ensemble by removing one of the four models I created. The lower score indicates that the model is more important to the ensemble. It can be seen that Swin makes a significant contribution to the ensemble.\n\n| Backbone | input size | train set | index set | public LB | private LB |\n|----------|-----------|----------|------------|---------|------------|\n| - ResNeSt101e | 512 | H50k->HID+H50k(a) | HID+H50k(b) | 0.8502 | 0.8551 |\n| - RegNetY120 | 512 | H50k→HID+H50k(a) | HID+H50k(b) | 0.8525 | 0.8594 |\n| - RegNetY120 | 512 | H50k→HID+H50k(b) | HID+H50k(b) | 0.8523 | 0.8589 |\n| - Swin | 384 | H50k→HID+H50k(b) | HID+H50k(b) | **0.8442** | **0.8508** |\n\n### Database Augmentation\n\nThe table below compares the effect of DBA after ensemble.\n\nHID+H50k(b) gives higher accuracy, but the number of images in the index set is large. This causes an out-of-memory error when trying to do DBA.\n\n| No | Index set | DBA method | public LB | private LB |\n|-----|----------|-----------|----------|--------------|\n| #1  | HID+H50k(a) | No QE&DBA | 0.8514 | 0.8557 |\n| #2 | HID+H50k(a) | QE&DBA*2 | 0.8527 | 0.8570 |\n| #3 | HID+H50k(a) | QE&LC-DBA*2 | **0.8576** | **0.8596** |\n| #4 | HID+H50k(b) | No QE&DBA | 0.8541 | 0.8595 |\n| #5 | HID+H50k(b) | QE&DBA*2 | OOM Error | OOM Error |\n| #6 | HID+H50k(b) | QE&LC-DBA*2 | OOM Error | OOM Error |\n\nFinally, my best submission is the simple average of (#3) and (#4). It scores 0.859 in public LB.\n\n## Things that did not work well\n\nI tried to add a chain classifier head, but it didn't work well. Since the number of images for many hotels is very small, I assumed that the information about the same chain would be useful in training the model. I think the reason why it did not work was that the diversity of images in the same chain was too large.\n\n* OSME block\n* chain_id Head\n* DELG descriptors\n* Reranking Transformer (My code might be wrong.)\n* Scoring similarities with using local feature matching\n* TransReID\n* Graph NN\n* Graph-based SSL",
      "votes": 50
    },
    {
      "id": 1325443,
      "postDate": "2021-05-27T18:47:04.880Z",
      "content": "<p>Congratulations and thank you for sharing your approach!</p>\n<p>Do I get it right, that LCDBA is applied on training where it only takes neighbors with the same label into account? And for  inference only DBA is applied?</p>",
      "rawMarkdown": "Congratulations and thank you for sharing your approach!\n\nDo I get it right, that LCDBA is applied on training where it only takes neighbors with the same label into account? And for  inference only DBA is applied?",
      "votes": 1
    },
    {
      "id": 1325246,
      "postDate": "2021-05-27T15:39:07.267Z",
      "content": "<p>Congratulations，it's my first time to join a competition, and my arcface loss seems to meet some problem.<br>\nCan I learn it from yours? thank you very much.</p>",
      "rawMarkdown": "Congratulations，it's my first time to join a competition, and my arcface loss seems to meet some problem.\nCan I learn it from yours? thank you very much.",
      "votes": 1
    },
    {
      "id": 1373116,
      "postDate": "2021-07-02T08:29:38.467Z",
      "content": "<p><a href=\"https://www.kaggle.com/confirm\" target=\"_blank\">@confirm</a></p>\n<ol>\n<li>does the model at <a href=\"https://www.kaggle.com/confirm/hotelid-kohei\" target=\"_blank\">https://www.kaggle.com/confirm/hotelid-kohei</a> to make an app that predicts the hotel by a picture? </li>\n<li>is it some policy to obfuscate submission.csv file, so that no dummy can copy, run and submit a successful kernel?</li>\n</ol>",
      "rawMarkdown": "@confirm\n1. does the model at https://www.kaggle.com/confirm/hotelid-kohei to make an app that predicts the hotel by a picture? \n2. is it some policy to obfuscate submission.csv file, so that no dummy can copy, run and submit a successful kernel?"
    },
    {
      "id": 1325177,
      "postDate": "2021-05-27T14:54:01.547Z",
      "content": "<p>Congratulations on winning, learned a lot！</p>",
      "rawMarkdown": "Congratulations on winning, learned a lot！"
    },
    {
      "id": 1325176,
      "postDate": "2021-05-27T14:53:56.107Z",
      "content": "<p>Congratulations!! Nice solution. I have one question about it. What is meant by cosine head? Thanks</p>",
      "rawMarkdown": "Congratulations!! Nice solution. I have one question about it. What is meant by cosine head? Thanks"
    },
    {
      "id": 1748692,
      "postDate": "2022-04-07T21:52:46.263Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1325443,
      "author_name": "Jo Tom",
      "author_url": "",
      "post_date": "2021-05-27T18:47:04.880000",
      "content": "<p>Congratulations and thank you for sharing your approach!</p>\n<p>Do I get it right, that LCDBA is applied on training where it only takes neighbors with the same label into account? And for  inference only DBA is applied?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1325246,
      "author_name": "Xingkui Zhu",
      "author_url": "",
      "post_date": "2021-05-27T15:39:07.267000",
      "content": "<p>Congratulations，it's my first time to join a competition, and my arcface loss seems to meet some problem.<br>\nCan I learn it from yours? thank you very much.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1373116,
      "author_name": "Siarhei Siniak",
      "author_url": "",
      "post_date": "2021-07-02T08:29:38.467000",
      "content": "<p><a href=\"https://www.kaggle.com/confirm\" target=\"_blank\">@confirm</a></p>\n<ol>\n<li>does the model at <a href=\"https://www.kaggle.com/confirm/hotelid-kohei\" target=\"_blank\">https://www.kaggle.com/confirm/hotelid-kohei</a> to make an app that predicts the hotel by a picture? </li>\n<li>is it some policy to obfuscate submission.csv file, so that no dummy can copy, run and submit a successful kernel?</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1325177,
      "author_name": "LoveLetter",
      "author_url": "",
      "post_date": "2021-05-27T14:54:01.547000",
      "content": "<p>Congratulations on winning, learned a lot！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1325176,
      "author_name": "adas_osus",
      "author_url": "",
      "post_date": "2021-05-27T14:53:56.107000",
      "content": "<p>Congratulations!! Nice solution. I have one question about it. What is meant by cosine head? Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1748692,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-07T21:52:46.263000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1325037": "Thanks to the organizers for this interesting challenge! I am very proud to have participated in such a socially meaningful competition.\n\nCode: https://github.com/smly/hotelid-2021-first-place-solution\n\n## Preprocessing: Rotation Correction\n\nMany Traffickcam images do not have the proper rotation angle. I created a classification model for correcting the rotation angle and applied it to all traffickcam images. This model was trained on the Hotels50k dataset. In the end, the rotation angles of 4859 out of 97554 training images were corrected.\n\n## Modeling\n\n* arch: Arcface (backbone: Swin, RegNetY, ResNeSt)\n* optimizer: AdamW, scheduler=CosineAnnealing\n* data augmentation: RandomResizedCrop, Resize, HorizontalFlip, RandomBrightness, RandomContrast, RandomGamma, ShiftScaleRotate and Cutout\n\nI solved this contest by metric learning and nearest neighbor search. The image representations were trained by Arcface. For the backbone, I used ResNeSt101e, RegNetY120, and Swin Transformer (Swin-L). All models I trained scores almost the same in public LB.\n\nFor imbalances, in the nearest neighbor search, only the TOP1 nearest neighbor similarity score for each hotel was used as confidence. When aggregated with the sum of topk scores, the result is worse.\n\n## Label-constrained Database Augmentation\n\nBy updating the image representations with Database augmentation (k=5), the LB score can be improved. However, DBA is easily affected by different classes of image representations because there are few similar images. Thus, I added the following label constraint into DBA. ([code](https://gist.github.com/smly/ce2457841941789cc4b9dcc670765dca))\n\n![](https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212650.png)\n\n## Database Expansion\n\nI performed a search using Hotels50k as an index with Hotel-ID as the query and mapped the hotels. Adding H50k images to the training set to the index set in nearest neighbor search further improved the classification accuracy.\n\n![](https://cdn-ak.f.st-hatena.com/images/fotolife/s/smly/20210527/20210527212607.png)\n\n## Ablation Study\n\n### Single model performance\n\nI used three backbone, ResNeSt101e, RegNetY120, and Swin Transformer, to create models with various combinations of training and index sets. The major improvement is from the difference in the training set and index set, not from the modeling part.\n\nHID+H50k(a) indicates the image set with added traffickcam images in H50k.  HID+H50k(b) indicates the image set with added traffickcam & travel site images in H50k. HID+H50k(a) and HID+H50k(b) correspond to `train_hotelidv3.csv` and `train_hotelidv4.csv` in https://www.kaggle.com/confirm/hotelid-models respectively.\n\n| Backbone | input size | train set | index set | public LB | private LB |\n|------|---------|------|--------|------------|----------|\n| ResNeSt101e | 512 | H50k->HID | HID only | 0.7852 | 0.8027 |\n| ResNeSt101e | 512 | H50k→HID+H50k(a) | HID+H50k(a) | 0.8197 | 0.8250 |\n| ResNeSt101e | 512 | H50k→HID+H50k(a) | HID+H50k(b) | tbd | tbd |\n| RegNetY120 | 512 | H50k→HID+H50k(a) | HID+H50k(a) | tbd | tbd |\n| RegNetY120 | 512 | H50k→HID+H50k(b) | HID+H50k(a) | best | best |\n| Swin | 384 | H50k→HID+H50k(b) | HID+H50k(a) | 0.8126 | 0.8218 |\n| Swin | 384 | H50k→HID+H50k(b) | HID+H50k(b) | 0.8181 | 0.8259 |\n\n### Ensemble\n\nI checked the performance contribution on the ensemble by removing one of the four models I created. The lower score indicates that the model is more important to the ensemble. It can be seen that Swin makes a significant contribution to the ensemble.\n\n| Backbone | input size | train set | index set | public LB | private LB |\n|----------|-----------|----------|------------|---------|------------|\n| - ResNeSt101e | 512 | H50k->HID+H50k(a) | HID+H50k(b) | 0.8502 | 0.8551 |\n| - RegNetY120 | 512 | H50k→HID+H50k(a) | HID+H50k(b) | 0.8525 | 0.8594 |\n| - RegNetY120 | 512 | H50k→HID+H50k(b) | HID+H50k(b) | 0.8523 | 0.8589 |\n| - Swin | 384 | H50k→HID+H50k(b) | HID+H50k(b) | **0.8442** | **0.8508** |\n\n### Database Augmentation\n\nThe table below compares the effect of DBA after ensemble.\n\nHID+H50k(b) gives higher accuracy, but the number of images in the index set is large. This causes an out-of-memory error when trying to do DBA.\n\n| No | Index set | DBA method | public LB | private LB |\n|-----|----------|-----------|----------|--------------|\n| #1  | HID+H50k(a) | No QE&DBA | 0.8514 | 0.8557 |\n| #2 | HID+H50k(a) | QE&DBA*2 | 0.8527 | 0.8570 |\n| #3 | HID+H50k(a) | QE&LC-DBA*2 | **0.8576** | **0.8596** |\n| #4 | HID+H50k(b) | No QE&DBA | 0.8541 | 0.8595 |\n| #5 | HID+H50k(b) | QE&DBA*2 | OOM Error | OOM Error |\n| #6 | HID+H50k(b) | QE&LC-DBA*2 | OOM Error | OOM Error |\n\nFinally, my best submission is the simple average of (#3) and (#4). It scores 0.859 in public LB.\n\n## Things that did not work well\n\nI tried to add a chain classifier head, but it didn't work well. Since the number of images for many hotels is very small, I assumed that the information about the same chain would be useful in training the model. I think the reason why it did not work was that the diversity of images in the same chain was too large.\n\n* OSME block\n* chain_id Head\n* DELG descriptors\n* Reranking Transformer (My code might be wrong.)\n* Scoring similarities with using local feature matching\n* TransReID\n* Graph NN\n* Graph-based SSL",
    "1325443": "Congratulations and thank you for sharing your approach!\n\nDo I get it right, that LCDBA is applied on training where it only takes neighbors with the same label into account? And for  inference only DBA is applied?",
    "1325246": "Congratulations，it's my first time to join a competition, and my arcface loss seems to meet some problem.\nCan I learn it from yours? thank you very much.",
    "1373116": "@confirm\n1. does the model at https://www.kaggle.com/confirm/hotelid-kohei to make an app that predicts the hotel by a picture? \n2. is it some policy to obfuscate submission.csv file, so that no dummy can copy, run and submit a successful kernel?",
    "1325177": "Congratulations on winning, learned a lot！",
    "1325176": "Congratulations!! Nice solution. I have one question about it. What is meant by cosine head? Thanks",
    "1748692": ""
  }
}