{
  "id": 329299,
  "title": "1st place solution",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/329299",
  "author_name": "Shuhao",
  "post_date": "2022-06-06T02:45:52.742000",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<ul>\n<li><p>Optimization on single model:</p>\n<table>\n<thead>\n<tr>\n<th>change</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>swin-B&nbsp;224&nbsp;baseline</td>\n<td>0.77223</td>\n<td>0.78442</td>\n</tr>\n<tr>\n<td>add&nbsp;cross-entropy&nbsp;losses&nbsp;on&nbsp;genera,species,&nbsp;family</td>\n<td>0.7829</td>\n<td>0.79544</td>\n</tr>\n<tr>\n<td>learning&nbsp;rate&nbsp;from&nbsp;2e-4&nbsp;to&nbsp;5e-4</td>\n<td>0.79444</td>\n<td>0.80501</td>\n</tr>\n<tr>\n<td>+5crop&nbsp;(resize&nbsp;256,&nbsp;288,&nbsp;320,&nbsp;384,&nbsp;448&nbsp;and&nbsp;center&nbsp;crop&nbsp;224)</td>\n<td>0.80088</td>\n<td>0.80981</td>\n</tr>\n<tr>\n<td>replacing&nbsp;cross-entropy&nbsp;by&nbsp;subcenter-arcface&nbsp;with&nbsp;<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">dynamic</a>&nbsp;margins</td>\n<td>0.81532</td>\n<td>0.82267</td>\n</tr>\n<tr>\n<td>+&nbsp;extra&nbsp;fc&nbsp;by&nbsp;cross-entropy,&nbsp;add&nbsp;to&nbsp;subcenter-arcface&nbsp;output&nbsp;after&nbsp;softmax</td>\n<td>0.82201</td>\n<td>0.82929</td>\n</tr>\n<tr>\n<td>+&nbsp;data&nbsp;augmentation&nbsp;the&nbsp;same&nbsp;with&nbsp;Swin-Transformer</td>\n<td>0.8315</td>\n<td>0.83554</td>\n</tr>\n<tr>\n<td>from&nbsp;swin-B224&nbsp;to&nbsp;SwinB384</td>\n<td>0.8468</td>\n<td>0.85245</td>\n</tr>\n<tr>\n<td>+5crop&nbsp;(resize&nbsp;400,&nbsp;416,&nbsp;448,&nbsp;480,&nbsp;512&nbsp;and&nbsp;center&nbsp;crop&nbsp;384)</td>\n<td>0.84895</td>\n<td>0.85654</td>\n</tr>\n<tr>\n<td>+Freeze&nbsp;layers&nbsp;(not&nbsp;update&nbsp;parameters)&nbsp;from&nbsp;100&nbsp;to&nbsp;0</td>\n<td>0.85499</td>\n<td>0.86055</td>\n</tr>\n<tr>\n<td>+square&nbsp;resize&nbsp;and&nbsp;test&nbsp;crop&nbsp;from&nbsp;384</td>\n<td>0.85563</td>\n<td>0.86201</td>\n</tr>\n<tr>\n<td>+from&nbsp;swin&nbsp;base&nbsp;to&nbsp;swin&nbsp;V2</td>\n<td>0.85851</td>\n<td>0.86282</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>Backbones:</p>\n<table>\n<thead>\n<tr>\n<th>id</th>\n<th>backbone</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>swin-B</td>\n<td>0.85499</td>\n<td>0.86055</td>\n</tr>\n<tr>\n<td>1</td>\n<td>convnext-B</td>\n<td>0.85464</td>\n<td>0.85956</td>\n</tr>\n<tr>\n<td>2</td>\n<td>deit-iii&nbsp;B</td>\n<td>0.84423</td>\n<td>0.8495</td>\n</tr>\n<tr>\n<td>3</td>\n<td>resnest-101</td>\n<td>0.83208</td>\n<td>0.83838</td>\n</tr>\n<tr>\n<td>4</td>\n<td>efficient-B6</td>\n<td>0.83841</td>\n<td>0.84321</td>\n</tr>\n<tr>\n<td>5</td>\n<td>cswin-L</td>\n<td>0.84373</td>\n<td>0.8501</td>\n</tr>\n<tr>\n<td>6</td>\n<td>swin-L</td>\n<td>0.85057</td>\n<td>0.85607</td>\n</tr>\n<tr>\n<td>7</td>\n<td>swinv2-B</td>\n<td>0.85851</td>\n<td>0.86282</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>Fusion：<br>\nwe utilize the scores on public dataset as the weight. </p>\n<table>\n<thead>\n<tr>\n<th>fusion&nbsp;approach</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fusion&nbsp;models&nbsp;(0,1,2)</td>\n<td>0.86366</td>\n<td>0.8679</td>\n</tr>\n<tr>\n<td>fusion&nbsp;models&nbsp;(0,1,2,3,4)</td>\n<td>0.86619</td>\n<td>0.87185</td>\n</tr>\n<tr>\n<td>fusion&nbsp;models&nbsp;(0,1,2,3,4,5,6,7)</td>\n<td>0.87193</td>\n<td>0.87662</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>Other settings:<br>\nBatch size is always 512 (256 for big models), trained with 100epochs.<br>\ntorch.amp is adopted.<br>\nFollowing <a href=\"https://www.kaggle.com/competitions/herbarium-2021-fgvc8/discussion/242233\" target=\"_blank\">solutions</a>， we adopt similar Post-Process and obtain 0.001 improvement.<br>\nWe use Swin-B pretrained on Imagenet22k for initialization, for other backbones, we try to use Imagenet22k pretrained weights if possible.<br>\nWe freeze partial layers to enlarge the batchsize during the training process, We then reduce the batch size and train the full parameters. </p></li>\n<li><p>Something less effective:<br>\nclass-aware sampling<br>\ndata cleaning                                              </p></li>\n</ul>",
  "messages": [
    {
      "id": 1812566,
      "postDate": "2022-06-06T02:45:52.743Z",
      "content": "<ul>\n<li><p>Optimization on single model:</p>\n<table>\n<thead>\n<tr>\n<th>change</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>swin-B&nbsp;224&nbsp;baseline</td>\n<td>0.77223</td>\n<td>0.78442</td>\n</tr>\n<tr>\n<td>add&nbsp;cross-entropy&nbsp;losses&nbsp;on&nbsp;genera,species,&nbsp;family</td>\n<td>0.7829</td>\n<td>0.79544</td>\n</tr>\n<tr>\n<td>learning&nbsp;rate&nbsp;from&nbsp;2e-4&nbsp;to&nbsp;5e-4</td>\n<td>0.79444</td>\n<td>0.80501</td>\n</tr>\n<tr>\n<td>+5crop&nbsp;(resize&nbsp;256,&nbsp;288,&nbsp;320,&nbsp;384,&nbsp;448&nbsp;and&nbsp;center&nbsp;crop&nbsp;224)</td>\n<td>0.80088</td>\n<td>0.80981</td>\n</tr>\n<tr>\n<td>replacing&nbsp;cross-entropy&nbsp;by&nbsp;subcenter-arcface&nbsp;with&nbsp;<a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757\" target=\"_blank\">dynamic</a>&nbsp;margins</td>\n<td>0.81532</td>\n<td>0.82267</td>\n</tr>\n<tr>\n<td>+&nbsp;extra&nbsp;fc&nbsp;by&nbsp;cross-entropy,&nbsp;add&nbsp;to&nbsp;subcenter-arcface&nbsp;output&nbsp;after&nbsp;softmax</td>\n<td>0.82201</td>\n<td>0.82929</td>\n</tr>\n<tr>\n<td>+&nbsp;data&nbsp;augmentation&nbsp;the&nbsp;same&nbsp;with&nbsp;Swin-Transformer</td>\n<td>0.8315</td>\n<td>0.83554</td>\n</tr>\n<tr>\n<td>from&nbsp;swin-B224&nbsp;to&nbsp;SwinB384</td>\n<td>0.8468</td>\n<td>0.85245</td>\n</tr>\n<tr>\n<td>+5crop&nbsp;(resize&nbsp;400,&nbsp;416,&nbsp;448,&nbsp;480,&nbsp;512&nbsp;and&nbsp;center&nbsp;crop&nbsp;384)</td>\n<td>0.84895</td>\n<td>0.85654</td>\n</tr>\n<tr>\n<td>+Freeze&nbsp;layers&nbsp;(not&nbsp;update&nbsp;parameters)&nbsp;from&nbsp;100&nbsp;to&nbsp;0</td>\n<td>0.85499</td>\n<td>0.86055</td>\n</tr>\n<tr>\n<td>+square&nbsp;resize&nbsp;and&nbsp;test&nbsp;crop&nbsp;from&nbsp;384</td>\n<td>0.85563</td>\n<td>0.86201</td>\n</tr>\n<tr>\n<td>+from&nbsp;swin&nbsp;base&nbsp;to&nbsp;swin&nbsp;V2</td>\n<td>0.85851</td>\n<td>0.86282</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>Backbones:</p>\n<table>\n<thead>\n<tr>\n<th>id</th>\n<th>backbone</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>swin-B</td>\n<td>0.85499</td>\n<td>0.86055</td>\n</tr>\n<tr>\n<td>1</td>\n<td>convnext-B</td>\n<td>0.85464</td>\n<td>0.85956</td>\n</tr>\n<tr>\n<td>2</td>\n<td>deit-iii&nbsp;B</td>\n<td>0.84423</td>\n<td>0.8495</td>\n</tr>\n<tr>\n<td>3</td>\n<td>resnest-101</td>\n<td>0.83208</td>\n<td>0.83838</td>\n</tr>\n<tr>\n<td>4</td>\n<td>efficient-B6</td>\n<td>0.83841</td>\n<td>0.84321</td>\n</tr>\n<tr>\n<td>5</td>\n<td>cswin-L</td>\n<td>0.84373</td>\n<td>0.8501</td>\n</tr>\n<tr>\n<td>6</td>\n<td>swin-L</td>\n<td>0.85057</td>\n<td>0.85607</td>\n</tr>\n<tr>\n<td>7</td>\n<td>swinv2-B</td>\n<td>0.85851</td>\n<td>0.86282</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>Fusion：<br>\nwe utilize the scores on public dataset as the weight. </p>\n<table>\n<thead>\n<tr>\n<th>fusion&nbsp;approach</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fusion&nbsp;models&nbsp;(0,1,2)</td>\n<td>0.86366</td>\n<td>0.8679</td>\n</tr>\n<tr>\n<td>fusion&nbsp;models&nbsp;(0,1,2,3,4)</td>\n<td>0.86619</td>\n<td>0.87185</td>\n</tr>\n<tr>\n<td>fusion&nbsp;models&nbsp;(0,1,2,3,4,5,6,7)</td>\n<td>0.87193</td>\n<td>0.87662</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>Other settings:<br>\nBatch size is always 512 (256 for big models), trained with 100epochs.<br>\ntorch.amp is adopted.<br>\nFollowing <a href=\"https://www.kaggle.com/competitions/herbarium-2021-fgvc8/discussion/242233\" target=\"_blank\">solutions</a>， we adopt similar Post-Process and obtain 0.001 improvement.<br>\nWe use Swin-B pretrained on Imagenet22k for initialization, for other backbones, we try to use Imagenet22k pretrained weights if possible.<br>\nWe freeze partial layers to enlarge the batchsize during the training process, We then reduce the batch size and train the full parameters. </p></li>\n<li><p>Something less effective:<br>\nclass-aware sampling<br>\ndata cleaning                                              </p></li>\n</ul>",
      "rawMarkdown": "- Optimization on single model:\n|change|public|private|\n|:--|:--|:--|\n|swin-B&nbsp;224&nbsp;baseline|0.77223|0.78442|\n|add&nbsp;cross-entropy&nbsp;losses&nbsp;on&nbsp;genera,species,&nbsp;family|0.7829|0.79544|\n|learning&nbsp;rate&nbsp;from&nbsp;2e-4&nbsp;to&nbsp;5e-4|0.79444|0.80501|\n|+5crop&nbsp;(resize&nbsp;256,&nbsp;288,&nbsp;320,&nbsp;384,&nbsp;448&nbsp;and&nbsp;center&nbsp;crop&nbsp;224)|0.80088|0.80981|\n|replacing&nbsp;cross-entropy&nbsp;by&nbsp;subcenter-arcface&nbsp;with&nbsp;[dynamic](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757)&nbsp;margins|0.81532|0.82267|\n|+&nbsp;extra&nbsp;fc&nbsp;by&nbsp;cross-entropy,&nbsp;add&nbsp;to&nbsp;subcenter-arcface&nbsp;output&nbsp;after&nbsp;softmax|0.82201|0.82929|\n|+&nbsp;data&nbsp;augmentation&nbsp;the&nbsp;same&nbsp;with&nbsp;Swin-Transformer|0.8315|0.83554|\n|from&nbsp;swin-B224&nbsp;to&nbsp;SwinB384|0.8468|0.85245|\n|+5crop&nbsp;(resize&nbsp;400,&nbsp;416,&nbsp;448,&nbsp;480,&nbsp;512&nbsp;and&nbsp;center&nbsp;crop&nbsp;384)|0.84895|0.85654|\n|+Freeze&nbsp;layers&nbsp;(not&nbsp;update&nbsp;parameters)&nbsp;from&nbsp;100&nbsp;to&nbsp;0|0.85499|0.86055|\n|+square&nbsp;resize&nbsp;and&nbsp;test&nbsp;crop&nbsp;from&nbsp;384|0.85563|0.86201|\n|+from&nbsp;swin&nbsp;base&nbsp;to&nbsp;swin&nbsp;V2|0.85851|0.86282|\n\n\n- Backbones:\n|id|backbone|public|private|\n|:--|:--|:--|:--|\n|0|swin-B|0.85499|0.86055|\n|1|convnext-B|0.85464|0.85956|\n|2|deit-iii&nbsp;B|0.84423|0.8495|\n|3|resnest-101|0.83208|0.83838|\n|4|efficient-B6|0.83841|0.84321|\n|5|cswin-L|0.84373|0.8501|\n|6|swin-L|0.85057|0.85607|\n|7|swinv2-B|0.85851|0.86282|\n\n\n- Fusion：\nwe utilize the scores on public dataset as the weight. \n|fusion&nbsp;approach|public|private|\n|:--|:--|:--|\n|fusion&nbsp;models&nbsp;(0,1,2)|0.86366|0.8679|\n|fusion&nbsp;models&nbsp;(0,1,2,3,4)|0.86619|0.87185|\n|fusion&nbsp;models&nbsp;(0,1,2,3,4,5,6,7)|0.87193|0.87662|\n\n\n- Other settings:\nBatch size is always 512 (256 for big models), trained with 100epochs.\ntorch.amp is adopted.\nFollowing [solutions](https://www.kaggle.com/competitions/herbarium-2021-fgvc8/discussion/242233)， we adopt similar Post-Process and obtain 0.001 improvement.\nWe use Swin-B pretrained on Imagenet22k for initialization, for other backbones, we try to use Imagenet22k pretrained weights if possible.\nWe freeze partial layers to enlarge the batchsize during the training process, We then reduce the batch size and train the full parameters. \n\n\n- Something less effective:\nclass-aware sampling\ndata cleaning                                              \n\n\n\n",
      "votes": 5
    },
    {
      "id": 1815325,
      "postDate": "2022-06-08T23:19:20.103Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shuhaocui\" target=\"_blank\">@shuhaocui</a>, congratulations and thank you very much for the detailed explanation of your methods. Would you please fill out this form as well? <a href=\"https://forms.gle/n8Bsr5ZfkK4VLP1f7\" target=\"_blank\">https://forms.gle/n8Bsr5ZfkK4VLP1f7</a> Thank you.</p>",
      "rawMarkdown": "Hi @shuhaocui, congratulations and thank you very much for the detailed explanation of your methods. Would you please fill out this form as well? https://forms.gle/n8Bsr5ZfkK4VLP1f7 Thank you."
    }
  ],
  "comments": [
    {
      "id": 1815325,
      "author_name": "John Park",
      "author_url": "",
      "post_date": "2022-06-08T23:19:20.103000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shuhaocui\" target=\"_blank\">@shuhaocui</a>, congratulations and thank you very much for the detailed explanation of your methods. Would you please fill out this form as well? <a href=\"https://forms.gle/n8Bsr5ZfkK4VLP1f7\" target=\"_blank\">https://forms.gle/n8Bsr5ZfkK4VLP1f7</a> Thank you.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1812566": "- Optimization on single model:\n|change|public|private|\n|:--|:--|:--|\n|swin-B&nbsp;224&nbsp;baseline|0.77223|0.78442|\n|add&nbsp;cross-entropy&nbsp;losses&nbsp;on&nbsp;genera,species,&nbsp;family|0.7829|0.79544|\n|learning&nbsp;rate&nbsp;from&nbsp;2e-4&nbsp;to&nbsp;5e-4|0.79444|0.80501|\n|+5crop&nbsp;(resize&nbsp;256,&nbsp;288,&nbsp;320,&nbsp;384,&nbsp;448&nbsp;and&nbsp;center&nbsp;crop&nbsp;224)|0.80088|0.80981|\n|replacing&nbsp;cross-entropy&nbsp;by&nbsp;subcenter-arcface&nbsp;with&nbsp;[dynamic](https://www.kaggle.com/c/landmark-recognition-2020/discussion/187757)&nbsp;margins|0.81532|0.82267|\n|+&nbsp;extra&nbsp;fc&nbsp;by&nbsp;cross-entropy,&nbsp;add&nbsp;to&nbsp;subcenter-arcface&nbsp;output&nbsp;after&nbsp;softmax|0.82201|0.82929|\n|+&nbsp;data&nbsp;augmentation&nbsp;the&nbsp;same&nbsp;with&nbsp;Swin-Transformer|0.8315|0.83554|\n|from&nbsp;swin-B224&nbsp;to&nbsp;SwinB384|0.8468|0.85245|\n|+5crop&nbsp;(resize&nbsp;400,&nbsp;416,&nbsp;448,&nbsp;480,&nbsp;512&nbsp;and&nbsp;center&nbsp;crop&nbsp;384)|0.84895|0.85654|\n|+Freeze&nbsp;layers&nbsp;(not&nbsp;update&nbsp;parameters)&nbsp;from&nbsp;100&nbsp;to&nbsp;0|0.85499|0.86055|\n|+square&nbsp;resize&nbsp;and&nbsp;test&nbsp;crop&nbsp;from&nbsp;384|0.85563|0.86201|\n|+from&nbsp;swin&nbsp;base&nbsp;to&nbsp;swin&nbsp;V2|0.85851|0.86282|\n\n\n- Backbones:\n|id|backbone|public|private|\n|:--|:--|:--|:--|\n|0|swin-B|0.85499|0.86055|\n|1|convnext-B|0.85464|0.85956|\n|2|deit-iii&nbsp;B|0.84423|0.8495|\n|3|resnest-101|0.83208|0.83838|\n|4|efficient-B6|0.83841|0.84321|\n|5|cswin-L|0.84373|0.8501|\n|6|swin-L|0.85057|0.85607|\n|7|swinv2-B|0.85851|0.86282|\n\n\n- Fusion：\nwe utilize the scores on public dataset as the weight. \n|fusion&nbsp;approach|public|private|\n|:--|:--|:--|\n|fusion&nbsp;models&nbsp;(0,1,2)|0.86366|0.8679|\n|fusion&nbsp;models&nbsp;(0,1,2,3,4)|0.86619|0.87185|\n|fusion&nbsp;models&nbsp;(0,1,2,3,4,5,6,7)|0.87193|0.87662|\n\n\n- Other settings:\nBatch size is always 512 (256 for big models), trained with 100epochs.\ntorch.amp is adopted.\nFollowing [solutions](https://www.kaggle.com/competitions/herbarium-2021-fgvc8/discussion/242233)， we adopt similar Post-Process and obtain 0.001 improvement.\nWe use Swin-B pretrained on Imagenet22k for initialization, for other backbones, we try to use Imagenet22k pretrained weights if possible.\nWe freeze partial layers to enlarge the batchsize during the training process, We then reduce the batch size and train the full parameters. \n\n\n- Something less effective:\nclass-aware sampling\ndata cleaning                                              \n\n\n\n",
    "1815325": "Hi @shuhaocui, congratulations and thank you very much for the detailed explanation of your methods. Would you please fill out this form as well? https://forms.gle/n8Bsr5ZfkK4VLP1f7 Thank you."
  }
}