{
  "id": 220640,
  "title": "Private 606th / Public 34th Solution",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/kozistr-private-606th-public-34th-solution",
  "author_name": "",
  "post_date": "2021-02-19T03:49:52.577Z",
  "votes": 12,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n<p>First of all, congratulations to all the winners and did a great job on all participants' hard works!</p>\n<h1>Overview</h1>\n<p>I selected two versions based on best CV &amp; LB scores because I'm really worried about a huge shakeup. (However, I did not make it out :cool_crying:)</p>\n<ul>\n<li>Public  LB:   34th, 0.907</li>\n<li>Private LB: 606th, 0.898</li>\n</ul>\n<h2>Base Model</h2>\n<p>I only got Kaggle GPU, so I couldn't experiment with many models with various recipes.<br>\nI tried 6 models (ResNeSt50,  ResNeSt50 4s2x40d, EffNet-B3, 4, ViT-L/16, DeiT-B/16) and decided to use <strong>ResNeSt50 4s2x40d</strong>, <strong>EffNet-B4</strong> based on LB &amp; CV scores.</p>\n<table>\n<thead>\n<tr>\n<th>arch</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNeSt50-fast-4s2x40d (tta4)</td>\n<td>0.891</td>\n<td>0.898</td>\n<td>0.895</td>\n</tr>\n<tr>\n<td>EfficientNet-B3 NS (tta4)</td>\n<td>0.895</td>\n<td>0.896</td>\n<td>0.895</td>\n</tr>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>0.891</td>\n<td>0.900</td>\n<td>0.893</td>\n</tr>\n</tbody>\n</table>\n<p>To increase the robustness, I trained the model on different training recipes with the same architecture. I chose to use <code>focal cosine loss</code> &amp; <code>cross-entropy w/ label smoothing 0.2</code>.</p>\n<table>\n<thead>\n<tr>\n<th>arch</th>\n<th>loss</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B4 NS (tta3)</td>\n<td>focal cosine</td>\n<td>0.890</td>\n<td>0.900</td>\n<td>0.895</td>\n</tr>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>cross entropy w/ label smooth</td>\n<td>0.891</td>\n<td>0.900</td>\n<td>0.893</td>\n</tr>\n</tbody>\n</table>\n<p>Lastly, the pseudo label boosts the CV score +0.007. (I generated a pseudo label based on my best CV score, 0.9051)</p>\n<table>\n<thead>\n<tr>\n<th>arch</th>\n<th>pseudo</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>no</td>\n<td>0.891</td>\n<td>0.900</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>yes</td>\n<td>0.898</td>\n<td>0.898</td>\n<td>0.893</td>\n</tr>\n</tbody>\n</table>\n<h2>Ensemble</h2>\n<h3>Weights</h3>\n<p>Single model's CV &amp; LB scores seem not good, But, ensembling raises the CV &amp; LB scores to 0.005. Also, TTA increases LB scores +0.002.</p>\n<p>I tuned the ensemble weights with <code>Optuna</code> &amp; <a href=\"https://www.kaggle.com/daisukelab/optimizing-ensemble-weights-using-simple\" target=\"_blank\"><code>Simple</code></a> on CV.</p>\n<p>Here is the script : <a href=\"https://www.kaggle.com/kozistr/valid-optimize-cv\" target=\"_blank\">optimive cv</a></p>\n<h3>Combinations</h3>\n<ul>\n<li>best LB : 2 x ResNeSt50-fast-4s2x40d + 2 x EffNet-B4</li>\n<li>best CV : 3 x ResNeSt50-fast-4s2x40d + 3 x EffNet-B4 + 1 x EffNet-B3</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>baseline</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best LB</td>\n<td>0.9003</td>\n<td>0.907</td>\n<td>0.897</td>\n</tr>\n<tr>\n<td>best CV</td>\n<td>0.9051</td>\n<td>0.905</td>\n<td>0.896</td>\n</tr>\n</tbody>\n</table>\n<p>Here is the script : <a href=\"https://www.kaggle.com/kozistr/inference-rns50-effnet-b4\" target=\"_blank\">inference</a></p>\n<h2>TTA</h2>\n<p>Simply using <code>CenterCrop + horizontal &amp; vertical flop</code> achieves better CV &amp; LB scores.</p>\n<h2>Training Recipe</h2>\n<ul>\n<li>Data : 2019 + 2020 datasets, only 2020, w/ pseudo labels</li>\n<li>Validation : stratified kfold - 5 folds</li>\n<li>Augmentations : FMix, CutMix, SnapMix, MixUp, Cutout, etc…</li>\n<li>Optimizer : AdamW, AdamP, RAdam</li>\n<li>Loss : focal cosine, cross-entropy w/ label smoothing, bi-tempered</li>\n<li>Scheduler : Cosine Anealling w/o warm-up (1e-4 to 1e-6)</li>\n<li>Epochs : 10 ~ 20</li>\n<li>Resolution : 512x512</li>\n<li>Auxiliary : freeze BN layer</li>\n<li>TTA : n_iters 4 is best in my case.</li>\n</ul>\n<p>Experiment : <a href=\"https://www.kaggle.com/kozistr/cassava-leaf-disease-experiments\" target=\"_blank\">note</a></p>\n<h3>Works</h3>\n<ol>\n<li>Label Smoothing (alpha = 0.2)</li>\n<li>TTA (n_iters 4 is best)</li>\n<li>Adam, AdamW to RAdam (slight improvement)</li>\n<li>external dataset (2019 dataset)</li>\n</ol>\n<h3>Not Works</h3>\n<ol>\n<li>noise-tolerant(?) loss like bi-tempered loss</li>\n</ol>\n<h2>Reflections</h2>\n<ol>\n<li>ensemble weights</li>\n</ol>\n<p>I guess tuning ensemble weights might be one of the reasons, especially where there're noisy labels. Usually, the tuned versions tend to have a lower score than the average one. (-0.001 ~ -0.002 gap on Private LB score). But, Ironically best Private score is the tuned version (not the selected one).</p>\n<ol>\n<li>multiple scales (image resolution)</li>\n</ol>\n<p>I think training &amp; inferencing with multiple scales would be helpful to improve the score. On several experiments, a lower resolution (384x384) works better than a higher resolution (512x512).</p>\n<p>Thank you very much!</p>",
  "messages": [
    {
      "id": "1209803",
      "postDate": "02/19/2021 03:44:23",
      "content": "<p>Hi everyone!</p>\n<p>First of all, congratulations to all the winners and did a great job on all participants' hard works!</p>\n<h1>Overview</h1>\n<p>I selected two versions based on best CV &amp; LB scores because I'm really worried about a huge shakeup. (However, I did not make it out :cool_crying:)</p>\n<ul>\n<li>Public  LB:   34th, 0.907</li>\n<li>Private LB: 606th, 0.898</li>\n</ul>\n<h2>Base Model</h2>\n<p>I only got Kaggle GPU, so I couldn't experiment with many models with various recipes.<br>\nI tried 6 models (ResNeSt50,  ResNeSt50 4s2x40d, EffNet-B3, 4, ViT-L/16, DeiT-B/16) and decided to use <strong>ResNeSt50 4s2x40d</strong>, <strong>EffNet-B4</strong> based on LB &amp; CV scores.</p>\n<table>\n<thead>\n<tr>\n<th>arch</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNeSt50-fast-4s2x40d (tta4)</td>\n<td>0.891</td>\n<td>0.898</td>\n<td>0.895</td>\n</tr>\n<tr>\n<td>EfficientNet-B3 NS (tta4)</td>\n<td>0.895</td>\n<td>0.896</td>\n<td>0.895</td>\n</tr>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>0.891</td>\n<td>0.900</td>\n<td>0.893</td>\n</tr>\n</tbody>\n</table>\n<p>To increase the robustness, I trained the model on different training recipes with the same architecture. I chose to use <code>focal cosine loss</code> &amp; <code>cross-entropy w/ label smoothing 0.2</code>.</p>\n<table>\n<thead>\n<tr>\n<th>arch</th>\n<th>loss</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B4 NS (tta3)</td>\n<td>focal cosine</td>\n<td>0.890</td>\n<td>0.900</td>\n<td>0.895</td>\n</tr>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>cross entropy w/ label smooth</td>\n<td>0.891</td>\n<td>0.900</td>\n<td>0.893</td>\n</tr>\n</tbody>\n</table>\n<p>Lastly, the pseudo label boosts the CV score +0.007. (I generated a pseudo label based on my best CV score, 0.9051)</p>\n<table>\n<thead>\n<tr>\n<th>arch</th>\n<th>pseudo</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>no</td>\n<td>0.891</td>\n<td>0.900</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>EfficientNet-B4 NS (tta4)</td>\n<td>yes</td>\n<td>0.898</td>\n<td>0.898</td>\n<td>0.893</td>\n</tr>\n</tbody>\n</table>\n<h2>Ensemble</h2>\n<h3>Weights</h3>\n<p>Single model's CV &amp; LB scores seem not good, But, ensembling raises the CV &amp; LB scores to 0.005. Also, TTA increases LB scores +0.002.</p>\n<p>I tuned the ensemble weights with <code>Optuna</code> &amp; <a href=\"https://www.kaggle.com/daisukelab/optimizing-ensemble-weights-using-simple\" target=\"_blank\"><code>Simple</code></a> on CV.</p>\n<p>Here is the script : <a href=\"https://www.kaggle.com/kozistr/valid-optimize-cv\" target=\"_blank\">optimive cv</a></p>\n<h3>Combinations</h3>\n<ul>\n<li>best LB : 2 x ResNeSt50-fast-4s2x40d + 2 x EffNet-B4</li>\n<li>best CV : 3 x ResNeSt50-fast-4s2x40d + 3 x EffNet-B4 + 1 x EffNet-B3</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>baseline</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best LB</td>\n<td>0.9003</td>\n<td>0.907</td>\n<td>0.897</td>\n</tr>\n<tr>\n<td>best CV</td>\n<td>0.9051</td>\n<td>0.905</td>\n<td>0.896</td>\n</tr>\n</tbody>\n</table>\n<p>Here is the script : <a href=\"https://www.kaggle.com/kozistr/inference-rns50-effnet-b4\" target=\"_blank\">inference</a></p>\n<h2>TTA</h2>\n<p>Simply using <code>CenterCrop + horizontal &amp; vertical flop</code> achieves better CV &amp; LB scores.</p>\n<h2>Training Recipe</h2>\n<ul>\n<li>Data : 2019 + 2020 datasets, only 2020, w/ pseudo labels</li>\n<li>Validation : stratified kfold - 5 folds</li>\n<li>Augmentations : FMix, CutMix, SnapMix, MixUp, Cutout, etc…</li>\n<li>Optimizer : AdamW, AdamP, RAdam</li>\n<li>Loss : focal cosine, cross-entropy w/ label smoothing, bi-tempered</li>\n<li>Scheduler : Cosine Anealling w/o warm-up (1e-4 to 1e-6)</li>\n<li>Epochs : 10 ~ 20</li>\n<li>Resolution : 512x512</li>\n<li>Auxiliary : freeze BN layer</li>\n<li>TTA : n_iters 4 is best in my case.</li>\n</ul>\n<p>Experiment : <a href=\"https://www.kaggle.com/kozistr/cassava-leaf-disease-experiments\" target=\"_blank\">note</a></p>\n<h3>Works</h3>\n<ol>\n<li>Label Smoothing (alpha = 0.2)</li>\n<li>TTA (n_iters 4 is best)</li>\n<li>Adam, AdamW to RAdam (slight improvement)</li>\n<li>external dataset (2019 dataset)</li>\n</ol>\n<h3>Not Works</h3>\n<ol>\n<li>noise-tolerant(?) loss like bi-tempered loss</li>\n</ol>\n<h2>Reflections</h2>\n<ol>\n<li>ensemble weights</li>\n</ol>\n<p>I guess tuning ensemble weights might be one of the reasons, especially where there're noisy labels. Usually, the tuned versions tend to have a lower score than the average one. (-0.001 ~ -0.002 gap on Private LB score). But, Ironically best Private score is the tuned version (not the selected one).</p>\n<ol>\n<li>multiple scales (image resolution)</li>\n</ol>\n<p>I think training &amp; inferencing with multiple scales would be helpful to improve the score. On several experiments, a lower resolution (384x384) works better than a higher resolution (512x512).</p>\n<p>Thank you very much!</p>",
      "rawMarkdown": "Hi everyone!\n\nFirst of all, congratulations to all the winners and did a great job on all participants' hard works!\n\n# Overview\n\nI selected two versions based on best CV & LB scores because I'm really worried about a huge shakeup. (However, I did not make it out :cool_crying:)\n\n* Public  LB:   34th, 0.907\n* Private LB: 606th, 0.898\n\n## Base Model\n\nI only got Kaggle GPU, so I couldn't experiment with many models with various recipes.\nI tried 6 models (ResNeSt50,  ResNeSt50 4s2x40d, EffNet-B3, 4, ViT-L/16, DeiT-B/16) and decided to use **ResNeSt50 4s2x40d**, **EffNet-B4** based on LB & CV scores.\n\n| arch | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| ResNeSt50-fast-4s2x40d (tta4) | 0.891 | 0.898 | 0.895 |\n| EfficientNet-B3 NS (tta4) | 0.895 | 0.896  | 0.895 |\n| EfficientNet-B4 NS (tta4) | 0.891 | 0.900 | 0.893 |\n\nTo increase the robustness, I trained the model on different training recipes with the same architecture. I chose to use `focal cosine loss` & `cross-entropy w/ label smoothing 0.2`.\n\n| arch | loss | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| EfficientNet-B4 NS (tta3) | focal cosine | 0.890 | 0.900  | 0.895 |\n| EfficientNet-B4 NS (tta4) | cross entropy w/ label smooth | 0.891 | 0.900 | 0.893 |\n\nLastly, the pseudo label boosts the CV score +0.007. (I generated a pseudo label based on my best CV score, 0.9051)\n\n| arch | pseudo | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: | :---: |\n| EfficientNet-B4 NS (tta4) | no | 0.891 | 0.900 | 0.893 |\n| EfficientNet-B4 NS (tta4) | yes | 0.898 | 0.898 | 0.893 |\n\n## Ensemble\n\n### Weights\n\nSingle model's CV & LB scores seem not good, But, ensembling raises the CV & LB scores to 0.005. Also, TTA increases LB scores +0.002.\n\nI tuned the ensemble weights with `Optuna` & [`Simple`] (https://www.kaggle.com/daisukelab/optimizing-ensemble-weights-using-simple) on CV.\n\nHere is the script : [optimive cv](https://www.kaggle.com/kozistr/valid-optimize-cv)\n\n### Combinations\n\n* best LB : 2 x ResNeSt50-fast-4s2x40d + 2 x EffNet-B4\n* best CV : 3 x ResNeSt50-fast-4s2x40d + 3 x EffNet-B4 + 1 x EffNet-B3\n\n| baseline | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| best LB | 0.9003 | 0.907 | 0.897 |\n| best CV | 0.9051 | 0.905 | 0.896 |\n\nHere is the script : [inference](https://www.kaggle.com/kozistr/inference-rns50-effnet-b4)\n\n## TTA\n\nSimply using `CenterCrop + horizontal & vertical flop` achieves better CV & LB scores.\n\n## Training Recipe\n\n* Data : 2019 + 2020 datasets, only 2020, w/ pseudo labels\n* Validation : stratified kfold - 5 folds\n* Augmentations : FMix, CutMix, SnapMix, MixUp, Cutout, etc...\n* Optimizer : AdamW, AdamP, RAdam\n* Loss : focal cosine, cross-entropy w/ label smoothing, bi-tempered\n* Scheduler : Cosine Anealling w/o warm-up (1e-4 to 1e-6)\n* Epochs : 10 ~ 20\n* Resolution : 512x512\n* Auxiliary : freeze BN layer\n* TTA : n_iters 4 is best in my case.\n\nExperiment : [note](https://www.kaggle.com/kozistr/cassava-leaf-disease-experiments)\n\n### Works\n\n1. Label Smoothing (alpha = 0.2)\n2. TTA (n_iters 4 is best)\n3. Adam, AdamW to RAdam (slight improvement)\n4. external dataset (2019 dataset)\n\n### Not Works\n\n1. noise-tolerant(?) loss like bi-tempered loss\n\n## Reflections\n\n1. ensemble weights\n\nI guess tuning ensemble weights might be one of the reasons, especially where there're noisy labels. Usually, the tuned versions tend to have a lower score than the average one. (-0.001 ~ -0.002 gap on Private LB score). But, Ironically best Private score is the tuned version (not the selected one).\n\n2. multiple scales (image resolution)\n\nI think training & inferencing with multiple scales would be helpful to improve the score. On several experiments, a lower resolution (384x384) works better than a higher resolution (512x512).\n\nThank you very much!",
      "votes": null
    },
    {
      "id": "1209811",
      "postDate": "02/19/2021 03:50:05",
      "content": "<p>Thanks for sharing. <br>\n고생많으셨습니다!!! 솔루션 너무 좋은데 운이 잘 안따라주신 것 같아서 아쉽네요 ㅠㅠ😪😪</p>",
      "rawMarkdown": "Thanks for sharing. \n고생많으셨습니다!!! 솔루션 너무 좋은데 운이 잘 안따라주신 것 같아서 아쉽네요 ㅠㅠ😪😪",
      "votes": null
    },
    {
      "id": "1209819",
      "postDate": "02/19/2021 03:56:29",
      "content": "<p>메달 정말 축하드리고 고생 많으셨어요!!<br>\n이번 대회는 아쉽지만, 더 열심히 공부해서 다음 대회를 노려봐야겠어요 ㅎㅎ</p>",
      "rawMarkdown": "메달 정말 축하드리고 고생 많으셨어요!!\n이번 대회는 아쉽지만, 더 열심히 공부해서 다음 대회를 노려봐야겠어요 ㅎㅎ",
      "votes": null
    },
    {
      "id": "1209847",
      "postDate": "02/19/2021 04:24:46",
      "content": "<p>흐으~ 다음번엔 금메달받아서 마스터가시져~!! 은메달도 이제 지겹네요 ㅋㅋㅋ</p>",
      "rawMarkdown": "흐으~ 다음번엔 금메달받아서 마스터가시져~!! 은메달도 이제 지겹네요 ㅋㅋㅋ",
      "votes": null
    },
    {
      "id": "1209893",
      "postDate": "02/19/2021 04:57:28",
      "content": "<p>ㅋㅋㅋㅋㅋ 금메달 가즈아ㅏㅏ 기회되면 다음에 같이 대회해요!</p>",
      "rawMarkdown": "ㅋㅋㅋㅋㅋ 금메달 가즈아ㅏㅏ 기회되면 다음에 같이 대회해요!",
      "votes": null
    },
    {
      "id": "1209902",
      "postDate": "02/19/2021 05:15:27",
      "content": "<p>Thanks for sharing and great effort for this competition.<br>\nI look forward to receiving your gold medal in the next competition. :)<br>\n고생하셨습니다.^^</p>\n<p><a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> </p>",
      "rawMarkdown": "Thanks for sharing and great effort for this competition.\nI look forward to receiving your gold medal in the next competition. :)\n고생하셨습니다.^^\n\n@kozistr",
      "votes": null
    },
    {
      "id": "1209952",
      "postDate": "02/19/2021 05:57:12",
      "content": "<p>감사합니다! Heroseo 님도 정말 수고하셨습니다!!</p>",
      "rawMarkdown": "감사합니다! Heroseo 님도 정말 수고하셨습니다!!",
      "votes": null
    },
    {
      "id": "1210292",
      "postDate": "02/19/2021 10:18:44",
      "content": "<p>Thanks for sharing great solution 👍<br>\n깔끔한 정리 보기 좋네요ㅎㅎ 수고하셨습니다~!</p>",
      "rawMarkdown": "Thanks for sharing great solution 👍\n깔끔한 정리 보기 좋네요ㅎㅎ 수고하셨습니다~!",
      "votes": null
    },
    {
      "id": "1211190",
      "postDate": "02/20/2021 03:40:29",
      "content": "<p>감사합니다 ㅎㅎ nevret 님도 수고하셨습니다!!</p>",
      "rawMarkdown": "감사합니다 ㅎㅎ nevret 님도 수고하셨습니다!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1209811,
      "author_name": "chocozzz",
      "author_url": "",
      "post_date": "02/19/2021 03:50:05",
      "content": "<p>Thanks for sharing. <br>\n고생많으셨습니다!!! 솔루션 너무 좋은데 운이 잘 안따라주신 것 같아서 아쉽네요 ㅠㅠ😪😪</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209819,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "02/19/2021 03:56:29",
          "content": "<p>메달 정말 축하드리고 고생 많으셨어요!!<br>\n이번 대회는 아쉽지만, 더 열심히 공부해서 다음 대회를 노려봐야겠어요 ㅎㅎ</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209847,
          "author_name": "chocozzz",
          "author_url": "",
          "post_date": "02/19/2021 04:24:46",
          "content": "<p>흐으~ 다음번엔 금메달받아서 마스터가시져~!! 은메달도 이제 지겹네요 ㅋㅋㅋ</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209893,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "02/19/2021 04:57:28",
          "content": "<p>ㅋㅋㅋㅋㅋ 금메달 가즈아ㅏㅏ 기회되면 다음에 같이 대회해요!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209902,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/19/2021 05:15:27",
      "content": "<p>Thanks for sharing and great effort for this competition.<br>\nI look forward to receiving your gold medal in the next competition. :)<br>\n고생하셨습니다.^^</p>\n<p><a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1209952,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "02/19/2021 05:57:12",
          "content": "<p>감사합니다! Heroseo 님도 정말 수고하셨습니다!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210292,
      "author_name": "nevret93",
      "author_url": "",
      "post_date": "02/19/2021 10:18:44",
      "content": "<p>Thanks for sharing great solution 👍<br>\n깔끔한 정리 보기 좋네요ㅎㅎ 수고하셨습니다~!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211190,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "02/20/2021 03:40:29",
          "content": "<p>감사합니다 ㅎㅎ nevret 님도 수고하셨습니다!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1209803": "Hi everyone!\n\nFirst of all, congratulations to all the winners and did a great job on all participants' hard works!\n\n# Overview\n\nI selected two versions based on best CV & LB scores because I'm really worried about a huge shakeup. (However, I did not make it out :cool_crying:)\n\n* Public  LB:   34th, 0.907\n* Private LB: 606th, 0.898\n\n## Base Model\n\nI only got Kaggle GPU, so I couldn't experiment with many models with various recipes.\nI tried 6 models (ResNeSt50,  ResNeSt50 4s2x40d, EffNet-B3, 4, ViT-L/16, DeiT-B/16) and decided to use **ResNeSt50 4s2x40d**, **EffNet-B4** based on LB & CV scores.\n\n| arch | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| ResNeSt50-fast-4s2x40d (tta4) | 0.891 | 0.898 | 0.895 |\n| EfficientNet-B3 NS (tta4) | 0.895 | 0.896  | 0.895 |\n| EfficientNet-B4 NS (tta4) | 0.891 | 0.900 | 0.893 |\n\nTo increase the robustness, I trained the model on different training recipes with the same architecture. I chose to use `focal cosine loss` & `cross-entropy w/ label smoothing 0.2`.\n\n| arch | loss | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| EfficientNet-B4 NS (tta3) | focal cosine | 0.890 | 0.900  | 0.895 |\n| EfficientNet-B4 NS (tta4) | cross entropy w/ label smooth | 0.891 | 0.900 | 0.893 |\n\nLastly, the pseudo label boosts the CV score +0.007. (I generated a pseudo label based on my best CV score, 0.9051)\n\n| arch | pseudo | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: | :---: |\n| EfficientNet-B4 NS (tta4) | no | 0.891 | 0.900 | 0.893 |\n| EfficientNet-B4 NS (tta4) | yes | 0.898 | 0.898 | 0.893 |\n\n## Ensemble\n\n### Weights\n\nSingle model's CV & LB scores seem not good, But, ensembling raises the CV & LB scores to 0.005. Also, TTA increases LB scores +0.002.\n\nI tuned the ensemble weights with `Optuna` & [`Simple`] (https://www.kaggle.com/daisukelab/optimizing-ensemble-weights-using-simple) on CV.\n\nHere is the script : [optimive cv](https://www.kaggle.com/kozistr/valid-optimize-cv)\n\n### Combinations\n\n* best LB : 2 x ResNeSt50-fast-4s2x40d + 2 x EffNet-B4\n* best CV : 3 x ResNeSt50-fast-4s2x40d + 3 x EffNet-B4 + 1 x EffNet-B3\n\n| baseline | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| best LB | 0.9003 | 0.907 | 0.897 |\n| best CV | 0.9051 | 0.905 | 0.896 |\n\nHere is the script : [inference](https://www.kaggle.com/kozistr/inference-rns50-effnet-b4)\n\n## TTA\n\nSimply using `CenterCrop + horizontal & vertical flop` achieves better CV & LB scores.\n\n## Training Recipe\n\n* Data : 2019 + 2020 datasets, only 2020, w/ pseudo labels\n* Validation : stratified kfold - 5 folds\n* Augmentations : FMix, CutMix, SnapMix, MixUp, Cutout, etc...\n* Optimizer : AdamW, AdamP, RAdam\n* Loss : focal cosine, cross-entropy w/ label smoothing, bi-tempered\n* Scheduler : Cosine Anealling w/o warm-up (1e-4 to 1e-6)\n* Epochs : 10 ~ 20\n* Resolution : 512x512\n* Auxiliary : freeze BN layer\n* TTA : n_iters 4 is best in my case.\n\nExperiment : [note](https://www.kaggle.com/kozistr/cassava-leaf-disease-experiments)\n\n### Works\n\n1. Label Smoothing (alpha = 0.2)\n2. TTA (n_iters 4 is best)\n3. Adam, AdamW to RAdam (slight improvement)\n4. external dataset (2019 dataset)\n\n### Not Works\n\n1. noise-tolerant(?) loss like bi-tempered loss\n\n## Reflections\n\n1. ensemble weights\n\nI guess tuning ensemble weights might be one of the reasons, especially where there're noisy labels. Usually, the tuned versions tend to have a lower score than the average one. (-0.001 ~ -0.002 gap on Private LB score). But, Ironically best Private score is the tuned version (not the selected one).\n\n2. multiple scales (image resolution)\n\nI think training & inferencing with multiple scales would be helpful to improve the score. On several experiments, a lower resolution (384x384) works better than a higher resolution (512x512).\n\nThank you very much!",
    "1209811": "Thanks for sharing. \n고생많으셨습니다!!! 솔루션 너무 좋은데 운이 잘 안따라주신 것 같아서 아쉽네요 ㅠㅠ😪😪",
    "1209819": "메달 정말 축하드리고 고생 많으셨어요!!\n이번 대회는 아쉽지만, 더 열심히 공부해서 다음 대회를 노려봐야겠어요 ㅎㅎ",
    "1209847": "흐으~ 다음번엔 금메달받아서 마스터가시져~!! 은메달도 이제 지겹네요 ㅋㅋㅋ",
    "1209893": "ㅋㅋㅋㅋㅋ 금메달 가즈아ㅏㅏ 기회되면 다음에 같이 대회해요!",
    "1209902": "Thanks for sharing and great effort for this competition.\nI look forward to receiving your gold medal in the next competition. :)\n고생하셨습니다.^^\n\n@kozistr",
    "1209952": "감사합니다! Heroseo 님도 정말 수고하셨습니다!!",
    "1210292": "Thanks for sharing great solution 👍\n깔끔한 정리 보기 좋네요ㅎㅎ 수고하셨습니다~!",
    "1211190": "감사합니다 ㅎㅎ nevret 님도 수고하셨습니다!!"
  },
  "source": "meta"
}