{
  "id": 242233,
  "title": "2nd place solution ",
  "url": "/competitions/herbarium-2021-fgvc8/discussion/242233",
  "author_name": "HaeC",
  "post_date": "2021-05-28T04:32:42.679000",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, Thanks to the organizers for preparing such a great competition </p>\n<p><strong>Difficult point of this competition:</strong></p>\n<ol>\n<li>The dataset for this competition was very large, a lot of GPU computation was required. So, I couldn't do much experimentation.</li>\n<li>So, I couldn't train the model for many epochs. This is not able to say that the model is completely converged. </li>\n</ol>\n<p><strong>Network</strong><br>\nThe network architecture was based on a configuration that had a good score in the previous year (2nd place). I added a few network layers to it.</p>\n<p>input -&gt; Backbone(Feature extractor) -&gt; GAP -&gt; FC (2048x2048) -&gt; BN -&gt; Leaky ReLU -&gt; FC2  (2048x512)-&gt; BN -&gt; Leaky ReLU -&gt;FC3(512x64500)</p>\n<p><strong>Loss</strong><br>\n+Metric Loss</p>\n<ul>\n<li>SoftTriple Loss was applied to the oupute of FC2. ( in my experiment, SoftTriple is better than ArcFace)</li>\n</ul>\n<p>+Classification Loss</p>\n<ul>\n<li>CrossEntropy was applied to the oupute of FC3 in 1st stage(30epoch).</li>\n<li><a href=\"https://papers.nips.cc/paper/2020/file/2ba61cc3a8f44143e1f2f13b2b729ab3-Paper.pdf\" target=\"_blank\">CrossEntropy with Balanced SoftMax </a>was applied to the oupute of FC3 in 2nd stage(5epoch).</li>\n</ul>\n<p><strong>Augmentation</strong><br>\n+Train Phase</p>\n<ul>\n<li>img size : 448x448</li>\n<li>AugMix without JSD Loss (custom implementation with Albu)</li>\n<li>pipeline is below</li>\n</ul>\n<pre><code>A.Resize(random.randint(CFG.img_size[0] + 32 , \n                   CFG.img_size[0] + 128), \n                   CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nRandomAugMix(severity=3, width=3, alpha=1.0, p=1.0),\nA.Cutout(p=0.5),\nToTensorV2()\n</code></pre>\n<p>+Test Phase </p>\n<ul>\n<li>TTA 5 fold</li>\n</ul>\n<pre><code>A.Resize(random.randint(CFG.img_size[0] + 16, \n                                CFG.img_size[0] + 64), \n                                CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nA.Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, \n                p=1.0)\n</code></pre>\n<p><strong>Optimizer and Other tricks</strong><br>\n-AdamP : lr=0.001, no weight decay<br>\n-CosineAnnealingLR<br>\n-Exponential Moving Average<br>\n-AMP (fp16)</p>\n<p><strong>Training and Validation Strategy</strong></p>\n<ul>\n<li>I trained the model for a total of 35 epochs on the 1st stage(30) and the 2nd stage(5).</li>\n<li>If there was enough time and models were trained with more epochs, it would have performed better.</li>\n<li>When I tried the 80:20 train-valid split, I saw the result of not overfitting during training (During 40 epoch). Therefore, model was trained without validation</li>\n</ul>\n<p><strong>Post-Process (switching)</strong><br>\nI switched the 1st and 2nd confidence class when satisfying each the 2 cases.<br>\ncase 1:</p>\n<ul>\n<li>top_2_conf &gt;= (top_1_conf * 0.7)</li>\n<li>top_2_class not in top_1_conf list</li>\n<li>top_1_class appears more than 2 in top_1_conf</li>\n</ul>\n<p>case 2:</p>\n<ul>\n<li>top_2_conf &gt;= (top_1_conf * 0.6)</li>\n<li>top_2_class not in top_1_conf list</li>\n<li>top_1_class appears more than 15 in top_1_conf</li>\n</ul>\n<p><strong>Model Score</strong></p>\n<ul>\n<li><p>TResNet-M-448 (2 stage and TTA 5 fold) [ Public : 0.72937, Private : 0.68392 ] </p></li>\n<li><p>TResNet-M-21k (2 stage and TTA 5 fold) [ sorry, no submission :( ] </p></li>\n<li><p>TResNet-L-448 (2 stage and TTA 5 fold) [ Public : 0.75006, Private : 0.70231 ] </p></li>\n<li><p>GENet-L (2 stage and TTA 5 fold) [ Public : 0.71862, Private : 0.67339 ] </p></li>\n<li><p>ECA-NFNet-L0 (2 stage and TTA 5 fold) [ Public : 0.73398, Private : 0.68774 ] </p></li>\n<li><p>Soft-Voting 5 model and Post-Process [ Public : 0.77625, Private : 0.73534 ] (Final Result)</p></li>\n</ul>\n<p>If you have any additional questions, please leave a comment. I hope my solution is helpful. :)</p>",
  "messages": [
    {
      "id": 1325859,
      "postDate": "2021-05-28T04:32:42.680Z",
      "content": "<p>First of all, Thanks to the organizers for preparing such a great competition </p>\n<p><strong>Difficult point of this competition:</strong></p>\n<ol>\n<li>The dataset for this competition was very large, a lot of GPU computation was required. So, I couldn't do much experimentation.</li>\n<li>So, I couldn't train the model for many epochs. This is not able to say that the model is completely converged. </li>\n</ol>\n<p><strong>Network</strong><br>\nThe network architecture was based on a configuration that had a good score in the previous year (2nd place). I added a few network layers to it.</p>\n<p>input -&gt; Backbone(Feature extractor) -&gt; GAP -&gt; FC (2048x2048) -&gt; BN -&gt; Leaky ReLU -&gt; FC2  (2048x512)-&gt; BN -&gt; Leaky ReLU -&gt;FC3(512x64500)</p>\n<p><strong>Loss</strong><br>\n+Metric Loss</p>\n<ul>\n<li>SoftTriple Loss was applied to the oupute of FC2. ( in my experiment, SoftTriple is better than ArcFace)</li>\n</ul>\n<p>+Classification Loss</p>\n<ul>\n<li>CrossEntropy was applied to the oupute of FC3 in 1st stage(30epoch).</li>\n<li><a href=\"https://papers.nips.cc/paper/2020/file/2ba61cc3a8f44143e1f2f13b2b729ab3-Paper.pdf\" target=\"_blank\">CrossEntropy with Balanced SoftMax </a>was applied to the oupute of FC3 in 2nd stage(5epoch).</li>\n</ul>\n<p><strong>Augmentation</strong><br>\n+Train Phase</p>\n<ul>\n<li>img size : 448x448</li>\n<li>AugMix without JSD Loss (custom implementation with Albu)</li>\n<li>pipeline is below</li>\n</ul>\n<pre><code>A.Resize(random.randint(CFG.img_size[0] + 32 , \n                   CFG.img_size[0] + 128), \n                   CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nRandomAugMix(severity=3, width=3, alpha=1.0, p=1.0),\nA.Cutout(p=0.5),\nToTensorV2()\n</code></pre>\n<p>+Test Phase </p>\n<ul>\n<li>TTA 5 fold</li>\n</ul>\n<pre><code>A.Resize(random.randint(CFG.img_size[0] + 16, \n                                CFG.img_size[0] + 64), \n                                CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nA.Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, \n                p=1.0)\n</code></pre>\n<p><strong>Optimizer and Other tricks</strong><br>\n-AdamP : lr=0.001, no weight decay<br>\n-CosineAnnealingLR<br>\n-Exponential Moving Average<br>\n-AMP (fp16)</p>\n<p><strong>Training and Validation Strategy</strong></p>\n<ul>\n<li>I trained the model for a total of 35 epochs on the 1st stage(30) and the 2nd stage(5).</li>\n<li>If there was enough time and models were trained with more epochs, it would have performed better.</li>\n<li>When I tried the 80:20 train-valid split, I saw the result of not overfitting during training (During 40 epoch). Therefore, model was trained without validation</li>\n</ul>\n<p><strong>Post-Process (switching)</strong><br>\nI switched the 1st and 2nd confidence class when satisfying each the 2 cases.<br>\ncase 1:</p>\n<ul>\n<li>top_2_conf &gt;= (top_1_conf * 0.7)</li>\n<li>top_2_class not in top_1_conf list</li>\n<li>top_1_class appears more than 2 in top_1_conf</li>\n</ul>\n<p>case 2:</p>\n<ul>\n<li>top_2_conf &gt;= (top_1_conf * 0.6)</li>\n<li>top_2_class not in top_1_conf list</li>\n<li>top_1_class appears more than 15 in top_1_conf</li>\n</ul>\n<p><strong>Model Score</strong></p>\n<ul>\n<li><p>TResNet-M-448 (2 stage and TTA 5 fold) [ Public : 0.72937, Private : 0.68392 ] </p></li>\n<li><p>TResNet-M-21k (2 stage and TTA 5 fold) [ sorry, no submission :( ] </p></li>\n<li><p>TResNet-L-448 (2 stage and TTA 5 fold) [ Public : 0.75006, Private : 0.70231 ] </p></li>\n<li><p>GENet-L (2 stage and TTA 5 fold) [ Public : 0.71862, Private : 0.67339 ] </p></li>\n<li><p>ECA-NFNet-L0 (2 stage and TTA 5 fold) [ Public : 0.73398, Private : 0.68774 ] </p></li>\n<li><p>Soft-Voting 5 model and Post-Process [ Public : 0.77625, Private : 0.73534 ] (Final Result)</p></li>\n</ul>\n<p>If you have any additional questions, please leave a comment. I hope my solution is helpful. :)</p>",
      "rawMarkdown": "First of all, Thanks to the organizers for preparing such a great competition \n\n\n**Difficult point of this competition:**\n1. The dataset for this competition was very large, a lot of GPU computation was required. So, I couldn't do much experimentation.\n2. So, I couldn't train the model for many epochs. This is not able to say that the model is completely converged. \n\n**Network**\nThe network architecture was based on a configuration that had a good score in the previous year (2nd place). I added a few network layers to it.\n\ninput -> Backbone(Feature extractor) -> GAP -> FC (2048x2048) -> BN -> Leaky ReLU -> FC2  (2048x512)-> BN -> Leaky ReLU ->FC3(512x64500)\n\n**Loss**\n+Metric Loss\n- SoftTriple Loss was applied to the oupute of FC2. ( in my experiment, SoftTriple is better than ArcFace)\n\n+Classification Loss\n- CrossEntropy was applied to the oupute of FC3 in 1st stage(30epoch).\n- [CrossEntropy with Balanced SoftMax ](https://papers.nips.cc/paper/2020/file/2ba61cc3a8f44143e1f2f13b2b729ab3-Paper.pdf)was applied to the oupute of FC3 in 2nd stage(5epoch).\n\n**Augmentation**\n+Train Phase\n- img size : 448x448\n- AugMix without JSD Loss (custom implementation with Albu)\n- pipeline is below\n```\nA.Resize(random.randint(CFG.img_size[0] + 32 , \n                   CFG.img_size[0] + 128), \n                   CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nRandomAugMix(severity=3, width=3, alpha=1.0, p=1.0),\nA.Cutout(p=0.5),\nToTensorV2()\n```\n+Test Phase \n- TTA 5 fold\n```\nA.Resize(random.randint(CFG.img_size[0] + 16, \n                                CFG.img_size[0] + 64), \n                                CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nA.Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, \n                p=1.0)\n```\n\n\n**Optimizer and Other tricks**\n-AdamP : lr=0.001, no weight decay\n-CosineAnnealingLR\n-Exponential Moving Average\n-AMP (fp16)\n\n**Training and Validation Strategy**\n- I trained the model for a total of 35 epochs on the 1st stage(30) and the 2nd stage(5).\n- If there was enough time and models were trained with more epochs, it would have performed better.\n- When I tried the 80:20 train-valid split, I saw the result of not overfitting during training (During 40 epoch). Therefore, model was trained without validation\n\n**Post-Process (switching)**\nI switched the 1st and 2nd confidence class when satisfying each the 2 cases.\ncase 1:\n- top_2_conf >= (top_1_conf * 0.7)\n- top_2_class not in top_1_conf list\n- top_1_class appears more than 2 in top_1_conf\n\ncase 2:\n- top_2_conf >= (top_1_conf * 0.6)\n- top_2_class not in top_1_conf list\n- top_1_class appears more than 15 in top_1_conf\n\n\n\n**Model Score**\n- TResNet-M-448 (2 stage and TTA 5 fold) [ Public : 0.72937, Private : 0.68392 ] \n- TResNet-M-21k (2 stage and TTA 5 fold) [ sorry, no submission :( ] \n- TResNet-L-448 (2 stage and TTA 5 fold) [ Public : 0.75006, Private : 0.70231 ] \n- GENet-L (2 stage and TTA 5 fold) [ Public : 0.71862, Private : 0.67339 ] \n- ECA-NFNet-L0 (2 stage and TTA 5 fold) [ Public : 0.73398, Private : 0.68774 ] \n\n- Soft-Voting 5 model and Post-Process [ Public : 0.77625, Private : 0.73534 ] (Final Result)\n\nIf you have any additional questions, please leave a comment. I hope my solution is helpful. :)\n",
      "votes": 13
    },
    {
      "id": 2413276,
      "postDate": "2023-08-28T18:57:00.007Z",
      "content": "<p>Hi HAEC,<br>\nWhat was the validation acc?</p>",
      "rawMarkdown": "Hi HAEC,\nWhat was the validation acc?"
    },
    {
      "id": 1329904,
      "postDate": "2021-05-31T13:06:17.757Z",
      "content": "<p>Hi HaeC,</p>\n<p>Congratulations on your second-place finish! And thank you for giving details to your solution that is very valuable!</p>\n<p>We will be collecting all of the findings from the top 10 teams for the FGVC workshop, it would be great if you could fill in this form: <br>\n<a href=\"https://forms.gle/qtyAc2QH8yYYjhGb8\" target=\"_blank\">https://forms.gle/qtyAc2QH8yYYjhGb8</a></p>\n<p>And if you have any general comments for the competition wrap-up you can do so on this thread: <br>\n<a href=\"https://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242914\" target=\"_blank\">https://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242914</a></p>\n<p>Best,<br>\nRiccardo</p>",
      "rawMarkdown": "Hi HaeC,\n\nCongratulations on your second-place finish! And thank you for giving details to your solution that is very valuable!\n\nWe will be collecting all of the findings from the top 10 teams for the FGVC workshop, it would be great if you could fill in this form: \nhttps://forms.gle/qtyAc2QH8yYYjhGb8\n\nAnd if you have any general comments for the competition wrap-up you can do so on this thread: \nhttps://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242914\n\nBest,\nRiccardo\n",
      "replies": [
        {
          "id": 1330634,
          "postDate": "2021-06-01T02:59:50.400Z",
          "content": "<p>Hi Riccardo,</p>\n<p>Thank you for your reply. Within today, I will fill out the Google Form!</p>\n<p>Best,<br>\nHaechan</p>",
          "rawMarkdown": "Hi Riccardo,\n\nThank you for your reply. Within today, I will fill out the Google Form!\n\nBest,\nHaechan"
        }
      ]
    },
    {
      "id": 1388515,
      "postDate": "2021-07-15T03:12:40.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2413276,
      "author_name": "Maryam Darei",
      "author_url": "",
      "post_date": "2023-08-28T18:57:00.007000",
      "content": "<p>Hi HAEC,<br>\nWhat was the validation acc?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1329904,
      "author_name": "Riccardo de Lutio",
      "author_url": "",
      "post_date": "2021-05-31T13:06:17.757000",
      "content": "<p>Hi HaeC,</p>\n<p>Congratulations on your second-place finish! And thank you for giving details to your solution that is very valuable!</p>\n<p>We will be collecting all of the findings from the top 10 teams for the FGVC workshop, it would be great if you could fill in this form: <br>\n<a href=\"https://forms.gle/qtyAc2QH8yYYjhGb8\" target=\"_blank\">https://forms.gle/qtyAc2QH8yYYjhGb8</a></p>\n<p>And if you have any general comments for the competition wrap-up you can do so on this thread: <br>\n<a href=\"https://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242914\" target=\"_blank\">https://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242914</a></p>\n<p>Best,<br>\nRiccardo</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1330634,
          "author_name": "HaeC",
          "author_url": "",
          "post_date": "2021-06-01T02:59:50.400000",
          "content": "<p>Hi Riccardo,</p>\n<p>Thank you for your reply. Within today, I will fill out the Google Form!</p>\n<p>Best,<br>\nHaechan</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1388515,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-15T03:12:40.367000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1325859": "First of all, Thanks to the organizers for preparing such a great competition \n\n\n**Difficult point of this competition:**\n1. The dataset for this competition was very large, a lot of GPU computation was required. So, I couldn't do much experimentation.\n2. So, I couldn't train the model for many epochs. This is not able to say that the model is completely converged. \n\n**Network**\nThe network architecture was based on a configuration that had a good score in the previous year (2nd place). I added a few network layers to it.\n\ninput -> Backbone(Feature extractor) -> GAP -> FC (2048x2048) -> BN -> Leaky ReLU -> FC2  (2048x512)-> BN -> Leaky ReLU ->FC3(512x64500)\n\n**Loss**\n+Metric Loss\n- SoftTriple Loss was applied to the oupute of FC2. ( in my experiment, SoftTriple is better than ArcFace)\n\n+Classification Loss\n- CrossEntropy was applied to the oupute of FC3 in 1st stage(30epoch).\n- [CrossEntropy with Balanced SoftMax ](https://papers.nips.cc/paper/2020/file/2ba61cc3a8f44143e1f2f13b2b729ab3-Paper.pdf)was applied to the oupute of FC3 in 2nd stage(5epoch).\n\n**Augmentation**\n+Train Phase\n- img size : 448x448\n- AugMix without JSD Loss (custom implementation with Albu)\n- pipeline is below\n```\nA.Resize(random.randint(CFG.img_size[0] + 32 , \n                   CFG.img_size[0] + 128), \n                   CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nRandomAugMix(severity=3, width=3, alpha=1.0, p=1.0),\nA.Cutout(p=0.5),\nToTensorV2()\n```\n+Test Phase \n- TTA 5 fold\n```\nA.Resize(random.randint(CFG.img_size[0] + 16, \n                                CFG.img_size[0] + 64), \n                                CFG.img_size[1]),\nA.RandomCrop(CFG.img_size[0], CFG.img_size[0]),\nA.HorizontalFlip(p=0.5),\nA.Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, \n                p=1.0)\n```\n\n\n**Optimizer and Other tricks**\n-AdamP : lr=0.001, no weight decay\n-CosineAnnealingLR\n-Exponential Moving Average\n-AMP (fp16)\n\n**Training and Validation Strategy**\n- I trained the model for a total of 35 epochs on the 1st stage(30) and the 2nd stage(5).\n- If there was enough time and models were trained with more epochs, it would have performed better.\n- When I tried the 80:20 train-valid split, I saw the result of not overfitting during training (During 40 epoch). Therefore, model was trained without validation\n\n**Post-Process (switching)**\nI switched the 1st and 2nd confidence class when satisfying each the 2 cases.\ncase 1:\n- top_2_conf >= (top_1_conf * 0.7)\n- top_2_class not in top_1_conf list\n- top_1_class appears more than 2 in top_1_conf\n\ncase 2:\n- top_2_conf >= (top_1_conf * 0.6)\n- top_2_class not in top_1_conf list\n- top_1_class appears more than 15 in top_1_conf\n\n\n\n**Model Score**\n- TResNet-M-448 (2 stage and TTA 5 fold) [ Public : 0.72937, Private : 0.68392 ] \n- TResNet-M-21k (2 stage and TTA 5 fold) [ sorry, no submission :( ] \n- TResNet-L-448 (2 stage and TTA 5 fold) [ Public : 0.75006, Private : 0.70231 ] \n- GENet-L (2 stage and TTA 5 fold) [ Public : 0.71862, Private : 0.67339 ] \n- ECA-NFNet-L0 (2 stage and TTA 5 fold) [ Public : 0.73398, Private : 0.68774 ] \n\n- Soft-Voting 5 model and Post-Process [ Public : 0.77625, Private : 0.73534 ] (Final Result)\n\nIf you have any additional questions, please leave a comment. I hope my solution is helpful. :)\n",
    "2413276": "Hi HAEC,\nWhat was the validation acc?",
    "1329904": "Hi HaeC,\n\nCongratulations on your second-place finish! And thank you for giving details to your solution that is very valuable!\n\nWe will be collecting all of the findings from the top 10 teams for the FGVC workshop, it would be great if you could fill in this form: \nhttps://forms.gle/qtyAc2QH8yYYjhGb8\n\nAnd if you have any general comments for the competition wrap-up you can do so on this thread: \nhttps://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242914\n\nBest,\nRiccardo\n",
    "1388515": ""
  }
}