{
  "id": 136799,
  "title": "55th place solution - A frugal approach",
  "url": "/competitions/bengaliai-cv19/writeups/one-hot-encoder-55th-place-solution-a-frugal-appro",
  "author_name": "",
  "post_date": "2020-03-19T07:38:33.877Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks Kaggle and Bengali.ai for hosting this contest. Congratulations to all the winners and all the participants who kept the discussion alive and encouraging throughout the contest with few special shoutouts to <a href=\"/haqishen\">@haqishen</a>, <a href=\"/pestipeti\">@pestipeti</a>, <a href=\"/hengck23\">@hengck23</a>, <a href=\"/ildoonet\">@ildoonet</a>, <a href=\"/bibek777\">@bibek777</a>, <a href=\"/machinelp\">@machinelp</a> for really interesting discussions inundating with ideas. Special thanks to <a href=\"/iafoss\">@iafoss</a> for his starter notebook and <a href=\"/drhabib\">@drhabib</a> for posting detailed code of previous competition solutions, which I  relied upon. </p>\n\n<p>This is my first competition medal on Kaggle in 2 years. The public leaderboard standing is from one single model. Due to lack of hardware resources, most of my code was trained on Kaggle Kernels (The reason why I call this a frugal approach ;p). </p>\n\n<h2>Preprocessing and Augmentations</h2>\n\n<p>I resized the images to <strong>240x240x1</strong> and zero-padded them. Simple augmentations such *as crop_resize, image_wrapping, rotations, and image_lightening* were used. </p>\n\n<h2>Validation</h2>\n\n<p>I experimented with <strong>5 fold CV</strong> and for submission notebook, I performed a <strong>80-20</strong> split of the dataset. </p>\n\n<h2>Backbone and architecture</h2>\n\n<p>I performed most of the ideas shared on discussion forums keeping EfficientNet-B1 as the model backbone, which I took from EfficientNet Pytorch, and carried out certain modifications inspired from <a href=\"/iafoss\">@iafoss</a>'s starter kernel. EfficientNet-B3 and above, being computationally expensive for Kaggle kernel, could not be trained properly in the limited kernel runtime of 32000s.</p>\n\n<h2>What didn't work</h2>\n\n<p>Although MixUp and Cutmix augmentation seemed really promising, they provided no effective improvement for me. They require higher epochs for training which was not at all feasible in my case.</p>\n\n<p>Based on some previous contest discussion advice, it was helpful to maintain experiment log files. I naively created the experiment log, but, it turned out to be confusing rather than helping when selecting the final solution for submission. </p>\n\n<p>Lack of hardware was a major turndown for me at the beginning of the contest and no improvement in the public lb was discouraging. The time that was wasted without any submission could have been used crucially. I timed every single epoch and then divided 32000s from that in order to train my model for longer and longer time. I used half precision and had to replace <em>Mish activation</em> in the code with <em>ReLU</em> for faster computation (It gave me almost one extra epoch). </p>\n\n<h3>What I learned during the competition ( note to self )</h3>\n\n<ul>\n<li>Prepare and maintain a proper experiment log file throughout the contest (Any help with this would be appreciated)</li>\n<li>Discussion forums are really interesting and full of ideas to experiment with but proper experiment documentation is also necessary.</li>\n<li>Proper selection of solution for submission is extremely important, this helped me through LB shakeup.</li>\n</ul>",
  "messages": [
    {
      "id": "777734",
      "postDate": "03/17/2020 22:08:04",
      "content": "<p>Thanks Kaggle and Bengali.ai for hosting this contest. Congratulations to all the winners and all the participants who kept the discussion alive and encouraging throughout the contest with few special shoutouts to <a href=\"/haqishen\">@haqishen</a>, <a href=\"/pestipeti\">@pestipeti</a>, <a href=\"/hengck23\">@hengck23</a>, <a href=\"/ildoonet\">@ildoonet</a>, <a href=\"/bibek777\">@bibek777</a>, <a href=\"/machinelp\">@machinelp</a> for really interesting discussions inundating with ideas. Special thanks to <a href=\"/iafoss\">@iafoss</a> for his starter notebook and <a href=\"/drhabib\">@drhabib</a> for posting detailed code of previous competition solutions, which I  relied upon. </p>\n\n<p>This is my first competition medal on Kaggle in 2 years. The public leaderboard standing is from one single model. Due to lack of hardware resources, most of my code was trained on Kaggle Kernels (The reason why I call this a frugal approach ;p). </p>\n\n<h2>Preprocessing and Augmentations</h2>\n\n<p>I resized the images to <strong>240x240x1</strong> and zero-padded them. Simple augmentations such *as crop_resize, image_wrapping, rotations, and image_lightening* were used. </p>\n\n<h2>Validation</h2>\n\n<p>I experimented with <strong>5 fold CV</strong> and for submission notebook, I performed a <strong>80-20</strong> split of the dataset. </p>\n\n<h2>Backbone and architecture</h2>\n\n<p>I performed most of the ideas shared on discussion forums keeping EfficientNet-B1 as the model backbone, which I took from EfficientNet Pytorch, and carried out certain modifications inspired from <a href=\"/iafoss\">@iafoss</a>'s starter kernel. EfficientNet-B3 and above, being computationally expensive for Kaggle kernel, could not be trained properly in the limited kernel runtime of 32000s.</p>\n\n<h2>What didn't work</h2>\n\n<p>Although MixUp and Cutmix augmentation seemed really promising, they provided no effective improvement for me. They require higher epochs for training which was not at all feasible in my case.</p>\n\n<p>Based on some previous contest discussion advice, it was helpful to maintain experiment log files. I naively created the experiment log, but, it turned out to be confusing rather than helping when selecting the final solution for submission. </p>\n\n<p>Lack of hardware was a major turndown for me at the beginning of the contest and no improvement in the public lb was discouraging. The time that was wasted without any submission could have been used crucially. I timed every single epoch and then divided 32000s from that in order to train my model for longer and longer time. I used half precision and had to replace <em>Mish activation</em> in the code with <em>ReLU</em> for faster computation (It gave me almost one extra epoch). </p>\n\n<h3>What I learned during the competition ( note to self )</h3>\n\n<ul>\n<li>Prepare and maintain a proper experiment log file throughout the contest (Any help with this would be appreciated)</li>\n<li>Discussion forums are really interesting and full of ideas to experiment with but proper experiment documentation is also necessary.</li>\n<li>Proper selection of solution for submission is extremely important, this helped me through LB shakeup.</li>\n</ul>",
      "rawMarkdown": "Thanks Kaggle and Bengali.ai for hosting this contest. Congratulations to all the winners and all the participants who kept the discussion alive and encouraging throughout the contest with few special shoutouts to @haqishen, @pestipeti, @hengck23, @ildoonet, @bibek777, @machinelp for really interesting discussions inundating with ideas. Special thanks to @iafoss for his starter notebook and @drhabib for posting detailed code of previous competition solutions, which I  relied upon. \n\nThis is my first competition medal on Kaggle in 2 years. The public leaderboard standing is from one single model. Due to lack of hardware resources, most of my code was trained on Kaggle Kernels (The reason why I call this a frugal approach ;p). \n\n## Preprocessing and Augmentations\nI resized the images to **240x240x1** and zero-padded them. Simple augmentations such *as crop_resize, image_wrapping, rotations, and image_lightening* were used. \n\n## Validation\nI experimented with **5 fold CV** and for submission notebook, I performed a **80-20** split of the dataset. \n\n## Backbone and architecture\nI performed most of the ideas shared on discussion forums keeping EfficientNet-B1 as the model backbone, which I took from EfficientNet Pytorch, and carried out certain modifications inspired from @iafoss's starter kernel. EfficientNet-B3 and above, being computationally expensive for Kaggle kernel, could not be trained properly in the limited kernel runtime of 32000s.\n\n## What didn't work\nAlthough MixUp and Cutmix augmentation seemed really promising, they provided no effective improvement for me. They require higher epochs for training which was not at all feasible in my case.\n\nBased on some previous contest discussion advice, it was helpful to maintain experiment log files. I naively created the experiment log, but, it turned out to be confusing rather than helping when selecting the final solution for submission. \n\nLack of hardware was a major turndown for me at the beginning of the contest and no improvement in the public lb was discouraging. The time that was wasted without any submission could have been used crucially. I timed every single epoch and then divided 32000s from that in order to train my model for longer and longer time. I used half precision and had to replace _Mish activation_ in the code with _ReLU_ for faster computation (It gave me almost one extra epoch). \n\n### What I learned during the competition ( note to self )\n* Prepare and maintain a proper experiment log file throughout the contest (Any help with this would be appreciated)\n* Discussion forums are really interesting and full of ideas to experiment with but proper experiment documentation is also necessary.\n* Proper selection of solution for submission is extremely important, this helped me through LB shakeup.",
      "votes": null
    },
    {
      "id": "777955",
      "postDate": "03/18/2020 03:24:28",
      "content": "<p>What're your public LB scores?</p>",
      "rawMarkdown": "What're your public LB scores?",
      "votes": null
    },
    {
      "id": "778163",
      "postDate": "03/18/2020 07:42:15",
      "content": "<p>Here are my two selected submissions stats:</p>\n\n<p><strong>Model backbone: EfficientNet-B1 with mixup</strong>\ntotal epochs: 50\nLearning rate: from 1e-6 to 2e-2\nVal score: 0.9779\nPublic LB score: 0.9638\nPrivate LB score: 0.9360</p>\n\n<p><strong>Model backbone: EfficientNet-B1</strong>\ntotal epochs: 18\nLearning rate: from 1e-4 to 1e-2\nCV score: 0.9717\nLB score: 0.9620\nPrivate LB score: 0.9385</p>",
      "rawMarkdown": "Here are my two selected submissions stats:\n\n**Model backbone: EfficientNet-B1 with mixup**\ntotal epochs: 50\nLearning rate: from 1e-6 to 2e-2\nVal score: 0.9779\nPublic LB score: 0.9638\nPrivate LB score: 0.9360\n\n**Model backbone: EfficientNet-B1**\ntotal epochs: 18\nLearning rate: from 1e-4 to 1e-2\nCV score: 0.9717\nLB score: 0.9620\nPrivate LB score: 0.9385",
      "votes": null
    },
    {
      "id": "778736",
      "postDate": "03/18/2020 17:26:30",
      "content": "<p><a href=\"/thanatoz\">@thanatoz</a> Thank you! That confirmed my feeling that efficientnet as a smaller model tends to underfit and perform well on private LB.</p>",
      "rawMarkdown": "thanatoz Thank you! That confirmed my feeling that efficientnet as a smaller model tends to underfit and perform well on private LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 777955,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "03/18/2020 03:24:28",
      "content": "<p>What're your public LB scores?</p>",
      "votes": null,
      "replies": [
        {
          "id": 778163,
          "author_name": "thanatoz",
          "author_url": "",
          "post_date": "03/18/2020 07:42:15",
          "content": "<p>Here are my two selected submissions stats:</p>\n\n<p><strong>Model backbone: EfficientNet-B1 with mixup</strong>\ntotal epochs: 50\nLearning rate: from 1e-6 to 2e-2\nVal score: 0.9779\nPublic LB score: 0.9638\nPrivate LB score: 0.9360</p>\n\n<p><strong>Model backbone: EfficientNet-B1</strong>\ntotal epochs: 18\nLearning rate: from 1e-4 to 1e-2\nCV score: 0.9717\nLB score: 0.9620\nPrivate LB score: 0.9385</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 778736,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "03/18/2020 17:26:30",
          "content": "<p><a href=\"/thanatoz\">@thanatoz</a> Thank you! That confirmed my feeling that efficientnet as a smaller model tends to underfit and perform well on private LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "777734": "Thanks Kaggle and Bengali.ai for hosting this contest. Congratulations to all the winners and all the participants who kept the discussion alive and encouraging throughout the contest with few special shoutouts to @haqishen, @pestipeti, @hengck23, @ildoonet, @bibek777, @machinelp for really interesting discussions inundating with ideas. Special thanks to @iafoss for his starter notebook and @drhabib for posting detailed code of previous competition solutions, which I  relied upon. \n\nThis is my first competition medal on Kaggle in 2 years. The public leaderboard standing is from one single model. Due to lack of hardware resources, most of my code was trained on Kaggle Kernels (The reason why I call this a frugal approach ;p). \n\n## Preprocessing and Augmentations\nI resized the images to **240x240x1** and zero-padded them. Simple augmentations such *as crop_resize, image_wrapping, rotations, and image_lightening* were used. \n\n## Validation\nI experimented with **5 fold CV** and for submission notebook, I performed a **80-20** split of the dataset. \n\n## Backbone and architecture\nI performed most of the ideas shared on discussion forums keeping EfficientNet-B1 as the model backbone, which I took from EfficientNet Pytorch, and carried out certain modifications inspired from @iafoss's starter kernel. EfficientNet-B3 and above, being computationally expensive for Kaggle kernel, could not be trained properly in the limited kernel runtime of 32000s.\n\n## What didn't work\nAlthough MixUp and Cutmix augmentation seemed really promising, they provided no effective improvement for me. They require higher epochs for training which was not at all feasible in my case.\n\nBased on some previous contest discussion advice, it was helpful to maintain experiment log files. I naively created the experiment log, but, it turned out to be confusing rather than helping when selecting the final solution for submission. \n\nLack of hardware was a major turndown for me at the beginning of the contest and no improvement in the public lb was discouraging. The time that was wasted without any submission could have been used crucially. I timed every single epoch and then divided 32000s from that in order to train my model for longer and longer time. I used half precision and had to replace _Mish activation_ in the code with _ReLU_ for faster computation (It gave me almost one extra epoch). \n\n### What I learned during the competition ( note to self )\n* Prepare and maintain a proper experiment log file throughout the contest (Any help with this would be appreciated)\n* Discussion forums are really interesting and full of ideas to experiment with but proper experiment documentation is also necessary.\n* Proper selection of solution for submission is extremely important, this helped me through LB shakeup.",
    "777955": "What're your public LB scores?",
    "778163": "Here are my two selected submissions stats:\n\n**Model backbone: EfficientNet-B1 with mixup**\ntotal epochs: 50\nLearning rate: from 1e-6 to 2e-2\nVal score: 0.9779\nPublic LB score: 0.9638\nPrivate LB score: 0.9360\n\n**Model backbone: EfficientNet-B1**\ntotal epochs: 18\nLearning rate: from 1e-4 to 1e-2\nCV score: 0.9717\nLB score: 0.9620\nPrivate LB score: 0.9385",
    "778736": "thanatoz Thank you! That confirmed my feeling that efficientnet as a smaller model tends to underfit and perform well on private LB."
  },
  "source": "meta"
}