{
  "id": 136084,
  "title": "[RECAP] 41st place solution with code",
  "url": "/competitions/bengaliai-cv19/writeups/cyr1l-recap-41st-place-solution-with-code",
  "author_name": "",
  "post_date": "2020-03-17T15:03:09.373Z",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n\n<h1>Story</h1>\n\n<p>I've joined this competition in the first days of January, with the purpose of learning computer vision from both theoretical and practical sides. I thought that it would be a good idea to finally learn how to write a dl pipeline using config files. I also wanted to dive deeper into Catalyst framework, so Catalyst Config API seemed to fit perfectly.</p>\n\n<p>For the first two weeks, I couldn't get a score higher than 0.955 whilst there were 0.96+ kernels. \nI'd thought that picking a dataset with cropped images is a good idea since I didn't want the model to memorize the position of the character to not overfit on public lb.</p>\n\n<p>Following pieces of advice on forum, I started training with mixup and cutmix. At that moment there was no Cutmix implementation in Catalyst and Mixup implementation didn't support multi-output.\nSo, I adopted code from forum and wrote MixupCutmixCallback, which later resulted in PR with Cutmix implementation to Catalyst library. It was my first time contributing something to open-source deep learning project.</p>\n\n<p>Mixup + Cutmix had given some improvement and I finally crossed 0.955 border. The next step was to go beyond high scoring public kernels (the best one had a score of 0.966 at that time). I tried different architectures (primarily sticked to Efficient Net series), optimizers, schedulers, lrs, losses, augmentations, but 0.966 was still far from me. I got very close to it after I added finetuning on 224x224 images (the first stage was on 128x128 cropped images).  </p>\n\n<p>On the very next day after I got close to 0.966 kernel a new 0.97 kernel came out, so I had to do the climb all over again and at that time I'd tried tweaking almost every possible thing in my pipeline. I <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/131358\">also</a> felt like I was doing some basic thing very wrong, but couldn't identify which exactly. I thought the problem was in misunderstanding some concept in Catalyst, so I decided to rewrite my whole pipeline to Pytorch. This is definitely not something you want to do taking into account there were 3.5 weeks before the end of the competition. Before I started rewriting my pipeline to Pytorch, I decided to train on uncropped images with Catalyst. On the next day I'd finished my Pytorch pipeline and decided to look at the results of training with uncropped images: it was the first time my CV crossed 0.97, resulting in 0.968 LB. Of course, I decided to stay with Catalyst😄 . I took a deeper model and finally beat the 0.97 kernel. </p>\n\n<p>Again, I was out of ideas and again I went on the forum. I skimmed through the main topics and even went through all the posts written by top-scorers individually. I decided to implement the tricks from <a href=\"https://arxiv.org/abs/1912.11370\">this</a> paper posted by <a href=\"/drhabib\">@drhabib</a> , but none of them seem to work. </p>\n\n<p>I started reading other papers, which had been posted on the forum and decided to improve my augmentations. I couldn't fit the geometric ones and thought to tweak <code>alpha</code> in Mixup and Cutmix. I decided to go aggressive with Mixup and set <code>alpha = 4</code>, which I think is pretty reasonable as there are unseen graphemes in the test set.</p>\n\n<p>In the last week, I'd switched to original images, trained 5-fold B3 and added 3-epoch finetuning with no augmentations, which yielded 0.985. I had a pretty small CV/LB gap and decided to choose my highest scoring model.</p>\n\n<h1>Configuration</h1>\n\n<ul>\n<li>Efnet B3 with weights from Noisy Student </li>\n<li>One model with three heads</li>\n<li>opt_level O2</li>\n<li>137x236 images</li>\n<li>CrossEntropyLoss. I guess I could spend some more time tweaking OHEM, it should give a better result.</li>\n<li>Mixup (alpha = 4) + Cutmix (alpha = 1), each with probability of 0.5</li>\n<li>loss weights: 7, 1, 2 respectively</li>\n<li>100 epochs</li>\n<li>AdamW optimizer</li>\n<li>OneCycleSchduler</li>\n<li>finetuning 3 epochs with no augs also on 137x236</li>\n</ul>\n\n<p>You can find my complete config files in the link to my repo provided below.</p>\n\n<h1>Forum</h1>\n\n<p>Thanks, to the people who shared actively on the forum. Namely: <a href=\"/drhabib\">@drhabib</a> , <a href=\"/pestipeti\">@pestipeti</a> , <a href=\"/hengck23\">@hengck23</a> ,  <a href=\"/machinelp\">@machinelp</a> , <a href=\"/haqishen\">@haqishen</a> . You are not just giving hints, but implicitly encourage others to participate in discussions and share as well, as a form of expressing gratitude for what we have learned from your posts and kernels (poor wording, but I hope you get the idea😅 ).</p>\n\n<p>The other important thing related to the forum is I think in this competition many people trapped themselves into hyper-parameters tweaking (myself included). Most posts for this comp were about architecture, number of epochs, augmentations, and even optimizers. I guess some of us can relate that they were jumping back and forth between parameter options after some high-ranked users posted some information about his/her pipeline. </p>\n\n<p>A very few people were discussing validation and strategies to deal with unseen graphemes. </p>\n\n<p>The challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter. </p>\n\n<h1>Takeaways</h1>\n\n<ul>\n<li>Lab book to track your experiments</li>\n<li>Cam activations to better understand what is going on with you model thx <a href=\"/cdeotte\">@cdeotte</a> </li>\n<li>Better error-analysis in general</li>\n<li>Data preprocessing (resize, file extension, etc) which allow you to conduct more experiments</li>\n<li>Use forum wisely</li>\n<li>Don't be afraid to ask questions on forum. The worst question is the question that hasn't' been asked.</li>\n</ul>\n\n<p>And congratulations to everyone who had taken part in this competition regardless of their score!</p>\n\n<p>You can find my code here:</p>\n\n<p><a href=\"https://github.com/LightnessOfBeing/kaggle-bengali-classification\">https://github.com/LightnessOfBeing/kaggle-bengali-classification</a></p>",
  "messages": [
    {
      "id": "776413",
      "postDate": "03/17/2020 11:24:49",
      "content": "<p>Hello everyone!</p>\n\n<h1>Story</h1>\n\n<p>I've joined this competition in the first days of January, with the purpose of learning computer vision from both theoretical and practical sides. I thought that it would be a good idea to finally learn how to write a dl pipeline using config files. I also wanted to dive deeper into Catalyst framework, so Catalyst Config API seemed to fit perfectly.</p>\n\n<p>For the first two weeks, I couldn't get a score higher than 0.955 whilst there were 0.96+ kernels. \nI'd thought that picking a dataset with cropped images is a good idea since I didn't want the model to memorize the position of the character to not overfit on public lb.</p>\n\n<p>Following pieces of advice on forum, I started training with mixup and cutmix. At that moment there was no Cutmix implementation in Catalyst and Mixup implementation didn't support multi-output.\nSo, I adopted code from forum and wrote MixupCutmixCallback, which later resulted in PR with Cutmix implementation to Catalyst library. It was my first time contributing something to open-source deep learning project.</p>\n\n<p>Mixup + Cutmix had given some improvement and I finally crossed 0.955 border. The next step was to go beyond high scoring public kernels (the best one had a score of 0.966 at that time). I tried different architectures (primarily sticked to Efficient Net series), optimizers, schedulers, lrs, losses, augmentations, but 0.966 was still far from me. I got very close to it after I added finetuning on 224x224 images (the first stage was on 128x128 cropped images).  </p>\n\n<p>On the very next day after I got close to 0.966 kernel a new 0.97 kernel came out, so I had to do the climb all over again and at that time I'd tried tweaking almost every possible thing in my pipeline. I <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/131358\">also</a> felt like I was doing some basic thing very wrong, but couldn't identify which exactly. I thought the problem was in misunderstanding some concept in Catalyst, so I decided to rewrite my whole pipeline to Pytorch. This is definitely not something you want to do taking into account there were 3.5 weeks before the end of the competition. Before I started rewriting my pipeline to Pytorch, I decided to train on uncropped images with Catalyst. On the next day I'd finished my Pytorch pipeline and decided to look at the results of training with uncropped images: it was the first time my CV crossed 0.97, resulting in 0.968 LB. Of course, I decided to stay with Catalyst😄 . I took a deeper model and finally beat the 0.97 kernel. </p>\n\n<p>Again, I was out of ideas and again I went on the forum. I skimmed through the main topics and even went through all the posts written by top-scorers individually. I decided to implement the tricks from <a href=\"https://arxiv.org/abs/1912.11370\">this</a> paper posted by <a href=\"/drhabib\">@drhabib</a> , but none of them seem to work. </p>\n\n<p>I started reading other papers, which had been posted on the forum and decided to improve my augmentations. I couldn't fit the geometric ones and thought to tweak <code>alpha</code> in Mixup and Cutmix. I decided to go aggressive with Mixup and set <code>alpha = 4</code>, which I think is pretty reasonable as there are unseen graphemes in the test set.</p>\n\n<p>In the last week, I'd switched to original images, trained 5-fold B3 and added 3-epoch finetuning with no augmentations, which yielded 0.985. I had a pretty small CV/LB gap and decided to choose my highest scoring model.</p>\n\n<h1>Configuration</h1>\n\n<ul>\n<li>Efnet B3 with weights from Noisy Student </li>\n<li>One model with three heads</li>\n<li>opt_level O2</li>\n<li>137x236 images</li>\n<li>CrossEntropyLoss. I guess I could spend some more time tweaking OHEM, it should give a better result.</li>\n<li>Mixup (alpha = 4) + Cutmix (alpha = 1), each with probability of 0.5</li>\n<li>loss weights: 7, 1, 2 respectively</li>\n<li>100 epochs</li>\n<li>AdamW optimizer</li>\n<li>OneCycleSchduler</li>\n<li>finetuning 3 epochs with no augs also on 137x236</li>\n</ul>\n\n<p>You can find my complete config files in the link to my repo provided below.</p>\n\n<h1>Forum</h1>\n\n<p>Thanks, to the people who shared actively on the forum. Namely: <a href=\"/drhabib\">@drhabib</a> , <a href=\"/pestipeti\">@pestipeti</a> , <a href=\"/hengck23\">@hengck23</a> ,  <a href=\"/machinelp\">@machinelp</a> , <a href=\"/haqishen\">@haqishen</a> . You are not just giving hints, but implicitly encourage others to participate in discussions and share as well, as a form of expressing gratitude for what we have learned from your posts and kernels (poor wording, but I hope you get the idea😅 ).</p>\n\n<p>The other important thing related to the forum is I think in this competition many people trapped themselves into hyper-parameters tweaking (myself included). Most posts for this comp were about architecture, number of epochs, augmentations, and even optimizers. I guess some of us can relate that they were jumping back and forth between parameter options after some high-ranked users posted some information about his/her pipeline. </p>\n\n<p>A very few people were discussing validation and strategies to deal with unseen graphemes. </p>\n\n<p>The challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter. </p>\n\n<h1>Takeaways</h1>\n\n<ul>\n<li>Lab book to track your experiments</li>\n<li>Cam activations to better understand what is going on with you model thx <a href=\"/cdeotte\">@cdeotte</a> </li>\n<li>Better error-analysis in general</li>\n<li>Data preprocessing (resize, file extension, etc) which allow you to conduct more experiments</li>\n<li>Use forum wisely</li>\n<li>Don't be afraid to ask questions on forum. The worst question is the question that hasn't' been asked.</li>\n</ul>\n\n<p>And congratulations to everyone who had taken part in this competition regardless of their score!</p>\n\n<p>You can find my code here:</p>\n\n<p><a href=\"https://github.com/LightnessOfBeing/kaggle-bengali-classification\">https://github.com/LightnessOfBeing/kaggle-bengali-classification</a></p>",
      "rawMarkdown": "Hello everyone!\n\n# Story\n\nI've joined this competition in the first days of January, with the purpose of learning computer vision from both theoretical and practical sides. I thought that it would be a good idea to finally learn how to write a dl pipeline using config files. I also wanted to dive deeper into Catalyst framework, so Catalyst Config API seemed to fit perfectly.\n\nFor the first two weeks, I couldn't get a score higher than 0.955 whilst there were 0.96+ kernels. \nI'd thought that picking a dataset with cropped images is a good idea since I didn't want the model to memorize the position of the character to not overfit on public lb.\n\nFollowing pieces of advice on forum, I started training with mixup and cutmix. At that moment there was no Cutmix implementation in Catalyst and Mixup implementation didn't support multi-output.\nSo, I adopted code from forum and wrote MixupCutmixCallback, which later resulted in PR with Cutmix implementation to Catalyst library. It was my first time contributing something to open-source deep learning project.\n\nMixup + Cutmix had given some improvement and I finally crossed 0.955 border. The next step was to go beyond high scoring public kernels (the best one had a score of 0.966 at that time). I tried different architectures (primarily sticked to Efficient Net series), optimizers, schedulers, lrs, losses, augmentations, but 0.966 was still far from me. I got very close to it after I added finetuning on 224x224 images (the first stage was on 128x128 cropped images).  \n\nOn the very next day after I got close to 0.966 kernel a new 0.97 kernel came out, so I had to do the climb all over again and at that time I'd tried tweaking almost every possible thing in my pipeline. I [also](https://www.kaggle.com/c/bengaliai-cv19/discussion/131358) felt like I was doing some basic thing very wrong, but couldn't identify which exactly. I thought the problem was in misunderstanding some concept in Catalyst, so I decided to rewrite my whole pipeline to Pytorch. This is definitely not something you want to do taking into account there were 3.5 weeks before the end of the competition. Before I started rewriting my pipeline to Pytorch, I decided to train on uncropped images with Catalyst. On the next day I'd finished my Pytorch pipeline and decided to look at the results of training with uncropped images: it was the first time my CV crossed 0.97, resulting in 0.968 LB. Of course, I decided to stay with Catalyst😄 . I took a deeper model and finally beat the 0.97 kernel. \n\nAgain, I was out of ideas and again I went on the forum. I skimmed through the main topics and even went through all the posts written by top-scorers individually. I decided to implement the tricks from [this](https://arxiv.org/abs/1912.11370) paper posted by @drhabib , but none of them seem to work. \n\nI started reading other papers, which had been posted on the forum and decided to improve my augmentations. I couldn't fit the geometric ones and thought to tweak `alpha` in Mixup and Cutmix. I decided to go aggressive with Mixup and set `alpha = 4`, which I think is pretty reasonable as there are unseen graphemes in the test set.\n\nIn the last week, I'd switched to original images, trained 5-fold B3 and added 3-epoch finetuning with no augmentations, which yielded 0.985. I had a pretty small CV/LB gap and decided to choose my highest scoring model.\n\n# Configuration\n\n-  Efnet B3 with weights from Noisy Student \n- One model with three heads\n- opt_level O2\n- 137x236 images\n- CrossEntropyLoss. I guess I could spend some more time tweaking OHEM, it should give a better result.\n- Mixup (alpha = 4) + Cutmix (alpha = 1), each with probability of 0.5\n- loss weights: 7, 1, 2 respectively\n- 100 epochs\n- AdamW optimizer\n- OneCycleSchduler\n- finetuning 3 epochs with no augs also on 137x236\n\nYou can find my complete config files in the link to my repo provided below.\n\n# Forum \n\n Thanks, to the people who shared actively on the forum. Namely: @drhabib , @pestipeti , @hengck23 ,  @machinelp , @haqishen . You are not just giving hints, but implicitly encourage others to participate in discussions and share as well, as a form of expressing gratitude for what we have learned from your posts and kernels (poor wording, but I hope you get the idea😅 ).\n\nThe other important thing related to the forum is I think in this competition many people trapped themselves into hyper-parameters tweaking (myself included). Most posts for this comp were about architecture, number of epochs, augmentations, and even optimizers. I guess some of us can relate that they were jumping back and forth between parameter options after some high-ranked users posted some information about his/her pipeline. \n\nA very few people were discussing validation and strategies to deal with unseen graphemes. \n\nThe challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter. \n\n# Takeaways\n\n- Lab book to track your experiments\n- Cam activations to better understand what is going on with you model thx @cdeotte \n- Better error-analysis in general\n- Data preprocessing (resize, file extension, etc) which allow you to conduct more experiments\n- Use forum wisely\n- Don't be afraid to ask questions on forum. The worst question is the question that hasn't' been asked.\n\nAnd congratulations to everyone who had taken part in this competition regardless of their score!\n\nYou can find my code here:\n\nhttps://github.com/LightnessOfBeing/kaggle-bengali-classification",
      "votes": null
    },
    {
      "id": "776582",
      "postDate": "03/17/2020 13:42:07",
      "content": "<p>Thanks for sharing. I have two questions: 1. how did you come up with a loss weight as (7, 1, 2)? 2. how much of the finetuning contributed to the score? Does finetuning have to be done on images without augmentation? </p>",
      "rawMarkdown": "Thanks for sharing. I have two questions: 1. how did you come up with a loss weight as (7, 1, 2)? 2. how much of the finetuning contributed to the score? Does finetuning have to be done on images without augmentation?",
      "votes": null
    },
    {
      "id": "776617",
      "postDate": "03/17/2020 14:06:20",
      "content": "<p>Hi!</p>\n\n<p>1). Originally I used (2, 1, 1) as in official metric. I had had a low score for <code>grapheme_root</code> and I decided to increase its contribution to the loss value, so the network could be punished more for missclassifying <code>grapheme_root</code>. There was an <a href=\"/iafoss\">@iafoss</a> kernel where he was using (0.7 0.1 0.2). This is how I picked (7, 1, 2)</p>\n\n<p>2). As for finetuning: without it my models score is in range (0.9820, 0.9825) CV. Finetuning added about 0.003 - 0.004, which yielded (0.9856 - 0.9867) CV range for single fold models. </p>\n\n<p>Regarding augmentations on finetuning stage you can do whatever you want, I tried some geometric ones, but it didn't work, so I decided to completely get rid of them.</p>",
      "rawMarkdown": "Hi!\n\n1). Originally I used (2, 1, 1) as in official metric. I had had a low score for `grapheme_root` and I decided to increase its contribution to the loss value, so the network could be punished more for missclassifying `grapheme_root`. There was an @iafoss kernel where he was using (0.7 0.1 0.2). This is how I picked (7, 1, 2)\n\n2). As for finetuning: without it my models score is in range (0.9820, 0.9825) CV. Finetuning added about 0.003 - 0.004, which yielded (0.9856 - 0.9867) CV range for single fold models. \n\nRegarding augmentations on finetuning stage you can do whatever you want, I tried some geometric ones, but it didn't work, so I decided to completely get rid of them.",
      "votes": null
    },
    {
      "id": "776684",
      "postDate": "03/17/2020 14:53:25",
      "content": "<p>Got it. Thanks for the elaboration! Does loss weight (7, 1, 2) help? I also tried a few experimentations ranging from (2, 1, 1) up to (5, 1, 1) but it did not work. But I never went up to 7 on the <code>root</code> part and didn't use different weights on <code>V</code> and <code>C</code>.</p>",
      "rawMarkdown": "Got it. Thanks for the elaboration! Does loss weight (7, 1, 2) help? I also tried a few experimentations ranging from (2, 1, 1) up to (5, 1, 1) but it did not work. But I never went up to 7 on the `root` part and didn't use different weights on `V` and `C`.",
      "votes": null
    },
    {
      "id": "776765",
      "postDate": "03/17/2020 16:13:24",
      "content": "<p>Yeah, it helped, but not too much.</p>",
      "rawMarkdown": "Yeah, it helped, but not too much.",
      "votes": null
    },
    {
      "id": "776811",
      "postDate": "03/17/2020 16:41:15",
      "content": "<p>Thank you for your sharing! I agree. I also fell into parameter turning. </p>\n\n<blockquote>\n  <p>The challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter.</p>\n</blockquote>",
      "rawMarkdown": "Thank you for your sharing! I agree. I also fell into parameter turning. \n&gt; The challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 776582,
      "author_name": "gxygomes",
      "author_url": "",
      "post_date": "03/17/2020 13:42:07",
      "content": "<p>Thanks for sharing. I have two questions: 1. how did you come up with a loss weight as (7, 1, 2)? 2. how much of the finetuning contributed to the score? Does finetuning have to be done on images without augmentation? </p>",
      "votes": null,
      "replies": [
        {
          "id": 776617,
          "author_name": "lightnezzofbeing",
          "author_url": "",
          "post_date": "03/17/2020 14:06:20",
          "content": "<p>Hi!</p>\n\n<p>1). Originally I used (2, 1, 1) as in official metric. I had had a low score for <code>grapheme_root</code> and I decided to increase its contribution to the loss value, so the network could be punished more for missclassifying <code>grapheme_root</code>. There was an <a href=\"/iafoss\">@iafoss</a> kernel where he was using (0.7 0.1 0.2). This is how I picked (7, 1, 2)</p>\n\n<p>2). As for finetuning: without it my models score is in range (0.9820, 0.9825) CV. Finetuning added about 0.003 - 0.004, which yielded (0.9856 - 0.9867) CV range for single fold models. </p>\n\n<p>Regarding augmentations on finetuning stage you can do whatever you want, I tried some geometric ones, but it didn't work, so I decided to completely get rid of them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 776684,
          "author_name": "gxygomes",
          "author_url": "",
          "post_date": "03/17/2020 14:53:25",
          "content": "<p>Got it. Thanks for the elaboration! Does loss weight (7, 1, 2) help? I also tried a few experimentations ranging from (2, 1, 1) up to (5, 1, 1) but it did not work. But I never went up to 7 on the <code>root</code> part and didn't use different weights on <code>V</code> and <code>C</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 776765,
          "author_name": "lightnezzofbeing",
          "author_url": "",
          "post_date": "03/17/2020 16:13:24",
          "content": "<p>Yeah, it helped, but not too much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776811,
      "author_name": "statsu",
      "author_url": "",
      "post_date": "03/17/2020 16:41:15",
      "content": "<p>Thank you for your sharing! I agree. I also fell into parameter turning. </p>\n\n<blockquote>\n  <p>The challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "776413": "Hello everyone!\n\n# Story\n\nI've joined this competition in the first days of January, with the purpose of learning computer vision from both theoretical and practical sides. I thought that it would be a good idea to finally learn how to write a dl pipeline using config files. I also wanted to dive deeper into Catalyst framework, so Catalyst Config API seemed to fit perfectly.\n\nFor the first two weeks, I couldn't get a score higher than 0.955 whilst there were 0.96+ kernels. \nI'd thought that picking a dataset with cropped images is a good idea since I didn't want the model to memorize the position of the character to not overfit on public lb.\n\nFollowing pieces of advice on forum, I started training with mixup and cutmix. At that moment there was no Cutmix implementation in Catalyst and Mixup implementation didn't support multi-output.\nSo, I adopted code from forum and wrote MixupCutmixCallback, which later resulted in PR with Cutmix implementation to Catalyst library. It was my first time contributing something to open-source deep learning project.\n\nMixup + Cutmix had given some improvement and I finally crossed 0.955 border. The next step was to go beyond high scoring public kernels (the best one had a score of 0.966 at that time). I tried different architectures (primarily sticked to Efficient Net series), optimizers, schedulers, lrs, losses, augmentations, but 0.966 was still far from me. I got very close to it after I added finetuning on 224x224 images (the first stage was on 128x128 cropped images).  \n\nOn the very next day after I got close to 0.966 kernel a new 0.97 kernel came out, so I had to do the climb all over again and at that time I'd tried tweaking almost every possible thing in my pipeline. I [also](https://www.kaggle.com/c/bengaliai-cv19/discussion/131358) felt like I was doing some basic thing very wrong, but couldn't identify which exactly. I thought the problem was in misunderstanding some concept in Catalyst, so I decided to rewrite my whole pipeline to Pytorch. This is definitely not something you want to do taking into account there were 3.5 weeks before the end of the competition. Before I started rewriting my pipeline to Pytorch, I decided to train on uncropped images with Catalyst. On the next day I'd finished my Pytorch pipeline and decided to look at the results of training with uncropped images: it was the first time my CV crossed 0.97, resulting in 0.968 LB. Of course, I decided to stay with Catalyst😄 . I took a deeper model and finally beat the 0.97 kernel. \n\nAgain, I was out of ideas and again I went on the forum. I skimmed through the main topics and even went through all the posts written by top-scorers individually. I decided to implement the tricks from [this](https://arxiv.org/abs/1912.11370) paper posted by @drhabib , but none of them seem to work. \n\nI started reading other papers, which had been posted on the forum and decided to improve my augmentations. I couldn't fit the geometric ones and thought to tweak `alpha` in Mixup and Cutmix. I decided to go aggressive with Mixup and set `alpha = 4`, which I think is pretty reasonable as there are unseen graphemes in the test set.\n\nIn the last week, I'd switched to original images, trained 5-fold B3 and added 3-epoch finetuning with no augmentations, which yielded 0.985. I had a pretty small CV/LB gap and decided to choose my highest scoring model.\n\n# Configuration\n\n-  Efnet B3 with weights from Noisy Student \n- One model with three heads\n- opt_level O2\n- 137x236 images\n- CrossEntropyLoss. I guess I could spend some more time tweaking OHEM, it should give a better result.\n- Mixup (alpha = 4) + Cutmix (alpha = 1), each with probability of 0.5\n- loss weights: 7, 1, 2 respectively\n- 100 epochs\n- AdamW optimizer\n- OneCycleSchduler\n- finetuning 3 epochs with no augs also on 137x236\n\nYou can find my complete config files in the link to my repo provided below.\n\n# Forum \n\n Thanks, to the people who shared actively on the forum. Namely: @drhabib , @pestipeti , @hengck23 ,  @machinelp , @haqishen . You are not just giving hints, but implicitly encourage others to participate in discussions and share as well, as a form of expressing gratitude for what we have learned from your posts and kernels (poor wording, but I hope you get the idea😅 ).\n\nThe other important thing related to the forum is I think in this competition many people trapped themselves into hyper-parameters tweaking (myself included). Most posts for this comp were about architecture, number of epochs, augmentations, and even optimizers. I guess some of us can relate that they were jumping back and forth between parameter options after some high-ranked users posted some information about his/her pipeline. \n\nA very few people were discussing validation and strategies to deal with unseen graphemes. \n\nThe challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter. \n\n# Takeaways\n\n- Lab book to track your experiments\n- Cam activations to better understand what is going on with you model thx @cdeotte \n- Better error-analysis in general\n- Data preprocessing (resize, file extension, etc) which allow you to conduct more experiments\n- Use forum wisely\n- Don't be afraid to ask questions on forum. The worst question is the question that hasn't' been asked.\n\nAnd congratulations to everyone who had taken part in this competition regardless of their score!\n\nYou can find my code here:\n\nhttps://github.com/LightnessOfBeing/kaggle-bengali-classification",
    "776582": "Thanks for sharing. I have two questions: 1. how did you come up with a loss weight as (7, 1, 2)? 2. how much of the finetuning contributed to the score? Does finetuning have to be done on images without augmentation?",
    "776617": "Hi!\n\n1). Originally I used (2, 1, 1) as in official metric. I had had a low score for `grapheme_root` and I decided to increase its contribution to the loss value, so the network could be punished more for missclassifying `grapheme_root`. There was an @iafoss kernel where he was using (0.7 0.1 0.2). This is how I picked (7, 1, 2)\n\n2). As for finetuning: without it my models score is in range (0.9820, 0.9825) CV. Finetuning added about 0.003 - 0.004, which yielded (0.9856 - 0.9867) CV range for single fold models. \n\nRegarding augmentations on finetuning stage you can do whatever you want, I tried some geometric ones, but it didn't work, so I decided to completely get rid of them.",
    "776684": "Got it. Thanks for the elaboration! Does loss weight (7, 1, 2) help? I also tried a few experimentations ranging from (2, 1, 1) up to (5, 1, 1) but it did not work. But I never went up to 7 on the `root` part and didn't use different weights on `V` and `C`.",
    "776765": "Yeah, it helped, but not too much.",
    "776811": "Thank you for your sharing! I agree. I also fell into parameter turning. \n&gt; The challenge here is the more information we have the harder to find the truth. So, it is very crucial to focus on things that matter."
  },
  "source": "meta"
}