{
  "id": 136075,
  "title": "Solution 38, some lessons learned",
  "url": "/competitions/bengaliai-cv19/writeups/kaz-cpmp-solution-38-some-lessons-learned",
  "author_name": "",
  "post_date": "2020-03-17T10:33:54.119308200Z",
  "votes": 42,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, let me thank my team mate Marios (Kazanova). Without him I'm not sure I would have a medal at all here.  Second, I want to thank all those who shared so much.  This was a goldmine for people with little experience in computer vision like me.  I wish I had paid more attention on some of it, for instance discussion about unseen graphemes in test, or predicting graphemes rather than components as a secondary target.  Last, but not least, many thanks to Bangali AI sponsor and Kaggle for providing a challenging image classification problems with rather small size dataset.  Hardware (access to a number of recent GPU) wasn't the main deciding factor here.  </p>\n\n<p>I am a bit disappointed to have dropped out of gold after being up to top 7th on public LB of course.  At the same time, my first sub is only from 13 days before competition end.  And I could team with Marios who had started earlier.  Still, issue with little time is that you have to be right in selecting what to try and what you decide to no try. And given training computer vision models takes time, you easily lose a day or two on ideas that don't work.  I won't enter computer visions less than 2 weeks before end again.  First lesson learned.</p>\n\n<p>We didn't looked at which targets where hard to predict.  In general, I didn't look at data enough.  Second lesson.</p>\n\n<p>I mentioned some of what we didn't try unfortunately, here is what we tried, with more or less success.</p>\n\n<p>Before teaming I worked on 64x64 images, and a bit on 128x128 images with seresnext50.  Small image sizes leads to high epoch throughput and allow for lots of experiments.  I didn't work on image augmentation and focused on learning pytorch, understanding losses and image preprocessing.  When I teamed, Marios had a efficientnet-b4 model with a good augmentation pipeline.  He used cutout, grid mask, mixup, albumentation, and found that cutmix wasn't effective.  He also used ohem loss, and cosine lr scheduler.  We didn't revisit this except for ohem loss after we teamed.</p>\n\n<p>After teaming we explored ways to add diversity, and ways to combine as many models as possible in the scoring kernel.  In hindsight it also helped overfit a bit more to the public test data.  Our CV LB score gap was rather stable, which also explains why we did't explore more the possible public/private distribution discrepancy.  Third lesson learned, pay more attention to this.</p>\n\n<p>We used an average of 14 efficientnet-b4 models predictions in the end.  Our scoring kernel could handle up to 18 models within 2 hours.  For each model we averaged weights of several checkpoints.  Models differ by:</p>\n\n<ul>\n<li>input data, original size, or crop resize with same image ratio</li>\n<li>fold or full train data</li>\n<li>loss, weighted loss for R, V, C, or weighted loss per class within each of R, V, C.</li>\n<li>ohem or unmodified cross entropy loss</li>\n<li>use of external data or not</li>\n<li>postprocessing </li>\n</ul>\n\n<p>We looked at postprocessing predictions the last few hours before end.  It was clearly promising, and one sub did improve a bit over our best blend.</p>\n\n<p>We only looked at external data, BanglaLekha-Isolated two days before end of competition.  We predicted targets using our best blend, then took the majority class and checked the mapping  by visual inspection between the BanglaLekha paper and the class_map csv.  Our mapping was wrong in few case.  Then we assigned the fixed for R target and the 0 target for V and C.  CV and Lb scores were not really better but it added dieversity to the blend.  Here again, we had to crop resize with same image ratio.</p>\n\n<p>All in all I learned a amazing lot in a short time.  This is where Kaggle is great, you can get a glimpse of state of the art practice in a very short time.  Again, this is only because of sharing from the community members.  I'm now looking forward to next computer vision competition to be honest!</p>\n\n<p>I hope I haven't missed too much, especially on what Marios did before we teamed.  I'll let him expand/correct me if need be.</p>",
  "messages": [
    {
      "id": "776362",
      "postDate": "03/17/2020 10:33:54",
      "content": "<p>First of all, let me thank my team mate Marios (Kazanova). Without him I'm not sure I would have a medal at all here.  Second, I want to thank all those who shared so much.  This was a goldmine for people with little experience in computer vision like me.  I wish I had paid more attention on some of it, for instance discussion about unseen graphemes in test, or predicting graphemes rather than components as a secondary target.  Last, but not least, many thanks to Bangali AI sponsor and Kaggle for providing a challenging image classification problems with rather small size dataset.  Hardware (access to a number of recent GPU) wasn't the main deciding factor here.  </p>\n\n<p>I am a bit disappointed to have dropped out of gold after being up to top 7th on public LB of course.  At the same time, my first sub is only from 13 days before competition end.  And I could team with Marios who had started earlier.  Still, issue with little time is that you have to be right in selecting what to try and what you decide to no try. And given training computer vision models takes time, you easily lose a day or two on ideas that don't work.  I won't enter computer visions less than 2 weeks before end again.  First lesson learned.</p>\n\n<p>We didn't looked at which targets where hard to predict.  In general, I didn't look at data enough.  Second lesson.</p>\n\n<p>I mentioned some of what we didn't try unfortunately, here is what we tried, with more or less success.</p>\n\n<p>Before teaming I worked on 64x64 images, and a bit on 128x128 images with seresnext50.  Small image sizes leads to high epoch throughput and allow for lots of experiments.  I didn't work on image augmentation and focused on learning pytorch, understanding losses and image preprocessing.  When I teamed, Marios had a efficientnet-b4 model with a good augmentation pipeline.  He used cutout, grid mask, mixup, albumentation, and found that cutmix wasn't effective.  He also used ohem loss, and cosine lr scheduler.  We didn't revisit this except for ohem loss after we teamed.</p>\n\n<p>After teaming we explored ways to add diversity, and ways to combine as many models as possible in the scoring kernel.  In hindsight it also helped overfit a bit more to the public test data.  Our CV LB score gap was rather stable, which also explains why we did't explore more the possible public/private distribution discrepancy.  Third lesson learned, pay more attention to this.</p>\n\n<p>We used an average of 14 efficientnet-b4 models predictions in the end.  Our scoring kernel could handle up to 18 models within 2 hours.  For each model we averaged weights of several checkpoints.  Models differ by:</p>\n\n<ul>\n<li>input data, original size, or crop resize with same image ratio</li>\n<li>fold or full train data</li>\n<li>loss, weighted loss for R, V, C, or weighted loss per class within each of R, V, C.</li>\n<li>ohem or unmodified cross entropy loss</li>\n<li>use of external data or not</li>\n<li>postprocessing </li>\n</ul>\n\n<p>We looked at postprocessing predictions the last few hours before end.  It was clearly promising, and one sub did improve a bit over our best blend.</p>\n\n<p>We only looked at external data, BanglaLekha-Isolated two days before end of competition.  We predicted targets using our best blend, then took the majority class and checked the mapping  by visual inspection between the BanglaLekha paper and the class_map csv.  Our mapping was wrong in few case.  Then we assigned the fixed for R target and the 0 target for V and C.  CV and Lb scores were not really better but it added dieversity to the blend.  Here again, we had to crop resize with same image ratio.</p>\n\n<p>All in all I learned a amazing lot in a short time.  This is where Kaggle is great, you can get a glimpse of state of the art practice in a very short time.  Again, this is only because of sharing from the community members.  I'm now looking forward to next computer vision competition to be honest!</p>\n\n<p>I hope I haven't missed too much, especially on what Marios did before we teamed.  I'll let him expand/correct me if need be.</p>",
      "rawMarkdown": "First of all, let me thank my team mate Marios (Kazanova). Without him I'm not sure I would have a medal at all here.  Second, I want to thank all those who shared so much.  This was a goldmine for people with little experience in computer vision like me.  I wish I had paid more attention on some of it, for instance discussion about unseen graphemes in test, or predicting graphemes rather than components as a secondary target.  Last, but not least, many thanks to Bangali AI sponsor and Kaggle for providing a challenging image classification problems with rather small size dataset.  Hardware (access to a number of recent GPU) wasn't the main deciding factor here.  \n\nI am a bit disappointed to have dropped out of gold after being up to top 7th on public LB of course.  At the same time, my first sub is only from 13 days before competition end.  And I could team with Marios who had started earlier.  Still, issue with little time is that you have to be right in selecting what to try and what you decide to no try. And given training computer vision models takes time, you easily lose a day or two on ideas that don't work.  I won't enter computer visions less than 2 weeks before end again.  First lesson learned.\n\nWe didn't looked at which targets where hard to predict.  In general, I didn't look at data enough.  Second lesson.\n\n I mentioned some of what we didn't try unfortunately, here is what we tried, with more or less success.\n\nBefore teaming I worked on 64x64 images, and a bit on 128x128 images with seresnext50.  Small image sizes leads to high epoch throughput and allow for lots of experiments.  I didn't work on image augmentation and focused on learning pytorch, understanding losses and image preprocessing.  When I teamed, Marios had a efficientnet-b4 model with a good augmentation pipeline.  He used cutout, grid mask, mixup, albumentation, and found that cutmix wasn't effective.  He also used ohem loss, and cosine lr scheduler.  We didn't revisit this except for ohem loss after we teamed.\n\nAfter teaming we explored ways to add diversity, and ways to combine as many models as possible in the scoring kernel.  In hindsight it also helped overfit a bit more to the public test data.  Our CV LB score gap was rather stable, which also explains why we did't explore more the possible public/private distribution discrepancy.  Third lesson learned, pay more attention to this.\n\nWe used an average of 14 efficientnet-b4 models predictions in the end.  Our scoring kernel could handle up to 18 models within 2 hours.  For each model we averaged weights of several checkpoints.  Models differ by:\n\n- input data, original size, or crop resize with same image ratio\n- fold or full train data\n- loss, weighted loss for R, V, C, or weighted loss per class within each of R, V, C.\n- ohem or unmodified cross entropy loss\n- use of external data or not\n- postprocessing \n\nWe looked at postprocessing predictions the last few hours before end.  It was clearly promising, and one sub did improve a bit over our best blend.\n\nWe only looked at external data, BanglaLekha-Isolated two days before end of competition.  We predicted targets using our best blend, then took the majority class and checked the mapping  by visual inspection between the BanglaLekha paper and the class_map csv.  Our mapping was wrong in few case.  Then we assigned the fixed for R target and the 0 target for V and C.  CV and Lb scores were not really better but it added dieversity to the blend.  Here again, we had to crop resize with same image ratio.\n\nAll in all I learned a amazing lot in a short time.  This is where Kaggle is great, you can get a glimpse of state of the art practice in a very short time.  Again, this is only because of sharing from the community members.  I'm now looking forward to next computer vision competition to be honest!\n\nI hope I haven't missed too much, especially on what Marios did before we teamed.  I'll let him expand/correct me if need be.",
      "votes": null
    },
    {
      "id": "776389",
      "postDate": "03/17/2020 11:03:13",
      "content": "<p>Thanks for sharing your knowledge, you have been inspiration throughout.</p>",
      "rawMarkdown": "Thanks for sharing your knowledge, you have been inspiration throughout.",
      "votes": null
    },
    {
      "id": "776456",
      "postDate": "03/17/2020 12:05:01",
      "content": "<p>That is a good summary , thank you <a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>Also congrats to the winners!</p>\n\n<p>To expand a little bit more: </p>\n\n<p>The most important part for getting a good score was to find the right mix of augmentations as many people had pointed out on the forums. </p>\n\n<p>In our case the combination that worked best was :\n-   30% mixup (with alpha 0.4)- <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\">https://www.kaggle.com/c/bengaliai-cv19/discussion/126504</a>\n-   50% gridmask with 1 grid. <a href=\"https://www.kaggle.com/haqishen/gridmask\">https://www.kaggle.com/haqishen/gridmask</a>\n-   And 20% cutout (one square piece of size [80,80])</p>\n\n<p>In other words, gridmask was probably the most important augmentation in our case.</p>\n\n<p>For those that may be wondering, Cross Entropy worked much better that Binary cross Entropy. E.g it was better to have 3 outputs with 3 softmax functions than one output with shape of 186. </p>\n\n<p>Other than that training using the full size (137, 236) gave the best results. </p>\n\n<p>Things that may have helped a little bit were:</p>\n\n<p>1)  Larger batch sizes, 128 and then 256 would give slightly better results over 64.\n2)  Ohme loss with a rate of 0.7 : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a>\n3)  Using these  custom loss weights : 6/3 (grapheme) and 2/3, 1/3\n4)  Training with cosine annealing (and hard restarts) for every 5 epochs.\n5)  Radam\n6)  Retraining with small learning rate (and disabling all augmentations) after cosine schedule stopped improving\n7)  Crop and resize, while maintaining aspect ratio. Standard resize was giving close results too. \n8)  Adjusting test predictions based on the distribution of the training data. Like if we expect 10% of class 1, but in test predictions we get only 9% class, we can adjust all predictions for class1 to 10/9 .</p>\n\n<p>Thing that did not manage to make them to work properly although there was a lot about them on the forums: \n-   Cutmix was consistently giving us worse results. I know people reported that it helps, so I will be keen to know what tweaks you had to do to make it work. We tried various alphas with no success. \n-   Augmix (tried various alphas here) and severities\n-   External data/datasets. They did show some promise. but we started using them very late to fully leverage that. </p>",
      "rawMarkdown": "That is a good summary , thank you @cpmpml \n\nAlso congrats to the winners!\n\nTo expand a little bit more: \n\nThe most important part for getting a good score was to find the right mix of augmentations as many people had pointed out on the forums. \n\nIn our case the combination that worked best was :\n-\t30% mixup (with alpha 0.4)- https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\n-\t50% gridmask with 1 grid. https://www.kaggle.com/haqishen/gridmask\n-\tAnd 20% cutout (one square piece of size [80,80])\n\nIn other words, gridmask was probably the most important augmentation in our case.\n\nFor those that may be wondering, Cross Entropy worked much better that Binary cross Entropy. E.g it was better to have 3 outputs with 3 softmax functions than one output with shape of 186. \n\nOther than that training using the full size (137, 236) gave the best results. \n\nThings that may have helped a little bit were:\n\n1)\tLarger batch sizes, 128 and then 256 would give slightly better results over 64.\n2)\tOhme loss with a rate of 0.7 : https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\n3)\tUsing these  custom loss weights : 6/3 (grapheme) and 2/3, 1/3\n4)\tTraining with cosine annealing (and hard restarts) for every 5 epochs.\n5)\tRadam\n6)\tRetraining with small learning rate (and disabling all augmentations) after cosine schedule stopped improving\n7)\tCrop and resize, while maintaining aspect ratio. Standard resize was giving close results too. \n8)\tAdjusting test predictions based on the distribution of the training data. Like if we expect 10% of class 1, but in test predictions we get only 9% class, we can adjust all predictions for class1 to 10/9 .\n\nThing that did not manage to make them to work properly although there was a lot about them on the forums: \n-\tCutmix was consistently giving us worse results. I know people reported that it helps, so I will be keen to know what tweaks you had to do to make it work. We tried various alphas with no success. \n-\tAugmix (tried various alphas here) and severities\n-\tExternal data/datasets. They did show some promise. but we started using them very late to fully leverage that.",
      "votes": null
    },
    {
      "id": "777497",
      "postDate": "03/17/2020 18:13:23",
      "content": "<p>Using the post processing from Chris Deotte's writeup we get private 0.9672 and public 0.9950.  Post processing was definitely one way to go!</p>",
      "rawMarkdown": "Using the post processing from Chris Deotte's writeup we get private 0.9672 and public 0.9950.  Post processing was definitely one way to go!",
      "votes": null
    },
    {
      "id": "777499",
      "postDate": "03/17/2020 18:14:09",
      "content": "<p>I doubt 9972 in private :D</p>",
      "rawMarkdown": "I doubt 9972 in private :D",
      "votes": null
    },
    {
      "id": "777501",
      "postDate": "03/17/2020 18:14:56",
      "content": "<p>Typo, fixed ;)</p>",
      "rawMarkdown": "Typo, fixed ;)",
      "votes": null
    },
    {
      "id": "777605",
      "postDate": "03/17/2020 20:13:15",
      "content": "<p>Great work Kaz and CPMP. I used my best cutout rectangle size as my size for CutMix and that improved upon cutout.</p>\n\n<p>I used two tricks to get CutMix to work. First (1) for the second image, I chose any image in the entire training set, not inside the batch. (2) I restricted the size of the replacement rectangle to roughly a quarter of total image area.</p>\n\n<p>Given image size (Y,X), I would randomly choose a rectangle with <code>7*X/16 &lt; W &lt; 9*X/16</code> by <code>7*Y/16 &lt; H &lt; 9*Y/16</code>. The center of this rectangle was uniform on <code>0 &lt; d0 &lt; H</code> and <code>0 &lt; d1 &lt; W</code>. I replaced this rectangle inside first image with second image. And updated the OHE label according based on area of replacement.</p>\n\n<p>I then converted this CutMix which was 19% - 31% image area into CAM CutMix of 15% - 25% image area and that improved upon CutMix. (Actually 50% CutMix 50% CAM CutMix)</p>",
      "rawMarkdown": "Great work Kaz and CPMP. I used my best cutout rectangle size as my size for CutMix and that improved upon cutout.\n\nI used two tricks to get CutMix to work. First (1) for the second image, I chose any image in the entire training set, not inside the batch. (2) I restricted the size of the replacement rectangle to roughly a quarter of total image area.\n\nGiven image size (Y,X), I would randomly choose a rectangle with `7*X/16 &lt; W &lt; 9*X/16` by `7*Y/16 &lt; H &lt; 9*Y/16`. The center of this rectangle was uniform on `0 &lt; d0 &lt; H` and `0 &lt; d1 &lt; W`. I replaced this rectangle inside first image with second image. And updated the OHE label according based on area of replacement.\n\nI then converted this CutMix which was 19% - 31% image area into CAM CutMix of 15% - 25% image area and that improved upon CutMix. (Actually 50% CutMix 50% CAM CutMix)",
      "votes": null
    },
    {
      "id": "777608",
      "postDate": "03/17/2020 20:15:11",
      "content": "<p>That's awesome CPMP. The effectiveness of post process makes me wonder how much of the top teams' scores are a result of advanced architecture to predict unseen graphemes or how much those techniques just balance the classes.</p>",
      "rawMarkdown": "That's awesome CPMP. The effectiveness of post process makes me wonder how much of the top teams' scores are a result of advanced architecture to predict unseen graphemes or how much those techniques just balance the classes.",
      "votes": null
    },
    {
      "id": "777613",
      "postDate": "03/17/2020 20:20:36",
      "content": "<p>Congrats CPMP and Kaz. Thanks for sharing.</p>",
      "rawMarkdown": "Congrats CPMP and Kaz. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "778138",
      "postDate": "03/18/2020 07:19:14",
      "content": "<p>Congrats <a href=\"/cpmpml\">@cpmpml</a> and nice write-up!</p>",
      "rawMarkdown": "Congrats @cpmpml and nice write-up!",
      "votes": null
    },
    {
      "id": "778258",
      "postDate": "03/18/2020 09:28:25",
      "content": "<p>My hunch is that models predicting grapheme are not suffering from class imbalance as models predicting components.  This si from reading writeups, not from proper experiment, hence may not be right.</p>",
      "rawMarkdown": "My hunch is that models predicting grapheme are not suffering from class imbalance as models predicting components.  This si from reading writeups, not from proper experiment, hence may not be right.",
      "votes": null
    },
    {
      "id": "778474",
      "postDate": "03/18/2020 13:32:01",
      "content": "<p>Congratulations and thank you for sharing your insights. </p>",
      "rawMarkdown": "Congratulations and thank you for sharing your insights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 776389,
      "author_name": "chandraroy",
      "author_url": "",
      "post_date": "03/17/2020 11:03:13",
      "content": "<p>Thanks for sharing your knowledge, you have been inspiration throughout.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776456,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "03/17/2020 12:05:01",
      "content": "<p>That is a good summary , thank you <a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>Also congrats to the winners!</p>\n\n<p>To expand a little bit more: </p>\n\n<p>The most important part for getting a good score was to find the right mix of augmentations as many people had pointed out on the forums. </p>\n\n<p>In our case the combination that worked best was :\n-   30% mixup (with alpha 0.4)- <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\">https://www.kaggle.com/c/bengaliai-cv19/discussion/126504</a>\n-   50% gridmask with 1 grid. <a href=\"https://www.kaggle.com/haqishen/gridmask\">https://www.kaggle.com/haqishen/gridmask</a>\n-   And 20% cutout (one square piece of size [80,80])</p>\n\n<p>In other words, gridmask was probably the most important augmentation in our case.</p>\n\n<p>For those that may be wondering, Cross Entropy worked much better that Binary cross Entropy. E.g it was better to have 3 outputs with 3 softmax functions than one output with shape of 186. </p>\n\n<p>Other than that training using the full size (137, 236) gave the best results. </p>\n\n<p>Things that may have helped a little bit were:</p>\n\n<p>1)  Larger batch sizes, 128 and then 256 would give slightly better results over 64.\n2)  Ohme loss with a rate of 0.7 : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a>\n3)  Using these  custom loss weights : 6/3 (grapheme) and 2/3, 1/3\n4)  Training with cosine annealing (and hard restarts) for every 5 epochs.\n5)  Radam\n6)  Retraining with small learning rate (and disabling all augmentations) after cosine schedule stopped improving\n7)  Crop and resize, while maintaining aspect ratio. Standard resize was giving close results too. \n8)  Adjusting test predictions based on the distribution of the training data. Like if we expect 10% of class 1, but in test predictions we get only 9% class, we can adjust all predictions for class1 to 10/9 .</p>\n\n<p>Thing that did not manage to make them to work properly although there was a lot about them on the forums: \n-   Cutmix was consistently giving us worse results. I know people reported that it helps, so I will be keen to know what tweaks you had to do to make it work. We tried various alphas with no success. \n-   Augmix (tried various alphas here) and severities\n-   External data/datasets. They did show some promise. but we started using them very late to fully leverage that. </p>",
      "votes": null,
      "replies": [
        {
          "id": 777605,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/17/2020 20:13:15",
          "content": "<p>Great work Kaz and CPMP. I used my best cutout rectangle size as my size for CutMix and that improved upon cutout.</p>\n\n<p>I used two tricks to get CutMix to work. First (1) for the second image, I chose any image in the entire training set, not inside the batch. (2) I restricted the size of the replacement rectangle to roughly a quarter of total image area.</p>\n\n<p>Given image size (Y,X), I would randomly choose a rectangle with <code>7*X/16 &lt; W &lt; 9*X/16</code> by <code>7*Y/16 &lt; H &lt; 9*Y/16</code>. The center of this rectangle was uniform on <code>0 &lt; d0 &lt; H</code> and <code>0 &lt; d1 &lt; W</code>. I replaced this rectangle inside first image with second image. And updated the OHE label according based on area of replacement.</p>\n\n<p>I then converted this CutMix which was 19% - 31% image area into CAM CutMix of 15% - 25% image area and that improved upon CutMix. (Actually 50% CutMix 50% CAM CutMix)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 777497,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/17/2020 18:13:23",
      "content": "<p>Using the post processing from Chris Deotte's writeup we get private 0.9672 and public 0.9950.  Post processing was definitely one way to go!</p>",
      "votes": null,
      "replies": [
        {
          "id": 777499,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/17/2020 18:14:09",
          "content": "<p>I doubt 9972 in private :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 777501,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/17/2020 18:14:56",
          "content": "<p>Typo, fixed ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 777608,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/17/2020 20:15:11",
          "content": "<p>That's awesome CPMP. The effectiveness of post process makes me wonder how much of the top teams' scores are a result of advanced architecture to predict unseen graphemes or how much those techniques just balance the classes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 778258,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/18/2020 09:28:25",
          "content": "<p>My hunch is that models predicting grapheme are not suffering from class imbalance as models predicting components.  This si from reading writeups, not from proper experiment, hence may not be right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 777613,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/17/2020 20:20:36",
      "content": "<p>Congrats CPMP and Kaz. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 778138,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "03/18/2020 07:19:14",
      "content": "<p>Congrats <a href=\"/cpmpml\">@cpmpml</a> and nice write-up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 778474,
      "author_name": "adityasingh177",
      "author_url": "",
      "post_date": "03/18/2020 13:32:01",
      "content": "<p>Congratulations and thank you for sharing your insights. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "776362": "First of all, let me thank my team mate Marios (Kazanova). Without him I'm not sure I would have a medal at all here.  Second, I want to thank all those who shared so much.  This was a goldmine for people with little experience in computer vision like me.  I wish I had paid more attention on some of it, for instance discussion about unseen graphemes in test, or predicting graphemes rather than components as a secondary target.  Last, but not least, many thanks to Bangali AI sponsor and Kaggle for providing a challenging image classification problems with rather small size dataset.  Hardware (access to a number of recent GPU) wasn't the main deciding factor here.  \n\nI am a bit disappointed to have dropped out of gold after being up to top 7th on public LB of course.  At the same time, my first sub is only from 13 days before competition end.  And I could team with Marios who had started earlier.  Still, issue with little time is that you have to be right in selecting what to try and what you decide to no try. And given training computer vision models takes time, you easily lose a day or two on ideas that don't work.  I won't enter computer visions less than 2 weeks before end again.  First lesson learned.\n\nWe didn't looked at which targets where hard to predict.  In general, I didn't look at data enough.  Second lesson.\n\n I mentioned some of what we didn't try unfortunately, here is what we tried, with more or less success.\n\nBefore teaming I worked on 64x64 images, and a bit on 128x128 images with seresnext50.  Small image sizes leads to high epoch throughput and allow for lots of experiments.  I didn't work on image augmentation and focused on learning pytorch, understanding losses and image preprocessing.  When I teamed, Marios had a efficientnet-b4 model with a good augmentation pipeline.  He used cutout, grid mask, mixup, albumentation, and found that cutmix wasn't effective.  He also used ohem loss, and cosine lr scheduler.  We didn't revisit this except for ohem loss after we teamed.\n\nAfter teaming we explored ways to add diversity, and ways to combine as many models as possible in the scoring kernel.  In hindsight it also helped overfit a bit more to the public test data.  Our CV LB score gap was rather stable, which also explains why we did't explore more the possible public/private distribution discrepancy.  Third lesson learned, pay more attention to this.\n\nWe used an average of 14 efficientnet-b4 models predictions in the end.  Our scoring kernel could handle up to 18 models within 2 hours.  For each model we averaged weights of several checkpoints.  Models differ by:\n\n- input data, original size, or crop resize with same image ratio\n- fold or full train data\n- loss, weighted loss for R, V, C, or weighted loss per class within each of R, V, C.\n- ohem or unmodified cross entropy loss\n- use of external data or not\n- postprocessing \n\nWe looked at postprocessing predictions the last few hours before end.  It was clearly promising, and one sub did improve a bit over our best blend.\n\nWe only looked at external data, BanglaLekha-Isolated two days before end of competition.  We predicted targets using our best blend, then took the majority class and checked the mapping  by visual inspection between the BanglaLekha paper and the class_map csv.  Our mapping was wrong in few case.  Then we assigned the fixed for R target and the 0 target for V and C.  CV and Lb scores were not really better but it added dieversity to the blend.  Here again, we had to crop resize with same image ratio.\n\nAll in all I learned a amazing lot in a short time.  This is where Kaggle is great, you can get a glimpse of state of the art practice in a very short time.  Again, this is only because of sharing from the community members.  I'm now looking forward to next computer vision competition to be honest!\n\nI hope I haven't missed too much, especially on what Marios did before we teamed.  I'll let him expand/correct me if need be.",
    "776389": "Thanks for sharing your knowledge, you have been inspiration throughout.",
    "776456": "That is a good summary , thank you @cpmpml \n\nAlso congrats to the winners!\n\nTo expand a little bit more: \n\nThe most important part for getting a good score was to find the right mix of augmentations as many people had pointed out on the forums. \n\nIn our case the combination that worked best was :\n-\t30% mixup (with alpha 0.4)- https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\n-\t50% gridmask with 1 grid. https://www.kaggle.com/haqishen/gridmask\n-\tAnd 20% cutout (one square piece of size [80,80])\n\nIn other words, gridmask was probably the most important augmentation in our case.\n\nFor those that may be wondering, Cross Entropy worked much better that Binary cross Entropy. E.g it was better to have 3 outputs with 3 softmax functions than one output with shape of 186. \n\nOther than that training using the full size (137, 236) gave the best results. \n\nThings that may have helped a little bit were:\n\n1)\tLarger batch sizes, 128 and then 256 would give slightly better results over 64.\n2)\tOhme loss with a rate of 0.7 : https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\n3)\tUsing these  custom loss weights : 6/3 (grapheme) and 2/3, 1/3\n4)\tTraining with cosine annealing (and hard restarts) for every 5 epochs.\n5)\tRadam\n6)\tRetraining with small learning rate (and disabling all augmentations) after cosine schedule stopped improving\n7)\tCrop and resize, while maintaining aspect ratio. Standard resize was giving close results too. \n8)\tAdjusting test predictions based on the distribution of the training data. Like if we expect 10% of class 1, but in test predictions we get only 9% class, we can adjust all predictions for class1 to 10/9 .\n\nThing that did not manage to make them to work properly although there was a lot about them on the forums: \n-\tCutmix was consistently giving us worse results. I know people reported that it helps, so I will be keen to know what tweaks you had to do to make it work. We tried various alphas with no success. \n-\tAugmix (tried various alphas here) and severities\n-\tExternal data/datasets. They did show some promise. but we started using them very late to fully leverage that.",
    "777497": "Using the post processing from Chris Deotte's writeup we get private 0.9672 and public 0.9950.  Post processing was definitely one way to go!",
    "777499": "I doubt 9972 in private :D",
    "777501": "Typo, fixed ;)",
    "777605": "Great work Kaz and CPMP. I used my best cutout rectangle size as my size for CutMix and that improved upon cutout.\n\nI used two tricks to get CutMix to work. First (1) for the second image, I chose any image in the entire training set, not inside the batch. (2) I restricted the size of the replacement rectangle to roughly a quarter of total image area.\n\nGiven image size (Y,X), I would randomly choose a rectangle with `7*X/16 &lt; W &lt; 9*X/16` by `7*Y/16 &lt; H &lt; 9*Y/16`. The center of this rectangle was uniform on `0 &lt; d0 &lt; H` and `0 &lt; d1 &lt; W`. I replaced this rectangle inside first image with second image. And updated the OHE label according based on area of replacement.\n\nI then converted this CutMix which was 19% - 31% image area into CAM CutMix of 15% - 25% image area and that improved upon CutMix. (Actually 50% CutMix 50% CAM CutMix)",
    "777608": "That's awesome CPMP. The effectiveness of post process makes me wonder how much of the top teams' scores are a result of advanced architecture to predict unseen graphemes or how much those techniques just balance the classes.",
    "777613": "Congrats CPMP and Kaz. Thanks for sharing.",
    "778138": "Congrats @cpmpml and nice write-up!",
    "778258": "My hunch is that models predicting grapheme are not suffering from class imbalance as models predicting components.  This si from reading writeups, not from proper experiment, hence may not be right.",
    "778474": "Congratulations and thank you for sharing your insights."
  },
  "source": "meta"
}