{
  "id": 136780,
  "title": "38th Place - ZoneMix - A domain specific augmentation",
  "url": "/competitions/bengaliai-cv19/writeups/dipam-ram-38th-place-zonemix-a-domain-specific-aug",
  "author_name": "",
  "post_date": "2020-03-17T20:31:50.885466600Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>A really big thanks to all the amazing participants for sharing the ideas. @haqishen @hengck23  @bibek777 @machinelp @quandapro and so many other awesome users without whose suggestions I would have been completely lost. And also big thanks to Bengali.AI organizers for this competition about my native language, the lessons learned from this competition will also be applicable to other languages of the Indian subcontinent like Hindi, Odia, Telegu, etc. </p>\n\n<p>I'll be honest, I joined the competition by at nearly the end of January, and for most of the time, I was trying to improve the score just by simple augmentations/model/hyperparameter tuning to improve the score. And having only Google Colab to train doesn't help with that either. Had I continued on that track I surely wouldn't have been in any medal zone. Thanks a lot to my teammate Ram for pointing out the important discusions about unseen graphemes by the organizers. And subsequently Qishen Ha's post on validating on unseen graphemes. </p>\n\n<p>With only 10 days remaining after realizing how poorly the model will do on unseen graphemes. Particularly - Classes with few combinations - Which is more than half of the grapheme roots, and consonant 3 and 6 - Generalize the worst.  (Check combination counting kernel <a href=\"https://www.kaggle.com/dipamc77/bengali-dataset-grapheme-combination-counts?scriptVersionId=30348992\">here</a>)</p>\n\n<p>I probed for the scores of some of these classes similar to the method described <a href=\"https://www.kaggle.com/philippsinger/checking-r-c-v-scores\">here</a>. From the gaps, I was sure that a huge shakeup was coming.</p>\n\n<p>After reading the top solutions I'm inspired to pursue ML-based  methods to improve generalization rather than a rule-based method like I describe below; I would still like to share it</p>\n\n<h1>ZoneMix - Splitting graphemes by zones to apply cutmix</h1>\n\n<p>As a native Bengali speaker, I can say that generally graphemes can be split into particular zones based on the combinations. Below is a diagram of the zones.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ffc1a8235909e0be649a656d3bece6637%2FZones.png?generation=1584474729279357&amp;alt=media\" alt=\"\"></p>\n\n<p>Some extra rules have to be made for some combinations, which was kind of a pain to write in code and made me question why I'm even working on such a tedious idea 😭 . I'll share the code surely, just right now its a huge mess.</p>\n\n<p>Each combination takes up some of these zones and cutting by the zone's coordinates based on the character's bounding box can roughly give the components.  Here is an example</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Fdf89e7805d9fb22cb46320f15a51cf63%2Fgoodexample.PNG?generation=1584475233238792&amp;alt=media\" alt=\"\"></p>\n\n<p>Of course - Real-world data doesn't always follow rules and failure cases are common</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F574f7c8d3cd2f39a37347ac672e5c4ab%2Fbadexample.PNG?generation=1584475938169427&amp;alt=media\" alt=\"\"></p>\n\n<p>Once the zone data is extracted, they can be placed on the other images by resizing based on the new zone coordinates. This can produce very realistic images, but also unrealistic ones.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F6af3d256dc7ef3fccc04024763b75102%2FGoodMix.PNG?generation=1584476645803135&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F4e649a79884043537ec5bc181b8e343c%2FBadMix.PNG?generation=1584476950698513&amp;alt=media\" alt=\"\"></p>\n\n<p>When all the setup was done I had about 3 days left to train with this method. And sadly it didn't work directly, I think its because the mistakes and unrealistic images are too much noise for the model to deal with.</p>\n\n<p>I was able to tone down the noise by doing the following tricks - The grapheme root zones are roughly correct most of the time, while the vowel and consonants are way off. I figured that since grapheme root and consonant 3 and 6 are the main classes to tackle, I should focus on those. </p>\n\n<p>Using zonemix I trained a model to only predict the grapheme root, this reduced the noisy gradients from the vowel and consonants and improved the recall as well as unseen recall.</p>\n\n<p>Consonant 3 and 6 have a particular property that they are basically a combination of Consonant 2 and 4 and Consonant 4 and 5 respectively. Dieter also explains this in his <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136129\">solution post</a> and I feel he found a much better way to deal with it. Anyway, so my approach was to use zonemix to place consonant 4 onto images with consonant 2 and 5. Examples below. With these I trained a new model for consonants.</p>\n\n<p>I literally trained both models parallely on two colab sessions on the last day and didn't even have time to tune anything, had to manually stop the runs because I had only 3 hours left to the end of the competition. </p>\n\n<p><strong>Final solution</strong> - Zonemix models for grapheme roots with less combinations (&lt;5), and consonant 3, 6, along the old trained ensemble for the rest (A pretty bad one with public LB only 0.982 and private LB 0.931). But I confidently picked the zonemix kernel knowing it will surely be good on private set 😄 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ff596e6402b241d8ee5cfe0be5646ce43%2FCapture.PNG?generation=1584474005219015&amp;alt=media\" alt=\"\"></p>\n\n<p>Maybe with some better tuning and validation it could have probably done much better. </p>\n\n<h1>Hyperparameters tuning that worked and didn't work for me</h1>\n\n<ol>\n<li>Wasted weeks on cropped and preprocessed images, didn't work at all. I'm still not convinced why.</li>\n<li>Did lots of experiments on resnet34 with 128x128 (just resize) - This was practical for fast experimentation on Google Colab.</li>\n<li>Balanced sampling - The metric is macro average, so helped to do balanced sampling I guess.</li>\n<li>Cutout/Cutmix - Gives better results than mixup.</li>\n<li>Interestingly Mixup + Gridmask gives much better results on Unseen CV</li>\n<li>Finally trained SeResNeXt50 with 128x128 which gave improvement on CV (Though same pathetic results on Unseen CV)</li>\n<li>Well tuned ReduceLRonPlateau converges as much as 1.5x - 2x faster than Cosine Decay. Didn't try other LR schedules.</li>\n</ol>\n\n<h1>Some ideas that I didn't get to try</h1>\n\n<ol>\n<li>CAM FMix - One of the top solution uses CAM Cutmix</li>\n<li>Correct zone segmentation via noisy student training methods on the noisy zones.</li>\n</ol>\n\n<h1>Final thoughts and important lessons learned</h1>\n\n<ol>\n<li><p>From the only other competition I took part in  <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection\">Severstal: Steel Defect Detection</a>  - I had this lesson:\nIf there is a gap in CV and LB, trust your CV more and don't overfit LB.</p></li>\n<li><p>I can change that lesson now - If there is a gap in CV and LB, try to understand why from the discussions, problem statement and the data. Don't dive into coding new ideas trying to reduce the gap without having a proper understanding of what is causing the gap.</p></li>\n<li>Try ML-based approaches for solving the problem, after seeing the top solutions I feel like I understand them very well, but it didn't occur to me to use these ideas. Maybe I'm biased to use rule-based ideas and that needs to change.</li>\n<li>Do lots of EDA from time to time, not just at the beginning of the competition. Seeing the results of trained models changes our intuition about the data.</li>\n<li>Don't ignore external datasets. Of course, working with them needs time and effort, but with correct usage, external data can work wonders.</li>\n<li>Lastly - It is important to keep hacks to a minimum, but at the end of the day, Its a competition, do what you need to do to improve your score, like the awesome metric hacking magic.</li>\n</ol>",
  "messages": [
    {
      "id": "777628",
      "postDate": "03/17/2020 20:31:50",
      "content": "<p>A really big thanks to all the amazing participants for sharing the ideas. @haqishen @hengck23  @bibek777 @machinelp @quandapro and so many other awesome users without whose suggestions I would have been completely lost. And also big thanks to Bengali.AI organizers for this competition about my native language, the lessons learned from this competition will also be applicable to other languages of the Indian subcontinent like Hindi, Odia, Telegu, etc. </p>\n\n<p>I'll be honest, I joined the competition by at nearly the end of January, and for most of the time, I was trying to improve the score just by simple augmentations/model/hyperparameter tuning to improve the score. And having only Google Colab to train doesn't help with that either. Had I continued on that track I surely wouldn't have been in any medal zone. Thanks a lot to my teammate Ram for pointing out the important discusions about unseen graphemes by the organizers. And subsequently Qishen Ha's post on validating on unseen graphemes. </p>\n\n<p>With only 10 days remaining after realizing how poorly the model will do on unseen graphemes. Particularly - Classes with few combinations - Which is more than half of the grapheme roots, and consonant 3 and 6 - Generalize the worst.  (Check combination counting kernel <a href=\"https://www.kaggle.com/dipamc77/bengali-dataset-grapheme-combination-counts?scriptVersionId=30348992\">here</a>)</p>\n\n<p>I probed for the scores of some of these classes similar to the method described <a href=\"https://www.kaggle.com/philippsinger/checking-r-c-v-scores\">here</a>. From the gaps, I was sure that a huge shakeup was coming.</p>\n\n<p>After reading the top solutions I'm inspired to pursue ML-based  methods to improve generalization rather than a rule-based method like I describe below; I would still like to share it</p>\n\n<h1>ZoneMix - Splitting graphemes by zones to apply cutmix</h1>\n\n<p>As a native Bengali speaker, I can say that generally graphemes can be split into particular zones based on the combinations. Below is a diagram of the zones.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ffc1a8235909e0be649a656d3bece6637%2FZones.png?generation=1584474729279357&amp;alt=media\" alt=\"\"></p>\n\n<p>Some extra rules have to be made for some combinations, which was kind of a pain to write in code and made me question why I'm even working on such a tedious idea 😭 . I'll share the code surely, just right now its a huge mess.</p>\n\n<p>Each combination takes up some of these zones and cutting by the zone's coordinates based on the character's bounding box can roughly give the components.  Here is an example</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Fdf89e7805d9fb22cb46320f15a51cf63%2Fgoodexample.PNG?generation=1584475233238792&amp;alt=media\" alt=\"\"></p>\n\n<p>Of course - Real-world data doesn't always follow rules and failure cases are common</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F574f7c8d3cd2f39a37347ac672e5c4ab%2Fbadexample.PNG?generation=1584475938169427&amp;alt=media\" alt=\"\"></p>\n\n<p>Once the zone data is extracted, they can be placed on the other images by resizing based on the new zone coordinates. This can produce very realistic images, but also unrealistic ones.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F6af3d256dc7ef3fccc04024763b75102%2FGoodMix.PNG?generation=1584476645803135&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F4e649a79884043537ec5bc181b8e343c%2FBadMix.PNG?generation=1584476950698513&amp;alt=media\" alt=\"\"></p>\n\n<p>When all the setup was done I had about 3 days left to train with this method. And sadly it didn't work directly, I think its because the mistakes and unrealistic images are too much noise for the model to deal with.</p>\n\n<p>I was able to tone down the noise by doing the following tricks - The grapheme root zones are roughly correct most of the time, while the vowel and consonants are way off. I figured that since grapheme root and consonant 3 and 6 are the main classes to tackle, I should focus on those. </p>\n\n<p>Using zonemix I trained a model to only predict the grapheme root, this reduced the noisy gradients from the vowel and consonants and improved the recall as well as unseen recall.</p>\n\n<p>Consonant 3 and 6 have a particular property that they are basically a combination of Consonant 2 and 4 and Consonant 4 and 5 respectively. Dieter also explains this in his <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136129\">solution post</a> and I feel he found a much better way to deal with it. Anyway, so my approach was to use zonemix to place consonant 4 onto images with consonant 2 and 5. Examples below. With these I trained a new model for consonants.</p>\n\n<p>I literally trained both models parallely on two colab sessions on the last day and didn't even have time to tune anything, had to manually stop the runs because I had only 3 hours left to the end of the competition. </p>\n\n<p><strong>Final solution</strong> - Zonemix models for grapheme roots with less combinations (&lt;5), and consonant 3, 6, along the old trained ensemble for the rest (A pretty bad one with public LB only 0.982 and private LB 0.931). But I confidently picked the zonemix kernel knowing it will surely be good on private set 😄 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ff596e6402b241d8ee5cfe0be5646ce43%2FCapture.PNG?generation=1584474005219015&amp;alt=media\" alt=\"\"></p>\n\n<p>Maybe with some better tuning and validation it could have probably done much better. </p>\n\n<h1>Hyperparameters tuning that worked and didn't work for me</h1>\n\n<ol>\n<li>Wasted weeks on cropped and preprocessed images, didn't work at all. I'm still not convinced why.</li>\n<li>Did lots of experiments on resnet34 with 128x128 (just resize) - This was practical for fast experimentation on Google Colab.</li>\n<li>Balanced sampling - The metric is macro average, so helped to do balanced sampling I guess.</li>\n<li>Cutout/Cutmix - Gives better results than mixup.</li>\n<li>Interestingly Mixup + Gridmask gives much better results on Unseen CV</li>\n<li>Finally trained SeResNeXt50 with 128x128 which gave improvement on CV (Though same pathetic results on Unseen CV)</li>\n<li>Well tuned ReduceLRonPlateau converges as much as 1.5x - 2x faster than Cosine Decay. Didn't try other LR schedules.</li>\n</ol>\n\n<h1>Some ideas that I didn't get to try</h1>\n\n<ol>\n<li>CAM FMix - One of the top solution uses CAM Cutmix</li>\n<li>Correct zone segmentation via noisy student training methods on the noisy zones.</li>\n</ol>\n\n<h1>Final thoughts and important lessons learned</h1>\n\n<ol>\n<li><p>From the only other competition I took part in  <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection\">Severstal: Steel Defect Detection</a>  - I had this lesson:\nIf there is a gap in CV and LB, trust your CV more and don't overfit LB.</p></li>\n<li><p>I can change that lesson now - If there is a gap in CV and LB, try to understand why from the discussions, problem statement and the data. Don't dive into coding new ideas trying to reduce the gap without having a proper understanding of what is causing the gap.</p></li>\n<li>Try ML-based approaches for solving the problem, after seeing the top solutions I feel like I understand them very well, but it didn't occur to me to use these ideas. Maybe I'm biased to use rule-based ideas and that needs to change.</li>\n<li>Do lots of EDA from time to time, not just at the beginning of the competition. Seeing the results of trained models changes our intuition about the data.</li>\n<li>Don't ignore external datasets. Of course, working with them needs time and effort, but with correct usage, external data can work wonders.</li>\n<li>Lastly - It is important to keep hacks to a minimum, but at the end of the day, Its a competition, do what you need to do to improve your score, like the awesome metric hacking magic.</li>\n</ol>",
      "rawMarkdown": "A really big thanks to all the amazing participants for sharing the ideas. @haqishen @hengck23  @bibek777 @machinelp @quandapro and so many other awesome users without whose suggestions I would have been completely lost. And also big thanks to Bengali.AI organizers for this competition about my native language, the lessons learned from this competition will also be applicable to other languages of the Indian subcontinent like Hindi, Odia, Telegu, etc. \n\nI'll be honest, I joined the competition by at nearly the end of January, and for most of the time, I was trying to improve the score just by simple augmentations/model/hyperparameter tuning to improve the score. And having only Google Colab to train doesn't help with that either. Had I continued on that track I surely wouldn't have been in any medal zone. Thanks a lot to my teammate Ram for pointing out the important discusions about unseen graphemes by the organizers. And subsequently Qishen Ha's post on validating on unseen graphemes. \n\nWith only 10 days remaining after realizing how poorly the model will do on unseen graphemes. Particularly - Classes with few combinations - Which is more than half of the grapheme roots, and consonant 3 and 6 - Generalize the worst.  (Check combination counting kernel [here](https://www.kaggle.com/dipamc77/bengali-dataset-grapheme-combination-counts?scriptVersionId=30348992))\n\nI probed for the scores of some of these classes similar to the method described [here](https://www.kaggle.com/philippsinger/checking-r-c-v-scores). From the gaps, I was sure that a huge shakeup was coming.\n\nAfter reading the top solutions I'm inspired to pursue ML-based  methods to improve generalization rather than a rule-based method like I describe below; I would still like to share it\n\n#ZoneMix - Splitting graphemes by zones to apply cutmix\n\nAs a native Bengali speaker, I can say that generally graphemes can be split into particular zones based on the combinations. Below is a diagram of the zones.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ffc1a8235909e0be649a656d3bece6637%2FZones.png?generation=1584474729279357&amp;alt=media)\n\n\nSome extra rules have to be made for some combinations, which was kind of a pain to write in code and made me question why I'm even working on such a tedious idea 😭 . I'll share the code surely, just right now its a huge mess.\n\nEach combination takes up some of these zones and cutting by the zone's coordinates based on the character's bounding box can roughly give the components.  Here is an example\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Fdf89e7805d9fb22cb46320f15a51cf63%2Fgoodexample.PNG?generation=1584475233238792&amp;alt=media)\n\n\nOf course - Real-world data doesn't always follow rules and failure cases are common\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F574f7c8d3cd2f39a37347ac672e5c4ab%2Fbadexample.PNG?generation=1584475938169427&amp;alt=media)\n\n\nOnce the zone data is extracted, they can be placed on the other images by resizing based on the new zone coordinates. This can produce very realistic images, but also unrealistic ones.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F6af3d256dc7ef3fccc04024763b75102%2FGoodMix.PNG?generation=1584476645803135&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F4e649a79884043537ec5bc181b8e343c%2FBadMix.PNG?generation=1584476950698513&amp;alt=media)\n\n\n\nWhen all the setup was done I had about 3 days left to train with this method. And sadly it didn't work directly, I think its because the mistakes and unrealistic images are too much noise for the model to deal with.\n\nI was able to tone down the noise by doing the following tricks - The grapheme root zones are roughly correct most of the time, while the vowel and consonants are way off. I figured that since grapheme root and consonant 3 and 6 are the main classes to tackle, I should focus on those. \n\nUsing zonemix I trained a model to only predict the grapheme root, this reduced the noisy gradients from the vowel and consonants and improved the recall as well as unseen recall.\n\nConsonant 3 and 6 have a particular property that they are basically a combination of Consonant 2 and 4 and Consonant 4 and 5 respectively. Dieter also explains this in his [solution post](https://www.kaggle.com/c/bengaliai-cv19/discussion/136129) and I feel he found a much better way to deal with it. Anyway, so my approach was to use zonemix to place consonant 4 onto images with consonant 2 and 5. Examples below. With these I trained a new model for consonants.\n\nI literally trained both models parallely on two colab sessions on the last day and didn't even have time to tune anything, had to manually stop the runs because I had only 3 hours left to the end of the competition. \n\n**Final solution** - Zonemix models for grapheme roots with less combinations (&lt;5), and consonant 3, 6, along the old trained ensemble for the rest (A pretty bad one with public LB only 0.982 and private LB 0.931). But I confidently picked the zonemix kernel knowing it will surely be good on private set 😄 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ff596e6402b241d8ee5cfe0be5646ce43%2FCapture.PNG?generation=1584474005219015&amp;alt=media)\n\nMaybe with some better tuning and validation it could have probably done much better. \n\n#Hyperparameters tuning that worked and didn't work for me\n\n0. Wasted weeks on cropped and preprocessed images, didn't work at all. I'm still not convinced why.\n1. Did lots of experiments on resnet34 with 128x128 (just resize) - This was practical for fast experimentation on Google Colab.\n2. Balanced sampling - The metric is macro average, so helped to do balanced sampling I guess.\n3. Cutout/Cutmix - Gives better results than mixup.\n4. Interestingly Mixup + Gridmask gives much better results on Unseen CV\n5. Finally trained SeResNeXt50 with 128x128 which gave improvement on CV (Though same pathetic results on Unseen CV)\n6. Well tuned ReduceLRonPlateau converges as much as 1.5x - 2x faster than Cosine Decay. Didn't try other LR schedules.\n\n#Some ideas that I didn't get to try\n1. CAM FMix - One of the top solution uses CAM Cutmix\n2. Correct zone segmentation via noisy student training methods on the noisy zones.\n\n#Final thoughts and important lessons learned\n\n0. From the only other competition I took part in  [Severstal: Steel Defect Detection](https://www.kaggle.com/c/severstal-steel-defect-detection)  - I had this lesson:\nIf there is a gap in CV and LB, trust your CV more and don't overfit LB.\n\n1. I can change that lesson now - If there is a gap in CV and LB, try to understand why from the discussions, problem statement and the data. Don't dive into coding new ideas trying to reduce the gap without having a proper understanding of what is causing the gap.\n2. Try ML-based approaches for solving the problem, after seeing the top solutions I feel like I understand them very well, but it didn't occur to me to use these ideas. Maybe I'm biased to use rule-based ideas and that needs to change.\n3. Do lots of EDA from time to time, not just at the beginning of the competition. Seeing the results of trained models changes our intuition about the data.\n4. Don't ignore external datasets. Of course, working with them needs time and effort, but with correct usage, external data can work wonders.\n5. Lastly - It is important to keep hacks to a minimum, but at the end of the day, Its a competition, do what you need to do to improve your score, like the awesome metric hacking magic.",
      "votes": null
    },
    {
      "id": "777646",
      "postDate": "03/17/2020 20:45:19",
      "content": "<p>Interesting&gt; Did you crop resize your images to make sure your zones would be relevant?  You say cropping does not work, but did you keep image ratio?</p>",
      "rawMarkdown": "Interesting&gt; Did you crop resize your images to make sure your zones would be relevant?  You say cropping does not work, but did you keep image ratio?",
      "votes": null
    },
    {
      "id": "777681",
      "postDate": "03/17/2020 21:12:18",
      "content": "<p>No, I find the bounding boxes and reverse calculate the zones coordinates based on the bounding box coordinates. When taking from one zone and placing onto the other, there is resizing, but for most cases maintaining aspect ratio of vowels and consonants should not be very important.</p>\n\n<p>During the first few weeks, I had tried publicly shared cropping methods that were maintaining aspect ratio. But I realized the bounding boxes were often wrong due to random lines on the sides that didn't go with simple rules. Finally, I wrote my own bounding box finding logic based on contours and that ignores contours on the sides, that was  more consistent. Though some images still have small ink spots and bounding boxes do come out wrong for those.</p>",
      "rawMarkdown": "No, I find the bounding boxes and reverse calculate the zones coordinates based on the bounding box coordinates. When taking from one zone and placing onto the other, there is resizing, but for most cases maintaining aspect ratio of vowels and consonants should not be very important.\n\nDuring the first few weeks, I had tried publicly shared cropping methods that were maintaining aspect ratio. But I realized the bounding boxes were often wrong due to random lines on the sides that didn't go with simple rules. Finally, I wrote my own bounding box finding logic based on contours and that ignores contours on the sides, that was  more consistent. Though some images still have small ink spots and bounding boxes do come out wrong for those.",
      "votes": null
    },
    {
      "id": "778065",
      "postDate": "03/18/2020 05:52:10",
      "content": "<p>This zone approach seems similar to CAM-cutmix. They are all trying to reduce the model dependency on the whole graphme. Is my understanding correct?</p>",
      "rawMarkdown": "This zone approach seems similar to CAM-cutmix. They are all trying to reduce the model dependency on the whole graphme. Is my understanding correct?",
      "votes": null
    },
    {
      "id": "778069",
      "postDate": "03/18/2020 05:55:41",
      "content": "<p>Also, I was wondering what exactly is the ML-solution that you are referring to? Thanks!</p>",
      "rawMarkdown": "Also, I was wondering what exactly is the ML-solution that you are referring to? Thanks!",
      "votes": null
    },
    {
      "id": "778260",
      "postDate": "03/18/2020 09:31:04",
      "content": "<p>None of the publicly shared cropping maintain image ratio, at least none of the ones I saw.</p>",
      "rawMarkdown": "None of the publicly shared cropping maintain image ratio, at least none of the ones I saw.",
      "votes": null
    },
    {
      "id": "778268",
      "postDate": "03/18/2020 09:39:35",
      "content": "<p>This kernel is a very popular one many people were using towards the beginning.\n<a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a>\nIt finds the bounding box and then pads the cropped image such that the aspect ratio of the cropped image is maintained.</p>\n\n<p>Is this what you mean by image ratio?</p>",
      "rawMarkdown": "This kernel is a very popular one many people were using towards the beginning.\nhttps://www.kaggle.com/iafoss/image-preprocessing-128x128\nIt finds the bounding box and then pads the cropped image such that the aspect ratio of the cropped image is maintained.\n\nIs this what you mean by image ratio?",
      "votes": null
    },
    {
      "id": "778751",
      "postDate": "03/18/2020 17:41:33",
      "content": "<p>Broadly speaking I mean understanding and how deep learning works and making creative solutions to exploit it. </p>\n\n<p>Two examples of this for me were the 1st and 5th place solutions. I understand cyclegan well but didn't think of using it to transfer handwritten to printed font and use a classifier on that. </p>\n\n<p>5th place one really felt magically creative to me. Because I too improved my consonant 3 and 6 by using zonemix. But the multilabel refactoring of the problem is so much more natural. </p>\n\n<p>Here I tried to improve the data with rule based thinking, I want to be more creative like those solutions and others. </p>",
      "rawMarkdown": "Broadly speaking I mean understanding and how deep learning works and making creative solutions to exploit it. \n\nTwo examples of this for me were the 1st and 5th place solutions. I understand cyclegan well but didn't think of using it to transfer handwritten to printed font and use a classifier on that. \n\n5th place one really felt magically creative to me. Because I too improved my consonant 3 and 6 by using zonemix. But the multilabel refactoring of the problem is so much more natural. \n\nHere I tried to improve the data with rule based thinking, I want to be more creative like those solutions and others.",
      "votes": null
    },
    {
      "id": "778755",
      "postDate": "03/18/2020 17:45:44",
      "content": "<p>Yes that is true, and frankly I believe it might be even better than CAM cutmix because of the CAM will most likely be wrong for classes with less combinations like I described. I too wanted to generate unseen graphemes with this method but it needed some polishing and I didn't get the time.</p>\n\n<p>Its an excuse on my part but my basic ensemble on seen graphemes wasn't good enough to begin with and that overshadowed this method I developed. Have to improve my basic DL skills for that.</p>",
      "rawMarkdown": "Yes that is true, and frankly I believe it might be even better than CAM cutmix because of the CAM will most likely be wrong for classes with less combinations like I described. I too wanted to generate unseen graphemes with this method but it needed some polishing and I didn't get the time.\n\nIts an excuse on my part but my basic ensemble on seen graphemes wasn't good enough to begin with and that overshadowed this method I developed. Have to improve my basic DL skills for that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 777646,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/17/2020 20:45:19",
      "content": "<p>Interesting&gt; Did you crop resize your images to make sure your zones would be relevant?  You say cropping does not work, but did you keep image ratio?</p>",
      "votes": null,
      "replies": [
        {
          "id": 777681,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "03/17/2020 21:12:18",
          "content": "<p>No, I find the bounding boxes and reverse calculate the zones coordinates based on the bounding box coordinates. When taking from one zone and placing onto the other, there is resizing, but for most cases maintaining aspect ratio of vowels and consonants should not be very important.</p>\n\n<p>During the first few weeks, I had tried publicly shared cropping methods that were maintaining aspect ratio. But I realized the bounding boxes were often wrong due to random lines on the sides that didn't go with simple rules. Finally, I wrote my own bounding box finding logic based on contours and that ignores contours on the sides, that was  more consistent. Though some images still have small ink spots and bounding boxes do come out wrong for those.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 778260,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/18/2020 09:31:04",
          "content": "<p>None of the publicly shared cropping maintain image ratio, at least none of the ones I saw.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 778268,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "03/18/2020 09:39:35",
          "content": "<p>This kernel is a very popular one many people were using towards the beginning.\n<a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a>\nIt finds the bounding box and then pads the cropped image such that the aspect ratio of the cropped image is maintained.</p>\n\n<p>Is this what you mean by image ratio?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 778065,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "03/18/2020 05:52:10",
      "content": "<p>This zone approach seems similar to CAM-cutmix. They are all trying to reduce the model dependency on the whole graphme. Is my understanding correct?</p>",
      "votes": null,
      "replies": [
        {
          "id": 778755,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "03/18/2020 17:45:44",
          "content": "<p>Yes that is true, and frankly I believe it might be even better than CAM cutmix because of the CAM will most likely be wrong for classes with less combinations like I described. I too wanted to generate unseen graphemes with this method but it needed some polishing and I didn't get the time.</p>\n\n<p>Its an excuse on my part but my basic ensemble on seen graphemes wasn't good enough to begin with and that overshadowed this method I developed. Have to improve my basic DL skills for that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 778069,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "03/18/2020 05:55:41",
      "content": "<p>Also, I was wondering what exactly is the ML-solution that you are referring to? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 778751,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "03/18/2020 17:41:33",
          "content": "<p>Broadly speaking I mean understanding and how deep learning works and making creative solutions to exploit it. </p>\n\n<p>Two examples of this for me were the 1st and 5th place solutions. I understand cyclegan well but didn't think of using it to transfer handwritten to printed font and use a classifier on that. </p>\n\n<p>5th place one really felt magically creative to me. Because I too improved my consonant 3 and 6 by using zonemix. But the multilabel refactoring of the problem is so much more natural. </p>\n\n<p>Here I tried to improve the data with rule based thinking, I want to be more creative like those solutions and others. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "777628": "A really big thanks to all the amazing participants for sharing the ideas. @haqishen @hengck23  @bibek777 @machinelp @quandapro and so many other awesome users without whose suggestions I would have been completely lost. And also big thanks to Bengali.AI organizers for this competition about my native language, the lessons learned from this competition will also be applicable to other languages of the Indian subcontinent like Hindi, Odia, Telegu, etc. \n\nI'll be honest, I joined the competition by at nearly the end of January, and for most of the time, I was trying to improve the score just by simple augmentations/model/hyperparameter tuning to improve the score. And having only Google Colab to train doesn't help with that either. Had I continued on that track I surely wouldn't have been in any medal zone. Thanks a lot to my teammate Ram for pointing out the important discusions about unseen graphemes by the organizers. And subsequently Qishen Ha's post on validating on unseen graphemes. \n\nWith only 10 days remaining after realizing how poorly the model will do on unseen graphemes. Particularly - Classes with few combinations - Which is more than half of the grapheme roots, and consonant 3 and 6 - Generalize the worst.  (Check combination counting kernel [here](https://www.kaggle.com/dipamc77/bengali-dataset-grapheme-combination-counts?scriptVersionId=30348992))\n\nI probed for the scores of some of these classes similar to the method described [here](https://www.kaggle.com/philippsinger/checking-r-c-v-scores). From the gaps, I was sure that a huge shakeup was coming.\n\nAfter reading the top solutions I'm inspired to pursue ML-based  methods to improve generalization rather than a rule-based method like I describe below; I would still like to share it\n\n#ZoneMix - Splitting graphemes by zones to apply cutmix\n\nAs a native Bengali speaker, I can say that generally graphemes can be split into particular zones based on the combinations. Below is a diagram of the zones.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ffc1a8235909e0be649a656d3bece6637%2FZones.png?generation=1584474729279357&amp;alt=media)\n\n\nSome extra rules have to be made for some combinations, which was kind of a pain to write in code and made me question why I'm even working on such a tedious idea 😭 . I'll share the code surely, just right now its a huge mess.\n\nEach combination takes up some of these zones and cutting by the zone's coordinates based on the character's bounding box can roughly give the components.  Here is an example\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Fdf89e7805d9fb22cb46320f15a51cf63%2Fgoodexample.PNG?generation=1584475233238792&amp;alt=media)\n\n\nOf course - Real-world data doesn't always follow rules and failure cases are common\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F574f7c8d3cd2f39a37347ac672e5c4ab%2Fbadexample.PNG?generation=1584475938169427&amp;alt=media)\n\n\nOnce the zone data is extracted, they can be placed on the other images by resizing based on the new zone coordinates. This can produce very realistic images, but also unrealistic ones.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F6af3d256dc7ef3fccc04024763b75102%2FGoodMix.PNG?generation=1584476645803135&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2F4e649a79884043537ec5bc181b8e343c%2FBadMix.PNG?generation=1584476950698513&amp;alt=media)\n\n\n\nWhen all the setup was done I had about 3 days left to train with this method. And sadly it didn't work directly, I think its because the mistakes and unrealistic images are too much noise for the model to deal with.\n\nI was able to tone down the noise by doing the following tricks - The grapheme root zones are roughly correct most of the time, while the vowel and consonants are way off. I figured that since grapheme root and consonant 3 and 6 are the main classes to tackle, I should focus on those. \n\nUsing zonemix I trained a model to only predict the grapheme root, this reduced the noisy gradients from the vowel and consonants and improved the recall as well as unseen recall.\n\nConsonant 3 and 6 have a particular property that they are basically a combination of Consonant 2 and 4 and Consonant 4 and 5 respectively. Dieter also explains this in his [solution post](https://www.kaggle.com/c/bengaliai-cv19/discussion/136129) and I feel he found a much better way to deal with it. Anyway, so my approach was to use zonemix to place consonant 4 onto images with consonant 2 and 5. Examples below. With these I trained a new model for consonants.\n\nI literally trained both models parallely on two colab sessions on the last day and didn't even have time to tune anything, had to manually stop the runs because I had only 3 hours left to the end of the competition. \n\n**Final solution** - Zonemix models for grapheme roots with less combinations (&lt;5), and consonant 3, 6, along the old trained ensemble for the rest (A pretty bad one with public LB only 0.982 and private LB 0.931). But I confidently picked the zonemix kernel knowing it will surely be good on private set 😄 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3468635%2Ff596e6402b241d8ee5cfe0be5646ce43%2FCapture.PNG?generation=1584474005219015&amp;alt=media)\n\nMaybe with some better tuning and validation it could have probably done much better. \n\n#Hyperparameters tuning that worked and didn't work for me\n\n0. Wasted weeks on cropped and preprocessed images, didn't work at all. I'm still not convinced why.\n1. Did lots of experiments on resnet34 with 128x128 (just resize) - This was practical for fast experimentation on Google Colab.\n2. Balanced sampling - The metric is macro average, so helped to do balanced sampling I guess.\n3. Cutout/Cutmix - Gives better results than mixup.\n4. Interestingly Mixup + Gridmask gives much better results on Unseen CV\n5. Finally trained SeResNeXt50 with 128x128 which gave improvement on CV (Though same pathetic results on Unseen CV)\n6. Well tuned ReduceLRonPlateau converges as much as 1.5x - 2x faster than Cosine Decay. Didn't try other LR schedules.\n\n#Some ideas that I didn't get to try\n1. CAM FMix - One of the top solution uses CAM Cutmix\n2. Correct zone segmentation via noisy student training methods on the noisy zones.\n\n#Final thoughts and important lessons learned\n\n0. From the only other competition I took part in  [Severstal: Steel Defect Detection](https://www.kaggle.com/c/severstal-steel-defect-detection)  - I had this lesson:\nIf there is a gap in CV and LB, trust your CV more and don't overfit LB.\n\n1. I can change that lesson now - If there is a gap in CV and LB, try to understand why from the discussions, problem statement and the data. Don't dive into coding new ideas trying to reduce the gap without having a proper understanding of what is causing the gap.\n2. Try ML-based approaches for solving the problem, after seeing the top solutions I feel like I understand them very well, but it didn't occur to me to use these ideas. Maybe I'm biased to use rule-based ideas and that needs to change.\n3. Do lots of EDA from time to time, not just at the beginning of the competition. Seeing the results of trained models changes our intuition about the data.\n4. Don't ignore external datasets. Of course, working with them needs time and effort, but with correct usage, external data can work wonders.\n5. Lastly - It is important to keep hacks to a minimum, but at the end of the day, Its a competition, do what you need to do to improve your score, like the awesome metric hacking magic.",
    "777646": "Interesting&gt; Did you crop resize your images to make sure your zones would be relevant?  You say cropping does not work, but did you keep image ratio?",
    "777681": "No, I find the bounding boxes and reverse calculate the zones coordinates based on the bounding box coordinates. When taking from one zone and placing onto the other, there is resizing, but for most cases maintaining aspect ratio of vowels and consonants should not be very important.\n\nDuring the first few weeks, I had tried publicly shared cropping methods that were maintaining aspect ratio. But I realized the bounding boxes were often wrong due to random lines on the sides that didn't go with simple rules. Finally, I wrote my own bounding box finding logic based on contours and that ignores contours on the sides, that was  more consistent. Though some images still have small ink spots and bounding boxes do come out wrong for those.",
    "778065": "This zone approach seems similar to CAM-cutmix. They are all trying to reduce the model dependency on the whole graphme. Is my understanding correct?",
    "778069": "Also, I was wondering what exactly is the ML-solution that you are referring to? Thanks!",
    "778260": "None of the publicly shared cropping maintain image ratio, at least none of the ones I saw.",
    "778268": "This kernel is a very popular one many people were using towards the beginning.\nhttps://www.kaggle.com/iafoss/image-preprocessing-128x128\nIt finds the bounding box and then pads the cropped image such that the aspect ratio of the cropped image is maintained.\n\nIs this what you mean by image ratio?",
    "778751": "Broadly speaking I mean understanding and how deep learning works and making creative solutions to exploit it. \n\nTwo examples of this for me were the 1st and 5th place solutions. I understand cyclegan well but didn't think of using it to transfer handwritten to printed font and use a classifier on that. \n\n5th place one really felt magically creative to me. Because I too improved my consonant 3 and 6 by using zonemix. But the multilabel refactoring of the problem is so much more natural. \n\nHere I tried to improve the data with rule based thinking, I want to be more creative like those solutions and others.",
    "778755": "Yes that is true, and frankly I believe it might be even better than CAM cutmix because of the CAM will most likely be wrong for classes with less combinations like I described. I too wanted to generate unseen graphemes with this method but it needed some polishing and I didn't get the time.\n\nIts an excuse on my part but my basic ensemble on seen graphemes wasn't good enough to begin with and that overshadowed this method I developed. Have to improve my basic DL skills for that."
  },
  "source": "meta"
}