{
  "id": 135960,
  "title": "7th Place Data-based Generalization Solution",
  "url": "/competitions/bengaliai-cv19/discussion/135960",
  "author_name": "Nicholas Lyu",
  "post_date": "2020-03-16T23:59:44.983000",
  "votes": 95,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. It is my first all-in kaggle competition, and needless to say none of this would have been possible without our awesome team Igor, Habib, Rinat, and Youhan. <a href=\"/cateek\">@cateek</a> <a href=\"/drhabib\">@drhabib</a> <a href=\"/trytolose\">@trytolose</a> <a href=\"/youhanlee\">@youhanlee</a> </p>\n\n<h3>Analysis: This competition has a distinctly stratified LB:</h3>\n\n<ul>\n<li>With a good pipeline, single model with baseline augmentations LB ~.97. </li>\n<li><strong>Removing crops, training on full-resolution, adding cutmix / cutout, and training enough</strong> should get model up to LB .985+, and with some tweaking ~.989+. These are well-covered in the discussions posts.</li>\n<li>Our best single model is from Rinat, PNASNet-5-Large from <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">Cadene's repository</a> which scores LB 0.9900, CV .9985. I believe most of the participants LB .9850-.9910 are using ensemble of models scoring around this range.</li>\n</ul>\n\n<p>At this stage (.9900-.9905) we found it very hard to further improve our LB. As <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134601\">this post</a> points out, higher CV does not necessary mean better LB anymore, so we were stuck for a while (a long while...).</p>\n\n<h3>Improving Generalization</h3>\n\n<p>With the CV &amp; LB relationship it is easy to see that the problem is that our model <strong>could not generalize to unseen graphemes</strong> (by grapheme I mean triplet combination of grapheme root, vowel and consonant diacritics). There are only ~1290 graphemes in training set, while there are 12936 possible combinations. Our models might be biased to output seen triplets; for whatever reason, as Qishen Ha's wonderful <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">post</a> pointed out, our models perform badly on these graphemes.</p>\n\n<p>In hindsight, it's no magic that the most convenient way to improve generalization is to <strong>add more data</strong>. When investigating the grapheme representations, I found that <strong>graphemes are actually encoded as sequence of unicode characters</strong>. The roots and diacritics also have their corresponding sequences. The most magical part is that <strong>the unicode sequence for the grapheme is a combination of the sequences of its roots</strong>. Best exemplified using this image, where the four rows are <em>grapheme, grapheme root, vowel, and conso</em>.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F82988244ae7fb6d4ca47f6c6094a3116%2Fimage_480.png?generation=1584399876994128&amp;alt=media\" alt=\"\"></p>\n\n<p>This observation works (with some minor exceptions) on all graphemes in the training set, and using a very basic algorithm I was able to generate correct labels for all 1295 graphemes in training set except for 2. Following Guanshuo Xu's post pointing to <a href=\"https://github.com/MinhasKamal/BengaliDictionary\">this repository</a>, we are able to access a comprehensive list of grapheme combinations. Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics. Our LB boost on the final day is from better generalization on these graphemes, which partially overlaps with train &amp; test set.</p>\n\n<p>With these graphemes, we rendered them using various fonts and obtained a cleanly-labeled synthetic dataset with ~47K images spanning these characters which can be found <a href=\"https://www.kaggle.com/roguekk007/bengaliai-synthetic-magic\">here</a>, all the ingredients for creating this synthetic dataset has been publicly available. Our hope was that by seeing synthetic images, our models should be able to at least learn the topological (if not stylistic) features of the grapheme. And it turned out they do. Here is a sample of synthetic image in our dataset.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Ff582156ca8f174802dbbc28a18fa1562%2F1_EkSharifa.png?generation=1584402247696903&amp;alt=media\" alt=\"\"></p>\n\n<p>At this point we have only 24 hours until the end of the competition so just finetuned our previous models with the synthetic data added to both original training &amp; validation, thus resulting in the exciting quantum jump on the last day😊 </p>\n\n<h3>Some Details</h3>\n\n<p>My teammates have wonderful pipelines, here are some approaches which have turned out useful:\n- Train on 128x128 data for ~100 epochs, then finetune on 224x224 data. It definitely yields faster training and maybe better generalization (used in Rinat's pipeline)\n- Anneal augmentations except for cutmix gradually to 0 at late stages of training (used in Habib's and Igor's pipeline)\n- A wonderful new architecture called MixNet, we found that it nicely complements PNAS5 for ensembling and performs pretty well.\n- <strong>Loss:</strong> Plain cross entropy. Tried arface (did not work), focal (did not work), normalized softmax+label smoothing (looked promising but ditched along with my pipeline)\n- <strong>Scheduler:</strong> ReduceLRonPleateau for Rinat's pipeline, and hand-monitored dropLR for Habib's and Igor's pipelines.</p>\n\n<h3>Finally</h3>\n\n<p>It has always been my dream to do a gold-medal solution write-up before 17th birthday😃 . In the end, more than happy to see hundreds of experiments and weeks of stress translate successfully into solid rank. With more time than 24 hours, we would have: \n1. Searched for a more comprehensive list of graphemes\n2. Created more synthetic data with more fonts\n3. Somehow make synthetic data more similar to handwritten data; we are using dimming max_val to 230 and gaussian-blur, but can definitely be improved by maybe CycleGAN\n4. Make flipping work as pointed out <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126761\">here</a>. It's really ironic I did not make my own idea work when Qishen Ha did; guess this is the difference🤓 \n5. Try more model-regularization methods, ShakeDrop was on the list and looked promising but was preceded by other priorities.</p>\n\n<p>Overall, this has a very exciting and memorable puzzle-solving competition. Again kudos to the team and:\nLove Kaggling</p>",
  "messages": [
    {
      "id": 775683,
      "postDate": "2020-03-16T23:59:44.983Z",
      "content": "<p>Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. It is my first all-in kaggle competition, and needless to say none of this would have been possible without our awesome team Igor, Habib, Rinat, and Youhan. <a href=\"/cateek\">@cateek</a> <a href=\"/drhabib\">@drhabib</a> <a href=\"/trytolose\">@trytolose</a> <a href=\"/youhanlee\">@youhanlee</a> </p>\n\n<h3>Analysis: This competition has a distinctly stratified LB:</h3>\n\n<ul>\n<li>With a good pipeline, single model with baseline augmentations LB ~.97. </li>\n<li><strong>Removing crops, training on full-resolution, adding cutmix / cutout, and training enough</strong> should get model up to LB .985+, and with some tweaking ~.989+. These are well-covered in the discussions posts.</li>\n<li>Our best single model is from Rinat, PNASNet-5-Large from <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">Cadene's repository</a> which scores LB 0.9900, CV .9985. I believe most of the participants LB .9850-.9910 are using ensemble of models scoring around this range.</li>\n</ul>\n\n<p>At this stage (.9900-.9905) we found it very hard to further improve our LB. As <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134601\">this post</a> points out, higher CV does not necessary mean better LB anymore, so we were stuck for a while (a long while...).</p>\n\n<h3>Improving Generalization</h3>\n\n<p>With the CV &amp; LB relationship it is easy to see that the problem is that our model <strong>could not generalize to unseen graphemes</strong> (by grapheme I mean triplet combination of grapheme root, vowel and consonant diacritics). There are only ~1290 graphemes in training set, while there are 12936 possible combinations. Our models might be biased to output seen triplets; for whatever reason, as Qishen Ha's wonderful <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">post</a> pointed out, our models perform badly on these graphemes.</p>\n\n<p>In hindsight, it's no magic that the most convenient way to improve generalization is to <strong>add more data</strong>. When investigating the grapheme representations, I found that <strong>graphemes are actually encoded as sequence of unicode characters</strong>. The roots and diacritics also have their corresponding sequences. The most magical part is that <strong>the unicode sequence for the grapheme is a combination of the sequences of its roots</strong>. Best exemplified using this image, where the four rows are <em>grapheme, grapheme root, vowel, and conso</em>.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F82988244ae7fb6d4ca47f6c6094a3116%2Fimage_480.png?generation=1584399876994128&amp;alt=media\" alt=\"\"></p>\n\n<p>This observation works (with some minor exceptions) on all graphemes in the training set, and using a very basic algorithm I was able to generate correct labels for all 1295 graphemes in training set except for 2. Following Guanshuo Xu's post pointing to <a href=\"https://github.com/MinhasKamal/BengaliDictionary\">this repository</a>, we are able to access a comprehensive list of grapheme combinations. Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics. Our LB boost on the final day is from better generalization on these graphemes, which partially overlaps with train &amp; test set.</p>\n\n<p>With these graphemes, we rendered them using various fonts and obtained a cleanly-labeled synthetic dataset with ~47K images spanning these characters which can be found <a href=\"https://www.kaggle.com/roguekk007/bengaliai-synthetic-magic\">here</a>, all the ingredients for creating this synthetic dataset has been publicly available. Our hope was that by seeing synthetic images, our models should be able to at least learn the topological (if not stylistic) features of the grapheme. And it turned out they do. Here is a sample of synthetic image in our dataset.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Ff582156ca8f174802dbbc28a18fa1562%2F1_EkSharifa.png?generation=1584402247696903&amp;alt=media\" alt=\"\"></p>\n\n<p>At this point we have only 24 hours until the end of the competition so just finetuned our previous models with the synthetic data added to both original training &amp; validation, thus resulting in the exciting quantum jump on the last day😊 </p>\n\n<h3>Some Details</h3>\n\n<p>My teammates have wonderful pipelines, here are some approaches which have turned out useful:\n- Train on 128x128 data for ~100 epochs, then finetune on 224x224 data. It definitely yields faster training and maybe better generalization (used in Rinat's pipeline)\n- Anneal augmentations except for cutmix gradually to 0 at late stages of training (used in Habib's and Igor's pipeline)\n- A wonderful new architecture called MixNet, we found that it nicely complements PNAS5 for ensembling and performs pretty well.\n- <strong>Loss:</strong> Plain cross entropy. Tried arface (did not work), focal (did not work), normalized softmax+label smoothing (looked promising but ditched along with my pipeline)\n- <strong>Scheduler:</strong> ReduceLRonPleateau for Rinat's pipeline, and hand-monitored dropLR for Habib's and Igor's pipelines.</p>\n\n<h3>Finally</h3>\n\n<p>It has always been my dream to do a gold-medal solution write-up before 17th birthday😃 . In the end, more than happy to see hundreds of experiments and weeks of stress translate successfully into solid rank. With more time than 24 hours, we would have: \n1. Searched for a more comprehensive list of graphemes\n2. Created more synthetic data with more fonts\n3. Somehow make synthetic data more similar to handwritten data; we are using dimming max_val to 230 and gaussian-blur, but can definitely be improved by maybe CycleGAN\n4. Make flipping work as pointed out <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126761\">here</a>. It's really ironic I did not make my own idea work when Qishen Ha did; guess this is the difference🤓 \n5. Try more model-regularization methods, ShakeDrop was on the list and looked promising but was preceded by other priorities.</p>\n\n<p>Overall, this has a very exciting and memorable puzzle-solving competition. Again kudos to the team and:\nLove Kaggling</p>",
      "rawMarkdown": "Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. It is my first all-in kaggle competition, and needless to say none of this would have been possible without our awesome team Igor, Habib, Rinat, and Youhan. @cateek @drhabib @trytolose @youhanlee \n\n### Analysis: This competition has a distinctly stratified LB:\n- With a good pipeline, single model with baseline augmentations LB ~.97. \n- **Removing crops, training on full-resolution, adding cutmix / cutout, and training enough** should get model up to LB .985+, and with some tweaking ~.989+. These are well-covered in the discussions posts.\n- Our best single model is from Rinat, PNASNet-5-Large from [Cadene's repository](https://github.com/Cadene/pretrained-models.pytorch) which scores LB 0.9900, CV .9985. I believe most of the participants LB .9850-.9910 are using ensemble of models scoring around this range.\n\nAt this stage (.9900-.9905) we found it very hard to further improve our LB. As [this post](https://www.kaggle.com/c/bengaliai-cv19/discussion/134601) points out, higher CV does not necessary mean better LB anymore, so we were stuck for a while (a long while...).\n\n### Improving Generalization\nWith the CV &amp; LB relationship it is easy to see that the problem is that our model **could not generalize to unseen graphemes** (by grapheme I mean triplet combination of grapheme root, vowel and consonant diacritics). There are only ~1290 graphemes in training set, while there are 12936 possible combinations. Our models might be biased to output seen triplets; for whatever reason, as Qishen Ha's wonderful [post](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434) pointed out, our models perform badly on these graphemes.\n\nIn hindsight, it's no magic that the most convenient way to improve generalization is to **add more data**. When investigating the grapheme representations, I found that **graphemes are actually encoded as sequence of unicode characters**. The roots and diacritics also have their corresponding sequences. The most magical part is that **the unicode sequence for the grapheme is a combination of the sequences of its roots**. Best exemplified using this image, where the four rows are *grapheme, grapheme root, vowel, and conso*.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F82988244ae7fb6d4ca47f6c6094a3116%2Fimage_480.png?generation=1584399876994128&amp;alt=media)\n\nThis observation works (with some minor exceptions) on all graphemes in the training set, and using a very basic algorithm I was able to generate correct labels for all 1295 graphemes in training set except for 2. Following Guanshuo Xu's post pointing to [this repository](https://github.com/MinhasKamal/BengaliDictionary), we are able to access a comprehensive list of grapheme combinations. Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics. Our LB boost on the final day is from better generalization on these graphemes, which partially overlaps with train &amp; test set.\n\nWith these graphemes, we rendered them using various fonts and obtained a cleanly-labeled synthetic dataset with ~47K images spanning these characters which can be found [here](https://www.kaggle.com/roguekk007/bengaliai-synthetic-magic), all the ingredients for creating this synthetic dataset has been publicly available. Our hope was that by seeing synthetic images, our models should be able to at least learn the topological (if not stylistic) features of the grapheme. And it turned out they do. Here is a sample of synthetic image in our dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Ff582156ca8f174802dbbc28a18fa1562%2F1_EkSharifa.png?generation=1584402247696903&amp;alt=media)\n\nAt this point we have only 24 hours until the end of the competition so just finetuned our previous models with the synthetic data added to both original training &amp; validation, thus resulting in the exciting quantum jump on the last day😊 \n\n### Some Details\nMy teammates have wonderful pipelines, here are some approaches which have turned out useful:\n- Train on 128x128 data for ~100 epochs, then finetune on 224x224 data. It definitely yields faster training and maybe better generalization (used in Rinat's pipeline)\n- Anneal augmentations except for cutmix gradually to 0 at late stages of training (used in Habib's and Igor's pipeline)\n- A wonderful new architecture called MixNet, we found that it nicely complements PNAS5 for ensembling and performs pretty well.\n- **Loss:** Plain cross entropy. Tried arface (did not work), focal (did not work), normalized softmax+label smoothing (looked promising but ditched along with my pipeline)\n- **Scheduler:** ReduceLRonPleateau for Rinat's pipeline, and hand-monitored dropLR for Habib's and Igor's pipelines.\n\n### Finally\nIt has always been my dream to do a gold-medal solution write-up before 17th birthday😃 . In the end, more than happy to see hundreds of experiments and weeks of stress translate successfully into solid rank. With more time than 24 hours, we would have: \n1. Searched for a more comprehensive list of graphemes\n2. Created more synthetic data with more fonts\n3. Somehow make synthetic data more similar to handwritten data; we are using dimming max_val to 230 and gaussian-blur, but can definitely be improved by maybe CycleGAN\n4. Make flipping work as pointed out [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/126761). It's really ironic I did not make my own idea work when Qishen Ha did; guess this is the difference🤓 \n5. Try more model-regularization methods, ShakeDrop was on the list and looked promising but was preceded by other priorities.\n\nOverall, this has a very exciting and memorable puzzle-solving competition. Again kudos to the team and:\nLove Kaggling",
      "votes": 94
    },
    {
      "id": 780442,
      "postDate": "2020-03-20T09:15:46.110Z",
      "content": "<p>Congratulation <a href=\"/roguekk007\">@roguekk007</a>. <br>\nhope to work with you one day. Anyway, good job 💙 🙂 </p>",
      "rawMarkdown": "Congratulation @roguekk007.  \nhope to work with you one day. Anyway, good job 💙 🙂 ",
      "votes": 1,
      "replies": [
        {
          "id": 783554,
          "postDate": "2020-03-23T13:20:41.050Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> hi, just curious to know, how you build <a href=\"https://www.kaggle.com/roguekk007/bengaliai-synthetic-magic\">this dataset</a> from a font?  Would you please share the implemented code? Thanks.</p>",
          "rawMarkdown": "@roguekk007 hi, just curious to know, how you build [this dataset](https://www.kaggle.com/roguekk007/bengaliai-synthetic-magic) from a font?  Would you please share the implemented code? Thanks."
        },
        {
          "id": 783750,
          "postDate": "2020-03-23T16:15:47.183Z",
          "content": "<p>Hi! \nyou can refer to this: \n<a href=\"https://www.kaggle.com/drhabib/generating-more-training-data\">https://www.kaggle.com/drhabib/generating-more-training-data</a>\nbut correct displaying using this discussion \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127938#775496\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127938#775496</a></p>\n\n<p>Hope it helps =) </p>",
          "rawMarkdown": "Hi! \nyou can refer to this: \nhttps://www.kaggle.com/drhabib/generating-more-training-data\nbut correct displaying using this discussion \nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/127938#775496\n\nHope it helps =) "
        },
        {
          "id": 784470,
          "postDate": "2020-03-24T08:31:53.803Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks a lot 💙 </p>",
          "rawMarkdown": "@drhabib thanks a lot 💙 "
        }
      ]
    },
    {
      "id": 777773,
      "postDate": "2020-03-17T23:29:53.953Z",
      "content": "<p>Very impressive idea! I never thought of generating synthetic fonts like this way :)</p>\n\n<p>I have 2 questions - </p>\n\n<p>1) for your synthetic images, did you generate extra graphemes from the given parts of 168 (r) + 11 (v) + 7 (c) , or you decompose those r v c from those extra graphemes fonts from <a href=\"https://github.com/MinhasKamal/BengaliDictionary\">https://github.com/MinhasKamal/BengaliDictionary</a>? </p>\n\n<p>2) You mentioned <code>Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics</code> - may I know how you decomposed those graphemes that are not in our 1295 graphemes? </p>\n\n<p>Thanks and congrats on your top ranking</p>",
      "rawMarkdown": "Very impressive idea! I never thought of generating synthetic fonts like this way :)\n\nI have 2 questions - \n\n1) for your synthetic images, did you generate extra graphemes from the given parts of 168 (r) + 11 (v) + 7 (c) , or you decompose those r v c from those extra graphemes fonts from https://github.com/MinhasKamal/BengaliDictionary? \n\n2) You mentioned `Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics` - may I know how you decomposed those graphemes that are not in our 1295 graphemes? \n\nThanks and congrats on your top ranking",
      "votes": 1,
      "replies": [
        {
          "id": 777979,
          "postDate": "2020-03-18T03:52:12.977Z",
          "content": "<p><a href=\"/weimin\">@weimin</a> Answering (2) first because it helps with (1)\n2) We can decompose each grapheme into a sequence of unicode character simply by calling character[i] after reading the character from .csv (or the dictionary). This sequence is generally (95%+ of the time) a combination of the grapheme's RCV in one order or another. I created a comprehensive list of 168x11x7x6 possible combinations this way and did simple matching. Some of the graphemes (~5%) are more complicated, there might be repeated characters in the sequence, so for graphemes that failed stage1 decomposition I simply did matching based on unique elements (set(seq)==set(element of comprehensive list)). Some items in the dictionary also have extraordinarily long sequences; I guess this is because they are compound characters; I simply assumed they will not appear in the testset and did not choose any grapheme with its sequence length &gt;10.</p>\n\n<p>1) I tried to decompose every graheme in the dictionary, and simply ignored graphemes that cannot be labeled through the method in (2) and ignored too-long graphemes. Because we had so short time left the method in (2) is only preliminary (has 2 failure-cases which I had to hardcode for graphemes in train). I'm sure that can be improved.</p>\n\n<p>Hope this helps</p>",
          "rawMarkdown": "@weimin Answering (2) first because it helps with (1)\n2) We can decompose each grapheme into a sequence of unicode character simply by calling character[i] after reading the character from .csv (or the dictionary). This sequence is generally (95%+ of the time) a combination of the grapheme's RCV in one order or another. I created a comprehensive list of 168x11x7x6 possible combinations this way and did simple matching. Some of the graphemes (~5%) are more complicated, there might be repeated characters in the sequence, so for graphemes that failed stage1 decomposition I simply did matching based on unique elements (set(seq)==set(element of comprehensive list)). Some items in the dictionary also have extraordinarily long sequences; I guess this is because they are compound characters; I simply assumed they will not appear in the testset and did not choose any grapheme with its sequence length &gt;10.\n\n1) I tried to decompose every graheme in the dictionary, and simply ignored graphemes that cannot be labeled through the method in (2) and ignored too-long graphemes. Because we had so short time left the method in (2) is only preliminary (has 2 failure-cases which I had to hardcode for graphemes in train). I'm sure that can be improved.\n\nHope this helps",
          "votes": 1
        },
        {
          "id": 777990,
          "postDate": "2020-03-18T04:05:40.537Z",
          "content": "<p>wow, I loaded the train.csv and did character[0] for one grapheme character, and immediately saw the majic! and it can be reverse-engineered in this way as well: character == character[0] + ... + character[5]. Can imagine it must have been the  'aha moment' when you found it out, congrats again :) </p>",
          "rawMarkdown": "wow, I loaded the train.csv and did character[0] for one grapheme character, and immediately saw the majic! and it can be reverse-engineered in this way as well: character == character[0] + ... + character[5]. Can imagine it must have been the  'aha moment' when you found it out, congrats again :) "
        },
        {
          "id": 778109,
          "postDate": "2020-03-18T06:43:34.843Z",
          "content": "<p>Happy to share it😊 </p>",
          "rawMarkdown": "Happy to share it😊 "
        }
      ]
    },
    {
      "id": 775821,
      "postDate": "2020-03-17T01:24:11.123Z",
      "content": "<p>Congrats! Wow, no wonder you guys get a gold medal. Good job! </p>",
      "rawMarkdown": "Congrats! Wow, no wonder you guys get a gold medal. Good job! ",
      "votes": 1
    },
    {
      "id": 775776,
      "postDate": "2020-03-17T00:52:59.820Z",
      "content": "<p>ingenious solution! well done and congrats!</p>",
      "rawMarkdown": "ingenious solution! well done and congrats!",
      "votes": 1
    },
    {
      "id": 776541,
      "postDate": "2020-03-17T13:16:54.537Z",
      "content": "<p>Congrats on the result and the method for generating additional data using various fonts!</p>",
      "rawMarkdown": "Congrats on the result and the method for generating additional data using various fonts!",
      "votes": 2
    },
    {
      "id": 775813,
      "postDate": "2020-03-17T01:19:53.313Z",
      "content": "<p>Congratulations！Thank you for sharing. Would you mind share the scores before and after using magic?</p>",
      "rawMarkdown": "Congratulations！Thank you for sharing. Would you mind share the scores before and after using magic?",
      "votes": 2,
      "replies": [
        {
          "id": 775817,
          "postDate": "2020-03-17T01:22:58.893Z",
          "content": "<p>before:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F21ddd8f6ddee6612a572781d6c469046%2FScreen%20Shot%202020-03-16%20at%209.21.37%20PM.png?generation=1584408145602961&amp;alt=media\" alt=\"\"></p>\n\n<p>after:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Ffc641882da8f41f5f567f02f33324cdc%2FScreen%20Shot%202020-03-16%20at%209.21.57%20PM.png?generation=1584408168087885&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "before:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F21ddd8f6ddee6612a572781d6c469046%2FScreen%20Shot%202020-03-16%20at%209.21.37%20PM.png?generation=1584408145602961&amp;alt=media)\n\nafter:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Ffc641882da8f41f5f567f02f33324cdc%2FScreen%20Shot%202020-03-16%20at%209.21.57%20PM.png?generation=1584408168087885&amp;alt=media)\n",
          "votes": 4
        },
        {
          "id": 775824,
          "postDate": "2020-03-17T01:27:44.777Z",
          "content": "<p>Good job, Thanks for sharing</p>",
          "rawMarkdown": "Good job, Thanks for sharing",
          "votes": 1
        },
        {
          "id": 775833,
          "postDate": "2020-03-17T01:30:59.767Z",
          "content": "<p>Thank you!\nfor our final submission we got score 3 min before competition deadline =) </p>",
          "rawMarkdown": "Thank you!\nfor our final submission we got score 3 min before competition deadline =) ",
          "votes": 3
        },
        {
          "id": 776237,
          "postDate": "2020-03-17T08:25:07.797Z",
          "content": "<p>I got similar cv / lb scores as yours before using the magic but couldn't find a proper way to generalize to unseen images :( You guys are really doing a great job in exploration.</p>",
          "rawMarkdown": "I got similar cv / lb scores as yours before using the magic but couldn't find a proper way to generalize to unseen images :( You guys are really doing a great job in exploration.",
          "votes": 1
        }
      ]
    },
    {
      "id": 775798,
      "postDate": "2020-03-17T01:07:39Z",
      "content": "<p>I'm confused about \"hand-monitored dropLR\", does it means you stop the training, adjust the learning rate, and resume the training?  Congrats by the way!</p>",
      "rawMarkdown": " I'm confused about \"hand-monitored dropLR\", does it means you stop the training, adjust the learning rate, and resume the training?  Congrats by the way!",
      "votes": 2,
      "replies": [
        {
          "id": 775819,
          "postDate": "2020-03-17T01:23:22.263Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes"
        }
      ]
    },
    {
      "id": 776197,
      "postDate": "2020-03-17T07:34:13.907Z",
      "content": "<p>can you share the code?</p>",
      "rawMarkdown": "can you share the code?",
      "votes": 1,
      "replies": [
        {
          "id": 776526,
          "postDate": "2020-03-17T13:12:03.230Z",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> Sorry our code is quite scattered :) My teammates should be able to handle this. But you should be able to achieve comparable performance with any normal high-scoring pipeline (.9970+CV locally) + simple adding of the synthetic data. You can find the data link herehttps://www.kaggle.com/roguekk007/bengaliai-synthetic-magic/kernels. It's quite clean. I suggest you to simply add this to your best CV pipeline and experiment😃 </p>",
          "rawMarkdown": "@kurianbenoy Sorry our code is quite scattered :) My teammates should be able to handle this. But you should be able to achieve comparable performance with any normal high-scoring pipeline (.9970+CV locally) + simple adding of the synthetic data. You can find the data link herehttps://www.kaggle.com/roguekk007/bengaliai-synthetic-magic/kernels. It's quite clean. I suggest you to simply add this to your best CV pipeline and experiment😃 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 775684,
      "postDate": "2020-03-17T00:01:46.340Z",
      "content": "<p>Happy that the shakeup has been kind to us :) </p>",
      "rawMarkdown": "Happy that the shakeup has been kind to us :) ",
      "votes": 1,
      "replies": [
        {
          "id": 775698,
          "postDate": "2020-03-17T00:09:35.437Z",
          "content": "<p>Great to see your solution, especially on how you handled the unseen graphemes with synthetic data</p>\n\n<p>Good that shakeup was kind to you, I lost a lot due to the shakeup!</p>",
          "rawMarkdown": "Great to see your solution, especially on how you handled the unseen graphemes with synthetic data\n\nGood that shakeup was kind to you, I lost a lot due to the shakeup!\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 777835,
      "postDate": "2020-03-18T01:40:37.973Z",
      "content": "<p>Congrats for winning medal with young age!</p>",
      "rawMarkdown": "Congrats for winning medal with young age!",
      "replies": [
        {
          "id": 777980,
          "postDate": "2020-03-18T03:52:48.637Z",
          "content": "<p>Your kernel helped me a lot during establishing baseline. Sorry about the shakeup</p>",
          "rawMarkdown": "Your kernel helped me a lot during establishing baseline. Sorry about the shakeup"
        }
      ]
    },
    {
      "id": 776191,
      "postDate": "2020-03-17T07:28:15.773Z",
      "content": "<p>Thanx, It might be helpful 👍 </p>",
      "rawMarkdown": "Thanx, It might be helpful 👍 "
    },
    {
      "id": 776187,
      "postDate": "2020-03-17T07:26:28.397Z",
      "content": "<p>Congrats! and Thank you for your sharing!\nWould you tell me about <code>Train on 128x128 data for ~100 epochs, then finetune on 224x224 data.</code></p>\n\n<p>Did you split data for pretraining and finetune? or use same data?</p>",
      "rawMarkdown": "Congrats! and Thank you for your sharing!\nWould you tell me about ```Train on 128x128 data for ~100 epochs, then finetune on 224x224 data.```\n\nDid you split data for pretraining and finetune? or use same data?",
      "replies": [
        {
          "id": 776527,
          "postDate": "2020-03-17T13:12:31.950Z",
          "content": "<p>We keep the split consistent in the two stages to prevent potential leaking</p>",
          "rawMarkdown": "We keep the split consistent in the two stages to prevent potential leaking",
          "votes": 1
        },
        {
          "id": 776615,
          "postDate": "2020-03-17T14:05:01.137Z",
          "content": "<p>Nice work! thank you ;)</p>",
          "rawMarkdown": "Nice work! thank you ;)"
        }
      ]
    },
    {
      "id": 776036,
      "postDate": "2020-03-17T05:01:03.537Z",
      "content": "<p>Congratulation for for the final quantum jump!</p>",
      "rawMarkdown": "Congratulation for for the final quantum jump!"
    },
    {
      "id": 775989,
      "postDate": "2020-03-17T04:12:08.013Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 775891,
      "postDate": "2020-03-17T02:21:57.163Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 775847,
      "postDate": "2020-03-17T01:47:11.763Z",
      "content": "<p>Congratulations. With the synthetic dataset, what validation split are you using? Do you use something like Qishen Ha's method for unseen? </p>",
      "rawMarkdown": "Congratulations. With the synthetic dataset, what validation split are you using? Do you use something like Qishen Ha's method for unseen? ",
      "replies": [
        {
          "id": 775878,
          "postDate": "2020-03-17T02:13:43.580Z",
          "content": "<p><a href=\"/murphy89\">@murphy89</a> Thank you. we are using simple stratified KFold split for synthetic dataset. We want the model to see as many graphemes as possible</p>",
          "rawMarkdown": "@murphy89 Thank you. we are using simple stratified KFold split for synthetic dataset. We want the model to see as many graphemes as possible",
          "votes": 1
        }
      ]
    },
    {
      "id": 775717,
      "postDate": "2020-03-17T00:24:01.640Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 775763,
      "postDate": "2020-03-17T00:44:44.940Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 780442,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-03-20T09:15:46.110000",
      "content": "<p>Congratulation <a href=\"/roguekk007\">@roguekk007</a>. <br>\nhope to work with you one day. Anyway, good job 💙 🙂 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 783554,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-23T13:20:41.050000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> hi, just curious to know, how you build <a href=\"https://www.kaggle.com/roguekk007/bengaliai-synthetic-magic\">this dataset</a> from a font?  Would you please share the implemented code? Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 783750,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-03-23T16:15:47.183000",
          "content": "<p>Hi! \nyou can refer to this: \n<a href=\"https://www.kaggle.com/drhabib/generating-more-training-data\">https://www.kaggle.com/drhabib/generating-more-training-data</a>\nbut correct displaying using this discussion \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127938#775496\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127938#775496</a></p>\n\n<p>Hope it helps =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 784470,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-24T08:31:53.803000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks a lot 💙 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 777773,
      "author_name": "Weimin Wang",
      "author_url": "",
      "post_date": "2020-03-17T23:29:53.953000",
      "content": "<p>Very impressive idea! I never thought of generating synthetic fonts like this way :)</p>\n\n<p>I have 2 questions - </p>\n\n<p>1) for your synthetic images, did you generate extra graphemes from the given parts of 168 (r) + 11 (v) + 7 (c) , or you decompose those r v c from those extra graphemes fonts from <a href=\"https://github.com/MinhasKamal/BengaliDictionary\">https://github.com/MinhasKamal/BengaliDictionary</a>? </p>\n\n<p>2) You mentioned <code>Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics</code> - may I know how you decomposed those graphemes that are not in our 1295 graphemes? </p>\n\n<p>Thanks and congrats on your top ranking</p>",
      "votes": 1,
      "replies": [
        {
          "id": 777979,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-18T03:52:12.977000",
          "content": "<p><a href=\"/weimin\">@weimin</a> Answering (2) first because it helps with (1)\n2) We can decompose each grapheme into a sequence of unicode character simply by calling character[i] after reading the character from .csv (or the dictionary). This sequence is generally (95%+ of the time) a combination of the grapheme's RCV in one order or another. I created a comprehensive list of 168x11x7x6 possible combinations this way and did simple matching. Some of the graphemes (~5%) are more complicated, there might be repeated characters in the sequence, so for graphemes that failed stage1 decomposition I simply did matching based on unique elements (set(seq)==set(element of comprehensive list)). Some items in the dictionary also have extraordinarily long sequences; I guess this is because they are compound characters; I simply assumed they will not appear in the testset and did not choose any grapheme with its sequence length &gt;10.</p>\n\n<p>1) I tried to decompose every graheme in the dictionary, and simply ignored graphemes that cannot be labeled through the method in (2) and ignored too-long graphemes. Because we had so short time left the method in (2) is only preliminary (has 2 failure-cases which I had to hardcode for graphemes in train). I'm sure that can be improved.</p>\n\n<p>Hope this helps</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 777990,
          "author_name": "Weimin Wang",
          "author_url": "",
          "post_date": "2020-03-18T04:05:40.537000",
          "content": "<p>wow, I loaded the train.csv and did character[0] for one grapheme character, and immediately saw the majic! and it can be reverse-engineered in this way as well: character == character[0] + ... + character[5]. Can imagine it must have been the  'aha moment' when you found it out, congrats again :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 778109,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-18T06:43:34.843000",
          "content": "<p>Happy to share it😊 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 775821,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-17T01:24:11.123000",
      "content": "<p>Congrats! Wow, no wonder you guys get a gold medal. Good job! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775776,
      "author_name": "YL",
      "author_url": "",
      "post_date": "2020-03-17T00:52:59.820000",
      "content": "<p>ingenious solution! well done and congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776541,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-03-17T13:16:54.537000",
      "content": "<p>Congrats on the result and the method for generating additional data using various fonts!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 775813,
      "author_name": "He",
      "author_url": "",
      "post_date": "2020-03-17T01:19:53.313000",
      "content": "<p>Congratulations！Thank you for sharing. Would you mind share the scores before and after using magic?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 775817,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-03-17T01:22:58.893000",
          "content": "<p>before:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F21ddd8f6ddee6612a572781d6c469046%2FScreen%20Shot%202020-03-16%20at%209.21.37%20PM.png?generation=1584408145602961&amp;alt=media\" alt=\"\"></p>\n\n<p>after:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Ffc641882da8f41f5f567f02f33324cdc%2FScreen%20Shot%202020-03-16%20at%209.21.57%20PM.png?generation=1584408168087885&amp;alt=media\" alt=\"\"></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 775824,
          "author_name": "He",
          "author_url": "",
          "post_date": "2020-03-17T01:27:44.777000",
          "content": "<p>Good job, Thanks for sharing</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 775833,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-03-17T01:30:59.767000",
          "content": "<p>Thank you!\nfor our final submission we got score 3 min before competition deadline =) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 776237,
          "author_name": "syoya",
          "author_url": "",
          "post_date": "2020-03-17T08:25:07.797000",
          "content": "<p>I got similar cv / lb scores as yours before using the magic but couldn't find a proper way to generalize to unseen images :( You guys are really doing a great job in exploration.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 775798,
      "author_name": "HOHOYAO",
      "author_url": "",
      "post_date": "2020-03-17T01:07:39",
      "content": "<p>I'm confused about \"hand-monitored dropLR\", does it means you stop the training, adjust the learning rate, and resume the training?  Congrats by the way!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 775819,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-03-17T01:23:22.263000",
          "content": "<p>yes</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776197,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-03-17T07:34:13.907000",
      "content": "<p>can you share the code?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776526,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-17T13:12:03.230000",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> Sorry our code is quite scattered :) My teammates should be able to handle this. But you should be able to achieve comparable performance with any normal high-scoring pipeline (.9970+CV locally) + simple adding of the synthetic data. You can find the data link herehttps://www.kaggle.com/roguekk007/bengaliai-synthetic-magic/kernels. It's quite clean. I suggest you to simply add this to your best CV pipeline and experiment😃 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 775684,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-03-17T00:01:46.340000",
      "content": "<p>Happy that the shakeup has been kind to us :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 775698,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-03-17T00:09:35.437000",
          "content": "<p>Great to see your solution, especially on how you handled the unseen graphemes with synthetic data</p>\n\n<p>Good that shakeup was kind to you, I lost a lot due to the shakeup!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 777835,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-03-18T01:40:37.973000",
      "content": "<p>Congrats for winning medal with young age!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 777980,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-18T03:52:48.637000",
          "content": "<p>Your kernel helped me a lot during establishing baseline. Sorry about the shakeup</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776191,
      "author_name": "Umer Nawaz",
      "author_url": "",
      "post_date": "2020-03-17T07:28:15.773000",
      "content": "<p>Thanx, It might be helpful 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776187,
      "author_name": "kaerururu",
      "author_url": "",
      "post_date": "2020-03-17T07:26:28.397000",
      "content": "<p>Congrats! and Thank you for your sharing!\nWould you tell me about <code>Train on 128x128 data for ~100 epochs, then finetune on 224x224 data.</code></p>\n\n<p>Did you split data for pretraining and finetune? or use same data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 776527,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-17T13:12:31.950000",
          "content": "<p>We keep the split consistent in the two stages to prevent potential leaking</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776615,
          "author_name": "kaerururu",
          "author_url": "",
          "post_date": "2020-03-17T14:05:01.137000",
          "content": "<p>Nice work! thank you ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776036,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-17T05:01:03.537000",
      "content": "<p>Congratulation for for the final quantum jump!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775989,
      "author_name": "cswwp",
      "author_url": "",
      "post_date": "2020-03-17T04:12:08.013000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775891,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2020-03-17T02:21:57.163000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775847,
      "author_name": "Morphy",
      "author_url": "",
      "post_date": "2020-03-17T01:47:11.763000",
      "content": "<p>Congratulations. With the synthetic dataset, what validation split are you using? Do you use something like Qishen Ha's method for unseen? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 775878,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-17T02:13:43.580000",
          "content": "<p><a href=\"/murphy89\">@murphy89</a> Thank you. we are using simple stratified KFold split for synthetic dataset. We want the model to see as many graphemes as possible</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 775717,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T00:24:01.640000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775763,
      "author_name": "ccchang",
      "author_url": "",
      "post_date": "2020-03-17T00:44:44.940000",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "775683": "Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. It is my first all-in kaggle competition, and needless to say none of this would have been possible without our awesome team Igor, Habib, Rinat, and Youhan. @cateek @drhabib @trytolose @youhanlee \n\n### Analysis: This competition has a distinctly stratified LB:\n- With a good pipeline, single model with baseline augmentations LB ~.97. \n- **Removing crops, training on full-resolution, adding cutmix / cutout, and training enough** should get model up to LB .985+, and with some tweaking ~.989+. These are well-covered in the discussions posts.\n- Our best single model is from Rinat, PNASNet-5-Large from [Cadene's repository](https://github.com/Cadene/pretrained-models.pytorch) which scores LB 0.9900, CV .9985. I believe most of the participants LB .9850-.9910 are using ensemble of models scoring around this range.\n\nAt this stage (.9900-.9905) we found it very hard to further improve our LB. As [this post](https://www.kaggle.com/c/bengaliai-cv19/discussion/134601) points out, higher CV does not necessary mean better LB anymore, so we were stuck for a while (a long while...).\n\n### Improving Generalization\nWith the CV &amp; LB relationship it is easy to see that the problem is that our model **could not generalize to unseen graphemes** (by grapheme I mean triplet combination of grapheme root, vowel and consonant diacritics). There are only ~1290 graphemes in training set, while there are 12936 possible combinations. Our models might be biased to output seen triplets; for whatever reason, as Qishen Ha's wonderful [post](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434) pointed out, our models perform badly on these graphemes.\n\nIn hindsight, it's no magic that the most convenient way to improve generalization is to **add more data**. When investigating the grapheme representations, I found that **graphemes are actually encoded as sequence of unicode characters**. The roots and diacritics also have their corresponding sequences. The most magical part is that **the unicode sequence for the grapheme is a combination of the sequences of its roots**. Best exemplified using this image, where the four rows are *grapheme, grapheme root, vowel, and conso*.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F82988244ae7fb6d4ca47f6c6094a3116%2Fimage_480.png?generation=1584399876994128&amp;alt=media)\n\nThis observation works (with some minor exceptions) on all graphemes in the training set, and using a very basic algorithm I was able to generate correct labels for all 1295 graphemes in training set except for 2. Following Guanshuo Xu's post pointing to [this repository](https://github.com/MinhasKamal/BengaliDictionary), we are able to access a comprehensive list of grapheme combinations. Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics. Our LB boost on the final day is from better generalization on these graphemes, which partially overlaps with train &amp; test set.\n\nWith these graphemes, we rendered them using various fonts and obtained a cleanly-labeled synthetic dataset with ~47K images spanning these characters which can be found [here](https://www.kaggle.com/roguekk007/bengaliai-synthetic-magic), all the ingredients for creating this synthetic dataset has been publicly available. Our hope was that by seeing synthetic images, our models should be able to at least learn the topological (if not stylistic) features of the grapheme. And it turned out they do. Here is a sample of synthetic image in our dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Ff582156ca8f174802dbbc28a18fa1562%2F1_EkSharifa.png?generation=1584402247696903&amp;alt=media)\n\nAt this point we have only 24 hours until the end of the competition so just finetuned our previous models with the synthetic data added to both original training &amp; validation, thus resulting in the exciting quantum jump on the last day😊 \n\n### Some Details\nMy teammates have wonderful pipelines, here are some approaches which have turned out useful:\n- Train on 128x128 data for ~100 epochs, then finetune on 224x224 data. It definitely yields faster training and maybe better generalization (used in Rinat's pipeline)\n- Anneal augmentations except for cutmix gradually to 0 at late stages of training (used in Habib's and Igor's pipeline)\n- A wonderful new architecture called MixNet, we found that it nicely complements PNAS5 for ensembling and performs pretty well.\n- **Loss:** Plain cross entropy. Tried arface (did not work), focal (did not work), normalized softmax+label smoothing (looked promising but ditched along with my pipeline)\n- **Scheduler:** ReduceLRonPleateau for Rinat's pipeline, and hand-monitored dropLR for Habib's and Igor's pipelines.\n\n### Finally\nIt has always been my dream to do a gold-medal solution write-up before 17th birthday😃 . In the end, more than happy to see hundreds of experiments and weeks of stress translate successfully into solid rank. With more time than 24 hours, we would have: \n1. Searched for a more comprehensive list of graphemes\n2. Created more synthetic data with more fonts\n3. Somehow make synthetic data more similar to handwritten data; we are using dimming max_val to 230 and gaussian-blur, but can definitely be improved by maybe CycleGAN\n4. Make flipping work as pointed out [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/126761). It's really ironic I did not make my own idea work when Qishen Ha did; guess this is the difference🤓 \n5. Try more model-regularization methods, ShakeDrop was on the list and looked promising but was preceded by other priorities.\n\nOverall, this has a very exciting and memorable puzzle-solving competition. Again kudos to the team and:\nLove Kaggling",
    "780442": "Congratulation @roguekk007.  \nhope to work with you one day. Anyway, good job 💙 🙂 ",
    "777773": "Very impressive idea! I never thought of generating synthetic fonts like this way :)\n\nI have 2 questions - \n\n1) for your synthetic images, did you generate extra graphemes from the given parts of 168 (r) + 11 (v) + 7 (c) , or you decompose those r v c from those extra graphemes fonts from https://github.com/MinhasKamal/BengaliDictionary? \n\n2) You mentioned `Using the decomposition algorithm we selected ~3000 graphemes that could be broken down into the given roots and diacritics` - may I know how you decomposed those graphemes that are not in our 1295 graphemes? \n\nThanks and congrats on your top ranking",
    "775821": "Congrats! Wow, no wonder you guys get a gold medal. Good job! ",
    "775776": "ingenious solution! well done and congrats!",
    "776541": "Congrats on the result and the method for generating additional data using various fonts!",
    "775813": "Congratulations！Thank you for sharing. Would you mind share the scores before and after using magic?",
    "775798": " I'm confused about \"hand-monitored dropLR\", does it means you stop the training, adjust the learning rate, and resume the training?  Congrats by the way!",
    "776197": "can you share the code?",
    "775684": "Happy that the shakeup has been kind to us :) ",
    "777835": "Congrats for winning medal with young age!",
    "776191": "Thanx, It might be helpful 👍 ",
    "776187": "Congrats! and Thank you for your sharing!\nWould you tell me about ```Train on 128x128 data for ~100 epochs, then finetune on 224x224 data.```\n\nDid you split data for pretraining and finetune? or use same data?",
    "776036": "Congratulation for for the final quantum jump!",
    "775989": "Congratulations",
    "775891": "Congratulations!",
    "775847": "Congratulations. With the synthetic dataset, what validation split are you using? Do you use something like Qishen Ha's method for unseen? ",
    "775717": "",
    "775763": "Congrats and thanks for sharing!"
  }
}