{
  "id": 136006,
  "title": "All Combinations 168 x 11 x 7 + Edge Cases",
  "url": "/competitions/bengaliai-cv19/discussion/136006",
  "author_name": "",
  "post_date": "2020-03-17T03:08:11.883755900Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>New to kaggle and it was fun learning a lot of things along the way. I quickly learnt that this competition was no brainer and required some definite hacks to get above 0.99. </p>\n\n<p>I came across Unicode kernels that helped me think for solid 2 days to generate all the possible combinations of root, vowel and consonant - all though most of them wouldn't make sense in the native language.</p>\n\n<p>Finally, with manual observation, I came up with the following rules (please consume the information with a grain of salt and let me know what native speakers of Bengali think about it)</p>\n\n<p><strong>Combination Rules</strong></p>\n\n<ol>\n<li>Consonant Label-1 has the highest priority on right and precedes vowel position.</li>\n<li>Consonant Label-2 has the highest priority on left and preceeds root and grapheme position.</li>\n<li>Consonant Label-3 has two components that wraps around the root with vowel taking the highest priority on left.</li>\n<li>Consonant Label-4,5,6 has right priority with the highest priority given to vowel if any.</li>\n</ol>\n\n<p>You can find details of the logic in this <strong><a href=\"https://www.kaggle.com/harshpatel1692/all-combinations-168-x-11-x-7-edge-cases\">kernel</a></strong></p>\n\n<p>I created 200 augmented (AugMix, CutMix, ShiftScaleRotate, IAAPiecewiseAffine, Distortion, GaussNoise, etc) images of each grapheme (14K times 200) to have perfect balanced data and was heartbroken that the model trained on that data showed negligible predictive power on test data. </p>\n\n<p>You can find all the combinations image set in 128x128 <a href=\"https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases\">here</a>.</p>\n\n<p><strong>My Best Score and Setup</strong>:\n- SeResNext50\n- Medium Augmix and CutMix as part of training data\n- Random rotation, brightness and crop+resize\n- ADAM LR 1e-3 to 1e-6\n- 2 Folds\n- Trained on TPUs</p>\n\n<p><strong>What didn't work for me</strong>\n- Training on all combinations (I definitely might have done something wrong in a hurry)\n- ArcFace on my pre-trained seresnext50. Didn't converge well\n- EfficientNetB5. For some reason, they were converging badly for me.\n- 3 Models for each classification channel.</p>\n\n<p>At last, I ended up with 0.96 on Public and 0.91 on Private with tons of knowledge.</p>\n\n<p>Congrats to the winners and excited to learn their architecture and tricks.</p>\n\n<p>Please vote this discussion and the kernel if it helped you!</p>",
  "messages": [
    {
      "id": "775936",
      "postDate": "03/17/2020 03:08:11",
      "content": "<p>New to kaggle and it was fun learning a lot of things along the way. I quickly learnt that this competition was no brainer and required some definite hacks to get above 0.99. </p>\n\n<p>I came across Unicode kernels that helped me think for solid 2 days to generate all the possible combinations of root, vowel and consonant - all though most of them wouldn't make sense in the native language.</p>\n\n<p>Finally, with manual observation, I came up with the following rules (please consume the information with a grain of salt and let me know what native speakers of Bengali think about it)</p>\n\n<p><strong>Combination Rules</strong></p>\n\n<ol>\n<li>Consonant Label-1 has the highest priority on right and precedes vowel position.</li>\n<li>Consonant Label-2 has the highest priority on left and preceeds root and grapheme position.</li>\n<li>Consonant Label-3 has two components that wraps around the root with vowel taking the highest priority on left.</li>\n<li>Consonant Label-4,5,6 has right priority with the highest priority given to vowel if any.</li>\n</ol>\n\n<p>You can find details of the logic in this <strong><a href=\"https://www.kaggle.com/harshpatel1692/all-combinations-168-x-11-x-7-edge-cases\">kernel</a></strong></p>\n\n<p>I created 200 augmented (AugMix, CutMix, ShiftScaleRotate, IAAPiecewiseAffine, Distortion, GaussNoise, etc) images of each grapheme (14K times 200) to have perfect balanced data and was heartbroken that the model trained on that data showed negligible predictive power on test data. </p>\n\n<p>You can find all the combinations image set in 128x128 <a href=\"https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases\">here</a>.</p>\n\n<p><strong>My Best Score and Setup</strong>:\n- SeResNext50\n- Medium Augmix and CutMix as part of training data\n- Random rotation, brightness and crop+resize\n- ADAM LR 1e-3 to 1e-6\n- 2 Folds\n- Trained on TPUs</p>\n\n<p><strong>What didn't work for me</strong>\n- Training on all combinations (I definitely might have done something wrong in a hurry)\n- ArcFace on my pre-trained seresnext50. Didn't converge well\n- EfficientNetB5. For some reason, they were converging badly for me.\n- 3 Models for each classification channel.</p>\n\n<p>At last, I ended up with 0.96 on Public and 0.91 on Private with tons of knowledge.</p>\n\n<p>Congrats to the winners and excited to learn their architecture and tricks.</p>\n\n<p>Please vote this discussion and the kernel if it helped you!</p>",
      "rawMarkdown": "New to kaggle and it was fun learning a lot of things along the way. I quickly learnt that this competition was no brainer and required some definite hacks to get above 0.99. \n\nI came across Unicode kernels that helped me think for solid 2 days to generate all the possible combinations of root, vowel and consonant - all though most of them wouldn't make sense in the native language.\n\nFinally, with manual observation, I came up with the following rules (please consume the information with a grain of salt and let me know what native speakers of Bengali think about it)\n\n**Combination Rules**\n\n1. Consonant Label-1 has the highest priority on right and precedes vowel position.\n2. Consonant Label-2 has the highest priority on left and preceeds root and grapheme position.\n3. Consonant Label-3 has two components that wraps around the root with vowel taking the highest priority on left.\n4. Consonant Label-4,5,6 has right priority with the highest priority given to vowel if any.\n\nYou can find details of the logic in this **[kernel](https://www.kaggle.com/harshpatel1692/all-combinations-168-x-11-x-7-edge-cases)**\n\nI created 200 augmented (AugMix, CutMix, ShiftScaleRotate, IAAPiecewiseAffine, Distortion, GaussNoise, etc) images of each grapheme (14K times 200) to have perfect balanced data and was heartbroken that the model trained on that data showed negligible predictive power on test data. \n\nYou can find all the combinations image set in 128x128 [here](https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases).\n\n**My Best Score and Setup**:\n- SeResNext50\n- Medium Augmix and CutMix as part of training data\n- Random rotation, brightness and crop+resize\n- ADAM LR 1e-3 to 1e-6\n- 2 Folds\n- Trained on TPUs\n\n**What didn't work for me**\n- Training on all combinations (I definitely might have done something wrong in a hurry)\n- ArcFace on my pre-trained seresnext50. Didn't converge well\n- EfficientNetB5. For some reason, they were converging badly for me.\n- 3 Models for each classification channel.\n\nAt last, I ended up with 0.96 on Public and 0.91 on Private with tons of knowledge.\n\nCongrats to the winners and excited to learn their architecture and tricks.\n\nPlease vote this discussion and the kernel if it helped you!",
      "votes": null
    },
    {
      "id": "776272",
      "postDate": "03/17/2020 08:58:12",
      "content": "<p>Hi Harsh, I had a similar approach. I used 30+ fonts to generate all legitimate grapheme combinations (~12k). There was some problem with the way PIL font writer was rendering the graphemes, and I see the same in the edge cases that you shared. It's hard to spot but since Bengali is my mother language, I saw the problem, I'd never write the characters the way it was rendered. </p>\n\n<p>I used pyvip library instead to render the fonts. Ultimately, due to lack of time / compute resource, I used only 6 of the 30 fonts to augment my images. Helped me jump +1350 places from the Public LB. So I can confirm that your approach works. I'll share details of my augmented dataset soon.</p>",
      "rawMarkdown": "Hi Harsh, I had a similar approach. I used 30+ fonts to generate all legitimate grapheme combinations (~12k). There was some problem with the way PIL font writer was rendering the graphemes, and I see the same in the edge cases that you shared. It's hard to spot but since Bengali is my mother language, I saw the problem, I'd never write the characters the way it was rendered. \n\nI used pyvip library instead to render the fonts. Ultimately, due to lack of time / compute resource, I used only 6 of the 30 fonts to augment my images. Helped me jump +1350 places from the Public LB. So I can confirm that your approach works. I'll share details of my augmented dataset soon.",
      "votes": null
    },
    {
      "id": "776755",
      "postDate": "03/17/2020 16:01:43",
      "content": "<p>I realized the problem with PIL and I even tried Tkinter to grab the canvas image but had no luck.\nI finally ended up building a selenium scraper that would post the character on a website and grab the exact image from the canvas from there. I have shared the full image dataset over <a href=\"https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases\">here</a>.</p>\n\n<p>Let me know if that helps!</p>",
      "rawMarkdown": "I realized the problem with PIL and I even tried Tkinter to grab the canvas image but had no luck.\nI finally ended up building a selenium scraper that would post the character on a website and grab the exact image from the canvas from there. I have shared the full image dataset over [here](https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases).\n\nLet me know if that helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 776272,
      "author_name": "sbpdev",
      "author_url": "",
      "post_date": "03/17/2020 08:58:12",
      "content": "<p>Hi Harsh, I had a similar approach. I used 30+ fonts to generate all legitimate grapheme combinations (~12k). There was some problem with the way PIL font writer was rendering the graphemes, and I see the same in the edge cases that you shared. It's hard to spot but since Bengali is my mother language, I saw the problem, I'd never write the characters the way it was rendered. </p>\n\n<p>I used pyvip library instead to render the fonts. Ultimately, due to lack of time / compute resource, I used only 6 of the 30 fonts to augment my images. Helped me jump +1350 places from the Public LB. So I can confirm that your approach works. I'll share details of my augmented dataset soon.</p>",
      "votes": null,
      "replies": [
        {
          "id": 776755,
          "author_name": "harshpatel1692",
          "author_url": "",
          "post_date": "03/17/2020 16:01:43",
          "content": "<p>I realized the problem with PIL and I even tried Tkinter to grab the canvas image but had no luck.\nI finally ended up building a selenium scraper that would post the character on a website and grab the exact image from the canvas from there. I have shared the full image dataset over <a href=\"https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases\">here</a>.</p>\n\n<p>Let me know if that helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "775936": "New to kaggle and it was fun learning a lot of things along the way. I quickly learnt that this competition was no brainer and required some definite hacks to get above 0.99. \n\nI came across Unicode kernels that helped me think for solid 2 days to generate all the possible combinations of root, vowel and consonant - all though most of them wouldn't make sense in the native language.\n\nFinally, with manual observation, I came up with the following rules (please consume the information with a grain of salt and let me know what native speakers of Bengali think about it)\n\n**Combination Rules**\n\n1. Consonant Label-1 has the highest priority on right and precedes vowel position.\n2. Consonant Label-2 has the highest priority on left and preceeds root and grapheme position.\n3. Consonant Label-3 has two components that wraps around the root with vowel taking the highest priority on left.\n4. Consonant Label-4,5,6 has right priority with the highest priority given to vowel if any.\n\nYou can find details of the logic in this **[kernel](https://www.kaggle.com/harshpatel1692/all-combinations-168-x-11-x-7-edge-cases)**\n\nI created 200 augmented (AugMix, CutMix, ShiftScaleRotate, IAAPiecewiseAffine, Distortion, GaussNoise, etc) images of each grapheme (14K times 200) to have perfect balanced data and was heartbroken that the model trained on that data showed negligible predictive power on test data. \n\nYou can find all the combinations image set in 128x128 [here](https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases).\n\n**My Best Score and Setup**:\n- SeResNext50\n- Medium Augmix and CutMix as part of training data\n- Random rotation, brightness and crop+resize\n- ADAM LR 1e-3 to 1e-6\n- 2 Folds\n- Trained on TPUs\n\n**What didn't work for me**\n- Training on all combinations (I definitely might have done something wrong in a hurry)\n- ArcFace on my pre-trained seresnext50. Didn't converge well\n- EfficientNetB5. For some reason, they were converging badly for me.\n- 3 Models for each classification channel.\n\nAt last, I ended up with 0.96 on Public and 0.91 on Private with tons of knowledge.\n\nCongrats to the winners and excited to learn their architecture and tricks.\n\nPlease vote this discussion and the kernel if it helped you!",
    "776272": "Hi Harsh, I had a similar approach. I used 30+ fonts to generate all legitimate grapheme combinations (~12k). There was some problem with the way PIL font writer was rendering the graphemes, and I see the same in the edge cases that you shared. It's hard to spot but since Bengali is my mother language, I saw the problem, I'd never write the characters the way it was rendered. \n\nI used pyvip library instead to render the fonts. Ultimately, due to lack of time / compute resource, I used only 6 of the 30 fonts to augment my images. Helped me jump +1350 places from the Public LB. So I can confirm that your approach works. I'll share details of my augmented dataset soon.",
    "776755": "I realized the problem with PIL and I even tried Tkinter to grab the canvas image but had no luck.\nI finally ended up building a selenium scraper that would post the character on a website and grab the exact image from the canvas from there. I have shared the full image dataset over [here](https://www.kaggle.com/harshpatel1692/bengaliai-all-combinations-168x11x7-edge-cases).\n\nLet me know if that helps!"
  },
  "source": "meta"
}