{
  "id": 122758,
  "title": "Competition Close to My Heart",
  "url": "/competitions/bengaliai-cv19/discussion/122758",
  "author_name": "",
  "post_date": "2019-12-22T18:05:46.723160200Z",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>As a Bengali Speaking person , this competition is very close to my heart . I am extremely happy to see a lot of people from all-over the world trying to crack the Bengali Handwritten Graphemes .  At the end of this competition which would mean a good advancement in the area of our mother tongue and lot of different application to prosper the language  in keyboards , document processing , exam paper verification , many more areas . That is indeed a happy news for me . \nApart from this prelude , I wanted to jot down few things in this thread which I find interesting .\nSomeone with more linguistic knowledge can correct me . I am not sure if this can help in this comp as I have not yet fully joined this comp.</p>\n\n<ol>\n<li><p>Grapheme Root 2 to 12 are \"SwaraBarna\" or vowels . The vowel_diacritics (0-10) can not co-exist with these .(Not sure , if that makes any difference in the comp).</p></li>\n<li><p>Except for Consonant Diacritic 1 = ঁ , no other consonant diacritic also might come together with Grapheme root 2 to 12 .</p></li>\n<li><p>Remaining Grapheme Roots are Single Consonants and Grapheme root formed by making multiple Consonants come together e.g. ক = Grapheme Root 13 .  (Pronounce as Ka)  and ক্ক = Grapheme root 14 = KKa  like \"Trekking\"</p></li>\n</ol>\n\n<p>4.Consonant Diacritic 2 and 3 would not normally go with grapheme root : 159 to 167</p>\n\n<p>I will put together other things , as i look into the data more. </p>",
  "messages": [
    {
      "id": "700851",
      "postDate": "12/22/2019 18:05:46",
      "content": "<p>As a Bengali Speaking person , this competition is very close to my heart . I am extremely happy to see a lot of people from all-over the world trying to crack the Bengali Handwritten Graphemes .  At the end of this competition which would mean a good advancement in the area of our mother tongue and lot of different application to prosper the language  in keyboards , document processing , exam paper verification , many more areas . That is indeed a happy news for me . \nApart from this prelude , I wanted to jot down few things in this thread which I find interesting .\nSomeone with more linguistic knowledge can correct me . I am not sure if this can help in this comp as I have not yet fully joined this comp.</p>\n\n<ol>\n<li><p>Grapheme Root 2 to 12 are \"SwaraBarna\" or vowels . The vowel_diacritics (0-10) can not co-exist with these .(Not sure , if that makes any difference in the comp).</p></li>\n<li><p>Except for Consonant Diacritic 1 = ঁ , no other consonant diacritic also might come together with Grapheme root 2 to 12 .</p></li>\n<li><p>Remaining Grapheme Roots are Single Consonants and Grapheme root formed by making multiple Consonants come together e.g. ক = Grapheme Root 13 .  (Pronounce as Ka)  and ক্ক = Grapheme root 14 = KKa  like \"Trekking\"</p></li>\n</ol>\n\n<p>4.Consonant Diacritic 2 and 3 would not normally go with grapheme root : 159 to 167</p>\n\n<p>I will put together other things , as i look into the data more. </p>",
      "rawMarkdown": "As a Bengali Speaking person , this competition is very close to my heart . I am extremely happy to see a lot of people from all-over the world trying to crack the Bengali Handwritten Graphemes .  At the end of this competition which would mean a good advancement in the area of our mother tongue and lot of different application to prosper the language  in keyboards , document processing , exam paper verification , many more areas . That is indeed a happy news for me . \nApart from this prelude , I wanted to jot down few things in this thread which I find interesting .\nSomeone with more linguistic knowledge can correct me . I am not sure if this can help in this comp as I have not yet fully joined this comp.\n\n1. Grapheme Root 2 to 12 are \"SwaraBarna\" or vowels . The vowel_diacritics (0-10) can not co-exist with these .(Not sure , if that makes any difference in the comp).\n\n2. Except for Consonant Diacritic 1 = ঁ , no other consonant diacritic also might come together with Grapheme root 2 to 12 .\n\n3. Remaining Grapheme Roots are Single Consonants and Grapheme root formed by making multiple Consonants come together e.g. ক = Grapheme Root 13 .  (Pronounce as Ka)  and ক্ক = Grapheme root 14 = KKa  like \"Trekking\"\n\n4.Consonant Diacritic 2 and 3 would not normally go with grapheme root : 159 to 167\n\nI will put together other things , as i look into the data more.",
      "votes": null
    },
    {
      "id": "731218",
      "postDate": "01/28/2020 12:51:26",
      "content": "<p>Thanks <a href=\"/phoenix9032\">@phoenix9032</a> , this is a great post! </p>\n\n<p>I was just searching online for just this kind of combination rules of Bengali Graphemes, i.e., compatibility of (root, vowel, consonant). </p>\n\n<p>I didn't look at what public kernels here are doing yet. But these facts seem to very useful to me\n- In test, it rules out incompatible prediction of (root, vowel, consonant) triples so that a compatible prediction can be chosen, if it has a slight lower predicted probability than an incompatible triple prediction. I thought using triple dictionary in train as this rule book, i.e., never predict grapheme unseen in train. However, the data description says <code>The test set includes some graphemes that do not exist in train but has no new grapheme components</code>. But, the rules like what you stated should always work unless the organizers are including insensible graphemes. Is this correct?\n- In train, we can include additional loss functions terms by adding heads that check if the predicted triple is compatible.</p>",
      "rawMarkdown": "Thanks @phoenix9032 , this is a great post! \n\nI was just searching online for just this kind of combination rules of Bengali Graphemes, i.e., compatibility of (root, vowel, consonant). \n\nI didn't look at what public kernels here are doing yet. But these facts seem to very useful to me\n- In test, it rules out incompatible prediction of (root, vowel, consonant) triples so that a compatible prediction can be chosen, if it has a slight lower predicted probability than an incompatible triple prediction. I thought using triple dictionary in train as this rule book, i.e., never predict grapheme unseen in train. However, the data description says `The test set includes some graphemes that do not exist in train but has no new grapheme components`. But, the rules like what you stated should always work unless the organizers are including insensible graphemes. Is this correct?\n- In train, we can include additional loss functions terms by adding heads that check if the predicted triple is compatible.",
      "votes": null
    },
    {
      "id": "731393",
      "postDate": "01/28/2020 15:21:47",
      "content": "<p>You are right . There are combinations that are not possible in Bengali Language . Unless someone just write something out of the ordinary to test us . I will probably try to create more examples soon in future.</p>",
      "rawMarkdown": "You are right . There are combinations that are not possible in Bengali Language . Unless someone just write something out of the ordinary to test us . I will probably try to create more examples soon in future.",
      "votes": null
    },
    {
      "id": "732920",
      "postDate": "01/30/2020 13:04:55",
      "content": "<p>Hi Nirjhar Roy. Thank you for the valuable information. <br>\nI was also interested in the (root, vowel, consonant) combination.\nTrain data has 1292 combinations. For the time being, I tried to create logic to select the second softmax if predicted something other than these combinations. This logic slightly improved the accuracy for Train data, but unfortunately Test data did not have the corresponding prediction and was ineffective. Nirjhar's information may be useful for creating logic.</p>",
      "rawMarkdown": "Hi Nirjhar Roy. Thank you for the valuable information. <br>\nI was also interested in the (root, vowel, consonant) combination.\nTrain data has 1292 combinations. For the time being, I tried to create logic to select the second softmax if predicted something other than these combinations. This logic slightly improved the accuracy for Train data, but unfortunately Test data did not have the corresponding prediction and was ineffective. Nirjhar's information may be useful for creating logic.",
      "votes": null
    },
    {
      "id": "732994",
      "postDate": "01/30/2020 14:40:24",
      "content": "<p>I have not been looking into this competition for few days because of Google Quest competition .However , as per my preliminary observation : It is not important what combination is not-possible . It is more important , if some combination is visually similar to other combination. \nIn that way , the person without the domain knowledge would be more in advantage , because , for years my eyes are trained to know the difference . \ne.g : I see my model make most mistakes in : <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff75b2db54bedb2a1821a83c0feae64b5%2F85.png?generation=1580394764389892&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F534bcf244aa86cd1dca5dc36e0b76923%2F62.png?generation=1580394768946394&amp;alt=media\" alt=\"\"></p>\n\n<p>Also , Grapheme_Root 0-12 are vowels and some special characters . They normally dont take all vowel and consonant diacritic . </p>\n\n<p>Below are the rules: \nGrapheme_root 0,1 : If you see any consonant or vowel diacritic then set as 0 . \nGrapheme root 2 and 9 they can have consonant diacritic 4 and vowel_diacritic 1  together . So , if you predict root 2 or 9 and consonant 4 , then surely vowel diacritic is 1 . and vice versa.\nGrapheme root: 3,4,5,6,7,8,11,12 etc can have no vowel diacritic and have only 0 or 1 as consonant diacritic </p>\n\n<p>However , as long as your model is not mistaking  a lot , the post processing this way can give marginal benefit only . I guess .</p>",
      "rawMarkdown": "I have not been looking into this competition for few days because of Google Quest competition .However , as per my preliminary observation : It is not important what combination is not-possible . It is more important , if some combination is visually similar to other combination. \nIn that way , the person without the domain knowledge would be more in advantage , because , for years my eyes are trained to know the difference . \ne.g : I see my model make most mistakes in : ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff75b2db54bedb2a1821a83c0feae64b5%2F85.png?generation=1580394764389892&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F534bcf244aa86cd1dca5dc36e0b76923%2F62.png?generation=1580394768946394&amp;alt=media)\n\n\nAlso , Grapheme_Root 0-12 are vowels and some special characters . They normally dont take all vowel and consonant diacritic . \n\nBelow are the rules: \nGrapheme_root 0,1 : If you see any consonant or vowel diacritic then set as 0 . \nGrapheme root 2 and 9 they can have consonant diacritic 4 and vowel_diacritic 1  together . So , if you predict root 2 or 9 and consonant 4 , then surely vowel diacritic is 1 . and vice versa.\nGrapheme root: 3,4,5,6,7,8,11,12 etc can have no vowel diacritic and have only 0 or 1 as consonant diacritic \n\nHowever , as long as your model is not mistaking  a lot , the post processing this way can give marginal benefit only . I guess .",
      "votes": null
    },
    {
      "id": "733687",
      "postDate": "01/31/2020 12:34:41",
      "content": "<p>Certainly, it is rare that the model's predict result corresponds to the rules. If the number of Test data is large, we may be able to detect miss-predict by the rules, but in this competition, there are only 12 characters, so it's not expected.\nThat said, your rules must be useful for a practical Bengali character recognition model.</p>",
      "rawMarkdown": "Certainly, it is rare that the model's predict result corresponds to the rules. If the number of Test data is large, we may be able to detect miss-predict by the rules, but in this competition, there are only 12 characters, so it's not expected.\nThat said, your rules must be useful for a practical Bengali character recognition model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 731218,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "01/28/2020 12:51:26",
      "content": "<p>Thanks <a href=\"/phoenix9032\">@phoenix9032</a> , this is a great post! </p>\n\n<p>I was just searching online for just this kind of combination rules of Bengali Graphemes, i.e., compatibility of (root, vowel, consonant). </p>\n\n<p>I didn't look at what public kernels here are doing yet. But these facts seem to very useful to me\n- In test, it rules out incompatible prediction of (root, vowel, consonant) triples so that a compatible prediction can be chosen, if it has a slight lower predicted probability than an incompatible triple prediction. I thought using triple dictionary in train as this rule book, i.e., never predict grapheme unseen in train. However, the data description says <code>The test set includes some graphemes that do not exist in train but has no new grapheme components</code>. But, the rules like what you stated should always work unless the organizers are including insensible graphemes. Is this correct?\n- In train, we can include additional loss functions terms by adding heads that check if the predicted triple is compatible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 731393,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "01/28/2020 15:21:47",
          "content": "<p>You are right . There are combinations that are not possible in Bengali Language . Unless someone just write something out of the ordinary to test us . I will probably try to create more examples soon in future.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 732920,
      "author_name": "amanooo",
      "author_url": "",
      "post_date": "01/30/2020 13:04:55",
      "content": "<p>Hi Nirjhar Roy. Thank you for the valuable information. <br>\nI was also interested in the (root, vowel, consonant) combination.\nTrain data has 1292 combinations. For the time being, I tried to create logic to select the second softmax if predicted something other than these combinations. This logic slightly improved the accuracy for Train data, but unfortunately Test data did not have the corresponding prediction and was ineffective. Nirjhar's information may be useful for creating logic.</p>",
      "votes": null,
      "replies": [
        {
          "id": 732994,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "01/30/2020 14:40:24",
          "content": "<p>I have not been looking into this competition for few days because of Google Quest competition .However , as per my preliminary observation : It is not important what combination is not-possible . It is more important , if some combination is visually similar to other combination. \nIn that way , the person without the domain knowledge would be more in advantage , because , for years my eyes are trained to know the difference . \ne.g : I see my model make most mistakes in : <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff75b2db54bedb2a1821a83c0feae64b5%2F85.png?generation=1580394764389892&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F534bcf244aa86cd1dca5dc36e0b76923%2F62.png?generation=1580394768946394&amp;alt=media\" alt=\"\"></p>\n\n<p>Also , Grapheme_Root 0-12 are vowels and some special characters . They normally dont take all vowel and consonant diacritic . </p>\n\n<p>Below are the rules: \nGrapheme_root 0,1 : If you see any consonant or vowel diacritic then set as 0 . \nGrapheme root 2 and 9 they can have consonant diacritic 4 and vowel_diacritic 1  together . So , if you predict root 2 or 9 and consonant 4 , then surely vowel diacritic is 1 . and vice versa.\nGrapheme root: 3,4,5,6,7,8,11,12 etc can have no vowel diacritic and have only 0 or 1 as consonant diacritic </p>\n\n<p>However , as long as your model is not mistaking  a lot , the post processing this way can give marginal benefit only . I guess .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 733687,
          "author_name": "amanooo",
          "author_url": "",
          "post_date": "01/31/2020 12:34:41",
          "content": "<p>Certainly, it is rare that the model's predict result corresponds to the rules. If the number of Test data is large, we may be able to detect miss-predict by the rules, but in this competition, there are only 12 characters, so it's not expected.\nThat said, your rules must be useful for a practical Bengali character recognition model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "700851": "As a Bengali Speaking person , this competition is very close to my heart . I am extremely happy to see a lot of people from all-over the world trying to crack the Bengali Handwritten Graphemes .  At the end of this competition which would mean a good advancement in the area of our mother tongue and lot of different application to prosper the language  in keyboards , document processing , exam paper verification , many more areas . That is indeed a happy news for me . \nApart from this prelude , I wanted to jot down few things in this thread which I find interesting .\nSomeone with more linguistic knowledge can correct me . I am not sure if this can help in this comp as I have not yet fully joined this comp.\n\n1. Grapheme Root 2 to 12 are \"SwaraBarna\" or vowels . The vowel_diacritics (0-10) can not co-exist with these .(Not sure , if that makes any difference in the comp).\n\n2. Except for Consonant Diacritic 1 = ঁ , no other consonant diacritic also might come together with Grapheme root 2 to 12 .\n\n3. Remaining Grapheme Roots are Single Consonants and Grapheme root formed by making multiple Consonants come together e.g. ক = Grapheme Root 13 .  (Pronounce as Ka)  and ক্ক = Grapheme root 14 = KKa  like \"Trekking\"\n\n4.Consonant Diacritic 2 and 3 would not normally go with grapheme root : 159 to 167\n\nI will put together other things , as i look into the data more.",
    "731218": "Thanks @phoenix9032 , this is a great post! \n\nI was just searching online for just this kind of combination rules of Bengali Graphemes, i.e., compatibility of (root, vowel, consonant). \n\nI didn't look at what public kernels here are doing yet. But these facts seem to very useful to me\n- In test, it rules out incompatible prediction of (root, vowel, consonant) triples so that a compatible prediction can be chosen, if it has a slight lower predicted probability than an incompatible triple prediction. I thought using triple dictionary in train as this rule book, i.e., never predict grapheme unseen in train. However, the data description says `The test set includes some graphemes that do not exist in train but has no new grapheme components`. But, the rules like what you stated should always work unless the organizers are including insensible graphemes. Is this correct?\n- In train, we can include additional loss functions terms by adding heads that check if the predicted triple is compatible.",
    "731393": "You are right . There are combinations that are not possible in Bengali Language . Unless someone just write something out of the ordinary to test us . I will probably try to create more examples soon in future.",
    "732920": "Hi Nirjhar Roy. Thank you for the valuable information. <br>\nI was also interested in the (root, vowel, consonant) combination.\nTrain data has 1292 combinations. For the time being, I tried to create logic to select the second softmax if predicted something other than these combinations. This logic slightly improved the accuracy for Train data, but unfortunately Test data did not have the corresponding prediction and was ineffective. Nirjhar's information may be useful for creating logic.",
    "732994": "I have not been looking into this competition for few days because of Google Quest competition .However , as per my preliminary observation : It is not important what combination is not-possible . It is more important , if some combination is visually similar to other combination. \nIn that way , the person without the domain knowledge would be more in advantage , because , for years my eyes are trained to know the difference . \ne.g : I see my model make most mistakes in : ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff75b2db54bedb2a1821a83c0feae64b5%2F85.png?generation=1580394764389892&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F534bcf244aa86cd1dca5dc36e0b76923%2F62.png?generation=1580394768946394&amp;alt=media)\n\n\nAlso , Grapheme_Root 0-12 are vowels and some special characters . They normally dont take all vowel and consonant diacritic . \n\nBelow are the rules: \nGrapheme_root 0,1 : If you see any consonant or vowel diacritic then set as 0 . \nGrapheme root 2 and 9 they can have consonant diacritic 4 and vowel_diacritic 1  together . So , if you predict root 2 or 9 and consonant 4 , then surely vowel diacritic is 1 . and vice versa.\nGrapheme root: 3,4,5,6,7,8,11,12 etc can have no vowel diacritic and have only 0 or 1 as consonant diacritic \n\nHowever , as long as your model is not mistaking  a lot , the post processing this way can give marginal benefit only . I guess .",
    "733687": "Certainly, it is rare that the model's predict result corresponds to the rules. If the number of Test data is large, we may be able to detect miss-predict by the rules, but in this competition, there are only 12 characters, so it's not expected.\nThat said, your rules must be useful for a practical Bengali character recognition model."
  },
  "source": "meta"
}