{
  "id": 123002,
  "title": "Consonant Conjuncts vs Consonant Diacritics",
  "url": "/competitions/bengaliai-cv19/discussion/123002",
  "author_name": "Ahmed Imtiaz Humayun",
  "post_date": "2019-12-24T03:56:51.226000",
  "votes": 28,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Both are ways of <strong>adding a consonant with another consonant.</strong> In our labeling scheme, consonant conjuncts fall under the grapheme root umbrella while consonant diacritics are a label themselves. There is a fine difference between these which depends on what the second consonant in the sum looks like in the resulting grapheme. </p>\n\n<p>Generally, consonant conjuncts have both the consonants being joined, retain their original graphical form partially. Like, ক (k) + ট (t) = ক্ট (kt) where we do see both the consonants retaining it's original graphical form in the resultant.</p>\n\n<p>Consonant diacritics on the other hand, being diacritics, are expressed by demarcations that are completely different graphically compared to it's original form. Like, ক (k) + র (r) = ক্র (kr) where, while parts of the ক are retained, a completely different graphical demarcation  ্র represents র (r). Similarly, ম (m) + র (r) = ম্র, প (p) + র (r) = প্র (pr). What's interesting is that consonant diacritics can be added to consonant conjuncts too! Like, ক্ট (kt) + ্র (r in diacritic form) = ক্ট্র (ktr).</p>\n\n<p>This is why they are considered different both in Bengali grammar and our labeling scheme. Consonant diacritics always have the same graphical form across different grapheme combinations, similar to vowel diacritics. Graphical representation of consonants in conjuncts however, can vary between combinations, like ক্ট, ক্ত both are conjuncts with a second consonant being added to ক (k). </p>\n\n<p>Hope this clears it out a little, let me know if you have questions?</p>",
  "messages": [
    {
      "id": 701925,
      "postDate": "2019-12-24T03:56:51.227Z",
      "content": "<p>Both are ways of <strong>adding a consonant with another consonant.</strong> In our labeling scheme, consonant conjuncts fall under the grapheme root umbrella while consonant diacritics are a label themselves. There is a fine difference between these which depends on what the second consonant in the sum looks like in the resulting grapheme. </p>\n\n<p>Generally, consonant conjuncts have both the consonants being joined, retain their original graphical form partially. Like, ক (k) + ট (t) = ক্ট (kt) where we do see both the consonants retaining it's original graphical form in the resultant.</p>\n\n<p>Consonant diacritics on the other hand, being diacritics, are expressed by demarcations that are completely different graphically compared to it's original form. Like, ক (k) + র (r) = ক্র (kr) where, while parts of the ক are retained, a completely different graphical demarcation  ্র represents র (r). Similarly, ম (m) + র (r) = ম্র, প (p) + র (r) = প্র (pr). What's interesting is that consonant diacritics can be added to consonant conjuncts too! Like, ক্ট (kt) + ্র (r in diacritic form) = ক্ট্র (ktr).</p>\n\n<p>This is why they are considered different both in Bengali grammar and our labeling scheme. Consonant diacritics always have the same graphical form across different grapheme combinations, similar to vowel diacritics. Graphical representation of consonants in conjuncts however, can vary between combinations, like ক্ট, ক্ত both are conjuncts with a second consonant being added to ক (k). </p>\n\n<p>Hope this clears it out a little, let me know if you have questions?</p>",
      "rawMarkdown": "Both are ways of **adding a consonant with another consonant.** In our labeling scheme, consonant conjuncts fall under the grapheme root umbrella while consonant diacritics are a label themselves. There is a fine difference between these which depends on what the second consonant in the sum looks like in the resulting grapheme. \n\nGenerally, consonant conjuncts have both the consonants being joined, retain their original graphical form partially. Like, ক (k) + ট (t) = ক্ট (kt) where we do see both the consonants retaining it's original graphical form in the resultant.\n\nConsonant diacritics on the other hand, being diacritics, are expressed by demarcations that are completely different graphically compared to it's original form. Like, ক (k) + র (r) = ক্র (kr) where, while parts of the ক are retained, a completely different graphical demarcation  ্র represents র (r). Similarly, ম (m) + র (r) = ম্র, প (p) + র (r) = প্র (pr). What's interesting is that consonant diacritics can be added to consonant conjuncts too! Like, ক্ট (kt) + ্র (r in diacritic form) = ক্ট্র (ktr).\n\nThis is why they are considered different both in Bengali grammar and our labeling scheme. Consonant diacritics always have the same graphical form across different grapheme combinations, similar to vowel diacritics. Graphical representation of consonants in conjuncts however, can vary between combinations, like ক্ট, ক্ত both are conjuncts with a second consonant being added to ক (k). \n\nHope this clears it out a little, let me know if you have questions?",
      "votes": 27
    },
    {
      "id": 728098,
      "postDate": "2020-01-24T11:57:24.597Z",
      "content": "<p>\"র্দ্র\", \"র্ত্রী\", \"র্ত্রে\" exist in the training data. However, they might be classified incorrectly with provided label triplet of grapheme root, vowel diacritic, and consonant diacritic, which makes them as the same as \"র্দ\", \"র্তী\", \"র্তে\". The affect is that there's only 1292 classes of grapheme with regard to label triplets, while there's 1295 classes with regard to graphemes.</p>\n\n<p>For example, the provided label triplet: (72,0,2) for \"র্দ্র\" actually represents another grapheme:\"র্দ\". \"র্দ্র\" seems more likely to be \"র্দ\"(root)+ ্র (diacritic5) or \"দ্র\"(root)+\"র্\"(diacritic2), while neither \"র্দ\" nor \"দ্র\" is served as a grapheme root in this competition. Although \"দ\" is a grapheme root, it still requires two consonant diacritics(2+5) to correctly represent \"র্দ্র\".\nSame thoughts are applicable for  \"র্ত্রী\" and \"র্ত্রে\" as well. \"র্ত\" or \"ত্র\" might be new consonant conjunctions regardless of existed \"ত\".</p>\n\n<p>It is also possible that they are viewed as the same in Bengali graphemes, but I am not familiar with this language. <a href=\"/imtiazprio\">@imtiazprio</a> could you help clarify this problem? Thanks a lot.</p>",
      "rawMarkdown": "\"র্দ্র\", \"র্ত্রী\", \"র্ত্রে\" exist in the training data. However, they might be classified incorrectly with provided label triplet of grapheme root, vowel diacritic, and consonant diacritic, which makes them as the same as \"র্দ\", \"র্তী\", \"র্তে\". The affect is that there's only 1292 classes of grapheme with regard to label triplets, while there's 1295 classes with regard to graphemes.\n\nFor example, the provided label triplet: (72,0,2) for \"র্দ্র\" actually represents another grapheme:\"র্দ\". \"র্দ্র\" seems more likely to be \"র্দ\"(root)+ ্র (diacritic5) or \"দ্র\"(root)+\"র্\"(diacritic2), while neither \"র্দ\" nor \"দ্র\" is served as a grapheme root in this competition. Although \"দ\" is a grapheme root, it still requires two consonant diacritics(2+5) to correctly represent \"র্দ্র\".\nSame thoughts are applicable for  \"র্ত্রী\" and \"র্ত্রে\" as well. \"র্ত\" or \"ত্র\" might be new consonant conjunctions regardless of existed \"ত\".\n\nIt is also possible that they are viewed as the same in Bengali graphemes, but I am not familiar with this language. @imtiazprio could you help clarify this problem? Thanks a lot.",
      "votes": 3,
      "replies": [
        {
          "id": 729782,
          "postDate": "2020-01-26T16:25:30.913Z",
          "content": "<p>Hey AllenChangTW, \nThis was brought to our attention and they indeed are different graphemes with a common triplet label. Expect a fix to be announced in the next few days. The consonant diacritics 2 and 5 occuring together is pretty unique because they both represent the phonetic pronunciation of র (r). </p>\n\n<p>Thanks for the elaborate explanation though, it looks like you are pretty familiar with bengali now 😁</p>",
          "rawMarkdown": "Hey AllenChangTW, \nThis was brought to our attention and they indeed are different graphemes with a common triplet label. Expect a fix to be announced in the next few days. The consonant diacritics 2 and 5 occuring together is pretty unique because they both represent the phonetic pronunciation of র (r). \n\nThanks for the elaborate explanation though, it looks like you are pretty familiar with bengali now 😁",
          "votes": 2
        }
      ]
    },
    {
      "id": 707668,
      "postDate": "2020-01-01T09:15:57.473Z",
      "content": "<p>Thank you so much <a href=\"/phoenix9032\">@phoenix9032</a> for taking the time and clarifying this.</p>\n\n<p><a href=\"/hengck23\">@hengck23</a> Indeed, there are different graphemes present in the test set which are made of the basic building blocks in our training set- the basic building blocks being the grapheme roots and diacritics. If you recall from the slides, the 1295 that we have in the training set are 'Commonly used graphemes' from the context of everyday bengali news. There are words in Bengali which has graphemes which are quite uncommon, it gets more complicated for transliterations and proper nouns. The 1295 gives us a way to collect the most important graphemes into a dataset from the span of ~13k graphemes that can be created using our basic building blocks. Nevertheless, we have more graphemes than the 1295 in our test dataset. The target of the competition is therefore to, disentangle and recognize the individual building blocks.</p>\n\n<p>To specifically answer your questions:\n1. It won't be correct to say this, but they are all made of the same roots and diacritics.\n2. Not true. If you recall from the slides same graphemes written in different ways are called Allographs. We have allographs all across both training and test and even we don't have metadata on which is which since the data was crowdsourced. Allographs generally happen for consonant conjuncts discussed upstairs, diacritics almost always retain their graphical form.</p>",
      "rawMarkdown": "Thank you so much @phoenix9032 for taking the time and clarifying this.\n\n@hengck23 Indeed, there are different graphemes present in the test set which are made of the basic building blocks in our training set- the basic building blocks being the grapheme roots and diacritics. If you recall from the slides, the 1295 that we have in the training set are 'Commonly used graphemes' from the context of everyday bengali news. There are words in Bengali which has graphemes which are quite uncommon, it gets more complicated for transliterations and proper nouns. The 1295 gives us a way to collect the most important graphemes into a dataset from the span of ~13k graphemes that can be created using our basic building blocks. Nevertheless, we have more graphemes than the 1295 in our test dataset. The target of the competition is therefore to, disentangle and recognize the individual building blocks.\n\nTo specifically answer your questions:\n1. It won't be correct to say this, but they are all made of the same roots and diacritics.\n2. Not true. If you recall from the slides same graphemes written in different ways are called Allographs. We have allographs all across both training and test and even we don't have metadata on which is which since the data was crowdsourced. Allographs generally happen for consonant conjuncts discussed upstairs, diacritics almost always retain their graphical form.",
      "votes": 4
    },
    {
      "id": 704537,
      "postDate": "2019-12-27T15:41:09.837Z",
      "content": "<p>Good Information on the general methodology of demarcation .</p>",
      "rawMarkdown": "Good Information on the general methodology of demarcation .",
      "votes": 4
    },
    {
      "id": 706975,
      "postDate": "2019-12-31T05:24:05.247Z",
      "content": "<p>according to the slides \"<a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf</a>\" and \"<a href=\"https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt\">https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt</a>\"</p>\n\n<p>There are 1295 graphemes.</p>\n\n<p>in the  data description page(<a href=\"https://www.kaggle.com/c/bengaliai-cv19/data\">https://www.kaggle.com/c/bengaliai-cv19/data</a>), it says that \"the test set includes some graphemes that do not exist in train but has no new grapheme components\"</p>\n\n<p>My questions are:</p>\n\n<ol>\n<li>is it correct to say that both test and train only contains graphemes from the \"1295 graphemes set\"?</li>\n<li>if the above is true, the extra graphemes in the test set are due to different way of writing?</li>\n</ol>",
      "rawMarkdown": "according to the slides \"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\" and \"https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt\"\n\nThere are 1295 graphemes.\n\nin the  data description page(https://www.kaggle.com/c/bengaliai-cv19/data), it says that \"the test set includes some graphemes that do not exist in train but has no new grapheme components\"\n\nMy questions are:\n\n1. is it correct to say that both test and train only contains graphemes from the \"1295 graphemes set\"?\n2. if the above is true, the extra graphemes in the test set are due to different way of writing?",
      "votes": 2,
      "replies": [
        {
          "id": 707019,
          "postDate": "2019-12-31T06:56:40.807Z",
          "content": "<p>Dear Heng , \nWhile the organizers can clarify , I would like to tell my take in this . What i understand is that there are no new Grapheme root but there are new combination of Grapheme , Vowel Diacritics and Consonant Diacritics . Lets see the train data . </p>\n\n<p>For Grapheme Root = 28  we have 4 available Vowel Diacritic in Training data (1,2,4,9) and 1 consonant diacritic(4) .(barring the 0) .</p>\n\n<p>However there would be cases in real life and might be in test dataset where 28 can be associated with other vowel diacritic and consonant diacritic . Example Vowel Diacritic 3 and 10 . and consonant diacritic 2 </p>\n\n<p>Available in Train Set : \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fb326c80a8d183f2205df3d9515a83450%2FGla.png?generation=1577775130346467&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F3702f8d867249a5194d78f45148e0706%2Fglo.png?generation=1577775131075528&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F0250510a179749882e749bbccb2eecdd%2Fglaa.png?generation=1577775131535573&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F4475d1454b72494362bd5d31c15c9be6%2Fgli.png?generation=1577775133460897&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F620599a30ca204a7d06dccfff9a439cf%2Fglu.png?generation=1577775134060384&amp;alt=media\" alt=\"\"></p>\n\n<p>Not Available in Train Set : </p>\n\n<p>Grapheme root 28 and Vowel Diacritic 3 and10 for example . Which looks like this in my handwriting : (If you want to use external data I will write few more LOL . You can do Hard Sampling with my bad handwriting .\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F89b00ecedf04d15c8c569bcec0a2f07f%2FGlow%20(2\" alt=\"\">.jpg?generation=1577775565278367&amp;alt=media)</p>",
          "rawMarkdown": "Dear Heng , \nWhile the organizers can clarify , I would like to tell my take in this . What i understand is that there are no new Grapheme root but there are new combination of Grapheme , Vowel Diacritics and Consonant Diacritics . Lets see the train data . \n\nFor Grapheme Root = 28  we have 4 available Vowel Diacritic in Training data (1,2,4,9) and 1 consonant diacritic(4) .(barring the 0) .\n\nHowever there would be cases in real life and might be in test dataset where 28 can be associated with other vowel diacritic and consonant diacritic . Example Vowel Diacritic 3 and 10 . and consonant diacritic 2 \n\nAvailable in Train Set : \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fb326c80a8d183f2205df3d9515a83450%2FGla.png?generation=1577775130346467&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F3702f8d867249a5194d78f45148e0706%2Fglo.png?generation=1577775131075528&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F0250510a179749882e749bbccb2eecdd%2Fglaa.png?generation=1577775131535573&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F4475d1454b72494362bd5d31c15c9be6%2Fgli.png?generation=1577775133460897&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F620599a30ca204a7d06dccfff9a439cf%2Fglu.png?generation=1577775134060384&amp;alt=media)\n\nNot Available in Train Set : \n\nGrapheme root 28 and Vowel Diacritic 3 and10 for example . Which looks like this in my handwriting : (If you want to use external data I will write few more LOL . You can do Hard Sampling with my bad handwriting .\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F89b00ecedf04d15c8c569bcec0a2f07f%2FGlow%20(2).jpg?generation=1577775565278367&amp;alt=media)\n\n",
          "votes": 8
        }
      ]
    },
    {
      "id": 704913,
      "postDate": "2019-12-28T06:24:14.717Z",
      "content": "<p>On a side note, can anyone tell me if consonant diacritics are used in Devnagari?</p>",
      "rawMarkdown": "On a side note, can anyone tell me if consonant diacritics are used in Devnagari?",
      "replies": [
        {
          "id": 705016,
          "postDate": "2019-12-28T09:55:25.220Z",
          "content": "<p>Hey, I think they are, but I would want to check Devnagari grammar to confirm whether they use the term consonant diacritic, or classify them as a form of consonant conjunct. Consonant diacritics on the other hand are loosely termed as <em>Fola</em> (ফলা) in Bengali.</p>\n\n<p>I checked Devnagari unicode and they do have an element used to join different consonants; it is also analogous to Bengali. They also have diacritic like demarcations for words of similar pronunciation- like প্রিয় (priy) in Bengali is प्रिय (priy) in Devnagari which seems to have '्र' (r) as a consonant diacritic analogy for the Bengali  ্র (r). </p>\n\n<p>Nevertheless, I don't think Devnagari follows the exact same set of rules as discussed above for consonant diacritics in Bengali.</p>",
          "rawMarkdown": "Hey, I think they are, but I would want to check Devnagari grammar to confirm whether they use the term consonant diacritic, or classify them as a form of consonant conjunct. Consonant diacritics on the other hand are loosely termed as *Fola* (ফলা) in Bengali.\n\nI checked Devnagari unicode and they do have an element used to join different consonants; it is also analogous to Bengali. They also have diacritic like demarcations for words of similar pronunciation- like প্রিয় (priy) in Bengali is प्रिय (priy) in Devnagari which seems to have '्र' (r) as a consonant diacritic analogy for the Bengali  ্র (r). \n\nNevertheless, I don't think Devnagari follows the exact same set of rules as discussed above for consonant diacritics in Bengali.",
          "votes": 2
        }
      ]
    },
    {
      "id": 710752,
      "postDate": "2020-01-05T07:04:01.290Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 710865,
          "postDate": "2020-01-05T10:51:05.340Z",
          "content": "<p>if you dont mind , can you please tell the mistake ? is it the order?  or the graphical representation ? </p>",
          "rawMarkdown": "if you dont mind , can you please tell the mistake ? is it the order?  or the graphical representation ? "
        },
        {
          "id": 712602,
          "postDate": "2020-01-07T12:46:56.997Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 712665,
          "postDate": "2020-01-07T13:44:13.030Z",
          "content": "<p>In slide the representation is wrong visually . However if you look at pronunciation then it probably makes sense . Raf is used like in case of equivalent word : PARK. it's a short sound of R before a consonant . The R in BAR or BAROMETER won't have Raf . But Park, Bark , curl would have . </p>",
          "rawMarkdown": "In slide the representation is wrong visually . However if you look at pronunciation then it probably makes sense . Raf is used like in case of equivalent word : PARK. it's a short sound of R before a consonant . The R in BAR or BAROMETER won't have Raf . But Park, Bark , curl would have . ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 728098,
      "author_name": "AllenChangTW",
      "author_url": "",
      "post_date": "2020-01-24T11:57:24.597000",
      "content": "<p>\"র্দ্র\", \"র্ত্রী\", \"র্ত্রে\" exist in the training data. However, they might be classified incorrectly with provided label triplet of grapheme root, vowel diacritic, and consonant diacritic, which makes them as the same as \"র্দ\", \"র্তী\", \"র্তে\". The affect is that there's only 1292 classes of grapheme with regard to label triplets, while there's 1295 classes with regard to graphemes.</p>\n\n<p>For example, the provided label triplet: (72,0,2) for \"র্দ্র\" actually represents another grapheme:\"র্দ\". \"র্দ্র\" seems more likely to be \"র্দ\"(root)+ ্র (diacritic5) or \"দ্র\"(root)+\"র্\"(diacritic2), while neither \"র্দ\" nor \"দ্র\" is served as a grapheme root in this competition. Although \"দ\" is a grapheme root, it still requires two consonant diacritics(2+5) to correctly represent \"র্দ্র\".\nSame thoughts are applicable for  \"র্ত্রী\" and \"র্ত্রে\" as well. \"র্ত\" or \"ত্র\" might be new consonant conjunctions regardless of existed \"ত\".</p>\n\n<p>It is also possible that they are viewed as the same in Bengali graphemes, but I am not familiar with this language. <a href=\"/imtiazprio\">@imtiazprio</a> could you help clarify this problem? Thanks a lot.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 729782,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2020-01-26T16:25:30.913000",
          "content": "<p>Hey AllenChangTW, \nThis was brought to our attention and they indeed are different graphemes with a common triplet label. Expect a fix to be announced in the next few days. The consonant diacritics 2 and 5 occuring together is pretty unique because they both represent the phonetic pronunciation of র (r). </p>\n\n<p>Thanks for the elaborate explanation though, it looks like you are pretty familiar with bengali now 😁</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 707668,
      "author_name": "Ahmed Imtiaz Humayun",
      "author_url": "",
      "post_date": "2020-01-01T09:15:57.473000",
      "content": "<p>Thank you so much <a href=\"/phoenix9032\">@phoenix9032</a> for taking the time and clarifying this.</p>\n\n<p><a href=\"/hengck23\">@hengck23</a> Indeed, there are different graphemes present in the test set which are made of the basic building blocks in our training set- the basic building blocks being the grapheme roots and diacritics. If you recall from the slides, the 1295 that we have in the training set are 'Commonly used graphemes' from the context of everyday bengali news. There are words in Bengali which has graphemes which are quite uncommon, it gets more complicated for transliterations and proper nouns. The 1295 gives us a way to collect the most important graphemes into a dataset from the span of ~13k graphemes that can be created using our basic building blocks. Nevertheless, we have more graphemes than the 1295 in our test dataset. The target of the competition is therefore to, disentangle and recognize the individual building blocks.</p>\n\n<p>To specifically answer your questions:\n1. It won't be correct to say this, but they are all made of the same roots and diacritics.\n2. Not true. If you recall from the slides same graphemes written in different ways are called Allographs. We have allographs all across both training and test and even we don't have metadata on which is which since the data was crowdsourced. Allographs generally happen for consonant conjuncts discussed upstairs, diacritics almost always retain their graphical form.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 704537,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2019-12-27T15:41:09.837000",
      "content": "<p>Good Information on the general methodology of demarcation .</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 706975,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-12-31T05:24:05.247000",
      "content": "<p>according to the slides \"<a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf</a>\" and \"<a href=\"https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt\">https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt</a>\"</p>\n\n<p>There are 1295 graphemes.</p>\n\n<p>in the  data description page(<a href=\"https://www.kaggle.com/c/bengaliai-cv19/data\">https://www.kaggle.com/c/bengaliai-cv19/data</a>), it says that \"the test set includes some graphemes that do not exist in train but has no new grapheme components\"</p>\n\n<p>My questions are:</p>\n\n<ol>\n<li>is it correct to say that both test and train only contains graphemes from the \"1295 graphemes set\"?</li>\n<li>if the above is true, the extra graphemes in the test set are due to different way of writing?</li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 707019,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2019-12-31T06:56:40.807000",
          "content": "<p>Dear Heng , \nWhile the organizers can clarify , I would like to tell my take in this . What i understand is that there are no new Grapheme root but there are new combination of Grapheme , Vowel Diacritics and Consonant Diacritics . Lets see the train data . </p>\n\n<p>For Grapheme Root = 28  we have 4 available Vowel Diacritic in Training data (1,2,4,9) and 1 consonant diacritic(4) .(barring the 0) .</p>\n\n<p>However there would be cases in real life and might be in test dataset where 28 can be associated with other vowel diacritic and consonant diacritic . Example Vowel Diacritic 3 and 10 . and consonant diacritic 2 </p>\n\n<p>Available in Train Set : \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fb326c80a8d183f2205df3d9515a83450%2FGla.png?generation=1577775130346467&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F3702f8d867249a5194d78f45148e0706%2Fglo.png?generation=1577775131075528&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F0250510a179749882e749bbccb2eecdd%2Fglaa.png?generation=1577775131535573&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F4475d1454b72494362bd5d31c15c9be6%2Fgli.png?generation=1577775133460897&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F620599a30ca204a7d06dccfff9a439cf%2Fglu.png?generation=1577775134060384&amp;alt=media\" alt=\"\"></p>\n\n<p>Not Available in Train Set : </p>\n\n<p>Grapheme root 28 and Vowel Diacritic 3 and10 for example . Which looks like this in my handwriting : (If you want to use external data I will write few more LOL . You can do Hard Sampling with my bad handwriting .\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F89b00ecedf04d15c8c569bcec0a2f07f%2FGlow%20(2\" alt=\"\">.jpg?generation=1577775565278367&amp;alt=media)</p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 704913,
      "author_name": "ibraheemmoosa",
      "author_url": "",
      "post_date": "2019-12-28T06:24:14.717000",
      "content": "<p>On a side note, can anyone tell me if consonant diacritics are used in Devnagari?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 705016,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2019-12-28T09:55:25.220000",
          "content": "<p>Hey, I think they are, but I would want to check Devnagari grammar to confirm whether they use the term consonant diacritic, or classify them as a form of consonant conjunct. Consonant diacritics on the other hand are loosely termed as <em>Fola</em> (ফলা) in Bengali.</p>\n\n<p>I checked Devnagari unicode and they do have an element used to join different consonants; it is also analogous to Bengali. They also have diacritic like demarcations for words of similar pronunciation- like প্রিয় (priy) in Bengali is प्रिय (priy) in Devnagari which seems to have '्र' (r) as a consonant diacritic analogy for the Bengali  ্র (r). </p>\n\n<p>Nevertheless, I don't think Devnagari follows the exact same set of rules as discussed above for consonant diacritics in Bengali.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 710752,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-05T07:04:01.290000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 710865,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-05T10:51:05.340000",
          "content": "<p>if you dont mind , can you please tell the mistake ? is it the order?  or the graphical representation ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712602,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-07T12:46:56.997000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712665,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-07T13:44:13.030000",
          "content": "<p>In slide the representation is wrong visually . However if you look at pronunciation then it probably makes sense . Raf is used like in case of equivalent word : PARK. it's a short sound of R before a consonant . The R in BAR or BAROMETER won't have Raf . But Park, Bark , curl would have . </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "701925": "Both are ways of **adding a consonant with another consonant.** In our labeling scheme, consonant conjuncts fall under the grapheme root umbrella while consonant diacritics are a label themselves. There is a fine difference between these which depends on what the second consonant in the sum looks like in the resulting grapheme. \n\nGenerally, consonant conjuncts have both the consonants being joined, retain their original graphical form partially. Like, ক (k) + ট (t) = ক্ট (kt) where we do see both the consonants retaining it's original graphical form in the resultant.\n\nConsonant diacritics on the other hand, being diacritics, are expressed by demarcations that are completely different graphically compared to it's original form. Like, ক (k) + র (r) = ক্র (kr) where, while parts of the ক are retained, a completely different graphical demarcation  ্র represents র (r). Similarly, ম (m) + র (r) = ম্র, প (p) + র (r) = প্র (pr). What's interesting is that consonant diacritics can be added to consonant conjuncts too! Like, ক্ট (kt) + ্র (r in diacritic form) = ক্ট্র (ktr).\n\nThis is why they are considered different both in Bengali grammar and our labeling scheme. Consonant diacritics always have the same graphical form across different grapheme combinations, similar to vowel diacritics. Graphical representation of consonants in conjuncts however, can vary between combinations, like ক্ট, ক্ত both are conjuncts with a second consonant being added to ক (k). \n\nHope this clears it out a little, let me know if you have questions?",
    "728098": "\"র্দ্র\", \"র্ত্রী\", \"র্ত্রে\" exist in the training data. However, they might be classified incorrectly with provided label triplet of grapheme root, vowel diacritic, and consonant diacritic, which makes them as the same as \"র্দ\", \"র্তী\", \"র্তে\". The affect is that there's only 1292 classes of grapheme with regard to label triplets, while there's 1295 classes with regard to graphemes.\n\nFor example, the provided label triplet: (72,0,2) for \"র্দ্র\" actually represents another grapheme:\"র্দ\". \"র্দ্র\" seems more likely to be \"র্দ\"(root)+ ্র (diacritic5) or \"দ্র\"(root)+\"র্\"(diacritic2), while neither \"র্দ\" nor \"দ্র\" is served as a grapheme root in this competition. Although \"দ\" is a grapheme root, it still requires two consonant diacritics(2+5) to correctly represent \"র্দ্র\".\nSame thoughts are applicable for  \"র্ত্রী\" and \"র্ত্রে\" as well. \"র্ত\" or \"ত্র\" might be new consonant conjunctions regardless of existed \"ত\".\n\nIt is also possible that they are viewed as the same in Bengali graphemes, but I am not familiar with this language. @imtiazprio could you help clarify this problem? Thanks a lot.",
    "707668": "Thank you so much @phoenix9032 for taking the time and clarifying this.\n\n@hengck23 Indeed, there are different graphemes present in the test set which are made of the basic building blocks in our training set- the basic building blocks being the grapheme roots and diacritics. If you recall from the slides, the 1295 that we have in the training set are 'Commonly used graphemes' from the context of everyday bengali news. There are words in Bengali which has graphemes which are quite uncommon, it gets more complicated for transliterations and proper nouns. The 1295 gives us a way to collect the most important graphemes into a dataset from the span of ~13k graphemes that can be created using our basic building blocks. Nevertheless, we have more graphemes than the 1295 in our test dataset. The target of the competition is therefore to, disentangle and recognize the individual building blocks.\n\nTo specifically answer your questions:\n1. It won't be correct to say this, but they are all made of the same roots and diacritics.\n2. Not true. If you recall from the slides same graphemes written in different ways are called Allographs. We have allographs all across both training and test and even we don't have metadata on which is which since the data was crowdsourced. Allographs generally happen for consonant conjuncts discussed upstairs, diacritics almost always retain their graphical form.",
    "704537": "Good Information on the general methodology of demarcation .",
    "706975": "according to the slides \"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\" and \"https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt\"\n\nThere are 1295 graphemes.\n\nin the  data description page(https://www.kaggle.com/c/bengaliai-cv19/data), it says that \"the test set includes some graphemes that do not exist in train but has no new grapheme components\"\n\nMy questions are:\n\n1. is it correct to say that both test and train only contains graphemes from the \"1295 graphemes set\"?\n2. if the above is true, the extra graphemes in the test set are due to different way of writing?",
    "704913": "On a side note, can anyone tell me if consonant diacritics are used in Devnagari?",
    "710752": ""
  }
}