{
  "id": 133735,
  "title": "Bengali AI Dataset - EDA Grapheme Combinations",
  "url": "/competitions/bengaliai-cv19/discussion/133735",
  "author_name": "",
  "post_date": "2020-03-04T02:53:34.102720800Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I have just started the Bengali AI competition and the first place I started was with an Exploratory Data Analysis on the training dataset.</p>\n\n<p>Whilst there are <code>168 * 11 * 7 = 12936</code> theoretical possible combinations, the dataset only contains <code>1295</code> unique roots/vowels/consonant combinations. Assuming that the training dataset is representative of common usage, then certian combinations may never (or rarely) be used in practice). </p>\n\n<p>I have an Unconfirmed Theory that the physics of the human mouth may make some combinations unpronouncable. If there are any native speakers of Bengali on the forum, I would be intrested in an explanation of how your language works both in theory and practice.</p>\n\n<h2>Findings</h2>\n\n<h3>Vowel / Consonant Combinations</h3>\n\n<ul>\n<li>Vowel #0 and Consonant #0 combine with everything</li>\n<li>Vowels #3, #5, #6, #8 have limited combinations with Consonants</li>\n<li>Consonant #3 is never combined except with Vowel #0</li>\n<li>Consonant #6 only combineds with Vowels #0 and #1</li>\n</ul>\n\n<h3>Grapheme Root Combinations</h3>\n\n<ul>\n<li>Vowel #0 and Consonant #0 combine with (nearly) everything</li>\n<li>ALL Roots combine with some Consonant #0</li>\n<li>Several Roots do NOT combine with Vowel #0 = [26, 28, 33, 34, 73, 82, 108, 114, 126, 152, 157, 158, 163]</li>\n<li>Several Roots do combine with ALL Vowels = [13, 23, 64, 72, 79, 81, 96, 107, 113, 115, 133, 147]}</li>\n<li>Only Root #107 combines with ALL Consonants</li>\n</ul>\n\n<h3>Combination Matrices</h3>\n\n<ul>\n<li>I wrote a combination_matrix() function to visualize the full list of which Grapheme Roots combine with which Vowels and Consonant Diacritics</li>\n</ul>\n\n<h2>Sanity Checking = Found Dataset BUG!</h2>\n\n<p>This combination_matrix lists 1292 unique grapheme combinations, which is 3 less than the 1295 unique graphemes listed in the training dataset. Something is WRONG!</p>\n\n<p>Found a discrepency BUG is in the dataset. The following root/vowel/consonant keys have multiple unicode graphemes renderings!</p>\n\n<p>{'64-3-2': ['র্তী', 'র্ত্রী'], '64-7-2': ['র্তে', 'র্ত্রে'], '72-0-2': ['র্দ্র', 'র্দ']}</p>\n\n<h3>EDA Notebook</h3>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/jamesmcguigan/bengali-ai-dataset-eda-grapheme-combinations\">https://www.kaggle.com/jamesmcguigan/bengali-ai-dataset-eda-grapheme-combinations</a></li>\n</ul>",
  "messages": [
    {
      "id": "763011",
      "postDate": "03/04/2020 02:53:34",
      "content": "<p>I have just started the Bengali AI competition and the first place I started was with an Exploratory Data Analysis on the training dataset.</p>\n\n<p>Whilst there are <code>168 * 11 * 7 = 12936</code> theoretical possible combinations, the dataset only contains <code>1295</code> unique roots/vowels/consonant combinations. Assuming that the training dataset is representative of common usage, then certian combinations may never (or rarely) be used in practice). </p>\n\n<p>I have an Unconfirmed Theory that the physics of the human mouth may make some combinations unpronouncable. If there are any native speakers of Bengali on the forum, I would be intrested in an explanation of how your language works both in theory and practice.</p>\n\n<h2>Findings</h2>\n\n<h3>Vowel / Consonant Combinations</h3>\n\n<ul>\n<li>Vowel #0 and Consonant #0 combine with everything</li>\n<li>Vowels #3, #5, #6, #8 have limited combinations with Consonants</li>\n<li>Consonant #3 is never combined except with Vowel #0</li>\n<li>Consonant #6 only combineds with Vowels #0 and #1</li>\n</ul>\n\n<h3>Grapheme Root Combinations</h3>\n\n<ul>\n<li>Vowel #0 and Consonant #0 combine with (nearly) everything</li>\n<li>ALL Roots combine with some Consonant #0</li>\n<li>Several Roots do NOT combine with Vowel #0 = [26, 28, 33, 34, 73, 82, 108, 114, 126, 152, 157, 158, 163]</li>\n<li>Several Roots do combine with ALL Vowels = [13, 23, 64, 72, 79, 81, 96, 107, 113, 115, 133, 147]}</li>\n<li>Only Root #107 combines with ALL Consonants</li>\n</ul>\n\n<h3>Combination Matrices</h3>\n\n<ul>\n<li>I wrote a combination_matrix() function to visualize the full list of which Grapheme Roots combine with which Vowels and Consonant Diacritics</li>\n</ul>\n\n<h2>Sanity Checking = Found Dataset BUG!</h2>\n\n<p>This combination_matrix lists 1292 unique grapheme combinations, which is 3 less than the 1295 unique graphemes listed in the training dataset. Something is WRONG!</p>\n\n<p>Found a discrepency BUG is in the dataset. The following root/vowel/consonant keys have multiple unicode graphemes renderings!</p>\n\n<p>{'64-3-2': ['র্তী', 'র্ত্রী'], '64-7-2': ['র্তে', 'র্ত্রে'], '72-0-2': ['র্দ্র', 'র্দ']}</p>\n\n<h3>EDA Notebook</h3>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/jamesmcguigan/bengali-ai-dataset-eda-grapheme-combinations\">https://www.kaggle.com/jamesmcguigan/bengali-ai-dataset-eda-grapheme-combinations</a></li>\n</ul>",
      "rawMarkdown": "I have just started the Bengali AI competition and the first place I started was with an Exploratory Data Analysis on the training dataset.\n\nWhilst there are `168 * 11 * 7 = 12936` theoretical possible combinations, the dataset only contains `1295` unique roots/vowels/consonant combinations. Assuming that the training dataset is representative of common usage, then certian combinations may never (or rarely) be used in practice). \n\nI have an Unconfirmed Theory that the physics of the human mouth may make some combinations unpronouncable. If there are any native speakers of Bengali on the forum, I would be intrested in an explanation of how your language works both in theory and practice.\n\n## Findings\n\n### Vowel / Consonant Combinations\n\n- Vowel #0 and Consonant #0 combine with everything\n- Vowels #3, #5, #6, #8 have limited combinations with Consonants\n- Consonant #3 is never combined except with Vowel #0\n- Consonant #6 only combineds with Vowels #0 and #1\n\n### Grapheme Root Combinations\n\n- Vowel #0 and Consonant #0 combine with (nearly) everything\n- ALL Roots combine with some Consonant #0\n- Several Roots do NOT combine with Vowel #0 = [26, 28, 33, 34, 73, 82, 108, 114, 126, 152, 157, 158, 163]\n- Several Roots do combine with ALL Vowels = [13, 23, 64, 72, 79, 81, 96, 107, 113, 115, 133, 147]}\n- Only Root #107 combines with ALL Consonants\n\n### Combination Matrices\n\n- I wrote a combination_matrix() function to visualize the full list of which Grapheme Roots combine with which Vowels and Consonant Diacritics\n\n## Sanity Checking = Found Dataset BUG!\n\nThis combination_matrix lists 1292 unique grapheme combinations, which is 3 less than the 1295 unique graphemes listed in the training dataset. Something is WRONG!\n\nFound a discrepency BUG is in the dataset. The following root/vowel/consonant keys have multiple unicode graphemes renderings!\n\n{'64-3-2': ['র্তী', 'র্ত্রী'], '64-7-2': ['র্তে', 'র্ত্রে'], '72-0-2': ['র্দ্র', 'র্দ']}\n\n### EDA Notebook\n- https://www.kaggle.com/jamesmcguigan/bengali-ai-dataset-eda-grapheme-combinations",
      "votes": null
    },
    {
      "id": "763133",
      "postDate": "03/04/2020 07:07:09",
      "content": "<p>Excellent notebook.</p>\n\n<p>Also The dset bug you've mentioned has been reported in two threads.</p>",
      "rawMarkdown": "Excellent notebook.\n\nAlso The dset bug you've mentioned has been reported in two threads.",
      "votes": null
    },
    {
      "id": "763288",
      "postDate": "03/04/2020 10:24:54",
      "content": "<p>Yes, you are right about your assumption that the most common combination of graphemes are represented in the training dataset and more  description of the concept you can find in the following <a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">slides</a></p>",
      "rawMarkdown": "Yes, you are right about your assumption that the most common combination of graphemes are represented in the training dataset and more  description of the concept you can find in the following [slides](https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf)",
      "votes": null
    },
    {
      "id": "763562",
      "postDate": "03/04/2020 15:49:57",
      "content": "<p>Interesting...</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F3aacda044ccde7227ebfa6ec7eecdb0c%2F2020-03-05%200.48.39.png?generation=1583336940871096&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Interesting...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F3aacda044ccde7227ebfa6ec7eecdb0c%2F2020-03-05%200.48.39.png?generation=1583336940871096&amp;alt=media)",
      "votes": null
    },
    {
      "id": "763564",
      "postDate": "03/04/2020 15:51:53",
      "content": "<p>Related posts:</p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127662\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127662</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123002\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123002</a></p>",
      "rawMarkdown": "Related posts:\n\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/127662\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123859\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123002",
      "votes": null
    },
    {
      "id": "763698",
      "postDate": "03/04/2020 18:55:31",
      "content": "<p>They are good observations. As <a href=\"/ren4yu\">@ren4yu</a> mentioned there have been similar discussions in the forum around it. The question is though, how to let your model know about this information? I have an idea of doing it using factored graphs but haven't had time to try it out. Nonetheless, the recall score is already so high that I honestly don't think implementing this kinda stuff helps that much. </p>",
      "rawMarkdown": "They are good observations. As @ren4yu mentioned there have been similar discussions in the forum around it. The question is though, how to let your model know about this information? I have an idea of doing it using factored graphs but haven't had time to try it out. Nonetheless, the recall score is already so high that I honestly don't think implementing this kinda stuff helps that much.",
      "votes": null
    },
    {
      "id": "763753",
      "postDate": "03/04/2020 20:17:59",
      "content": "<p>That is curious, it appears that graphemes in Bengali may have different regional representations that the unicode developers have chosen to encode as having duplicate diacritics.</p>\n\n<p>In English we have 0 and Ø, which have two different Unicode representations for zero.</p>\n\n<p>In old English, we have Æ which is the closest example I can think of for mixing diacritics within a grapheme.</p>",
      "rawMarkdown": "That is curious, it appears that graphemes in Bengali may have different regional representations that the unicode developers have chosen to encode as having duplicate diacritics.\n\nIn English we have 0 and Ø, which have two different Unicode representations for zero.\n\nIn old English, we have Æ which is the closest example I can think of for mixing diacritics within a grapheme.",
      "votes": null
    },
    {
      "id": "764356",
      "postDate": "03/05/2020 11:37:30",
      "content": "<p>I need to read up on factored graphs. </p>\n\n<p>However a crude implementation would be to write a custom tensorflow activation layer function. There are three one-hot-encoded output variables. Run softmax on each output group (first stage prediction).</p>\n\n<p>Multiply each output layer group by the probability from the combination matrix. Use the root to calculate probability of vowel/consonant, then use vowel to update root/consonant probability, and consonant to update vowel/root. Any combinations with a probability of 0 will be hard excluded. Rerun softmax on your modified output layers, and compute loss and back propagation as normal.</p>",
      "rawMarkdown": "I need to read up on factored graphs. \n\nHowever a crude implementation would be to write a custom tensorflow activation layer function. There are three one-hot-encoded output variables. Run softmax on each output group (first stage prediction).\n\nMultiply each output layer group by the probability from the combination matrix. Use the root to calculate probability of vowel/consonant, then use vowel to update root/consonant probability, and consonant to update vowel/root. Any combinations with a probability of 0 will be hard excluded. Rerun softmax on your modified output layers, and compute loss and back propagation as normal.",
      "votes": null
    },
    {
      "id": "766195",
      "postDate": "03/07/2020 20:43:01",
      "content": "<p>Following on from <a href=\"/ren4yu\">@ren4yu</a> discovery of the bengali multibyte unicode encoding, I have discovered that the suposed root_graphemes are in fact composed of smaller set of 62 base_graphemes.</p>\n\n<p>Using a set analyis, I have identified the set of unicode base_graphemes for each of the vowels, consonants and roots.</p>\n\n<p>I have also performed a full visualization of the bengali alphabet</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet\">https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet</a></li>\n</ul>",
      "rawMarkdown": "Following on from @ren4yu discovery of the bengali multibyte unicode encoding, I have discovered that the suposed root\\_graphemes are in fact composed of smaller set of 62 base\\_graphemes.\n\nUsing a set analyis, I have identified the set of unicode base\\_graphemes for each of the vowels, consonants and roots.\n\nI have also performed a full visualization of the bengali alphabet\n\n- https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet",
      "votes": null
    },
    {
      "id": "766294",
      "postDate": "03/08/2020 00:28:04",
      "content": "<p>Nice Notebook</p>",
      "rawMarkdown": "Nice Notebook",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 763133,
      "author_name": "authman",
      "author_url": "",
      "post_date": "03/04/2020 07:07:09",
      "content": "<p>Excellent notebook.</p>\n\n<p>Also The dset bug you've mentioned has been reported in two threads.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 763288,
      "author_name": "onslaught",
      "author_url": "",
      "post_date": "03/04/2020 10:24:54",
      "content": "<p>Yes, you are right about your assumption that the most common combination of graphemes are represented in the training dataset and more  description of the concept you can find in the following <a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">slides</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 763562,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "03/04/2020 15:49:57",
      "content": "<p>Interesting...</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F3aacda044ccde7227ebfa6ec7eecdb0c%2F2020-03-05%200.48.39.png?generation=1583336940871096&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 763753,
          "author_name": "jamesmcguigan",
          "author_url": "",
          "post_date": "03/04/2020 20:17:59",
          "content": "<p>That is curious, it appears that graphemes in Bengali may have different regional representations that the unicode developers have chosen to encode as having duplicate diacritics.</p>\n\n<p>In English we have 0 and Ø, which have two different Unicode representations for zero.</p>\n\n<p>In old English, we have Æ which is the closest example I can think of for mixing diacritics within a grapheme.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763564,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "03/04/2020 15:51:53",
      "content": "<p>Related posts:</p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127662\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127662</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123002\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123002</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 763698,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "03/04/2020 18:55:31",
      "content": "<p>They are good observations. As <a href=\"/ren4yu\">@ren4yu</a> mentioned there have been similar discussions in the forum around it. The question is though, how to let your model know about this information? I have an idea of doing it using factored graphs but haven't had time to try it out. Nonetheless, the recall score is already so high that I honestly don't think implementing this kinda stuff helps that much. </p>",
      "votes": null,
      "replies": [
        {
          "id": 764356,
          "author_name": "jamesmcguigan",
          "author_url": "",
          "post_date": "03/05/2020 11:37:30",
          "content": "<p>I need to read up on factored graphs. </p>\n\n<p>However a crude implementation would be to write a custom tensorflow activation layer function. There are three one-hot-encoded output variables. Run softmax on each output group (first stage prediction).</p>\n\n<p>Multiply each output layer group by the probability from the combination matrix. Use the root to calculate probability of vowel/consonant, then use vowel to update root/consonant probability, and consonant to update vowel/root. Any combinations with a probability of 0 will be hard excluded. Rerun softmax on your modified output layers, and compute loss and back propagation as normal.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 766195,
      "author_name": "jamesmcguigan",
      "author_url": "",
      "post_date": "03/07/2020 20:43:01",
      "content": "<p>Following on from <a href=\"/ren4yu\">@ren4yu</a> discovery of the bengali multibyte unicode encoding, I have discovered that the suposed root_graphemes are in fact composed of smaller set of 62 base_graphemes.</p>\n\n<p>Using a set analyis, I have identified the set of unicode base_graphemes for each of the vowels, consonants and roots.</p>\n\n<p>I have also performed a full visualization of the bengali alphabet</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet\">https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 766294,
      "author_name": "mikofarinucla",
      "author_url": "",
      "post_date": "03/08/2020 00:28:04",
      "content": "<p>Nice Notebook</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "763011": "I have just started the Bengali AI competition and the first place I started was with an Exploratory Data Analysis on the training dataset.\n\nWhilst there are `168 * 11 * 7 = 12936` theoretical possible combinations, the dataset only contains `1295` unique roots/vowels/consonant combinations. Assuming that the training dataset is representative of common usage, then certian combinations may never (or rarely) be used in practice). \n\nI have an Unconfirmed Theory that the physics of the human mouth may make some combinations unpronouncable. If there are any native speakers of Bengali on the forum, I would be intrested in an explanation of how your language works both in theory and practice.\n\n## Findings\n\n### Vowel / Consonant Combinations\n\n- Vowel #0 and Consonant #0 combine with everything\n- Vowels #3, #5, #6, #8 have limited combinations with Consonants\n- Consonant #3 is never combined except with Vowel #0\n- Consonant #6 only combineds with Vowels #0 and #1\n\n### Grapheme Root Combinations\n\n- Vowel #0 and Consonant #0 combine with (nearly) everything\n- ALL Roots combine with some Consonant #0\n- Several Roots do NOT combine with Vowel #0 = [26, 28, 33, 34, 73, 82, 108, 114, 126, 152, 157, 158, 163]\n- Several Roots do combine with ALL Vowels = [13, 23, 64, 72, 79, 81, 96, 107, 113, 115, 133, 147]}\n- Only Root #107 combines with ALL Consonants\n\n### Combination Matrices\n\n- I wrote a combination_matrix() function to visualize the full list of which Grapheme Roots combine with which Vowels and Consonant Diacritics\n\n## Sanity Checking = Found Dataset BUG!\n\nThis combination_matrix lists 1292 unique grapheme combinations, which is 3 less than the 1295 unique graphemes listed in the training dataset. Something is WRONG!\n\nFound a discrepency BUG is in the dataset. The following root/vowel/consonant keys have multiple unicode graphemes renderings!\n\n{'64-3-2': ['র্তী', 'র্ত্রী'], '64-7-2': ['র্তে', 'র্ত্রে'], '72-0-2': ['র্দ্র', 'র্দ']}\n\n### EDA Notebook\n- https://www.kaggle.com/jamesmcguigan/bengali-ai-dataset-eda-grapheme-combinations",
    "763133": "Excellent notebook.\n\nAlso The dset bug you've mentioned has been reported in two threads.",
    "763288": "Yes, you are right about your assumption that the most common combination of graphemes are represented in the training dataset and more  description of the concept you can find in the following [slides](https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf)",
    "763562": "Interesting...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F3aacda044ccde7227ebfa6ec7eecdb0c%2F2020-03-05%200.48.39.png?generation=1583336940871096&amp;alt=media)",
    "763564": "Related posts:\n\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/127662\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123859\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123002",
    "763698": "They are good observations. As @ren4yu mentioned there have been similar discussions in the forum around it. The question is though, how to let your model know about this information? I have an idea of doing it using factored graphs but haven't had time to try it out. Nonetheless, the recall score is already so high that I honestly don't think implementing this kinda stuff helps that much.",
    "763753": "That is curious, it appears that graphemes in Bengali may have different regional representations that the unicode developers have chosen to encode as having duplicate diacritics.\n\nIn English we have 0 and Ø, which have two different Unicode representations for zero.\n\nIn old English, we have Æ which is the closest example I can think of for mixing diacritics within a grapheme.",
    "764356": "I need to read up on factored graphs. \n\nHowever a crude implementation would be to write a custom tensorflow activation layer function. There are three one-hot-encoded output variables. Run softmax on each output group (first stage prediction).\n\nMultiply each output layer group by the probability from the combination matrix. Use the root to calculate probability of vowel/consonant, then use vowel to update root/consonant probability, and consonant to update vowel/root. Any combinations with a probability of 0 will be hard excluded. Rerun softmax on your modified output layers, and compute loss and back propagation as normal.",
    "766195": "Following on from @ren4yu discovery of the bengali multibyte unicode encoding, I have discovered that the suposed root\\_graphemes are in fact composed of smaller set of 62 base\\_graphemes.\n\nUsing a set analyis, I have identified the set of unicode base\\_graphemes for each of the vowels, consonants and roots.\n\nI have also performed a full visualization of the bengali alphabet\n\n- https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet",
    "766294": "Nice Notebook"
  },
  "source": "meta"
}