{
  "id": 127662,
  "title": "I'm confused by Bengali Grapheme",
  "url": "/competitions/bengaliai-cv19/discussion/127662",
  "author_name": "K.Amano",
  "post_date": "2020-01-25T14:31:08.313000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Are there multiple graphemes for the same combination of <code>grapheme_root</code>, <code>vowel_diacritic</code>, and <code>consonant_diacritic</code>?</p>\n\n<p>For example, in train.csv\n<code>\nimage_id=378  :  (64, 3, 2)   :  র্তী\nimage_id=1482 :  (64, 3, 2)   :  র্ত্রী\n</code>\nThere are <strong>1295</strong> types of grapheme in train.csv, but there are only <strong>1292</strong> combinations.\nPlease tell me about Bengali Grapheme.</p>",
  "messages": [
    {
      "id": 730068,
      "postDate": "2020-01-27T04:00:34.220Z",
      "content": "<p>I've reported this problem and got a response from one of competetion hosts saying that a fix is expected recently.\nJust check this link: (hope it help)\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123002#728098\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123002#728098</a></p>",
      "rawMarkdown": "I've reported this problem and got a response from one of competetion hosts saying that a fix is expected recently.\nJust check this link: (hope it help)\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123002#728098",
      "votes": 1,
      "replies": [
        {
          "id": 730425,
          "postDate": "2020-01-27T13:24:23.183Z",
          "content": "<p>Hi AllenChangTW.\nThank you for the information. I was missing valuable information.<br>\nBengali characters are very complicated, but 漢字(kanji) used in Japan and China is also a character that combines 'hen', 'kanmuri', 'tsukuri', etc. It has similarities with Bengali characters and has a close affinity.</p>",
          "rawMarkdown": "Hi AllenChangTW.\nThank you for the information. I was missing valuable information.<br>\nBengali characters are very complicated, but 漢字(kanji) used in Japan and China is also a character that combines 'hen', 'kanmuri', 'tsukuri', etc. It has similarities with Bengali characters and has a close affinity."
        }
      ]
    },
    {
      "id": 729028,
      "postDate": "2020-01-25T15:55:47.070Z",
      "content": "<p>You are right there are 1295 distinct graphemes but only 1291 combinations of the 3 components. \nBut you can notice that all the 3 components are kind of the same, just a little different shape? Maybe there was some sort of a change in the fonts they used while preparing the dataset.</p>\n\n<p>|      | image_id   |   grapheme_root |   vowel_diacritic |   consonant_diacritic | grapheme   |    mix |\n|-----:|:-----------|----------------:|------------------:|----------------------:|:-----------|-------:|\n|  387 | Train_387  |              64 |                 3 |                     2 | র্তী        | 64_3_2 |\n| 1482 | Train_1482 |              64 |                 3 |                     2 | র্ত্রী       | 64_3_2 |\n|  577 | Train_577  |              64 |                 7 |                     2 | র্তে        | 64_7_2 |\n| 1874 | Train_1874 |              64 |                 7 |                     2 | র্ত্রে       | 64_7_2 |\n|  532 | Train_532  |              72 |                 0 |                     2 | র্দ্র        | 72_0_2 |\n| 1117 | Train_1117 |              72 |                 0 |                     2 | র্দ         | 72_0_2 |</p>\n\n<p>|     | component_type      |   label | component   |\n|----:|:--------------------|--------:|:------------|\n| 64 | grapheme_root    |      64 | ত           |\n| 72 | grapheme_root    |      72 | দ           |\n| 171 | vowel_diacritic     |       3 | ী           |\n| 175 | vowel_diacritic     |       7 | ে           |\n| 168 | vowel_diacritic     |       0 | 0           |\n| 181 | consonant_diacritic |       2 | র্           |</p>",
      "rawMarkdown": "You are right there are 1295 distinct graphemes but only 1291 combinations of the 3 components. \nBut you can notice that all the 3 components are kind of the same, just a little different shape? Maybe there was some sort of a change in the fonts they used while preparing the dataset.\n\n|      | image_id   |   grapheme_root |   vowel_diacritic |   consonant_diacritic | grapheme   |    mix |\n|-----:|:-----------|----------------:|------------------:|----------------------:|:-----------|-------:|\n|  387 | Train_387  |              64 |                 3 |                     2 | র্তী        | 64_3_2 |\n| 1482 | Train_1482 |              64 |                 3 |                     2 | র্ত্রী       | 64_3_2 |\n|  577 | Train_577  |              64 |                 7 |                     2 | র্তে        | 64_7_2 |\n| 1874 | Train_1874 |              64 |                 7 |                     2 | র্ত্রে       | 64_7_2 |\n|  532 | Train_532  |              72 |                 0 |                     2 | র্দ্র        | 72_0_2 |\n| 1117 | Train_1117 |              72 |                 0 |                     2 | র্দ         | 72_0_2 |\n\n|     | component_type      |   label | component   |\n|----:|:--------------------|--------:|:------------|\n| 64 | grapheme_root    |      64 | ত           |\n| 72 | grapheme_root    |      72 | দ           |\n| 171 | vowel_diacritic     |       3 | ী           |\n| 175 | vowel_diacritic     |       7 | ে           |\n| 168 | vowel_diacritic     |       0 | 0           |\n| 181 | consonant_diacritic |       2 | র্           |\n",
      "votes": 1
    },
    {
      "id": 728979,
      "postDate": "2020-01-25T14:31:08.313Z",
      "content": "<p>Are there multiple graphemes for the same combination of <code>grapheme_root</code>, <code>vowel_diacritic</code>, and <code>consonant_diacritic</code>?</p>\n\n<p>For example, in train.csv\n<code>\nimage_id=378  :  (64, 3, 2)   :  র্তী\nimage_id=1482 :  (64, 3, 2)   :  র্ত্রী\n</code>\nThere are <strong>1295</strong> types of grapheme in train.csv, but there are only <strong>1292</strong> combinations.\nPlease tell me about Bengali Grapheme.</p>",
      "rawMarkdown": "Are there multiple graphemes for the same combination of `grapheme_root`, `vowel_diacritic`, and `consonant_diacritic`?\n\nFor example, in train.csv\n```\nimage_id=378  :  (64, 3, 2)   :  র্তী\nimage_id=1482 :  (64, 3, 2)   :  র্ত্রী\n```\nThere are **1295** types of grapheme in train.csv, but there are only **1292** combinations.\nPlease tell me about Bengali Grapheme.",
      "votes": 2
    },
    {
      "id": 729058,
      "postDate": "2020-01-25T17:14:34.183Z",
      "content": "<p>Hi Datta!\nর্তী and র্ত্রী are different typefaces, but they are the same character.\nThank you very much.</p>",
      "rawMarkdown": "Hi Datta!\nর্তী and র্ত্রী are different typefaces, but they are the same character.\nThank you very much.",
      "replies": [
        {
          "id": 730603,
          "postDate": "2020-01-27T17:13:29.823Z",
          "content": "<p>I thought so too, but have a look here. <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a></p>",
          "rawMarkdown": "I thought so too, but have a look here. https://www.kaggle.com/c/bengaliai-cv19/discussion/123859"
        }
      ]
    },
    {
      "id": 729000,
      "postDate": "2020-01-25T15:04:04.910Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 730068,
      "author_name": "AllenChangTW",
      "author_url": "",
      "post_date": "2020-01-27T04:00:34.220000",
      "content": "<p>I've reported this problem and got a response from one of competetion hosts saying that a fix is expected recently.\nJust check this link: (hope it help)\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123002#728098\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123002#728098</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 730425,
          "author_name": "K.Amano",
          "author_url": "",
          "post_date": "2020-01-27T13:24:23.183000",
          "content": "<p>Hi AllenChangTW.\nThank you for the information. I was missing valuable information.<br>\nBengali characters are very complicated, but 漢字(kanji) used in Japan and China is also a character that combines 'hen', 'kanmuri', 'tsukuri', etc. It has similarities with Bengali characters and has a close affinity.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 729028,
      "author_name": "datta",
      "author_url": "",
      "post_date": "2020-01-25T15:55:47.070000",
      "content": "<p>You are right there are 1295 distinct graphemes but only 1291 combinations of the 3 components. \nBut you can notice that all the 3 components are kind of the same, just a little different shape? Maybe there was some sort of a change in the fonts they used while preparing the dataset.</p>\n\n<p>|      | image_id   |   grapheme_root |   vowel_diacritic |   consonant_diacritic | grapheme   |    mix |\n|-----:|:-----------|----------------:|------------------:|----------------------:|:-----------|-------:|\n|  387 | Train_387  |              64 |                 3 |                     2 | র্তী        | 64_3_2 |\n| 1482 | Train_1482 |              64 |                 3 |                     2 | র্ত্রী       | 64_3_2 |\n|  577 | Train_577  |              64 |                 7 |                     2 | র্তে        | 64_7_2 |\n| 1874 | Train_1874 |              64 |                 7 |                     2 | র্ত্রে       | 64_7_2 |\n|  532 | Train_532  |              72 |                 0 |                     2 | র্দ্র        | 72_0_2 |\n| 1117 | Train_1117 |              72 |                 0 |                     2 | র্দ         | 72_0_2 |</p>\n\n<p>|     | component_type      |   label | component   |\n|----:|:--------------------|--------:|:------------|\n| 64 | grapheme_root    |      64 | ত           |\n| 72 | grapheme_root    |      72 | দ           |\n| 171 | vowel_diacritic     |       3 | ী           |\n| 175 | vowel_diacritic     |       7 | ে           |\n| 168 | vowel_diacritic     |       0 | 0           |\n| 181 | consonant_diacritic |       2 | র্           |</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 729058,
      "author_name": "K.Amano",
      "author_url": "",
      "post_date": "2020-01-25T17:14:34.183000",
      "content": "<p>Hi Datta!\nর্তী and র্ত্রী are different typefaces, but they are the same character.\nThank you very much.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 730603,
          "author_name": "datta",
          "author_url": "",
          "post_date": "2020-01-27T17:13:29.823000",
          "content": "<p>I thought so too, but have a look here. <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 729000,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-25T15:04:04.910000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "730068": "I've reported this problem and got a response from one of competetion hosts saying that a fix is expected recently.\nJust check this link: (hope it help)\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123002#728098",
    "729028": "You are right there are 1295 distinct graphemes but only 1291 combinations of the 3 components. \nBut you can notice that all the 3 components are kind of the same, just a little different shape? Maybe there was some sort of a change in the fonts they used while preparing the dataset.\n\n|      | image_id   |   grapheme_root |   vowel_diacritic |   consonant_diacritic | grapheme   |    mix |\n|-----:|:-----------|----------------:|------------------:|----------------------:|:-----------|-------:|\n|  387 | Train_387  |              64 |                 3 |                     2 | র্তী        | 64_3_2 |\n| 1482 | Train_1482 |              64 |                 3 |                     2 | র্ত্রী       | 64_3_2 |\n|  577 | Train_577  |              64 |                 7 |                     2 | র্তে        | 64_7_2 |\n| 1874 | Train_1874 |              64 |                 7 |                     2 | র্ত্রে       | 64_7_2 |\n|  532 | Train_532  |              72 |                 0 |                     2 | র্দ্র        | 72_0_2 |\n| 1117 | Train_1117 |              72 |                 0 |                     2 | র্দ         | 72_0_2 |\n\n|     | component_type      |   label | component   |\n|----:|:--------------------|--------:|:------------|\n| 64 | grapheme_root    |      64 | ত           |\n| 72 | grapheme_root    |      72 | দ           |\n| 171 | vowel_diacritic     |       3 | ী           |\n| 175 | vowel_diacritic     |       7 | ে           |\n| 168 | vowel_diacritic     |       0 | 0           |\n| 181 | consonant_diacritic |       2 | র্           |\n",
    "728979": "Are there multiple graphemes for the same combination of `grapheme_root`, `vowel_diacritic`, and `consonant_diacritic`?\n\nFor example, in train.csv\n```\nimage_id=378  :  (64, 3, 2)   :  র্তী\nimage_id=1482 :  (64, 3, 2)   :  র্ত্রী\n```\nThere are **1295** types of grapheme in train.csv, but there are only **1292** combinations.\nPlease tell me about Bengali Grapheme.",
    "729058": "Hi Datta!\nর্তী and র্ত্রী are different typefaces, but they are the same character.\nThank you very much.",
    "729000": ""
  }
}