{
  "id": 134258,
  "title": "Handling multi consonant graphemes",
  "url": "/competitions/bengaliai-cv19/discussion/134258",
  "author_name": "",
  "post_date": "2020-03-06T23:54:31.026347Z",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<p>As noted in <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">this discussion</a>, two of the consonant diacritics র্ (class 2) and ্র ( class 5) can coexist in the same grapheme.\nThe original labeling scheme did not account for this possibility so these cases are currently labeled as class (2). We are releasing an updated class map that adds this special case as class 7, and a list of the affected rows in the training set. Approximately 450 rows were affected in each of the train and test sets. </p>\n\n<p>Given the small number of rows affected and the limited time left in the competition, we are NOT changing the solution file. Please do not make any predictions of class 7!</p>\n\n<p>We really appreciate your patience while we were sorting out this issue. Thanks again to @mnpinto for flagging the issue!</p>",
  "messages": [
    {
      "id": "765658",
      "postDate": "03/06/2020 23:54:31",
      "content": "<p>As noted in <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">this discussion</a>, two of the consonant diacritics র্ (class 2) and ্র ( class 5) can coexist in the same grapheme.\nThe original labeling scheme did not account for this possibility so these cases are currently labeled as class (2). We are releasing an updated class map that adds this special case as class 7, and a list of the affected rows in the training set. Approximately 450 rows were affected in each of the train and test sets. </p>\n\n<p>Given the small number of rows affected and the limited time left in the competition, we are NOT changing the solution file. Please do not make any predictions of class 7!</p>\n\n<p>We really appreciate your patience while we were sorting out this issue. Thanks again to @mnpinto for flagging the issue!</p>",
      "rawMarkdown": "As noted in [this discussion](https://www.kaggle.com/c/bengaliai-cv19/discussion/123859), two of the consonant diacritics র্ (class 2) and ্র ( class 5) can coexist in the same grapheme.\nThe original labeling scheme did not account for this possibility so these cases are currently labeled as class (2). We are releasing an updated class map that adds this special case as class 7, and a list of the affected rows in the training set. Approximately 450 rows were affected in each of the train and test sets. \n\nGiven the small number of rows affected and the limited time left in the competition, we are NOT changing the solution file. Please do not make any predictions of class 7!\n\nWe really appreciate your patience while we were sorting out this issue. Thanks again to @mnpinto for flagging the issue!",
      "votes": null
    },
    {
      "id": "765920",
      "postDate": "03/07/2020 11:09:49",
      "content": "<p>So the model is suppose to classify these as 2 then? </p>",
      "rawMarkdown": "So the model is suppose to classify these as 2 then?",
      "votes": null
    },
    {
      "id": "765994",
      "postDate": "03/07/2020 14:03:13",
      "content": "<p>Yes!</p>",
      "rawMarkdown": "Yes!",
      "votes": null
    },
    {
      "id": "765995",
      "postDate": "03/07/2020 14:07:26",
      "content": "<p>Ok, so in a nutshell, nothing changes :)</p>",
      "rawMarkdown": "Ok, so in a nutshell, nothing changes :)",
      "votes": null
    },
    {
      "id": "766055",
      "postDate": "03/07/2020 15:51:40",
      "content": "<p>Yeah ! so why updating the class map then... that's really confusing I think a few days before the end of the challenge.</p>",
      "rawMarkdown": "Yeah ! so why updating the class map then... that's really confusing I think a few days before the end of the challenge.",
      "votes": null
    },
    {
      "id": "766196",
      "postDate": "03/07/2020 20:44:31",
      "content": "<p>I have discovered that the suposed root_graphemes are in fact composed of smaller set of 62 base_graphemes.</p>\n\n<p>Using a set analyis, I have identified the set of unicode base_graphemes for each of the vowels, consonants and roots.</p>\n\n<p>I have also performed a full visualization of the bengali alphabet</p>\n\n<p><a href=\"https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet\">https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet</a></p>",
      "rawMarkdown": "I have discovered that the suposed root\\_graphemes are in fact composed of smaller set of 62 base\\_graphemes.\n\nUsing a set analyis, I have identified the set of unicode base\\_graphemes for each of the vowels, consonants and roots.\n\nI have also performed a full visualization of the bengali alphabet\n\nhttps://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet",
      "votes": null
    },
    {
      "id": "766214",
      "postDate": "03/07/2020 21:17:49",
      "content": "<p><a href=\"/ogrellier\">@ogrellier</a> said true, why now at the very end of the competition, this should address much earlier as such problems were informed long time ago. </p>",
      "rawMarkdown": "ogrellier said true, why now at the very end of the competition, this should address much earlier as such problems were informed long time ago.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 765920,
      "author_name": "maxjon",
      "author_url": "",
      "post_date": "03/07/2020 11:09:49",
      "content": "<p>So the model is suppose to classify these as 2 then? </p>",
      "votes": null,
      "replies": [
        {
          "id": 765994,
          "author_name": "reasat",
          "author_url": "",
          "post_date": "03/07/2020 14:03:13",
          "content": "<p>Yes!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765995,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/07/2020 14:07:26",
          "content": "<p>Ok, so in a nutshell, nothing changes :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766055,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "03/07/2020 15:51:40",
          "content": "<p>Yeah ! so why updating the class map then... that's really confusing I think a few days before the end of the challenge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766214,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "03/07/2020 21:17:49",
          "content": "<p><a href=\"/ogrellier\">@ogrellier</a> said true, why now at the very end of the competition, this should address much earlier as such problems were informed long time ago. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 766196,
      "author_name": "jamesmcguigan",
      "author_url": "",
      "post_date": "03/07/2020 20:44:31",
      "content": "<p>I have discovered that the suposed root_graphemes are in fact composed of smaller set of 62 base_graphemes.</p>\n\n<p>Using a set analyis, I have identified the set of unicode base_graphemes for each of the vowels, consonants and roots.</p>\n\n<p>I have also performed a full visualization of the bengali alphabet</p>\n\n<p><a href=\"https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet\">https://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "765658": "As noted in [this discussion](https://www.kaggle.com/c/bengaliai-cv19/discussion/123859), two of the consonant diacritics র্ (class 2) and ্র ( class 5) can coexist in the same grapheme.\nThe original labeling scheme did not account for this possibility so these cases are currently labeled as class (2). We are releasing an updated class map that adds this special case as class 7, and a list of the affected rows in the training set. Approximately 450 rows were affected in each of the train and test sets. \n\nGiven the small number of rows affected and the limited time left in the competition, we are NOT changing the solution file. Please do not make any predictions of class 7!\n\nWe really appreciate your patience while we were sorting out this issue. Thanks again to @mnpinto for flagging the issue!",
    "765920": "So the model is suppose to classify these as 2 then?",
    "765994": "Yes!",
    "765995": "Ok, so in a nutshell, nothing changes :)",
    "766055": "Yeah ! so why updating the class map then... that's really confusing I think a few days before the end of the challenge.",
    "766196": "I have discovered that the suposed root\\_graphemes are in fact composed of smaller set of 62 base\\_graphemes.\n\nUsing a set analyis, I have identified the set of unicode base\\_graphemes for each of the vowels, consonants and roots.\n\nI have also performed a full visualization of the bengali alphabet\n\nhttps://www.kaggle.com/jamesmcguigan/unicode-visualization-of-the-bengali-alphabet",
    "766214": "ogrellier said true, why now at the very end of the competition, this should address much earlier as such problems were informed long time ago."
  },
  "source": "meta"
}