{
  "id": 126833,
  "title": "Wrong Train Data Labels -Organizers plz clarify about Private Set",
  "url": "/competitions/bengaliai-cv19/discussion/126833",
  "author_name": "Nirjhar Roy",
  "post_date": "2020-01-20T15:19:15.117000",
  "votes": 35,
  "comment_count": 14,
  "views": 0,
  "content": "<p>There are certain wrong labels in each category of grapheme_root, vowel_diacritic, consonant_diacritic  in the train data . I did a very small sampling by running train data through a good model and found those few cases . There can be few others , will there be such cases in private test set also ? Since there are two months to go for the competition , and by the progress so far , we can expect high accuracy score. However , if there are noisy data or wrong label in private set , then some model might get the favor of luck in their side .  Please correct if my analysis is wrong.\nExample : \nWrong Grapheme Root:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F895ee0ee2f2de182ddcb5c68786573af%2FWrong_Grapheme_root_Label.PNG?generation=1579533479727553&amp;alt=media\" alt=\"\">\nWrong Vowel: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F24d5492262755d16a668d56cf4b93b00%2FWrong_Vowel_Label.PNG?generation=1579533513327925&amp;alt=media\" alt=\"\">\nWrong Consonant Label: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F6293a328362b30c194a4351fa15fd086%2FWrong_consonant_label.PNG?generation=1579533548428764&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 723890,
      "postDate": "2020-01-20T15:19:15.117Z",
      "content": "<p>There are certain wrong labels in each category of grapheme_root, vowel_diacritic, consonant_diacritic  in the train data . I did a very small sampling by running train data through a good model and found those few cases . There can be few others , will there be such cases in private test set also ? Since there are two months to go for the competition , and by the progress so far , we can expect high accuracy score. However , if there are noisy data or wrong label in private set , then some model might get the favor of luck in their side .  Please correct if my analysis is wrong.\nExample : \nWrong Grapheme Root:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F895ee0ee2f2de182ddcb5c68786573af%2FWrong_Grapheme_root_Label.PNG?generation=1579533479727553&amp;alt=media\" alt=\"\">\nWrong Vowel: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F24d5492262755d16a668d56cf4b93b00%2FWrong_Vowel_Label.PNG?generation=1579533513327925&amp;alt=media\" alt=\"\">\nWrong Consonant Label: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F6293a328362b30c194a4351fa15fd086%2FWrong_consonant_label.PNG?generation=1579533548428764&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "There are certain wrong labels in each category of grapheme_root, vowel_diacritic, consonant_diacritic  in the train data . I did a very small sampling by running train data through a good model and found those few cases . There can be few others , will there be such cases in private test set also ? Since there are two months to go for the competition , and by the progress so far , we can expect high accuracy score. However , if there are noisy data or wrong label in private set , then some model might get the favor of luck in their side .  Please correct if my analysis is wrong.\nExample : \nWrong Grapheme Root:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F895ee0ee2f2de182ddcb5c68786573af%2FWrong_Grapheme_root_Label.PNG?generation=1579533479727553&amp;alt=media)\nWrong Vowel: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F24d5492262755d16a668d56cf4b93b00%2FWrong_Vowel_Label.PNG?generation=1579533513327925&amp;alt=media)\nWrong Consonant Label: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F6293a328362b30c194a4351fa15fd086%2FWrong_consonant_label.PNG?generation=1579533548428764&amp;alt=media)\n",
      "votes": 35
    },
    {
      "id": 725340,
      "postDate": "2020-01-22T02:22:12.940Z",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> Thanks for pointing them out. Although we went through a rigorous curation process, some noisy samples seem to have passed through. As you can imagine, this can happen due to the volume of the data. </p>\n\n<p>We have checked the test set before the launch of the competition and corrected the labels (if there's any). We are making another check on the test data just to be sure. Will get back to you soon. </p>\n\n<p>In the meantime, if you find any issues in the train set, let us know. Thanks!</p>",
      "rawMarkdown": "@phoenix9032 Thanks for pointing them out. Although we went through a rigorous curation process, some noisy samples seem to have passed through. As you can imagine, this can happen due to the volume of the data. \n\nWe have checked the test set before the launch of the competition and corrected the labels (if there's any). We are making another check on the test data just to be sure. Will get back to you soon. \n\nIn the meantime, if you find any issues in the train set, let us know. Thanks!",
      "votes": 7,
      "replies": [
        {
          "id": 725609,
          "postDate": "2020-01-22T09:34:13.713Z",
          "content": "<p>Thank you so much for your response .  I don't think there are any major issues . There has been lot of care taken to prepare dataset and it shows .</p>",
          "rawMarkdown": "Thank you so much for your response .  I don't think there are any major issues . There has been lot of care taken to prepare dataset and it shows .",
          "votes": 1
        },
        {
          "id": 727374,
          "postDate": "2020-01-23T17:06:58.757Z",
          "content": "<p>I grouped by the labels dataframe from train.csv with 3-target columns, as shown below as \"b\";\nalso I grouped by it with grapheme images provided in the csv file, as shown below as \"a\".\n<code>label_csv = pd.read_csv('/kaggle/input/bengaliai-cv19/train.csv')</code>\n<code>a=label_csv.groupby(['grapheme'], sort=False).size()</code>\n<code>b=label_csv.groupby(['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'], sort=False).size()</code>\n\"a\" got a correct size count which is <strong>1295 classes</strong>, also mentioned in the slides provided from competition hosts.\nHowever, size count of \"b\" is only <strong>1292 classes</strong>, which means some labels are incorrectly provided in csv file.\nBtw, the minimum size of all classes are identical as 118 in \"a\" and \"b\", but the maximum size is 283, 303 respectively.</p>\n\n<p>Any competition host could fix it? (mistakes in train.csv)</p>",
          "rawMarkdown": "I grouped by the labels dataframe from train.csv with 3-target columns, as shown below as \"b\";\nalso I grouped by it with grapheme images provided in the csv file, as shown below as \"a\".\n`label_csv = pd.read_csv('/kaggle/input/bengaliai-cv19/train.csv')`\n`a=label_csv.groupby(['grapheme'], sort=False).size()`\n`b=label_csv.groupby(['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'], sort=False).size()`\n\"a\" got a correct size count which is **1295 classes**, also mentioned in the slides provided from competition hosts.\nHowever, size count of \"b\" is only **1292 classes**, which means some labels are incorrectly provided in csv file.\nBtw, the minimum size of all classes are identical as 118 in \"a\" and \"b\", but the maximum size is 283, 303 respectively.\n\nAny competition host could fix it? (mistakes in train.csv)"
        },
        {
          "id": 727391,
          "postDate": "2020-01-23T17:21:44.830Z",
          "content": "<p>you've found the same thing as this: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a>\nwhich haven't been solved yet</p>",
          "rawMarkdown": "you've found the same thing as this: https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\nwhich haven't been solved yet",
          "votes": 3
        },
        {
          "id": 727436,
          "postDate": "2020-01-23T18:11:46.450Z",
          "content": "<p><a href=\"/haqishen\">@haqishen</a>, I found the 3 missing graphemes as the same in that link: \"র্দ্র\", \"র্ত্রী\", \"র্ত্রে\"\nThey are new classes but have been recognized as old ones with existed combination in train.csv.\nThe combination of ['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'] for these three should be re-classified and then replace those mistaken labels with new ones.</p>\n\n<p>In case anyone would like to take a look at them, the index of missing graphemes appearing first time in the dataset is provided here as 532, 1482, 1874 respectively. Also, there are 151, 145, 150 data respectively in these three classes.</p>",
          "rawMarkdown": "@haqishen, I found the 3 missing graphemes as the same in that link: \"র্দ্র\", \"র্ত্রী\", \"র্ত্রে\"\nThey are new classes but have been recognized as old ones with existed combination in train.csv.\nThe combination of ['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'] for these three should be re-classified and then replace those mistaken labels with new ones.\n\nIn case anyone would like to take a look at them, the index of missing graphemes appearing first time in the dataset is provided here as 532, 1482, 1874 respectively. Also, there are 151, 145, 150 data respectively in these three classes.",
          "votes": 1
        },
        {
          "id": 728474,
          "postDate": "2020-01-24T19:00:48.757Z",
          "content": "<p>Thanks <a href=\"/haqishen\">@haqishen</a>  and <a href=\"/allenchang9201\">@allenchang9201</a>. We are aware of the linked discussion and the issues with the three graphemes. We are trying to decide between a few possible solutions. We shall come up with the updates soon.</p>",
          "rawMarkdown": "Thanks @haqishen  and @allenchang9201. We are aware of the linked discussion and the issues with the three graphemes. We are trying to decide between a few possible solutions. We shall come up with the updates soon.",
          "votes": 2
        }
      ]
    },
    {
      "id": 725349,
      "postDate": "2020-01-22T02:34:54.703Z",
      "content": "<p>we have approximately 200k test samples, and each sample have 3 labels.\nso first 4 digits of LB score may not be affected even if there are roughly 30 incorrect labels in test set?\n(just a quick calculatition: 30 / 200k / 3 = 5e-05)</p>",
      "rawMarkdown": "we have approximately 200k test samples, and each sample have 3 labels.\nso first 4 digits of LB score may not be affected even if there are roughly 30 incorrect labels in test set?\n(just a quick calculatition: 30 / 200k / 3 = 5e-05)\n",
      "votes": 4,
      "replies": [
        {
          "id": 725608,
          "postDate": "2020-01-22T09:33:16.883Z",
          "content": "<p>Thank you </p>",
          "rawMarkdown": "Thank you "
        }
      ]
    },
    {
      "id": 726697,
      "postDate": "2020-01-23T07:17:40.117Z",
      "content": "<p>Is there any native speaker/reader of Bengali? When I look at my train/validation images with max errors, most of the time I cannot tell if they are correct or not. I assume it is within the rules to collaborate on cleaning the training dataset. Anyone?</p>",
      "rawMarkdown": "Is there any native speaker/reader of Bengali? When I look at my train/validation images with max errors, most of the time I cannot tell if they are correct or not. I assume it is within the rules to collaborate on cleaning the training dataset. Anyone?",
      "votes": 1,
      "replies": [
        {
          "id": 726842,
          "postDate": "2020-01-23T09:13:29.657Z",
          "content": "<p>You can always post those images with the ground truth and prediction and someone will surely help.</p>",
          "rawMarkdown": "You can always post those images with the ground truth and prediction and someone will surely help."
        }
      ]
    },
    {
      "id": 724067,
      "postDate": "2020-01-20T18:55:25.813Z",
      "content": "<p>Dear <a href=\"/reasat\">@reasat</a> and <a href=\"/imtiazprio\">@imtiazprio</a>, have a look here. Its very important issue!</p>",
      "rawMarkdown": "Dear @reasat and @imtiazprio, have a look here. Its very important issue!",
      "votes": 2
    },
    {
      "id": 724066,
      "postDate": "2020-01-20T18:53:36.027Z",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> thanks for the observation, it will be quite helpful in my model building and training. Bengali seems your native language, By random sampling can you tell us what fraction of data is mis labeled?</p>",
      "rawMarkdown": "@phoenix9032 thanks for the observation, it will be quite helpful in my model building and training. Bengali seems your native language, By random sampling can you tell us what fraction of data is mis labeled?"
    },
    {
      "id": 727370,
      "postDate": "2020-01-23T17:03:08.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 725722,
      "postDate": "2020-01-22T12:37:58.960Z",
      "content": "<p>Thanks for sharing this with us !! </p>",
      "rawMarkdown": "Thanks for sharing this with us !! ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 725340,
      "author_name": "Tahsin",
      "author_url": "",
      "post_date": "2020-01-22T02:22:12.940000",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> Thanks for pointing them out. Although we went through a rigorous curation process, some noisy samples seem to have passed through. As you can imagine, this can happen due to the volume of the data. </p>\n\n<p>We have checked the test set before the launch of the competition and corrected the labels (if there's any). We are making another check on the test data just to be sure. Will get back to you soon. </p>\n\n<p>In the meantime, if you find any issues in the train set, let us know. Thanks!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 725609,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-22T09:34:13.713000",
          "content": "<p>Thank you so much for your response .  I don't think there are any major issues . There has been lot of care taken to prepare dataset and it shows .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 727374,
          "author_name": "AllenChangTW",
          "author_url": "",
          "post_date": "2020-01-23T17:06:58.757000",
          "content": "<p>I grouped by the labels dataframe from train.csv with 3-target columns, as shown below as \"b\";\nalso I grouped by it with grapheme images provided in the csv file, as shown below as \"a\".\n<code>label_csv = pd.read_csv('/kaggle/input/bengaliai-cv19/train.csv')</code>\n<code>a=label_csv.groupby(['grapheme'], sort=False).size()</code>\n<code>b=label_csv.groupby(['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'], sort=False).size()</code>\n\"a\" got a correct size count which is <strong>1295 classes</strong>, also mentioned in the slides provided from competition hosts.\nHowever, size count of \"b\" is only <strong>1292 classes</strong>, which means some labels are incorrectly provided in csv file.\nBtw, the minimum size of all classes are identical as 118 in \"a\" and \"b\", but the maximum size is 283, 303 respectively.</p>\n\n<p>Any competition host could fix it? (mistakes in train.csv)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 727391,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-23T17:21:44.830000",
          "content": "<p>you've found the same thing as this: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a>\nwhich haven't been solved yet</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 727436,
          "author_name": "AllenChangTW",
          "author_url": "",
          "post_date": "2020-01-23T18:11:46.450000",
          "content": "<p><a href=\"/haqishen\">@haqishen</a>, I found the 3 missing graphemes as the same in that link: \"র্দ্র\", \"র্ত্রী\", \"র্ত্রে\"\nThey are new classes but have been recognized as old ones with existed combination in train.csv.\nThe combination of ['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'] for these three should be re-classified and then replace those mistaken labels with new ones.</p>\n\n<p>In case anyone would like to take a look at them, the index of missing graphemes appearing first time in the dataset is provided here as 532, 1482, 1874 respectively. Also, there are 151, 145, 150 data respectively in these three classes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 728474,
          "author_name": "Tahsin",
          "author_url": "",
          "post_date": "2020-01-24T19:00:48.757000",
          "content": "<p>Thanks <a href=\"/haqishen\">@haqishen</a>  and <a href=\"/allenchang9201\">@allenchang9201</a>. We are aware of the linked discussion and the issues with the three graphemes. We are trying to decide between a few possible solutions. We shall come up with the updates soon.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 725349,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-01-22T02:34:54.703000",
      "content": "<p>we have approximately 200k test samples, and each sample have 3 labels.\nso first 4 digits of LB score may not be affected even if there are roughly 30 incorrect labels in test set?\n(just a quick calculatition: 30 / 200k / 3 = 5e-05)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 725608,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-22T09:33:16.883000",
          "content": "<p>Thank you </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 726697,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "2020-01-23T07:17:40.117000",
      "content": "<p>Is there any native speaker/reader of Bengali? When I look at my train/validation images with max errors, most of the time I cannot tell if they are correct or not. I assume it is within the rules to collaborate on cleaning the training dataset. Anyone?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 726842,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-23T09:13:29.657000",
          "content": "<p>You can always post those images with the ground truth and prediction and someone will surely help.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 724067,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2020-01-20T18:55:25.813000",
      "content": "<p>Dear <a href=\"/reasat\">@reasat</a> and <a href=\"/imtiazprio\">@imtiazprio</a>, have a look here. Its very important issue!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 724066,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2020-01-20T18:53:36.027000",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> thanks for the observation, it will be quite helpful in my model building and training. Bengali seems your native language, By random sampling can you tell us what fraction of data is mis labeled?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727370,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T17:03:08.317000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 725722,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-22T12:37:58.960000",
      "content": "<p>Thanks for sharing this with us !! </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "723890": "There are certain wrong labels in each category of grapheme_root, vowel_diacritic, consonant_diacritic  in the train data . I did a very small sampling by running train data through a good model and found those few cases . There can be few others , will there be such cases in private test set also ? Since there are two months to go for the competition , and by the progress so far , we can expect high accuracy score. However , if there are noisy data or wrong label in private set , then some model might get the favor of luck in their side .  Please correct if my analysis is wrong.\nExample : \nWrong Grapheme Root:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F895ee0ee2f2de182ddcb5c68786573af%2FWrong_Grapheme_root_Label.PNG?generation=1579533479727553&amp;alt=media)\nWrong Vowel: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F24d5492262755d16a668d56cf4b93b00%2FWrong_Vowel_Label.PNG?generation=1579533513327925&amp;alt=media)\nWrong Consonant Label: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F6293a328362b30c194a4351fa15fd086%2FWrong_consonant_label.PNG?generation=1579533548428764&amp;alt=media)\n",
    "725340": "@phoenix9032 Thanks for pointing them out. Although we went through a rigorous curation process, some noisy samples seem to have passed through. As you can imagine, this can happen due to the volume of the data. \n\nWe have checked the test set before the launch of the competition and corrected the labels (if there's any). We are making another check on the test data just to be sure. Will get back to you soon. \n\nIn the meantime, if you find any issues in the train set, let us know. Thanks!",
    "725349": "we have approximately 200k test samples, and each sample have 3 labels.\nso first 4 digits of LB score may not be affected even if there are roughly 30 incorrect labels in test set?\n(just a quick calculatition: 30 / 200k / 3 = 5e-05)\n",
    "726697": "Is there any native speaker/reader of Bengali? When I look at my train/validation images with max errors, most of the time I cannot tell if they are correct or not. I assume it is within the rules to collaborate on cleaning the training dataset. Anyone?",
    "724067": "Dear @reasat and @imtiazprio, have a look here. Its very important issue!",
    "724066": "@phoenix9032 thanks for the observation, it will be quite helpful in my model building and training. Bengali seems your native language, By random sampling can you tell us what fraction of data is mis labeled?",
    "727370": "",
    "725722": "Thanks for sharing this with us !! "
  }
}