{
  "id": 434714,
  "title": "I found a bad annotation in train data, and have a question about test data",
  "url": "/competitions/bengaliai-speech/discussion/434714",
  "author_name": "",
  "post_date": "2023-08-26T08:26:48.351218500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I found that some of sentences in the train data whose split is 'valid' don't end with punctuations.<br>\nThe example I found is as following.</p>\n<p>id:03ef414022ac<br>\nsentence:তিনি নলডাঙ্গা উপজেলার ব্যাপক উন্নতি করেছে</p>\n<p>According to Google Translation, the sentence means \"He has greatly improved Naldanga upazila\".</p>\n<p>I don't know Bengali well, but I think that this sentence  should end with dari(।).<br>\nAre there any annotation in test data which don't end with dari(।) though dari(।) is needed grammatically?<br>\nI think test data should not contain such an annotation.</p>",
  "messages": [
    {
      "id": "2409409",
      "postDate": "08/26/2023 08:26:48",
      "content": "<p>I found that some of sentences in the train data whose split is 'valid' don't end with punctuations.<br>\nThe example I found is as following.</p>\n<p>id:03ef414022ac<br>\nsentence:তিনি নলডাঙ্গা উপজেলার ব্যাপক উন্নতি করেছে</p>\n<p>According to Google Translation, the sentence means \"He has greatly improved Naldanga upazila\".</p>\n<p>I don't know Bengali well, but I think that this sentence  should end with dari(।).<br>\nAre there any annotation in test data which don't end with dari(।) though dari(।) is needed grammatically?<br>\nI think test data should not contain such an annotation.</p>",
      "rawMarkdown": "I found that some of sentences in the train data whose split is 'valid' don't end with punctuations.\nThe example I found is as following.\n\nid:03ef414022ac\nsentence:তিনি নলডাঙ্গা উপজেলার ব্যাপক উন্নতি করেছে\n\nAccording to Google Translation, the sentence means \"He has greatly improved Naldanga upazila\".\n\nI don't know Bengali well, but I think that this sentence  should end with dari(।).\nAre there any annotation in test data which don't end with dari(।) though dari(।) is needed grammatically?\nI think test data should not contain such an annotation.",
      "votes": null
    },
    {
      "id": "2409566",
      "postDate": "08/26/2023 10:09:35",
      "content": "<p>Hi moto, great question!</p>\n<p>First thing first, the test data is normalized using bnunicodenormalizer, so unicode specific errors and conflicts in punctuation types are more or less resolved by that. We didn't normalize the training data to provide training data in its rawest form.</p>\n<p>Secondly, if there are punctuations in test data, it is very likely because a human annotator could hear its context in the audio. The test data was quite rigorously validated so we can expect label noise (including noise related to punctuation labels) should be very low.</p>",
      "rawMarkdown": "Hi moto, great question!\n\nFirst thing first, the test data is normalized using bnunicodenormalizer, so unicode specific errors and conflicts in punctuation types are more or less resolved by that. We didn't normalize the training data to provide training data in its rawest form.\n\nSecondly, if there are punctuations in test data, it is very likely because a human annotator could hear its context in the audio. The test data was quite rigorously validated so we can expect label noise (including noise related to punctuation labels) should be very low.",
      "votes": null
    },
    {
      "id": "2410125",
      "postDate": "08/26/2023 16:41:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yamashitamotokazu\" target=\"_blank\">@yamashitamotokazu</a> , great find! I have ran an analysis and here are my findings. </p>\n<blockquote>\n  <p>2454 sentences in the train set and 214 sentences in the valid split don't contain any punctuations ( '।': '?', '!' ) in the end. <br>\n  My hypothesis is either these are mistakes by the annotators or some of these are not complete sentences. </p>\n</blockquote>",
      "rawMarkdown": "Hi @yamashitamotokazu , great find! I have ran an analysis and here are my findings. \n> 2454 sentences in the train set and 214 sentences in the valid split don't contain any punctuations ( '।': '?', '!' ) in the end. \nMy hypothesis is either these are mistakes by the annotators or some of these are not complete sentences.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2409566,
      "author_name": "imtiazprio",
      "author_url": "",
      "post_date": "08/26/2023 10:09:35",
      "content": "<p>Hi moto, great question!</p>\n<p>First thing first, the test data is normalized using bnunicodenormalizer, so unicode specific errors and conflicts in punctuation types are more or less resolved by that. We didn't normalize the training data to provide training data in its rawest form.</p>\n<p>Secondly, if there are punctuations in test data, it is very likely because a human annotator could hear its context in the audio. The test data was quite rigorously validated so we can expect label noise (including noise related to punctuation labels) should be very low.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2410125,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "08/26/2023 16:41:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yamashitamotokazu\" target=\"_blank\">@yamashitamotokazu</a> , great find! I have ran an analysis and here are my findings. </p>\n<blockquote>\n  <p>2454 sentences in the train set and 214 sentences in the valid split don't contain any punctuations ( '।': '?', '!' ) in the end. <br>\n  My hypothesis is either these are mistakes by the annotators or some of these are not complete sentences. </p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2409409": "I found that some of sentences in the train data whose split is 'valid' don't end with punctuations.\nThe example I found is as following.\n\nid:03ef414022ac\nsentence:তিনি নলডাঙ্গা উপজেলার ব্যাপক উন্নতি করেছে\n\nAccording to Google Translation, the sentence means \"He has greatly improved Naldanga upazila\".\n\nI don't know Bengali well, but I think that this sentence  should end with dari(।).\nAre there any annotation in test data which don't end with dari(।) though dari(।) is needed grammatically?\nI think test data should not contain such an annotation.",
    "2409566": "Hi moto, great question!\n\nFirst thing first, the test data is normalized using bnunicodenormalizer, so unicode specific errors and conflicts in punctuation types are more or less resolved by that. We didn't normalize the training data to provide training data in its rawest form.\n\nSecondly, if there are punctuations in test data, it is very likely because a human annotator could hear its context in the audio. The test data was quite rigorously validated so we can expect label noise (including noise related to punctuation labels) should be very low.",
    "2410125": "Hi @yamashitamotokazu , great find! I have ran an analysis and here are my findings. \n> 2454 sentences in the train set and 214 sentences in the valid split don't contain any punctuations ( '।': '?', '!' ) in the end. \nMy hypothesis is either these are mistakes by the annotators or some of these are not complete sentences."
  },
  "source": "meta"
}