{
  "id": 392304,
  "title": " Potential leakage in the test data",
  "url": "/competitions/ml-olympiad-dialectrecognition/discussion/392304",
  "author_name": "",
  "post_date": "2023-03-04T16:47:03.527276300Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Dear organizers,</p>\n<p>I wanted to bring to your attention a potential leakage issue in the test data. Upon further exploration, I found that the <code>FileName</code>and <code>ShowName</code>features in the test data are identical to those in the training data. This could allow a model to learn the dialect of the speaker from the file name rather than the audio content itself, resulting in performance-overestimation.</p>\n<p>I am reporting this issue to seek guidance on how to proceed. It would be helpful if a list of allowed features is provided clearly to ensure that all participants are working within the constraints of the competition.</p>\n<p>Best wishes!</p>",
  "messages": [
    {
      "id": "2168946",
      "postDate": "03/04/2023 16:47:03",
      "content": "<p>Dear organizers,</p>\n<p>I wanted to bring to your attention a potential leakage issue in the test data. Upon further exploration, I found that the <code>FileName</code>and <code>ShowName</code>features in the test data are identical to those in the training data. This could allow a model to learn the dialect of the speaker from the file name rather than the audio content itself, resulting in performance-overestimation.</p>\n<p>I am reporting this issue to seek guidance on how to proceed. It would be helpful if a list of allowed features is provided clearly to ensure that all participants are working within the constraints of the competition.</p>\n<p>Best wishes!</p>",
      "rawMarkdown": "Dear organizers,\n\nI wanted to bring to your attention a potential leakage issue in the test data. Upon further exploration, I found that the `FileName `and `ShowName `features in the test data are identical to those in the training data. This could allow a model to learn the dialect of the speaker from the file name rather than the audio content itself, resulting in performance-overestimation.\n\nI am reporting this issue to seek guidance on how to proceed. It would be helpful if a list of allowed features is provided clearly to ensure that all participants are working within the constraints of the competition.\n\nBest wishes!",
      "votes": null
    },
    {
      "id": "2169555",
      "postDate": "03/05/2023 07:47:58",
      "content": "<p>There are possible other data leaks also , however I think the idea of the competition is to use the text to predict ( as the main feature). </p>\n<p>And the main goal ( I think) is to develope solutions that can generalize well ( on text or voice !) outside the provided data or the original data.</p>",
      "rawMarkdown": "There are possible other data leaks also , however I think the idea of the competition is to use the text to predict ( as the main feature). \n\nAnd the main goal ( I think) is to develope solutions that can generalize well ( on text or voice !) outside the provided data or the original data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2169555,
      "author_name": "asalhi",
      "author_url": "",
      "post_date": "03/05/2023 07:47:58",
      "content": "<p>There are possible other data leaks also , however I think the idea of the competition is to use the text to predict ( as the main feature). </p>\n<p>And the main goal ( I think) is to develope solutions that can generalize well ( on text or voice !) outside the provided data or the original data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2168946": "Dear organizers,\n\nI wanted to bring to your attention a potential leakage issue in the test data. Upon further exploration, I found that the `FileName `and `ShowName `features in the test data are identical to those in the training data. This could allow a model to learn the dialect of the speaker from the file name rather than the audio content itself, resulting in performance-overestimation.\n\nI am reporting this issue to seek guidance on how to proceed. It would be helpful if a list of allowed features is provided clearly to ensure that all participants are working within the constraints of the competition.\n\nBest wishes!",
    "2169555": "There are possible other data leaks also , however I think the idea of the competition is to use the text to predict ( as the main feature). \n\nAnd the main goal ( I think) is to develope solutions that can generalize well ( on text or voice !) outside the provided data or the original data."
  },
  "source": "meta"
}