{
  "id": 209493,
  "title": "Should we expect a clean dataset?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/209493",
  "author_name": "",
  "post_date": "2021-01-07T18:10:28.453756600Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Continuing the topic of noise in data, I would like to ask the organizers a specific question: should we expect a clean dataset?</p>\n<p>I started noticing posts about numerous pictures with incorrect labels on the forum a couple of weeks ago. This is quite normal practice for Kaggle competitions. A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. </p>\n<p>Based on the above, it is obvious that I'd like to know if we'll ever get the clean dataset. The future course of solving this competition for many people (or not solving it) depends on this. At the moment, half of all my ideas and methods are aimed at dealing with the noise.</p>\n<p>Special loss functions, label smoothing, knowledge distillation are good methods. But none of them is comparable to the manual relabeling of the data. And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?</p>\n<p>It follows that a person who has the specialists for this purpose or the ability to hire them, can ensure a good place in the competition and at the same time devalue other people's solutions? Or is this exactly what the organizers expect us to do so that we just clean their dataset?</p>\n<p>Thank you in advance.</p>",
  "messages": [
    {
      "id": "1143025",
      "postDate": "01/07/2021 18:10:28",
      "content": "<p>Continuing the topic of noise in data, I would like to ask the organizers a specific question: should we expect a clean dataset?</p>\n<p>I started noticing posts about numerous pictures with incorrect labels on the forum a couple of weeks ago. This is quite normal practice for Kaggle competitions. A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. </p>\n<p>Based on the above, it is obvious that I'd like to know if we'll ever get the clean dataset. The future course of solving this competition for many people (or not solving it) depends on this. At the moment, half of all my ideas and methods are aimed at dealing with the noise.</p>\n<p>Special loss functions, label smoothing, knowledge distillation are good methods. But none of them is comparable to the manual relabeling of the data. And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?</p>\n<p>It follows that a person who has the specialists for this purpose or the ability to hire them, can ensure a good place in the competition and at the same time devalue other people's solutions? Or is this exactly what the organizers expect us to do so that we just clean their dataset?</p>\n<p>Thank you in advance.</p>",
      "rawMarkdown": "Continuing the topic of noise in data, I would like to ask the organizers a specific question: should we expect a clean dataset?\n\nI started noticing posts about numerous pictures with incorrect labels on the forum a couple of weeks ago. This is quite normal practice for Kaggle competitions. A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. \n\nBased on the above, it is obvious that I'd like to know if we'll ever get the clean dataset. The future course of solving this competition for many people (or not solving it) depends on this. At the moment, half of all my ideas and methods are aimed at dealing with the noise.\n\nSpecial loss functions, label smoothing, knowledge distillation are good methods. But none of them is comparable to the manual relabeling of the data. And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?\n\nIt follows that a person who has the specialists for this purpose or the ability to hire them, can ensure a good place in the competition and at the same time devalue other people's solutions? Or is this exactly what the organizers expect us to do so that we just clean their dataset?\n\nThank you in advance.",
      "votes": null
    },
    {
      "id": "1143380",
      "postDate": "01/07/2021 21:47:42",
      "content": "<p>Not all competitions publish clean dataset after the start of the competition. There are some exceptional cases though.<br>\nCompetitions with noisy dataset is not new on Kaggle. If test datasets contains same noise then it's okay I think. Moreover, huge amount of allotted time for this competition have already passed. Release of new dataset in such stage wash the people's work who tried to handle this noise in some way or other.</p>",
      "rawMarkdown": "Not all competitions publish clean dataset after the start of the competition. There are some exceptional cases though.\nCompetitions with noisy dataset is not new on Kaggle. If test datasets contains same noise then it's okay I think. Moreover, huge amount of allotted time for this competition have already passed. Release of new dataset in such stage wash the people's work who tried to handle this noise in some way or other.",
      "votes": null
    },
    {
      "id": "1144120",
      "postDate": "01/08/2021 08:44:37",
      "content": "<blockquote>\n  <p>A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. </p>\n</blockquote>\n<p>Changing the competition training data just because their are label noises ? Do you have examples ? </p>\n<blockquote>\n  <p>And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?</p>\n</blockquote>\n<p>Noisy training data are the most common in real life examples. And making your model robust on them is part of the challenge. And I don't think you need to manually change all these labels ( At least for me , it would be very boring and exhausting task  and somewhat biased by your labelling expertise, while not guaranteeing the model will be robust on unseen test data). </p>",
      "rawMarkdown": ">  A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. \n\nChanging the competition training data just because their are label noises ? Do you have examples ? \n\n\n> And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?\n\nNoisy training data are the most common in real life examples. And making your model robust on them is part of the challenge. And I don't think you need to manually change all these labels ( At least for me , it would be very boring and exhausting task  and somewhat biased by your labelling expertise, while not guaranteeing the model will be robust on unseen test data).",
      "votes": null
    },
    {
      "id": "1144270",
      "postDate": "01/08/2021 11:03:35",
      "content": "<p>At least this: <br>\nhttps//<a href=\"http://www.kaggle.com/c/global-wheat-detection/discussion/163291\" target=\"_blank\">www.kaggle.com/c/global-wheat-detection/discussion/163291</a></p>",
      "rawMarkdown": "At least this: \nhttps//www.kaggle.com/c/global-wheat-detection/discussion/163291",
      "votes": null
    },
    {
      "id": "1144398",
      "postDate": "01/08/2021 12:46:52",
      "content": "<p>I read the topic .  They just changed the incorrect boxes in private test set. <br>\nDespite the \"flaw\" in the training set.</p>\n<p>Here we have much bigger data and no threshold based metric, thus much less sensitive to noisy data points. And we don't know the situation in private data. </p>",
      "rawMarkdown": "I read the topic .  They just changed the incorrect boxes in private test set. \nDespite the \"flaw\" in the training set.\n\nHere we have much bigger data and no threshold based metric, thus much less sensitive to noisy data points. And we don't know the situation in private data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143380,
      "author_name": "zaber666",
      "author_url": "",
      "post_date": "01/07/2021 21:47:42",
      "content": "<p>Not all competitions publish clean dataset after the start of the competition. There are some exceptional cases though.<br>\nCompetitions with noisy dataset is not new on Kaggle. If test datasets contains same noise then it's okay I think. Moreover, huge amount of allotted time for this competition have already passed. Release of new dataset in such stage wash the people's work who tried to handle this noise in some way or other.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144120,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/08/2021 08:44:37",
      "content": "<blockquote>\n  <p>A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. </p>\n</blockquote>\n<p>Changing the competition training data just because their are label noises ? Do you have examples ? </p>\n<blockquote>\n  <p>And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?</p>\n</blockquote>\n<p>Noisy training data are the most common in real life examples. And making your model robust on them is part of the challenge. And I don't think you need to manually change all these labels ( At least for me , it would be very boring and exhausting task  and somewhat biased by your labelling expertise, while not guaranteeing the model will be robust on unseen test data). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1144270,
          "author_name": "vadimtimakin",
          "author_url": "",
          "post_date": "01/08/2021 11:03:35",
          "content": "<p>At least this: <br>\nhttps//<a href=\"http://www.kaggle.com/c/global-wheat-detection/discussion/163291\" target=\"_blank\">www.kaggle.com/c/global-wheat-detection/discussion/163291</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144398,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "01/08/2021 12:46:52",
          "content": "<p>I read the topic .  They just changed the incorrect boxes in private test set. <br>\nDespite the \"flaw\" in the training set.</p>\n<p>Here we have much bigger data and no threshold based metric, thus much less sensitive to noisy data points. And we don't know the situation in private data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1143025": "Continuing the topic of noise in data, I would like to ask the organizers a specific question: should we expect a clean dataset?\n\nI started noticing posts about numerous pictures with incorrect labels on the forum a couple of weeks ago. This is quite normal practice for Kaggle competitions. A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. \n\nBased on the above, it is obvious that I'd like to know if we'll ever get the clean dataset. The future course of solving this competition for many people (or not solving it) depends on this. At the moment, half of all my ideas and methods are aimed at dealing with the noise.\n\nSpecial loss functions, label smoothing, knowledge distillation are good methods. But none of them is comparable to the manual relabeling of the data. And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?\n\nIt follows that a person who has the specialists for this purpose or the ability to hire them, can ensure a good place in the competition and at the same time devalue other people's solutions? Or is this exactly what the organizers expect us to do so that we just clean their dataset?\n\nThank you in advance.",
    "1143380": "Not all competitions publish clean dataset after the start of the competition. There are some exceptional cases though.\nCompetitions with noisy dataset is not new on Kaggle. If test datasets contains same noise then it's okay I think. Moreover, huge amount of allotted time for this competition have already passed. Release of new dataset in such stage wash the people's work who tried to handle this noise in some way or other.",
    "1144120": ">  A couple of weeks after the start of the competition, errors in the markup are detected, and after a while, the organizers usually share a clean dataset. All this time I was waiting for the publication of a clean dataset from the organizers, but this doesn't happen. \n\nChanging the competition training data just because their are label noises ? Do you have examples ? \n\n\n> And even this is normal practice for competitions. But how should I manually change the labels of thousands of plants that are difficult to distinguish from each other?\n\nNoisy training data are the most common in real life examples. And making your model robust on them is part of the challenge. And I don't think you need to manually change all these labels ( At least for me , it would be very boring and exhausting task  and somewhat biased by your labelling expertise, while not guaranteeing the model will be robust on unseen test data).",
    "1144270": "At least this: \nhttps//www.kaggle.com/c/global-wheat-detection/discussion/163291",
    "1144398": "I read the topic .  They just changed the incorrect boxes in private test set. \nDespite the \"flaw\" in the training set.\n\nHere we have much bigger data and no threshold based metric, thus much less sensitive to noisy data points. And we don't know the situation in private data."
  },
  "source": "meta"
}