{
  "id": 200901,
  "title": "erroneous labeling",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200901",
  "author_name": "",
  "post_date": "2020-12-02T10:38:45.905102500Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, i have a question about labels.<br>\nit seems there are a lot of erroneous labels in training dataset.<br>\nfor example:<br>\n2782668721.jpg 4<br>\n3238704279.jpg 4<br>\n1365612235.jpg 4<br>\n1649500149.jpg 4<br>\n1236952675.jpg 4<br>\n3085440105.jpg 4<br>\n2929245875.jpg 4<br>\n2509491848.jpg 4<br>\n1227531167.jpg 4<br>\n…</p>\n<p>how about the test dataset?<br>\nare they so dirty as well?<br>\nif so, how do we (or AI/ML) handel erroneous labels?</p>",
  "messages": [
    {
      "id": "1099398",
      "postDate": "12/02/2020 10:38:45",
      "content": "<p>Hi, i have a question about labels.<br>\nit seems there are a lot of erroneous labels in training dataset.<br>\nfor example:<br>\n2782668721.jpg 4<br>\n3238704279.jpg 4<br>\n1365612235.jpg 4<br>\n1649500149.jpg 4<br>\n1236952675.jpg 4<br>\n3085440105.jpg 4<br>\n2929245875.jpg 4<br>\n2509491848.jpg 4<br>\n1227531167.jpg 4<br>\n…</p>\n<p>how about the test dataset?<br>\nare they so dirty as well?<br>\nif so, how do we (or AI/ML) handel erroneous labels?</p>",
      "rawMarkdown": "Hi, i have a question about labels.\nit seems there are a lot of erroneous labels in training dataset.\nfor example:\n2782668721.jpg 4\n3238704279.jpg 4\n1365612235.jpg 4\n1649500149.jpg 4\n1236952675.jpg 4\n3085440105.jpg 4\n2929245875.jpg 4\n2509491848.jpg 4\n1227531167.jpg 4\n...\n\nhow about the test dataset?\nare they so dirty as well?\nif so, how do we (or AI/ML) handel erroneous labels?",
      "votes": null
    },
    {
      "id": "1099425",
      "postDate": "12/02/2020 10:52:26",
      "content": "<p>i'm thinking that,<br>\nif 10% pictures are wrongly labeled in \"train\" dataset, it should be ok that smart ML can learn well because there are still 90% pictures correctly labeled.<br>\nBut if the \"test\" dataset have 10% wrong labels, and only god knows what labels they are, then how do we solve it to increase score?</p>",
      "rawMarkdown": "i'm thinking that,\nif 10% pictures are wrongly labeled in \"train\" dataset, it should be ok that smart ML can learn well because there are still 90% pictures correctly labeled.\nBut if the \"test\" dataset have 10% wrong labels, and only god knows what labels they are, then how do we solve it to increase score?",
      "votes": null
    },
    {
      "id": "1100446",
      "postDate": "12/03/2020 05:03:30",
      "content": "<p>If the test set also has wrong labels, then the idea of ​​using pseudo-labels does not seem to work, because as you said, the distribution of test and training data is almost the same, so deleting data with low confidence is not very useful.</p>",
      "rawMarkdown": "If the test set also has wrong labels, then the idea of ​​using pseudo-labels does not seem to work, because as you said, the distribution of test and training data is almost the same, so deleting data with low confidence is not very useful.",
      "votes": null
    },
    {
      "id": "1100451",
      "postDate": "12/03/2020 05:06:09",
      "content": "<p>There is another possibility that some key features of the test set may not be available in the training set. Then the use of pseudo-labels will work, but the premise is that your training set must contain pseudo-labels with the test set to learn the key features of the test set .</p>",
      "rawMarkdown": "There is another possibility that some key features of the test set may not be available in the training set. Then the use of pseudo-labels will work, but the premise is that your training set must contain pseudo-labels with the test set to learn the key features of the test set .",
      "votes": null
    },
    {
      "id": "1100452",
      "postDate": "12/03/2020 05:08:04",
      "content": "<p>If the features in the test set are in the training set, and the test set has no wrong labels, then just delete the low-confidence labels in the training set that seem to be wrong. This is what I am doing now.</p>",
      "rawMarkdown": "If the features in the test set are in the training set, and the test set has no wrong labels, then just delete the low-confidence labels in the training set that seem to be wrong. This is what I am doing now.",
      "votes": null
    },
    {
      "id": "1100645",
      "postDate": "12/03/2020 08:03:47",
      "content": "<p>Models are usually robust to non systematic label noise. In fact you train a model run prediction on full set check false negative healthy image and you will notice that model is capable of learning non healthy images correctly even when some of them are labelled as healthy. The bigger problem is the label noise in the hidden test set, there is no way know that.</p>",
      "rawMarkdown": "Models are usually robust to non systematic label noise. In fact you train a model run prediction on full set check false negative healthy image and you will notice that model is capable of learning non healthy images correctly even when some of them are labelled as healthy. The bigger problem is the label noise in the hidden test set, there is no way know that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1099425,
      "author_name": "shihyung",
      "author_url": "",
      "post_date": "12/02/2020 10:52:26",
      "content": "<p>i'm thinking that,<br>\nif 10% pictures are wrongly labeled in \"train\" dataset, it should be ok that smart ML can learn well because there are still 90% pictures correctly labeled.<br>\nBut if the \"test\" dataset have 10% wrong labels, and only god knows what labels they are, then how do we solve it to increase score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1100446,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/03/2020 05:03:30",
      "content": "<p>If the test set also has wrong labels, then the idea of ​​using pseudo-labels does not seem to work, because as you said, the distribution of test and training data is almost the same, so deleting data with low confidence is not very useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1100451,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/03/2020 05:06:09",
      "content": "<p>There is another possibility that some key features of the test set may not be available in the training set. Then the use of pseudo-labels will work, but the premise is that your training set must contain pseudo-labels with the test set to learn the key features of the test set .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1100452,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/03/2020 05:08:04",
      "content": "<p>If the features in the test set are in the training set, and the test set has no wrong labels, then just delete the low-confidence labels in the training set that seem to be wrong. This is what I am doing now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1100645,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "12/03/2020 08:03:47",
      "content": "<p>Models are usually robust to non systematic label noise. In fact you train a model run prediction on full set check false negative healthy image and you will notice that model is capable of learning non healthy images correctly even when some of them are labelled as healthy. The bigger problem is the label noise in the hidden test set, there is no way know that.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1099398": "Hi, i have a question about labels.\nit seems there are a lot of erroneous labels in training dataset.\nfor example:\n2782668721.jpg 4\n3238704279.jpg 4\n1365612235.jpg 4\n1649500149.jpg 4\n1236952675.jpg 4\n3085440105.jpg 4\n2929245875.jpg 4\n2509491848.jpg 4\n1227531167.jpg 4\n...\n\nhow about the test dataset?\nare they so dirty as well?\nif so, how do we (or AI/ML) handel erroneous labels?",
    "1099425": "i'm thinking that,\nif 10% pictures are wrongly labeled in \"train\" dataset, it should be ok that smart ML can learn well because there are still 90% pictures correctly labeled.\nBut if the \"test\" dataset have 10% wrong labels, and only god knows what labels they are, then how do we solve it to increase score?",
    "1100446": "If the test set also has wrong labels, then the idea of ​​using pseudo-labels does not seem to work, because as you said, the distribution of test and training data is almost the same, so deleting data with low confidence is not very useful.",
    "1100451": "There is another possibility that some key features of the test set may not be available in the training set. Then the use of pseudo-labels will work, but the premise is that your training set must contain pseudo-labels with the test set to learn the key features of the test set .",
    "1100452": "If the features in the test set are in the training set, and the test set has no wrong labels, then just delete the low-confidence labels in the training set that seem to be wrong. This is what I am doing now.",
    "1100645": "Models are usually robust to non systematic label noise. In fact you train a model run prediction on full set check false negative healthy image and you will notice that model is capable of learning non healthy images correctly even when some of them are labelled as healthy. The bigger problem is the label noise in the hidden test set, there is no way know that."
  },
  "source": "meta"
}