{
  "id": 202837,
  "title": "What about the noise in the private test data???",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202837",
  "author_name": "",
  "post_date": "2020-12-12T08:23:42.073392100Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have come across many discussion threads like these <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673\" target=\"_blank\">here</a> signifying the presence of the noise in terms of wrongly labeled images. The claim seems to be strong though, and it could be possible because the labeling is done by hand. It could also be observed that there are some images in the test set as well which could be noisy. We could handle that too. But what about the images in the private test data? Even if we are able to tackle the problem of noise from the training and public test data, we might end up predicting wrong values for the private test dataset as it was meant to be a noise according to us but was termed correct by the competition organizers i.e. the experts at the Makerere University.</p>\n<p>Is there any way to evade that misclassification? Else, if we handle the noisy images too well, we might see a huge shake-up!</p>",
  "messages": [
    {
      "id": "1109935",
      "postDate": "12/12/2020 08:23:42",
      "content": "<p>I have come across many discussion threads like these <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673\" target=\"_blank\">here</a> signifying the presence of the noise in terms of wrongly labeled images. The claim seems to be strong though, and it could be possible because the labeling is done by hand. It could also be observed that there are some images in the test set as well which could be noisy. We could handle that too. But what about the images in the private test data? Even if we are able to tackle the problem of noise from the training and public test data, we might end up predicting wrong values for the private test dataset as it was meant to be a noise according to us but was termed correct by the competition organizers i.e. the experts at the Makerere University.</p>\n<p>Is there any way to evade that misclassification? Else, if we handle the noisy images too well, we might see a huge shake-up!</p>",
      "rawMarkdown": "I have come across many discussion threads like these [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017) and [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673) signifying the presence of the noise in terms of wrongly labeled images. The claim seems to be strong though, and it could be possible because the labeling is done by hand. It could also be observed that there are some images in the test set as well which could be noisy. We could handle that too. But what about the images in the private test data? Even if we are able to tackle the problem of noise from the training and public test data, we might end up predicting wrong values for the private test dataset as it was meant to be a noise according to us but was termed correct by the competition organizers i.e. the experts at the Makerere University.\n\nIs there any way to evade that misclassification? Else, if we handle the noisy images too well, we might see a huge shake-up!",
      "votes": null
    },
    {
      "id": "1109956",
      "postDate": "12/12/2020 08:44:58",
      "content": "<p>noise estimation and how to make your model robust to the wrong label is part of the competition (and maybe part of the real life)</p>\n<p>each kaggle competition has its own data problem. no two competitions are the same</p>",
      "rawMarkdown": "noise estimation and how to make your model robust to the wrong label is part of the competition (and maybe part of the real life)\n\neach kaggle competition has its own data problem. no two competitions are the same",
      "votes": null
    },
    {
      "id": "1109960",
      "postDate": "12/12/2020 08:54:16",
      "content": "<p>Perhaps, you're right!</p>",
      "rawMarkdown": "Perhaps, you're right!",
      "votes": null
    },
    {
      "id": "1110844",
      "postDate": "12/13/2020 06:12:01",
      "content": "<p>Label smoothing and augs like cutmix,mixup and fmix could help to make the model robust I suppose</p>",
      "rawMarkdown": "Label smoothing and augs like cutmix,mixup and fmix could help to make the model robust I suppose",
      "votes": null
    },
    {
      "id": "1148993",
      "postDate": "01/11/2021 14:17:56",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  do we know if test is also having same issue as train ?</p>",
      "rawMarkdown": "hengck23  do we know if test is also having same issue as train ?",
      "votes": null
    },
    {
      "id": "1180412",
      "postDate": "02/01/2021 08:32:02",
      "content": "<p>It is time to talk about this….</p>",
      "rawMarkdown": "It is time to talk about this....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1109956,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/12/2020 08:44:58",
      "content": "<p>noise estimation and how to make your model robust to the wrong label is part of the competition (and maybe part of the real life)</p>\n<p>each kaggle competition has its own data problem. no two competitions are the same</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109960,
          "author_name": "divyansh22",
          "author_url": "",
          "post_date": "12/12/2020 08:54:16",
          "content": "<p>Perhaps, you're right!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148993,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/11/2021 14:17:56",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  do we know if test is also having same issue as train ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1110844,
      "author_name": "kingofarmy",
      "author_url": "",
      "post_date": "12/13/2020 06:12:01",
      "content": "<p>Label smoothing and augs like cutmix,mixup and fmix could help to make the model robust I suppose</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1180412,
      "author_name": "zekunn",
      "author_url": "",
      "post_date": "02/01/2021 08:32:02",
      "content": "<p>It is time to talk about this….</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1109935": "I have come across many discussion threads like these [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017) and [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673) signifying the presence of the noise in terms of wrongly labeled images. The claim seems to be strong though, and it could be possible because the labeling is done by hand. It could also be observed that there are some images in the test set as well which could be noisy. We could handle that too. But what about the images in the private test data? Even if we are able to tackle the problem of noise from the training and public test data, we might end up predicting wrong values for the private test dataset as it was meant to be a noise according to us but was termed correct by the competition organizers i.e. the experts at the Makerere University.\n\nIs there any way to evade that misclassification? Else, if we handle the noisy images too well, we might see a huge shake-up!",
    "1109956": "noise estimation and how to make your model robust to the wrong label is part of the competition (and maybe part of the real life)\n\neach kaggle competition has its own data problem. no two competitions are the same",
    "1109960": "Perhaps, you're right!",
    "1110844": "Label smoothing and augs like cutmix,mixup and fmix could help to make the model robust I suppose",
    "1148993": "hengck23  do we know if test is also having same issue as train ?",
    "1180412": "It is time to talk about this...."
  },
  "source": "meta"
}