{
  "id": 240471,
  "title": "Dataset annotation review",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/240471",
  "author_name": "",
  "post_date": "2021-05-20T02:45:03.459754600Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello fellow scholars,</p>\n<p>During the training and validation sessions on the Plant Pathology dataset I´ve found something strange. The model was predicting the right classification (as per my judgment) but the actual label on train.csv had the wrong annotation.</p>\n<p>I´ve decided to take a further look on the dataset and I´ve found several cases that were in this condition. So I´ve created the following notebook to compile some findings.</p>\n<p><a href=\"https://www.kaggle.com/dandierlean/dataset-evaluation\" target=\"_blank\">https://www.kaggle.com/dandierlean/dataset-evaluation</a></p>\n<p>I´m no Botanist or Apple Leaf Disease specialist, but a few inputs from the community would be great.</p>",
  "messages": [
    {
      "id": "1315623",
      "postDate": "05/20/2021 02:45:03",
      "content": "<p>Hello fellow scholars,</p>\n<p>During the training and validation sessions on the Plant Pathology dataset I´ve found something strange. The model was predicting the right classification (as per my judgment) but the actual label on train.csv had the wrong annotation.</p>\n<p>I´ve decided to take a further look on the dataset and I´ve found several cases that were in this condition. So I´ve created the following notebook to compile some findings.</p>\n<p><a href=\"https://www.kaggle.com/dandierlean/dataset-evaluation\" target=\"_blank\">https://www.kaggle.com/dandierlean/dataset-evaluation</a></p>\n<p>I´m no Botanist or Apple Leaf Disease specialist, but a few inputs from the community would be great.</p>",
      "rawMarkdown": "Hello fellow scholars,\n\nDuring the training and validation sessions on the Plant Pathology dataset I´ve found something strange. The model was predicting the right classification (as per my judgment) but the actual label on train.csv had the wrong annotation.\n\nI´ve decided to take a further look on the dataset and I´ve found several cases that were in this condition. So I´ve created the following notebook to compile some findings.\n\nhttps://www.kaggle.com/dandierlean/dataset-evaluation\n\nI´m no Botanist or Apple Leaf Disease specialist, but a few inputs from the community would be great.",
      "votes": null
    },
    {
      "id": "1316626",
      "postDate": "05/20/2021 18:08:29",
      "content": "<p>Yep, there are a lot of miss-labeled samples. The label is noisy. I also found duplicates that were labeled differently, so there is quite a bit of noise in the labels in the dateset.</p>",
      "rawMarkdown": "Yep, there are a lot of miss-labeled samples. The label is noisy. I also found duplicates that were labeled differently, so there is quite a bit of noise in the labels in the dateset.",
      "votes": null
    },
    {
      "id": "1316643",
      "postDate": "05/20/2021 18:30:13",
      "content": "<p>Exactly. But the implication is that the hidden Test dataset is also noisy. So a lot of classification errors might be that the label is incorrect instead of the model prediction. </p>\n<p>This applies to obvious misclassifications (leaf with rust labelled as healthy) or for the most difficult ones (uncertainty between a healthy leaf and one with a initial scab contamination).</p>",
      "rawMarkdown": "Exactly. But the implication is that the hidden Test dataset is also noisy. So a lot of classification errors might be that the label is incorrect instead of the model prediction. \n\nThis applies to obvious misclassifications (leaf with rust labelled as healthy) or for the most difficult ones (uncertainty between a healthy leaf and one with a initial scab contamination).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1316626,
      "author_name": "datasciencegeek",
      "author_url": "",
      "post_date": "05/20/2021 18:08:29",
      "content": "<p>Yep, there are a lot of miss-labeled samples. The label is noisy. I also found duplicates that were labeled differently, so there is quite a bit of noise in the labels in the dateset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1316643,
          "author_name": "dandierlean",
          "author_url": "",
          "post_date": "05/20/2021 18:30:13",
          "content": "<p>Exactly. But the implication is that the hidden Test dataset is also noisy. So a lot of classification errors might be that the label is incorrect instead of the model prediction. </p>\n<p>This applies to obvious misclassifications (leaf with rust labelled as healthy) or for the most difficult ones (uncertainty between a healthy leaf and one with a initial scab contamination).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1315623": "Hello fellow scholars,\n\nDuring the training and validation sessions on the Plant Pathology dataset I´ve found something strange. The model was predicting the right classification (as per my judgment) but the actual label on train.csv had the wrong annotation.\n\nI´ve decided to take a further look on the dataset and I´ve found several cases that were in this condition. So I´ve created the following notebook to compile some findings.\n\nhttps://www.kaggle.com/dandierlean/dataset-evaluation\n\nI´m no Botanist or Apple Leaf Disease specialist, but a few inputs from the community would be great.",
    "1316626": "Yep, there are a lot of miss-labeled samples. The label is noisy. I also found duplicates that were labeled differently, so there is quite a bit of noise in the labels in the dateset.",
    "1316643": "Exactly. But the implication is that the hidden Test dataset is also noisy. So a lot of classification errors might be that the label is incorrect instead of the model prediction. \n\nThis applies to obvious misclassifications (leaf with rust labelled as healthy) or for the most difficult ones (uncertainty between a healthy leaf and one with a initial scab contamination)."
  },
  "source": "meta"
}