{
  "id": 20242,
  "title": "More on mislabeled training data",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20242",
  "author_name": "",
  "post_date": "2016-04-18T20:13:50.617Z",
  "votes": 1,
  "comment_count": 2,
  "views": 671,
  "content": "<p>It sure looks (to me at least) that some of the classes are badly mislabeled. I sampled some of the training data and, at for some of the categories there are a lot of mistakes -- I get around 20% in the c9 (talking to passenger) category, for example. Anyone else running into this problem?</p>",
  "messages": [
    {
      "id": "115473",
      "postDate": "04/18/2016 20:13:50",
      "content": "<p>It sure looks (to me at least) that some of the classes are badly mislabeled. I sampled some of the training data and, at for some of the categories there are a lot of mistakes -- I get around 20% in the c9 (talking to passenger) category, for example. Anyone else running into this problem?</p>",
      "rawMarkdown": "It sure looks (to me at least) that some of the classes are badly mislabeled. I sampled some of the training data and, at for some of the categories there are a lot of mistakes -- I get around 20% in the c9 (talking to passenger) category, for example. Anyone else running into this problem?",
      "votes": null
    },
    {
      "id": "115490",
      "postDate": "04/18/2016 22:49:54",
      "content": "<p>I am also having the same concern. Althogh, we can exclude or re-label the mislabeled images in the training set. I am a bit worried, about a possible misslabeled images in the evaluation set, that might lead to the improper judgment of submissions.</p>",
      "rawMarkdown": "I am also having the same concern. Althogh, we can exclude or re-label the mislabeled images in the training set. I am a bit worried, about a possible misslabeled images in the evaluation set, that might lead to the improper judgment of submissions.",
      "votes": null
    },
    {
      "id": "115492",
      "postDate": "04/18/2016 23:23:46",
      "content": "<p>It's even worse than I thought -- if you look at the c0 images for the woman labeled as p081, her face is turned to the right and she is clearly talking to someone in nearly 50% of the images -- very similar to the images of her that are classified as c9. And that isn't the only example. At this point, I don't have a lot of confidence that the test dataset is properly labeled. </p>",
      "rawMarkdown": "It's even worse than I thought -- if you look at the c0 images for the woman labeled as p081, her face is turned to the right and she is clearly talking to someone in nearly 50% of the images -- very similar to the images of her that are classified as c9. And that isn't the only example. At this point, I don't have a lot of confidence that the test dataset is properly labeled.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115490,
      "author_name": "malzantot",
      "author_url": "",
      "post_date": "04/18/2016 22:49:54",
      "content": "<p>I am also having the same concern. Althogh, we can exclude or re-label the mislabeled images in the training set. I am a bit worried, about a possible misslabeled images in the evaluation set, that might lead to the improper judgment of submissions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115492,
      "author_name": "spammy",
      "author_url": "",
      "post_date": "04/18/2016 23:23:46",
      "content": "<p>It's even worse than I thought -- if you look at the c0 images for the woman labeled as p081, her face is turned to the right and she is clearly talking to someone in nearly 50% of the images -- very similar to the images of her that are classified as c9. And that isn't the only example. At this point, I don't have a lot of confidence that the test dataset is properly labeled. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115473": "It sure looks (to me at least) that some of the classes are badly mislabeled. I sampled some of the training data and, at for some of the categories there are a lot of mistakes -- I get around 20% in the c9 (talking to passenger) category, for example. Anyone else running into this problem?",
    "115490": "I am also having the same concern. Althogh, we can exclude or re-label the mislabeled images in the training set. I am a bit worried, about a possible misslabeled images in the evaluation set, that might lead to the improper judgment of submissions.",
    "115492": "It's even worse than I thought -- if you look at the c0 images for the woman labeled as p081, her face is turned to the right and she is clearly talking to someone in nearly 50% of the images -- very similar to the images of her that are classified as c9. And that isn't the only example. At this point, I don't have a lot of confidence that the test dataset is properly labeled."
  },
  "source": "meta"
}