{
  "id": 41223,
  "title": "Corrupted scans and incorrect labels",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/41223",
  "author_name": "",
  "post_date": "2017-10-14T23:21:02.932071300Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I noticed a couple of issues in the dataset:</p>\n\n<p>A. <code>42181583618ce4bbfbc0c4c300108bf5.aps</code> is corrupted. It looks like the result of a data race, if I had to make a wild guess (Perhaps something was read before it was fully written). It's a problem if this happens with stage 2.</p>\n\n<p>B. About 1-3% of the scans include incorrect labels, in my estimation. Some are glaringly obvious as operator errors. With this rate of mislabeled scans, the \"perfect\" model will get a score of only about 0.04 <code>(-log(1e-15) * 0.02/17)</code></p>\n\n<p>There are three possible ways for Kaggle to address this:</p>\n\n<ol>\n<li>Train a model and examine its \"mistakes\" on stage 2 data (This is biased in favor of similar models)</li>\n<li>Wait for after stage 2 and manually examine some submissions' \"mistakes\" (Biased against other submissions)</li>\n<li>Reexamine all stage 2 labels. I have a PyQt script that I'd be happy to share with Kaggle. (Email me if you want it) It should make it convenient to spot the mislabels in about 6 hours. You would still have to gamify this, with 2+ people competing to find the most mislabels, or something like that. Otherwise, it's very hard not to succumb to the tedium and just scroll through everything.</li>\n</ol>",
  "messages": [
    {
      "id": "231452",
      "postDate": "10/14/2017 23:21:02",
      "content": "<p>I noticed a couple of issues in the dataset:</p>\n\n<p>A. <code>42181583618ce4bbfbc0c4c300108bf5.aps</code> is corrupted. It looks like the result of a data race, if I had to make a wild guess (Perhaps something was read before it was fully written). It's a problem if this happens with stage 2.</p>\n\n<p>B. About 1-3% of the scans include incorrect labels, in my estimation. Some are glaringly obvious as operator errors. With this rate of mislabeled scans, the \"perfect\" model will get a score of only about 0.04 <code>(-log(1e-15) * 0.02/17)</code></p>\n\n<p>There are three possible ways for Kaggle to address this:</p>\n\n<ol>\n<li>Train a model and examine its \"mistakes\" on stage 2 data (This is biased in favor of similar models)</li>\n<li>Wait for after stage 2 and manually examine some submissions' \"mistakes\" (Biased against other submissions)</li>\n<li>Reexamine all stage 2 labels. I have a PyQt script that I'd be happy to share with Kaggle. (Email me if you want it) It should make it convenient to spot the mislabels in about 6 hours. You would still have to gamify this, with 2+ people competing to find the most mislabels, or something like that. Otherwise, it's very hard not to succumb to the tedium and just scroll through everything.</li>\n</ol>",
      "rawMarkdown": "I noticed a couple of issues in the dataset:\n\nA. `42181583618ce4bbfbc0c4c300108bf5.aps` is corrupted. It looks like the result of a data race, if I had to make a wild guess (Perhaps something was read before it was fully written). It's a problem if this happens with stage 2.\n\nB. About 1-3% of the scans include incorrect labels, in my estimation. Some are glaringly obvious as operator errors. With this rate of mislabeled scans, the \"perfect\" model will get a score of only about 0.04 `(-log(1e-15) * 0.02/17)`\n\nThere are three possible ways for Kaggle to address this:\n\n 1. Train a model and examine its \"mistakes\" on stage 2 data (This is biased in favor of similar models)\n 2. Wait for after stage 2 and manually examine some submissions' \"mistakes\" (Biased against other submissions)\n 3. Reexamine all stage 2 labels. I have a PyQt script that I'd be happy to share with Kaggle. (Email me if you want it) It should make it convenient to spot the mislabels in about 6 hours. You would still have to gamify this, with 2+ people competing to find the most mislabels, or something like that. Otherwise, it's very hard not to succumb to the tedium and just scroll through everything.",
      "votes": null
    },
    {
      "id": "232096",
      "postDate": "10/16/2017 20:23:48",
      "content": "<p>I am also concerned about this.  I am not sure about percentage but 1-3% estimate is on the low side. Mislabelings in the train set are something we can handle, however if this will be the case for the stage 2 we are in trouble. There is still time for Kaggle to communicate with the sponsor and ask them for a thorough review of stage 2 data.  I hope they do that.</p>",
      "rawMarkdown": "I am also concerned about this.  I am not sure about percentage but 1-3% estimate is on the low side. Mislabelings in the train set are something we can handle, however if this will be the case for the stage 2 we are in trouble. There is still time for Kaggle to communicate with the sponsor and ask them for a thorough review of stage 2 data.  I hope they do that.",
      "votes": null
    },
    {
      "id": "232659",
      "postDate": "10/18/2017 02:58:25",
      "content": "<p>In the LIDC/LUNA16 dataset, they used multiple experts, who looked for a consensus:</p>\n\n<blockquote>\n  <p>In the initial blinded-read phase, each radiologist independently\n  reviewed each CT scan [...]. In the subsequent unblinded-read phase,\n  each radiologist independently reviewed their own marks along with the\n  anonymized marks of the three other radiologists to render a final\n  opinion.</p>\n</blockquote>\n\n<p>Perhaps Kaggle/DHS could do something similar for the test data here, to eliminate random human error?</p>",
      "rawMarkdown": "In the LIDC/LUNA16 dataset, they used multiple experts, who looked for a consensus:\n\n&gt; In the initial blinded-read phase, each radiologist independently\n&gt; reviewed each CT scan [...]. In the subsequent unblinded-read phase,\n&gt; each radiologist independently reviewed their own marks along with the\n&gt; anonymized marks of the three other radiologists to render a final\n&gt; opinion.\n\nPerhaps Kaggle/DHS could do something similar for the test data here, to eliminate random human error?",
      "votes": null
    },
    {
      "id": "233031",
      "postDate": "10/19/2017 02:48:33",
      "content": "<p>I have not seen 1-3% incorrect labels in the data, apart from the error mentioned in <a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615</a>. Exactly what incorrect labels have you seen?</p>",
      "rawMarkdown": "I have not seen 1-3% incorrect labels in the data, apart from the error mentioned in https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615. Exactly what incorrect labels have you seen?",
      "votes": null
    },
    {
      "id": "233409",
      "postDate": "10/20/2017 06:01:03",
      "content": "<p>-</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "233450",
      "postDate": "10/20/2017 08:09:43",
      "content": "<p>There are 1,247 scans in this dataset.  1-3% implies 12 - 37 mislabels.  I highly doubt there could be that many.  I'm pretty sure the number is closer to 5-7ish so maybe 0.5% realistically.  I also found the corrupted image and did a search for more, that was the only one.  Corrupted images could be caught during stage 2 and hopefully removed from scoring.</p>\n\n<p>Regardless of the exact percentage, I do agree that the mislables are a problem particularly given that LogLoss scoring very heavily penalizes these kind of mistakes.  I'd really like either a thorough review of the stage 2 labels by a panel of humans, or better, allowing stage 2 scores to be corrected after-the-fact if mislabels are discovered (would be quick and easy to find them in a few hours if ground truth is released after stage 2).</p>",
      "rawMarkdown": "There are 1,247 scans in this dataset.  1-3% implies 12 - 37 mislabels.  I highly doubt there could be that many.  I'm pretty sure the number is closer to 5-7ish so maybe 0.5% realistically.  I also found the corrupted image and did a search for more, that was the only one.  Corrupted images could be caught during stage 2 and hopefully removed from scoring.\n\nRegardless of the exact percentage, I do agree that the mislables are a problem particularly given that LogLoss scoring very heavily penalizes these kind of mistakes.  I'd really like either a thorough review of the stage 2 labels by a panel of humans, or better, allowing stage 2 scores to be corrected after-the-fact if mislabels are discovered (would be quick and easy to find them in a few hours if ground truth is released after stage 2).",
      "votes": null
    },
    {
      "id": "233533",
      "postDate": "10/20/2017 13:29:33",
      "content": "<p>If you train a NN on a dataset with mistakes in it, it should automatically learn to hedge against the mistakes by never being too confident. If the mistakes are deleted from the test set, as a surprise, such hedging becomes a handicap. Should Kaggle penalize models that are, in some ML sense, optimal? Whichever they decide, I hope they communicate their intentions to the players.</p>\n\n<p>Do you agree with e4b560b0f6d2c44535610f38a787df93_Zone15,1 ?</p>",
      "rawMarkdown": "If you train a NN on a dataset with mistakes in it, it should automatically learn to hedge against the mistakes by never being too confident. If the mistakes are deleted from the test set, as a surprise, such hedging becomes a handicap. Should Kaggle penalize models that are, in some ML sense, optimal? Whichever they decide, I hope they communicate their intentions to the players.\n\nDo you agree with e4b560b0f6d2c44535610f38a787df93_Zone15,1 ?",
      "votes": null
    },
    {
      "id": "233610",
      "postDate": "10/20/2017 17:17:46",
      "content": "<p>I can't say for sure if e4b560b0f6d2c44535610f38a787df93 is mislabeled...  I kind of see something dark there that could conceivably be a threat, its definitely not obvious to me as a human at least (nor should it be to the algorithm), so I'm not concerned about this type of mislabel if it is one.</p>\n\n<p>I'm more concerned about gross mislabels on zones that don't have mislabels in the stage 1 set.  I have no idea if the mislabeled areas would be consistent between the stages.  What if 2 scans with totally different zones are swapped in stage 2?  It might be that those zones had perfect labels in the training so there could be no way to \"hedge against it\".</p>",
      "rawMarkdown": "I can't say for sure if e4b560b0f6d2c44535610f38a787df93 is mislabeled...  I kind of see something dark there that could conceivably be a threat, its definitely not obvious to me as a human at least (nor should it be to the algorithm), so I'm not concerned about this type of mislabel if it is one.\n\nI'm more concerned about gross mislabels on zones that don't have mislabels in the stage 1 set.  I have no idea if the mislabeled areas would be consistent between the stages.  What if 2 scans with totally different zones are swapped in stage 2?  It might be that those zones had perfect labels in the training so there could be no way to \"hedge against it\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 232096,
      "author_name": "sarkadiu",
      "author_url": "",
      "post_date": "10/16/2017 20:23:48",
      "content": "<p>I am also concerned about this.  I am not sure about percentage but 1-3% estimate is on the low side. Mislabelings in the train set are something we can handle, however if this will be the case for the stage 2 we are in trouble. There is still time for Kaggle to communicate with the sponsor and ask them for a thorough review of stage 2 data.  I hope they do that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 232659,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "10/18/2017 02:58:25",
          "content": "<p>In the LIDC/LUNA16 dataset, they used multiple experts, who looked for a consensus:</p>\n\n<blockquote>\n  <p>In the initial blinded-read phase, each radiologist independently\n  reviewed each CT scan [...]. In the subsequent unblinded-read phase,\n  each radiologist independently reviewed their own marks along with the\n  anonymized marks of the three other radiologists to render a final\n  opinion.</p>\n</blockquote>\n\n<p>Perhaps Kaggle/DHS could do something similar for the test data here, to eliminate random human error?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 233031,
      "author_name": "suchir",
      "author_url": "",
      "post_date": "10/19/2017 02:48:33",
      "content": "<p>I have not seen 1-3% incorrect labels in the data, apart from the error mentioned in <a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615</a>. Exactly what incorrect labels have you seen?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 233409,
      "author_name": "sandshift",
      "author_url": "",
      "post_date": "10/20/2017 06:01:03",
      "content": "<p>-</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 233450,
      "author_name": "hackerpoet",
      "author_url": "",
      "post_date": "10/20/2017 08:09:43",
      "content": "<p>There are 1,247 scans in this dataset.  1-3% implies 12 - 37 mislabels.  I highly doubt there could be that many.  I'm pretty sure the number is closer to 5-7ish so maybe 0.5% realistically.  I also found the corrupted image and did a search for more, that was the only one.  Corrupted images could be caught during stage 2 and hopefully removed from scoring.</p>\n\n<p>Regardless of the exact percentage, I do agree that the mislables are a problem particularly given that LogLoss scoring very heavily penalizes these kind of mistakes.  I'd really like either a thorough review of the stage 2 labels by a panel of humans, or better, allowing stage 2 scores to be corrected after-the-fact if mislabels are discovered (would be quick and easy to find them in a few hours if ground truth is released after stage 2).</p>",
      "votes": null,
      "replies": [
        {
          "id": 233533,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "10/20/2017 13:29:33",
          "content": "<p>If you train a NN on a dataset with mistakes in it, it should automatically learn to hedge against the mistakes by never being too confident. If the mistakes are deleted from the test set, as a surprise, such hedging becomes a handicap. Should Kaggle penalize models that are, in some ML sense, optimal? Whichever they decide, I hope they communicate their intentions to the players.</p>\n\n<p>Do you agree with e4b560b0f6d2c44535610f38a787df93_Zone15,1 ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233610,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "10/20/2017 17:17:46",
          "content": "<p>I can't say for sure if e4b560b0f6d2c44535610f38a787df93 is mislabeled...  I kind of see something dark there that could conceivably be a threat, its definitely not obvious to me as a human at least (nor should it be to the algorithm), so I'm not concerned about this type of mislabel if it is one.</p>\n\n<p>I'm more concerned about gross mislabels on zones that don't have mislabels in the stage 1 set.  I have no idea if the mislabeled areas would be consistent between the stages.  What if 2 scans with totally different zones are swapped in stage 2?  It might be that those zones had perfect labels in the training so there could be no way to \"hedge against it\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "231452": "I noticed a couple of issues in the dataset:\n\nA. `42181583618ce4bbfbc0c4c300108bf5.aps` is corrupted. It looks like the result of a data race, if I had to make a wild guess (Perhaps something was read before it was fully written). It's a problem if this happens with stage 2.\n\nB. About 1-3% of the scans include incorrect labels, in my estimation. Some are glaringly obvious as operator errors. With this rate of mislabeled scans, the \"perfect\" model will get a score of only about 0.04 `(-log(1e-15) * 0.02/17)`\n\nThere are three possible ways for Kaggle to address this:\n\n 1. Train a model and examine its \"mistakes\" on stage 2 data (This is biased in favor of similar models)\n 2. Wait for after stage 2 and manually examine some submissions' \"mistakes\" (Biased against other submissions)\n 3. Reexamine all stage 2 labels. I have a PyQt script that I'd be happy to share with Kaggle. (Email me if you want it) It should make it convenient to spot the mislabels in about 6 hours. You would still have to gamify this, with 2+ people competing to find the most mislabels, or something like that. Otherwise, it's very hard not to succumb to the tedium and just scroll through everything.",
    "232096": "I am also concerned about this.  I am not sure about percentage but 1-3% estimate is on the low side. Mislabelings in the train set are something we can handle, however if this will be the case for the stage 2 we are in trouble. There is still time for Kaggle to communicate with the sponsor and ask them for a thorough review of stage 2 data.  I hope they do that.",
    "232659": "In the LIDC/LUNA16 dataset, they used multiple experts, who looked for a consensus:\n\n&gt; In the initial blinded-read phase, each radiologist independently\n&gt; reviewed each CT scan [...]. In the subsequent unblinded-read phase,\n&gt; each radiologist independently reviewed their own marks along with the\n&gt; anonymized marks of the three other radiologists to render a final\n&gt; opinion.\n\nPerhaps Kaggle/DHS could do something similar for the test data here, to eliminate random human error?",
    "233031": "I have not seen 1-3% incorrect labels in the data, apart from the error mentioned in https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615. Exactly what incorrect labels have you seen?",
    "233409": "",
    "233450": "There are 1,247 scans in this dataset.  1-3% implies 12 - 37 mislabels.  I highly doubt there could be that many.  I'm pretty sure the number is closer to 5-7ish so maybe 0.5% realistically.  I also found the corrupted image and did a search for more, that was the only one.  Corrupted images could be caught during stage 2 and hopefully removed from scoring.\n\nRegardless of the exact percentage, I do agree that the mislables are a problem particularly given that LogLoss scoring very heavily penalizes these kind of mistakes.  I'd really like either a thorough review of the stage 2 labels by a panel of humans, or better, allowing stage 2 scores to be corrected after-the-fact if mislabels are discovered (would be quick and easy to find them in a few hours if ground truth is released after stage 2).",
    "233533": "If you train a NN on a dataset with mistakes in it, it should automatically learn to hedge against the mistakes by never being too confident. If the mistakes are deleted from the test set, as a surprise, such hedging becomes a handicap. Should Kaggle penalize models that are, in some ML sense, optimal? Whichever they decide, I hope they communicate their intentions to the players.\n\nDo you agree with e4b560b0f6d2c44535610f38a787df93_Zone15,1 ?",
    "233610": "I can't say for sure if e4b560b0f6d2c44535610f38a787df93 is mislabeled...  I kind of see something dark there that could conceivably be a threat, its definitely not obvious to me as a human at least (nor should it be to the algorithm), so I'm not concerned about this type of mislabel if it is one.\n\nI'm more concerned about gross mislabels on zones that don't have mislabels in the stage 1 set.  I have no idea if the mislabeled areas would be consistent between the stages.  What if 2 scans with totally different zones are swapped in stage 2?  It might be that those zones had perfect labels in the training so there could be no way to \"hedge against it\"."
  },
  "source": "meta"
}