{
  "id": 45801,
  "title": "Errors in the training and test labels and data",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/45801",
  "author_name": "",
  "post_date": "2017-12-16T01:20:29.192653600Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Training labels have 17 errors and training data has 1 error (see bellow). Total 18 errors for 1147 records. Stage 2 test dataset has 1388 records. If it has the same error ratio, then it would have 22 errors. 22 errors are enough to put winning model out of the winning places (each error contributes up to 0.0015 to the log loss). So, what do you think, maybe community should verify test labels for the stage 2 before you announce winners?</p>\n\n<p>Label errors (original threats – actual threats):</p>\n\n<ul>\n<li><p>496ec724cc1f2886aac5840cf890988a: 5, 11, 16 – 3 (4 errors)</p></li>\n<li><p>52c8235df3f0552e6c134529ca85d958: 4, 9, 15 – 4, 15 (1 error)</p></li>\n<li><p>56b9c0086836fe2fca86d773cacaf783: 4 – 2 (2 errors)</p></li>\n<li><p>6cbda8596c5c9b1e31d4fab9b5a9e02b: 9 – None (1 error)</p></li>\n<li><p>623c761b4db398ea2157e6c5cd6c8c58: 3 – 5, 11, 16 (4 errors)</p></li>\n<li><p>7235e754185d3321c4b6883d001a35ad: 4, 9 – 4 (1 error)</p></li>\n<li><p>cbc6f0a3be3d802fc3d2bd45c183a49d: 4, 7, 13 – 2, 7, 13 (2 errors)</p></li>\n<li><p>d904d73f5e53eed05fef89ce0032fc1c: 2, 9, 11 – 2, 11, 12 (2 errors)</p></li>\n</ul>\n\n<p>Data errors:</p>\n\n<ul>\n<li>42181583618ce4bbfbc0c4c300108bf5: broken data of the aps images #15 and #16 (1 error)\n </li>\n</ul>",
  "messages": [
    {
      "id": "258364",
      "postDate": "12/16/2017 01:20:29",
      "content": "<p>Training labels have 17 errors and training data has 1 error (see bellow). Total 18 errors for 1147 records. Stage 2 test dataset has 1388 records. If it has the same error ratio, then it would have 22 errors. 22 errors are enough to put winning model out of the winning places (each error contributes up to 0.0015 to the log loss). So, what do you think, maybe community should verify test labels for the stage 2 before you announce winners?</p>\n\n<p>Label errors (original threats – actual threats):</p>\n\n<ul>\n<li><p>496ec724cc1f2886aac5840cf890988a: 5, 11, 16 – 3 (4 errors)</p></li>\n<li><p>52c8235df3f0552e6c134529ca85d958: 4, 9, 15 – 4, 15 (1 error)</p></li>\n<li><p>56b9c0086836fe2fca86d773cacaf783: 4 – 2 (2 errors)</p></li>\n<li><p>6cbda8596c5c9b1e31d4fab9b5a9e02b: 9 – None (1 error)</p></li>\n<li><p>623c761b4db398ea2157e6c5cd6c8c58: 3 – 5, 11, 16 (4 errors)</p></li>\n<li><p>7235e754185d3321c4b6883d001a35ad: 4, 9 – 4 (1 error)</p></li>\n<li><p>cbc6f0a3be3d802fc3d2bd45c183a49d: 4, 7, 13 – 2, 7, 13 (2 errors)</p></li>\n<li><p>d904d73f5e53eed05fef89ce0032fc1c: 2, 9, 11 – 2, 11, 12 (2 errors)</p></li>\n</ul>\n\n<p>Data errors:</p>\n\n<ul>\n<li>42181583618ce4bbfbc0c4c300108bf5: broken data of the aps images #15 and #16 (1 error)\n </li>\n</ul>",
      "rawMarkdown": "Training labels have 17 errors and training data has 1 error (see bellow). Total 18 errors for 1147 records. Stage 2 test dataset has 1388 records. If it has the same error ratio, then it would have 22 errors. 22 errors are enough to put winning model out of the winning places (each error contributes up to 0.0015 to the log loss). So, what do you think, maybe community should verify test labels for the stage 2 before you announce winners?\n\nLabel errors (original threats – actual threats):\n\n - 496ec724cc1f2886aac5840cf890988a: 5, 11, 16 – 3 (4 errors)\n\n - 52c8235df3f0552e6c134529ca85d958: 4, 9, 15 – 4, 15 (1 error)\n\n - 56b9c0086836fe2fca86d773cacaf783: 4 – 2 (2 errors)\n\n - 6cbda8596c5c9b1e31d4fab9b5a9e02b: 9 – None (1 error)\n\n - 623c761b4db398ea2157e6c5cd6c8c58: 3 – 5, 11, 16 (4 errors)\n\n - 7235e754185d3321c4b6883d001a35ad: 4, 9 – 4 (1 error)\n\n - cbc6f0a3be3d802fc3d2bd45c183a49d: 4, 7, 13 – 2, 7, 13 (2 errors)\n\n - d904d73f5e53eed05fef89ce0032fc1c: 2, 9, 11 – 2, 11, 12 (2 errors)\n\nData errors:\n\n - 42181583618ce4bbfbc0c4c300108bf5: broken data of the aps images #15 and #16 (1 error)",
      "votes": null
    },
    {
      "id": "258414",
      "postDate": "12/16/2017 03:35:34",
      "content": "<p>The organizers had stated (as I understood it) that the truth values were whatever the DHS said they were.</p>\n\n<p>Based on that, the optimal thing to do was never to predict 0 or 1. For example, when your model thinks the correct label is 1, and you think there's a 0.001 probability of a mislabel there, you should output 0.999.</p>\n\n<p>I'm fairly certain that most of the top entrants did something of this nature.</p>\n\n<p>With this in place, I think the mislabels had a relatively minor, although unfortunate, influence.</p>",
      "rawMarkdown": "The organizers had stated (as I understood it) that the truth values were whatever the DHS said they were.\n\nBased on that, the optimal thing to do was never to predict 0 or 1. For example, when your model thinks the correct label is 1, and you think there's a 0.001 probability of a mislabel there, you should output 0.999.\n\nI'm fairly certain that most of the top entrants did something of this nature.\n\nWith this in place, I think the mislabels had a relatively minor, although unfortunate, influence.",
      "votes": null
    },
    {
      "id": "258416",
      "postDate": "12/16/2017 03:38:48",
      "content": "<p>Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.</p>",
      "rawMarkdown": "Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.",
      "votes": null
    },
    {
      "id": "258421",
      "postDate": "12/16/2017 03:48:12",
      "content": "<blockquote>\n  <p><strong>Dmitry Kovba wrote</strong></p>\n  \n  <blockquote>\n    <p>Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.</p>\n  </blockquote>\n</blockquote>\n\n<p>I think you are miscounting: <code>-log(0.001) / 17 / 1388  =  0.00029</code></p>\n\n<p>Further, everyone's in the same boat. It's not like someone got the clean labels to score against, while others get the above penalty.</p>",
      "rawMarkdown": "&gt; **Dmitry Kovba wrote**\n&gt; \n&gt; &gt; Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.\n\nI think you are miscounting: `-log(0.001) / 17 / 1388  =  0.00029`\n\nFurther, everyone's in the same boat. It's not like someone got the clean labels to score against, while others get the above penalty.",
      "votes": null
    },
    {
      "id": "258426",
      "postDate": "12/16/2017 03:58:29",
      "content": "<p>Thanks for the correction. You are right, the problem happens when the model gives predictions near 0 and 1. Errors in the test data can actually make scores much worse than they should be in this case. Up to 0.0015 per error: <code>-log(1e-15) / 17 / 1388 = 0.0015</code></p>",
      "rawMarkdown": "Thanks for the correction. You are right, the problem happens when the model gives predictions near 0 and 1. Errors in the test data can actually make scores much worse than they should be in this case. Up to 0.0015 per error: `-log(1e-15) / 17 / 1388 = 0.0015`",
      "votes": null
    },
    {
      "id": "258460",
      "postDate": "12/16/2017 05:51:10",
      "content": "<p>Like Oleg said, no model (even an optimal one which always predicts the \"real\" ground truth) should predict 1e-15 given the presence of mislabels, and being this confident has a very small benefit even if all the labels are correct (versus predicting 1e-3, the improvement in loss is only ~0.001).</p>",
      "rawMarkdown": "Like Oleg said, no model (even an optimal one which always predicts the \"real\" ground truth) should predict 1e-15 given the presence of mislabels, and being this confident has a very small benefit even if all the labels are correct (versus predicting 1e-3, the improvement in loss is only ~0.001).",
      "votes": null
    },
    {
      "id": "258645",
      "postDate": "12/16/2017 17:36:22",
      "content": "<p>Thanks for sharing. If you put it that way it looks like the winners score (0.024) is entirely made by the label errors. That is, they made 0 errors and only got the label errors contributing to their score.</p>\n\n<p>How did you find these label errors? You run your model in 5 folds and check where it deviates the most from the ground truth? And manually check those? I found, even manually looking at the images, it always hard to find the threat. With the label as a guidance I could often say \"yeah I see something there\", but I was never so confident that I could say \"it definitely not there but there\".</p>",
      "rawMarkdown": "Thanks for sharing. If you put it that way it looks like the winners score (0.024) is entirely made by the label errors. That is, they made 0 errors and only got the label errors contributing to their score.\n\nHow did you find these label errors? You run your model in 5 folds and check where it deviates the most from the ground truth? And manually check those? I found, even manually looking at the images, it always hard to find the threat. With the label as a guidance I could often say \"yeah I see something there\", but I was never so confident that I could say \"it definitely not there but there\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 258414,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "12/16/2017 03:35:34",
      "content": "<p>The organizers had stated (as I understood it) that the truth values were whatever the DHS said they were.</p>\n\n<p>Based on that, the optimal thing to do was never to predict 0 or 1. For example, when your model thinks the correct label is 1, and you think there's a 0.001 probability of a mislabel there, you should output 0.999.</p>\n\n<p>I'm fairly certain that most of the top entrants did something of this nature.</p>\n\n<p>With this in place, I think the mislabels had a relatively minor, although unfortunate, influence.</p>",
      "votes": null,
      "replies": [
        {
          "id": 258416,
          "author_name": "dmitrykovba",
          "author_url": "",
          "post_date": "12/16/2017 03:38:48",
          "content": "<p>Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258421,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/16/2017 03:48:12",
          "content": "<blockquote>\n  <p><strong>Dmitry Kovba wrote</strong></p>\n  \n  <blockquote>\n    <p>Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.</p>\n  </blockquote>\n</blockquote>\n\n<p>I think you are miscounting: <code>-log(0.001) / 17 / 1388  =  0.00029</code></p>\n\n<p>Further, everyone's in the same boat. It's not like someone got the clean labels to score against, while others get the above penalty.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258426,
          "author_name": "dmitrykovba",
          "author_url": "",
          "post_date": "12/16/2017 03:58:29",
          "content": "<p>Thanks for the correction. You are right, the problem happens when the model gives predictions near 0 and 1. Errors in the test data can actually make scores much worse than they should be in this case. Up to 0.0015 per error: <code>-log(1e-15) / 17 / 1388 = 0.0015</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258460,
          "author_name": "suchir",
          "author_url": "",
          "post_date": "12/16/2017 05:51:10",
          "content": "<p>Like Oleg said, no model (even an optimal one which always predicts the \"real\" ground truth) should predict 1e-15 given the presence of mislabels, and being this confident has a very small benefit even if all the labels are correct (versus predicting 1e-3, the improvement in loss is only ~0.001).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258645,
      "author_name": "bastiaanbergman",
      "author_url": "",
      "post_date": "12/16/2017 17:36:22",
      "content": "<p>Thanks for sharing. If you put it that way it looks like the winners score (0.024) is entirely made by the label errors. That is, they made 0 errors and only got the label errors contributing to their score.</p>\n\n<p>How did you find these label errors? You run your model in 5 folds and check where it deviates the most from the ground truth? And manually check those? I found, even manually looking at the images, it always hard to find the threat. With the label as a guidance I could often say \"yeah I see something there\", but I was never so confident that I could say \"it definitely not there but there\".</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "258364": "Training labels have 17 errors and training data has 1 error (see bellow). Total 18 errors for 1147 records. Stage 2 test dataset has 1388 records. If it has the same error ratio, then it would have 22 errors. 22 errors are enough to put winning model out of the winning places (each error contributes up to 0.0015 to the log loss). So, what do you think, maybe community should verify test labels for the stage 2 before you announce winners?\n\nLabel errors (original threats – actual threats):\n\n - 496ec724cc1f2886aac5840cf890988a: 5, 11, 16 – 3 (4 errors)\n\n - 52c8235df3f0552e6c134529ca85d958: 4, 9, 15 – 4, 15 (1 error)\n\n - 56b9c0086836fe2fca86d773cacaf783: 4 – 2 (2 errors)\n\n - 6cbda8596c5c9b1e31d4fab9b5a9e02b: 9 – None (1 error)\n\n - 623c761b4db398ea2157e6c5cd6c8c58: 3 – 5, 11, 16 (4 errors)\n\n - 7235e754185d3321c4b6883d001a35ad: 4, 9 – 4 (1 error)\n\n - cbc6f0a3be3d802fc3d2bd45c183a49d: 4, 7, 13 – 2, 7, 13 (2 errors)\n\n - d904d73f5e53eed05fef89ce0032fc1c: 2, 9, 11 – 2, 11, 12 (2 errors)\n\nData errors:\n\n - 42181583618ce4bbfbc0c4c300108bf5: broken data of the aps images #15 and #16 (1 error)",
    "258414": "The organizers had stated (as I understood it) that the truth values were whatever the DHS said they were.\n\nBased on that, the optimal thing to do was never to predict 0 or 1. For example, when your model thinks the correct label is 1, and you think there's a 0.001 probability of a mislabel there, you should output 0.999.\n\nI'm fairly certain that most of the top entrants did something of this nature.\n\nWith this in place, I think the mislabels had a relatively minor, although unfortunate, influence.",
    "258416": "Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.",
    "258421": "&gt; **Dmitry Kovba wrote**\n&gt; \n&gt; &gt; Prediction of 0.001 where the truth is 1 adds 0.005 to your final log loss. 12 errors like this and you are out of the winning places even if you had the first place.\n\nI think you are miscounting: `-log(0.001) / 17 / 1388  =  0.00029`\n\nFurther, everyone's in the same boat. It's not like someone got the clean labels to score against, while others get the above penalty.",
    "258426": "Thanks for the correction. You are right, the problem happens when the model gives predictions near 0 and 1. Errors in the test data can actually make scores much worse than they should be in this case. Up to 0.0015 per error: `-log(1e-15) / 17 / 1388 = 0.0015`",
    "258460": "Like Oleg said, no model (even an optimal one which always predicts the \"real\" ground truth) should predict 1e-15 given the presence of mislabels, and being this confident has a very small benefit even if all the labels are correct (versus predicting 1e-3, the improvement in loss is only ~0.001).",
    "258645": "Thanks for sharing. If you put it that way it looks like the winners score (0.024) is entirely made by the label errors. That is, they made 0 errors and only got the label errors contributing to their score.\n\nHow did you find these label errors? You run your model in 5 folds and check where it deviates the most from the ground truth? And manually check those? I found, even manually looking at the images, it always hard to find the threat. With the label as a guidance I could often say \"yeah I see something there\", but I was never so confident that I could say \"it definitely not there but there\"."
  },
  "source": "meta"
}