{
  "id": 43710,
  "title": "What is kaggle's policy on mislabels in the stage 2 data?",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/43710",
  "author_name": "",
  "post_date": "2017-11-18T08:41:21.741049600Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Given the presence of at least a few unambiguous mislabels in the training data (e.g <a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615</a>), I'm wondering what kaggle's policy on mislabels on the test / stage 2 data is? Since different assumptions about this policy will probably lead to competitors tweaking their models in different ways (and could result in pretty significant changes in the final loss given how log-loss penalizes mislabels), it seems important to have this be clear before the stage 2 data is released. Should we expect:</p>\n\n<ol>\n<li>the labels are the ones provided by the DHS, completely unaltered?</li>\n<li>the labels are the ones provided by the DHS, but have been reviewed more thoroughly than the stage 1 data and potentially revised by Kaggle or the DHS?</li>\n<li>Kaggle will release the stage 2 labels after submissions have been made, and allow competitors to submit revisions to the labels, regrading the leaderboard if necessary?</li>\n<li><p>other?</p>\n\n<p>I'm guessing it's 1.), but I'm new here and am not sure if there's any precedent for this stuff, so it would be nice if someone could clear this up :)</p></li>\n</ol>",
  "messages": [
    {
      "id": "245380",
      "postDate": "11/18/2017 08:41:21",
      "content": "<p>Given the presence of at least a few unambiguous mislabels in the training data (e.g <a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615</a>), I'm wondering what kaggle's policy on mislabels on the test / stage 2 data is? Since different assumptions about this policy will probably lead to competitors tweaking their models in different ways (and could result in pretty significant changes in the final loss given how log-loss penalizes mislabels), it seems important to have this be clear before the stage 2 data is released. Should we expect:</p>\n\n<ol>\n<li>the labels are the ones provided by the DHS, completely unaltered?</li>\n<li>the labels are the ones provided by the DHS, but have been reviewed more thoroughly than the stage 1 data and potentially revised by Kaggle or the DHS?</li>\n<li>Kaggle will release the stage 2 labels after submissions have been made, and allow competitors to submit revisions to the labels, regrading the leaderboard if necessary?</li>\n<li><p>other?</p>\n\n<p>I'm guessing it's 1.), but I'm new here and am not sure if there's any precedent for this stuff, so it would be nice if someone could clear this up :)</p></li>\n</ol>",
      "rawMarkdown": "Given the presence of at least a few unambiguous mislabels in the training data (e.g https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615), I'm wondering what kaggle's policy on mislabels on the test / stage 2 data is? Since different assumptions about this policy will probably lead to competitors tweaking their models in different ways (and could result in pretty significant changes in the final loss given how log-loss penalizes mislabels), it seems important to have this be clear before the stage 2 data is released. Should we expect:\n\n 1. the labels are the ones provided by the DHS, completely unaltered?\n 2. the labels are the ones provided by the DHS, but have been reviewed more thoroughly than the stage 1 data and potentially revised by Kaggle or the DHS?\n 3. Kaggle will release the stage 2 labels after submissions have been made, and allow competitors to submit revisions to the labels, regrading the leaderboard if necessary?\n 4. other?\n\n I'm guessing it's 1.), but I'm new here and am not sure if there's any precedent for this stuff, so it would be nice if someone could clear this up :)",
      "votes": null
    },
    {
      "id": "245779",
      "postDate": "11/19/2017 17:02:09",
      "content": "<p>Mislabels span a continuum, from a few definite ones, to mislabels that are likely, but are less certain.</p>\n\n<p>Whatever the organizers decide to do about them, I hope we know the details, such as what percentage of labels were removed or altered, if any.</p>\n\n<p>If stage 2 labels will be be error-free, it makes sense to train one's model on cleaned up stage 1 data.</p>\n\n<p>On the other hand, if stage 2 labels are expected to have the same proportion of errors, this approach may be problematic.</p>",
      "rawMarkdown": "Mislabels span a continuum, from a few definite ones, to mislabels that are likely, but are less certain.\n\nWhatever the organizers decide to do about them, I hope we know the details, such as what percentage of labels were removed or altered, if any.\n\nIf stage 2 labels will be be error-free, it makes sense to train one's model on cleaned up stage 1 data.\n\nOn the other hand, if stage 2 labels are expected to have the same proportion of errors, this approach may be problematic.",
      "votes": null
    },
    {
      "id": "247231",
      "postDate": "11/22/2017 16:28:55",
      "content": "<p>Upvoting on this. It would be very helpful if Kaggle/TSA could clarify whether the labeling (truthing) process changed or not for stage 2. In other words, can we expect approximately the same mislabel error rate in stages 1 and 2?</p>",
      "rawMarkdown": "Upvoting on this. It would be very helpful if Kaggle/TSA could clarify whether the labeling (truthing) process changed or not for stage 2. In other words, can we expect approximately the same mislabel error rate in stages 1 and 2?",
      "votes": null
    },
    {
      "id": "247243",
      "postDate": "11/22/2017 16:40:03",
      "content": "<p>We processed stage 1 and 2 data in the exact same manner, so you should expect (1) the labels are the ones provided by the DHS, completely unaltered. Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).</p>",
      "rawMarkdown": "We processed stage 1 and 2 data in the exact same manner, so you should expect (1) the labels are the ones provided by the DHS, completely unaltered. Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).",
      "votes": null
    },
    {
      "id": "250652",
      "postDate": "11/30/2017 07:13:02",
      "content": "<blockquote>\n  <p>Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).</p>\n</blockquote>\n\n<p>They do arbitrarily constrain how good models can be though. If the goal here is to get the best model to the TSA, then we do that by manually ensuring our data is correct, which means correcting the few errors present in the training set. This strategy is punished if the stage 2 data is not accurate.</p>",
      "rawMarkdown": "&gt; Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).\n\nThey do arbitrarily constrain how good models can be though. If the goal here is to get the best model to the TSA, then we do that by manually ensuring our data is correct, which means correcting the few errors present in the training set. This strategy is punished if the stage 2 data is not accurate.",
      "votes": null
    },
    {
      "id": "252039",
      "postDate": "12/02/2017 05:48:12",
      "content": "<p>You can't get any advantage in prediction of random errors by training on the data with random errors.</p>",
      "rawMarkdown": "You can't get any advantage in prediction of random errors by training on the data with random errors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 245779,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "11/19/2017 17:02:09",
      "content": "<p>Mislabels span a continuum, from a few definite ones, to mislabels that are likely, but are less certain.</p>\n\n<p>Whatever the organizers decide to do about them, I hope we know the details, such as what percentage of labels were removed or altered, if any.</p>\n\n<p>If stage 2 labels will be be error-free, it makes sense to train one's model on cleaned up stage 1 data.</p>\n\n<p>On the other hand, if stage 2 labels are expected to have the same proportion of errors, this approach may be problematic.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 247231,
      "author_name": "serg14",
      "author_url": "",
      "post_date": "11/22/2017 16:28:55",
      "content": "<p>Upvoting on this. It would be very helpful if Kaggle/TSA could clarify whether the labeling (truthing) process changed or not for stage 2. In other words, can we expect approximately the same mislabel error rate in stages 1 and 2?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 247243,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "11/22/2017 16:40:03",
      "content": "<p>We processed stage 1 and 2 data in the exact same manner, so you should expect (1) the labels are the ones provided by the DHS, completely unaltered. Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250652,
      "author_name": "",
      "author_url": "",
      "post_date": "11/30/2017 07:13:02",
      "content": "<blockquote>\n  <p>Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).</p>\n</blockquote>\n\n<p>They do arbitrarily constrain how good models can be though. If the goal here is to get the best model to the TSA, then we do that by manually ensuring our data is correct, which means correcting the few errors present in the training set. This strategy is punished if the stage 2 data is not accurate.</p>",
      "votes": null,
      "replies": [
        {
          "id": 252039,
          "author_name": "dmitrykovba",
          "author_url": "",
          "post_date": "12/02/2017 05:48:12",
          "content": "<p>You can't get any advantage in prediction of random errors by training on the data with random errors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "245380": "Given the presence of at least a few unambiguous mislabels in the training data (e.g https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/37615), I'm wondering what kaggle's policy on mislabels on the test / stage 2 data is? Since different assumptions about this policy will probably lead to competitors tweaking their models in different ways (and could result in pretty significant changes in the final loss given how log-loss penalizes mislabels), it seems important to have this be clear before the stage 2 data is released. Should we expect:\n\n 1. the labels are the ones provided by the DHS, completely unaltered?\n 2. the labels are the ones provided by the DHS, but have been reviewed more thoroughly than the stage 1 data and potentially revised by Kaggle or the DHS?\n 3. Kaggle will release the stage 2 labels after submissions have been made, and allow competitors to submit revisions to the labels, regrading the leaderboard if necessary?\n 4. other?\n\n I'm guessing it's 1.), but I'm new here and am not sure if there's any precedent for this stuff, so it would be nice if someone could clear this up :)",
    "245779": "Mislabels span a continuum, from a few definite ones, to mislabels that are likely, but are less certain.\n\nWhatever the organizers decide to do about them, I hope we know the details, such as what percentage of labels were removed or altered, if any.\n\nIf stage 2 labels will be be error-free, it makes sense to train one's model on cleaned up stage 1 data.\n\nOn the other hand, if stage 2 labels are expected to have the same proportion of errors, this approach may be problematic.",
    "247231": "Upvoting on this. It would be very helpful if Kaggle/TSA could clarify whether the labeling (truthing) process changed or not for stage 2. In other words, can we expect approximately the same mislabel error rate in stages 1 and 2?",
    "247243": "We processed stage 1 and 2 data in the exact same manner, so you should expect (1) the labels are the ones provided by the DHS, completely unaltered. Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).",
    "250652": "&gt; Keep in mind that almost all competitions have labeling noise, and that such mistakes do not confer an advantage to any team (i.e. everyone is punished equally by errors, up to a factor of statistical luck).\n\nThey do arbitrarily constrain how good models can be though. If the goal here is to get the best model to the TSA, then we do that by manually ensuring our data is correct, which means correcting the few errors present in the training set. This strategy is punished if the stage 2 data is not accurate.",
    "252039": "You can't get any advantage in prediction of random errors by training on the data with random errors."
  },
  "source": "meta"
}