{
  "id": 37615,
  "title": "Annotation issue",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/37615",
  "author_name": "",
  "post_date": "2017-08-06T00:16:28.757442800Z",
  "votes": 10,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Can someone confirm whether dataset 623c761b4db398ea2157e6c5cd6c8c58 is correctly annotated?  The labels are:</p>\n\n<p><code>\n623c761b4db398ea2157e6c5cd6c8c58_Zone1,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone10,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone11,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone12,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone13,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone14,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone15,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone16,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone17,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone2,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone3,1\n623c761b4db398ea2157e6c5cd6c8c58_Zone4,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone5,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone6,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone7,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone8,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone9,0\n</code></p>\n\n<p>But I see objects in zones 11, 5 and 16? </p>\n\n<p>[Attachments deleted by Admin]</p>",
  "messages": [
    {
      "id": "210496",
      "postDate": "08/06/2017 00:16:28",
      "content": "<p>Can someone confirm whether dataset 623c761b4db398ea2157e6c5cd6c8c58 is correctly annotated?  The labels are:</p>\n\n<p><code>\n623c761b4db398ea2157e6c5cd6c8c58_Zone1,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone10,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone11,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone12,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone13,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone14,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone15,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone16,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone17,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone2,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone3,1\n623c761b4db398ea2157e6c5cd6c8c58_Zone4,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone5,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone6,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone7,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone8,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone9,0\n</code></p>\n\n<p>But I see objects in zones 11, 5 and 16? </p>\n\n<p>[Attachments deleted by Admin]</p>",
      "rawMarkdown": "Can someone confirm whether dataset 623c761b4db398ea2157e6c5cd6c8c58 is correctly annotated?  The labels are:\n\n```\n623c761b4db398ea2157e6c5cd6c8c58_Zone1,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone10,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone11,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone12,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone13,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone14,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone15,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone16,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone17,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone2,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone3,1\n623c761b4db398ea2157e6c5cd6c8c58_Zone4,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone5,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone6,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone7,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone8,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone9,0\n```\n\nBut I see objects in zones 11, 5 and 16? \n\n[Attachments deleted by Admin]",
      "votes": null
    },
    {
      "id": "210927",
      "postDate": "08/07/2017 16:30:53",
      "content": "<p>Thanks for such a great post! Unfortunately, the attachments will publicly release images, which we can't have for the competition, and I had to remove the links. Thanks for understanding!</p>",
      "rawMarkdown": "Thanks for such a great post! Unfortunately, the attachments will publicly release images, which we can't have for the competition, and I had to remove the links. Thanks for understanding!",
      "votes": null
    },
    {
      "id": "211746",
      "postDate": "08/09/2017 21:17:57",
      "content": "<p>I decided to check this out and I agree it is very obvious that this is an incorrect label, and I believe your suggested labeling is correct.  I also found a couple other questionable labels, but I'm not confident enough to say for sure like I am with this one.</p>",
      "rawMarkdown": "I decided to check this out and I agree it is very obvious that this is an incorrect label, and I believe your suggested labeling is correct.  I also found a couple other questionable labels, but I'm not confident enough to say for sure like I am with this one.",
      "votes": null
    },
    {
      "id": "211770",
      "postDate": "08/09/2017 22:12:58",
      "content": "<p>If someone at Kaggle could please confirm the correct ground truth here and update the CSV it would be greatly appreciated.  Thanks!</p>",
      "rawMarkdown": "If someone at Kaggle could please confirm the correct ground truth here and update the CSV it would be greatly appreciated.  Thanks!",
      "votes": null
    },
    {
      "id": "211807",
      "postDate": "08/09/2017 23:39:55",
      "content": "<p>Yes, I also agree. Ground truth need to be updated.</p>",
      "rawMarkdown": "Yes, I also agree. Ground truth need to be updated.",
      "votes": null
    },
    {
      "id": "212039",
      "postDate": "08/10/2017 14:58:18",
      "content": "<p>We generally do not update competition datasets unless there is a systematic problem. Doing so would generate widespread confusion and spin off an endless stream of dataset versions as errors are uncovered.</p>\n\n<p><a href=\"https://www.kaggle.com/wiki/ANoteOnDataQuality\">Most datasets have errors like this</a>. When you spot them, you're free to correct them and use them to your advantage.</p>",
      "rawMarkdown": "We generally do not update competition datasets unless there is a systematic problem. Doing so would generate widespread confusion and spin off an endless stream of dataset versions as errors are uncovered.\n\n[Most datasets have errors like this](https://www.kaggle.com/wiki/ANoteOnDataQuality). When you spot them, you're free to correct them and use them to your advantage.",
      "votes": null
    },
    {
      "id": "212046",
      "postDate": "08/10/2017 15:20:24",
      "content": "<p>Hi William,</p>\n\n<p>Thanks for the response!  Understood and my apologies for the request.  I'm new here and still learning how these things are handled.  I mistakenly assumed that the 'hand labeling' prohibition meant that we weren't really supposed to be verifying the images manually (hence why the 'ground truth' was provided in the first place).  Does that rule only apply to the test set in Phase 2?  Is hand labeling allowed for the test set in Phase 1? (i.e. Would doing so in Phase 1 (to have more data for training) disqualify a team from winning?)</p>\n\n<p>Thanks!\n-Adam</p>",
      "rawMarkdown": "Hi William,\n\nThanks for the response!  Understood and my apologies for the request.  I'm new here and still learning how these things are handled.  I mistakenly assumed that the 'hand labeling' prohibition meant that we weren't really supposed to be verifying the images manually (hence why the 'ground truth' was provided in the first place).  Does that rule only apply to the test set in Phase 2?  Is hand labeling allowed for the test set in Phase 1? (i.e. Would doing so in Phase 1 (to have more data for training) disqualify a team from winning?)\n\nThanks!\n-Adam",
      "votes": null
    },
    {
      "id": "212049",
      "postDate": "08/10/2017 15:25:06",
      "content": "<p>Hand labeling is allowed on the training set(s), but not on the test set(s). Per the rules:</p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>",
      "rawMarkdown": "Hand labeling is allowed on the training set(s), but not on the test set(s). Per the rules:\n\n&gt;  Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.",
      "votes": null
    },
    {
      "id": "212080",
      "postDate": "08/10/2017 16:22:57",
      "content": "<p>Have seen A LOT of mislabel or no label as well. Think \"ground truth\" should at least be looked at. But, if we correct labeling, then will get a worse score. Since it only matters if our predictions match what they have; even though will have better detection in general. In a sense we are training models to make same mistakes.</p>",
      "rawMarkdown": "Have seen A LOT of mislabel or no label as well. Think \"ground truth\" should at least be looked at. But, if we correct labeling, then will get a worse score. Since it only matters if our predictions match what they have; even though will have better detection in general. In a sense we are training models to make same mistakes.",
      "votes": null
    },
    {
      "id": "212103",
      "postDate": "08/10/2017 17:19:09",
      "content": "<p>Hi William,</p>\n\n<p>Alright, I don't mean to be pedantic, but in an effort to make sure I understand the rules:</p>\n\n<ul>\n<li><p>I assume that since there is no 'formal' validation dataset here, that the human prediction clause does not apply in a case where we are doing, say, k-fold cross-validation using the training dataset -- even though we would be effectively using 'human predicted' (or at least confirmed) training data as validation data in that case.  Is that correct?</p></li>\n<li><p>Is it correct that anyone who hand-labels Phase 1 test data and uses those human predictions to train their models (so as to improve performance on Phase 1 and ultimately Phase 2) will be disqualified from the competition?</p></li>\n</ul>\n\n<p>Again, sorry if this seems 'nit picky' - I'm just trying to make sure I don't mistakenly do something that will have me disqualified.</p>\n\n<p>Thanks!\n-Adam </p>",
      "rawMarkdown": "Hi William,\n\nAlright, I don't mean to be pedantic, but in an effort to make sure I understand the rules:\n\n- I assume that since there is no 'formal' validation dataset here, that the human prediction clause does not apply in a case where we are doing, say, k-fold cross-validation using the training dataset -- even though we would be effectively using 'human predicted' (or at least confirmed) training data as validation data in that case.  Is that correct?\n\n-  Is it correct that anyone who hand-labels Phase 1 test data and uses those human predictions to train their models (so as to improve performance on Phase 1 and ultimately Phase 2) will be disqualified from the competition?\n\nAgain, sorry if this seems 'nit picky' - I'm just trying to make sure I don't mistakenly do something that will have me disqualified.\n\nThanks!\n-Adam",
      "votes": null
    },
    {
      "id": "212116",
      "postDate": "08/10/2017 18:07:00",
      "content": "<p>No worries, happy to clarify!</p>\n\n<ul>\n<li>You can hand label anything in stage1_labels.csv.</li>\n<li>We will give out the stage 1 test labels prior to stage 2. You are allowed to retrain your models using the combination of stage 1 training images and stage 1 test images (this is done to level the playing field, as people can probe the leaderboard to get the stage 1 test set labels during the first stage).</li>\n</ul>\n\n<p>The thing not to do is hand label stage 2 test images. If you do this, finish in the prizes, and your (previously locked-in) model does not reproduce your labels, you would be DQ'd.</p>",
      "rawMarkdown": "No worries, happy to clarify!\n\n - You can hand label anything in stage1_labels.csv.\n - We will give out the stage 1 test labels prior to stage 2. You are allowed to retrain your models using the combination of stage 1 training images and stage 1 test images (this is done to level the playing field, as people can probe the leaderboard to get the stage 1 test set labels during the first stage).\n\nThe thing not to do is hand label stage 2 test images. If you do this, finish in the prizes, and your (previously locked-in) model does not reproduce your labels, you would be DQ'd.",
      "votes": null
    },
    {
      "id": "212118",
      "postDate": "08/10/2017 18:09:20",
      "content": "<p>Also, folks who are new to two stage comps may benefit from reading this (it's a complicated process, even though we try to make it as simple as we can) - <a href=\"https://www.kaggle.com/two-stage-faq\">https://www.kaggle.com/two-stage-faq</a></p>",
      "rawMarkdown": "Also, folks who are new to two stage comps may benefit from reading this (it's a complicated process, even though we try to make it as simple as we can) - https://www.kaggle.com/two-stage-faq",
      "votes": null
    },
    {
      "id": "212146",
      "postDate": "08/10/2017 19:32:00",
      "content": "<p>How does one \"probe the leaderboard to get the stage 1 test labels\"? Other than person with best score has best labels, so if share predictions then get best labels; but not actual. Thanks.</p>",
      "rawMarkdown": "How does one \"probe the leaderboard to get the stage 1 test labels\"? Other than person with best score has best labels, so if share predictions then get best labels; but not actual. Thanks.",
      "votes": null
    },
    {
      "id": "212178",
      "postDate": "08/10/2017 21:17:25",
      "content": "<p>Thanks William!  Got it.</p>",
      "rawMarkdown": "Thanks William!  Got it.",
      "votes": null
    },
    {
      "id": "218059",
      "postDate": "09/02/2017 02:30:57",
      "content": "<p>So today, I discovered another mislabel: 496ec724cc1f2886aac5840cf890988a</p>\n\n<p>Current Labels: [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 1, 0]</p>\n\n<p>Proposed Labels: [0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]</p>\n\n<p>In case you missed it, this appears to be <strong>exactly</strong> swapped with 623c761b4db398ea2157e6c5cd6c8c58, and that would explain how the mislabel happened... Amazing!  This gives me even more confidence that the proposed labels for both of these are correct.</p>",
      "rawMarkdown": "So today, I discovered another mislabel: 496ec724cc1f2886aac5840cf890988a\n\nCurrent Labels: [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 1, 0]\n\nProposed Labels: [0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]\n\nIn case you missed it, this appears to be **exactly** swapped with 623c761b4db398ea2157e6c5cd6c8c58, and that would explain how the mislabel happened... Amazing!  This gives me even more confidence that the proposed labels for both of these are correct.",
      "votes": null
    },
    {
      "id": "218569",
      "postDate": "09/04/2017 23:52:13",
      "content": "<blockquote>\n  <p>We will give out the stage 1 test labels prior to stage 2.</p>\n</blockquote>\n\n<p>I do not see this statement in the rules, I take it William works at Kaggle and it's official now?</p>",
      "rawMarkdown": "&gt; We will give out the stage 1 test labels prior to stage 2.\n\nI do not see this statement in the rules, I take it William works at Kaggle and it's official now?",
      "votes": null
    },
    {
      "id": "219071",
      "postDate": "09/06/2017 21:55:46",
      "content": "<p>Kevin,</p>\n\n<p>Thanks for this find. I concur with your evaluation.</p>",
      "rawMarkdown": "Kevin,\n\nThanks for this find. I concur with your evaluation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 210927,
      "author_name": "addisonhoward",
      "author_url": "",
      "post_date": "08/07/2017 16:30:53",
      "content": "<p>Thanks for such a great post! Unfortunately, the attachments will publicly release images, which we can't have for the competition, and I had to remove the links. Thanks for understanding!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 211746,
      "author_name": "hackerpoet",
      "author_url": "",
      "post_date": "08/09/2017 21:17:57",
      "content": "<p>I decided to check this out and I agree it is very obvious that this is an incorrect label, and I believe your suggested labeling is correct.  I also found a couple other questionable labels, but I'm not confident enough to say for sure like I am with this one.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 211770,
      "author_name": "adam314",
      "author_url": "",
      "post_date": "08/09/2017 22:12:58",
      "content": "<p>If someone at Kaggle could please confirm the correct ground truth here and update the CSV it would be greatly appreciated.  Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 212039,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "08/10/2017 14:58:18",
          "content": "<p>We generally do not update competition datasets unless there is a systematic problem. Doing so would generate widespread confusion and spin off an endless stream of dataset versions as errors are uncovered.</p>\n\n<p><a href=\"https://www.kaggle.com/wiki/ANoteOnDataQuality\">Most datasets have errors like this</a>. When you spot them, you're free to correct them and use them to your advantage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212046,
          "author_name": "adam314",
          "author_url": "",
          "post_date": "08/10/2017 15:20:24",
          "content": "<p>Hi William,</p>\n\n<p>Thanks for the response!  Understood and my apologies for the request.  I'm new here and still learning how these things are handled.  I mistakenly assumed that the 'hand labeling' prohibition meant that we weren't really supposed to be verifying the images manually (hence why the 'ground truth' was provided in the first place).  Does that rule only apply to the test set in Phase 2?  Is hand labeling allowed for the test set in Phase 1? (i.e. Would doing so in Phase 1 (to have more data for training) disqualify a team from winning?)</p>\n\n<p>Thanks!\n-Adam</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212049,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "08/10/2017 15:25:06",
          "content": "<p>Hand labeling is allowed on the training set(s), but not on the test set(s). Per the rules:</p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212103,
          "author_name": "adam314",
          "author_url": "",
          "post_date": "08/10/2017 17:19:09",
          "content": "<p>Hi William,</p>\n\n<p>Alright, I don't mean to be pedantic, but in an effort to make sure I understand the rules:</p>\n\n<ul>\n<li><p>I assume that since there is no 'formal' validation dataset here, that the human prediction clause does not apply in a case where we are doing, say, k-fold cross-validation using the training dataset -- even though we would be effectively using 'human predicted' (or at least confirmed) training data as validation data in that case.  Is that correct?</p></li>\n<li><p>Is it correct that anyone who hand-labels Phase 1 test data and uses those human predictions to train their models (so as to improve performance on Phase 1 and ultimately Phase 2) will be disqualified from the competition?</p></li>\n</ul>\n\n<p>Again, sorry if this seems 'nit picky' - I'm just trying to make sure I don't mistakenly do something that will have me disqualified.</p>\n\n<p>Thanks!\n-Adam </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212116,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "08/10/2017 18:07:00",
          "content": "<p>No worries, happy to clarify!</p>\n\n<ul>\n<li>You can hand label anything in stage1_labels.csv.</li>\n<li>We will give out the stage 1 test labels prior to stage 2. You are allowed to retrain your models using the combination of stage 1 training images and stage 1 test images (this is done to level the playing field, as people can probe the leaderboard to get the stage 1 test set labels during the first stage).</li>\n</ul>\n\n<p>The thing not to do is hand label stage 2 test images. If you do this, finish in the prizes, and your (previously locked-in) model does not reproduce your labels, you would be DQ'd.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212118,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "08/10/2017 18:09:20",
          "content": "<p>Also, folks who are new to two stage comps may benefit from reading this (it's a complicated process, even though we try to make it as simple as we can) - <a href=\"https://www.kaggle.com/two-stage-faq\">https://www.kaggle.com/two-stage-faq</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212146,
          "author_name": "srdhaeayautnon",
          "author_url": "",
          "post_date": "08/10/2017 19:32:00",
          "content": "<p>How does one \"probe the leaderboard to get the stage 1 test labels\"? Other than person with best score has best labels, so if share predictions then get best labels; but not actual. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212178,
          "author_name": "adam314",
          "author_url": "",
          "post_date": "08/10/2017 21:17:25",
          "content": "<p>Thanks William!  Got it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 211807,
      "author_name": "edvin1983",
      "author_url": "",
      "post_date": "08/09/2017 23:39:55",
      "content": "<p>Yes, I also agree. Ground truth need to be updated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 212080,
      "author_name": "srdhaeayautnon",
      "author_url": "",
      "post_date": "08/10/2017 16:22:57",
      "content": "<p>Have seen A LOT of mislabel or no label as well. Think \"ground truth\" should at least be looked at. But, if we correct labeling, then will get a worse score. Since it only matters if our predictions match what they have; even though will have better detection in general. In a sense we are training models to make same mistakes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 218569,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "09/04/2017 23:52:13",
          "content": "<blockquote>\n  <p>We will give out the stage 1 test labels prior to stage 2.</p>\n</blockquote>\n\n<p>I do not see this statement in the rules, I take it William works at Kaggle and it's official now?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 218059,
      "author_name": "hackerpoet",
      "author_url": "",
      "post_date": "09/02/2017 02:30:57",
      "content": "<p>So today, I discovered another mislabel: 496ec724cc1f2886aac5840cf890988a</p>\n\n<p>Current Labels: [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 1, 0]</p>\n\n<p>Proposed Labels: [0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]</p>\n\n<p>In case you missed it, this appears to be <strong>exactly</strong> swapped with 623c761b4db398ea2157e6c5cd6c8c58, and that would explain how the mislabel happened... Amazing!  This gives me even more confidence that the proposed labels for both of these are correct.</p>",
      "votes": null,
      "replies": [
        {
          "id": 219071,
          "author_name": "rluethy",
          "author_url": "",
          "post_date": "09/06/2017 21:55:46",
          "content": "<p>Kevin,</p>\n\n<p>Thanks for this find. I concur with your evaluation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "210496": "Can someone confirm whether dataset 623c761b4db398ea2157e6c5cd6c8c58 is correctly annotated?  The labels are:\n\n```\n623c761b4db398ea2157e6c5cd6c8c58_Zone1,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone10,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone11,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone12,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone13,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone14,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone15,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone16,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone17,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone2,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone3,1\n623c761b4db398ea2157e6c5cd6c8c58_Zone4,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone5,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone6,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone7,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone8,0\n623c761b4db398ea2157e6c5cd6c8c58_Zone9,0\n```\n\nBut I see objects in zones 11, 5 and 16? \n\n[Attachments deleted by Admin]",
    "210927": "Thanks for such a great post! Unfortunately, the attachments will publicly release images, which we can't have for the competition, and I had to remove the links. Thanks for understanding!",
    "211746": "I decided to check this out and I agree it is very obvious that this is an incorrect label, and I believe your suggested labeling is correct.  I also found a couple other questionable labels, but I'm not confident enough to say for sure like I am with this one.",
    "211770": "If someone at Kaggle could please confirm the correct ground truth here and update the CSV it would be greatly appreciated.  Thanks!",
    "211807": "Yes, I also agree. Ground truth need to be updated.",
    "212039": "We generally do not update competition datasets unless there is a systematic problem. Doing so would generate widespread confusion and spin off an endless stream of dataset versions as errors are uncovered.\n\n[Most datasets have errors like this](https://www.kaggle.com/wiki/ANoteOnDataQuality). When you spot them, you're free to correct them and use them to your advantage.",
    "212046": "Hi William,\n\nThanks for the response!  Understood and my apologies for the request.  I'm new here and still learning how these things are handled.  I mistakenly assumed that the 'hand labeling' prohibition meant that we weren't really supposed to be verifying the images manually (hence why the 'ground truth' was provided in the first place).  Does that rule only apply to the test set in Phase 2?  Is hand labeling allowed for the test set in Phase 1? (i.e. Would doing so in Phase 1 (to have more data for training) disqualify a team from winning?)\n\nThanks!\n-Adam",
    "212049": "Hand labeling is allowed on the training set(s), but not on the test set(s). Per the rules:\n\n&gt;  Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.",
    "212080": "Have seen A LOT of mislabel or no label as well. Think \"ground truth\" should at least be looked at. But, if we correct labeling, then will get a worse score. Since it only matters if our predictions match what they have; even though will have better detection in general. In a sense we are training models to make same mistakes.",
    "212103": "Hi William,\n\nAlright, I don't mean to be pedantic, but in an effort to make sure I understand the rules:\n\n- I assume that since there is no 'formal' validation dataset here, that the human prediction clause does not apply in a case where we are doing, say, k-fold cross-validation using the training dataset -- even though we would be effectively using 'human predicted' (or at least confirmed) training data as validation data in that case.  Is that correct?\n\n-  Is it correct that anyone who hand-labels Phase 1 test data and uses those human predictions to train their models (so as to improve performance on Phase 1 and ultimately Phase 2) will be disqualified from the competition?\n\nAgain, sorry if this seems 'nit picky' - I'm just trying to make sure I don't mistakenly do something that will have me disqualified.\n\nThanks!\n-Adam",
    "212116": "No worries, happy to clarify!\n\n - You can hand label anything in stage1_labels.csv.\n - We will give out the stage 1 test labels prior to stage 2. You are allowed to retrain your models using the combination of stage 1 training images and stage 1 test images (this is done to level the playing field, as people can probe the leaderboard to get the stage 1 test set labels during the first stage).\n\nThe thing not to do is hand label stage 2 test images. If you do this, finish in the prizes, and your (previously locked-in) model does not reproduce your labels, you would be DQ'd.",
    "212118": "Also, folks who are new to two stage comps may benefit from reading this (it's a complicated process, even though we try to make it as simple as we can) - https://www.kaggle.com/two-stage-faq",
    "212146": "How does one \"probe the leaderboard to get the stage 1 test labels\"? Other than person with best score has best labels, so if share predictions then get best labels; but not actual. Thanks.",
    "212178": "Thanks William!  Got it.",
    "218059": "So today, I discovered another mislabel: 496ec724cc1f2886aac5840cf890988a\n\nCurrent Labels: [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 1, 0]\n\nProposed Labels: [0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]\n\nIn case you missed it, this appears to be **exactly** swapped with 623c761b4db398ea2157e6c5cd6c8c58, and that would explain how the mislabel happened... Amazing!  This gives me even more confidence that the proposed labels for both of these are correct.",
    "218569": "&gt; We will give out the stage 1 test labels prior to stage 2.\n\nI do not see this statement in the rules, I take it William works at Kaggle and it's official now?",
    "219071": "Kevin,\n\nThanks for this find. I concur with your evaluation."
  },
  "source": "meta"
}