{
  "id": 68335,
  "title": "Danger: Hand Labeling",
  "url": "/competitions/airbus-ship-detection/discussion/68335",
  "author_name": "",
  "post_date": "2018-10-11T16:27:13.654146200Z",
  "votes": 11,
  "comment_count": 11,
  "views": 0,
  "content": "<p>According to this nice kernel (<a href=\"https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb\">https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb</a>), it seems it is easy to eliminate the images with no ship or hard to find ship. So out of 18k images, ships are segmented for only 3.3k images. One can divide this 3.3k images to remaining 33 days and label 100 images by hand every day, which is not really hard for human. This may cause a lot of unfair advantages in the leaderboard. I think either the test data should have been bigger like before the reset or the task should have been not easy for hand labeling. How does Kaggle plan to prevent it?</p>",
  "messages": [
    {
      "id": "402395",
      "postDate": "10/11/2018 16:27:13",
      "content": "<p>According to this nice kernel (<a href=\"https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb\">https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb</a>), it seems it is easy to eliminate the images with no ship or hard to find ship. So out of 18k images, ships are segmented for only 3.3k images. One can divide this 3.3k images to remaining 33 days and label 100 images by hand every day, which is not really hard for human. This may cause a lot of unfair advantages in the leaderboard. I think either the test data should have been bigger like before the reset or the task should have been not easy for hand labeling. How does Kaggle plan to prevent it?</p>",
      "rawMarkdown": "According to this nice kernel (https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb), it seems it is easy to eliminate the images with no ship or hard to find ship. So out of 18k images, ships are segmented for only 3.3k images. One can divide this 3.3k images to remaining 33 days and label 100 images by hand every day, which is not really hard for human. This may cause a lot of unfair advantages in the leaderboard. I think either the test data should have been bigger like before the reset or the task should have been not easy for hand labeling. How does Kaggle plan to prevent it?",
      "votes": null
    },
    {
      "id": "402483",
      "postDate": "10/11/2018 19:33:08",
      "content": "<p>I did checking of model errors yesterday and found that the main problem is misalignment of small ships (I would say ~98-99% of ships are identified as masks by my main model, but they may get low score because of low IoU). The model outperforms humans for such task: if you have a ~50 pixel blurry object, you really never know where to put the box. Though the issue of hand labeling should be considered.</p>",
      "rawMarkdown": "I did checking of model errors yesterday and found that the main problem is misalignment of small ships (I would say ~98-99% of ships are identified as masks by my main model, but they may get low score because of low IoU). The model outperforms humans for such task: if you have a ~50 pixel blurry object, you really never know where to put the box. Though the issue of hand labeling should be considered.",
      "votes": null
    },
    {
      "id": "402487",
      "postDate": "10/11/2018 19:38:44",
      "content": "<p>Absolutely agree. Test data seems too easy to hand labeling.</p>",
      "rawMarkdown": "Absolutely agree. Test data seems too easy to hand labeling.",
      "votes": null
    },
    {
      "id": "402508",
      "postDate": "10/11/2018 20:03:56",
      "content": "<p>The organizers said that they included a number of un-scored files to discourage hand labeling. Also if you win the competition you need to submit your model and would get disqualified for hand labeling. Regarding medals, I agree with you in that there is no way to check. <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/68042\">https://www.kaggle.com/c/airbus-ship-detection/discussion/68042</a></p>",
      "rawMarkdown": "The organizers said that they included a number of un-scored files to discourage hand labeling. Also if you win the competition you need to submit your model and would get disqualified for hand labeling. Regarding medals, I agree with you in that there is no way to check. https://www.kaggle.com/c/airbus-ship-detection/discussion/68042",
      "votes": null
    },
    {
      "id": "403628",
      "postDate": "10/14/2018 06:22:28",
      "content": "<p>Posted some further thoughts on this topic here\n<a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/68538\">https://www.kaggle.com/c/airbus-ship-detection/discussion/68538</a></p>",
      "rawMarkdown": "Posted some further thoughts on this topic here\nhttps://www.kaggle.com/c/airbus-ship-detection/discussion/68538",
      "votes": null
    },
    {
      "id": "403777",
      "postDate": "10/14/2018 15:28:15",
      "content": "<p>I don't think the issue is submission of hand-labeled test set (without a real model that generates good results). I think that the main issue here is that it's really easy to hand-label the test set, and then use it as a training set for your model. If the only way to evaluate models is based on the score on a small, publicly available test set, then there is practically no way to disqualify such model (as long as you only need to provide a trained model).</p>",
      "rawMarkdown": "I don't think the issue is submission of hand-labeled test set (without a real model that generates good results). I think that the main issue here is that it's really easy to hand-label the test set, and then use it as a training set for your model. If the only way to evaluate models is based on the score on a small, publicly available test set, then there is practically no way to disqualify such model (as long as you only need to provide a trained model).",
      "votes": null
    },
    {
      "id": "405466",
      "postDate": "10/17/2018 14:38:10",
      "content": "<p>Yes, as long as you don't submit code there is no way to check.</p>",
      "rawMarkdown": "Yes, as long as you don't submit code there is no way to check.",
      "votes": null
    },
    {
      "id": "406089",
      "postDate": "10/18/2018 16:10:38",
      "content": "<p>@YaG320 -</p>\n\n<p>So there's no misunderstanding, winners are required to provide not only the winning models, but also the code necessary to generate the models. </p>",
      "rawMarkdown": "YaG320 -\n\nSo there's no misunderstanding, winners are required to provide not only the winning models, but also the code necessary to generate the models.",
      "votes": null
    },
    {
      "id": "406091",
      "postDate": "10/18/2018 16:14:44",
      "content": "<p>Hopefully I'm pointing out the obvious here, but the rules for this competition, which all competitors agreed to when entering the competition, forbid any hand labeling of the test images. Regardless of of how easy someone believes it may be to do, or whether or not someone believes it will be detected, it is the wrong thing to do.</p>",
      "rawMarkdown": "Hopefully I'm pointing out the obvious here, but the rules for this competition, which all competitors agreed to when entering the competition, forbid any hand labeling of the test images. Regardless of of how easy someone believes it may be to do, or whether or not someone believes it will be detected, it is the wrong thing to do.",
      "votes": null
    },
    {
      "id": "407321",
      "postDate": "10/20/2018 21:51:34",
      "content": "<p>Great,</p>\n\n<p>I think it's a really good idea to have this requirement.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Great,\n\nI think it's a really good idea to have this requirement.\n\nThanks!",
      "votes": null
    },
    {
      "id": "407347",
      "postDate": "10/21/2018 00:15:15",
      "content": "<p>However, if the same code is executed again,  the result may be slightly different. There are many things affecting it. For example, depending on how images are selecting into batches or which weights are doped in dropout, the result may slightly change. In my training I have never specified the initial state for such things.  But even I did it, the random numbers are not guaranteed to be the same depending on the platform and version of python and other libraries. So, it may be difficult to exactly reproduce the results having the code. They will be about the same, but even ~0.01 score can determine the winer.</p>",
      "rawMarkdown": "However, if the same code is executed again,  the result may be slightly different. There are many things affecting it. For example, depending on how images are selecting into batches or which weights are doped in dropout, the result may slightly change. In my training I have never specified the initial state for such things.  But even I did it, the random numbers are not guaranteed to be the same depending on the platform and version of python and other libraries. So, it may be difficult to exactly reproduce the results having the code. They will be about the same, but even ~0.01 score can determine the winer.",
      "votes": null
    },
    {
      "id": "407848",
      "postDate": "10/21/2018 22:07:27",
      "content": "<p>If somebody won the competition by using a model that was trained with a hand-labeled test set, then the deviations in a re-trained model (using valid data-set) would be very large (And hopefully would raise the organizers' suspicion). </p>\n\n<p>I think <a href=\"/inversion\">@inversion</a> only mentioned this requirement in order to verify that the winner didn't cheat, and not to re-evaluate the results after the winner was announced.</p>",
      "rawMarkdown": "If somebody won the competition by using a model that was trained with a hand-labeled test set, then the deviations in a re-trained model (using valid data-set) would be very large (And hopefully would raise the organizers' suspicion). \n\nI think @inversion only mentioned this requirement in order to verify that the winner didn't cheat, and not to re-evaluate the results after the winner was announced.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 402483,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "10/11/2018 19:33:08",
      "content": "<p>I did checking of model errors yesterday and found that the main problem is misalignment of small ships (I would say ~98-99% of ships are identified as masks by my main model, but they may get low score because of low IoU). The model outperforms humans for such task: if you have a ~50 pixel blurry object, you really never know where to put the box. Though the issue of hand labeling should be considered.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 402487,
      "author_name": "donchuk",
      "author_url": "",
      "post_date": "10/11/2018 19:38:44",
      "content": "<p>Absolutely agree. Test data seems too easy to hand labeling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 402508,
      "author_name": "msmelguizo",
      "author_url": "",
      "post_date": "10/11/2018 20:03:56",
      "content": "<p>The organizers said that they included a number of un-scored files to discourage hand labeling. Also if you win the competition you need to submit your model and would get disqualified for hand labeling. Regarding medals, I agree with you in that there is no way to check. <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/68042\">https://www.kaggle.com/c/airbus-ship-detection/discussion/68042</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 403777,
          "author_name": "yag320",
          "author_url": "",
          "post_date": "10/14/2018 15:28:15",
          "content": "<p>I don't think the issue is submission of hand-labeled test set (without a real model that generates good results). I think that the main issue here is that it's really easy to hand-label the test set, and then use it as a training set for your model. If the only way to evaluate models is based on the score on a small, publicly available test set, then there is practically no way to disqualify such model (as long as you only need to provide a trained model).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 405466,
          "author_name": "msmelguizo",
          "author_url": "",
          "post_date": "10/17/2018 14:38:10",
          "content": "<p>Yes, as long as you don't submit code there is no way to check.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406089,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "10/18/2018 16:10:38",
          "content": "<p>@YaG320 -</p>\n\n<p>So there's no misunderstanding, winners are required to provide not only the winning models, but also the code necessary to generate the models. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407321,
          "author_name": "yag320",
          "author_url": "",
          "post_date": "10/20/2018 21:51:34",
          "content": "<p>Great,</p>\n\n<p>I think it's a really good idea to have this requirement.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407347,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/21/2018 00:15:15",
          "content": "<p>However, if the same code is executed again,  the result may be slightly different. There are many things affecting it. For example, depending on how images are selecting into batches or which weights are doped in dropout, the result may slightly change. In my training I have never specified the initial state for such things.  But even I did it, the random numbers are not guaranteed to be the same depending on the platform and version of python and other libraries. So, it may be difficult to exactly reproduce the results having the code. They will be about the same, but even ~0.01 score can determine the winer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407848,
          "author_name": "yag320",
          "author_url": "",
          "post_date": "10/21/2018 22:07:27",
          "content": "<p>If somebody won the competition by using a model that was trained with a hand-labeled test set, then the deviations in a re-trained model (using valid data-set) would be very large (And hopefully would raise the organizers' suspicion). </p>\n\n<p>I think <a href=\"/inversion\">@inversion</a> only mentioned this requirement in order to verify that the winner didn't cheat, and not to re-evaluate the results after the winner was announced.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 403628,
      "author_name": "snakers41",
      "author_url": "",
      "post_date": "10/14/2018 06:22:28",
      "content": "<p>Posted some further thoughts on this topic here\n<a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/68538\">https://www.kaggle.com/c/airbus-ship-detection/discussion/68538</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 406091,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "10/18/2018 16:14:44",
      "content": "<p>Hopefully I'm pointing out the obvious here, but the rules for this competition, which all competitors agreed to when entering the competition, forbid any hand labeling of the test images. Regardless of of how easy someone believes it may be to do, or whether or not someone believes it will be detected, it is the wrong thing to do.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "402395": "According to this nice kernel (https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb), it seems it is easy to eliminate the images with no ship or hard to find ship. So out of 18k images, ships are segmented for only 3.3k images. One can divide this 3.3k images to remaining 33 days and label 100 images by hand every day, which is not really hard for human. This may cause a lot of unfair advantages in the leaderboard. I think either the test data should have been bigger like before the reset or the task should have been not easy for hand labeling. How does Kaggle plan to prevent it?",
    "402483": "I did checking of model errors yesterday and found that the main problem is misalignment of small ships (I would say ~98-99% of ships are identified as masks by my main model, but they may get low score because of low IoU). The model outperforms humans for such task: if you have a ~50 pixel blurry object, you really never know where to put the box. Though the issue of hand labeling should be considered.",
    "402487": "Absolutely agree. Test data seems too easy to hand labeling.",
    "402508": "The organizers said that they included a number of un-scored files to discourage hand labeling. Also if you win the competition you need to submit your model and would get disqualified for hand labeling. Regarding medals, I agree with you in that there is no way to check. https://www.kaggle.com/c/airbus-ship-detection/discussion/68042",
    "403628": "Posted some further thoughts on this topic here\nhttps://www.kaggle.com/c/airbus-ship-detection/discussion/68538",
    "403777": "I don't think the issue is submission of hand-labeled test set (without a real model that generates good results). I think that the main issue here is that it's really easy to hand-label the test set, and then use it as a training set for your model. If the only way to evaluate models is based on the score on a small, publicly available test set, then there is practically no way to disqualify such model (as long as you only need to provide a trained model).",
    "405466": "Yes, as long as you don't submit code there is no way to check.",
    "406089": "YaG320 -\n\nSo there's no misunderstanding, winners are required to provide not only the winning models, but also the code necessary to generate the models.",
    "406091": "Hopefully I'm pointing out the obvious here, but the rules for this competition, which all competitors agreed to when entering the competition, forbid any hand labeling of the test images. Regardless of of how easy someone believes it may be to do, or whether or not someone believes it will be detected, it is the wrong thing to do.",
    "407321": "Great,\n\nI think it's a really good idea to have this requirement.\n\nThanks!",
    "407347": "However, if the same code is executed again,  the result may be slightly different. There are many things affecting it. For example, depending on how images are selecting into batches or which weights are doped in dropout, the result may slightly change. In my training I have never specified the initial state for such things.  But even I did it, the random numbers are not guaranteed to be the same depending on the platform and version of python and other libraries. So, it may be difficult to exactly reproduce the results having the code. They will be about the same, but even ~0.01 score can determine the winer.",
    "407848": "If somebody won the competition by using a model that was trained with a hand-labeled test set, then the deviations in a re-trained model (using valid data-set) would be very large (And hopefully would raise the organizers' suspicion). \n\nI think @inversion only mentioned this requirement in order to verify that the winner didn't cheat, and not to re-evaluate the results after the winner was announced."
  },
  "source": "meta"
}