{
  "id": 55436,
  "title": "Human annotation rules",
  "url": "/competitions/avito-demand-prediction/discussion/55436",
  "author_name": "",
  "post_date": "2018-04-26T14:11:39.701919Z",
  "votes": 4,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi, </p>\n\n<p>are there any rules about using my own custom annotations for training model ? Is it allowed ?</p>\n\n<p>For example I will create my own derived dataset from images provided with my own annotation of \"niceness\" of image, later train feature extractor with this self-annotated data and use it as feature. Is it allowed according to the rules ? </p>\n\n<p>Thanks\nJan</p>",
  "messages": [
    {
      "id": "319647",
      "postDate": "04/26/2018 14:11:39",
      "content": "<p>Hi, </p>\n\n<p>are there any rules about using my own custom annotations for training model ? Is it allowed ?</p>\n\n<p>For example I will create my own derived dataset from images provided with my own annotation of \"niceness\" of image, later train feature extractor with this self-annotated data and use it as feature. Is it allowed according to the rules ? </p>\n\n<p>Thanks\nJan</p>",
      "rawMarkdown": "Hi, \n\nare there any rules about using my own custom annotations for training model ? Is it allowed ?\n\nFor example I will create my own derived dataset from images provided with my own annotation of \"niceness\" of image, later train feature extractor with this self-annotated data and use it as feature. Is it allowed according to the rules ? \n\nThanks\nJan",
      "votes": null
    },
    {
      "id": "319974",
      "postDate": "04/27/2018 08:05:53",
      "content": "<p>This is a good question. But given that there are 155,197 unique images in the train and test set to annotate, that would be quite some dedication! Even if each image took 5sec to annotate, you're looking at 215h 33m of non-stop work!</p>",
      "rawMarkdown": "This is a good question. But given that there are 155,197 unique images in the train and test set to annotate, that would be quite some dedication! Even if each image took 5sec to annotate, you're looking at 215h 33m of non-stop work!",
      "votes": null
    },
    {
      "id": "319981",
      "postDate": "04/27/2018 08:26:40",
      "content": "<p>You don`t need to annotate all the images to have reasonably working feature extractor in my opinion. Of course the more annotations you have the better.</p>\n\n<p>What is your interpretation of the rules ? Do you think it is allowed ? </p>",
      "rawMarkdown": "You don`t need to annotate all the images to have reasonably working feature extractor in my opinion. Of course the more annotations you have the better.\n\nWhat is your interpretation of the rules ? Do you think it is allowed ?",
      "votes": null
    },
    {
      "id": "320038",
      "postDate": "04/27/2018 10:35:16",
      "content": "<p>Actually this is <strong>explicitly against the rules</strong>. I just did a search through the rules and it says \"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\"</p>",
      "rawMarkdown": "Actually this is **explicitly against the rules**. I just did a search through the rules and it says \"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\"",
      "votes": null
    },
    {
      "id": "320100",
      "postDate": "04/27/2018 13:31:58",
      "content": "<p>Does it apply to labeling training data ? </p>",
      "rawMarkdown": "Does it apply to labeling training data ?",
      "votes": null
    },
    {
      "id": "320141",
      "postDate": "04/27/2018 16:01:03",
      "content": "<p>Why not use deal_probability as proxy for 'niceness'?</p>",
      "rawMarkdown": "Why not use deal_probability as proxy for 'niceness'?",
      "votes": null
    },
    {
      "id": "320195",
      "postDate": "04/27/2018 20:41:17",
      "content": "<p>That's what I was thinking...</p>",
      "rawMarkdown": "That's what I was thinking...",
      "votes": null
    },
    {
      "id": "320196",
      "postDate": "04/27/2018 20:41:51",
      "content": "<p>Yes, it applies to any models, train data, test data, or anything else.  No hand labeling or human prediction allowed, in any aspect, whatsoever.</p>",
      "rawMarkdown": "Yes, it applies to any models, train data, test data, or anything else.  No hand labeling or human prediction allowed, in any aspect, whatsoever.",
      "votes": null
    },
    {
      "id": "321012",
      "postDate": "04/30/2018 11:25:15",
      "content": "<p>To my understanding, this doesn't mean you can't use intermediate labels in your models. This would mean that you can't label yourself the output of your global classifier, to submit these handmade labels. \nAll the more since external data is permitted in the  competition, that wouldn't make sense to prevent competitors from using intermediate hand labels to train their models. (for the use of external data see the competition rules in \"Competition-specific terms/External data\")</p>",
      "rawMarkdown": "To my understanding, this doesn't mean you can't use intermediate labels in your models. This would mean that you can't label yourself the output of your global classifier, to submit these handmade labels. \nAll the more since external data is permitted in the  competition, that wouldn't make sense to prevent competitors from using intermediate hand labels to train their models. (for the use of external data see the competition rules in \"Competition-specific terms/External data\")",
      "votes": null
    },
    {
      "id": "321019",
      "postDate": "04/30/2018 12:15:57",
      "content": "<p>In \"data science bowl 2018\" competition, Kaggle admins allowed to propose curated (hand-labeled) ground truth images through @WendyKan post <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/50518#288558\">here</a></p>\n\n<p>I guess you can propose to Kaggle admins hand labeled data that you both make publicly available and announce in the official external data thread (which I am still looking for).</p>\n\n<p>Hopefully they will be OK with that.</p>",
      "rawMarkdown": "In \"data science bowl 2018\" competition, Kaggle admins allowed to propose curated (hand-labeled) ground truth images through @WendyKan post [here][1]\n\nI guess you can propose to Kaggle admins hand labeled data that you both make publicly available and announce in the official external data thread (which I am still looking for).\n\nHopefully they will be OK with that.\n\n\n  [1]: https://www.kaggle.com/c/data-science-bowl-2018/discussion/50518#288558",
      "votes": null
    },
    {
      "id": "321062",
      "postDate": "04/30/2018 14:09:59",
      "content": "<p>@Hoel Plantec, I agree with you now, as that makes sense. We could really use admin clarification here.</p>",
      "rawMarkdown": "Hoel Plantec, I agree with you now, as that makes sense. We could really use admin clarification here.",
      "votes": null
    },
    {
      "id": "321177",
      "postDate": "04/30/2018 19:21:26",
      "content": "<p>Hand labeling on the train set is fine. Hand labeling on the test set is prohibited. </p>",
      "rawMarkdown": "Hand labeling on the train set is fine. Hand labeling on the test set is prohibited.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 319974,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "04/27/2018 08:05:53",
      "content": "<p>This is a good question. But given that there are 155,197 unique images in the train and test set to annotate, that would be quite some dedication! Even if each image took 5sec to annotate, you're looking at 215h 33m of non-stop work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 319981,
          "author_name": "jantkacik",
          "author_url": "",
          "post_date": "04/27/2018 08:26:40",
          "content": "<p>You don`t need to annotate all the images to have reasonably working feature extractor in my opinion. Of course the more annotations you have the better.</p>\n\n<p>What is your interpretation of the rules ? Do you think it is allowed ? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320038,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "04/27/2018 10:35:16",
      "content": "<p>Actually this is <strong>explicitly against the rules</strong>. I just did a search through the rules and it says \"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 320100,
          "author_name": "jantkacik",
          "author_url": "",
          "post_date": "04/27/2018 13:31:58",
          "content": "<p>Does it apply to labeling training data ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320196,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "04/27/2018 20:41:51",
          "content": "<p>Yes, it applies to any models, train data, test data, or anything else.  No hand labeling or human prediction allowed, in any aspect, whatsoever.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321012,
          "author_name": "hoelplantec",
          "author_url": "",
          "post_date": "04/30/2018 11:25:15",
          "content": "<p>To my understanding, this doesn't mean you can't use intermediate labels in your models. This would mean that you can't label yourself the output of your global classifier, to submit these handmade labels. \nAll the more since external data is permitted in the  competition, that wouldn't make sense to prevent competitors from using intermediate hand labels to train their models. (for the use of external data see the competition rules in \"Competition-specific terms/External data\")</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321019,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "04/30/2018 12:15:57",
          "content": "<p>In \"data science bowl 2018\" competition, Kaggle admins allowed to propose curated (hand-labeled) ground truth images through @WendyKan post <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/50518#288558\">here</a></p>\n\n<p>I guess you can propose to Kaggle admins hand labeled data that you both make publicly available and announce in the official external data thread (which I am still looking for).</p>\n\n<p>Hopefully they will be OK with that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321062,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "04/30/2018 14:09:59",
          "content": "<p>@Hoel Plantec, I agree with you now, as that makes sense. We could really use admin clarification here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320141,
      "author_name": "alexfir",
      "author_url": "",
      "post_date": "04/27/2018 16:01:03",
      "content": "<p>Why not use deal_probability as proxy for 'niceness'?</p>",
      "votes": null,
      "replies": [
        {
          "id": 320195,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "04/27/2018 20:41:17",
          "content": "<p>That's what I was thinking...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321177,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/30/2018 19:21:26",
      "content": "<p>Hand labeling on the train set is fine. Hand labeling on the test set is prohibited. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "319647": "Hi, \n\nare there any rules about using my own custom annotations for training model ? Is it allowed ?\n\nFor example I will create my own derived dataset from images provided with my own annotation of \"niceness\" of image, later train feature extractor with this self-annotated data and use it as feature. Is it allowed according to the rules ? \n\nThanks\nJan",
    "319974": "This is a good question. But given that there are 155,197 unique images in the train and test set to annotate, that would be quite some dedication! Even if each image took 5sec to annotate, you're looking at 215h 33m of non-stop work!",
    "319981": "You don`t need to annotate all the images to have reasonably working feature extractor in my opinion. Of course the more annotations you have the better.\n\nWhat is your interpretation of the rules ? Do you think it is allowed ?",
    "320038": "Actually this is **explicitly against the rules**. I just did a search through the rules and it says \"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\"",
    "320100": "Does it apply to labeling training data ?",
    "320141": "Why not use deal_probability as proxy for 'niceness'?",
    "320195": "That's what I was thinking...",
    "320196": "Yes, it applies to any models, train data, test data, or anything else.  No hand labeling or human prediction allowed, in any aspect, whatsoever.",
    "321012": "To my understanding, this doesn't mean you can't use intermediate labels in your models. This would mean that you can't label yourself the output of your global classifier, to submit these handmade labels. \nAll the more since external data is permitted in the  competition, that wouldn't make sense to prevent competitors from using intermediate hand labels to train their models. (for the use of external data see the competition rules in \"Competition-specific terms/External data\")",
    "321019": "In \"data science bowl 2018\" competition, Kaggle admins allowed to propose curated (hand-labeled) ground truth images through @WendyKan post [here][1]\n\nI guess you can propose to Kaggle admins hand labeled data that you both make publicly available and announce in the official external data thread (which I am still looking for).\n\nHopefully they will be OK with that.\n\n\n  [1]: https://www.kaggle.com/c/data-science-bowl-2018/discussion/50518#288558",
    "321062": "Hoel Plantec, I agree with you now, as that makes sense. We could really use admin clarification here.",
    "321177": "Hand labeling on the train set is fine. Hand labeling on the test set is prohibited."
  },
  "source": "meta"
}