{
  "id": 19965,
  "title": "Manual labelling of training data?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/19965",
  "author_name": "",
  "post_date": "2016-04-06T07:51:44.223Z",
  "votes": 4,
  "comment_count": 26,
  "views": 5235,
  "content": "<p>Following on from clarification external data rules (i.e. no external data or pre-trained nets etc allowed).</p>\n\n<p>To aid with training, are we allowed to annotate the <em>training</em> images by hand, or must all processing be automated and provided as part of the training model? </p>\n\n<p>E.g. is it ok to add segmentation data for hands/face to training images (and <em>not</em> to test images)? I cannot see any competition rule that explicitly states this is not allowed, and it is not technically external data, nor is it manual labelling of test images, it is derived from the training data. However, this clearly goes beyond usual data analysis.</p>\n\n<p>I am not at this stage planning to do any manual labelling (seems far too much work!), but I can imagine a few pipeline designs where it would be useful. So it is worth clarifying before someone puts the effort in.</p>",
  "messages": [
    {
      "id": "113921",
      "postDate": "04/06/2016 07:51:44",
      "content": "<p>Following on from clarification external data rules (i.e. no external data or pre-trained nets etc allowed).</p>\n\n<p>To aid with training, are we allowed to annotate the <em>training</em> images by hand, or must all processing be automated and provided as part of the training model? </p>\n\n<p>E.g. is it ok to add segmentation data for hands/face to training images (and <em>not</em> to test images)? I cannot see any competition rule that explicitly states this is not allowed, and it is not technically external data, nor is it manual labelling of test images, it is derived from the training data. However, this clearly goes beyond usual data analysis.</p>\n\n<p>I am not at this stage planning to do any manual labelling (seems far too much work!), but I can imagine a few pipeline designs where it would be useful. So it is worth clarifying before someone puts the effort in.</p>",
      "rawMarkdown": "Following on from clarification external data rules (i.e. no external data or pre-trained nets etc allowed).\r\n\r\nTo aid with training, are we allowed to annotate the *training* images by hand, or must all processing be automated and provided as part of the training model? \r\n\r\nE.g. is it ok to add segmentation data for hands/face to training images (and *not* to test images)? I cannot see any competition rule that explicitly states this is not allowed, and it is not technically external data, nor is it manual labelling of test images, it is derived from the training data. However, this clearly goes beyond usual data analysis.\r\n\r\nI am not at this stage planning to do any manual labelling (seems far too much work!), but I can imagine a few pipeline designs where it would be useful. So it is worth clarifying before someone puts the effort in.",
      "votes": null
    },
    {
      "id": "113938",
      "postDate": "04/06/2016 10:52:57",
      "content": "<p>In general, you can do whatever you want to the training images (crop, label, etc), but you can't modify or label the test images.</p>\n\n<p>According to my reading of the rules, William's comment also applies to this contest:</p>\n\n<p>&quot;Use this rule of thumb - If you were to get a totally new test set tomorrow, your method should be able to classify it with comparable performance, without manual intervention.&quot;</p>\n\n<p><a href=\"https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743#post64743\">https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743#post64743</a></p>",
      "rawMarkdown": "In general, you can do whatever you want to the training images (crop, label, etc), but you can't modify or label the test images.\r\n\r\nAccording to my reading of the rules, William's comment also applies to this contest:\r\n\r\n\"Use this rule of thumb - If you were to get a totally new test set tomorrow, your method should be able to classify it with comparable performance, without manual intervention.\"\r\n\r\nhttps://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743#post64743",
      "votes": null
    },
    {
      "id": "113962",
      "postDate": "04/06/2016 13:25:11",
      "content": "<p>If that is the case, then a potentially workable strategy would be to crop out many examples of hands and heads from the training examples (plus of course negative examples without hands or heads in them), create identifiers for hands and heads, then use those to pick out parts of the image to run a final stage image classifier on.</p>\n\n<p>I guess it would be a few hours (up to a few days) work to generate enough heads/hands samples.</p>\n\n<p>A similar issue might be worthwhile to predict camera angle from some sub-part of the image, by picking out fixed parts of the car. However, that <em>might</em> be possible with more old-school non-ML approaches, because there are some nice regular shapes to work from.</p>",
      "rawMarkdown": "If that is the case, then a potentially workable strategy would be to crop out many examples of hands and heads from the training examples (plus of course negative examples without hands or heads in them), create identifiers for hands and heads, then use those to pick out parts of the image to run a final stage image classifier on.\r\n\r\nI guess it would be a few hours (up to a few days) work to generate enough heads/hands samples.\r\n\r\nA similar issue might be worthwhile to predict camera angle from some sub-part of the image, by picking out fixed parts of the car. However, that *might* be possible with more old-school non-ML approaches, because there are some nice regular shapes to work from.",
      "votes": null
    },
    {
      "id": "113964",
      "postDate": "04/06/2016 14:23:31",
      "content": "<p>A similar strategy was used in <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\">benchmark code</a> for the Right Whales competition.</p>\n\n<p>Step 1: Find the whale in the image</p>\n\n<p>Step 2: Classify the whale</p>\n\n<p>I think this will be a fun competition to learn more about the tools and strategies used in image recognition.</p>",
      "rawMarkdown": "A similar strategy was used in [benchmark code][1] for the Right Whales competition.\r\n\r\nStep 1: Find the whale in the image\r\n\r\nStep 2: Classify the whale\r\n\r\nI think this will be a fun competition to learn more about the tools and strategies used in image recognition.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales",
      "votes": null
    },
    {
      "id": "114037",
      "postDate": "04/07/2016 06:26:33",
      "content": "<blockquote>\n  <p><em>but you can't modify or label the test images.</em></p>\n</blockquote>\n\n<p>The way how that's worded indicates that you can't even convert the test images to gray scale.</p>\n\n<p>Shouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?</p>",
      "rawMarkdown": "> *but you can't modify or label the test images.*\r\n\r\nThe way how that's worded indicates that you can't even convert the test images to gray scale.\r\n\r\nShouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?",
      "votes": null
    },
    {
      "id": "114045",
      "postDate": "04/07/2016 08:02:05",
      "content": "<p>[quote=Gerard Toonstra;114037]</p>\n\n<blockquote>\n  <p><em>but you can't modify or label the test images.</em></p>\n</blockquote>\n\n<p>The way how that's worded indicates that you can't even convert the test images to gray scale.</p>\n\n<p>[/quote]</p>\n\n<p>Only on a very literal out-of-context reading. The quote is not rules legalese, it is just someone trying to help you understand the competition limits.</p>\n\n<p>The restriction is on processing steps using human judgement per-image. Any fully automated processing driven by data in the training and test sets is clearly fine. Otherwise, even normalising the pixel values to feed into a neural net would not be possible, which is nonsense.</p>",
      "rawMarkdown": "[quote=Gerard Toonstra;114037]\r\n\r\n> *but you can't modify or label the test images.*\r\n\r\nThe way how that's worded indicates that you can't even convert the test images to gray scale.\r\n\r\n[/quote]\r\n\r\nOnly on a very literal out-of-context reading. The quote is not rules legalese, it is just someone trying to help you understand the competition limits.\r\n\r\nThe restriction is on processing steps using human judgement per-image. Any fully automated processing driven by data in the training and test sets is clearly fine. Otherwise, even normalising the pixel values to feed into a neural net would not be possible, which is nonsense.",
      "votes": null
    },
    {
      "id": "114076",
      "postDate": "04/07/2016 13:04:53",
      "content": "<p>[quote=Gerard Toonstra;114037]</p>\n\n<blockquote>\n  <p><em>but you can't modify or label the test images.</em></p>\n</blockquote>\n\n<p>The way how that's worded indicates that you can't even convert the test images to gray scale.</p>\n\n<p>Shouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?</p>\n\n<p>[/quote]</p>\n\n<p>More accurate would be - you can't modify or label the test images <em>manually</em>. You can do whatever you want to the images, as long as the process is automated.</p>",
      "rawMarkdown": "[quote=Gerard Toonstra;114037]\r\n\r\n> *but you can't modify or label the test images.*\r\n\r\nThe way how that's worded indicates that you can't even convert the test images to gray scale.\r\n\r\nShouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?\r\n\r\n[/quote]\r\n\r\nMore accurate would be - you can't modify or label the test images _manually_. You can do whatever you want to the images, as long as the process is automated.",
      "votes": null
    },
    {
      "id": "114088",
      "postDate": "04/07/2016 14:38:05",
      "content": "<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>",
      "rawMarkdown": "Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.",
      "votes": null
    },
    {
      "id": "114090",
      "postDate": "04/07/2016 14:40:55",
      "content": "<p>[quote=DavidGbodiOdaibo;114088]</p>\n\n<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>\n\n<p>[/quote]</p>\n\n<p>I think there are 22,424 images</p>",
      "rawMarkdown": "[quote=DavidGbodiOdaibo;114088]\r\n\r\nJust out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.\r\n\r\n[/quote]\r\n \r\nI think there are 22,424 images",
      "votes": null
    },
    {
      "id": "114095",
      "postDate": "04/07/2016 14:48:02",
      "content": "<p>[quote=DavidGbodiOdaibo;114088]</p>\n\n<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>\n\n<p>[/quote]</p>\n\n<p>FYI, there is nothing in the rules that prevents an <em>open-sourced</em> annotation library of the (training!) images. This was done in the Right Whales competition.</p>\n\n<p><a href=\"https://github.com/Smerity/right_whale_hunt\">https://github.com/Smerity/right_whale_hunt</a></p>",
      "rawMarkdown": "[quote=DavidGbodiOdaibo;114088]\r\n\r\nJust out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.\r\n\r\n[/quote]\r\n\r\nFYI, there is nothing in the rules that prevents an _open-sourced_ annotation library of the (training!) images. This was done in the Right Whales competition.\r\n\r\nhttps://github.com/Smerity/right_whale_hunt",
      "votes": null
    },
    {
      "id": "114100",
      "postDate": "04/07/2016 14:55:30",
      "content": "<p>[quote=DavidGbodiOdaibo;114088]</p>\n\n<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>\n\n<p>[/quote]</p>\n\n<p>Train: 22424</p>\n\n<p>Test: 79726</p>\n\n<p>But hand labelling is <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">also forbidden</a>.</p>\n\n<p>[quote=<a>https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules]</a></p>\n\n<p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "[quote=DavidGbodiOdaibo;114088]\r\n\r\nJust out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.\r\n\r\n[/quote]\r\n\r\nTrain: 22424\r\n\r\nTest: 79726\r\n\r\nBut hand labelling is [also forbidden][1].\r\n\r\n[quote=https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules]\r\n\r\nSubmissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules",
      "votes": null
    },
    {
      "id": "114102",
      "postDate": "04/07/2016 15:00:29",
      "content": "<p>[quote=Scott Lowe;114100]</p>\n\n<p>But hand labelling is <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">also forbidden</a>.</p>\n\n<p>[/quote]</p>\n\n<p>From the rules:</p>\n\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction <strong>of the validation dataset or test data records</strong>.&quot;</p>\n\n<p>It says nothing about hand labeling the training data.</p>",
      "rawMarkdown": "[quote=Scott Lowe;114100]\r\n\r\nBut hand labelling is [also forbidden][1].\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\r\n\r\n[/quote]\r\n\r\nFrom the rules:\r\n\r\n\"Submissions may not use or incorporate information from hand labeling or human prediction **of the validation dataset or test data records**.\"\r\n\r\nIt says nothing about hand labeling the training data.",
      "votes": null
    },
    {
      "id": "114113",
      "postDate": "04/07/2016 15:32:26",
      "content": "<p>[quote=inversion;114102]</p>\n\n<p>From the rules:</p>\n\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction <strong>of the validation dataset or test data records</strong>.&quot;</p>\n\n<p>It says nothing about hand labeling the training data.</p>\n\n<p>[/quote]</p>\n\n<p>Which is why I asked. From feedback here about previous competitions, it looks like it could be allowed. </p>\n\n<p>It would still be nice to get an official ruling on that, before someone goes ahead and puts in hours of labelling effort on training data. And for future competitions, maybe cover whether the approach is acceptable by default on the rules page or as a sticky forum thread with clarifications.</p>",
      "rawMarkdown": "[quote=inversion;114102]\r\n\r\nFrom the rules:\r\n\r\n\"Submissions may not use or incorporate information from hand labeling or human prediction **of the validation dataset or test data records**.\"\r\n\r\nIt says nothing about hand labeling the training data.\r\n\r\n[/quote]\r\n\r\nWhich is why I asked. From feedback here about previous competitions, it looks like it could be allowed. \r\n\r\nIt would still be nice to get an official ruling on that, before someone goes ahead and puts in hours of labelling effort on training data. And for future competitions, maybe cover whether the approach is acceptable by default on the rules page or as a sticky forum thread with clarifications.",
      "votes": null
    },
    {
      "id": "114152",
      "postDate": "04/07/2016 20:07:45",
      "content": "<p>[quote=inversion;114102]</p>\n\n<p>[quote=Scott Lowe;114100]</p>\n\n<p>But hand labelling is <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">also forbidden</a>.</p>\n\n<p>[/quote]</p>\n\n<p>From the rules:</p>\n\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction <strong>of the validation dataset or test data records</strong>.&quot;</p>\n\n<p>It says nothing about hand labeling the training data.</p>\n\n<p>[/quote]</p>\n\n<p>Ah, yes, I should have paid more attention to what I was quoting. (I think absent mindedly misread <em>validation</em> as <em>training</em> originally.)</p>\n\n<p>Thanks then to Neil for pointing this out. It would certainly be good to get some clarification on this from the admins.</p>",
      "rawMarkdown": "[quote=inversion;114102]\r\n\r\n[quote=Scott Lowe;114100]\r\n\r\nBut hand labelling is [also forbidden][1].\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\r\n\r\n[/quote]\r\n\r\nFrom the rules:\r\n\r\n\"Submissions may not use or incorporate information from hand labeling or human prediction **of the validation dataset or test data records**.\"\r\n\r\nIt says nothing about hand labeling the training data.\r\n\r\n[/quote]\r\n\r\nAh, yes, I should have paid more attention to what I was quoting. (I think absent mindedly misread *validation* as *training* originally.)\r\n\r\nThanks then to Neil for pointing this out. It would certainly be good to get some clarification on this from the admins.",
      "votes": null
    },
    {
      "id": "114172",
      "postDate": "04/07/2016 23:26:34",
      "content": "<p>Similar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good. </p>",
      "rawMarkdown": "Similar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good.",
      "votes": null
    },
    {
      "id": "114285",
      "postDate": "04/08/2016 21:39:13",
      "content": "<p>Maybe I'm missing something, but can we use external data/pre trained models to produce labeling of the training data and then build a model which only consumes the training data and the labels? Asumming that everyone is allowed to use some kind of labels for training data, there should be no constraints on how the labeling is obtained. Especially if there is no way to prove that labeling was done by hand or by some other means?</p>",
      "rawMarkdown": "Maybe I'm missing something, but can we use external data/pre trained models to produce labeling of the training data and then build a model which only consumes the training data and the labels? Asumming that everyone is allowed to use some kind of labels for training data, there should be no constraints on how the labeling is obtained. Especially if there is no way to prove that labeling was done by hand or by some other means?",
      "votes": null
    },
    {
      "id": "114286",
      "postDate": "04/08/2016 21:55:35",
      "content": "<p>As far as I understand it you can preprocess training and test data in any way you want. There is only one restriction. TEST data processing must be AUTOMATED. That doesn't apply to training data which could be processed/labeled manually as well.</p>",
      "rawMarkdown": "As far as I understand it you can preprocess training and test data in any way you want. There is only one restriction. TEST data processing must be AUTOMATED. That doesn't apply to training data which could be processed/labeled manually as well.",
      "votes": null
    },
    {
      "id": "114612",
      "postDate": "04/12/2016 13:29:24",
      "content": "<p>[quote=Wendy Kan;114172]</p>\n\n<p>Similar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good. </p>\n\n<p>[/quote]</p>\n\n<p>If hand annotation is aided by a pre-trained model, does that count as using &quot;external data&quot;?</p>\n\n<p>A more concrete example would be something similar to what <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20024/opencv-face-detection-external-data/114351#post114351\">Neil Slater mention here</a></p>\n\n<p>[quote]</p>\n\n<p>However the following would be allowed IMO:</p>\n\n<ul>\n<li>Use opencv to extract all driver faces in training set (but not the test set).</li>\n<li>Train your own face recogniser from those extracted faces.</li>\n<li>Use your new face recogniser to identify face position in the test images as part of a ML pipeline.</li>\n</ul>\n\n<p>[/quote]</p>",
      "rawMarkdown": "[quote=Wendy Kan;114172]\r\n\r\nSimilar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good. \r\n\r\n[/quote]\r\n\r\nIf hand annotation is aided by a pre-trained model, does that count as using \"external data\"?\r\n\r\nA more concrete example would be something similar to what [Neil Slater mention here](https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20024/opencv-face-detection-external-data/114351#post114351)\r\n\r\n[quote]\r\n\r\nHowever the following would be allowed IMO:\r\n\r\n- Use opencv to extract all driver faces in training set (but not the test set).\r\n- Train your own face recogniser from those extracted faces.\r\n- Use your new face recogniser to identify face position in the test images as part of a ML pipeline.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "114670",
      "postDate": "04/12/2016 20:09:01",
      "content": "<p>[quote=felixlaumon;114612]</p>\n\n<p>If hand annotation is aided by a pre-trained model, does that count as using &quot;external data&quot;?</p>\n\n<p>[/quote]</p>\n\n<p>Hand annotation of training images is fine, therefore using pre-trained models for that purpose is fine. A similar thinking is that your hand labeling has some pre-trained model (your brain) built in. But these pre-trained models should not be applied in your test dataset. </p>",
      "rawMarkdown": "[quote=felixlaumon;114612]\r\n\r\nIf hand annotation is aided by a pre-trained model, does that count as using \"external data\"?\r\n\r\n[/quote]\r\n\r\nHand annotation of training images is fine, therefore using pre-trained models for that purpose is fine. A similar thinking is that your hand labeling has some pre-trained model (your brain) built in. But these pre-trained models should not be applied in your test dataset.",
      "votes": null
    },
    {
      "id": "114673",
      "postDate": "04/12/2016 20:24:22",
      "content": "<p>[quote=Wendy Kan;114670]</p>\n\n<p>Hand annotation of training images is fine . . . </p>\n\n<p>[/quote]</p>\n\n<p>Can we annotate other body parts too, like the head?</p>",
      "rawMarkdown": "[quote=Wendy Kan;114670]\r\n\r\nHand annotation of training images is fine . . . \r\n\r\n[/quote]\r\n\r\nCan we annotate other body parts too, like the head?",
      "votes": null
    },
    {
      "id": "114685",
      "postDate": "04/12/2016 21:28:53",
      "content": "<p>[quote=inversion;114673]</p>\n\n<p>[quote=Wendy Kan;114670]</p>\n\n<p>Hand annotation of training images is fine . . . </p>\n\n<p>[/quote]</p>\n\n<p>Can we annotate other body parts too, like the head?</p>\n\n<p>[/quote]</p>\n\n<p>Yes. </p>",
      "rawMarkdown": "[quote=inversion;114673]\r\n\r\n[quote=Wendy Kan;114670]\r\n\r\nHand annotation of training images is fine . . . \r\n\r\n[/quote]\r\n\r\nCan we annotate other body parts too, like the head?\r\n\r\n[/quote]\r\n\r\nYes.",
      "votes": null
    },
    {
      "id": "114696",
      "postDate": "04/12/2016 23:01:35",
      "content": "<p>Can I use the OpenCV human detector for training?\nIs this pre-training model?</p>",
      "rawMarkdown": "Can I use the OpenCV human detector for training?\r\nIs this pre-training model?",
      "votes": null
    },
    {
      "id": "114713",
      "postDate": "04/13/2016 04:39:50",
      "content": "<p>Wendy, Thanks for the reply. I like your analogy between human brain to pre-trained model :).</p>\n\n<p>Since questions about hand-labelling comes up in almost every computer vision competition, may I suggest adding an FAQ about hand annotation to the Kaggle wiki?</p>\n\n<p>You and other admins have made various clarifications about what's allowed but they are all spread out in different forums. Having them in one place will certainly help!</p>",
      "rawMarkdown": "Wendy, Thanks for the reply. I like your analogy between human brain to pre-trained model :).\r\n\r\nSince questions about hand-labelling comes up in almost every computer vision competition, may I suggest adding an FAQ about hand annotation to the Kaggle wiki?\r\n\r\nYou and other admins have made various clarifications about what's allowed but they are all spread out in different forums. Having them in one place will certainly help!",
      "votes": null
    },
    {
      "id": "114721",
      "postDate": "04/13/2016 07:10:44",
      "content": "<p>[quote=tereka;114696]</p>\n\n<p>Can I use the OpenCV human detector for training?\nIs this pre-training model?</p>\n\n<p>[/quote]</p>\n\n<p>Yes that is a pre-trained model that used external data - pretty much all off-the-shelf object detectors are, except the really simple line, circle, connected region filters. </p>\n\n<p>So it looks like you could use the human detector to help manage the training data, but it should not be run on the test images as part of prediction code.</p>",
      "rawMarkdown": "[quote=tereka;114696]\r\n\r\nCan I use the OpenCV human detector for training?\r\nIs this pre-training model?\r\n\r\n[/quote]\r\n\r\nYes that is a pre-trained model that used external data - pretty much all off-the-shelf object detectors are, except the really simple line, circle, connected region filters. \r\n\r\nSo it looks like you could use the human detector to help manage the training data, but it should not be run on the test images as part of prediction code.",
      "votes": null
    },
    {
      "id": "114731",
      "postDate": "04/13/2016 10:52:16",
      "content": "<p>Neil. Thanks for your reply.\nI only use to help putting the training label.</p>",
      "rawMarkdown": "Neil. Thanks for your reply.\r\nI only use to help putting the training label.",
      "votes": null
    },
    {
      "id": "120204",
      "postDate": "05/16/2016 10:41:42",
      "content": "<p>[quote=Wendy Kan;114685]</p>\n\n<p>[quote=inversion;114673]</p>\n\n<p>[quote=Wendy Kan;114670]</p>\n\n<p>Hand annotation of training images is fine . . . </p>\n\n<p>[/quote]</p>\n\n<p>Can we annotate other body parts too, like the head?</p>\n\n<p>[/quote]</p>\n\n<p>Yes. </p>\n\n<p>[/quote]\nThen, should we share the annotations? and where will be found them?</p>",
      "rawMarkdown": "[quote=Wendy Kan;114685]\r\n\r\n[quote=inversion;114673]\r\n\r\n[quote=Wendy Kan;114670]\r\n\r\nHand annotation of training images is fine . . . \r\n\r\n[/quote]\r\n\r\nCan we annotate other body parts too, like the head?\r\n\r\n[/quote]\r\n\r\nYes. \r\n\r\n[/quote]\r\nThen, should we share the annotations? and where will be found them?",
      "votes": null
    },
    {
      "id": "120214",
      "postDate": "05/16/2016 12:58:39",
      "content": "<p>[quote=long long;120204]</p>\n\n<p>Then, should we share the annotations? and where will be found them?</p>\n\n<p>[/quote]</p>\n\n<p>You are under no obligation to share your own personal annotations. (But, you certainly may if you want.) Some may choose to crowd source annotations. If that is the case, they must be shared to the forum, since crowd sourcing annotations is sharing across teams. (Thus, it cannot be done privately.)</p>",
      "rawMarkdown": "[quote=long long;120204]\r\n\r\nThen, should we share the annotations? and where will be found them?\r\n\r\n[/quote]\r\n\r\nYou are under no obligation to share your own personal annotations. (But, you certainly may if you want.) Some may choose to crowd source annotations. If that is the case, they must be shared to the forum, since crowd sourcing annotations is sharing across teams. (Thus, it cannot be done privately.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 113938,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/06/2016 10:52:57",
      "content": "<p>In general, you can do whatever you want to the training images (crop, label, etc), but you can't modify or label the test images.</p>\n\n<p>According to my reading of the rules, William's comment also applies to this contest:</p>\n\n<p>&quot;Use this rule of thumb - If you were to get a totally new test set tomorrow, your method should be able to classify it with comparable performance, without manual intervention.&quot;</p>\n\n<p><a href=\"https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743#post64743\">https://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743#post64743</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 113962,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/06/2016 13:25:11",
      "content": "<p>If that is the case, then a potentially workable strategy would be to crop out many examples of hands and heads from the training examples (plus of course negative examples without hands or heads in them), create identifiers for hands and heads, then use those to pick out parts of the image to run a final stage image classifier on.</p>\n\n<p>I guess it would be a few hours (up to a few days) work to generate enough heads/hands samples.</p>\n\n<p>A similar issue might be worthwhile to predict camera angle from some sub-part of the image, by picking out fixed parts of the car. However, that <em>might</em> be possible with more old-school non-ML approaches, because there are some nice regular shapes to work from.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 113964,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/06/2016 14:23:31",
      "content": "<p>A similar strategy was used in <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales\">benchmark code</a> for the Right Whales competition.</p>\n\n<p>Step 1: Find the whale in the image</p>\n\n<p>Step 2: Classify the whale</p>\n\n<p>I think this will be a fun competition to learn more about the tools and strategies used in image recognition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114037,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "04/07/2016 06:26:33",
      "content": "<blockquote>\n  <p><em>but you can't modify or label the test images.</em></p>\n</blockquote>\n\n<p>The way how that's worded indicates that you can't even convert the test images to gray scale.</p>\n\n<p>Shouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114045,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/07/2016 08:02:05",
      "content": "<p>[quote=Gerard Toonstra;114037]</p>\n\n<blockquote>\n  <p><em>but you can't modify or label the test images.</em></p>\n</blockquote>\n\n<p>The way how that's worded indicates that you can't even convert the test images to gray scale.</p>\n\n<p>[/quote]</p>\n\n<p>Only on a very literal out-of-context reading. The quote is not rules legalese, it is just someone trying to help you understand the competition limits.</p>\n\n<p>The restriction is on processing steps using human judgement per-image. Any fully automated processing driven by data in the training and test sets is clearly fine. Otherwise, even normalising the pixel values to feed into a neural net would not be possible, which is nonsense.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114076,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/07/2016 13:04:53",
      "content": "<p>[quote=Gerard Toonstra;114037]</p>\n\n<blockquote>\n  <p><em>but you can't modify or label the test images.</em></p>\n</blockquote>\n\n<p>The way how that's worded indicates that you can't even convert the test images to gray scale.</p>\n\n<p>Shouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?</p>\n\n<p>[/quote]</p>\n\n<p>More accurate would be - you can't modify or label the test images <em>manually</em>. You can do whatever you want to the images, as long as the process is automated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114088,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "04/07/2016 14:38:05",
      "content": "<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114090,
      "author_name": "whamfish",
      "author_url": "",
      "post_date": "04/07/2016 14:40:55",
      "content": "<p>[quote=DavidGbodiOdaibo;114088]</p>\n\n<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>\n\n<p>[/quote]</p>\n\n<p>I think there are 22,424 images</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114095,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/07/2016 14:48:02",
      "content": "<p>[quote=DavidGbodiOdaibo;114088]</p>\n\n<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>\n\n<p>[/quote]</p>\n\n<p>FYI, there is nothing in the rules that prevents an <em>open-sourced</em> annotation library of the (training!) images. This was done in the Right Whales competition.</p>\n\n<p><a href=\"https://github.com/Smerity/right_whale_hunt\">https://github.com/Smerity/right_whale_hunt</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114100,
      "author_name": "scottclowe",
      "author_url": "",
      "post_date": "04/07/2016 14:55:30",
      "content": "<p>[quote=DavidGbodiOdaibo;114088]</p>\n\n<p>Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.</p>\n\n<p>[/quote]</p>\n\n<p>Train: 22424</p>\n\n<p>Test: 79726</p>\n\n<p>But hand labelling is <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">also forbidden</a>.</p>\n\n<p>[quote=<a>https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules]</a></p>\n\n<p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114102,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/07/2016 15:00:29",
      "content": "<p>[quote=Scott Lowe;114100]</p>\n\n<p>But hand labelling is <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">also forbidden</a>.</p>\n\n<p>[/quote]</p>\n\n<p>From the rules:</p>\n\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction <strong>of the validation dataset or test data records</strong>.&quot;</p>\n\n<p>It says nothing about hand labeling the training data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114113,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/07/2016 15:32:26",
      "content": "<p>[quote=inversion;114102]</p>\n\n<p>From the rules:</p>\n\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction <strong>of the validation dataset or test data records</strong>.&quot;</p>\n\n<p>It says nothing about hand labeling the training data.</p>\n\n<p>[/quote]</p>\n\n<p>Which is why I asked. From feedback here about previous competitions, it looks like it could be allowed. </p>\n\n<p>It would still be nice to get an official ruling on that, before someone goes ahead and puts in hours of labelling effort on training data. And for future competitions, maybe cover whether the approach is acceptable by default on the rules page or as a sticky forum thread with clarifications.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114152,
      "author_name": "scottclowe",
      "author_url": "",
      "post_date": "04/07/2016 20:07:45",
      "content": "<p>[quote=inversion;114102]</p>\n\n<p>[quote=Scott Lowe;114100]</p>\n\n<p>But hand labelling is <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\">also forbidden</a>.</p>\n\n<p>[/quote]</p>\n\n<p>From the rules:</p>\n\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction <strong>of the validation dataset or test data records</strong>.&quot;</p>\n\n<p>It says nothing about hand labeling the training data.</p>\n\n<p>[/quote]</p>\n\n<p>Ah, yes, I should have paid more attention to what I was quoting. (I think absent mindedly misread <em>validation</em> as <em>training</em> originally.)</p>\n\n<p>Thanks then to Neil for pointing this out. It would certainly be good to get some clarification on this from the admins.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114172,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/07/2016 23:26:34",
      "content": "<p>Similar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114285,
      "author_name": "jureso",
      "author_url": "",
      "post_date": "04/08/2016 21:39:13",
      "content": "<p>Maybe I'm missing something, but can we use external data/pre trained models to produce labeling of the training data and then build a model which only consumes the training data and the labels? Asumming that everyone is allowed to use some kind of labels for training data, there should be no constraints on how the labeling is obtained. Especially if there is no way to prove that labeling was done by hand or by some other means?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114286,
      "author_name": "wlodekf",
      "author_url": "",
      "post_date": "04/08/2016 21:55:35",
      "content": "<p>As far as I understand it you can preprocess training and test data in any way you want. There is only one restriction. TEST data processing must be AUTOMATED. That doesn't apply to training data which could be processed/labeled manually as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114612,
      "author_name": "felixlaumon",
      "author_url": "",
      "post_date": "04/12/2016 13:29:24",
      "content": "<p>[quote=Wendy Kan;114172]</p>\n\n<p>Similar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good. </p>\n\n<p>[/quote]</p>\n\n<p>If hand annotation is aided by a pre-trained model, does that count as using &quot;external data&quot;?</p>\n\n<p>A more concrete example would be something similar to what <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20024/opencv-face-detection-external-data/114351#post114351\">Neil Slater mention here</a></p>\n\n<p>[quote]</p>\n\n<p>However the following would be allowed IMO:</p>\n\n<ul>\n<li>Use opencv to extract all driver faces in training set (but not the test set).</li>\n<li>Train your own face recogniser from those extracted faces.</li>\n<li>Use your new face recogniser to identify face position in the test images as part of a ML pipeline.</li>\n</ul>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114670,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/12/2016 20:09:01",
      "content": "<p>[quote=felixlaumon;114612]</p>\n\n<p>If hand annotation is aided by a pre-trained model, does that count as using &quot;external data&quot;?</p>\n\n<p>[/quote]</p>\n\n<p>Hand annotation of training images is fine, therefore using pre-trained models for that purpose is fine. A similar thinking is that your hand labeling has some pre-trained model (your brain) built in. But these pre-trained models should not be applied in your test dataset. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114673,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/12/2016 20:24:22",
      "content": "<p>[quote=Wendy Kan;114670]</p>\n\n<p>Hand annotation of training images is fine . . . </p>\n\n<p>[/quote]</p>\n\n<p>Can we annotate other body parts too, like the head?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114685,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/12/2016 21:28:53",
      "content": "<p>[quote=inversion;114673]</p>\n\n<p>[quote=Wendy Kan;114670]</p>\n\n<p>Hand annotation of training images is fine . . . </p>\n\n<p>[/quote]</p>\n\n<p>Can we annotate other body parts too, like the head?</p>\n\n<p>[/quote]</p>\n\n<p>Yes. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114696,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "04/12/2016 23:01:35",
      "content": "<p>Can I use the OpenCV human detector for training?\nIs this pre-training model?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114713,
      "author_name": "felixlaumon",
      "author_url": "",
      "post_date": "04/13/2016 04:39:50",
      "content": "<p>Wendy, Thanks for the reply. I like your analogy between human brain to pre-trained model :).</p>\n\n<p>Since questions about hand-labelling comes up in almost every computer vision competition, may I suggest adding an FAQ about hand annotation to the Kaggle wiki?</p>\n\n<p>You and other admins have made various clarifications about what's allowed but they are all spread out in different forums. Having them in one place will certainly help!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114721,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/13/2016 07:10:44",
      "content": "<p>[quote=tereka;114696]</p>\n\n<p>Can I use the OpenCV human detector for training?\nIs this pre-training model?</p>\n\n<p>[/quote]</p>\n\n<p>Yes that is a pre-trained model that used external data - pretty much all off-the-shelf object detectors are, except the really simple line, circle, connected region filters. </p>\n\n<p>So it looks like you could use the human detector to help manage the training data, but it should not be run on the test images as part of prediction code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114731,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "04/13/2016 10:52:16",
      "content": "<p>Neil. Thanks for your reply.\nI only use to help putting the training label.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120204,
      "author_name": "lolong",
      "author_url": "",
      "post_date": "05/16/2016 10:41:42",
      "content": "<p>[quote=Wendy Kan;114685]</p>\n\n<p>[quote=inversion;114673]</p>\n\n<p>[quote=Wendy Kan;114670]</p>\n\n<p>Hand annotation of training images is fine . . . </p>\n\n<p>[/quote]</p>\n\n<p>Can we annotate other body parts too, like the head?</p>\n\n<p>[/quote]</p>\n\n<p>Yes. </p>\n\n<p>[/quote]\nThen, should we share the annotations? and where will be found them?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120214,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/16/2016 12:58:39",
      "content": "<p>[quote=long long;120204]</p>\n\n<p>Then, should we share the annotations? and where will be found them?</p>\n\n<p>[/quote]</p>\n\n<p>You are under no obligation to share your own personal annotations. (But, you certainly may if you want.) Some may choose to crowd source annotations. If that is the case, they must be shared to the forum, since crowd sourcing annotations is sharing across teams. (Thus, it cannot be done privately.)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "113921": "Following on from clarification external data rules (i.e. no external data or pre-trained nets etc allowed).\r\n\r\nTo aid with training, are we allowed to annotate the *training* images by hand, or must all processing be automated and provided as part of the training model? \r\n\r\nE.g. is it ok to add segmentation data for hands/face to training images (and *not* to test images)? I cannot see any competition rule that explicitly states this is not allowed, and it is not technically external data, nor is it manual labelling of test images, it is derived from the training data. However, this clearly goes beyond usual data analysis.\r\n\r\nI am not at this stage planning to do any manual labelling (seems far too much work!), but I can imagine a few pipeline designs where it would be useful. So it is worth clarifying before someone puts the effort in.",
    "113938": "In general, you can do whatever you want to the training images (crop, label, etc), but you can't modify or label the test images.\r\n\r\nAccording to my reading of the rules, William's comment also applies to this contest:\r\n\r\n\"Use this rule of thumb - If you were to get a totally new test set tomorrow, your method should be able to classify it with comparable performance, without manual intervention.\"\r\n\r\nhttps://www.kaggle.com/c/datasciencebowl/forums/t/12587/manual-vs-auto-feature-selection/64743#post64743",
    "113962": "If that is the case, then a potentially workable strategy would be to crop out many examples of hands and heads from the training examples (plus of course negative examples without hands or heads in them), create identifiers for hands and heads, then use those to pick out parts of the image to run a final stage image classifier on.\r\n\r\nI guess it would be a few hours (up to a few days) work to generate enough heads/hands samples.\r\n\r\nA similar issue might be worthwhile to predict camera angle from some sub-part of the image, by picking out fixed parts of the car. However, that *might* be possible with more old-school non-ML approaches, because there are some nice regular shapes to work from.",
    "113964": "A similar strategy was used in [benchmark code][1] for the Right Whales competition.\r\n\r\nStep 1: Find the whale in the image\r\n\r\nStep 2: Classify the whale\r\n\r\nI think this will be a fun competition to learn more about the tools and strategies used in image recognition.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/details/creating-a-face-detector-for-whales",
    "114037": "> *but you can't modify or label the test images.*\r\n\r\nThe way how that's worded indicates that you can't even convert the test images to gray scale.\r\n\r\nShouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?",
    "114045": "[quote=Gerard Toonstra;114037]\r\n\r\n> *but you can't modify or label the test images.*\r\n\r\nThe way how that's worded indicates that you can't even convert the test images to gray scale.\r\n\r\n[/quote]\r\n\r\nOnly on a very literal out-of-context reading. The quote is not rules legalese, it is just someone trying to help you understand the competition limits.\r\n\r\nThe restriction is on processing steps using human judgement per-image. Any fully automated processing driven by data in the training and test sets is clearly fine. Otherwise, even normalising the pixel values to feed into a neural net would not be possible, which is nonsense.",
    "114076": "[quote=Gerard Toonstra;114037]\r\n\r\n> *but you can't modify or label the test images.*\r\n\r\nThe way how that's worded indicates that you can't even convert the test images to gray scale.\r\n\r\nShouldn't we be able to do any operation that's fully automated without manual intervention or which introduces a process for retraining a network?\r\n\r\n[/quote]\r\n\r\nMore accurate would be - you can't modify or label the test images _manually_. You can do whatever you want to the images, as long as the process is automated.",
    "114088": "Just out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.",
    "114090": "[quote=DavidGbodiOdaibo;114088]\r\n\r\nJust out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.\r\n\r\n[/quote]\r\n \r\nI think there are 22,424 images",
    "114095": "[quote=DavidGbodiOdaibo;114088]\r\n\r\nJust out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.\r\n\r\n[/quote]\r\n\r\nFYI, there is nothing in the rules that prevents an _open-sourced_ annotation library of the (training!) images. This was done in the Right Whales competition.\r\n\r\nhttps://github.com/Smerity/right_whale_hunt",
    "114100": "[quote=DavidGbodiOdaibo;114088]\r\n\r\nJust out of curiosity how many training samples are there in the dataset? I am pondering whether to even get into this competition, hand labeling will be the only competitive solution since no pre-training is allowed, I am not sure I am that crazy to hand label tons of images.\r\n\r\n[/quote]\r\n\r\nTrain: 22424\r\n\r\nTest: 79726\r\n\r\nBut hand labelling is [also forbidden][1].\r\n\r\n[quote=https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules]\r\n\r\nSubmissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules",
    "114102": "[quote=Scott Lowe;114100]\r\n\r\nBut hand labelling is [also forbidden][1].\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\r\n\r\n[/quote]\r\n\r\nFrom the rules:\r\n\r\n\"Submissions may not use or incorporate information from hand labeling or human prediction **of the validation dataset or test data records**.\"\r\n\r\nIt says nothing about hand labeling the training data.",
    "114113": "[quote=inversion;114102]\r\n\r\nFrom the rules:\r\n\r\n\"Submissions may not use or incorporate information from hand labeling or human prediction **of the validation dataset or test data records**.\"\r\n\r\nIt says nothing about hand labeling the training data.\r\n\r\n[/quote]\r\n\r\nWhich is why I asked. From feedback here about previous competitions, it looks like it could be allowed. \r\n\r\nIt would still be nice to get an official ruling on that, before someone goes ahead and puts in hours of labelling effort on training data. And for future competitions, maybe cover whether the approach is acceptable by default on the rules page or as a sticky forum thread with clarifications.",
    "114152": "[quote=inversion;114102]\r\n\r\n[quote=Scott Lowe;114100]\r\n\r\nBut hand labelling is [also forbidden][1].\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/rules\r\n\r\n[/quote]\r\n\r\nFrom the rules:\r\n\r\n\"Submissions may not use or incorporate information from hand labeling or human prediction **of the validation dataset or test data records**.\"\r\n\r\nIt says nothing about hand labeling the training data.\r\n\r\n[/quote]\r\n\r\nAh, yes, I should have paid more attention to what I was quoting. (I think absent mindedly misread *validation* as *training* originally.)\r\n\r\nThanks then to Neil for pointing this out. It would certainly be good to get some clarification on this from the admins.",
    "114172": "Similar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good.",
    "114285": "Maybe I'm missing something, but can we use external data/pre trained models to produce labeling of the training data and then build a model which only consumes the training data and the labels? Asumming that everyone is allowed to use some kind of labels for training data, there should be no constraints on how the labeling is obtained. Especially if there is no way to prove that labeling was done by hand or by some other means?",
    "114286": "As far as I understand it you can preprocess training and test data in any way you want. There is only one restriction. TEST data processing must be AUTOMATED. That doesn't apply to training data which could be processed/labeled manually as well.",
    "114612": "[quote=Wendy Kan;114172]\r\n\r\nSimilar to the right whales competition, annotating the training images by hand is ok. As long as the final processing of the test set is fully automated and doesn't require any manual annotation, you're good. \r\n\r\n[/quote]\r\n\r\nIf hand annotation is aided by a pre-trained model, does that count as using \"external data\"?\r\n\r\nA more concrete example would be something similar to what [Neil Slater mention here](https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20024/opencv-face-detection-external-data/114351#post114351)\r\n\r\n[quote]\r\n\r\nHowever the following would be allowed IMO:\r\n\r\n- Use opencv to extract all driver faces in training set (but not the test set).\r\n- Train your own face recogniser from those extracted faces.\r\n- Use your new face recogniser to identify face position in the test images as part of a ML pipeline.\r\n\r\n[/quote]",
    "114670": "[quote=felixlaumon;114612]\r\n\r\nIf hand annotation is aided by a pre-trained model, does that count as using \"external data\"?\r\n\r\n[/quote]\r\n\r\nHand annotation of training images is fine, therefore using pre-trained models for that purpose is fine. A similar thinking is that your hand labeling has some pre-trained model (your brain) built in. But these pre-trained models should not be applied in your test dataset.",
    "114673": "[quote=Wendy Kan;114670]\r\n\r\nHand annotation of training images is fine . . . \r\n\r\n[/quote]\r\n\r\nCan we annotate other body parts too, like the head?",
    "114685": "[quote=inversion;114673]\r\n\r\n[quote=Wendy Kan;114670]\r\n\r\nHand annotation of training images is fine . . . \r\n\r\n[/quote]\r\n\r\nCan we annotate other body parts too, like the head?\r\n\r\n[/quote]\r\n\r\nYes.",
    "114696": "Can I use the OpenCV human detector for training?\r\nIs this pre-training model?",
    "114713": "Wendy, Thanks for the reply. I like your analogy between human brain to pre-trained model :).\r\n\r\nSince questions about hand-labelling comes up in almost every computer vision competition, may I suggest adding an FAQ about hand annotation to the Kaggle wiki?\r\n\r\nYou and other admins have made various clarifications about what's allowed but they are all spread out in different forums. Having them in one place will certainly help!",
    "114721": "[quote=tereka;114696]\r\n\r\nCan I use the OpenCV human detector for training?\r\nIs this pre-training model?\r\n\r\n[/quote]\r\n\r\nYes that is a pre-trained model that used external data - pretty much all off-the-shelf object detectors are, except the really simple line, circle, connected region filters. \r\n\r\nSo it looks like you could use the human detector to help manage the training data, but it should not be run on the test images as part of prediction code.",
    "114731": "Neil. Thanks for your reply.\r\nI only use to help putting the training label.",
    "120204": "[quote=Wendy Kan;114685]\r\n\r\n[quote=inversion;114673]\r\n\r\n[quote=Wendy Kan;114670]\r\n\r\nHand annotation of training images is fine . . . \r\n\r\n[/quote]\r\n\r\nCan we annotate other body parts too, like the head?\r\n\r\n[/quote]\r\n\r\nYes. \r\n\r\n[/quote]\r\nThen, should we share the annotations? and where will be found them?",
    "120214": "[quote=long long;120204]\r\n\r\nThen, should we share the annotations? and where will be found them?\r\n\r\n[/quote]\r\n\r\nYou are under no obligation to share your own personal annotations. (But, you certainly may if you want.) Some may choose to crowd source annotations. If that is the case, they must be shared to the forum, since crowd sourcing annotations is sharing across teams. (Thus, it cannot be done privately.)"
  },
  "source": "meta"
}