{
  "id": 20162,
  "title": "Mislabeled training images",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20162",
  "author_name": "",
  "post_date": "2016-04-16T02:24:27.703Z",
  "votes": 3,
  "comment_count": 10,
  "views": 1838,
  "content": "<p>Can anything be done with mislabeled images? Here are a few random examples from &quot;c0: normal driving&quot; category, which normally requires both hands on the steering wheel and forward gaze:</p>\n\n<ul>\n<li>img_3600.jpg: only one hand on the wheel, the other on soda bottle\nand sideways gaze </li>\n<li>img_2151.jpg: both hands off the wheel</li>\n<li>img_6002.jpg: talking to passenger, head turned 90 degrees</li>\n<li>img_12470.jpg: right hand off the wheel, holding the cell phone</li>\n</ul>\n\n<p>The solutions will have higher value for the sponsor, if some of this noise is removed. </p>\n\n<p>Since it is very early in the contest, I don't think it will hurt anyone to do it now. There could be a short deadline established for reporting errors in training data. Organizers would have a final word in accepting or rejecting proposed corrections.</p>",
  "messages": [
    {
      "id": "115072",
      "postDate": "04/16/2016 02:24:27",
      "content": "<p>Can anything be done with mislabeled images? Here are a few random examples from &quot;c0: normal driving&quot; category, which normally requires both hands on the steering wheel and forward gaze:</p>\n\n<ul>\n<li>img_3600.jpg: only one hand on the wheel, the other on soda bottle\nand sideways gaze </li>\n<li>img_2151.jpg: both hands off the wheel</li>\n<li>img_6002.jpg: talking to passenger, head turned 90 degrees</li>\n<li>img_12470.jpg: right hand off the wheel, holding the cell phone</li>\n</ul>\n\n<p>The solutions will have higher value for the sponsor, if some of this noise is removed. </p>\n\n<p>Since it is very early in the contest, I don't think it will hurt anyone to do it now. There could be a short deadline established for reporting errors in training data. Organizers would have a final word in accepting or rejecting proposed corrections.</p>",
      "rawMarkdown": "Can anything be done with mislabeled images? Here are a few random examples from \"c0: normal driving\" category, which normally requires both hands on the steering wheel and forward gaze:\r\n\r\n - img_3600.jpg: only one hand on the wheel, the other on soda bottle\r\n   and sideways gaze \r\n - img_2151.jpg: both hands off the wheel\r\n - img_6002.jpg: talking to passenger, head turned 90 degrees\r\n - img_12470.jpg: right hand off the wheel, holding the cell phone\r\n\r\nThe solutions will have higher value for the sponsor, if some of this noise is removed. \r\n\r\nSince it is very early in the contest, I don't think it will hurt anyone to do it now. There could be a short deadline established for reporting errors in training data. Organizers would have a final word in accepting or rejecting proposed corrections.",
      "votes": null
    },
    {
      "id": "115079",
      "postDate": "04/16/2016 04:11:16",
      "content": "<p>This is a very good point, there are quite few of those. What would be also very valuable is to provide definitions what constitutes different activities (like &quot;<em>normal driving requires both hands on the steering wheel and forward gaze</em>&quot; above). Some are quite clear, like talking on the phone, but others are much more vague. This will not only help to weed out the errors in the train data, but possibly to help us to devise some targeted tests to identify them.</p>\n\n<p>Furthermore, it would be nice if sponsor checked how test images were categorized as well ;) Since the split between training, private and public LB was probably  random (we hope :) ) I would expect some mis-categorization in the test set as well. Unfortunately while we can check training images, we can not check test ones ...</p>",
      "rawMarkdown": "This is a very good point, there are quite few of those. What would be also very valuable is to provide definitions what constitutes different activities (like \"*normal driving requires both hands on the steering wheel and forward gaze*\" above). Some are quite clear, like talking on the phone, but others are much more vague. This will not only help to weed out the errors in the train data, but possibly to help us to devise some targeted tests to identify them.\r\n\r\nFurthermore, it would be nice if sponsor checked how test images were categorized as well ;) Since the split between training, private and public LB was probably  random (we hope :) ) I would expect some mis-categorization in the test set as well. Unfortunately while we can check training images, we can not check test ones ...",
      "votes": null
    },
    {
      "id": "115080",
      "postDate": "04/16/2016 04:14:32",
      "content": "<p>Actually, it is not clear to me how to categorize <strong>img_2151</strong> at all. It is clearly <strong>not</strong> a normal driving with both hands off the wheel, but at the same time it is not texting, nor talking on the phone, nor any other categorized activity.</p>",
      "rawMarkdown": "Actually, it is not clear to me how to categorize **img_2151** at all. It is clearly **not** a normal driving with both hands off the wheel, but at the same time it is not texting, nor talking on the phone, nor any other categorized activity.",
      "votes": null
    },
    {
      "id": "115458",
      "postDate": "04/18/2016 18:47:53",
      "content": "<p>It would be extremely helpful if the sponsor published the guidelines that were used to classify the images. The problem, as mentioned earlier in this thread, is that there's a lot of ambiguity even for humans to classify these images. Take img_3600, for instance: although the driver's head is turned to the right and she is only driving with one hand on the steering wheel and her index finger is on a water bottle, if you look closely you can see that her eyes are looking at the road. How should that be classified? </p>",
      "rawMarkdown": "It would be extremely helpful if the sponsor published the guidelines that were used to classify the images. The problem, as mentioned earlier in this thread, is that there's a lot of ambiguity even for humans to classify these images. Take img_3600, for instance: although the driver's head is turned to the right and she is only driving with one hand on the steering wheel and her index finger is on a water bottle, if you look closely you can see that her eyes are looking at the road. How should that be classified?",
      "votes": null
    },
    {
      "id": "115702",
      "postDate": "04/19/2016 17:18:24",
      "content": "<p>I would like to chime in and comment on your concerns. These are all excellent points and your concerns are shared by State Farm! Trying to identify distracted driving itself is a hard problem. But trying to formal define what is &#8220;normal&#8221; or &#8220;safe&#8221; driving is even harder. </p>\n\n<p>I want to remind competitors that we intend for this dataset to reflect real world data collection. For real world problems, noise in the data is to be expected and should be addressed accordingly. Throwing out the noisy data may be the easiest route, and you could expect a model built on a nice, clean dataset would perform with highest accuracy. However, when retraining/reevaluating the model with new real world data, the labels are going to be noisy again. State Farm is looking for models that perform well on the real-world noisy data. Please take a look at <a>this link</a> for more information on data quality for Kaggle competition datasets.</p>\n\n<p>Here is some background in the experiment. A number of volunteers were asked to perform specific driving actions while being recorded. The drivers were given some basic instructions on the action to take. Each action had a timeframe assigned throughout the video stream. Frames from that video at incremental time-steps were extracted to be used as a classification problem. The process was automated, at the process consisted of tens of thousands of images. It is expected that some frames could contain &#8220;misclassified&#8221; images as the driver could be in a transition state for the action.  </p>\n\n<p>The actions selected for this experiment have (for the most part) distinct features that can be identified, e.g. hand holding phone and driver looking down, or arm reached behind. Some features will have some classification crossover that makes this problem even more difficult. </p>\n\n<p>The most difficult action was &#8220;normal&#8221; driving, because it is so ambiguous. Plus, certain features of the driver could easily be confused with other classes, such as looking in blind spot (reaching behind or talking to passenger), checking rearview mirror (grooming), etc. The guidelines for &#8220;normal&#8221; driving were not strict other than simply &#8220;Do not do any of the other actions.&#8221;. It was recommended that they keep both hands on the wheel, but whether it be subconscious or on purpose, the volunteers may have bent the guidelines slightly, but not to the point that they explicitly performed other &#8220;distracted driving&#8221; actions. </p>\n\n<p>The second difficult one was &#8220;talking to passenger&#8221;. Drivers tend to move their head from looking straight ahead to looking 90 degrees right to talk to the passenger. Although the head orientation is a feature that could be used in classifying these images, there are additional features that should be considered as well. For example, mouth movement, hand gestures, facial expressions, etc. </p>\n\n<p>Please feel free to interpret the training data in any way you see fit. If you have any further questions, I will be happy to assist.</p>\n\n<p>Thanks,\nDan, State Farm</p>",
      "rawMarkdown": "I would like to chime in and comment on your concerns. These are all excellent points and your concerns are shared by State Farm! Trying to identify distracted driving itself is a hard problem. But trying to formal define what is “normal” or “safe” driving is even harder. \r\n \r\nI want to remind competitors that we intend for this dataset to reflect real world data collection. For real world problems, noise in the data is to be expected and should be addressed accordingly. Throwing out the noisy data may be the easiest route, and you could expect a model built on a nice, clean dataset would perform with highest accuracy. However, when retraining/reevaluating the model with new real world data, the labels are going to be noisy again. State Farm is looking for models that perform well on the real-world noisy data. Please take a look at [this link][1] for more information on data quality for Kaggle competition datasets.\r\n \r\nHere is some background in the experiment. A number of volunteers were asked to perform specific driving actions while being recorded. The drivers were given some basic instructions on the action to take. Each action had a timeframe assigned throughout the video stream. Frames from that video at incremental time-steps were extracted to be used as a classification problem. The process was automated, at the process consisted of tens of thousands of images. It is expected that some frames could contain “misclassified” images as the driver could be in a transition state for the action.  \r\n \r\nThe actions selected for this experiment have (for the most part) distinct features that can be identified, e.g. hand holding phone and driver looking down, or arm reached behind. Some features will have some classification crossover that makes this problem even more difficult. \r\n \r\nThe most difficult action was “normal” driving, because it is so ambiguous. Plus, certain features of the driver could easily be confused with other classes, such as looking in blind spot (reaching behind or talking to passenger), checking rearview mirror (grooming), etc. The guidelines for “normal” driving were not strict other than simply “Do not do any of the other actions.”. It was recommended that they keep both hands on the wheel, but whether it be subconscious or on purpose, the volunteers may have bent the guidelines slightly, but not to the point that they explicitly performed other “distracted driving” actions. \r\n \r\nThe second difficult one was “talking to passenger”. Drivers tend to move their head from looking straight ahead to looking 90 degrees right to talk to the passenger. Although the head orientation is a feature that could be used in classifying these images, there are additional features that should be considered as well. For example, mouth movement, hand gestures, facial expressions, etc. \r\n \r\nPlease feel free to interpret the training data in any way you see fit. If you have any further questions, I will be happy to assist.\r\n \r\nThanks,\r\nDan, State Farm\r\n\r\n\r\n  [1]: http://%20https://www.kaggle.com/wiki/ANoteOnDataQuality",
      "votes": null
    },
    {
      "id": "115710",
      "postDate": "04/19/2016 17:26:33",
      "content": "<p>[quote=Dan Holman;115702]</p>\n\n<p>Some features will have some classification crossover that makes this problem even more difficult.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks Dan. Great clarification.</p>\n\n<p>My favorite example of an ambiguous image so far is 75152, where I can't tell if the driver is holding an imaginary cell phone, or in the process of making a &quot;hand signal&quot; to another driver.  :-)</p>",
      "rawMarkdown": "[quote=Dan Holman;115702]\r\n\r\nSome features will have some classification crossover that makes this problem even more difficult.\r\n\r\n  [1]: http://%20https://www.kaggle.com/wiki/ANoteOnDataQuality\r\n\r\n[/quote]\r\n\r\nThanks Dan. Great clarification.\r\n\r\nMy favorite example of an ambiguous image so far is 75152, where I can't tell if the driver is holding an imaginary cell phone, or in the process of making a \"hand signal\" to another driver.  :-)",
      "votes": null
    },
    {
      "id": "116127",
      "postDate": "04/22/2016 06:31:55",
      "content": "<p>Thank you Dan for the explanation. It is very interesting to peek behind the curtain ;) Let me repeat my understanding of the process to make sure I understood it correctly. </p>\n\n<p>So, some parts of a video stream (let's say for example, T0 to T1 sec in the video stream) have been assigned to certain activity (let's say A0 - 'normal driving') and all images from that time slot are assigned the same activity A0? So humans did not look at the images and all labeling was done automatically based on the recorded time slot T0 to T1? Therefore, if the driver engaged in any other activity (possibly accidentally or unconsciously ) during this time slot it is still going to be categorized as  A0? For example, if during normal driving the person has asked the question of the passenger, it is still going to be categorized as &quot;normal driving&quot;. Or if during the &quot;conversation with the passenger&quot; activity the driver gazed straight ahead this frame is going to be categorized as &quot;talking to the passenger&quot;. And I assume the same process applies to both part of the test set, and for the train labels. </p>\n\n<p>Is my understanding of the process correct? </p>",
      "rawMarkdown": "Thank you Dan for the explanation. It is very interesting to peek behind the curtain ;) Let me repeat my understanding of the process to make sure I understood it correctly. \r\n\r\nSo, some parts of a video stream (let's say for example, T0 to T1 sec in the video stream) have been assigned to certain activity (let's say A0 - 'normal driving') and all images from that time slot are assigned the same activity A0? So humans did not look at the images and all labeling was done automatically based on the recorded time slot T0 to T1? Therefore, if the driver engaged in any other activity (possibly accidentally or unconsciously ) during this time slot it is still going to be categorized as  A0? For example, if during normal driving the person has asked the question of the passenger, it is still going to be categorized as \"normal driving\". Or if during the \"conversation with the passenger\" activity the driver gazed straight ahead this frame is going to be categorized as \"talking to the passenger\". And I assume the same process applies to both part of the test set, and for the train labels. \r\n\r\nIs my understanding of the process correct?",
      "votes": null
    },
    {
      "id": "116238",
      "postDate": "04/22/2016 16:26:51",
      "content": "<p>Yes, this is correct. Human labeling is difficult, and very time consuming. Plus, there is some degree of subjectivity in both human labeling of images and what a driver would consider &quot;normal driving&quot;. Automating the data collection process does introduce noise to the data set, but it's also what makes this problem so challenging and fun!</p>",
      "rawMarkdown": "Yes, this is correct. Human labeling is difficult, and very time consuming. Plus, there is some degree of subjectivity in both human labeling of images and what a driver would consider \"normal driving\". Automating the data collection process does introduce noise to the data set, but it's also what makes this problem so challenging and fun!",
      "votes": null
    },
    {
      "id": "116244",
      "postDate": "04/22/2016 17:16:27",
      "content": "<p>Am I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.</p>\n\n<p>Since the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.</p>",
      "rawMarkdown": "Am I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.\r\n\r\nSince the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.",
      "votes": null
    },
    {
      "id": "116245",
      "postDate": "04/22/2016 17:23:59",
      "content": "<blockquote>\n  <p>&quot;having one hand off the wheel is totally normal, unless of course, you\n  have 3 hands&quot;</p>\n</blockquote>\n\n<p>In which case it is totally normal to have two hands off the wheel ;)</p>",
      "rawMarkdown": "> \"having one hand off the wheel is totally normal, unless of course, you\r\n> have 3 hands\"\r\n\r\nIn which case it is totally normal to have two hands off the wheel ;)",
      "votes": null
    },
    {
      "id": "116330",
      "postDate": "04/23/2016 09:53:11",
      "content": "<p>[quote=NxGTR;116244]</p>\n\n<p>Am I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.</p>\n\n<p>Since the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, but then your hand is typically placed on the knob, not holding soda/phone. </p>",
      "rawMarkdown": "[quote=NxGTR;116244]\r\n\r\nAm I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.\r\n\r\nSince the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.\r\n\r\n[/quote]\r\n\r\nYes, but then your hand is typically placed on the knob, not holding soda/phone.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115079,
      "author_name": "alexlzzz",
      "author_url": "",
      "post_date": "04/16/2016 04:11:16",
      "content": "<p>This is a very good point, there are quite few of those. What would be also very valuable is to provide definitions what constitutes different activities (like &quot;<em>normal driving requires both hands on the steering wheel and forward gaze</em>&quot; above). Some are quite clear, like talking on the phone, but others are much more vague. This will not only help to weed out the errors in the train data, but possibly to help us to devise some targeted tests to identify them.</p>\n\n<p>Furthermore, it would be nice if sponsor checked how test images were categorized as well ;) Since the split between training, private and public LB was probably  random (we hope :) ) I would expect some mis-categorization in the test set as well. Unfortunately while we can check training images, we can not check test ones ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115080,
      "author_name": "alexlzzz",
      "author_url": "",
      "post_date": "04/16/2016 04:14:32",
      "content": "<p>Actually, it is not clear to me how to categorize <strong>img_2151</strong> at all. It is clearly <strong>not</strong> a normal driving with both hands off the wheel, but at the same time it is not texting, nor talking on the phone, nor any other categorized activity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115458,
      "author_name": "spammy",
      "author_url": "",
      "post_date": "04/18/2016 18:47:53",
      "content": "<p>It would be extremely helpful if the sponsor published the guidelines that were used to classify the images. The problem, as mentioned earlier in this thread, is that there's a lot of ambiguity even for humans to classify these images. Take img_3600, for instance: although the driver's head is turned to the right and she is only driving with one hand on the steering wheel and her index finger is on a water bottle, if you look closely you can see that her eyes are looking at the road. How should that be classified? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115702,
      "author_name": "danholman",
      "author_url": "",
      "post_date": "04/19/2016 17:18:24",
      "content": "<p>I would like to chime in and comment on your concerns. These are all excellent points and your concerns are shared by State Farm! Trying to identify distracted driving itself is a hard problem. But trying to formal define what is &#8220;normal&#8221; or &#8220;safe&#8221; driving is even harder. </p>\n\n<p>I want to remind competitors that we intend for this dataset to reflect real world data collection. For real world problems, noise in the data is to be expected and should be addressed accordingly. Throwing out the noisy data may be the easiest route, and you could expect a model built on a nice, clean dataset would perform with highest accuracy. However, when retraining/reevaluating the model with new real world data, the labels are going to be noisy again. State Farm is looking for models that perform well on the real-world noisy data. Please take a look at <a>this link</a> for more information on data quality for Kaggle competition datasets.</p>\n\n<p>Here is some background in the experiment. A number of volunteers were asked to perform specific driving actions while being recorded. The drivers were given some basic instructions on the action to take. Each action had a timeframe assigned throughout the video stream. Frames from that video at incremental time-steps were extracted to be used as a classification problem. The process was automated, at the process consisted of tens of thousands of images. It is expected that some frames could contain &#8220;misclassified&#8221; images as the driver could be in a transition state for the action.  </p>\n\n<p>The actions selected for this experiment have (for the most part) distinct features that can be identified, e.g. hand holding phone and driver looking down, or arm reached behind. Some features will have some classification crossover that makes this problem even more difficult. </p>\n\n<p>The most difficult action was &#8220;normal&#8221; driving, because it is so ambiguous. Plus, certain features of the driver could easily be confused with other classes, such as looking in blind spot (reaching behind or talking to passenger), checking rearview mirror (grooming), etc. The guidelines for &#8220;normal&#8221; driving were not strict other than simply &#8220;Do not do any of the other actions.&#8221;. It was recommended that they keep both hands on the wheel, but whether it be subconscious or on purpose, the volunteers may have bent the guidelines slightly, but not to the point that they explicitly performed other &#8220;distracted driving&#8221; actions. </p>\n\n<p>The second difficult one was &#8220;talking to passenger&#8221;. Drivers tend to move their head from looking straight ahead to looking 90 degrees right to talk to the passenger. Although the head orientation is a feature that could be used in classifying these images, there are additional features that should be considered as well. For example, mouth movement, hand gestures, facial expressions, etc. </p>\n\n<p>Please feel free to interpret the training data in any way you see fit. If you have any further questions, I will be happy to assist.</p>\n\n<p>Thanks,\nDan, State Farm</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115710,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/19/2016 17:26:33",
      "content": "<p>[quote=Dan Holman;115702]</p>\n\n<p>Some features will have some classification crossover that makes this problem even more difficult.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks Dan. Great clarification.</p>\n\n<p>My favorite example of an ambiguous image so far is 75152, where I can't tell if the driver is holding an imaginary cell phone, or in the process of making a &quot;hand signal&quot; to another driver.  :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116127,
      "author_name": "alexlzzz",
      "author_url": "",
      "post_date": "04/22/2016 06:31:55",
      "content": "<p>Thank you Dan for the explanation. It is very interesting to peek behind the curtain ;) Let me repeat my understanding of the process to make sure I understood it correctly. </p>\n\n<p>So, some parts of a video stream (let's say for example, T0 to T1 sec in the video stream) have been assigned to certain activity (let's say A0 - 'normal driving') and all images from that time slot are assigned the same activity A0? So humans did not look at the images and all labeling was done automatically based on the recorded time slot T0 to T1? Therefore, if the driver engaged in any other activity (possibly accidentally or unconsciously ) during this time slot it is still going to be categorized as  A0? For example, if during normal driving the person has asked the question of the passenger, it is still going to be categorized as &quot;normal driving&quot;. Or if during the &quot;conversation with the passenger&quot; activity the driver gazed straight ahead this frame is going to be categorized as &quot;talking to the passenger&quot;. And I assume the same process applies to both part of the test set, and for the train labels. </p>\n\n<p>Is my understanding of the process correct? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116238,
      "author_name": "danholman",
      "author_url": "",
      "post_date": "04/22/2016 16:26:51",
      "content": "<p>Yes, this is correct. Human labeling is difficult, and very time consuming. Plus, there is some degree of subjectivity in both human labeling of images and what a driver would consider &quot;normal driving&quot;. Automating the data collection process does introduce noise to the data set, but it's also what makes this problem so challenging and fun!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116244,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "04/22/2016 17:16:27",
      "content": "<p>Am I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.</p>\n\n<p>Since the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116245,
      "author_name": "alexlzzz",
      "author_url": "",
      "post_date": "04/22/2016 17:23:59",
      "content": "<blockquote>\n  <p>&quot;having one hand off the wheel is totally normal, unless of course, you\n  have 3 hands&quot;</p>\n</blockquote>\n\n<p>In which case it is totally normal to have two hands off the wheel ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116330,
      "author_name": "inoryy",
      "author_url": "",
      "post_date": "04/23/2016 09:53:11",
      "content": "<p>[quote=NxGTR;116244]</p>\n\n<p>Am I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.</p>\n\n<p>Since the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, but then your hand is typically placed on the knob, not holding soda/phone. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115072": "Can anything be done with mislabeled images? Here are a few random examples from \"c0: normal driving\" category, which normally requires both hands on the steering wheel and forward gaze:\r\n\r\n - img_3600.jpg: only one hand on the wheel, the other on soda bottle\r\n   and sideways gaze \r\n - img_2151.jpg: both hands off the wheel\r\n - img_6002.jpg: talking to passenger, head turned 90 degrees\r\n - img_12470.jpg: right hand off the wheel, holding the cell phone\r\n\r\nThe solutions will have higher value for the sponsor, if some of this noise is removed. \r\n\r\nSince it is very early in the contest, I don't think it will hurt anyone to do it now. There could be a short deadline established for reporting errors in training data. Organizers would have a final word in accepting or rejecting proposed corrections.",
    "115079": "This is a very good point, there are quite few of those. What would be also very valuable is to provide definitions what constitutes different activities (like \"*normal driving requires both hands on the steering wheel and forward gaze*\" above). Some are quite clear, like talking on the phone, but others are much more vague. This will not only help to weed out the errors in the train data, but possibly to help us to devise some targeted tests to identify them.\r\n\r\nFurthermore, it would be nice if sponsor checked how test images were categorized as well ;) Since the split between training, private and public LB was probably  random (we hope :) ) I would expect some mis-categorization in the test set as well. Unfortunately while we can check training images, we can not check test ones ...",
    "115080": "Actually, it is not clear to me how to categorize **img_2151** at all. It is clearly **not** a normal driving with both hands off the wheel, but at the same time it is not texting, nor talking on the phone, nor any other categorized activity.",
    "115458": "It would be extremely helpful if the sponsor published the guidelines that were used to classify the images. The problem, as mentioned earlier in this thread, is that there's a lot of ambiguity even for humans to classify these images. Take img_3600, for instance: although the driver's head is turned to the right and she is only driving with one hand on the steering wheel and her index finger is on a water bottle, if you look closely you can see that her eyes are looking at the road. How should that be classified?",
    "115702": "I would like to chime in and comment on your concerns. These are all excellent points and your concerns are shared by State Farm! Trying to identify distracted driving itself is a hard problem. But trying to formal define what is “normal” or “safe” driving is even harder. \r\n \r\nI want to remind competitors that we intend for this dataset to reflect real world data collection. For real world problems, noise in the data is to be expected and should be addressed accordingly. Throwing out the noisy data may be the easiest route, and you could expect a model built on a nice, clean dataset would perform with highest accuracy. However, when retraining/reevaluating the model with new real world data, the labels are going to be noisy again. State Farm is looking for models that perform well on the real-world noisy data. Please take a look at [this link][1] for more information on data quality for Kaggle competition datasets.\r\n \r\nHere is some background in the experiment. A number of volunteers were asked to perform specific driving actions while being recorded. The drivers were given some basic instructions on the action to take. Each action had a timeframe assigned throughout the video stream. Frames from that video at incremental time-steps were extracted to be used as a classification problem. The process was automated, at the process consisted of tens of thousands of images. It is expected that some frames could contain “misclassified” images as the driver could be in a transition state for the action.  \r\n \r\nThe actions selected for this experiment have (for the most part) distinct features that can be identified, e.g. hand holding phone and driver looking down, or arm reached behind. Some features will have some classification crossover that makes this problem even more difficult. \r\n \r\nThe most difficult action was “normal” driving, because it is so ambiguous. Plus, certain features of the driver could easily be confused with other classes, such as looking in blind spot (reaching behind or talking to passenger), checking rearview mirror (grooming), etc. The guidelines for “normal” driving were not strict other than simply “Do not do any of the other actions.”. It was recommended that they keep both hands on the wheel, but whether it be subconscious or on purpose, the volunteers may have bent the guidelines slightly, but not to the point that they explicitly performed other “distracted driving” actions. \r\n \r\nThe second difficult one was “talking to passenger”. Drivers tend to move their head from looking straight ahead to looking 90 degrees right to talk to the passenger. Although the head orientation is a feature that could be used in classifying these images, there are additional features that should be considered as well. For example, mouth movement, hand gestures, facial expressions, etc. \r\n \r\nPlease feel free to interpret the training data in any way you see fit. If you have any further questions, I will be happy to assist.\r\n \r\nThanks,\r\nDan, State Farm\r\n\r\n\r\n  [1]: http://%20https://www.kaggle.com/wiki/ANoteOnDataQuality",
    "115710": "[quote=Dan Holman;115702]\r\n\r\nSome features will have some classification crossover that makes this problem even more difficult.\r\n\r\n  [1]: http://%20https://www.kaggle.com/wiki/ANoteOnDataQuality\r\n\r\n[/quote]\r\n\r\nThanks Dan. Great clarification.\r\n\r\nMy favorite example of an ambiguous image so far is 75152, where I can't tell if the driver is holding an imaginary cell phone, or in the process of making a \"hand signal\" to another driver.  :-)",
    "116127": "Thank you Dan for the explanation. It is very interesting to peek behind the curtain ;) Let me repeat my understanding of the process to make sure I understood it correctly. \r\n\r\nSo, some parts of a video stream (let's say for example, T0 to T1 sec in the video stream) have been assigned to certain activity (let's say A0 - 'normal driving') and all images from that time slot are assigned the same activity A0? So humans did not look at the images and all labeling was done automatically based on the recorded time slot T0 to T1? Therefore, if the driver engaged in any other activity (possibly accidentally or unconsciously ) during this time slot it is still going to be categorized as  A0? For example, if during normal driving the person has asked the question of the passenger, it is still going to be categorized as \"normal driving\". Or if during the \"conversation with the passenger\" activity the driver gazed straight ahead this frame is going to be categorized as \"talking to the passenger\". And I assume the same process applies to both part of the test set, and for the train labels. \r\n\r\nIs my understanding of the process correct?",
    "116238": "Yes, this is correct. Human labeling is difficult, and very time consuming. Plus, there is some degree of subjectivity in both human labeling of images and what a driver would consider \"normal driving\". Automating the data collection process does introduce noise to the data set, but it's also what makes this problem so challenging and fun!",
    "116244": "Am I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.\r\n\r\nSince the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.",
    "116245": "> \"having one hand off the wheel is totally normal, unless of course, you\r\n> have 3 hands\"\r\n\r\nIn which case it is totally normal to have two hands off the wheel ;)",
    "116330": "[quote=NxGTR;116244]\r\n\r\nAm I the only one here that drives a manual car?, having one hand off the wheel is totally normal, unless of course, you have 3 hands.\r\n\r\nSince the pictures are fake (not real driving), car transmission and hand position is something to consider for real life performance.\r\n\r\n[/quote]\r\n\r\nYes, but then your hand is typically placed on the knob, not holding soda/phone."
  },
  "source": "meta"
}