{
  "id": 22631,
  "title": "#3 BR POWER - Solution",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22631",
  "author_name": "Giba",
  "post_date": "2016-08-02T13:09:51.550000",
  "votes": 98,
  "comment_count": 33,
  "views": 8895,
  "content": "<p>First of all congratulations to Jacobkie and Z_B_C for winning such amazing competition. @Z_B_C we almost draw.\nAlso I would like to thanks Kaggle and StateFarm for such a unique competition. We leaned a lot and it is my first DeepNets competition win ;-D</p>\n\n<p>Our solution is most based in the &quot;Time&quot; feature. Just watch the movies in the link and try to figure out why: <a href=\"https://www.kaggle.com/titericz/state-farm-distracted-driver-detection/just-relax-and-watch-some-cool-movies\">https://www.kaggle.com/titericz/state-farm-distracted-driver-detection/just-relax-and-watch-some-cool-movies</a></p>\n\n<p>All pictures in trainset are taken in sequence so we explored that characteristic. Trainset is given in the correct sequence, but Testset not. So taking into account that subsequent images are very close to each one (driver position, light, shadows, vehicle external objects, etc...), we decided to try nearest neighbors on all images. And for our surprise it presented very good results catching the nearest images in the correct trainset sequence. The first neighbor have a subject hit ratio of 100% and class hit rate of 99,5% in trainset. So we used a blend of the 20 nearest neighbors in our solution.</p>\n\n<p>We trained 8 CNN models, but for our final submission we used only 4. All models trained over 5 folds CV:</p>\n\n<p>1)- resnet-152 caffe, CV: 0.31 LB: 0.27. Original dataset, no augmentation</p>\n\n<p>2)- resnet-152 caffe, CV: 0.36 LB: 0.31. Original dataset, some augmentation (under-tuned?)</p>\n\n<p>3)- resnet-152 torch, CV: 0.223 LB: 0.181. Modified images 1, no augmentation</p>\n\n<p>4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation</p>\n\n<p>=&gt;Modified Images 1: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:</p>\n\n<p>Channel R:  (current image + nearest1)/2</p>\n\n<p>Channel G:  (nearest2 + nearest3)/2</p>\n\n<p>Channel B:  (nearest4 + nearest5)/2</p>\n\n<p>It built very redundant images. Also on these images we tried to center the steering wheel via regression.</p>\n\n<p>=&gt;Modified Images 2: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:</p>\n\n<p>Channel R:  current image - nearest1</p>\n\n<p>Channel G: nearest2 - nearest3</p>\n\n<p>Channel B:  nearest4 - nearest5</p>\n\n<p>It build very redundant images and tries to catch drivers movements. </p>\n\n<p>So these 4 models presented high diversity between then and predictions are very stable at Level 1 training.</p>\n\n<p>So we used Level 1 prediction of all 4 models to ensemble at Level 2. But this time merging all 20 neighbors predictions of each image. So each image have its own prediction + 20 neighbors predictions. For Level 2 training we used scipy minimize function and created a custom geometric average function to minimize logloss of all models. That architecture improved CV of each model:</p>\n\n<p>1)- resnet-152 caffe, CV: 0.31 =&gt;  0.192</p>\n\n<p>2)- resnet-152 caffe, CV: 0.36 =&gt; 0.276</p>\n\n<p>3)- resnet-152 torch, CV: 0.22 =&gt; 0.180</p>\n\n<p>4)- VGG-16 Keras, CV: 0.30 =&gt; 0.192</p>\n\n<p>Our final solution is and weighted geometric average of these 4 models and CV score is about 0.116 and Class hit Ratio of 96.4%</p>\n\n<p>Also we found via cross-validation that replacing all prediction &lt; 0.00001 to zero improved our scores in CV, but we didn't used that in our last submission. If we had choosen it we would finished #1  :-/</p>\n\n<p>That's it and...</p>\n\n<p>Congratulation again Jacobkie for its late huge jump to #1. And congrats to all top10 teams that didn't overfitted and to everyone that spent much time in this one. Lots of learnings to all.</p>\n\n<p>Thanks again!\nGiba</p>\n\n<p>obs. don't forget to upvote the post and script if you like it ;-P </p>",
  "messages": [
    {
      "id": 129818,
      "postDate": "2016-08-02T13:09:51.550Z",
      "content": "<p>First of all congratulations to Jacobkie and Z_B_C for winning such amazing competition. @Z_B_C we almost draw.\nAlso I would like to thanks Kaggle and StateFarm for such a unique competition. We leaned a lot and it is my first DeepNets competition win ;-D</p>\n\n<p>Our solution is most based in the &quot;Time&quot; feature. Just watch the movies in the link and try to figure out why: <a href=\"https://www.kaggle.com/titericz/state-farm-distracted-driver-detection/just-relax-and-watch-some-cool-movies\">https://www.kaggle.com/titericz/state-farm-distracted-driver-detection/just-relax-and-watch-some-cool-movies</a></p>\n\n<p>All pictures in trainset are taken in sequence so we explored that characteristic. Trainset is given in the correct sequence, but Testset not. So taking into account that subsequent images are very close to each one (driver position, light, shadows, vehicle external objects, etc...), we decided to try nearest neighbors on all images. And for our surprise it presented very good results catching the nearest images in the correct trainset sequence. The first neighbor have a subject hit ratio of 100% and class hit rate of 99,5% in trainset. So we used a blend of the 20 nearest neighbors in our solution.</p>\n\n<p>We trained 8 CNN models, but for our final submission we used only 4. All models trained over 5 folds CV:</p>\n\n<p>1)- resnet-152 caffe, CV: 0.31 LB: 0.27. Original dataset, no augmentation</p>\n\n<p>2)- resnet-152 caffe, CV: 0.36 LB: 0.31. Original dataset, some augmentation (under-tuned?)</p>\n\n<p>3)- resnet-152 torch, CV: 0.223 LB: 0.181. Modified images 1, no augmentation</p>\n\n<p>4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation</p>\n\n<p>=&gt;Modified Images 1: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:</p>\n\n<p>Channel R:  (current image + nearest1)/2</p>\n\n<p>Channel G:  (nearest2 + nearest3)/2</p>\n\n<p>Channel B:  (nearest4 + nearest5)/2</p>\n\n<p>It built very redundant images. Also on these images we tried to center the steering wheel via regression.</p>\n\n<p>=&gt;Modified Images 2: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:</p>\n\n<p>Channel R:  current image - nearest1</p>\n\n<p>Channel G: nearest2 - nearest3</p>\n\n<p>Channel B:  nearest4 - nearest5</p>\n\n<p>It build very redundant images and tries to catch drivers movements. </p>\n\n<p>So these 4 models presented high diversity between then and predictions are very stable at Level 1 training.</p>\n\n<p>So we used Level 1 prediction of all 4 models to ensemble at Level 2. But this time merging all 20 neighbors predictions of each image. So each image have its own prediction + 20 neighbors predictions. For Level 2 training we used scipy minimize function and created a custom geometric average function to minimize logloss of all models. That architecture improved CV of each model:</p>\n\n<p>1)- resnet-152 caffe, CV: 0.31 =&gt;  0.192</p>\n\n<p>2)- resnet-152 caffe, CV: 0.36 =&gt; 0.276</p>\n\n<p>3)- resnet-152 torch, CV: 0.22 =&gt; 0.180</p>\n\n<p>4)- VGG-16 Keras, CV: 0.30 =&gt; 0.192</p>\n\n<p>Our final solution is and weighted geometric average of these 4 models and CV score is about 0.116 and Class hit Ratio of 96.4%</p>\n\n<p>Also we found via cross-validation that replacing all prediction &lt; 0.00001 to zero improved our scores in CV, but we didn't used that in our last submission. If we had choosen it we would finished #1  :-/</p>\n\n<p>That's it and...</p>\n\n<p>Congratulation again Jacobkie for its late huge jump to #1. And congrats to all top10 teams that didn't overfitted and to everyone that spent much time in this one. Lots of learnings to all.</p>\n\n<p>Thanks again!\nGiba</p>\n\n<p>obs. don't forget to upvote the post and script if you like it ;-P </p>",
      "rawMarkdown": "First of all congratulations to Jacobkie and Z_B_C for winning such amazing competition. @Z_B_C we almost draw.\r\nAlso I would like to thanks Kaggle and StateFarm for such a unique competition. We leaned a lot and it is my first DeepNets competition win ;-D\r\n\r\nOur solution is most based in the \"Time\" feature. Just watch the movies in the link and try to figure out why: https://www.kaggle.com/titericz/state-farm-distracted-driver-detection/just-relax-and-watch-some-cool-movies\r\n\r\nAll pictures in trainset are taken in sequence so we explored that characteristic. Trainset is given in the correct sequence, but Testset not. So taking into account that subsequent images are very close to each one (driver position, light, shadows, vehicle external objects, etc...), we decided to try nearest neighbors on all images. And for our surprise it presented very good results catching the nearest images in the correct trainset sequence. The first neighbor have a subject hit ratio of 100% and class hit rate of 99,5% in trainset. So we used a blend of the 20 nearest neighbors in our solution.\r\n\r\nWe trained 8 CNN models, but for our final submission we used only 4. All models trained over 5 folds CV:\r\n\r\n1)- resnet-152 caffe, CV: 0.31 LB: 0.27. Original dataset, no augmentation\r\n\r\n2)- resnet-152 caffe, CV: 0.36 LB: 0.31. Original dataset, some augmentation (under-tuned?)\r\n\r\n3)- resnet-152 torch, CV: 0.223 LB: 0.181. Modified images 1, no augmentation\r\n\r\n4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation\r\n\r\n=>Modified Images 1: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:\r\n\r\nChannel R:  (current image + nearest1)/2\r\n\r\nChannel G:  (nearest2 + nearest3)/2\r\n\r\nChannel B:  (nearest4 + nearest5)/2\r\n\r\nIt built very redundant images. Also on these images we tried to center the steering wheel via regression.\r\n\r\n\r\n=>Modified Images 2: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:\r\n\r\nChannel R:  current image - nearest1\r\n\r\nChannel G: nearest2 - nearest3\r\n\r\nChannel B:  nearest4 - nearest5\r\n\r\nIt build very redundant images and tries to catch drivers movements. \r\n\r\nSo these 4 models presented high diversity between then and predictions are very stable at Level 1 training.\r\n\r\n\r\nSo we used Level 1 prediction of all 4 models to ensemble at Level 2. But this time merging all 20 neighbors predictions of each image. So each image have its own prediction + 20 neighbors predictions. For Level 2 training we used scipy minimize function and created a custom geometric average function to minimize logloss of all models. That architecture improved CV of each model:\r\n\r\n1)- resnet-152 caffe, CV: 0.31 =>  0.192\r\n\r\n2)- resnet-152 caffe, CV: 0.36 => 0.276\r\n\r\n3)- resnet-152 torch, CV: 0.22 => 0.180\r\n\r\n4)- VGG-16 Keras, CV: 0.30 => 0.192\r\n\r\n\r\nOur final solution is and weighted geometric average of these 4 models and CV score is about 0.116 and Class hit Ratio of 96.4%\r\n\r\nAlso we found via cross-validation that replacing all prediction < 0.00001 to zero improved our scores in CV, but we didn't used that in our last submission. If we had choosen it we would finished #1  :-/\r\n\r\n\r\nThat's it and...\r\n\r\nCongratulation again Jacobkie for its late huge jump to #1. And congrats to all top10 teams that didn't overfitted and to everyone that spent much time in this one. Lots of learnings to all.\r\n\r\nThanks again!\r\nGiba\r\n\r\nobs. don't forget to upvote the post and script if you like it ;-P ",
      "votes": 98
    },
    {
      "id": 129851,
      "postDate": "2016-08-02T15:07:00.273Z",
      "content": "<p>Congratulations! This is very interesting work. Could you provide short answers to some of my questions? Appreciate!</p>\n\n<ul>\n<li>When building Modified Images 1 and Modified Images 2, what is idea behind encoding image sequence information into different color channels? Is it only for compatibility with the pre-trained models?</li>\n<li>In these scenarios, you are predicting on short sequences instead of single images. Do you think a short sequence provides a more stable prediction?</li>\n<li>When you center the steeling wheel, did you manually cropped the training data, and built a bounding-box regressor?</li>\n<li>For Level 2 training, you mentioned &quot;each image have its own prediction + 20 neighbors predictions&quot;. Did you concatenate those predictions to form a 210-D (1x10+20x10) vector for each image?</li>\n<li>When performing Level 2 training, did you use the same CV splits as used in the Level 1 training, because otherwise the Level 2 training would be prune to overfitting?</li>\n</ul>",
      "rawMarkdown": "Congratulations! This is very interesting work. Could you provide short answers to some of my questions? Appreciate!\r\n\r\n - When building Modified Images 1 and Modified Images 2, what is idea behind encoding image sequence information into different color channels? Is it only for compatibility with the pre-trained models?\r\n - In these scenarios, you are predicting on short sequences instead of single images. Do you think a short sequence provides a more stable prediction?\r\n - When you center the steeling wheel, did you manually cropped the training data, and built a bounding-box regressor?\r\n - For Level 2 training, you mentioned \"each image have its own prediction + 20 neighbors predictions\". Did you concatenate those predictions to form a 210-D (1x10+20x10) vector for each image?\r\n - When performing Level 2 training, did you use the same CV splits as used in the Level 1 training, because otherwise the Level 2 training would be prune to overfitting?\r\n\r\n\r\n\r\n",
      "votes": 8
    },
    {
      "id": 328883,
      "postDate": "2018-05-15T08:58:59.907Z",
      "content": "<p>HI, am newbie to kaggle, may i know about leader board scores? IT should be high or low? If we see in above explanation they got LB score like 0.31,0.27 etc. So am confused. When i submit my results i got 0.31 LB score. I thought my results are very bad.</p>",
      "rawMarkdown": "HI, am newbie to kaggle, may i know about leader board scores? IT should be high or low? If we see in above explanation they got LB score like 0.31,0.27 etc. So am confused. When i submit my results i got 0.31 LB score. I thought my results are very bad.",
      "votes": 1
    },
    {
      "id": 129903,
      "postDate": "2016-08-02T19:59:09.447Z",
      "content": "<p>StateFarm will unlikely implement any submitted solution exactly as it is presented here anyway.  Kaggle is essentially a fun game and it's understood that most anything goes.   As competitors, we know this.  The prizes are there mostly to encourage ferreting out the best solutions...emphasis on plural!  If they have the hardware to ensemble 100 models from video real-time, that's up to them.  For $65k, they used up several thousands of hours of our time and what sounds to be even more of expensive GPU time.  Money well spent if they want a problem worked on with success.</p>\n\n<p>Big congrats to all the best performers in this comp!  Very well played.  Also, thanks for the many cool scripts that educated many of us in some new areas.  Dark data! I totally wanted to experiment with this...thanks for confirming that it does boost!  </p>\n\n<p>To further address the &quot;not real world&quot;, etc....it's not out of the realm of possibility to hope for a competition in the future to have a different set of requirements like:</p>\n\n<p>-leakage-proofed.  Difficult as it seems there are some of you who seem to specialize in creative ways of discovering it!</p>\n\n<p>-execution speed points(or limit). (balancing hardware, ensembling, etc)</p>\n\n<p>-and, what I would like to see(just for fun), a leaderboard where we enter our own cv scores. yes, totally gameable...but that's the point.  That, and no leaderboard probing.</p>\n\n<p>Currently I'm working on a real-world problem and I simply cannot use ensemble as it would end up costing too much or taking too long.  In fact, I was forced to use googlenet instead of a slightly better performing vgg19 as well...for speed and memory.  And there's no leaderboard...I wish there was...it'd make flicking the switch this week less intimidating. As I keep telling my boss &quot;I'm <em>pretty</em> sure this will work...it worked on my test dataset!&quot;.</p>",
      "rawMarkdown": "StateFarm will unlikely implement any submitted solution exactly as it is presented here anyway.  Kaggle is essentially a fun game and it's understood that most anything goes.   As competitors, we know this.  The prizes are there mostly to encourage ferreting out the best solutions...emphasis on plural!  If they have the hardware to ensemble 100 models from video real-time, that's up to them.  For $65k, they used up several thousands of hours of our time and what sounds to be even more of expensive GPU time.  Money well spent if they want a problem worked on with success.\r\n\r\nBig congrats to all the best performers in this comp!  Very well played.  Also, thanks for the many cool scripts that educated many of us in some new areas.  Dark data! I totally wanted to experiment with this...thanks for confirming that it does boost!  \r\n\r\nTo further address the \"not real world\", etc....it's not out of the realm of possibility to hope for a competition in the future to have a different set of requirements like:\r\n\r\n-leakage-proofed.  Difficult as it seems there are some of you who seem to specialize in creative ways of discovering it!\r\n\r\n-execution speed points(or limit). (balancing hardware, ensembling, etc)\r\n\r\n-and, what I would like to see(just for fun), a leaderboard where we enter our own cv scores. yes, totally gameable...but that's the point.  That, and no leaderboard probing.\r\n\r\n\r\nCurrently I'm working on a real-world problem and I simply cannot use ensemble as it would end up costing too much or taking too long.  In fact, I was forced to use googlenet instead of a slightly better performing vgg19 as well...for speed and memory.  And there's no leaderboard...I wish there was...it'd make flicking the switch this week less intimidating. As I keep telling my boss \"I'm *pretty* sure this will work...it worked on my test dataset!\".\r\n\r\n\r\n\r\n",
      "votes": 6
    },
    {
      "id": 129877,
      "postDate": "2016-08-02T17:42:05.187Z",
      "content": "<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>",
      "rawMarkdown": "We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... ",
      "votes": 6
    },
    {
      "id": 256111,
      "postDate": "2017-12-11T07:09:24.840Z",
      "content": "<p>can you sharing your demo?</p>",
      "rawMarkdown": "can you sharing your demo?",
      "votes": 1
    },
    {
      "id": 129885,
      "postDate": "2016-08-02T18:27:59.290Z",
      "content": "<p>&quot;semi-supervised&quot; may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You &quot;reverse-engineered&quot; those image in some way and I believe this is essentially not this competition looking for.</p>\n\n<p>State farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.</p>\n\n<p>To be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.</p>\n\n<p>Most of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. </p>\n\n<p>Overall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.</p>\n\n<p>[quote=Gilberto Titericz Junior;129877]</p>\n\n<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "\"semi-supervised\" may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You \"reverse-engineered\" those image in some way and I believe this is essentially not this competition looking for.\r\n\r\nState farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.\r\n\r\nTo be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.\r\n\r\nMost of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. \r\n\r\nOverall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.\r\n\r\n[quote=Gilberto Titericz Junior;129877]\r\n\r\nWe didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... \r\n\r\n[/quote]\r\n",
      "votes": 3
    },
    {
      "id": 129869,
      "postDate": "2016-08-02T17:00:15.257Z",
      "content": "<p>Although this is a great solution to achieve the very good position in LB, I beg this is not something state farm is really looking for. They are very unprofessional to collect dataset based on video and thus this kind of semi-supervised technique to overfit  the test dataset works well for this competition. But I am afraid it is not the real technique to be robust in real life.</p>",
      "rawMarkdown": "Although this is a great solution to achieve the very good position in LB, I beg this is not something state farm is really looking for. They are very unprofessional to collect dataset based on video and thus this kind of semi-supervised technique to overfit  the test dataset works well for this competition. But I am afraid it is not the real technique to be robust in real life.",
      "votes": 3
    },
    {
      "id": 129996,
      "postDate": "2016-08-03T08:51:36.977Z",
      "content": "<p>If you look carefully at modern architectures like ResNet or Inception_v3, you'll find that they use GAP layer by default, so, there is nothing special in adding GAP layer to nets nowadays. What this post really proposes, is the augmentation technique, based on CAMs. I can add, that we also tried this augmentation technique, but it gave similar results with classic techniques for ONE model. So, I cannot agree, that this is some kind of a silver bullet for this competition.</p>\n\n<p>[quote=ShiweiSheng;129888]</p>\n\n<p>If any one is still interesting, I would highly recommend this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>The CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.</p>\n\n<p>This is an example of what I mean the competition is really looking for.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "If you look carefully at modern architectures like ResNet or Inception_v3, you'll find that they use GAP layer by default, so, there is nothing special in adding GAP layer to nets nowadays. What this post really proposes, is the augmentation technique, based on CAMs. I can add, that we also tried this augmentation technique, but it gave similar results with classic techniques for ONE model. So, I cannot agree, that this is some kind of a silver bullet for this competition.\r\n\r\n[quote=ShiweiSheng;129888]\r\n\r\nIf any one is still interesting, I would highly recommend this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nThe CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.\r\n\r\nThis is an example of what I mean the competition is really looking for.\r\n\r\n[/quote]",
      "votes": 1
    },
    {
      "id": 129976,
      "postDate": "2016-08-03T05:32:38.027Z",
      "content": "<p>Congratulation, Thank you for sharing!!   great work.</p>",
      "rawMarkdown": "Congratulation, Thank you for sharing!!   great work.",
      "votes": 1
    },
    {
      "id": 129944,
      "postDate": "2016-08-03T00:42:55.250Z",
      "content": "<p>[quote=Gilberto Titericz Junior;129818]\n we decided to try nearest neighbors on all images.\n....\n4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation\n....</p>\n\n<p>It built very redundant images. Also on these images we tried to center the steering wheel via regression.</p>\n\n<p>[/quote]\nThank you. </p>\n\n<p>But I still have few questions:</p>\n\n<ol>\n<li>How did you measure distance between different pictures?</li>\n<li>What do you mean by 2x augmentation?</li>\n<li>How exactly did you center the steering wheel? </li>\n</ol>",
      "rawMarkdown": "[quote=Gilberto Titericz Junior;129818]\r\n we decided to try nearest neighbors on all images.\r\n....\r\n4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation\r\n....\r\n\r\nIt built very redundant images. Also on these images we tried to center the steering wheel via regression.\r\n\r\n[/quote]\r\nThank you. \r\n\r\nBut I still have few questions:\r\n\r\n 1. How did you measure distance between different pictures?\r\n 2.  What do you mean by 2x augmentation?\r\n 3.  How exactly did you center the steering wheel? ",
      "votes": 1
    },
    {
      "id": 129920,
      "postDate": "2016-08-02T21:20:22.687Z",
      "content": "<p>Much appreciated, no rush, very interested in these techniques &amp; how to get the most out of them.</p>\n\n<p>[quote=Heng CherKeng;129919]</p>\n\n<p>@tetmin\nWe are open sourcing the code . It should be ready within a month, after some &quot;administrative process&quot;. Please wait.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Much appreciated, no rush, very interested in these techniques & how to get the most out of them.\r\n\r\n[quote=Heng CherKeng;129919]\r\n\r\n@tetmin\r\nWe are open sourcing the code . It should be ready within a month, after some \"administrative process\". Please wait.\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 129919,
      "postDate": "2016-08-02T21:12:14.347Z",
      "content": "<p>@tetmin\nWe are open sourcing the code . It should be ready within a month, after some &quot;administrative process&quot;. Please wait.</p>",
      "rawMarkdown": "@tetmin\r\nWe are open sourcing the code . It should be ready within a month, after some \"administrative process\". Please wait.",
      "votes": 1
    },
    {
      "id": 129918,
      "postDate": "2016-08-02T21:09:20.490Z",
      "content": "<p>Speaking of the CAM post, I was never able to recreate those results. Would someone who managed to get around 0.3 LB with that network please share their exact training parameters &amp; data augmentation techniques?</p>\n\n<p>[quote=ShiweiSheng;129888]</p>\n\n<p>If any one is still interesting, I would highly recommend this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>The CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.</p>\n\n<p>This is an example of what I mean the competition is really looking for.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Speaking of the CAM post, I was never able to recreate those results. Would someone who managed to get around 0.3 LB with that network please share their exact training parameters & data augmentation techniques?\r\n\r\n\r\n[quote=ShiweiSheng;129888]\r\n\r\nIf any one is still interesting, I would highly recommend this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nThe CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.\r\n\r\nThis is an example of what I mean the competition is really looking for.\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 129914,
      "postDate": "2016-08-02T20:44:09.330Z",
      "content": "<p>Clever way to modify images~ Congrats!</p>",
      "rawMarkdown": "Clever way to modify images~ Congrats!",
      "votes": 1
    },
    {
      "id": 129893,
      "postDate": "2016-08-02T19:12:24.307Z",
      "content": "<p>I don't mean use a series of image is faulty. I believe a series of image would be more informative than a single one and it would be great when series of image is available.</p>\n\n<p>The problem is, in this competition the host didn't provide a series of images as the data set. They provide each individual image as the data set. I think a more fair way should be everyone is forced to use single image as input, or officially the host provides a group of series of image and assigns one label for each series.</p>\n\n<p>Right now they assign one label per single image so I think the host actually wants kagglers to detect distracted driver based on one single image. Otherwise why don't they simply provide their recorded video as data set? They actually ask drivers to follow certain timeline and take the snapshot at certain time to capture a single image for each class. I think their intention should be use only single image to detect distracted drivers. </p>\n\n<p>For example, just like normal imagenet data set, multiple image for one kind of cat is NOT taken from one cat's continuous action but it is really multiple shot of the same kind of cats. However, the host has only around 20 drivers for train set, it has to generate 'fake' multiple images which don't contain enough information.</p>\n\n<p>Actually this weird method causes a lot of noise for those who use single image as CNN input because of the automatic but rather in accurate capture. </p>\n\n<p>[quote=Luis Andre Dutra e Silva;129890]</p>\n\n<p>@ShiweiSheng,</p>\n\n<p>In &quot;real life&quot; a detector that takes only one frame to analyze a behaviour is a faulty one.\nIn &quot;real life&quot; our solution would take less than 0.34s to respond and I can prove it mathematically and creating a prototype</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I don't mean use a series of image is faulty. I believe a series of image would be more informative than a single one and it would be great when series of image is available.\r\n\r\nThe problem is, in this competition the host didn't provide a series of images as the data set. They provide each individual image as the data set. I think a more fair way should be everyone is forced to use single image as input, or officially the host provides a group of series of image and assigns one label for each series.\r\n\r\nRight now they assign one label per single image so I think the host actually wants kagglers to detect distracted driver based on one single image. Otherwise why don't they simply provide their recorded video as data set? They actually ask drivers to follow certain timeline and take the snapshot at certain time to capture a single image for each class. I think their intention should be use only single image to detect distracted drivers. \r\n\r\nFor example, just like normal imagenet data set, multiple image for one kind of cat is NOT taken from one cat's continuous action but it is really multiple shot of the same kind of cats. However, the host has only around 20 drivers for train set, it has to generate 'fake' multiple images which don't contain enough information.\r\n\r\nActually this weird method causes a lot of noise for those who use single image as CNN input because of the automatic but rather in accurate capture. \r\n\r\n\r\n[quote=Luis Andre Dutra e Silva;129890]\r\n\r\n@ShiweiSheng,\r\n\r\nIn \"real life\" a detector that takes only one frame to analyze a behaviour is a faulty one.\r\nIn \"real life\" our solution would take less than 0.34s to respond and I can prove it mathematically and creating a prototype\r\n\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 129888,
      "postDate": "2016-08-02T18:35:14.737Z",
      "content": "<p>If any one is still interesting, I would highly recommend this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>The CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.</p>\n\n<p>This is an example of what I mean the competition is really looking for.</p>",
      "rawMarkdown": "If any one is still interesting, I would highly recommend this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nThe CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.\r\n\r\nThis is an example of what I mean the competition is really looking for.",
      "votes": 1
    },
    {
      "id": 129831,
      "postDate": "2016-08-02T13:52:47.093Z",
      "content": "<p>Great work, congratulations on coming in 3rd place.</p>\n\n<p>Now the competition is over, it would be great to get some clearer pointers on the exact parameters and procedure used to fine-tune VGG-16 for this task.</p>",
      "rawMarkdown": "Great work, congratulations on coming in 3rd place.\r\n\r\nNow the competition is over, it would be great to get some clearer pointers on the exact parameters and procedure used to fine-tune VGG-16 for this task.",
      "votes": 1
    },
    {
      "id": 129823,
      "postDate": "2016-08-02T13:21:47.253Z",
      "content": "<p>Thanks for sharing your solution. Also congratz to the winners!</p>",
      "rawMarkdown": "Thanks for sharing your solution. Also congratz to the winners!",
      "votes": 1
    },
    {
      "id": 129997,
      "postDate": "2016-08-03T08:52:57.847Z",
      "content": "<p>If you didn't notice some things about the dataset, it doesn't mean that it is &quot;unfair&quot; to use this information. I think that solution proposed here is elegant and what is more is quite simple to implement in &quot;real life&quot;. The authors didn't use temporal information in explicit manner, they just found 20 most similar images to the sample for prediction, and there is no restriction to use several images to make prediction for one image. I find it a lot more cheating to use test set for semi-supervised approaches like pseudo-labbeling etc. And 20 models is not a great number for Kaggle competitions, sometimes there are 1000+ =)</p>\n\n<p>Thanks for sharing, I always find it inspiring to read solutions that are based on good idea.</p>\n\n<p>[quote=frankman;129935]</p>\n\n<p>I agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\n[quote=ShiweiSheng;129885]</p>\n\n<p>&quot;semi-supervised&quot; may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You &quot;reverse-engineered&quot; those image in some way and I believe this is essentially not this competition looking for.</p>\n\n<p>State farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.</p>\n\n<p>To be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.</p>\n\n<p>Most of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. </p>\n\n<p>Overall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.</p>\n\n<p>[quote=Gilberto Titericz Junior;129877]</p>\n\n<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "If you didn't notice some things about the dataset, it doesn't mean that it is \"unfair\" to use this information. I think that solution proposed here is elegant and what is more is quite simple to implement in \"real life\". The authors didn't use temporal information in explicit manner, they just found 20 most similar images to the sample for prediction, and there is no restriction to use several images to make prediction for one image. I find it a lot more cheating to use test set for semi-supervised approaches like pseudo-labbeling etc. And 20 models is not a great number for Kaggle competitions, sometimes there are 1000+ =)\r\n\r\nThanks for sharing, I always find it inspiring to read solutions that are based on good idea.\r\n\r\n[quote=frankman;129935]\r\n\r\nI agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\r\n[quote=ShiweiSheng;129885]\r\n\r\n\"semi-supervised\" may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You \"reverse-engineered\" those image in some way and I believe this is essentially not this competition looking for.\r\n\r\nState farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.\r\n\r\nTo be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.\r\n\r\nMost of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. \r\n\r\nOverall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.\r\n\r\n[quote=Gilberto Titericz Junior;129877]\r\n\r\nWe didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... \r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]",
      "votes": 2
    },
    {
      "id": 129954,
      "postDate": "2016-08-03T03:24:42.750Z",
      "content": "<p>Thanks for the sharing. \nModel ensemble is so important as I see in your solution.\nI didn't take part in this competition. But I also have a try on ResNet using this data set. I gave up very quickly because I found that my local test score was just around 0.33 while top score of other uses on leader board was around 0.16 at that time. </p>",
      "rawMarkdown": "Thanks for the sharing. \r\nModel ensemble is so important as I see in your solution.\r\nI didn't take part in this competition. But I also have a try on ResNet using this data set. I gave up very quickly because I found that my local test score was just around 0.33 while top score of other uses on leader board was around 0.16 at that time. \r\n",
      "votes": 2
    },
    {
      "id": 129854,
      "postDate": "2016-08-02T15:13:00.077Z",
      "content": "<p>Great work, really insightful.</p>",
      "rawMarkdown": "Great work, really insightful.\r\n\r\n",
      "votes": 2
    },
    {
      "id": 129835,
      "postDate": "2016-08-02T14:16:50.863Z",
      "content": "<p>Congratulations for your results and thank you for the post!</p>\n\n<p>Could you elaborate a bit on how you did nearest neighbor? Just pixelwise?</p>",
      "rawMarkdown": "Congratulations for your results and thank you for the post!\r\n\r\nCould you elaborate a bit on how you did nearest neighbor? Just pixelwise?",
      "votes": 2
    },
    {
      "id": 129935,
      "postDate": "2016-08-02T23:38:38.973Z",
      "content": "<p>I agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\n[quote=ShiweiSheng;129885]</p>\n\n<p>&quot;semi-supervised&quot; may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You &quot;reverse-engineered&quot; those image in some way and I believe this is essentially not this competition looking for.</p>\n\n<p>State farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.</p>\n\n<p>To be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.</p>\n\n<p>Most of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. </p>\n\n<p>Overall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.</p>\n\n<p>[quote=Gilberto Titericz Junior;129877]</p>\n\n<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\r\n[quote=ShiweiSheng;129885]\r\n\r\n\"semi-supervised\" may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You \"reverse-engineered\" those image in some way and I believe this is essentially not this competition looking for.\r\n\r\nState farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.\r\n\r\nTo be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.\r\n\r\nMost of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. \r\n\r\nOverall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.\r\n\r\n[quote=Gilberto Titericz Junior;129877]\r\n\r\nWe didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... \r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 129872,
      "postDate": "2016-08-02T17:12:15.713Z",
      "content": "<p>@ShiweiSheng I also have to agree, I will add that in a practical setting if you need ensemble of 24 large convnets to make a single prediction the latency will make the solution impractical. I would place priority on a single high performing model, than ensemble of 20. The 5 kfold here is really 5 models. I wonder if kaggle would start placing restrictions on ensembling beyond a certain number of models.</p>",
      "rawMarkdown": "@ShiweiSheng I also have to agree, I will add that in a practical setting if you need ensemble of 24 large convnets to make a single prediction the latency will make the solution impractical. I would place priority on a single high performing model, than ensemble of 20. The 5 kfold here is really 5 models. I wonder if kaggle would start placing restrictions on ensembling beyond a certain number of models."
    },
    {
      "id": 129878,
      "postDate": "2016-08-02T17:45:44.443Z",
      "content": "<p>@Guanshuo Xu:\n1)- We used the 3 channels of images because we used pretrained models. Also using multiple neighbors images per channel improves stability, adds redundancy and removes outliers from images.\nall other questions: YES</p>",
      "rawMarkdown": "@Guanshuo Xu:\r\n1)- We used the 3 channels of images because we used pretrained models. Also using multiple neighbors images per channel improves stability, adds redundancy and removes outliers from images.\r\nall other questions: YES"
    },
    {
      "id": 129838,
      "postDate": "2016-08-02T14:22:01.070Z",
      "content": "<p>@Lucian Ionita. Yes, I just resized the images to something about 80x60 and got the euclidean distance 20 nearest neighbors</p>",
      "rawMarkdown": "@Lucian Ionita. Yes, I just resized the images to something about 80x60 and got the euclidean distance 20 nearest neighbors"
    },
    {
      "id": 129896,
      "postDate": "2016-08-02T19:30:11.970Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 129890,
      "postDate": "2016-08-02T18:36:03.557Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 310555,
      "postDate": "2018-04-08T00:44:33.030Z",
      "content": "<p>Nice work, thanks for sharing!</p>",
      "rawMarkdown": "Nice work, thanks for sharing!",
      "votes": 1
    },
    {
      "id": 299185,
      "postDate": "2018-03-20T16:44:34.623Z",
      "content": "<p>Thanks for the write-up and nice work.</p>",
      "rawMarkdown": "Thanks for the write-up and nice work.",
      "votes": 1
    },
    {
      "id": 181544,
      "postDate": "2017-05-09T23:11:49.077Z",
      "content": "<p>Very interesting, thanks!</p>",
      "rawMarkdown": "Very interesting, thanks!",
      "votes": 1
    },
    {
      "id": 152560,
      "postDate": "2016-12-27T06:52:34.343Z",
      "content": "<p>thank you for sharing!</p>",
      "rawMarkdown": "thank you for sharing!",
      "votes": 1
    },
    {
      "id": 129833,
      "postDate": "2016-08-02T14:15:15.327Z",
      "content": "<p>Thanks for sharing. </p>",
      "rawMarkdown": "Thanks for sharing. ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 129851,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2016-08-02T15:07:00.273000",
      "content": "<p>Congratulations! This is very interesting work. Could you provide short answers to some of my questions? Appreciate!</p>\n\n<ul>\n<li>When building Modified Images 1 and Modified Images 2, what is idea behind encoding image sequence information into different color channels? Is it only for compatibility with the pre-trained models?</li>\n<li>In these scenarios, you are predicting on short sequences instead of single images. Do you think a short sequence provides a more stable prediction?</li>\n<li>When you center the steeling wheel, did you manually cropped the training data, and built a bounding-box regressor?</li>\n<li>For Level 2 training, you mentioned &quot;each image have its own prediction + 20 neighbors predictions&quot;. Did you concatenate those predictions to form a 210-D (1x10+20x10) vector for each image?</li>\n<li>When performing Level 2 training, did you use the same CV splits as used in the Level 1 training, because otherwise the Level 2 training would be prune to overfitting?</li>\n</ul>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 328883,
      "author_name": "Jhansi Anumula",
      "author_url": "",
      "post_date": "2018-05-15T08:58:59.907000",
      "content": "<p>HI, am newbie to kaggle, may i know about leader board scores? IT should be high or low? If we see in above explanation they got LB score like 0.31,0.27 etc. So am confused. When i submit my results i got 0.31 LB score. I thought my results are very bad.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129903,
      "author_name": "zero zero",
      "author_url": "",
      "post_date": "2016-08-02T19:59:09.447000",
      "content": "<p>StateFarm will unlikely implement any submitted solution exactly as it is presented here anyway.  Kaggle is essentially a fun game and it's understood that most anything goes.   As competitors, we know this.  The prizes are there mostly to encourage ferreting out the best solutions...emphasis on plural!  If they have the hardware to ensemble 100 models from video real-time, that's up to them.  For $65k, they used up several thousands of hours of our time and what sounds to be even more of expensive GPU time.  Money well spent if they want a problem worked on with success.</p>\n\n<p>Big congrats to all the best performers in this comp!  Very well played.  Also, thanks for the many cool scripts that educated many of us in some new areas.  Dark data! I totally wanted to experiment with this...thanks for confirming that it does boost!  </p>\n\n<p>To further address the &quot;not real world&quot;, etc....it's not out of the realm of possibility to hope for a competition in the future to have a different set of requirements like:</p>\n\n<p>-leakage-proofed.  Difficult as it seems there are some of you who seem to specialize in creative ways of discovering it!</p>\n\n<p>-execution speed points(or limit). (balancing hardware, ensembling, etc)</p>\n\n<p>-and, what I would like to see(just for fun), a leaderboard where we enter our own cv scores. yes, totally gameable...but that's the point.  That, and no leaderboard probing.</p>\n\n<p>Currently I'm working on a real-world problem and I simply cannot use ensemble as it would end up costing too much or taking too long.  In fact, I was forced to use googlenet instead of a slightly better performing vgg19 as well...for speed and memory.  And there's no leaderboard...I wish there was...it'd make flicking the switch this week less intimidating. As I keep telling my boss &quot;I'm <em>pretty</em> sure this will work...it worked on my test dataset!&quot;.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 129877,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2016-08-02T17:42:05.187000",
      "content": "<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 256111,
      "author_name": "penmily",
      "author_url": "",
      "post_date": "2017-12-11T07:09:24.840000",
      "content": "<p>can you sharing your demo?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129885,
      "author_name": "ShiweiSheng",
      "author_url": "",
      "post_date": "2016-08-02T18:27:59.290000",
      "content": "<p>&quot;semi-supervised&quot; may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You &quot;reverse-engineered&quot; those image in some way and I believe this is essentially not this competition looking for.</p>\n\n<p>State farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.</p>\n\n<p>To be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.</p>\n\n<p>Most of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. </p>\n\n<p>Overall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.</p>\n\n<p>[quote=Gilberto Titericz Junior;129877]</p>\n\n<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>\n\n<p>[/quote]</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 129869,
      "author_name": "ShiweiSheng",
      "author_url": "",
      "post_date": "2016-08-02T17:00:15.257000",
      "content": "<p>Although this is a great solution to achieve the very good position in LB, I beg this is not something state farm is really looking for. They are very unprofessional to collect dataset based on video and thus this kind of semi-supervised technique to overfit  the test dataset works well for this competition. But I am afraid it is not the real technique to be robust in real life.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 129996,
      "author_name": "Andrey Rykov",
      "author_url": "",
      "post_date": "2016-08-03T08:51:36.977000",
      "content": "<p>If you look carefully at modern architectures like ResNet or Inception_v3, you'll find that they use GAP layer by default, so, there is nothing special in adding GAP layer to nets nowadays. What this post really proposes, is the augmentation technique, based on CAMs. I can add, that we also tried this augmentation technique, but it gave similar results with classic techniques for ONE model. So, I cannot agree, that this is some kind of a silver bullet for this competition.</p>\n\n<p>[quote=ShiweiSheng;129888]</p>\n\n<p>If any one is still interesting, I would highly recommend this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>The CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.</p>\n\n<p>This is an example of what I mean the competition is really looking for.</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129976,
      "author_name": "Anil Kumar",
      "author_url": "",
      "post_date": "2016-08-03T05:32:38.027000",
      "content": "<p>Congratulation, Thank you for sharing!!   great work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129944,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2016-08-03T00:42:55.250000",
      "content": "<p>[quote=Gilberto Titericz Junior;129818]\n we decided to try nearest neighbors on all images.\n....\n4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation\n....</p>\n\n<p>It built very redundant images. Also on these images we tried to center the steering wheel via regression.</p>\n\n<p>[/quote]\nThank you. </p>\n\n<p>But I still have few questions:</p>\n\n<ol>\n<li>How did you measure distance between different pictures?</li>\n<li>What do you mean by 2x augmentation?</li>\n<li>How exactly did you center the steering wheel? </li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129920,
      "author_name": "tetmin",
      "author_url": "",
      "post_date": "2016-08-02T21:20:22.687000",
      "content": "<p>Much appreciated, no rush, very interested in these techniques &amp; how to get the most out of them.</p>\n\n<p>[quote=Heng CherKeng;129919]</p>\n\n<p>@tetmin\nWe are open sourcing the code . It should be ready within a month, after some &quot;administrative process&quot;. Please wait.</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129919,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-08-02T21:12:14.347000",
      "content": "<p>@tetmin\nWe are open sourcing the code . It should be ready within a month, after some &quot;administrative process&quot;. Please wait.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129918,
      "author_name": "tetmin",
      "author_url": "",
      "post_date": "2016-08-02T21:09:20.490000",
      "content": "<p>Speaking of the CAM post, I was never able to recreate those results. Would someone who managed to get around 0.3 LB with that network please share their exact training parameters &amp; data augmentation techniques?</p>\n\n<p>[quote=ShiweiSheng;129888]</p>\n\n<p>If any one is still interesting, I would highly recommend this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>The CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.</p>\n\n<p>This is an example of what I mean the competition is really looking for.</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129914,
      "author_name": "mining_machine",
      "author_url": "",
      "post_date": "2016-08-02T20:44:09.330000",
      "content": "<p>Clever way to modify images~ Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129893,
      "author_name": "ShiweiSheng",
      "author_url": "",
      "post_date": "2016-08-02T19:12:24.307000",
      "content": "<p>I don't mean use a series of image is faulty. I believe a series of image would be more informative than a single one and it would be great when series of image is available.</p>\n\n<p>The problem is, in this competition the host didn't provide a series of images as the data set. They provide each individual image as the data set. I think a more fair way should be everyone is forced to use single image as input, or officially the host provides a group of series of image and assigns one label for each series.</p>\n\n<p>Right now they assign one label per single image so I think the host actually wants kagglers to detect distracted driver based on one single image. Otherwise why don't they simply provide their recorded video as data set? They actually ask drivers to follow certain timeline and take the snapshot at certain time to capture a single image for each class. I think their intention should be use only single image to detect distracted drivers. </p>\n\n<p>For example, just like normal imagenet data set, multiple image for one kind of cat is NOT taken from one cat's continuous action but it is really multiple shot of the same kind of cats. However, the host has only around 20 drivers for train set, it has to generate 'fake' multiple images which don't contain enough information.</p>\n\n<p>Actually this weird method causes a lot of noise for those who use single image as CNN input because of the automatic but rather in accurate capture. </p>\n\n<p>[quote=Luis Andre Dutra e Silva;129890]</p>\n\n<p>@ShiweiSheng,</p>\n\n<p>In &quot;real life&quot; a detector that takes only one frame to analyze a behaviour is a faulty one.\nIn &quot;real life&quot; our solution would take less than 0.34s to respond and I can prove it mathematically and creating a prototype</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129888,
      "author_name": "ShiweiSheng",
      "author_url": "",
      "post_date": "2016-08-02T18:35:14.737000",
      "content": "<p>If any one is still interesting, I would highly recommend this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>The CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.</p>\n\n<p>This is an example of what I mean the competition is really looking for.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129831,
      "author_name": "tetmin",
      "author_url": "",
      "post_date": "2016-08-02T13:52:47.093000",
      "content": "<p>Great work, congratulations on coming in 3rd place.</p>\n\n<p>Now the competition is over, it would be great to get some clearer pointers on the exact parameters and procedure used to fine-tune VGG-16 for this task.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129823,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-08-02T13:21:47.253000",
      "content": "<p>Thanks for sharing your solution. Also congratz to the winners!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129997,
      "author_name": "Andrey Rykov",
      "author_url": "",
      "post_date": "2016-08-03T08:52:57.847000",
      "content": "<p>If you didn't notice some things about the dataset, it doesn't mean that it is &quot;unfair&quot; to use this information. I think that solution proposed here is elegant and what is more is quite simple to implement in &quot;real life&quot;. The authors didn't use temporal information in explicit manner, they just found 20 most similar images to the sample for prediction, and there is no restriction to use several images to make prediction for one image. I find it a lot more cheating to use test set for semi-supervised approaches like pseudo-labbeling etc. And 20 models is not a great number for Kaggle competitions, sometimes there are 1000+ =)</p>\n\n<p>Thanks for sharing, I always find it inspiring to read solutions that are based on good idea.</p>\n\n<p>[quote=frankman;129935]</p>\n\n<p>I agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\n[quote=ShiweiSheng;129885]</p>\n\n<p>&quot;semi-supervised&quot; may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You &quot;reverse-engineered&quot; those image in some way and I believe this is essentially not this competition looking for.</p>\n\n<p>State farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.</p>\n\n<p>To be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.</p>\n\n<p>Most of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. </p>\n\n<p>Overall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.</p>\n\n<p>[quote=Gilberto Titericz Junior;129877]</p>\n\n<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 129954,
      "author_name": "YcdoiT",
      "author_url": "",
      "post_date": "2016-08-03T03:24:42.750000",
      "content": "<p>Thanks for the sharing. \nModel ensemble is so important as I see in your solution.\nI didn't take part in this competition. But I also have a try on ResNet using this data set. I gave up very quickly because I found that my local test score was just around 0.33 while top score of other uses on leader board was around 0.16 at that time. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 129854,
      "author_name": "Wind Bear",
      "author_url": "",
      "post_date": "2016-08-02T15:13:00.077000",
      "content": "<p>Great work, really insightful.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 129835,
      "author_name": "Lucian Ionita",
      "author_url": "",
      "post_date": "2016-08-02T14:16:50.863000",
      "content": "<p>Congratulations for your results and thank you for the post!</p>\n\n<p>Could you elaborate a bit on how you did nearest neighbor? Just pixelwise?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 129935,
      "author_name": "frankman",
      "author_url": "",
      "post_date": "2016-08-02T23:38:38.973000",
      "content": "<p>I agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\n[quote=ShiweiSheng;129885]</p>\n\n<p>&quot;semi-supervised&quot; may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You &quot;reverse-engineered&quot; those image in some way and I believe this is essentially not this competition looking for.</p>\n\n<p>State farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.</p>\n\n<p>To be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.</p>\n\n<p>Most of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. </p>\n\n<p>Overall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.</p>\n\n<p>[quote=Gilberto Titericz Junior;129877]</p>\n\n<p>We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.</p>\n\n<p>@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.</p>\n\n<p>Also blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129872,
      "author_name": "DavidGbodiOdaibo",
      "author_url": "",
      "post_date": "2016-08-02T17:12:15.713000",
      "content": "<p>@ShiweiSheng I also have to agree, I will add that in a practical setting if you need ensemble of 24 large convnets to make a single prediction the latency will make the solution impractical. I would place priority on a single high performing model, than ensemble of 20. The 5 kfold here is really 5 models. I wonder if kaggle would start placing restrictions on ensembling beyond a certain number of models.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129878,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2016-08-02T17:45:44.443000",
      "content": "<p>@Guanshuo Xu:\n1)- We used the 3 channels of images because we used pretrained models. Also using multiple neighbors images per channel improves stability, adds redundancy and removes outliers from images.\nall other questions: YES</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129838,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2016-08-02T14:22:01.070000",
      "content": "<p>@Lucian Ionita. Yes, I just resized the images to something about 80x60 and got the euclidean distance 20 nearest neighbors</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129896,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-08-02T19:30:11.970000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129890,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-08-02T18:36:03.557000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 310555,
      "author_name": "qclijun",
      "author_url": "",
      "post_date": "2018-04-08T00:44:33.030000",
      "content": "<p>Nice work, thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 299185,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-20T16:44:34.623000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 181544,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-05-09T23:11:49.077000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 152560,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-27T06:52:34.343000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129833,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-08-02T14:15:15.327000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129818": "First of all congratulations to Jacobkie and Z_B_C for winning such amazing competition. @Z_B_C we almost draw.\r\nAlso I would like to thanks Kaggle and StateFarm for such a unique competition. We leaned a lot and it is my first DeepNets competition win ;-D\r\n\r\nOur solution is most based in the \"Time\" feature. Just watch the movies in the link and try to figure out why: https://www.kaggle.com/titericz/state-farm-distracted-driver-detection/just-relax-and-watch-some-cool-movies\r\n\r\nAll pictures in trainset are taken in sequence so we explored that characteristic. Trainset is given in the correct sequence, but Testset not. So taking into account that subsequent images are very close to each one (driver position, light, shadows, vehicle external objects, etc...), we decided to try nearest neighbors on all images. And for our surprise it presented very good results catching the nearest images in the correct trainset sequence. The first neighbor have a subject hit ratio of 100% and class hit rate of 99,5% in trainset. So we used a blend of the 20 nearest neighbors in our solution.\r\n\r\nWe trained 8 CNN models, but for our final submission we used only 4. All models trained over 5 folds CV:\r\n\r\n1)- resnet-152 caffe, CV: 0.31 LB: 0.27. Original dataset, no augmentation\r\n\r\n2)- resnet-152 caffe, CV: 0.36 LB: 0.31. Original dataset, some augmentation (under-tuned?)\r\n\r\n3)- resnet-152 torch, CV: 0.223 LB: 0.181. Modified images 1, no augmentation\r\n\r\n4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation\r\n\r\n=>Modified Images 1: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:\r\n\r\nChannel R:  (current image + nearest1)/2\r\n\r\nChannel G:  (nearest2 + nearest3)/2\r\n\r\nChannel B:  (nearest4 + nearest5)/2\r\n\r\nIt built very redundant images. Also on these images we tried to center the steering wheel via regression.\r\n\r\n\r\n=>Modified Images 2: Took current image and 5 nearest neighbor images. For each one of the 3 RGB channels I replaced the image by:\r\n\r\nChannel R:  current image - nearest1\r\n\r\nChannel G: nearest2 - nearest3\r\n\r\nChannel B:  nearest4 - nearest5\r\n\r\nIt build very redundant images and tries to catch drivers movements. \r\n\r\nSo these 4 models presented high diversity between then and predictions are very stable at Level 1 training.\r\n\r\n\r\nSo we used Level 1 prediction of all 4 models to ensemble at Level 2. But this time merging all 20 neighbors predictions of each image. So each image have its own prediction + 20 neighbors predictions. For Level 2 training we used scipy minimize function and created a custom geometric average function to minimize logloss of all models. That architecture improved CV of each model:\r\n\r\n1)- resnet-152 caffe, CV: 0.31 =>  0.192\r\n\r\n2)- resnet-152 caffe, CV: 0.36 => 0.276\r\n\r\n3)- resnet-152 torch, CV: 0.22 => 0.180\r\n\r\n4)- VGG-16 Keras, CV: 0.30 => 0.192\r\n\r\n\r\nOur final solution is and weighted geometric average of these 4 models and CV score is about 0.116 and Class hit Ratio of 96.4%\r\n\r\nAlso we found via cross-validation that replacing all prediction < 0.00001 to zero improved our scores in CV, but we didn't used that in our last submission. If we had choosen it we would finished #1  :-/\r\n\r\n\r\nThat's it and...\r\n\r\nCongratulation again Jacobkie for its late huge jump to #1. And congrats to all top10 teams that didn't overfitted and to everyone that spent much time in this one. Lots of learnings to all.\r\n\r\nThanks again!\r\nGiba\r\n\r\nobs. don't forget to upvote the post and script if you like it ;-P ",
    "129851": "Congratulations! This is very interesting work. Could you provide short answers to some of my questions? Appreciate!\r\n\r\n - When building Modified Images 1 and Modified Images 2, what is idea behind encoding image sequence information into different color channels? Is it only for compatibility with the pre-trained models?\r\n - In these scenarios, you are predicting on short sequences instead of single images. Do you think a short sequence provides a more stable prediction?\r\n - When you center the steeling wheel, did you manually cropped the training data, and built a bounding-box regressor?\r\n - For Level 2 training, you mentioned \"each image have its own prediction + 20 neighbors predictions\". Did you concatenate those predictions to form a 210-D (1x10+20x10) vector for each image?\r\n - When performing Level 2 training, did you use the same CV splits as used in the Level 1 training, because otherwise the Level 2 training would be prune to overfitting?\r\n\r\n\r\n\r\n",
    "328883": "HI, am newbie to kaggle, may i know about leader board scores? IT should be high or low? If we see in above explanation they got LB score like 0.31,0.27 etc. So am confused. When i submit my results i got 0.31 LB score. I thought my results are very bad.",
    "129903": "StateFarm will unlikely implement any submitted solution exactly as it is presented here anyway.  Kaggle is essentially a fun game and it's understood that most anything goes.   As competitors, we know this.  The prizes are there mostly to encourage ferreting out the best solutions...emphasis on plural!  If they have the hardware to ensemble 100 models from video real-time, that's up to them.  For $65k, they used up several thousands of hours of our time and what sounds to be even more of expensive GPU time.  Money well spent if they want a problem worked on with success.\r\n\r\nBig congrats to all the best performers in this comp!  Very well played.  Also, thanks for the many cool scripts that educated many of us in some new areas.  Dark data! I totally wanted to experiment with this...thanks for confirming that it does boost!  \r\n\r\nTo further address the \"not real world\", etc....it's not out of the realm of possibility to hope for a competition in the future to have a different set of requirements like:\r\n\r\n-leakage-proofed.  Difficult as it seems there are some of you who seem to specialize in creative ways of discovering it!\r\n\r\n-execution speed points(or limit). (balancing hardware, ensembling, etc)\r\n\r\n-and, what I would like to see(just for fun), a leaderboard where we enter our own cv scores. yes, totally gameable...but that's the point.  That, and no leaderboard probing.\r\n\r\n\r\nCurrently I'm working on a real-world problem and I simply cannot use ensemble as it would end up costing too much or taking too long.  In fact, I was forced to use googlenet instead of a slightly better performing vgg19 as well...for speed and memory.  And there's no leaderboard...I wish there was...it'd make flicking the switch this week less intimidating. As I keep telling my boss \"I'm *pretty* sure this will work...it worked on my test dataset!\".\r\n\r\n\r\n\r\n",
    "129877": "We didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... ",
    "256111": "can you sharing your demo?",
    "129885": "\"semi-supervised\" may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You \"reverse-engineered\" those image in some way and I believe this is essentially not this competition looking for.\r\n\r\nState farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.\r\n\r\nTo be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.\r\n\r\nMost of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. \r\n\r\nOverall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.\r\n\r\n[quote=Gilberto Titericz Junior;129877]\r\n\r\nWe didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... \r\n\r\n[/quote]\r\n",
    "129869": "Although this is a great solution to achieve the very good position in LB, I beg this is not something state farm is really looking for. They are very unprofessional to collect dataset based on video and thus this kind of semi-supervised technique to overfit  the test dataset works well for this competition. But I am afraid it is not the real technique to be robust in real life.",
    "129996": "If you look carefully at modern architectures like ResNet or Inception_v3, you'll find that they use GAP layer by default, so, there is nothing special in adding GAP layer to nets nowadays. What this post really proposes, is the augmentation technique, based on CAMs. I can add, that we also tried this augmentation technique, but it gave similar results with classic techniques for ONE model. So, I cannot agree, that this is some kind of a silver bullet for this competition.\r\n\r\n[quote=ShiweiSheng;129888]\r\n\r\nIf any one is still interesting, I would highly recommend this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nThe CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.\r\n\r\nThis is an example of what I mean the competition is really looking for.\r\n\r\n[/quote]",
    "129976": "Congratulation, Thank you for sharing!!   great work.",
    "129944": "[quote=Gilberto Titericz Junior;129818]\r\n we decided to try nearest neighbors on all images.\r\n....\r\n4)- VGG-16 Keras, CV: 0.30 LB: 0.28, . Modified images 2, 2x augmentation\r\n....\r\n\r\nIt built very redundant images. Also on these images we tried to center the steering wheel via regression.\r\n\r\n[/quote]\r\nThank you. \r\n\r\nBut I still have few questions:\r\n\r\n 1. How did you measure distance between different pictures?\r\n 2.  What do you mean by 2x augmentation?\r\n 3.  How exactly did you center the steering wheel? ",
    "129920": "Much appreciated, no rush, very interested in these techniques & how to get the most out of them.\r\n\r\n[quote=Heng CherKeng;129919]\r\n\r\n@tetmin\r\nWe are open sourcing the code . It should be ready within a month, after some \"administrative process\". Please wait.\r\n\r\n[/quote]\r\n",
    "129919": "@tetmin\r\nWe are open sourcing the code . It should be ready within a month, after some \"administrative process\". Please wait.",
    "129918": "Speaking of the CAM post, I was never able to recreate those results. Would someone who managed to get around 0.3 LB with that network please share their exact training parameters & data augmentation techniques?\r\n\r\n\r\n[quote=ShiweiSheng;129888]\r\n\r\nIf any one is still interesting, I would highly recommend this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nThe CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.\r\n\r\nThis is an example of what I mean the competition is really looking for.\r\n\r\n[/quote]\r\n",
    "129914": "Clever way to modify images~ Congrats!",
    "129893": "I don't mean use a series of image is faulty. I believe a series of image would be more informative than a single one and it would be great when series of image is available.\r\n\r\nThe problem is, in this competition the host didn't provide a series of images as the data set. They provide each individual image as the data set. I think a more fair way should be everyone is forced to use single image as input, or officially the host provides a group of series of image and assigns one label for each series.\r\n\r\nRight now they assign one label per single image so I think the host actually wants kagglers to detect distracted driver based on one single image. Otherwise why don't they simply provide their recorded video as data set? They actually ask drivers to follow certain timeline and take the snapshot at certain time to capture a single image for each class. I think their intention should be use only single image to detect distracted drivers. \r\n\r\nFor example, just like normal imagenet data set, multiple image for one kind of cat is NOT taken from one cat's continuous action but it is really multiple shot of the same kind of cats. However, the host has only around 20 drivers for train set, it has to generate 'fake' multiple images which don't contain enough information.\r\n\r\nActually this weird method causes a lot of noise for those who use single image as CNN input because of the automatic but rather in accurate capture. \r\n\r\n\r\n[quote=Luis Andre Dutra e Silva;129890]\r\n\r\n@ShiweiSheng,\r\n\r\nIn \"real life\" a detector that takes only one frame to analyze a behaviour is a faulty one.\r\nIn \"real life\" our solution would take less than 0.34s to respond and I can prove it mathematically and creating a prototype\r\n\r\n\r\n[/quote]\r\n",
    "129888": "If any one is still interesting, I would highly recommend this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nThe CAM technique in this paper really improves the performance over vanilla CNN model. It seems the global average pooling layer forces the CNN to pay attention to where really matters on particular image. Therefore CNN improves its performance by looking at the correct part.\r\n\r\nThis is an example of what I mean the competition is really looking for.",
    "129831": "Great work, congratulations on coming in 3rd place.\r\n\r\nNow the competition is over, it would be great to get some clearer pointers on the exact parameters and procedure used to fine-tune VGG-16 for this task.",
    "129823": "Thanks for sharing your solution. Also congratz to the winners!",
    "129997": "If you didn't notice some things about the dataset, it doesn't mean that it is \"unfair\" to use this information. I think that solution proposed here is elegant and what is more is quite simple to implement in \"real life\". The authors didn't use temporal information in explicit manner, they just found 20 most similar images to the sample for prediction, and there is no restriction to use several images to make prediction for one image. I find it a lot more cheating to use test set for semi-supervised approaches like pseudo-labbeling etc. And 20 models is not a great number for Kaggle competitions, sometimes there are 1000+ =)\r\n\r\nThanks for sharing, I always find it inspiring to read solutions that are based on good idea.\r\n\r\n[quote=frankman;129935]\r\n\r\nI agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\r\n[quote=ShiweiSheng;129885]\r\n\r\n\"semi-supervised\" may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You \"reverse-engineered\" those image in some way and I believe this is essentially not this competition looking for.\r\n\r\nState farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.\r\n\r\nTo be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.\r\n\r\nMost of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. \r\n\r\nOverall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.\r\n\r\n[quote=Gilberto Titericz Junior;129877]\r\n\r\nWe didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... \r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]",
    "129954": "Thanks for the sharing. \r\nModel ensemble is so important as I see in your solution.\r\nI didn't take part in this competition. But I also have a try on ResNet using this data set. I gave up very quickly because I found that my local test score was just around 0.33 while top score of other uses on leader board was around 0.16 at that time. \r\n",
    "129854": "Great work, really insightful.\r\n\r\n",
    "129835": "Congratulations for your results and thank you for the post!\r\n\r\nCould you elaborate a bit on how you did nearest neighbor? Just pixelwise?",
    "129935": "I agree with what you have posted.  I think this practice is unfair to other competitors who use static images to predict a label. This solution uses temporal information to improve the result, which defineatly will improve the result over spatial information only, as human beheviour is more well-defined on videos. But state farm obviously does't expect us to use it as they provide each individually labeled image only. If they expect us to use temporal information, they can provide a labelled image sequence or short video. I hope the other two winners use static image only to get such a good score.\r\n[quote=ShiweiSheng;129885]\r\n\r\n\"semi-supervised\" may be not accurate. To be more clearly, what I mean is you do take advantage that each image in this data set is taken from a video clip. You \"reverse-engineered\" those image in some way and I believe this is essentially not this competition looking for.\r\n\r\nState farm's method, ie, shooting video and taking snapshot of certain frame, causes this dataset containing insufficient information. Therefore this dataset exhibits a strong behavior that train set cannot really be generalized to the test set. Your solution to explore the inherent structure of this particular data set is a great solution for this competition. But in real life, let's say we want to detect lots of distracted drivers but each drivers has only one image, I am afraid your technique to enhance  data set won't have any advantage.\r\n\r\nTo be more clearly, a series of action from 20 drivers cannot represent millions of drivers real action in one moment.\r\n\r\nMost of kagglers don't realized that each image has a close neighbor (and this is caused by state farm's collecting method and should not appear in a well-collected dataset). We treat each image alone and try to fit on a very limited-sized train set to generalize on much larger test set. Finally, the quality of this dataset leads to rather, well, I don't know how to say, reults. \r\n\r\nOverall, we don't really get a good new CNN structure, or a new effective layer, or a new kind of loss to coupe with distracted driver detection problem. Rather, we get a smart way to reverse-engineer the dataset.\r\n\r\n[quote=Gilberto Titericz Junior;129877]\r\n\r\nWe didn't over fitted the testset. We didn't used semisupervised learning on our solution. What I proposed is in a real life problem, use last 20 collected driver images to make the prediction and not only the last one. Once the system starts capturing images and calculating probabilities it need only to predict the last image probabilities in a moving window scheme. The other 19 probabilities are already calculated in the past. So there are no delay or excessive computing doing it.\r\n\r\n@David If you built a better single model than any one of my models, congrats! Probably many teams did it. But single models usually don't win competitions.\r\n\r\nAlso blending models is not a problem but a good and stable technique that generalizes very well and perform MUCH better than single models, but only if you really know how to do it... \r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "129872": "@ShiweiSheng I also have to agree, I will add that in a practical setting if you need ensemble of 24 large convnets to make a single prediction the latency will make the solution impractical. I would place priority on a single high performing model, than ensemble of 20. The 5 kfold here is really 5 models. I wonder if kaggle would start placing restrictions on ensembling beyond a certain number of models.",
    "129878": "@Guanshuo Xu:\r\n1)- We used the 3 channels of images because we used pretrained models. Also using multiple neighbors images per channel improves stability, adds redundancy and removes outliers from images.\r\nall other questions: YES",
    "129838": "@Lucian Ionita. Yes, I just resized the images to something about 80x60 and got the euclidean distance 20 nearest neighbors",
    "129896": "",
    "129890": "",
    "310555": "Nice work, thanks for sharing!",
    "299185": "Thanks for the write-up and nice work.",
    "181544": "Very interesting, thanks!",
    "152560": "thank you for sharing!",
    "129833": "Thanks for sharing. "
  }
}