{
  "id": 22659,
  "title": "summary of methods of top 20 methods",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22659",
  "author_name": "",
  "post_date": "2016-08-03T04:41:10.663Z",
  "votes": 24,
  "comment_count": 11,
  "views": 3369,
  "content": "<p>If kagglers can help me to fill up this below, it would be great. Thanks!</p>\n\n<p>format:</p>\n\n<p>[rank] team (public LB--&gt;private LB)   </p>\n\n<ul>\n<li>key methods</li>\n</ul>\n\n<p>...</p>\n\n<hr>\n\n<p>[01]  jacobkie (0.08690-&gt;0.08740)</p>\n\n<ul>\n<li><p>... to be updated ...</p>\n\n<p>@jacobkie:  <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22906/a-brief-summary\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22906/a-brief-summary</a></p></li>\n</ul>\n\n<hr>\n\n<p>[02] Z_B_C (0.08868-&gt;0.09058)</p>\n\n<ul>\n<li>???\n<hr></li>\n</ul>\n\n<p>[03 ]    &#127463;&#127479; BRAZIL POWER &#127463;&#127479; Team *( 0.08877--&gt;0.09058)</p>\n\n<ul>\n<li>ensemble of 4 models of resNet152, vgg16</li>\n<li>use synthetic <strong>test</strong> image = image + nearest neighbor</li>\n<li>multiple image prediction</li>\n</ul>\n\n<p>@Gilberto Titericz Junior : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22631/3-br-power-solution\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22631/3-br-power-solution</a></p>\n\n<hr>\n\n<p>[04 ]     MakeAmericaGreatAgain  (0.08690 --&gt; 0.10065)</p>\n\n<ul>\n<li>???\n<hr></li>\n</ul>\n\n<p>[05 ]  DZS Team (0.10252--&gt;0.12144)</p>\n\n<ul>\n<li>ensemble of ???</li>\n<li>synthetic train image = half+half &quot;cut and paste&quot;. 5 million synthetic image to train googlenet_v3 from scratch using 10 TitianX (single model 0.15)</li>\n</ul>\n\n<p>@DavidGbodiOdaibo : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb</a></p>\n\n<hr>\n\n<p>[06 ]   TitanX &amp;&amp; 1080 Team (0.10050--&gt;0.12673)</p>\n\n<ul>\n<li>ensemble of ???</li>\n<li>semi supervised training using dark knowledge. Use trained model to give pseudo label to test images which is fused with train images for further training (0.125 with dark knowledge, 0.175 without)</li>\n<li>Faster R-CNN with VGG-16 as classifier, treat objection detection task as region sub-region classifier (RPN_POSITIVE_OVERLAP threshold from 0.7 to 0.5 gave a large decrease in loss (~0.22 to 0.175) )</li>\n</ul>\n\n<p>@bobutis : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb</a></p>\n\n<hr>\n\n<p>[07 ]     Vinh Nguyen  (0.13677--&gt; 0.13672)</p>\n\n<ul>\n<li>???\n<hr></li>\n</ul>\n\n<p>[08 ]     nash (0.21812--&gt; 0.13836 ??? )</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[09] iwiwi (0.13780 --&gt;0.14866)</p>\n\n<ul>\n<li>ensemble of resNet101, resNet152</li>\n<li>resDrop training (Deep Networks with Stochastic Depth), pedsuo Label</li>\n<li>graph-based Semi-supervised learning</li>\n</ul>\n\n<p><a href=\"https://twitter.com/iwiwi/status/760281411093868544\">https://twitter.com/iwiwi/status/760281411093868544</a></p>\n\n<hr>\n\n<p>[10] toshi_k (0.14354--&gt;0.14911)</p>\n\n<ul>\n<li>20 models for ensembling</li>\n<li>fully convolution network to detect driver body pixel (semantic segmentation)</li>\n<li>crop driver region (largest bounding box) and use another classifier on this region</li>\n<li>fully supervised, single image prediction</li>\n</ul>\n\n<p>@toshi_k :  <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22666/10th-place-solution\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22666/10th-place-solution</a></p>\n\n<hr>\n\n<p>[11] Alexey Matveev (0.16140--&gt;0.15006)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[12] John Seamons  (0.15085--&gt;0.15033)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[13]  Esper Team  (0.86233--&gt;0.15209 ???)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[14]   XuleiYang (0.12000--&gt;0.15575)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[15 ]   I'mpossible Team (0.15592--&gt;0.15985)</p>\n\n<ul>\n<li><p>ensemble of 150+ models:</p>\n\n<ul><li><p>Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles).</p></li>\n<li><p>Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer).</p></li>\n<li><p>Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)</p></li>\n<li><p>Each trained 16 ~ 32 models, totally more than 150</p></li></ul></li>\n</ul>\n\n<p>@Guanshuo Xu : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22614/how-to-do-cross-validation\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22614/how-to-do-cross-validation</a></p>\n\n<hr>\n\n<p>[16] Jianmin Sun (0.14547--&gt;0.16136)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[17] 4Fun Team  (0.13898--&gt;0.16415)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[18] Balbesy Team  (0.15718--&gt;0.16421)</p>\n\n<ul>\n<li><p>about 7 models: 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101</p></li>\n<li><p>ResNet-101trained only on faces finded by easy haar-like algorithm. We trained it to recognize only 0-8-9 classes.</p></li>\n<li><p>Augmentation: rotation, crop</p></li>\n<li><p>Semi-supervised learning</p></li>\n<li><p>single ResNet-100 gives 0.21 on public LB (batch 12, augmentation). After semi-supervised learning it get us about 0.17</p></li>\n</ul>\n\n<p>@AntonMaltsev : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb</a></p>\n\n<hr>\n\n<p>[19]  Ehsan  (0.14059--&gt;0.16475)</p>\n\n<ul>\n<li><p>Ensemble of 26 models (VGG16 &amp; VGG19)</p></li>\n<li><p>Different shapes (160x160, 192x192, 224x224, 256x256)</p></li>\n<li><p>Semi-supervised with pseudo labeling and entropy regularization.</p></li>\n</ul>\n\n<hr>\n\n<p>[20] King's Hand Team  (0.16820 --&gt; 0.16522)</p>\n\n<ul>\n<li>ensemble of 25 models of googlenetv1/v2, vgg16/19 and resNet50, using class-activation map framework</li>\n<li>use synthetic train image = image + &quot;cut and paste&quot; of random crop of another image</li>\n<li>single image prediction</li>\n<li>fully supervised</li>\n</ul>\n\n<p>@Heng Cher Keng : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<hr>\n\n<p>[26] frankman  (0.16596--&gt; 0.16961)</p>\n\n<ul>\n<li>vgg16 and resnet50 on resized 224x224 images</li>\n<li>a pose estimation model to predict head, left hand, right hand position, shoulders and elbows</li>\n<li>a pure joint position based logistic regression model</li>\n<li>3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand</li>\n<li>1 resenet 50 &amp; 1 resnet 101 &amp; 1 vgg19 model on head part</li>\n<li>1 resnet 50 model on right hand part</li>\n<li>assemble all above models using svm to get final result</li>\n</ul>\n\n<hr>\n\n<p>[60] kyv(0.23746--&gt; 0.19884)</p>\n\n<ul>\n<li>Fine-tuning VGG16 with no dense layers(global average layer instead) with Adam for 10 iterations.  </li>\n<li>Fine tune already fine-tuned model over 8 folds with only 1 iteration.</li>\n<li>Mean ensemble.  </li>\n</ul>",
  "messages": [
    {
      "id": "129970",
      "postDate": "08/03/2016 04:41:10",
      "content": "<p>If kagglers can help me to fill up this below, it would be great. Thanks!</p>\n\n<p>format:</p>\n\n<p>[rank] team (public LB--&gt;private LB)   </p>\n\n<ul>\n<li>key methods</li>\n</ul>\n\n<p>...</p>\n\n<hr>\n\n<p>[01]  jacobkie (0.08690-&gt;0.08740)</p>\n\n<ul>\n<li><p>... to be updated ...</p>\n\n<p>@jacobkie:  <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22906/a-brief-summary\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22906/a-brief-summary</a></p></li>\n</ul>\n\n<hr>\n\n<p>[02] Z_B_C (0.08868-&gt;0.09058)</p>\n\n<ul>\n<li>???\n<hr></li>\n</ul>\n\n<p>[03 ]    &#127463;&#127479; BRAZIL POWER &#127463;&#127479; Team *( 0.08877--&gt;0.09058)</p>\n\n<ul>\n<li>ensemble of 4 models of resNet152, vgg16</li>\n<li>use synthetic <strong>test</strong> image = image + nearest neighbor</li>\n<li>multiple image prediction</li>\n</ul>\n\n<p>@Gilberto Titericz Junior : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22631/3-br-power-solution\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22631/3-br-power-solution</a></p>\n\n<hr>\n\n<p>[04 ]     MakeAmericaGreatAgain  (0.08690 --&gt; 0.10065)</p>\n\n<ul>\n<li>???\n<hr></li>\n</ul>\n\n<p>[05 ]  DZS Team (0.10252--&gt;0.12144)</p>\n\n<ul>\n<li>ensemble of ???</li>\n<li>synthetic train image = half+half &quot;cut and paste&quot;. 5 million synthetic image to train googlenet_v3 from scratch using 10 TitianX (single model 0.15)</li>\n</ul>\n\n<p>@DavidGbodiOdaibo : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb</a></p>\n\n<hr>\n\n<p>[06 ]   TitanX &amp;&amp; 1080 Team (0.10050--&gt;0.12673)</p>\n\n<ul>\n<li>ensemble of ???</li>\n<li>semi supervised training using dark knowledge. Use trained model to give pseudo label to test images which is fused with train images for further training (0.125 with dark knowledge, 0.175 without)</li>\n<li>Faster R-CNN with VGG-16 as classifier, treat objection detection task as region sub-region classifier (RPN_POSITIVE_OVERLAP threshold from 0.7 to 0.5 gave a large decrease in loss (~0.22 to 0.175) )</li>\n</ul>\n\n<p>@bobutis : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb</a></p>\n\n<hr>\n\n<p>[07 ]     Vinh Nguyen  (0.13677--&gt; 0.13672)</p>\n\n<ul>\n<li>???\n<hr></li>\n</ul>\n\n<p>[08 ]     nash (0.21812--&gt; 0.13836 ??? )</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[09] iwiwi (0.13780 --&gt;0.14866)</p>\n\n<ul>\n<li>ensemble of resNet101, resNet152</li>\n<li>resDrop training (Deep Networks with Stochastic Depth), pedsuo Label</li>\n<li>graph-based Semi-supervised learning</li>\n</ul>\n\n<p><a href=\"https://twitter.com/iwiwi/status/760281411093868544\">https://twitter.com/iwiwi/status/760281411093868544</a></p>\n\n<hr>\n\n<p>[10] toshi_k (0.14354--&gt;0.14911)</p>\n\n<ul>\n<li>20 models for ensembling</li>\n<li>fully convolution network to detect driver body pixel (semantic segmentation)</li>\n<li>crop driver region (largest bounding box) and use another classifier on this region</li>\n<li>fully supervised, single image prediction</li>\n</ul>\n\n<p>@toshi_k :  <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22666/10th-place-solution\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22666/10th-place-solution</a></p>\n\n<hr>\n\n<p>[11] Alexey Matveev (0.16140--&gt;0.15006)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[12] John Seamons  (0.15085--&gt;0.15033)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[13]  Esper Team  (0.86233--&gt;0.15209 ???)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[14]   XuleiYang (0.12000--&gt;0.15575)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[15 ]   I'mpossible Team (0.15592--&gt;0.15985)</p>\n\n<ul>\n<li><p>ensemble of 150+ models:</p>\n\n<ul><li><p>Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles).</p></li>\n<li><p>Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer).</p></li>\n<li><p>Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)</p></li>\n<li><p>Each trained 16 ~ 32 models, totally more than 150</p></li></ul></li>\n</ul>\n\n<p>@Guanshuo Xu : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22614/how-to-do-cross-validation\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22614/how-to-do-cross-validation</a></p>\n\n<hr>\n\n<p>[16] Jianmin Sun (0.14547--&gt;0.16136)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[17] 4Fun Team  (0.13898--&gt;0.16415)</p>\n\n<ul>\n<li>???</li>\n</ul>\n\n<hr>\n\n<p>[18] Balbesy Team  (0.15718--&gt;0.16421)</p>\n\n<ul>\n<li><p>about 7 models: 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101</p></li>\n<li><p>ResNet-101trained only on faces finded by easy haar-like algorithm. We trained it to recognize only 0-8-9 classes.</p></li>\n<li><p>Augmentation: rotation, crop</p></li>\n<li><p>Semi-supervised learning</p></li>\n<li><p>single ResNet-100 gives 0.21 on public LB (batch 12, augmentation). After semi-supervised learning it get us about 0.17</p></li>\n</ul>\n\n<p>@AntonMaltsev : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb</a></p>\n\n<hr>\n\n<p>[19]  Ehsan  (0.14059--&gt;0.16475)</p>\n\n<ul>\n<li><p>Ensemble of 26 models (VGG16 &amp; VGG19)</p></li>\n<li><p>Different shapes (160x160, 192x192, 224x224, 256x256)</p></li>\n<li><p>Semi-supervised with pseudo labeling and entropy regularization.</p></li>\n</ul>\n\n<hr>\n\n<p>[20] King's Hand Team  (0.16820 --&gt; 0.16522)</p>\n\n<ul>\n<li>ensemble of 25 models of googlenetv1/v2, vgg16/19 and resNet50, using class-activation map framework</li>\n<li>use synthetic train image = image + &quot;cut and paste&quot; of random crop of another image</li>\n<li>single image prediction</li>\n<li>fully supervised</li>\n</ul>\n\n<p>@Heng Cher Keng : <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<hr>\n\n<p>[26] frankman  (0.16596--&gt; 0.16961)</p>\n\n<ul>\n<li>vgg16 and resnet50 on resized 224x224 images</li>\n<li>a pose estimation model to predict head, left hand, right hand position, shoulders and elbows</li>\n<li>a pure joint position based logistic regression model</li>\n<li>3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand</li>\n<li>1 resenet 50 &amp; 1 resnet 101 &amp; 1 vgg19 model on head part</li>\n<li>1 resnet 50 model on right hand part</li>\n<li>assemble all above models using svm to get final result</li>\n</ul>\n\n<hr>\n\n<p>[60] kyv(0.23746--&gt; 0.19884)</p>\n\n<ul>\n<li>Fine-tuning VGG16 with no dense layers(global average layer instead) with Adam for 10 iterations.  </li>\n<li>Fine tune already fine-tuned model over 8 folds with only 1 iteration.</li>\n<li>Mean ensemble.  </li>\n</ul>",
      "rawMarkdown": "If kagglers can help me to fill up this below, it would be great. Thanks!\r\n\r\nformat:\r\n\r\n[rank] team (public LB-->private LB)   \r\n\r\n- key methods\r\n\r\n   \r\n\r\n...\r\n   \r\n ----------------------------------------------------------------------\r\n\r\n[01]  jacobkie (0.08690->0.08740)\r\n\r\n -  ... to be updated ...\r\n\r\n @jacobkie:  https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22906/a-brief-summary\r\n\r\n ----------------------------------------------------------------------\r\n[02] Z_B_C (0.08868->0.09058)\r\n \r\n- ???\r\n ----------------------------------------------------------------------\r\n[03 ]    🇧🇷 BRAZIL POWER 🇧🇷 Team *( 0.08877-->0.09058)\r\n\r\n- ensemble of 4 models of resNet152, vgg16\r\n- use synthetic **test** image = image + nearest neighbor\r\n- multiple image prediction\r\n\r\n@Gilberto Titericz Junior : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22631/3-br-power-solution\r\n ----------------------------------------------------------------------\r\n[04 ]     MakeAmericaGreatAgain  (0.08690 --> 0.10065)\r\n\r\n- ???\r\n ----------------------------------------------------------------------\r\n[05 ]  DZS Team (0.10252-->0.12144)\r\n\r\n- ensemble of ???\r\n- synthetic train image = half+half \"cut and paste\". 5 million synthetic image to train googlenet_v3 from scratch using 10 TitianX (single model 0.15)\r\n\r\n@DavidGbodiOdaibo : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\r\n\r\n ----------------------------------------------------------------------\r\n\r\n[06 ]   TitanX && 1080 Team (0.10050-->0.12673)\r\n\r\n- ensemble of ???\r\n- semi supervised training using dark knowledge. Use trained model to give pseudo label to test images which is fused with train images for further training (0.125 with dark knowledge, 0.175 without)\r\n- Faster R-CNN with VGG-16 as classifier, treat objection detection task as region sub-region classifier (RPN_POSITIVE_OVERLAP threshold from 0.7 to 0.5 gave a large decrease in loss (~0.22 to 0.175) )\r\n\r\n@bobutis : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\r\n ----------------------------------------------------------------------\r\n[07 ]     Vinh Nguyen  (0.13677--> 0.13672)\r\n\r\n- ???\r\n ----------------------------------------------------------------------\r\n[08 ]     nash (0.21812--> 0.13836 ??? )\r\n\r\n- ???\r\n\r\n----------------------------------------------------------------------\r\n[09] iwiwi (0.13780 -->0.14866)\r\n\r\n- ensemble of resNet101, resNet152\r\n- resDrop training (Deep Networks with Stochastic Depth), pedsuo Label\r\n- graph-based Semi-supervised learning\r\n\r\nhttps://twitter.com/iwiwi/status/760281411093868544\r\n\r\n----------------------------------------------------------------------\r\n[10] toshi_k (0.14354-->0.14911)\r\n\r\n- 20 models for ensembling\r\n- fully convolution network to detect driver body pixel (semantic segmentation)\r\n- crop driver region (largest bounding box) and use another classifier on this region\r\n- fully supervised, single image prediction\r\n\r\n@toshi_k :  https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22666/10th-place-solution\r\n\r\n----------------------------------------------------------------------\r\n[11] Alexey Matveev (0.16140-->0.15006)\r\n\r\n- ???\r\n\r\n---------------------------------------------------------------------- \r\n[12] John Seamons  (0.15085-->0.15033)\r\n\r\n - ???\r\n\r\n----------------------------------------------------------------------   \r\n[13]  Esper Team  (0.86233-->0.15209 ???)\r\n\r\n - ???\r\n\r\n----------------------------------------------------------------------   \r\n\r\n[14]   XuleiYang (0.12000-->0.15575)\r\n   \r\n - ???\r\n\r\n----------------------------------------------------------------------   \r\n[15 ]   I'mpossible Team (0.15592-->0.15985)\r\n\r\n -  ensemble of 150+ models:\r\n\r\n- Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles).\r\n\r\n- Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer).\r\n\r\n- Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)\r\n\r\n- Each trained 16 ~ 32 models, totally more than 150\r\n\r\n\r\n@Guanshuo Xu : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22614/how-to-do-cross-validation\r\n\r\n----------------------------------------------------------------------    \r\n[16] Jianmin Sun (0.14547-->0.16136)\r\n   \r\n - ???\r\n\r\n----------------------------------------------------------------------    \r\n[17] 4Fun Team  (0.13898-->0.16415)\r\n     \r\n - ???\r\n\r\n----------------------------------------------------------------------     \r\n\r\n \r\n[18] Balbesy Team  (0.15718-->0.16421)\r\n\r\n-  about 7 models: 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101\r\n\r\n-  ResNet-101trained only on faces finded by easy haar-like algorithm. We trained it to recognize only 0-8-9 classes.\r\n\r\n-  Augmentation: rotation, crop\r\n\r\n-  Semi-supervised learning\r\n\r\n- single ResNet-100 gives 0.21 on public LB (batch 12, augmentation). After semi-supervised learning it get us about 0.17\r\n\r\n \r\n\r\n\r\n\r\n\r\n@AntonMaltsev : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\r\n\r\n\r\n----------------------------------------------------------------------     \r\n\r\n [19]  Ehsan  (0.14059-->0.16475)\r\n\r\n- Ensemble of 26 models (VGG16 & VGG19)\r\n\r\n- Different shapes (160x160, 192x192, 224x224, 256x256)\r\n\r\n- Semi-supervised with pseudo labeling and entropy regularization.\r\n\r\n----------------------------------------------------------------------     \r\n[20] King's Hand Team  (0.16820 --> 0.16522)\r\n\r\n- ensemble of 25 models of googlenetv1/v2, vgg16/19 and resNet50, using class-activation map framework\r\n- use synthetic train image = image + \"cut and paste\" of random crop of another image\r\n- single image prediction\r\n- fully supervised\r\n\r\n@Heng Cher Keng : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\n----------------------------------------------------------------------     \r\n[26] frankman  (0.16596--> 0.16961)\r\n\r\n-  vgg16 and resnet50 on resized 224x224 images\r\n-  a pose estimation model to predict head, left hand, right hand position, shoulders and elbows\r\n-  a pure joint position based logistic regression model\r\n-  3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand\r\n-  1 resenet 50 & 1 resnet 101 & 1 vgg19 model on head part\r\n-  1 resnet 50 model on right hand part\r\n-  assemble all above models using svm to get final result\r\n\r\n----------------------------------------------------------------------    \r\n \r\n[60] kyv(0.23746--> 0.19884)\r\n \r\n- Fine-tuning VGG16 with no dense layers(global average layer instead) with Adam for 10 iterations.  \r\n- Fine tune already fine-tuned model over 8 folds with only 1 iteration.\r\n- Mean ensemble.",
      "votes": null
    },
    {
      "id": "129983",
      "postDate": "08/03/2016 07:23:07",
      "content": "<p>I want to share mine, though my ranking is only 26 in private LB\n[26] frankmanbb (0.16596--&gt; 0.16961)</p>\n\n<ul>\n<li>vgg16 and resnet50 on resized 224x224 images</li>\n<li>a pose estimation model to predict head, left hand, right hand position, shoulders and elbows</li>\n<li>a pure joint position based logistic regression model </li>\n<li>3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand </li>\n<li>1 resenet 50 &amp; 1 resnet 101 &amp; 1 vgg19  model on head part</li>\n<li>1 resnet 50 model on right hand part</li>\n<li>assemble all above models using svm to get final result</li>\n</ul>",
      "rawMarkdown": "I want to share mine, though my ranking is only 26 in private LB\r\n[26] frankmanbb (0.16596--> 0.16961)\r\n\r\n-  vgg16 and resnet50 on resized 224x224 images\r\n-  a pose estimation model to predict head, left hand, right hand position, shoulders and elbows\r\n- a pure joint position based logistic regression model \r\n-  3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand \r\n-  1 resenet 50 & 1 resnet 101 & 1 vgg19  model on head part\r\n- 1 resnet 50 model on right hand part\r\n-  assemble all above models using svm to get final result",
      "votes": null
    },
    {
      "id": "129987",
      "postDate": "08/03/2016 08:02:28",
      "content": "<p>@frankman thanks a lot!</p>",
      "rawMarkdown": "frankman thanks a lot!",
      "votes": null
    },
    {
      "id": "130044",
      "postDate": "08/03/2016 13:41:32",
      "content": "<p>Ehsan (0.14059--&gt;0.16475)</p>\n\n<ul>\n<li>Ensemble of 26 models (VGG16 &amp; VGG19)</li>\n<li>Different shapes (160x160, 192x192, 224x224, 256x256)</li>\n<li>Semi-supervised with pseudo labeling and entropy regularization. </li>\n</ul>",
      "rawMarkdown": "Ehsan (0.14059-->0.16475)\r\n\r\n - Ensemble of 26 models (VGG16 & VGG19)\r\n - Different shapes (160x160, 192x192, 224x224, 256x256)\r\n - Semi-supervised with pseudo labeling and entropy regularization.",
      "votes": null
    },
    {
      "id": "130077",
      "postDate": "08/03/2016 16:54:52",
      "content": "<p>@Ehsan. Thanks a lot!</p>",
      "rawMarkdown": "Ehsan. Thanks a lot!",
      "votes": null
    },
    {
      "id": "130094",
      "postDate": "08/03/2016 18:35:37",
      "content": "<p>Some more details ...</p>\n\n<ul>\n<li>Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles). </li>\n<li>Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer). </li>\n<li>Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)</li>\n<li>Each trained 16 ~ 32 models, totally more than 150</li>\n</ul>",
      "rawMarkdown": "Some more details ...\r\n\r\n - Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles). \r\n - Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer). \r\n - Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)\r\n - Each trained 16 ~ 32 models, totally more than 150",
      "votes": null
    },
    {
      "id": "130099",
      "postDate": "08/03/2016 20:33:40",
      "content": "<p>0.23746 =&gt; 0.19884</p>\n\n<p>I want to share mine, my position is only 60, but It is possible that It least computationally complex.</p>\n\n<ol>\n<li>Fine-tuning VGG16 with no dense layers(global average  layer instead) with Adam for 10 iterations.  Save model</li>\n<li>Fine tune already fine-tuned model over 8 folds with only 1 iteration. </li>\n<li>Mean ensemble.\nSo total 10+8 forward/backward iterations over simplified VGG16. It was 600s per epoch = 180 min for full model.</li>\n</ol>",
      "rawMarkdown": "0.23746\t=> 0.19884\r\n\r\nI want to share mine, my position is only 60, but It is possible that It least computationally complex.\r\n \r\n1. Fine-tuning VGG16 with no dense layers(global average  layer instead) with Adam for 10 iterations.  Save model\r\n2. Fine tune already fine-tuned model over 8 folds with only 1 iteration. \r\n3. Mean ensemble.\r\nSo total 10+8 forward/backward iterations over simplified VGG16. It was 600s per epoch = 180 min for full model.",
      "votes": null
    },
    {
      "id": "130125",
      "postDate": "08/04/2016 04:13:39",
      "content": "<p>@Guanshuo Xu  and @kyv thanks alot!</p>",
      "rawMarkdown": "Guanshuo Xu  and @kyv thanks alot!",
      "votes": null
    },
    {
      "id": "130139",
      "postDate": "08/04/2016 09:49:43",
      "content": "<p>[18] Balbesy Team</p>\n\n<p>We have about 7 models. 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101 trained only on faces finded by easy haar-like algorithm. We trained it to recognise only 0-8-9 classes. </p>\n\n<p>We also use:</p>\n\n<ul>\n<li>Agumentation: rotation, crop</li>\n<li>Semi-supervised learning</li>\n</ul>",
      "rawMarkdown": "[18] Balbesy Team\r\n\r\nWe have about 7 models. 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101 trained only on faces finded by easy haar-like algorithm. We trained it to recognise only 0-8-9 classes. \r\n\r\nWe also use:\r\n\r\n - Agumentation: rotation, crop\r\n - Semi-supervised learning",
      "votes": null
    },
    {
      "id": "130141",
      "postDate": "08/04/2016 10:14:46",
      "content": "<p>@AntonMaltsev thanks a lot!</p>",
      "rawMarkdown": "AntonMaltsev thanks a lot!",
      "votes": null
    },
    {
      "id": "130174",
      "postDate": "08/04/2016 14:11:35",
      "content": "<p>I want to share my approach too, although I only got to 28th on private LB [28] RaceToTheTop (0.15841--&gt; 0.17005)</p>\n\n<p>I ensembled predictions from the following models.</p>\n\n<p>vgg16, vgg19, googlenet, vgg16 with average pooling + softmax</p>\n\n<p>I used image size of 224x224.</p>\n\n<p>I cropped the training data and test data using faster RCNN to crop images to reduce noise..</p>\n\n<p>Some of the cropped regions using RCNN was wrong so I replaced those with non cropped images. </p>\n\n<p>I trained the model using both the original images and cropped images and averaged the prediction of them.</p>\n\n<p>I (over) used semi supervised learning.</p>\n\n<p>I also used your <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>In hindsight I should have hand labeled the regions of some of the training images to fine tune the RCNN.  </p>\n\n<p>The submission was the arithmetic mean of the predictions from each models.</p>",
      "rawMarkdown": "I want to share my approach too, although I only got to 28th on private LB [28] RaceToTheTop (0.15841--> 0.17005)\r\n\r\nI ensembled predictions from the following models.\r\n\r\nvgg16, vgg19, googlenet, vgg16 with average pooling + softmax\r\n\r\nI used image size of 224x224.\r\n\r\nI cropped the training data and test data using faster RCNN to crop images to reduce noise..\r\n\r\nSome of the cropped regions using RCNN was wrong so I replaced those with non cropped images. \r\n\r\nI trained the model using both the original images and cropped images and averaged the prediction of them.\r\n\r\nI (over) used semi supervised learning.\r\n\r\nI also used your https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nIn hindsight I should have hand labeled the regions of some of the training images to fine tune the RCNN.  \r\n\r\nThe submission was the arithmetic mean of the predictions from each models.",
      "votes": null
    },
    {
      "id": "130181",
      "postDate": "08/04/2016 14:24:25",
      "content": "<p>Hi Heng CherKeng,</p>\n\n<p>Thank you for describing my solution.<br>\nI'm looking forward to seeing 1st and 2nd place solution.</p>",
      "rawMarkdown": "Hi Heng CherKeng,\r\n\r\nThank you for describing my solution.<br>\r\nI'm looking forward to seeing 1st and 2nd place solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129983,
      "author_name": "frankmanbb",
      "author_url": "",
      "post_date": "08/03/2016 07:23:07",
      "content": "<p>I want to share mine, though my ranking is only 26 in private LB\n[26] frankmanbb (0.16596--&gt; 0.16961)</p>\n\n<ul>\n<li>vgg16 and resnet50 on resized 224x224 images</li>\n<li>a pose estimation model to predict head, left hand, right hand position, shoulders and elbows</li>\n<li>a pure joint position based logistic regression model </li>\n<li>3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand </li>\n<li>1 resenet 50 &amp; 1 resnet 101 &amp; 1 vgg19  model on head part</li>\n<li>1 resnet 50 model on right hand part</li>\n<li>assemble all above models using svm to get final result</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129987,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/03/2016 08:02:28",
      "content": "<p>@frankman thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130044,
      "author_name": "ehsanma",
      "author_url": "",
      "post_date": "08/03/2016 13:41:32",
      "content": "<p>Ehsan (0.14059--&gt;0.16475)</p>\n\n<ul>\n<li>Ensemble of 26 models (VGG16 &amp; VGG19)</li>\n<li>Different shapes (160x160, 192x192, 224x224, 256x256)</li>\n<li>Semi-supervised with pseudo labeling and entropy regularization. </li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130077,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/03/2016 16:54:52",
      "content": "<p>@Ehsan. Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130094,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "08/03/2016 18:35:37",
      "content": "<p>Some more details ...</p>\n\n<ul>\n<li>Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles). </li>\n<li>Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer). </li>\n<li>Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)</li>\n<li>Each trained 16 ~ 32 models, totally more than 150</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130099,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "08/03/2016 20:33:40",
      "content": "<p>0.23746 =&gt; 0.19884</p>\n\n<p>I want to share mine, my position is only 60, but It is possible that It least computationally complex.</p>\n\n<ol>\n<li>Fine-tuning VGG16 with no dense layers(global average  layer instead) with Adam for 10 iterations.  Save model</li>\n<li>Fine tune already fine-tuned model over 8 folds with only 1 iteration. </li>\n<li>Mean ensemble.\nSo total 10+8 forward/backward iterations over simplified VGG16. It was 600s per epoch = 180 min for full model.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130125,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/04/2016 04:13:39",
      "content": "<p>@Guanshuo Xu  and @kyv thanks alot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130139,
      "author_name": "zlodeibaal",
      "author_url": "",
      "post_date": "08/04/2016 09:49:43",
      "content": "<p>[18] Balbesy Team</p>\n\n<p>We have about 7 models. 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101 trained only on faces finded by easy haar-like algorithm. We trained it to recognise only 0-8-9 classes. </p>\n\n<p>We also use:</p>\n\n<ul>\n<li>Agumentation: rotation, crop</li>\n<li>Semi-supervised learning</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130141,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/04/2016 10:14:46",
      "content": "<p>@AntonMaltsev thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130174,
      "author_name": "kenkenp",
      "author_url": "",
      "post_date": "08/04/2016 14:11:35",
      "content": "<p>I want to share my approach too, although I only got to 28th on private LB [28] RaceToTheTop (0.15841--&gt; 0.17005)</p>\n\n<p>I ensembled predictions from the following models.</p>\n\n<p>vgg16, vgg19, googlenet, vgg16 with average pooling + softmax</p>\n\n<p>I used image size of 224x224.</p>\n\n<p>I cropped the training data and test data using faster RCNN to crop images to reduce noise..</p>\n\n<p>Some of the cropped regions using RCNN was wrong so I replaced those with non cropped images. </p>\n\n<p>I trained the model using both the original images and cropped images and averaged the prediction of them.</p>\n\n<p>I (over) used semi supervised learning.</p>\n\n<p>I also used your <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output</a></p>\n\n<p>In hindsight I should have hand labeled the regions of some of the training images to fine tune the RCNN.  </p>\n\n<p>The submission was the arithmetic mean of the predictions from each models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130181,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "08/04/2016 14:24:25",
      "content": "<p>Hi Heng CherKeng,</p>\n\n<p>Thank you for describing my solution.<br>\nI'm looking forward to seeing 1st and 2nd place solution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129970": "If kagglers can help me to fill up this below, it would be great. Thanks!\r\n\r\nformat:\r\n\r\n[rank] team (public LB-->private LB)   \r\n\r\n- key methods\r\n\r\n   \r\n\r\n...\r\n   \r\n ----------------------------------------------------------------------\r\n\r\n[01]  jacobkie (0.08690->0.08740)\r\n\r\n -  ... to be updated ...\r\n\r\n @jacobkie:  https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22906/a-brief-summary\r\n\r\n ----------------------------------------------------------------------\r\n[02] Z_B_C (0.08868->0.09058)\r\n \r\n- ???\r\n ----------------------------------------------------------------------\r\n[03 ]    🇧🇷 BRAZIL POWER 🇧🇷 Team *( 0.08877-->0.09058)\r\n\r\n- ensemble of 4 models of resNet152, vgg16\r\n- use synthetic **test** image = image + nearest neighbor\r\n- multiple image prediction\r\n\r\n@Gilberto Titericz Junior : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22631/3-br-power-solution\r\n ----------------------------------------------------------------------\r\n[04 ]     MakeAmericaGreatAgain  (0.08690 --> 0.10065)\r\n\r\n- ???\r\n ----------------------------------------------------------------------\r\n[05 ]  DZS Team (0.10252-->0.12144)\r\n\r\n- ensemble of ???\r\n- synthetic train image = half+half \"cut and paste\". 5 million synthetic image to train googlenet_v3 from scratch using 10 TitianX (single model 0.15)\r\n\r\n@DavidGbodiOdaibo : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\r\n\r\n ----------------------------------------------------------------------\r\n\r\n[06 ]   TitanX && 1080 Team (0.10050-->0.12673)\r\n\r\n- ensemble of ???\r\n- semi supervised training using dark knowledge. Use trained model to give pseudo label to test images which is fused with train images for further training (0.125 with dark knowledge, 0.175 without)\r\n- Faster R-CNN with VGG-16 as classifier, treat objection detection task as region sub-region classifier (RPN_POSITIVE_OVERLAP threshold from 0.7 to 0.5 gave a large decrease in loss (~0.22 to 0.175) )\r\n\r\n@bobutis : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\r\n ----------------------------------------------------------------------\r\n[07 ]     Vinh Nguyen  (0.13677--> 0.13672)\r\n\r\n- ???\r\n ----------------------------------------------------------------------\r\n[08 ]     nash (0.21812--> 0.13836 ??? )\r\n\r\n- ???\r\n\r\n----------------------------------------------------------------------\r\n[09] iwiwi (0.13780 -->0.14866)\r\n\r\n- ensemble of resNet101, resNet152\r\n- resDrop training (Deep Networks with Stochastic Depth), pedsuo Label\r\n- graph-based Semi-supervised learning\r\n\r\nhttps://twitter.com/iwiwi/status/760281411093868544\r\n\r\n----------------------------------------------------------------------\r\n[10] toshi_k (0.14354-->0.14911)\r\n\r\n- 20 models for ensembling\r\n- fully convolution network to detect driver body pixel (semantic segmentation)\r\n- crop driver region (largest bounding box) and use another classifier on this region\r\n- fully supervised, single image prediction\r\n\r\n@toshi_k :  https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22666/10th-place-solution\r\n\r\n----------------------------------------------------------------------\r\n[11] Alexey Matveev (0.16140-->0.15006)\r\n\r\n- ???\r\n\r\n---------------------------------------------------------------------- \r\n[12] John Seamons  (0.15085-->0.15033)\r\n\r\n - ???\r\n\r\n----------------------------------------------------------------------   \r\n[13]  Esper Team  (0.86233-->0.15209 ???)\r\n\r\n - ???\r\n\r\n----------------------------------------------------------------------   \r\n\r\n[14]   XuleiYang (0.12000-->0.15575)\r\n   \r\n - ???\r\n\r\n----------------------------------------------------------------------   \r\n[15 ]   I'mpossible Team (0.15592-->0.15985)\r\n\r\n -  ensemble of 150+ models:\r\n\r\n- Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles).\r\n\r\n- Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer).\r\n\r\n- Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)\r\n\r\n- Each trained 16 ~ 32 models, totally more than 150\r\n\r\n\r\n@Guanshuo Xu : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22614/how-to-do-cross-validation\r\n\r\n----------------------------------------------------------------------    \r\n[16] Jianmin Sun (0.14547-->0.16136)\r\n   \r\n - ???\r\n\r\n----------------------------------------------------------------------    \r\n[17] 4Fun Team  (0.13898-->0.16415)\r\n     \r\n - ???\r\n\r\n----------------------------------------------------------------------     \r\n\r\n \r\n[18] Balbesy Team  (0.15718-->0.16421)\r\n\r\n-  about 7 models: 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101\r\n\r\n-  ResNet-101trained only on faces finded by easy haar-like algorithm. We trained it to recognize only 0-8-9 classes.\r\n\r\n-  Augmentation: rotation, crop\r\n\r\n-  Semi-supervised learning\r\n\r\n- single ResNet-100 gives 0.21 on public LB (batch 12, augmentation). After semi-supervised learning it get us about 0.17\r\n\r\n \r\n\r\n\r\n\r\n\r\n@AntonMaltsev : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/22627/share-your-best-single-model-score-on-public-lb\r\n\r\n\r\n----------------------------------------------------------------------     \r\n\r\n [19]  Ehsan  (0.14059-->0.16475)\r\n\r\n- Ensemble of 26 models (VGG16 & VGG19)\r\n\r\n- Different shapes (160x160, 192x192, 224x224, 256x256)\r\n\r\n- Semi-supervised with pseudo labeling and entropy regularization.\r\n\r\n----------------------------------------------------------------------     \r\n[20] King's Hand Team  (0.16820 --> 0.16522)\r\n\r\n- ensemble of 25 models of googlenetv1/v2, vgg16/19 and resNet50, using class-activation map framework\r\n- use synthetic train image = image + \"cut and paste\" of random crop of another image\r\n- single image prediction\r\n- fully supervised\r\n\r\n@Heng Cher Keng : https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\n----------------------------------------------------------------------     \r\n[26] frankman  (0.16596--> 0.16961)\r\n\r\n-  vgg16 and resnet50 on resized 224x224 images\r\n-  a pose estimation model to predict head, left hand, right hand position, shoulders and elbows\r\n-  a pure joint position based logistic regression model\r\n-  3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand\r\n-  1 resenet 50 & 1 resnet 101 & 1 vgg19 model on head part\r\n-  1 resnet 50 model on right hand part\r\n-  assemble all above models using svm to get final result\r\n\r\n----------------------------------------------------------------------    \r\n \r\n[60] kyv(0.23746--> 0.19884)\r\n \r\n- Fine-tuning VGG16 with no dense layers(global average layer instead) with Adam for 10 iterations.  \r\n- Fine tune already fine-tuned model over 8 folds with only 1 iteration.\r\n- Mean ensemble.",
    "129983": "I want to share mine, though my ranking is only 26 in private LB\r\n[26] frankmanbb (0.16596--> 0.16961)\r\n\r\n-  vgg16 and resnet50 on resized 224x224 images\r\n-  a pose estimation model to predict head, left hand, right hand position, shoulders and elbows\r\n- a pure joint position based logistic regression model \r\n-  3 vgg 16 models trained on 224x224 cropped on head, left hand, right hand \r\n-  1 resenet 50 & 1 resnet 101 & 1 vgg19  model on head part\r\n- 1 resnet 50 model on right hand part\r\n-  assemble all above models using svm to get final result",
    "129987": "frankman thanks a lot!",
    "130044": "Ehsan (0.14059-->0.16475)\r\n\r\n - Ensemble of 26 models (VGG16 & VGG19)\r\n - Different shapes (160x160, 192x192, 224x224, 256x256)\r\n - Semi-supervised with pseudo labeling and entropy regularization.",
    "130077": "Ehsan. Thanks a lot!",
    "130094": "Some more details ...\r\n\r\n - Trained bbox regressors to cut heads and st wheels (expanded the bboxes to around 256x256 to include more valuable information such as phones and bottles). \r\n - Input to the CNNs are composed of the original images together with their corresponding heads and st wheel images. These three parts are separately initialized by pre-trained nets (I call them 3-way nets, the CNNs only involve original images I call them 1-way nets). They are fused at the last fully connected layers (usually the global pooling layer). \r\n - Included pre-trained CNNs are: resnet50 (1-way and 3-way), resnet110 (3-way), blvc_inception (1-way and 3-way), bn_inception (3-way), princeton_inception (3-way)\r\n - Each trained 16 ~ 32 models, totally more than 150",
    "130099": "0.23746\t=> 0.19884\r\n\r\nI want to share mine, my position is only 60, but It is possible that It least computationally complex.\r\n \r\n1. Fine-tuning VGG16 with no dense layers(global average  layer instead) with Adam for 10 iterations.  Save model\r\n2. Fine tune already fine-tuned model over 8 folds with only 1 iteration. \r\n3. Mean ensemble.\r\nSo total 10+8 forward/backward iterations over simplified VGG16. It was 600s per epoch = 180 min for full model.",
    "130125": "Guanshuo Xu  and @kyv thanks alot!",
    "130139": "[18] Balbesy Team\r\n\r\nWe have about 7 models. 4 - ResNet-100, 2- ResNet-50 and one - ResNet-101 trained only on faces finded by easy haar-like algorithm. We trained it to recognise only 0-8-9 classes. \r\n\r\nWe also use:\r\n\r\n - Agumentation: rotation, crop\r\n - Semi-supervised learning",
    "130141": "AntonMaltsev thanks a lot!",
    "130174": "I want to share my approach too, although I only got to 28th on private LB [28] RaceToTheTop (0.15841--> 0.17005)\r\n\r\nI ensembled predictions from the following models.\r\n\r\nvgg16, vgg19, googlenet, vgg16 with average pooling + softmax\r\n\r\nI used image size of 224x224.\r\n\r\nI cropped the training data and test data using faster RCNN to crop images to reduce noise..\r\n\r\nSome of the cropped regions using RCNN was wrong so I replaced those with non cropped images. \r\n\r\nI trained the model using both the original images and cropped images and averaged the prediction of them.\r\n\r\nI (over) used semi supervised learning.\r\n\r\nI also used your https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output\r\n\r\nIn hindsight I should have hand labeled the regions of some of the training images to fine tune the RCNN.  \r\n\r\nThe submission was the arithmetic mean of the predictions from each models.",
    "130181": "Hi Heng CherKeng,\r\n\r\nThank you for describing my solution.<br>\r\nI'm looking forward to seeing 1st and 2nd place solution."
  },
  "source": "meta"
}