{
  "id": 22615,
  "title": "Code and strategy sharing",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22615",
  "author_name": "",
  "post_date": "2016-08-02T00:15:16.607Z",
  "votes": 15,
  "comment_count": 7,
  "views": 841,
  "content": "<p>Now that the competition has ended, time to share !</p>\n\n<p>What worked:</p>\n\n<ul>\n<li>Pretrained VGG16 VGG19 ResNet50</li>\n<li>Data augmentation (although I'd have a hard time quantifying how much)</li>\n<li>Model averaging (simple geometric, harmonic or standard averaging)</li>\n<li>The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss</li>\n</ul>\n\n<p>What did not work:</p>\n\n<ul>\n<li>Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)</li>\n<li>Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).</li>\n<li>Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.</li>\n<li>Training from scratch on external CelebA dataset and using the trained weights for the comp.</li>\n<li>Adversarial pre-training (pre-training a net to distinguish between train and test images)</li>\n<li>Also tried a clustering approach to ensemble (using t-sne / largevis on the predicted probabilities from several models)</li>\n<li>Ensembling by voting</li>\n</ul>\n\n<p>I uploaded most of my code in my <a href=\"https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition\">Kaggle folder</a></p>",
  "messages": [
    {
      "id": "129725",
      "postDate": "08/02/2016 00:15:16",
      "content": "<p>Now that the competition has ended, time to share !</p>\n\n<p>What worked:</p>\n\n<ul>\n<li>Pretrained VGG16 VGG19 ResNet50</li>\n<li>Data augmentation (although I'd have a hard time quantifying how much)</li>\n<li>Model averaging (simple geometric, harmonic or standard averaging)</li>\n<li>The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss</li>\n</ul>\n\n<p>What did not work:</p>\n\n<ul>\n<li>Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)</li>\n<li>Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).</li>\n<li>Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.</li>\n<li>Training from scratch on external CelebA dataset and using the trained weights for the comp.</li>\n<li>Adversarial pre-training (pre-training a net to distinguish between train and test images)</li>\n<li>Also tried a clustering approach to ensemble (using t-sne / largevis on the predicted probabilities from several models)</li>\n<li>Ensembling by voting</li>\n</ul>\n\n<p>I uploaded most of my code in my <a href=\"https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition\">Kaggle folder</a></p>",
      "rawMarkdown": "Now that the competition has ended, time to share !\r\n\r\nWhat worked:\r\n\r\n- Pretrained VGG16 VGG19 ResNet50\r\n- Data augmentation (although I'd have a hard time quantifying how much)\r\n- Model averaging (simple geometric, harmonic or standard averaging)\r\n- The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss\r\n\r\nWhat did not work:\r\n\r\n- Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)\r\n- Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).\r\n- Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.\r\n- Training from scratch on external CelebA dataset and using the trained weights for the comp.\r\n- Adversarial pre-training (pre-training a net to distinguish between train and test images)\r\n- Also tried a clustering approach to ensemble (using t-sne / largevis on the predicted probabilities from several models)\r\n- Ensembling by voting\r\n\r\nI uploaded most of my code in my [Kaggle folder][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition",
      "votes": null
    },
    {
      "id": "129733",
      "postDate": "08/02/2016 00:55:09",
      "content": "<p>Excellent work and write up, tmain! </p>\n\n<p>Thanks for sharing. </p>\n\n<p>[quote=tmain;129725]</p>\n\n<p>Now that the competition has ended, time to share !</p>\n\n<p>What worked:</p>\n\n<ul>\n<li>Pretrained VGG16 VGG19 ResNet50</li>\n<li>Data augmentation (although I'd have a hard time quantifying how much)</li>\n<li>Model averaging (simple geometric, harmonic or standard averaging)</li>\n<li>The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss</li>\n</ul>\n\n<p>What did not work:</p>\n\n<ul>\n<li>Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)</li>\n<li>Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).</li>\n<li>Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.</li>\n<li>Training from scratch on external CelebA dataset and using the trained weights for the comp.</li>\n<li>Adversarial pre-training (pre-training a net to distinguish between train and test images)</li>\n</ul>\n\n<p>I uploaded most of my code in my <a href=\"https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition\">Kaggle folder</a></p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Excellent work and write up, tmain! \r\n\r\nThanks for sharing. \r\n\r\n[quote=tmain;129725]\r\n\r\nNow that the competition has ended, time to share !\r\n\r\nWhat worked:\r\n\r\n- Pretrained VGG16 VGG19 ResNet50\r\n- Data augmentation (although I'd have a hard time quantifying how much)\r\n- Model averaging (simple geometric, harmonic or standard averaging)\r\n- The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss\r\n\r\nWhat did not work:\r\n\r\n- Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)\r\n- Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).\r\n- Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.\r\n- Training from scratch on external CelebA dataset and using the trained weights for the comp.\r\n- Adversarial pre-training (pre-training a net to distinguish between train and test images)\r\n\r\nI uploaded most of my code in my [Kaggle folder][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "129738",
      "postDate": "08/02/2016 01:16:13",
      "content": "<p>Thanks for the great sharing. </p>\n\n<p>My only doubt is &quot;semi-supervised learning&quot; seems cross the border. It is a bit like peek the answer beforehand. In reality, it means the model has no practical use unless hire someone to judge the result manually like LB did.</p>\n\n<p>I am not criticizing this strategy, but using testing data in this way is more like for kaggle's own sake.</p>",
      "rawMarkdown": "Thanks for the great sharing. \r\n\r\nMy only doubt is \"semi-supervised learning\" seems cross the border. It is a bit like peek the answer beforehand. In reality, it means the model has no practical use unless hire someone to judge the result manually like LB did.\r\n\r\nI am not criticizing this strategy, but using testing data in this way is more like for kaggle's own sake.",
      "votes": null
    },
    {
      "id": "129747",
      "postDate": "08/02/2016 03:11:23",
      "content": "<p>Wind Bear, it's not a peek at the answer, it's a peek at the question.</p>",
      "rawMarkdown": "Wind Bear, it's not a peek at the answer, it's a peek at the question.",
      "votes": null
    },
    {
      "id": "129751",
      "postDate": "08/02/2016 03:31:16",
      "content": "<p>[quote=Lucian Ionita;129747]</p>\n\n<p>Wind Bear, it's not a peek at the answer, it's a peek at the question.</p>\n\n<p>[/quote]</p>\n\n<p>You are right. :-) </p>",
      "rawMarkdown": "[quote=Lucian Ionita;129747]\r\n\r\nWind Bear, it's not a peek at the answer, it's a peek at the question.\r\n\r\n[/quote]\r\n\r\nYou are right. :-)",
      "votes": null
    },
    {
      "id": "129752",
      "postDate": "08/02/2016 03:55:20",
      "content": "<p>Very valuable information. Thanks for sharing!</p>\n\n<p>I still don't understand how semi-supervised works here. Is it combining predicted sample with training data, and train another model on both?</p>",
      "rawMarkdown": "Very valuable information. Thanks for sharing!\r\n\r\nI still don't understand how semi-supervised works here. Is it combining predicted sample with training data, and train another model on both?",
      "votes": null
    },
    {
      "id": "129753",
      "postDate": "08/02/2016 04:26:05",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "129789",
      "postDate": "08/02/2016 10:36:58",
      "content": "<p>My humble approach was how far would a &quot;simple&quot; convnet  go using low resolution images (64x64)  no pretrain and very limited data augmentation (10% zooming-displacement in horizontal axis, making train 4 times bigger).</p>\n\n<p>With that constraints... \nWhat worked:\n- dropping. Big way, p=0.8. In hidden layer and, to a smaller extent, in higher level last conv. layer\n(I think the reason was  overfitting was most trickiest issue with such small number of drivers)\n-relu everywhere\n-small kernels never bigger than 3x3</p>\n\n<p>What didnt work:\n- kernels bigger than 3x3\n- inception-like layers for filter number reduction, not so much for reducing computing power but aiming to help with overfitting by reducing parameter numbers. I could not find any use of inception (&quot;filter level pooling&quot;) in my context.</p>\n\n<p>In the end, 1.19864 in private leaderboard, not far from my public result and of course a world away from the top. So that is the farthest I could go with a 3 convolutional (+pooling) layer plus hidden layer with HUGE dropping. In case anyone with similar resolution 64x64 and no pretrain did significantly better than that I would love to know about net architecture differences. </p>",
      "rawMarkdown": "My humble approach was how far would a \"simple\" convnet  go using low resolution images (64x64)  no pretrain and very limited data augmentation (10% zooming-displacement in horizontal axis, making train 4 times bigger).\r\n\r\nWith that constraints... \r\nWhat worked:\r\n- dropping. Big way, p=0.8. In hidden layer and, to a smaller extent, in higher level last conv. layer\r\n(I think the reason was  overfitting was most trickiest issue with such small number of drivers)\r\n-relu everywhere\r\n-small kernels never bigger than 3x3\r\n\r\nWhat didnt work:\r\n- kernels bigger than 3x3\r\n- inception-like layers for filter number reduction, not so much for reducing computing power but aiming to help with overfitting by reducing parameter numbers. I could not find any use of inception (\"filter level pooling\") in my context.\r\n\r\nIn the end, 1.19864 in private leaderboard, not far from my public result and of course a world away from the top. So that is the farthest I could go with a 3 convolutional (+pooling) layer plus hidden layer with HUGE dropping. In case anyone with similar resolution 64x64 and no pretrain did significantly better than that I would love to know about net architecture differences.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129733,
      "author_name": "vinhnguyen",
      "author_url": "",
      "post_date": "08/02/2016 00:55:09",
      "content": "<p>Excellent work and write up, tmain! </p>\n\n<p>Thanks for sharing. </p>\n\n<p>[quote=tmain;129725]</p>\n\n<p>Now that the competition has ended, time to share !</p>\n\n<p>What worked:</p>\n\n<ul>\n<li>Pretrained VGG16 VGG19 ResNet50</li>\n<li>Data augmentation (although I'd have a hard time quantifying how much)</li>\n<li>Model averaging (simple geometric, harmonic or standard averaging)</li>\n<li>The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss</li>\n</ul>\n\n<p>What did not work:</p>\n\n<ul>\n<li>Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)</li>\n<li>Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).</li>\n<li>Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.</li>\n<li>Training from scratch on external CelebA dataset and using the trained weights for the comp.</li>\n<li>Adversarial pre-training (pre-training a net to distinguish between train and test images)</li>\n</ul>\n\n<p>I uploaded most of my code in my <a href=\"https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition\">Kaggle folder</a></p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129738,
      "author_name": "yangyang",
      "author_url": "",
      "post_date": "08/02/2016 01:16:13",
      "content": "<p>Thanks for the great sharing. </p>\n\n<p>My only doubt is &quot;semi-supervised learning&quot; seems cross the border. It is a bit like peek the answer beforehand. In reality, it means the model has no practical use unless hire someone to judge the result manually like LB did.</p>\n\n<p>I am not criticizing this strategy, but using testing data in this way is more like for kaggle's own sake.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129747,
      "author_name": "ilucian",
      "author_url": "",
      "post_date": "08/02/2016 03:11:23",
      "content": "<p>Wind Bear, it's not a peek at the answer, it's a peek at the question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129751,
      "author_name": "yangyang",
      "author_url": "",
      "post_date": "08/02/2016 03:31:16",
      "content": "<p>[quote=Lucian Ionita;129747]</p>\n\n<p>Wind Bear, it's not a peek at the answer, it's a peek at the question.</p>\n\n<p>[/quote]</p>\n\n<p>You are right. :-) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129752,
      "author_name": "ferris",
      "author_url": "",
      "post_date": "08/02/2016 03:55:20",
      "content": "<p>Very valuable information. Thanks for sharing!</p>\n\n<p>I still don't understand how semi-supervised works here. Is it combining predicted sample with training data, and train another model on both?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129753,
      "author_name": "sagarverma",
      "author_url": "",
      "post_date": "08/02/2016 04:26:05",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129789,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "08/02/2016 10:36:58",
      "content": "<p>My humble approach was how far would a &quot;simple&quot; convnet  go using low resolution images (64x64)  no pretrain and very limited data augmentation (10% zooming-displacement in horizontal axis, making train 4 times bigger).</p>\n\n<p>With that constraints... \nWhat worked:\n- dropping. Big way, p=0.8. In hidden layer and, to a smaller extent, in higher level last conv. layer\n(I think the reason was  overfitting was most trickiest issue with such small number of drivers)\n-relu everywhere\n-small kernels never bigger than 3x3</p>\n\n<p>What didnt work:\n- kernels bigger than 3x3\n- inception-like layers for filter number reduction, not so much for reducing computing power but aiming to help with overfitting by reducing parameter numbers. I could not find any use of inception (&quot;filter level pooling&quot;) in my context.</p>\n\n<p>In the end, 1.19864 in private leaderboard, not far from my public result and of course a world away from the top. So that is the farthest I could go with a 3 convolutional (+pooling) layer plus hidden layer with HUGE dropping. In case anyone with similar resolution 64x64 and no pretrain did significantly better than that I would love to know about net architecture differences. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129725": "Now that the competition has ended, time to share !\r\n\r\nWhat worked:\r\n\r\n- Pretrained VGG16 VGG19 ResNet50\r\n- Data augmentation (although I'd have a hard time quantifying how much)\r\n- Model averaging (simple geometric, harmonic or standard averaging)\r\n- The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss\r\n\r\nWhat did not work:\r\n\r\n- Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)\r\n- Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).\r\n- Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.\r\n- Training from scratch on external CelebA dataset and using the trained weights for the comp.\r\n- Adversarial pre-training (pre-training a net to distinguish between train and test images)\r\n- Also tried a clustering approach to ensemble (using t-sne / largevis on the predicted probabilities from several models)\r\n- Ensembling by voting\r\n\r\nI uploaded most of my code in my [Kaggle folder][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition",
    "129733": "Excellent work and write up, tmain! \r\n\r\nThanks for sharing. \r\n\r\n[quote=tmain;129725]\r\n\r\nNow that the competition has ended, time to share !\r\n\r\nWhat worked:\r\n\r\n- Pretrained VGG16 VGG19 ResNet50\r\n- Data augmentation (although I'd have a hard time quantifying how much)\r\n- Model averaging (simple geometric, harmonic or standard averaging)\r\n- The real key was semi-supervised learning (make predictions on test set and use predictions as soft targets). This had a real impact on LB and an excellent regularising effect on my valid loss\r\n\r\nWhat did not work:\r\n\r\n- Model stacking (for some reason, simple average always beat any clever ensemble I tried to train)\r\n- Test time augmentation (i.e. combining test predictions on several perturbed versions of the test images).\r\n- Neural style to create stylized train image with the hope it would have a regularizing effect but without much success.\r\n- Training from scratch on external CelebA dataset and using the trained weights for the comp.\r\n- Adversarial pre-training (pre-training a net to distinguish between train and test images)\r\n\r\nI uploaded most of my code in my [Kaggle folder][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/Kaggle/tree/master/StateFarm/Competition\r\n\r\n[/quote]",
    "129738": "Thanks for the great sharing. \r\n\r\nMy only doubt is \"semi-supervised learning\" seems cross the border. It is a bit like peek the answer beforehand. In reality, it means the model has no practical use unless hire someone to judge the result manually like LB did.\r\n\r\nI am not criticizing this strategy, but using testing data in this way is more like for kaggle's own sake.",
    "129747": "Wind Bear, it's not a peek at the answer, it's a peek at the question.",
    "129751": "[quote=Lucian Ionita;129747]\r\n\r\nWind Bear, it's not a peek at the answer, it's a peek at the question.\r\n\r\n[/quote]\r\n\r\nYou are right. :-)",
    "129752": "Very valuable information. Thanks for sharing!\r\n\r\nI still don't understand how semi-supervised works here. Is it combining predicted sample with training data, and train another model on both?",
    "129753": "Thanks for sharing.",
    "129789": "My humble approach was how far would a \"simple\" convnet  go using low resolution images (64x64)  no pretrain and very limited data augmentation (10% zooming-displacement in horizontal axis, making train 4 times bigger).\r\n\r\nWith that constraints... \r\nWhat worked:\r\n- dropping. Big way, p=0.8. In hidden layer and, to a smaller extent, in higher level last conv. layer\r\n(I think the reason was  overfitting was most trickiest issue with such small number of drivers)\r\n-relu everywhere\r\n-small kernels never bigger than 3x3\r\n\r\nWhat didnt work:\r\n- kernels bigger than 3x3\r\n- inception-like layers for filter number reduction, not so much for reducing computing power but aiming to help with overfitting by reducing parameter numbers. I could not find any use of inception (\"filter level pooling\") in my context.\r\n\r\nIn the end, 1.19864 in private leaderboard, not far from my public result and of course a world away from the top. So that is the farthest I could go with a 3 convolutional (+pooling) layer plus hidden layer with HUGE dropping. In case anyone with similar resolution 64x64 and no pretrain did significantly better than that I would love to know about net architecture differences."
  },
  "source": "meta"
}