{
  "id": 49367,
  "title": "1st place solution",
  "url": "/competitions/sp-society-camera-model-identification/writeups/ods-ai-stamp-1st-place-solution",
  "author_name": "",
  "post_date": "2018-02-10T14:01:14.493Z",
  "votes": 54,
  "comment_count": 29,
  "views": 0,
  "content": "<h2>Models</h2>\n\n<p>Here is the list of our best model. Half of them is based on Andres code (forked before adding any restrictions on usage) and half on its PyTorch implementation:</p>\n\n<ul>\n<li>984_densenet201_antorsaegen_29_0.98624</li>\n<li>976_densenet201_antorsaegen_62_0.98271</li>\n<li>977_resnet50_antorsaegen_119_val_0.9815</li>\n<li>976_DenseNet201_do0.3_doc0.0_avg-epoch072-val_acc0.981250</li>\n<li>967_InceptionResNetV2_do0.1_avg-epoch154-val_acc0.965625</li>\n<li>962_Xception_do0.3_avg-epoch079-val_acc0.991667 (leaky validation, pls ignore)</li>\n</ul>\n\n<p>All of the models had a full-size crop (512) + TTA 8. Choosing this crop was our biggest flaw since smaller crops allow faster training, more TTA options and produce almost the same accuracy.  </p>\n\n<h2>Data</h2>\n\n<p>In total, we collected 300 Gb of photos from Flickr and Yandex.Foto. But due to the huge crop size, we could not fully utilize the entire data set and used only 20,000 photos for training. The rest 50,000 was used for blending. \nData was filtered by resolution and camera type, also all photos with non-default software in Exif  (i.e. processed somehow) and with poor quality (lower than 95 as identified by ImageMagick) were excluded.</p>\n\n<h2>Hardware</h2>\n\n<p>Nothing special for a team with five members: 5 x 1080ti + 1070</p>\n\n<h2>Submissions</h2>\n\n<p>We had 2 submission strategies:</p>\n\n<ul>\n<li>Averaging all TTA prediction from all models by power mean - powers of 1,2,4 as a parameter. Additionally, we explicitly made all classes equally distributed (132 manip and 132 unalt photos for each class) by using Hungarian algorithm on probabilities. As expected, this resulted in a huge LB overfit: 0.991 score on public and 0.985 on private.</li>\n<li>Blending models by using hold-out predictions. There were 52 models with slightly different parameters and equal weights: XGBoost (20) + LightGBM (20) + Keras (12). This gave us 0.986 score on public and 0.989 on private. No class equalization was used with this approach except some weight tuning for Nexus 5 during training (all models had a smaller number of predictions for this phone). </li>\n</ul>\n\n<h2>Key takeaways</h2>\n\n<ul>\n<li>Collect as much data as possible but do not forget to clean it (GIGO)</li>\n<li>Try smaller crops first and all architecture types</li>\n<li>LB probing is evil, trust your CV</li>\n</ul>",
  "messages": [
    {
      "id": "280404",
      "postDate": "02/09/2018 22:33:39",
      "content": "<h2>Models</h2>\n\n<p>Here is the list of our best model. Half of them is based on Andres code (forked before adding any restrictions on usage) and half on its PyTorch implementation:</p>\n\n<ul>\n<li>984_densenet201_antorsaegen_29_0.98624</li>\n<li>976_densenet201_antorsaegen_62_0.98271</li>\n<li>977_resnet50_antorsaegen_119_val_0.9815</li>\n<li>976_DenseNet201_do0.3_doc0.0_avg-epoch072-val_acc0.981250</li>\n<li>967_InceptionResNetV2_do0.1_avg-epoch154-val_acc0.965625</li>\n<li>962_Xception_do0.3_avg-epoch079-val_acc0.991667 (leaky validation, pls ignore)</li>\n</ul>\n\n<p>All of the models had a full-size crop (512) + TTA 8. Choosing this crop was our biggest flaw since smaller crops allow faster training, more TTA options and produce almost the same accuracy.  </p>\n\n<h2>Data</h2>\n\n<p>In total, we collected 300 Gb of photos from Flickr and Yandex.Foto. But due to the huge crop size, we could not fully utilize the entire data set and used only 20,000 photos for training. The rest 50,000 was used for blending. \nData was filtered by resolution and camera type, also all photos with non-default software in Exif  (i.e. processed somehow) and with poor quality (lower than 95 as identified by ImageMagick) were excluded.</p>\n\n<h2>Hardware</h2>\n\n<p>Nothing special for a team with five members: 5 x 1080ti + 1070</p>\n\n<h2>Submissions</h2>\n\n<p>We had 2 submission strategies:</p>\n\n<ul>\n<li>Averaging all TTA prediction from all models by power mean - powers of 1,2,4 as a parameter. Additionally, we explicitly made all classes equally distributed (132 manip and 132 unalt photos for each class) by using Hungarian algorithm on probabilities. As expected, this resulted in a huge LB overfit: 0.991 score on public and 0.985 on private.</li>\n<li>Blending models by using hold-out predictions. There were 52 models with slightly different parameters and equal weights: XGBoost (20) + LightGBM (20) + Keras (12). This gave us 0.986 score on public and 0.989 on private. No class equalization was used with this approach except some weight tuning for Nexus 5 during training (all models had a smaller number of predictions for this phone). </li>\n</ul>\n\n<h2>Key takeaways</h2>\n\n<ul>\n<li>Collect as much data as possible but do not forget to clean it (GIGO)</li>\n<li>Try smaller crops first and all architecture types</li>\n<li>LB probing is evil, trust your CV</li>\n</ul>",
      "rawMarkdown": "## Models ##\nHere is the list of our best model. Half of them is based on Andres code (forked before adding any restrictions on usage) and half on its PyTorch implementation:\n\n- 984_densenet201_antorsaegen_29_0.98624\n- 976_densenet201_antorsaegen_62_0.98271\n- 977_resnet50_antorsaegen_119_val_0.9815\n- 976_DenseNet201_do0.3_doc0.0_avg-epoch072-val_acc0.981250\n- 967_InceptionResNetV2_do0.1_avg-epoch154-val_acc0.965625\n- 962_Xception_do0.3_avg-epoch079-val_acc0.991667 (leaky validation, pls ignore)\n\nAll of the models had a full-size crop (512) + TTA 8. Choosing this crop was our biggest flaw since smaller crops allow faster training, more TTA options and produce almost the same accuracy.  \n\n## Data ##\nIn total, we collected 300 Gb of photos from Flickr and Yandex.Foto. But due to the huge crop size, we could not fully utilize the entire data set and used only 20,000 photos for training. The rest 50,000 was used for blending. \nData was filtered by resolution and camera type, also all photos with non-default software in Exif  (i.e. processed somehow) and with poor quality (lower than 95 as identified by ImageMagick) were excluded.\n\n## Hardware ##\nNothing special for a team with five members: 5 x 1080ti + 1070\n\n## Submissions ##\nWe had 2 submission strategies:\n\n - Averaging all TTA prediction from all models by power mean - powers of 1,2,4 as a parameter. Additionally, we explicitly made all classes equally distributed (132 manip and 132 unalt photos for each class) by using Hungarian algorithm on probabilities. As expected, this resulted in a huge LB overfit: 0.991 score on public and 0.985 on private.\n - Blending models by using hold-out predictions. There were 52 models with slightly different parameters and equal weights: XGBoost (20) + LightGBM (20) + Keras (12). This gave us 0.986 score on public and 0.989 on private. No class equalization was used with this approach except some weight tuning for Nexus 5 during training (all models had a smaller number of predictions for this phone). \n\n## Key takeaways ##\n\n- Collect as much data as possible but do not forget to clean it (GIGO)\n- Try smaller crops first and all architecture types\n- LB probing is evil, trust your CV",
      "votes": null
    },
    {
      "id": "280437",
      "postDate": "02/10/2018 01:27:38",
      "content": "<p>I didn't know what blending was. Here's my understanding, please correct me:</p>\n\n<p>Instead of averaging the predictions of the models in your ensemble, which is just a way of combing the predictions, <strong>use machine learning to learn an optimal way to combine the predictions</strong>.</p>\n\n<p>In blending, you use your validation set to train your combiner. This training doesn't take much time, since you're working with probabilities instead of images. This faster training time allowed Pavel's team to use the extra 50,000 images to train their combiner, despite them not having enough time to use those images to train their base models. To speed up this training, you can cache the base models' predictions for those 50,000 images, since the predictions won't change, since you'll be training the combiner and not the base models.</p>\n\n<p>Where I learned this: <a href=\"https://mlwave.com/kaggle-ensembling-guide/\">https://mlwave.com/kaggle-ensembling-guide/</a></p>\n\n<p>A rule I'm going to try: \"If you don't have time to use newly released data, use them for blending.\"</p>",
      "rawMarkdown": "I didn't know what blending was. Here's my understanding, please correct me:\n\nInstead of averaging the predictions of the models in your ensemble, which is just a way of combing the predictions, **use machine learning to learn an optimal way to combine the predictions**.\n\nIn blending, you use your validation set to train your combiner. This training doesn't take much time, since you're working with probabilities instead of images. This faster training time allowed Pavel's team to use the extra 50,000 images to train their combiner, despite them not having enough time to use those images to train their base models. To speed up this training, you can cache the base models' predictions for those 50,000 images, since the predictions won't change, since you'll be training the combiner and not the base models.\n\nWhere I learned this: https://mlwave.com/kaggle-ensembling-guide/\n\nA rule I'm going to try: \"If you don't have time to use newly released data, use them for blending.\"",
      "votes": null
    },
    {
      "id": "280447",
      "postDate": "02/10/2018 02:10:00",
      "content": "<p>We were lucky that predictions did not take a lot of time to make in this competition. It is not always the case though.</p>",
      "rawMarkdown": "We were lucky that predictions did not take a lot of time to make in this competition. It is not always the case though.",
      "votes": null
    },
    {
      "id": "280448",
      "postDate": "02/10/2018 02:24:46",
      "content": "<p>Did the fast-prediction time allow you to use TTA when blending?</p>",
      "rawMarkdown": "Did the fast-prediction time allow you to use TTA when blending?",
      "votes": null
    },
    {
      "id": "280460",
      "postDate": "02/10/2018 03:09:49",
      "content": "<p>Dear Pleskov,\n          Thanks for your sharing. It is help for me. About the architecture of the models, did you modify it?  such as adding the high-pass filter into first layer?  By the way,  what is the meaning of 'antorsaegen'?</p>",
      "rawMarkdown": "Dear Pleskov,\n          Thanks for your sharing. It is help for me. About the architecture of the models, did you modify it?  such as adding the high-pass filter into first layer?  By the way,  what is the meaning of 'antorsaegen'?",
      "votes": null
    },
    {
      "id": "280462",
      "postDate": "02/10/2018 03:42:26",
      "content": "<p>Congratulations !!! </p>",
      "rawMarkdown": "Congratulations !!!",
      "votes": null
    },
    {
      "id": "280463",
      "postDate": "02/10/2018 03:44:42",
      "content": "<blockquote>\n  <p>what is the meaning of 'antorsaegen'?</p>\n</blockquote>\n\n<p>This is a reference to Ivan's PyTorch implementation of Andres's code: <a href=\"https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/train.py#L71\">https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/train.py#L71</a></p>\n\n<p>Andres's code: <a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">https://github.com/antorsae/sp-society-camera-model-identification</a></p>",
      "rawMarkdown": "&gt; what is the meaning of 'antorsaegen'?\n\nThis is a reference to Ivan's PyTorch implementation of Andres's code: https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/train.py#L71\n\nAndres's code: https://github.com/antorsae/sp-society-camera-model-identification",
      "votes": null
    },
    {
      "id": "280488",
      "postDate": "02/10/2018 05:06:35",
      "content": "<blockquote>\n  <p>XGBoost (20) + LightGBM (20) + Keras (12)</p>\n</blockquote>\n\n<p>Wait a second, how did you use XGBoost for an image classification problem? <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49366\">Bojan's team (10th place)</a> used XGBoost on extracted features related to noise patterns, but you didn't mention noise patterns. Did you use XGBoost on CNN activations?</p>",
      "rawMarkdown": "&gt; XGBoost (20) + LightGBM (20) + Keras (12)\n\nWait a second, how did you use XGBoost for an image classification problem? [Bojan's team (10th place)][1] used XGBoost on extracted features related to noise patterns, but you didn't mention noise patterns. Did you use XGBoost on CNN activations?\n\n\n  [1]: https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49366",
      "votes": null
    },
    {
      "id": "280496",
      "postDate": "02/10/2018 05:40:38",
      "content": "<p>We used it for model predictions on hold-out set. Each model predicts probabilities of 10 classes and they become an input for XGBoost and etc.</p>",
      "rawMarkdown": "We used it for model predictions on hold-out set. Each model predicts probabilities of 10 classes and they become an input for XGBoost and etc.",
      "votes": null
    },
    {
      "id": "280499",
      "postDate": "02/10/2018 05:42:57",
      "content": "<p>For XGBoost it is critical to have a lot of data points, so predicting time matters. The faster you predict the more test time augmentations you can produce and fit models better. </p>",
      "rawMarkdown": "For XGBoost it is critical to have a lot of data points, so predicting time matters. The faster you predict the more test time augmentations you can produce and fit models better.",
      "votes": null
    },
    {
      "id": "280503",
      "postDate": "02/10/2018 05:57:14",
      "content": "<p>Oh, okay. You used it as the stacker/blender/combiner/level-2-generalizer/meta-model.</p>\n\n<p>Not yet sure how you produced 20 XGBoost models, though. Maybe you used different subsets of the base models, or used different subsets of the hold-out set. Either way, it seems that you used a meta-stacker/meta-blender/meta-combiner/level-3-generalizer/meta-meta-model to combine these XGBoost models.</p>\n\n<p>This is all very cool. I can see now why making one's pipeline efficient is a good thing to focus on. There are many ideas to test.</p>",
      "rawMarkdown": "Oh, okay. You used it as the stacker/blender/combiner/level-2-generalizer/meta-model.\n\nNot yet sure how you produced 20 XGBoost models, though. Maybe you used different subsets of the base models, or used different subsets of the hold-out set. Either way, it seems that you used a meta-stacker/meta-blender/meta-combiner/level-3-generalizer/meta-meta-model to combine these XGBoost models.\n\nThis is all very cool. I can see now why making one's pipeline efficient is a good thing to focus on. There are many ideas to test.",
      "votes": null
    },
    {
      "id": "280509",
      "postDate": "02/10/2018 06:26:55",
      "content": "<p>Below is a train of thought that leads to a hypothesis. I invite readers to discuss this topic with me:</p>\n\n<hr>\n\n<p>Ah so instead of using TTA on the models that feed into XGBoost, one should use simple data augmentation, if one is pursuing more data points for XGBoost. Using TTA on the models would average multiple data points together before they reach XGBoost, leaving one with fewer data points.</p>\n\n<p>However, at test time XGBoost will receive TTA-based predictions, not individual predictions based on individual data augmentations. But maybe the difference in training conditions and testing conditions is worth giving XGBoost additional data points. Wait, maybe you can do TTA with respect to the output of XGBoost rather than the output of the base models. This way you could get the extra data points from the individual data augmentations while still getting the benefit of TTA.</p>\n\n<hr>\n\n<p>Hypothesis: When using TTA with an XGBoost-blended set of models, one should average the predictions coming from XGBoost, not the base models. </p>\n\n<p>Elaboration: This allows one to increase the set of samples XGBoost is trained on, via data augmentation, while preserving the benefit of TTA with respect to those augmentations. The alternative is to average the predictions coming from the base models, which would unfortunately decrease the set of samples XGBoost trains on.</p>",
      "rawMarkdown": "Below is a train of thought that leads to a hypothesis. I invite readers to discuss this topic with me:\n\n----------\n\nAh so instead of using TTA on the models that feed into XGBoost, one should use simple data augmentation, if one is pursuing more data points for XGBoost. Using TTA on the models would average multiple data points together before they reach XGBoost, leaving one with fewer data points.\n\nHowever, at test time XGBoost will receive TTA-based predictions, not individual predictions based on individual data augmentations. But maybe the difference in training conditions and testing conditions is worth giving XGBoost additional data points. Wait, maybe you can do TTA with respect to the output of XGBoost rather than the output of the base models. This way you could get the extra data points from the individual data augmentations while still getting the benefit of TTA.\n\n---\n\nHypothesis: When using TTA with an XGBoost-blended set of models, one should average the predictions coming from XGBoost, not the base models. \n\nElaboration: This allows one to increase the set of samples XGBoost is trained on, via data augmentation, while preserving the benefit of TTA with respect to those augmentations. The alternative is to average the predictions coming from the base models, which would unfortunately decrease the set of samples XGBoost trains on.",
      "votes": null
    },
    {
      "id": "280549",
      "postDate": "02/10/2018 08:10:38",
      "content": "<p>Pavel and team, congratulations!</p>\n\n<p>How did you get the images by camera model from Yandex.Foto? Is there an API?</p>\n\n<p>The flickr API doesn't have a specific API call for camera model but I found an undocumented way :-).</p>\n\n<p>Also, my dataset explicitly queried images with CC license, was it possible with Yandex.Foto?</p>\n\n<p>Thanks! </p>",
      "rawMarkdown": "Pavel and team, congratulations!\n\nHow did you get the images by camera model from Yandex.Foto? Is there an API?\n\nThe flickr API doesn't have a specific API call for camera model but I found an undocumented way :-).\n\nAlso, my dataset explicitly queried images with CC license, was it possible with Yandex.Foto?\n\nThanks!",
      "votes": null
    },
    {
      "id": "280630",
      "postDate": "02/10/2018 14:05:55",
      "content": "<p>To produce a bunch of XGBoost models we used different combinations of training parameters, i.e. number of folds, optimizer, patience</p>",
      "rawMarkdown": "To produce a bunch of XGBoost models we used different combinations of training parameters, i.e. number of folds, optimizer, patience",
      "votes": null
    },
    {
      "id": "280631",
      "postDate": "02/10/2018 14:07:27",
      "content": "<p>We did not touch architecture at all, except for dropout rate</p>",
      "rawMarkdown": "We did not touch architecture at all, except for dropout rate",
      "votes": null
    },
    {
      "id": "280632",
      "postDate": "02/10/2018 14:10:13",
      "content": "<p>There was slightly different run parameters for each XGBoost/LightGBM/Keras run: Depth, LR, Class weights, neural net structure etc. </p>",
      "rawMarkdown": "There was slightly different run parameters for each XGBoost/LightGBM/Keras run: Depth, LR, Class weights, neural net structure etc.",
      "votes": null
    },
    {
      "id": "280633",
      "postDate": "02/10/2018 14:16:22",
      "content": "<p>We used <a href=\"http://www.seleniumhq.org/\">http://www.seleniumhq.org/</a> for scraping image preview links, then extracted links to the original files from it</p>",
      "rawMarkdown": "We used http://www.seleniumhq.org/ for scraping image preview links, then extracted links to the original files from it",
      "votes": null
    },
    {
      "id": "280716",
      "postDate": "02/10/2018 18:28:00",
      "content": "<p>great stuff </p>",
      "rawMarkdown": "great stuff",
      "votes": null
    },
    {
      "id": "280723",
      "postDate": "02/10/2018 19:02:10",
      "content": "<p>Good Work</p>",
      "rawMarkdown": "Good Work",
      "votes": null
    },
    {
      "id": "280841",
      "postDate": "02/11/2018 07:04:21",
      "content": "<p>Got it. Thanks very much</p>",
      "rawMarkdown": "Got it. Thanks very much",
      "votes": null
    },
    {
      "id": "281251",
      "postDate": "02/12/2018 10:09:20",
      "content": "<blockquote>\n  <p>LB probing is evil, trust your CV</p>\n</blockquote>\n\n<p>Love it!  congrats for your win.</p>",
      "rawMarkdown": "&gt; LB probing is evil, trust your CV\n\nLove it!  congrats for your win.",
      "votes": null
    },
    {
      "id": "282088",
      "postDate": "02/13/2018 16:01:36",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "282398",
      "postDate": "02/13/2018 23:44:05",
      "content": "<p>Great work Pavel, congrats</p>",
      "rawMarkdown": "Great work Pavel, congrats",
      "votes": null
    },
    {
      "id": "282436",
      "postDate": "02/14/2018 01:29:50",
      "content": "<p>Definitely a good submission &amp; great notes to read up on!\nCongratulations!</p>",
      "rawMarkdown": "Definitely a good submission &amp; great notes to read up on!\nCongratulations!",
      "votes": null
    },
    {
      "id": "283378",
      "postDate": "02/15/2018 11:55:41",
      "content": "<p>Congratulations to the team !</p>",
      "rawMarkdown": "Congratulations to the team !",
      "votes": null
    },
    {
      "id": "284600",
      "postDate": "02/17/2018 18:17:31",
      "content": "<p>Congratulations !!!</p>",
      "rawMarkdown": "Congratulations !!!",
      "votes": null
    },
    {
      "id": "284639",
      "postDate": "02/17/2018 20:36:55",
      "content": "<p>wow nice pytorch usage</p>",
      "rawMarkdown": "wow nice pytorch usage",
      "votes": null
    },
    {
      "id": "284643",
      "postDate": "02/17/2018 20:44:16",
      "content": "<p>Congrats!!</p>",
      "rawMarkdown": "Congrats!!",
      "votes": null
    },
    {
      "id": "284884",
      "postDate": "02/18/2018 17:50:02",
      "content": "<p>congrats</p>",
      "rawMarkdown": "congrats",
      "votes": null
    },
    {
      "id": "328067",
      "postDate": "05/13/2018 09:13:19",
      "content": "<p>We have published a video with 1st and 2nd places solutions with English subtitles from Moscow ML trainings meetup. \n<a href=\"https://youtu.be/ETh8bJ_xKGA\">https://youtu.be/ETh8bJ_xKGA</a></p>",
      "rawMarkdown": "We have published a video with 1st and 2nd places solutions with English subtitles from Moscow ML trainings meetup. \nhttps://youtu.be/ETh8bJ_xKGA",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 280437,
      "author_name": "kleinsmith",
      "author_url": "",
      "post_date": "02/10/2018 01:27:38",
      "content": "<p>I didn't know what blending was. Here's my understanding, please correct me:</p>\n\n<p>Instead of averaging the predictions of the models in your ensemble, which is just a way of combing the predictions, <strong>use machine learning to learn an optimal way to combine the predictions</strong>.</p>\n\n<p>In blending, you use your validation set to train your combiner. This training doesn't take much time, since you're working with probabilities instead of images. This faster training time allowed Pavel's team to use the extra 50,000 images to train their combiner, despite them not having enough time to use those images to train their base models. To speed up this training, you can cache the base models' predictions for those 50,000 images, since the predictions won't change, since you'll be training the combiner and not the base models.</p>\n\n<p>Where I learned this: <a href=\"https://mlwave.com/kaggle-ensembling-guide/\">https://mlwave.com/kaggle-ensembling-guide/</a></p>\n\n<p>A rule I'm going to try: \"If you don't have time to use newly released data, use them for blending.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 280447,
          "author_name": "",
          "author_url": "",
          "post_date": "02/10/2018 02:10:00",
          "content": "<p>We were lucky that predictions did not take a lot of time to make in this competition. It is not always the case though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280448,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "02/10/2018 02:24:46",
          "content": "<p>Did the fast-prediction time allow you to use TTA when blending?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280499,
          "author_name": "",
          "author_url": "",
          "post_date": "02/10/2018 05:42:57",
          "content": "<p>For XGBoost it is critical to have a lot of data points, so predicting time matters. The faster you predict the more test time augmentations you can produce and fit models better. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280509,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "02/10/2018 06:26:55",
          "content": "<p>Below is a train of thought that leads to a hypothesis. I invite readers to discuss this topic with me:</p>\n\n<hr>\n\n<p>Ah so instead of using TTA on the models that feed into XGBoost, one should use simple data augmentation, if one is pursuing more data points for XGBoost. Using TTA on the models would average multiple data points together before they reach XGBoost, leaving one with fewer data points.</p>\n\n<p>However, at test time XGBoost will receive TTA-based predictions, not individual predictions based on individual data augmentations. But maybe the difference in training conditions and testing conditions is worth giving XGBoost additional data points. Wait, maybe you can do TTA with respect to the output of XGBoost rather than the output of the base models. This way you could get the extra data points from the individual data augmentations while still getting the benefit of TTA.</p>\n\n<hr>\n\n<p>Hypothesis: When using TTA with an XGBoost-blended set of models, one should average the predictions coming from XGBoost, not the base models. </p>\n\n<p>Elaboration: This allows one to increase the set of samples XGBoost is trained on, via data augmentation, while preserving the benefit of TTA with respect to those augmentations. The alternative is to average the predictions coming from the base models, which would unfortunately decrease the set of samples XGBoost trains on.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280460,
      "author_name": "meprobjtu",
      "author_url": "",
      "post_date": "02/10/2018 03:09:49",
      "content": "<p>Dear Pleskov,\n          Thanks for your sharing. It is help for me. About the architecture of the models, did you modify it?  such as adding the high-pass filter into first layer?  By the way,  what is the meaning of 'antorsaegen'?</p>",
      "votes": null,
      "replies": [
        {
          "id": 280463,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "02/10/2018 03:44:42",
          "content": "<blockquote>\n  <p>what is the meaning of 'antorsaegen'?</p>\n</blockquote>\n\n<p>This is a reference to Ivan's PyTorch implementation of Andres's code: <a href=\"https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/train.py#L71\">https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/train.py#L71</a></p>\n\n<p>Andres's code: <a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">https://github.com/antorsae/sp-society-camera-model-identification</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280631,
          "author_name": "",
          "author_url": "",
          "post_date": "02/10/2018 14:07:27",
          "content": "<p>We did not touch architecture at all, except for dropout rate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280841,
          "author_name": "meprobjtu",
          "author_url": "",
          "post_date": "02/11/2018 07:04:21",
          "content": "<p>Got it. Thanks very much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280462,
      "author_name": "deeraen",
      "author_url": "",
      "post_date": "02/10/2018 03:42:26",
      "content": "<p>Congratulations !!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280488,
      "author_name": "kleinsmith",
      "author_url": "",
      "post_date": "02/10/2018 05:06:35",
      "content": "<blockquote>\n  <p>XGBoost (20) + LightGBM (20) + Keras (12)</p>\n</blockquote>\n\n<p>Wait a second, how did you use XGBoost for an image classification problem? <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49366\">Bojan's team (10th place)</a> used XGBoost on extracted features related to noise patterns, but you didn't mention noise patterns. Did you use XGBoost on CNN activations?</p>",
      "votes": null,
      "replies": [
        {
          "id": 280496,
          "author_name": "",
          "author_url": "",
          "post_date": "02/10/2018 05:40:38",
          "content": "<p>We used it for model predictions on hold-out set. Each model predicts probabilities of 10 classes and they become an input for XGBoost and etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280503,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "02/10/2018 05:57:14",
          "content": "<p>Oh, okay. You used it as the stacker/blender/combiner/level-2-generalizer/meta-model.</p>\n\n<p>Not yet sure how you produced 20 XGBoost models, though. Maybe you used different subsets of the base models, or used different subsets of the hold-out set. Either way, it seems that you used a meta-stacker/meta-blender/meta-combiner/level-3-generalizer/meta-meta-model to combine these XGBoost models.</p>\n\n<p>This is all very cool. I can see now why making one's pipeline efficient is a good thing to focus on. There are many ideas to test.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280630,
          "author_name": "",
          "author_url": "",
          "post_date": "02/10/2018 14:05:55",
          "content": "<p>To produce a bunch of XGBoost models we used different combinations of training parameters, i.e. number of folds, optimizer, patience</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280632,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "02/10/2018 14:10:13",
          "content": "<p>There was slightly different run parameters for each XGBoost/LightGBM/Keras run: Depth, LR, Class weights, neural net structure etc. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280549,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/10/2018 08:10:38",
      "content": "<p>Pavel and team, congratulations!</p>\n\n<p>How did you get the images by camera model from Yandex.Foto? Is there an API?</p>\n\n<p>The flickr API doesn't have a specific API call for camera model but I found an undocumented way :-).</p>\n\n<p>Also, my dataset explicitly queried images with CC license, was it possible with Yandex.Foto?</p>\n\n<p>Thanks! </p>",
      "votes": null,
      "replies": [
        {
          "id": 280633,
          "author_name": "",
          "author_url": "",
          "post_date": "02/10/2018 14:16:22",
          "content": "<p>We used <a href=\"http://www.seleniumhq.org/\">http://www.seleniumhq.org/</a> for scraping image preview links, then extracted links to the original files from it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280716,
      "author_name": "caynosadler",
      "author_url": "",
      "post_date": "02/10/2018 18:28:00",
      "content": "<p>great stuff </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280723,
      "author_name": "penedon",
      "author_url": "",
      "post_date": "02/10/2018 19:02:10",
      "content": "<p>Good Work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 281251,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/12/2018 10:09:20",
      "content": "<blockquote>\n  <p>LB probing is evil, trust your CV</p>\n</blockquote>\n\n<p>Love it!  congrats for your win.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 282088,
      "author_name": "kasplat",
      "author_url": "",
      "post_date": "02/13/2018 16:01:36",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 282398,
      "author_name": "cchahuas",
      "author_url": "",
      "post_date": "02/13/2018 23:44:05",
      "content": "<p>Great work Pavel, congrats</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 282436,
      "author_name": "ktdevmasters",
      "author_url": "",
      "post_date": "02/14/2018 01:29:50",
      "content": "<p>Definitely a good submission &amp; great notes to read up on!\nCongratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 283378,
      "author_name": "faizunnabi",
      "author_url": "",
      "post_date": "02/15/2018 11:55:41",
      "content": "<p>Congratulations to the team !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 284600,
      "author_name": "gomes555",
      "author_url": "",
      "post_date": "02/17/2018 18:17:31",
      "content": "<p>Congratulations !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 284639,
      "author_name": "incognito124",
      "author_url": "",
      "post_date": "02/17/2018 20:36:55",
      "content": "<p>wow nice pytorch usage</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 284643,
      "author_name": "kps143",
      "author_url": "",
      "post_date": "02/17/2018 20:44:16",
      "content": "<p>Congrats!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 284884,
      "author_name": "jatinder8889",
      "author_url": "",
      "post_date": "02/18/2018 17:50:02",
      "content": "<p>congrats</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 328067,
      "author_name": "emilkayumov",
      "author_url": "",
      "post_date": "05/13/2018 09:13:19",
      "content": "<p>We have published a video with 1st and 2nd places solutions with English subtitles from Moscow ML trainings meetup. \n<a href=\"https://youtu.be/ETh8bJ_xKGA\">https://youtu.be/ETh8bJ_xKGA</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "280404": "## Models ##\nHere is the list of our best model. Half of them is based on Andres code (forked before adding any restrictions on usage) and half on its PyTorch implementation:\n\n- 984_densenet201_antorsaegen_29_0.98624\n- 976_densenet201_antorsaegen_62_0.98271\n- 977_resnet50_antorsaegen_119_val_0.9815\n- 976_DenseNet201_do0.3_doc0.0_avg-epoch072-val_acc0.981250\n- 967_InceptionResNetV2_do0.1_avg-epoch154-val_acc0.965625\n- 962_Xception_do0.3_avg-epoch079-val_acc0.991667 (leaky validation, pls ignore)\n\nAll of the models had a full-size crop (512) + TTA 8. Choosing this crop was our biggest flaw since smaller crops allow faster training, more TTA options and produce almost the same accuracy.  \n\n## Data ##\nIn total, we collected 300 Gb of photos from Flickr and Yandex.Foto. But due to the huge crop size, we could not fully utilize the entire data set and used only 20,000 photos for training. The rest 50,000 was used for blending. \nData was filtered by resolution and camera type, also all photos with non-default software in Exif  (i.e. processed somehow) and with poor quality (lower than 95 as identified by ImageMagick) were excluded.\n\n## Hardware ##\nNothing special for a team with five members: 5 x 1080ti + 1070\n\n## Submissions ##\nWe had 2 submission strategies:\n\n - Averaging all TTA prediction from all models by power mean - powers of 1,2,4 as a parameter. Additionally, we explicitly made all classes equally distributed (132 manip and 132 unalt photos for each class) by using Hungarian algorithm on probabilities. As expected, this resulted in a huge LB overfit: 0.991 score on public and 0.985 on private.\n - Blending models by using hold-out predictions. There were 52 models with slightly different parameters and equal weights: XGBoost (20) + LightGBM (20) + Keras (12). This gave us 0.986 score on public and 0.989 on private. No class equalization was used with this approach except some weight tuning for Nexus 5 during training (all models had a smaller number of predictions for this phone). \n\n## Key takeaways ##\n\n- Collect as much data as possible but do not forget to clean it (GIGO)\n- Try smaller crops first and all architecture types\n- LB probing is evil, trust your CV",
    "280437": "I didn't know what blending was. Here's my understanding, please correct me:\n\nInstead of averaging the predictions of the models in your ensemble, which is just a way of combing the predictions, **use machine learning to learn an optimal way to combine the predictions**.\n\nIn blending, you use your validation set to train your combiner. This training doesn't take much time, since you're working with probabilities instead of images. This faster training time allowed Pavel's team to use the extra 50,000 images to train their combiner, despite them not having enough time to use those images to train their base models. To speed up this training, you can cache the base models' predictions for those 50,000 images, since the predictions won't change, since you'll be training the combiner and not the base models.\n\nWhere I learned this: https://mlwave.com/kaggle-ensembling-guide/\n\nA rule I'm going to try: \"If you don't have time to use newly released data, use them for blending.\"",
    "280447": "We were lucky that predictions did not take a lot of time to make in this competition. It is not always the case though.",
    "280448": "Did the fast-prediction time allow you to use TTA when blending?",
    "280460": "Dear Pleskov,\n          Thanks for your sharing. It is help for me. About the architecture of the models, did you modify it?  such as adding the high-pass filter into first layer?  By the way,  what is the meaning of 'antorsaegen'?",
    "280462": "Congratulations !!!",
    "280463": "&gt; what is the meaning of 'antorsaegen'?\n\nThis is a reference to Ivan's PyTorch implementation of Andres's code: https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/train.py#L71\n\nAndres's code: https://github.com/antorsae/sp-society-camera-model-identification",
    "280488": "&gt; XGBoost (20) + LightGBM (20) + Keras (12)\n\nWait a second, how did you use XGBoost for an image classification problem? [Bojan's team (10th place)][1] used XGBoost on extracted features related to noise patterns, but you didn't mention noise patterns. Did you use XGBoost on CNN activations?\n\n\n  [1]: https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/49366",
    "280496": "We used it for model predictions on hold-out set. Each model predicts probabilities of 10 classes and they become an input for XGBoost and etc.",
    "280499": "For XGBoost it is critical to have a lot of data points, so predicting time matters. The faster you predict the more test time augmentations you can produce and fit models better.",
    "280503": "Oh, okay. You used it as the stacker/blender/combiner/level-2-generalizer/meta-model.\n\nNot yet sure how you produced 20 XGBoost models, though. Maybe you used different subsets of the base models, or used different subsets of the hold-out set. Either way, it seems that you used a meta-stacker/meta-blender/meta-combiner/level-3-generalizer/meta-meta-model to combine these XGBoost models.\n\nThis is all very cool. I can see now why making one's pipeline efficient is a good thing to focus on. There are many ideas to test.",
    "280509": "Below is a train of thought that leads to a hypothesis. I invite readers to discuss this topic with me:\n\n----------\n\nAh so instead of using TTA on the models that feed into XGBoost, one should use simple data augmentation, if one is pursuing more data points for XGBoost. Using TTA on the models would average multiple data points together before they reach XGBoost, leaving one with fewer data points.\n\nHowever, at test time XGBoost will receive TTA-based predictions, not individual predictions based on individual data augmentations. But maybe the difference in training conditions and testing conditions is worth giving XGBoost additional data points. Wait, maybe you can do TTA with respect to the output of XGBoost rather than the output of the base models. This way you could get the extra data points from the individual data augmentations while still getting the benefit of TTA.\n\n---\n\nHypothesis: When using TTA with an XGBoost-blended set of models, one should average the predictions coming from XGBoost, not the base models. \n\nElaboration: This allows one to increase the set of samples XGBoost is trained on, via data augmentation, while preserving the benefit of TTA with respect to those augmentations. The alternative is to average the predictions coming from the base models, which would unfortunately decrease the set of samples XGBoost trains on.",
    "280549": "Pavel and team, congratulations!\n\nHow did you get the images by camera model from Yandex.Foto? Is there an API?\n\nThe flickr API doesn't have a specific API call for camera model but I found an undocumented way :-).\n\nAlso, my dataset explicitly queried images with CC license, was it possible with Yandex.Foto?\n\nThanks!",
    "280630": "To produce a bunch of XGBoost models we used different combinations of training parameters, i.e. number of folds, optimizer, patience",
    "280631": "We did not touch architecture at all, except for dropout rate",
    "280632": "There was slightly different run parameters for each XGBoost/LightGBM/Keras run: Depth, LR, Class weights, neural net structure etc.",
    "280633": "We used http://www.seleniumhq.org/ for scraping image preview links, then extracted links to the original files from it",
    "280716": "great stuff",
    "280723": "Good Work",
    "280841": "Got it. Thanks very much",
    "281251": "&gt; LB probing is evil, trust your CV\n\nLove it!  congrats for your win.",
    "282088": "Congratulations!",
    "282398": "Great work Pavel, congrats",
    "282436": "Definitely a good submission &amp; great notes to read up on!\nCongratulations!",
    "283378": "Congratulations to the team !",
    "284600": "Congratulations !!!",
    "284639": "wow nice pytorch usage",
    "284643": "Congrats!!",
    "284884": "congrats",
    "328067": "We have published a video with 1st and 2nd places solutions with English subtitles from Moscow ML trainings meetup. \nhttps://youtu.be/ETh8bJ_xKGA"
  },
  "source": "meta"
}