{
  "id": 49366,
  "title": "10th Place Solution",
  "url": "/competitions/sp-society-camera-model-identification/writeups/make-ensemble-great-again-10th-place-solution",
  "author_name": "",
  "post_date": "2018-02-09T21:47:38.939963200Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Here is a brief overview of our modeling and solution. Please feel free to ask any questions.</p>\n\n<ol>\n<li><p>We only used the data provided by Kaggle and Gleb´s links. We had some concerns about the Gleb's dataset license of use for some of the images, to be honest, but most of them seem to be good. </p></li>\n<li><p>We first tried some CNN models from scratch, but discovered that the well known pretrained models converge faster and gave better accuracy. Our best single model was a DenseNet201 trained  on patches of size 336x336 scoring 0.963/0.967 (public/private LB). As the accuracy was pretty high, we decide to give a try to pseudo-labeled images from the testset using the predictions from this model.</p></li>\n<li><p>We applied a threshold of p&gt;=0.999 and used only unaltered images. This gives us 1116 new images to train with.</p></li>\n<li><p>The same DenseNet201 model, trained now with this extra dataset, scored 0.975/0.973. After intense TTA, it gave us 0.980/0.975.</p></li>\n<li><p>We trained a few more models (another DenseNet201, InceptionResnetV2, etc.) , but we couldn't improve much by just blending their predictions (and also being worried about overfitting to the LB), we decide to try stacking predictions from different models, and to train a second layer model with them.</p></li>\n<li><p>Since we only had about a week to train them, and since we didn't have too many GPU resources, we decided to do stacking with just 2-folds.</p></li>\n<li><p>While we were training NN models, we tried to extract image features to train a non-NN model. We basically followed the ideas from this paper:</p></li>\n</ol>\n\n<p><a href=\"https://www.semanticscholar.org/paper/Using-sensor-pattern-noise-for-camera-model-identi-Filler-Fridrich/e9d9d89d81e49c8fcbb8b2cc960efa4c91434319\">\"Using sensor pattern noise for camera model identification\" Thomas Filler, Jessica Fridrich, Miroslav Goljan</a></p>\n\n<ol>\n<li><p>We extracted approx. 3000 features from image linear pattern and noise pattern. We trained XGB, LightGBM and MultilayerPerceptron after applying PCA, scoring 0.898/0.881. (We found out that PCA reduced the training times of the models, but the raw dataset was trainable as well, achieving similar performance, if not better.)</p></li>\n<li><p>Our final model is an XGB model of stacked predictions from 9 L1 models:\n    - From images: DenseNet201, VGG16, Xception, DenseNet121, InceptionV3, ResNet50 (we tried others but we discarded them)\n    - From extracted features: XGB, LightGBM, MultilayerPerceptron</p></li>\n<li><p>This model (A13) scored 0.9877/0.9866. We trained the same model but excluding pseudolabeled data. We thought that they could introduce a strong bias toward these images, and after all, they were already used for all ours L1 models. This second model (A13a) scored  0.9798/0.9847, and we made the random decision of select this submission for the final, finishing 10th instead of 5th. (Our 2nd final submissions was a optimized weighted average of several submissions scoring 0.9885/0.9835, but we weren't very confident about it).</p></li>\n</ol>",
  "messages": [
    {
      "id": "280391",
      "postDate": "02/09/2018 21:47:38",
      "content": "<p>Here is a brief overview of our modeling and solution. Please feel free to ask any questions.</p>\n\n<ol>\n<li><p>We only used the data provided by Kaggle and Gleb´s links. We had some concerns about the Gleb's dataset license of use for some of the images, to be honest, but most of them seem to be good. </p></li>\n<li><p>We first tried some CNN models from scratch, but discovered that the well known pretrained models converge faster and gave better accuracy. Our best single model was a DenseNet201 trained  on patches of size 336x336 scoring 0.963/0.967 (public/private LB). As the accuracy was pretty high, we decide to give a try to pseudo-labeled images from the testset using the predictions from this model.</p></li>\n<li><p>We applied a threshold of p&gt;=0.999 and used only unaltered images. This gives us 1116 new images to train with.</p></li>\n<li><p>The same DenseNet201 model, trained now with this extra dataset, scored 0.975/0.973. After intense TTA, it gave us 0.980/0.975.</p></li>\n<li><p>We trained a few more models (another DenseNet201, InceptionResnetV2, etc.) , but we couldn't improve much by just blending their predictions (and also being worried about overfitting to the LB), we decide to try stacking predictions from different models, and to train a second layer model with them.</p></li>\n<li><p>Since we only had about a week to train them, and since we didn't have too many GPU resources, we decided to do stacking with just 2-folds.</p></li>\n<li><p>While we were training NN models, we tried to extract image features to train a non-NN model. We basically followed the ideas from this paper:</p></li>\n</ol>\n\n<p><a href=\"https://www.semanticscholar.org/paper/Using-sensor-pattern-noise-for-camera-model-identi-Filler-Fridrich/e9d9d89d81e49c8fcbb8b2cc960efa4c91434319\">\"Using sensor pattern noise for camera model identification\" Thomas Filler, Jessica Fridrich, Miroslav Goljan</a></p>\n\n<ol>\n<li><p>We extracted approx. 3000 features from image linear pattern and noise pattern. We trained XGB, LightGBM and MultilayerPerceptron after applying PCA, scoring 0.898/0.881. (We found out that PCA reduced the training times of the models, but the raw dataset was trainable as well, achieving similar performance, if not better.)</p></li>\n<li><p>Our final model is an XGB model of stacked predictions from 9 L1 models:\n    - From images: DenseNet201, VGG16, Xception, DenseNet121, InceptionV3, ResNet50 (we tried others but we discarded them)\n    - From extracted features: XGB, LightGBM, MultilayerPerceptron</p></li>\n<li><p>This model (A13) scored 0.9877/0.9866. We trained the same model but excluding pseudolabeled data. We thought that they could introduce a strong bias toward these images, and after all, they were already used for all ours L1 models. This second model (A13a) scored  0.9798/0.9847, and we made the random decision of select this submission for the final, finishing 10th instead of 5th. (Our 2nd final submissions was a optimized weighted average of several submissions scoring 0.9885/0.9835, but we weren't very confident about it).</p></li>\n</ol>",
      "rawMarkdown": "Here is a brief overview of our modeling and solution. Please feel free to ask any questions.\n\n1. We only used the data provided by Kaggle and Gleb´s links. We had some concerns about the Gleb's dataset license of use for some of the images, to be honest, but most of them seem to be good. \n\n2. We first tried some CNN models from scratch, but discovered that the well known pretrained models converge faster and gave better accuracy. Our best single model was a DenseNet201 trained  on patches of size 336x336 scoring 0.963/0.967 (public/private LB). As the accuracy was pretty high, we decide to give a try to pseudo-labeled images from the testset using the predictions from this model.\n\n3. We applied a threshold of p&gt;=0.999 and used only unaltered images. This gives us 1116 new images to train with.\n\n4. The same DenseNet201 model, trained now with this extra dataset, scored 0.975/0.973. After intense TTA, it gave us 0.980/0.975.\n\n5. We trained a few more models (another DenseNet201, InceptionResnetV2, etc.) , but we couldn't improve much by just blending their predictions (and also being worried about overfitting to the LB), we decide to try stacking predictions from different models, and to train a second layer model with them.\n\n6.  Since we only had about a week to train them, and since we didn't have too many GPU resources, we decided to do stacking with just 2-folds.\n\n7. While we were training NN models, we tried to extract image features to train a non-NN model. We basically followed the ideas from this paper:\n\n[\"Using sensor pattern noise for camera model identification\" Thomas Filler, Jessica Fridrich, Miroslav Goljan][1]\n\n\n8. We extracted approx. 3000 features from image linear pattern and noise pattern. We trained XGB, LightGBM and MultilayerPerceptron after applying PCA, scoring 0.898/0.881. (We found out that PCA reduced the training times of the models, but the raw dataset was trainable as well, achieving similar performance, if not better.)\n\n9. Our final model is an XGB model of stacked predictions from 9 L1 models:\n        - From images: DenseNet201, VGG16, Xception, DenseNet121, InceptionV3, ResNet50 (we tried others but we discarded them)\n        - From extracted features: XGB, LightGBM, MultilayerPerceptron\n\n10. This model (A13) scored 0.9877/0.9866. We trained the same model but excluding pseudolabeled data. We thought that they could introduce a strong bias toward these images, and after all, they were already used for all ours L1 models. This second model (A13a) scored  0.9798/0.9847, and we made the random decision of select this submission for the final, finishing 10th instead of 5th. (Our 2nd final submissions was a optimized weighted average of several submissions scoring 0.9885/0.9835, but we weren't very confident about it).\n\n\n  [1]: https://www.semanticscholar.org/paper/Using-sensor-pattern-noise-for-camera-model-identi-Filler-Fridrich/e9d9d89d81e49c8fcbb8b2cc960efa4c91434319",
      "votes": null
    },
    {
      "id": "280395",
      "postDate": "02/09/2018 22:02:21",
      "content": "<p>@Bojan: well done, as usual! Thanks for sharing your insights :-)</p>",
      "rawMarkdown": "Bojan: well done, as usual! Thanks for sharing your insights :-)",
      "votes": null
    },
    {
      "id": "280405",
      "postDate": "02/09/2018 22:33:46",
      "content": "<p>Great work. I specially like the faux labels.</p>\n\n<p>When selecting faux labels w/ e.g.  <code>p&gt;=0.999</code>  is that from a single model at a given checkpoint, a single model at different checkpoints (I believe flipping <code>p</code>s even at sequential checkpoints are an indication of doubtful labels) or multiple models?</p>",
      "rawMarkdown": "Great work. I specially like the faux labels.\n\nWhen selecting faux labels w/ e.g.  `p&gt;=0.999`  is that from a single model at a given checkpoint, a single model at different checkpoints (I believe flipping `p`s even at sequential checkpoints are an indication of doubtful labels) or multiple models?",
      "votes": null
    },
    {
      "id": "280407",
      "postDate": "02/09/2018 22:42:15",
      "content": "<p>Those were from our best single DenseNet201 that was scoring 0.963. When we teamed up, we all retrained our best models with the same pseudolabels, and all of them showed varying but substantial improvements. We all felt pretty confident that the pseudolabels were working. </p>",
      "rawMarkdown": "Those were from our best single DenseNet201 that was scoring 0.963. When we teamed up, we all retrained our best models with the same pseudolabels, and all of them showed varying but substantial improvements. We all felt pretty confident that the pseudolabels were working.",
      "votes": null
    },
    {
      "id": "280805",
      "postDate": "02/11/2018 04:45:13",
      "content": "<p>How do you get the probability?</p>",
      "rawMarkdown": "How do you get the probability?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 280395,
      "author_name": "gvyshnya",
      "author_url": "",
      "post_date": "02/09/2018 22:02:21",
      "content": "<p>@Bojan: well done, as usual! Thanks for sharing your insights :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280405,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/09/2018 22:33:46",
      "content": "<p>Great work. I specially like the faux labels.</p>\n\n<p>When selecting faux labels w/ e.g.  <code>p&gt;=0.999</code>  is that from a single model at a given checkpoint, a single model at different checkpoints (I believe flipping <code>p</code>s even at sequential checkpoints are an indication of doubtful labels) or multiple models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 280407,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "02/09/2018 22:42:15",
          "content": "<p>Those were from our best single DenseNet201 that was scoring 0.963. When we teamed up, we all retrained our best models with the same pseudolabels, and all of them showed varying but substantial improvements. We all felt pretty confident that the pseudolabels were working. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280805,
          "author_name": "zhaoyangma",
          "author_url": "",
          "post_date": "02/11/2018 04:45:13",
          "content": "<p>How do you get the probability?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "280391": "Here is a brief overview of our modeling and solution. Please feel free to ask any questions.\n\n1. We only used the data provided by Kaggle and Gleb´s links. We had some concerns about the Gleb's dataset license of use for some of the images, to be honest, but most of them seem to be good. \n\n2. We first tried some CNN models from scratch, but discovered that the well known pretrained models converge faster and gave better accuracy. Our best single model was a DenseNet201 trained  on patches of size 336x336 scoring 0.963/0.967 (public/private LB). As the accuracy was pretty high, we decide to give a try to pseudo-labeled images from the testset using the predictions from this model.\n\n3. We applied a threshold of p&gt;=0.999 and used only unaltered images. This gives us 1116 new images to train with.\n\n4. The same DenseNet201 model, trained now with this extra dataset, scored 0.975/0.973. After intense TTA, it gave us 0.980/0.975.\n\n5. We trained a few more models (another DenseNet201, InceptionResnetV2, etc.) , but we couldn't improve much by just blending their predictions (and also being worried about overfitting to the LB), we decide to try stacking predictions from different models, and to train a second layer model with them.\n\n6.  Since we only had about a week to train them, and since we didn't have too many GPU resources, we decided to do stacking with just 2-folds.\n\n7. While we were training NN models, we tried to extract image features to train a non-NN model. We basically followed the ideas from this paper:\n\n[\"Using sensor pattern noise for camera model identification\" Thomas Filler, Jessica Fridrich, Miroslav Goljan][1]\n\n\n8. We extracted approx. 3000 features from image linear pattern and noise pattern. We trained XGB, LightGBM and MultilayerPerceptron after applying PCA, scoring 0.898/0.881. (We found out that PCA reduced the training times of the models, but the raw dataset was trainable as well, achieving similar performance, if not better.)\n\n9. Our final model is an XGB model of stacked predictions from 9 L1 models:\n        - From images: DenseNet201, VGG16, Xception, DenseNet121, InceptionV3, ResNet50 (we tried others but we discarded them)\n        - From extracted features: XGB, LightGBM, MultilayerPerceptron\n\n10. This model (A13) scored 0.9877/0.9866. We trained the same model but excluding pseudolabeled data. We thought that they could introduce a strong bias toward these images, and after all, they were already used for all ours L1 models. This second model (A13a) scored  0.9798/0.9847, and we made the random decision of select this submission for the final, finishing 10th instead of 5th. (Our 2nd final submissions was a optimized weighted average of several submissions scoring 0.9885/0.9835, but we weren't very confident about it).\n\n\n  [1]: https://www.semanticscholar.org/paper/Using-sensor-pattern-noise-for-camera-model-identi-Filler-Fridrich/e9d9d89d81e49c8fcbb8b2cc960efa4c91434319",
    "280395": "Bojan: well done, as usual! Thanks for sharing your insights :-)",
    "280405": "Great work. I specially like the faux labels.\n\nWhen selecting faux labels w/ e.g.  `p&gt;=0.999`  is that from a single model at a given checkpoint, a single model at different checkpoints (I believe flipping `p`s even at sequential checkpoints are an indication of doubtful labels) or multiple models?",
    "280407": "Those were from our best single DenseNet201 that was scoring 0.963. When we teamed up, we all retrained our best models with the same pseudolabels, and all of them showed varying but substantial improvements. We all felt pretty confident that the pseudolabels were working.",
    "280805": "How do you get the probability?"
  },
  "source": "meta"
}