{
  "id": 49602,
  "title": "3rd place solution",
  "url": "/competitions/sp-society-camera-model-identification/writeups/fiigo-spcup-eligible-3rd-place-solution",
  "author_name": "",
  "post_date": "2018-02-13T11:16:26.087856400Z",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>Data</strong></p>\n\n<p>The training set has been augmented with both private images and images downloaded from Flickr. \nAll images were preliminary sifted based on the EXIF header meta-data information,\nto keep only images with meta-data coherent with the corresponding camera of the training set. \nAll images that were processed by an editing software or with restrictive license were removed. </p>\n\n<p>Finally we obtained a dataset of 450 images per model (about 15 GB). 400 images have been used for training, 50 images for validation.</p>\n\n<p><strong>Training</strong></p>\n\n<p>During training patches were randomly extracted from the central part of image (1024x1024 pixels) and randomly rotated with steps of 90 degrees. </p>\n\n<p>We trained separate networks for the model identification of unaltered and manipulated images. \nFor the first problem, we included only unaltered images in the training set. \nOn the contrary, the networks for manipulated images were trained using both unaltered and manipulated images, as we found this solution to provide consistently better results.</p>\n\n<p><strong>Solution overview</strong></p>\n\n<p>Fusion of different CNNs models (arithmetic mean of the scores). </p>\n\n<p>For unaltered images we considered 5 networks trained at different patch size: XceptionNet (96), XceptionNet (299), Inception v3 (139), InceptionResNet v2 (299), DenseNet121 (224). </p>\n\n<p>For manipulated images we considered the same networks above but used various input patch sizes, for a total of 14 models. \nWe tested many alternative solutions to improve performance, two of them were eventually applied: 1) network training on small patches and fine-tuning on larger ones, 2) identification of problematic classes (JPEG 70 and Resizing 0.5) and design of dedicated detectors.</p>\n\n<p><strong>Hardware</strong></p>\n\n<p>NVIDIA Tesla P100 GPU with 16GB of RAM.</p>\n\n<p><strong>Special mention</strong> </p>\n\n<p>It is interesting that using only a single model for unaltered (XceptionNet on patches of dimension 96) and manipulated images (XceptionNet on patches of dimension 96 fine tuned on 299) gave a public score = 0.986666 and a private score = 0.981309. </p>\n\n<p><strong>Lessons learnt</strong></p>\n\n<p>Do not trust the public LB too much. Use a better validation set!</p>\n\n<p>This was a great experience for us (our first time in a Kaggle competition) and we really learnt a lot. Many thanks to the organization and to all the teams sharing their work.</p>",
  "messages": [
    {
      "id": "281960",
      "postDate": "02/13/2018 11:16:26",
      "content": "<p><strong>Data</strong></p>\n\n<p>The training set has been augmented with both private images and images downloaded from Flickr. \nAll images were preliminary sifted based on the EXIF header meta-data information,\nto keep only images with meta-data coherent with the corresponding camera of the training set. \nAll images that were processed by an editing software or with restrictive license were removed. </p>\n\n<p>Finally we obtained a dataset of 450 images per model (about 15 GB). 400 images have been used for training, 50 images for validation.</p>\n\n<p><strong>Training</strong></p>\n\n<p>During training patches were randomly extracted from the central part of image (1024x1024 pixels) and randomly rotated with steps of 90 degrees. </p>\n\n<p>We trained separate networks for the model identification of unaltered and manipulated images. \nFor the first problem, we included only unaltered images in the training set. \nOn the contrary, the networks for manipulated images were trained using both unaltered and manipulated images, as we found this solution to provide consistently better results.</p>\n\n<p><strong>Solution overview</strong></p>\n\n<p>Fusion of different CNNs models (arithmetic mean of the scores). </p>\n\n<p>For unaltered images we considered 5 networks trained at different patch size: XceptionNet (96), XceptionNet (299), Inception v3 (139), InceptionResNet v2 (299), DenseNet121 (224). </p>\n\n<p>For manipulated images we considered the same networks above but used various input patch sizes, for a total of 14 models. \nWe tested many alternative solutions to improve performance, two of them were eventually applied: 1) network training on small patches and fine-tuning on larger ones, 2) identification of problematic classes (JPEG 70 and Resizing 0.5) and design of dedicated detectors.</p>\n\n<p><strong>Hardware</strong></p>\n\n<p>NVIDIA Tesla P100 GPU with 16GB of RAM.</p>\n\n<p><strong>Special mention</strong> </p>\n\n<p>It is interesting that using only a single model for unaltered (XceptionNet on patches of dimension 96) and manipulated images (XceptionNet on patches of dimension 96 fine tuned on 299) gave a public score = 0.986666 and a private score = 0.981309. </p>\n\n<p><strong>Lessons learnt</strong></p>\n\n<p>Do not trust the public LB too much. Use a better validation set!</p>\n\n<p>This was a great experience for us (our first time in a Kaggle competition) and we really learnt a lot. Many thanks to the organization and to all the teams sharing their work.</p>",
      "rawMarkdown": "**Data**\n\nThe training set has been augmented with both private images and images downloaded from Flickr. \nAll images were preliminary sifted based on the EXIF header meta-data information,\nto keep only images with meta-data coherent with the corresponding camera of the training set. \nAll images that were processed by an editing software or with restrictive license were removed. \n\nFinally we obtained a dataset of 450 images per model (about 15 GB). 400 images have been used for training, 50 images for validation.\n\n**Training**\n\nDuring training patches were randomly extracted from the central part of image (1024x1024 pixels) and randomly rotated with steps of 90 degrees. \n\nWe trained separate networks for the model identification of unaltered and manipulated images. \nFor the first problem, we included only unaltered images in the training set. \nOn the contrary, the networks for manipulated images were trained using both unaltered and manipulated images, as we found this solution to provide consistently better results.\n\n**Solution overview**\n\nFusion of different CNNs models (arithmetic mean of the scores). \n\nFor unaltered images we considered 5 networks trained at different patch size: XceptionNet (96), XceptionNet (299), Inception v3 (139), InceptionResNet v2 (299), DenseNet121 (224). \n\nFor manipulated images we considered the same networks above but used various input patch sizes, for a total of 14 models. \nWe tested many alternative solutions to improve performance, two of them were eventually applied: 1) network training on small patches and fine-tuning on larger ones, 2) identification of problematic classes (JPEG 70 and Resizing 0.5) and design of dedicated detectors.\n\n**Hardware**\n\nNVIDIA Tesla P100 GPU with 16GB of RAM.\n\n**Special mention** \n\nIt is interesting that using only a single model for unaltered (XceptionNet on patches of dimension 96) and manipulated images (XceptionNet on patches of dimension 96 fine tuned on 299) gave a public score = 0.986666 and a private score = 0.981309. \n\n**Lessons learnt**\n\nDo not trust the public LB too much. Use a better validation set!\n\nThis was a great experience for us (our first time in a Kaggle competition) and we really learnt a lot. Many thanks to the organization and to all the teams sharing their work.",
      "votes": null
    },
    {
      "id": "282614",
      "postDate": "02/14/2018 09:03:59",
      "content": "<p>Very impressive results considering you only had 450 images for each class. Congratulations.</p>\n\n<p>Can you elaborate on:</p>\n\n<blockquote>\n  <p>2) identification of problematic classes (JPEG 70 and Resizing 0.5)\n  and design of dedicated detectors.</p>\n</blockquote>",
      "rawMarkdown": "Very impressive results considering you only had 450 images for each class. Congratulations.\n\nCan you elaborate on:\n\n&gt; 2) identification of problematic classes (JPEG 70 and Resizing 0.5)\n&gt; and design of dedicated detectors.",
      "votes": null
    },
    {
      "id": "282936",
      "postDate": "02/14/2018 19:47:11",
      "content": "<p>We noticed on our internal validation set that the accuracy was different based on the type of manipulations.\nFor example, with reference to XceptionNet working on patches of dimension 96, we obtained the following results:</p>\n\n<p>Accuracy</p>\n\n<p>0.982  gamma = 0.8</p>\n\n<p>0.982  gamma = 1.2</p>\n\n<p>0.978  JPEG 90</p>\n\n<p>0.940  JPEG 70</p>\n\n<p>0.982  resize = 0.8</p>\n\n<p>0.966  resize = 0.5</p>\n\n<p>0.969  resize = 1.5</p>\n\n<p>0.954  resize = 2.0</p>\n\n<p>Hence, we decided to train specific networks for the classes achieving the worst results:\ncompression (JPEG = 70) and resizing (0.5 and 2).\nThis actually means that in the training set only images manipulated using that specific class are present.\nWhen fusing with the other detectors only two of them helped us increasing the score on the public LB,\nhence we discarded the one designed for detecting resizing by a factor of 2.</p>",
      "rawMarkdown": "We noticed on our internal validation set that the accuracy was different based on the type of manipulations.\nFor example, with reference to XceptionNet working on patches of dimension 96, we obtained the following results:\n\nAccuracy\n\n0.982  gamma = 0.8\n\n0.982  gamma = 1.2\n\n0.978  JPEG 90\n\n0.940  JPEG 70\n\n0.982  resize = 0.8\n\n0.966  resize = 0.5\n\n0.969  resize = 1.5\n\n0.954  resize = 2.0\n\nHence, we decided to train specific networks for the classes achieving the worst results:\ncompression (JPEG = 70) and resizing (0.5 and 2).\nThis actually means that in the training set only images manipulated using that specific class are present.\nWhen fusing with the other detectors only two of them helped us increasing the score on the public LB,\nhence we discarded the one designed for detecting resizing by a factor of 2.",
      "votes": null
    },
    {
      "id": "282948",
      "postDate": "02/14/2018 19:57:12",
      "content": "<p>So you also designed different networks to <strong>detect</strong> the manipulation type? \nIf so, how accurate where the detectors?</p>\n\n<p>FWIW, I tried a similar approach but all in one net, the net would detect the manipulation type, and feed that prediction concatenated with the feature extractor output (DenseNet201) to the fully-connected head responsible to classify the camera.</p>",
      "rawMarkdown": "So you also designed different networks to **detect** the manipulation type? \nIf so, how accurate where the detectors?\n\nFWIW, I tried a similar approach but all in one net, the net would detect the manipulation type, and feed that prediction concatenated with the feature extractor output (DenseNet201) to the fully-connected head responsible to classify the camera.",
      "votes": null
    },
    {
      "id": "282979",
      "postDate": "02/14/2018 20:31:23",
      "content": "<p>No. We tried to detect the type of manipulations, but we were not able to achieve good results.\nSo we decided to give up along this direction.\nWhat we did was only to verify that the accuracy on the detection of the manipulations was not equally good\nand added specific detectors for these problematic classes.</p>",
      "rawMarkdown": "No. We tried to detect the type of manipulations, but we were not able to achieve good results.\nSo we decided to give up along this direction.\nWhat we did was only to verify that the accuracy on the detection of the manipulations was not equally good\nand added specific detectors for these problematic classes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 282614,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/14/2018 09:03:59",
      "content": "<p>Very impressive results considering you only had 450 images for each class. Congratulations.</p>\n\n<p>Can you elaborate on:</p>\n\n<blockquote>\n  <p>2) identification of problematic classes (JPEG 70 and Resizing 0.5)\n  and design of dedicated detectors.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 282936,
      "author_name": "verdoliv",
      "author_url": "",
      "post_date": "02/14/2018 19:47:11",
      "content": "<p>We noticed on our internal validation set that the accuracy was different based on the type of manipulations.\nFor example, with reference to XceptionNet working on patches of dimension 96, we obtained the following results:</p>\n\n<p>Accuracy</p>\n\n<p>0.982  gamma = 0.8</p>\n\n<p>0.982  gamma = 1.2</p>\n\n<p>0.978  JPEG 90</p>\n\n<p>0.940  JPEG 70</p>\n\n<p>0.982  resize = 0.8</p>\n\n<p>0.966  resize = 0.5</p>\n\n<p>0.969  resize = 1.5</p>\n\n<p>0.954  resize = 2.0</p>\n\n<p>Hence, we decided to train specific networks for the classes achieving the worst results:\ncompression (JPEG = 70) and resizing (0.5 and 2).\nThis actually means that in the training set only images manipulated using that specific class are present.\nWhen fusing with the other detectors only two of them helped us increasing the score on the public LB,\nhence we discarded the one designed for detecting resizing by a factor of 2.</p>",
      "votes": null,
      "replies": [
        {
          "id": 282948,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/14/2018 19:57:12",
          "content": "<p>So you also designed different networks to <strong>detect</strong> the manipulation type? \nIf so, how accurate where the detectors?</p>\n\n<p>FWIW, I tried a similar approach but all in one net, the net would detect the manipulation type, and feed that prediction concatenated with the feature extractor output (DenseNet201) to the fully-connected head responsible to classify the camera.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 282979,
          "author_name": "verdoliv",
          "author_url": "",
          "post_date": "02/14/2018 20:31:23",
          "content": "<p>No. We tried to detect the type of manipulations, but we were not able to achieve good results.\nSo we decided to give up along this direction.\nWhat we did was only to verify that the accuracy on the detection of the manipulations was not equally good\nand added specific detectors for these problematic classes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "281960": "**Data**\n\nThe training set has been augmented with both private images and images downloaded from Flickr. \nAll images were preliminary sifted based on the EXIF header meta-data information,\nto keep only images with meta-data coherent with the corresponding camera of the training set. \nAll images that were processed by an editing software or with restrictive license were removed. \n\nFinally we obtained a dataset of 450 images per model (about 15 GB). 400 images have been used for training, 50 images for validation.\n\n**Training**\n\nDuring training patches were randomly extracted from the central part of image (1024x1024 pixels) and randomly rotated with steps of 90 degrees. \n\nWe trained separate networks for the model identification of unaltered and manipulated images. \nFor the first problem, we included only unaltered images in the training set. \nOn the contrary, the networks for manipulated images were trained using both unaltered and manipulated images, as we found this solution to provide consistently better results.\n\n**Solution overview**\n\nFusion of different CNNs models (arithmetic mean of the scores). \n\nFor unaltered images we considered 5 networks trained at different patch size: XceptionNet (96), XceptionNet (299), Inception v3 (139), InceptionResNet v2 (299), DenseNet121 (224). \n\nFor manipulated images we considered the same networks above but used various input patch sizes, for a total of 14 models. \nWe tested many alternative solutions to improve performance, two of them were eventually applied: 1) network training on small patches and fine-tuning on larger ones, 2) identification of problematic classes (JPEG 70 and Resizing 0.5) and design of dedicated detectors.\n\n**Hardware**\n\nNVIDIA Tesla P100 GPU with 16GB of RAM.\n\n**Special mention** \n\nIt is interesting that using only a single model for unaltered (XceptionNet on patches of dimension 96) and manipulated images (XceptionNet on patches of dimension 96 fine tuned on 299) gave a public score = 0.986666 and a private score = 0.981309. \n\n**Lessons learnt**\n\nDo not trust the public LB too much. Use a better validation set!\n\nThis was a great experience for us (our first time in a Kaggle competition) and we really learnt a lot. Many thanks to the organization and to all the teams sharing their work.",
    "282614": "Very impressive results considering you only had 450 images for each class. Congratulations.\n\nCan you elaborate on:\n\n&gt; 2) identification of problematic classes (JPEG 70 and Resizing 0.5)\n&gt; and design of dedicated detectors.",
    "282936": "We noticed on our internal validation set that the accuracy was different based on the type of manipulations.\nFor example, with reference to XceptionNet working on patches of dimension 96, we obtained the following results:\n\nAccuracy\n\n0.982  gamma = 0.8\n\n0.982  gamma = 1.2\n\n0.978  JPEG 90\n\n0.940  JPEG 70\n\n0.982  resize = 0.8\n\n0.966  resize = 0.5\n\n0.969  resize = 1.5\n\n0.954  resize = 2.0\n\nHence, we decided to train specific networks for the classes achieving the worst results:\ncompression (JPEG = 70) and resizing (0.5 and 2).\nThis actually means that in the training set only images manipulated using that specific class are present.\nWhen fusing with the other detectors only two of them helped us increasing the score on the public LB,\nhence we discarded the one designed for detecting resizing by a factor of 2.",
    "282948": "So you also designed different networks to **detect** the manipulation type? \nIf so, how accurate where the detectors?\n\nFWIW, I tried a similar approach but all in one net, the net would detect the manipulation type, and feed that prediction concatenated with the feature extractor output (DenseNet201) to the fully-connected head responsible to classify the camera.",
    "282979": "No. We tried to detect the type of manipulations, but we were not able to achieve good results.\nSo we decided to give up along this direction.\nWhat we did was only to verify that the accuracy on the detection of the manipulations was not equally good\nand added specific detectors for these problematic classes."
  },
  "source": "meta"
}