{
  "id": 49198,
  "title": "Your Best Results without additional data ? ",
  "url": "/competitions/sp-society-camera-model-identification/discussion/49198",
  "author_name": "",
  "post_date": "2018-02-07T21:43:26.523914100Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Dear Kagglers, </p>\n\n<p>Did anyone of you try to solve the problem without using additional data ? </p>\n\n<p>I am about to submitt my very last results ; it is far from the best ones but it gave me the occasion to learn a lot about camera identification. I am about to be around 95% on the public LB with a single model (resnet-18) and without using additional data. To get there, I trained actually two resnets, one for altered and one for original images on small 64x64 crops. This gives an accuracy of 88% on public LB. It gives also a 99% cross validation accuracy. So the idea is to see how easy is to transfer the knowledge from one camera to another model of the same type.</p>\n\n<p>By fine tuning these nets on pseudo labels and using a standard PRNU based camera identification method, I can get up to 95% (98% on original images and 86% on altered images). Noise residues are estimated using locally adaptative dct. </p>\n\n<p>Like many of you I noticed that it is very easy to obtain 99% cross validation accuracy. So what is the most elegant way to transfer this knowledge on just one other camera ; without requiring to learn a few hundreds different camera models ?</p>\n\n<p>I am looking forward to reading the students solutions. I am looking forward to seeing the solution of the FIIGO team because they are producing very interesting research material on the subject. It would be awesome if they did it using unsupervised clustering.</p>\n\n<p>Edit, a recent paper of the leading team: <a href=\"http://ieeexplore.ieee.org/document/7919241/\">http://ieeexplore.ieee.org/document/7919241/</a></p>",
  "messages": [
    {
      "id": "279355",
      "postDate": "02/07/2018 21:43:26",
      "content": "<p>Dear Kagglers, </p>\n\n<p>Did anyone of you try to solve the problem without using additional data ? </p>\n\n<p>I am about to submitt my very last results ; it is far from the best ones but it gave me the occasion to learn a lot about camera identification. I am about to be around 95% on the public LB with a single model (resnet-18) and without using additional data. To get there, I trained actually two resnets, one for altered and one for original images on small 64x64 crops. This gives an accuracy of 88% on public LB. It gives also a 99% cross validation accuracy. So the idea is to see how easy is to transfer the knowledge from one camera to another model of the same type.</p>\n\n<p>By fine tuning these nets on pseudo labels and using a standard PRNU based camera identification method, I can get up to 95% (98% on original images and 86% on altered images). Noise residues are estimated using locally adaptative dct. </p>\n\n<p>Like many of you I noticed that it is very easy to obtain 99% cross validation accuracy. So what is the most elegant way to transfer this knowledge on just one other camera ; without requiring to learn a few hundreds different camera models ?</p>\n\n<p>I am looking forward to reading the students solutions. I am looking forward to seeing the solution of the FIIGO team because they are producing very interesting research material on the subject. It would be awesome if they did it using unsupervised clustering.</p>\n\n<p>Edit, a recent paper of the leading team: <a href=\"http://ieeexplore.ieee.org/document/7919241/\">http://ieeexplore.ieee.org/document/7919241/</a></p>",
      "rawMarkdown": "Dear Kagglers, \n\nDid anyone of you try to solve the problem without using additional data ? \n\nI am about to submitt my very last results ; it is far from the best ones but it gave me the occasion to learn a lot about camera identification. I am about to be around 95% on the public LB with a single model (resnet-18) and without using additional data. To get there, I trained actually two resnets, one for altered and one for original images on small 64x64 crops. This gives an accuracy of 88% on public LB. It gives also a 99% cross validation accuracy. So the idea is to see how easy is to transfer the knowledge from one camera to another model of the same type.\n\nBy fine tuning these nets on pseudo labels and using a standard PRNU based camera identification method, I can get up to 95% (98% on original images and 86% on altered images). Noise residues are estimated using locally adaptative dct. \n\nLike many of you I noticed that it is very easy to obtain 99% cross validation accuracy. So what is the most elegant way to transfer this knowledge on just one other camera ; without requiring to learn a few hundreds different camera models ?\n\nI am looking forward to reading the students solutions. I am looking forward to seeing the solution of the FIIGO team because they are producing very interesting research material on the subject. It would be awesome if they did it using unsupervised clustering.\n\nEdit, a recent paper of the leading team: http://ieeexplore.ieee.org/document/7919241/",
      "votes": null
    },
    {
      "id": "279989",
      "postDate": "02/09/2018 03:16:35",
      "content": "<p>I 'm eager to know best model without extra data , I am curious about SVM\\PCA , good solutions.</p>",
      "rawMarkdown": "I 'm eager to know best model without extra data , I am curious about SVM\\PCA , good solutions.",
      "votes": null
    },
    {
      "id": "280016",
      "postDate": "02/09/2018 05:05:50",
      "content": "<p>Did you try the majority voting over your 64x64 blocks classification labels to define the 512x512 image's source? can you tell me a guideline to make me understand your \"noise residues estimated using locally adaptative dct\"? can you also teach me how to \"fine tune with pseudo-labels?\"</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Did you try the majority voting over your 64x64 blocks classification labels to define the 512x512 image's source? can you tell me a guideline to make me understand your \"noise residues estimated using locally adaptative dct\"? can you also teach me how to \"fine tune with pseudo-labels?\"\n\nThanks!",
      "votes": null
    },
    {
      "id": "280040",
      "postDate": "02/09/2018 06:30:36",
      "content": "<p>Yes I extact 256 patches from an image and assign the majority class to this image. </p>\n\n<p>The noise residues are estimated as follows:\nresidues = image - denoised_image</p>\n\n<p>The denoised image is obtain with a locally adaptative dct implemented in opencv:\n<a href=\"http://www.ipol.im/pub/art/2011/ys-dct/\">http://www.ipol.im/pub/art/2011/ys-dct/</a>\n<a href=\"https://docs.opencv.org/3.0-beta/modules/xphoto/doc/denoising/denoising.html\">https://docs.opencv.org/3.0-beta/modules/xphoto/doc/denoising/denoising.html</a></p>\n\n<p>There are of course better methods NLmeans, BM3D or even a denoising autoencoder network. </p>\n\n<p>Fine tuning with pseudo labels:\n- basically you estimate the labels of the test set and consider them as ground truth. The point is that you do not retrain the whole network with these additional data but only fine tune the FC layers. </p>",
      "rawMarkdown": "Yes I extact 256 patches from an image and assign the majority class to this image. \n\nThe noise residues are estimated as follows:\nresidues = image - denoised_image\n\nThe denoised image is obtain with a locally adaptative dct implemented in opencv:\nhttp://www.ipol.im/pub/art/2011/ys-dct/\nhttps://docs.opencv.org/3.0-beta/modules/xphoto/doc/denoising/denoising.html\n\nThere are of course better methods NLmeans, BM3D or even a denoising autoencoder network. \n\n\nFine tuning with pseudo labels:\n- basically you estimate the labels of the test set and consider them as ground truth. The point is that you do not retrain the whole network with these additional data but only fine tune the FC layers.",
      "votes": null
    },
    {
      "id": "280042",
      "postDate": "02/09/2018 06:34:41",
      "content": "<p>wow! i did not know that! fine tuning only fully connected layers! is there any reason for that?</p>\n\n<p>you use 64x64 overlapping blocks, correct?</p>",
      "rawMarkdown": "wow! i did not know that! fine tuning only fully connected layers! is there any reason for that?\n\nyou use 64x64 overlapping blocks, correct?",
      "votes": null
    },
    {
      "id": "280043",
      "postDate": "02/09/2018 06:42:28",
      "content": "<p>Yes there is a reason for only fine tuning the FC layers. You do not want your network to overfit on the new data that you provide. You know that your network already generalizes well on the test data. So you select the most probable labels of the test data and guide your network to adapt itself on these data. Actually I only fine tuned the very last linear layer. </p>",
      "rawMarkdown": "Yes there is a reason for only fine tuning the FC layers. You do not want your network to overfit on the new data that you provide. You know that your network already generalizes well on the test data. So you select the most probable labels of the test data and guide your network to adapt itself on these data. Actually I only fine tuned the very last linear layer.",
      "votes": null
    },
    {
      "id": "280046",
      "postDate": "02/09/2018 06:43:54",
      "content": "<p>Hello jeandebleau, sorry for repeating the same question in two topics :-(</p>\n\n<p>Do you believe your noise residues would be a better input for your CNN approach? because we are eliminating image content this way. Or do you believe it is a bad idea considering that a common solution is doing data augmentation using the same challenge's manipulations? </p>\n\n<p>Thank you so much.</p>",
      "rawMarkdown": "Hello jeandebleau, sorry for repeating the same question in two topics :-(\n\nDo you believe your noise residues would be a better input for your CNN approach? because we are eliminating image content this way. Or do you believe it is a bad idea considering that a common solution is doing data augmentation using the same challenge's manipulations? \n\nThank you so much.",
      "votes": null
    },
    {
      "id": "280126",
      "postDate": "02/09/2018 12:26:27",
      "content": "<p>This was my first competition.  I was able to achieve 87% on the public and private leaderboards (a few hours after the competition closed).  I used a single resnext50 model.  I was conservative with my approach as this was my first competition so I wanted to stay within the bounds of the rules.  I used the following for training: \n- Trained on initially provided training dataset only (-15% of images set aside for validation).\n- I set up my training and validation images based 1024x1024 center crops with 1 unaltered image and 8 more processed per the competition test set.  From this set I used random 512x512 crops from within the 1024x1024 center crops for training and 512x512 center crops for validation.  My model is set up to do random rotation (0,90,180,270) during training and validation.  I used TTA on my test set submissions which included the rotation only.\n- I trained the FC layers to get to ~54% val accuracy initially and then trained for 7 epocs across all layers to get to ~98.5% val accuracy and 87% on the leaderboard.  .  </p>",
      "rawMarkdown": "This was my first competition.  I was able to achieve 87% on the public and private leaderboards (a few hours after the competition closed).  I used a single resnext50 model.  I was conservative with my approach as this was my first competition so I wanted to stay within the bounds of the rules.  I used the following for training: \n- Trained on initially provided training dataset only (-15% of images set aside for validation).\n- I set up my training and validation images based 1024x1024 center crops with 1 unaltered image and 8 more processed per the competition test set.  From this set I used random 512x512 crops from within the 1024x1024 center crops for training and 512x512 center crops for validation.  My model is set up to do random rotation (0,90,180,270) during training and validation.  I used TTA on my test set submissions which included the rotation only.\n- I trained the FC layers to get to ~54% val accuracy initially and then trained for 7 epocs across all layers to get to ~98.5% val accuracy and 87% on the leaderboard.  .",
      "votes": null
    },
    {
      "id": "280807",
      "postDate": "02/11/2018 04:50:26",
      "content": "<p>@jeandebleau\n\"Like many of you I noticed that it is very easy to obtain 99% cross validation accuracy.\" I never get a validation accuracy score above 93%. How do you choose the training set and validation set?</p>",
      "rawMarkdown": "jeandebleau\n\"Like many of you I noticed that it is very easy to obtain 99% cross validation accuracy.\" I never get a validation accuracy score above 93%. How do you choose the training set and validation set?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 279989,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "02/09/2018 03:16:35",
      "content": "<p>I 'm eager to know best model without extra data , I am curious about SVM\\PCA , good solutions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280016,
      "author_name": "anselmoferreira35",
      "author_url": "",
      "post_date": "02/09/2018 05:05:50",
      "content": "<p>Did you try the majority voting over your 64x64 blocks classification labels to define the 512x512 image's source? can you tell me a guideline to make me understand your \"noise residues estimated using locally adaptative dct\"? can you also teach me how to \"fine tune with pseudo-labels?\"</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 280040,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "02/09/2018 06:30:36",
          "content": "<p>Yes I extact 256 patches from an image and assign the majority class to this image. </p>\n\n<p>The noise residues are estimated as follows:\nresidues = image - denoised_image</p>\n\n<p>The denoised image is obtain with a locally adaptative dct implemented in opencv:\n<a href=\"http://www.ipol.im/pub/art/2011/ys-dct/\">http://www.ipol.im/pub/art/2011/ys-dct/</a>\n<a href=\"https://docs.opencv.org/3.0-beta/modules/xphoto/doc/denoising/denoising.html\">https://docs.opencv.org/3.0-beta/modules/xphoto/doc/denoising/denoising.html</a></p>\n\n<p>There are of course better methods NLmeans, BM3D or even a denoising autoencoder network. </p>\n\n<p>Fine tuning with pseudo labels:\n- basically you estimate the labels of the test set and consider them as ground truth. The point is that you do not retrain the whole network with these additional data but only fine tune the FC layers. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280042,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/09/2018 06:34:41",
          "content": "<p>wow! i did not know that! fine tuning only fully connected layers! is there any reason for that?</p>\n\n<p>you use 64x64 overlapping blocks, correct?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280043,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "02/09/2018 06:42:28",
          "content": "<p>Yes there is a reason for only fine tuning the FC layers. You do not want your network to overfit on the new data that you provide. You know that your network already generalizes well on the test data. So you select the most probable labels of the test data and guide your network to adapt itself on these data. Actually I only fine tuned the very last linear layer. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280046,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/09/2018 06:43:54",
          "content": "<p>Hello jeandebleau, sorry for repeating the same question in two topics :-(</p>\n\n<p>Do you believe your noise residues would be a better input for your CNN approach? because we are eliminating image content this way. Or do you believe it is a bad idea considering that a common solution is doing data augmentation using the same challenge's manipulations? </p>\n\n<p>Thank you so much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280126,
      "author_name": "matdmiller",
      "author_url": "",
      "post_date": "02/09/2018 12:26:27",
      "content": "<p>This was my first competition.  I was able to achieve 87% on the public and private leaderboards (a few hours after the competition closed).  I used a single resnext50 model.  I was conservative with my approach as this was my first competition so I wanted to stay within the bounds of the rules.  I used the following for training: \n- Trained on initially provided training dataset only (-15% of images set aside for validation).\n- I set up my training and validation images based 1024x1024 center crops with 1 unaltered image and 8 more processed per the competition test set.  From this set I used random 512x512 crops from within the 1024x1024 center crops for training and 512x512 center crops for validation.  My model is set up to do random rotation (0,90,180,270) during training and validation.  I used TTA on my test set submissions which included the rotation only.\n- I trained the FC layers to get to ~54% val accuracy initially and then trained for 7 epocs across all layers to get to ~98.5% val accuracy and 87% on the leaderboard.  .  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280807,
      "author_name": "zhaoyangma",
      "author_url": "",
      "post_date": "02/11/2018 04:50:26",
      "content": "<p>@jeandebleau\n\"Like many of you I noticed that it is very easy to obtain 99% cross validation accuracy.\" I never get a validation accuracy score above 93%. How do you choose the training set and validation set?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "279355": "Dear Kagglers, \n\nDid anyone of you try to solve the problem without using additional data ? \n\nI am about to submitt my very last results ; it is far from the best ones but it gave me the occasion to learn a lot about camera identification. I am about to be around 95% on the public LB with a single model (resnet-18) and without using additional data. To get there, I trained actually two resnets, one for altered and one for original images on small 64x64 crops. This gives an accuracy of 88% on public LB. It gives also a 99% cross validation accuracy. So the idea is to see how easy is to transfer the knowledge from one camera to another model of the same type.\n\nBy fine tuning these nets on pseudo labels and using a standard PRNU based camera identification method, I can get up to 95% (98% on original images and 86% on altered images). Noise residues are estimated using locally adaptative dct. \n\nLike many of you I noticed that it is very easy to obtain 99% cross validation accuracy. So what is the most elegant way to transfer this knowledge on just one other camera ; without requiring to learn a few hundreds different camera models ?\n\nI am looking forward to reading the students solutions. I am looking forward to seeing the solution of the FIIGO team because they are producing very interesting research material on the subject. It would be awesome if they did it using unsupervised clustering.\n\nEdit, a recent paper of the leading team: http://ieeexplore.ieee.org/document/7919241/",
    "279989": "I 'm eager to know best model without extra data , I am curious about SVM\\PCA , good solutions.",
    "280016": "Did you try the majority voting over your 64x64 blocks classification labels to define the 512x512 image's source? can you tell me a guideline to make me understand your \"noise residues estimated using locally adaptative dct\"? can you also teach me how to \"fine tune with pseudo-labels?\"\n\nThanks!",
    "280040": "Yes I extact 256 patches from an image and assign the majority class to this image. \n\nThe noise residues are estimated as follows:\nresidues = image - denoised_image\n\nThe denoised image is obtain with a locally adaptative dct implemented in opencv:\nhttp://www.ipol.im/pub/art/2011/ys-dct/\nhttps://docs.opencv.org/3.0-beta/modules/xphoto/doc/denoising/denoising.html\n\nThere are of course better methods NLmeans, BM3D or even a denoising autoencoder network. \n\n\nFine tuning with pseudo labels:\n- basically you estimate the labels of the test set and consider them as ground truth. The point is that you do not retrain the whole network with these additional data but only fine tune the FC layers.",
    "280042": "wow! i did not know that! fine tuning only fully connected layers! is there any reason for that?\n\nyou use 64x64 overlapping blocks, correct?",
    "280043": "Yes there is a reason for only fine tuning the FC layers. You do not want your network to overfit on the new data that you provide. You know that your network already generalizes well on the test data. So you select the most probable labels of the test data and guide your network to adapt itself on these data. Actually I only fine tuned the very last linear layer.",
    "280046": "Hello jeandebleau, sorry for repeating the same question in two topics :-(\n\nDo you believe your noise residues would be a better input for your CNN approach? because we are eliminating image content this way. Or do you believe it is a bad idea considering that a common solution is doing data augmentation using the same challenge's manipulations? \n\nThank you so much.",
    "280126": "This was my first competition.  I was able to achieve 87% on the public and private leaderboards (a few hours after the competition closed).  I used a single resnext50 model.  I was conservative with my approach as this was my first competition so I wanted to stay within the bounds of the rules.  I used the following for training: \n- Trained on initially provided training dataset only (-15% of images set aside for validation).\n- I set up my training and validation images based 1024x1024 center crops with 1 unaltered image and 8 more processed per the competition test set.  From this set I used random 512x512 crops from within the 1024x1024 center crops for training and 512x512 center crops for validation.  My model is set up to do random rotation (0,90,180,270) during training and validation.  I used TTA on my test set submissions which included the rotation only.\n- I trained the FC layers to get to ~54% val accuracy initially and then trained for 7 epocs across all layers to get to ~98.5% val accuracy and 87% on the leaderboard.  .",
    "280807": "jeandebleau\n\"Like many of you I noticed that it is very easy to obtain 99% cross validation accuracy.\" I never get a validation accuracy score above 93%. How do you choose the training set and validation set?"
  },
  "source": "meta"
}