{
  "id": 47764,
  "title": "Simple CNN makes no sense ",
  "url": "/competitions/sp-society-camera-model-identification/discussion/47764",
  "author_name": "",
  "post_date": "2018-01-18T16:25:53.088439500Z",
  "votes": 3,
  "comment_count": 18,
  "views": 0,
  "content": "<p>CNNs are great but I am not able to digest the fact that people are applying simple \nCNNs for making predictions. IMO, training architectures like VGG16 and ResNet50 are of very little use. Those architectures, as far as I know, cannot capture so dense information  about the image. They are good for object classification, object detection, etc but when you want to predict the camera model from an image, you are making too much of assumptions that the network can learn by itself.</p>",
  "messages": [
    {
      "id": "270628",
      "postDate": "01/18/2018 16:25:53",
      "content": "<p>CNNs are great but I am not able to digest the fact that people are applying simple \nCNNs for making predictions. IMO, training architectures like VGG16 and ResNet50 are of very little use. Those architectures, as far as I know, cannot capture so dense information  about the image. They are good for object classification, object detection, etc but when you want to predict the camera model from an image, you are making too much of assumptions that the network can learn by itself.</p>",
      "rawMarkdown": "CNNs are great but I am not able to digest the fact that people are applying simple \nCNNs for making predictions. IMO, training architectures like VGG16 and ResNet50 are of very little use. Those architectures, as far as I know, cannot capture so dense information  about the image. They are good for object classification, object detection, etc but when you want to predict the camera model from an image, you are making too much of assumptions that the network can learn by itself.",
      "votes": null
    },
    {
      "id": "270756",
      "postDate": "01/18/2018 21:33:21",
      "content": "<p>So, what do you suggest for this task?</p>",
      "rawMarkdown": "So, what do you suggest for this task?",
      "votes": null
    },
    {
      "id": "270888",
      "postDate": "01/19/2018 06:48:03",
      "content": "<p>The way I like to think about this is and I might be totally wrong, we know that NNs are universal approximators and here we have a task to approximate whatever complex composition of functions that are happening inside the cameras include demosacing, denoising, compression etc. Now we could do this with a fully connected network but it would be difficult to train with SGD and related optimization algorithms. The only prior assumptions with using convolution layers are local relationships, translation equivariance for the functions learned (and some translation invariance due to pooling) which are presumably correct or do no harm for the task at hand.  Using convolutions and pooling are also lowering the memory footprint significantly and some computation and time cost so a win-win for all related grid based tasks including time series data, 1D conv for audio etc. The use of existing architecture is giving a proven arrangement of layers, some useful shared weights from imagenet (no idea how) and possibly a better than random initialisation if you proceed with training all the layers.</p>",
      "rawMarkdown": "The way I like to think about this is and I might be totally wrong, we know that NNs are universal approximators and here we have a task to approximate whatever complex composition of functions that are happening inside the cameras include demosacing, denoising, compression etc. Now we could do this with a fully connected network but it would be difficult to train with SGD and related optimization algorithms. The only prior assumptions with using convolution layers are local relationships, translation equivariance for the functions learned (and some translation invariance due to pooling) which are presumably correct or do no harm for the task at hand.  Using convolutions and pooling are also lowering the memory footprint significantly and some computation and time cost so a win-win for all related grid based tasks including time series data, 1D conv for audio etc. The use of existing architecture is giving a proven arrangement of layers, some useful shared weights from imagenet (no idea how) and possibly a better than random initialisation if you proceed with training all the layers.",
      "votes": null
    },
    {
      "id": "270958",
      "postDate": "01/19/2018 10:33:43",
      "content": "<p>However, it seems that CNN indeed worked. Please check this discussion <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/46688\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/46688</a>, where the author has ranked 3rd. </p>",
      "rawMarkdown": "However, it seems that CNN indeed worked. Please check this discussion https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/46688, where the author has ranked 3rd.",
      "votes": null
    },
    {
      "id": "270971",
      "postDate": "01/19/2018 10:50:58",
      "content": "<p>I have 97.2% accuracy on LB on unaltered pictures with one CNN. They do work but it is not the only possible approach. Abhishek does give an explanation of why they could perform</p>",
      "rawMarkdown": "I have 97.2% accuracy on LB on unaltered pictures with one CNN. They do work but it is not the only possible approach. Abhishek does give an explanation of why they could perform",
      "votes": null
    },
    {
      "id": "271000",
      "postDate": "01/19/2018 12:29:58",
      "content": "<p>Hi, Max. You said \"I have 97.2% accuracy on LB on unaltered pictures\". In the test set, there are both unaltered and manipulated images, how did you know you accuracy only on unaltered ones?</p>",
      "rawMarkdown": "Hi, Max. You said \"I have 97.2% accuracy on LB on unaltered pictures\". In the test set, there are both unaltered and manipulated images, how did you know you accuracy only on unaltered ones?",
      "votes": null
    },
    {
      "id": "271066",
      "postDate": "01/19/2018 16:14:26",
      "content": "<p>Wait for Private LB</p>",
      "rawMarkdown": "Wait for Private LB",
      "votes": null
    },
    {
      "id": "271070",
      "postDate": "01/19/2018 16:18:41",
      "content": "<p>I know that NN can approximate any mathematical function. The thing is you are trying to explain NNs to me which I know very well. Also, you are treating a CNN like a black box, more precisely a magical box. I understand convolution, pooling, etc, I am trying to understand how a CNN can learn things like compression, focal length, etc from a photo</p>",
      "rawMarkdown": "I know that NN can approximate any mathematical function. The thing is you are trying to explain NNs to me which I know very well. Also, you are treating a CNN like a black box, more precisely a magical box. I understand convolution, pooling, etc, I am trying to understand how a CNN can learn things like compression, focal length, etc from a photo",
      "votes": null
    },
    {
      "id": "271071",
      "postDate": "01/19/2018 16:20:47",
      "content": "<p>I am not interested in ranking. The only thing I want is a logical explanation about how a CNN can learn camera based stats from the images taken by it. Had it been so simple, don't you think people would already have done that and there would have been no need to host this competition?</p>",
      "rawMarkdown": "I am not interested in ranking. The only thing I want is a logical explanation about how a CNN can learn camera based stats from the images taken by it. Had it been so simple, don't you think people would already have done that and there would have been no need to host this competition?",
      "votes": null
    },
    {
      "id": "271147",
      "postDate": "01/19/2018 19:24:50",
      "content": "<p>Each camera uses different optics, and those create very specific artifacts like shape of bokeh (google it up, it depends for example on aperture and focal length). Then you can have noise on different sensitivities (ISO). You may have different rolling shutter (how the image is captured when the camera is moving) - that you see the best for example when recording from inside the running train or the running train itself from the outside, normally straight lines get skewed. Also depending on the sensor each camera may have different color scale usually from 12 to 16 bits. Some of them may have build in stabilization and some not, that will also affect quality of picture. </p>\n\n<p>That's all about the theory, from the top of my head. :)</p>",
      "rawMarkdown": "Each camera uses different optics, and those create very specific artifacts like shape of bokeh (google it up, it depends for example on aperture and focal length). Then you can have noise on different sensitivities (ISO). You may have different rolling shutter (how the image is captured when the camera is moving) - that you see the best for example when recording from inside the running train or the running train itself from the outside, normally straight lines get skewed. Also depending on the sensor each camera may have different color scale usually from 12 to 16 bits. Some of them may have build in stabilization and some not, that will also affect quality of picture. \n\nThat's all about the theory, from the top of my head. :)",
      "votes": null
    },
    {
      "id": "271149",
      "postDate": "01/19/2018 19:27:06",
      "content": "<p>I would have forgotten, also different cameras usually produce different quality jpegs, but also raws if they can output them. </p>",
      "rawMarkdown": "I would have forgotten, also different cameras usually produce different quality jpegs, but also raws if they can output them.",
      "votes": null
    },
    {
      "id": "271159",
      "postDate": "01/19/2018 19:51:02",
      "content": "<p>Exactly. Now, the question comes down to this \"Can a CNN learn such filters?\" If yes, then, what changes are required in a normal CNN architecture to capture this kind of information and how do these filters look like?</p>",
      "rawMarkdown": "Exactly. Now, the question comes down to this \"Can a CNN learn such filters?\" If yes, then, what changes are required in a normal CNN architecture to capture this kind of information and how do these filters look like?",
      "votes": null
    },
    {
      "id": "271169",
      "postDate": "01/19/2018 20:30:26",
      "content": "<p>Hi @NAIN, a partial answer to why a CNN can work here.</p>\n\n<p>I guess we can agree that a convnet can detect a pattern that a human can. And a human can easily spot differences in focal length, for example. Straigth lines, paralel lines, horizon, will be affected by focal length so there is at least one \"macroscopic\" footprint of a lense. </p>\n\n<p>I am not saying that this CNNs are actually using this, just , in the line of reasoning you suggest, putting it as an example of something that a CNN can possibly learn, sure there can be more subtle and better patterns explaining that performance, but I wouldn't say it is so weird that CNNs are working :-) </p>",
      "rawMarkdown": "Hi @NAIN, a partial answer to why a CNN can work here.\n\n I guess we can agree that a convnet can detect a pattern that a human can. And a human can easily spot differences in focal length, for example. Straigth lines, paralel lines, horizon, will be affected by focal length so there is at least one \"macroscopic\" footprint of a lense. \n\nI am not saying that this CNNs are actually using this, just , in the line of reasoning you suggest, putting it as an example of something that a CNN can possibly learn, sure there can be more subtle and better patterns explaining that performance, but I wouldn't say it is so weird that CNNs are working :-)",
      "votes": null
    },
    {
      "id": "271356",
      "postDate": "01/20/2018 08:42:25",
      "content": "<p>Hi @Miguel, \nI am not saying that CNN wouldn't work here. My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.</p>",
      "rawMarkdown": "Hi @Miguel, \nI am not saying that CNN wouldn't work here. My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.",
      "votes": null
    },
    {
      "id": "271772",
      "postDate": "01/21/2018 11:42:00",
      "content": "<p>To know accuracy of one type -- hardcode the other type to say 'iPhone-6' and submit .. since the dataset is scored 0.3x + 0.7y... and since the data is balanced your hardcoded values will fetch you 0.03 or 0.07 respectively (depending on which set you hardcoded) and the rest is your actual score on the other set. </p>",
      "rawMarkdown": "To know accuracy of one type -- hardcode the other type to say 'iPhone-6' and submit .. since the dataset is scored 0.3x + 0.7y... and since the data is balanced your hardcoded values will fetch you 0.03 or 0.07 respectively (depending on which set you hardcoded) and the rest is your actual score on the other set.",
      "votes": null
    },
    {
      "id": "271793",
      "postDate": "01/21/2018 13:35:01",
      "content": "<p>More precisely, you can know the score of each ones by submitting blanks for the other part.</p>",
      "rawMarkdown": "More precisely, you can know the score of each ones by submitting blanks for the other part.",
      "votes": null
    },
    {
      "id": "272220",
      "postDate": "01/22/2018 15:06:05",
      "content": "<p>I agree with you. I tried using Resnet/Inception and got some results, but now I'm trying out my own architecture. In my thoughts this information coming from the filters of the camera are very sensitive to use very deep networks and does not make sense fine tuning with Imagenet weigths.</p>\n\n<p>But I didn't get good results so maybe I'm totally wrong xD</p>",
      "rawMarkdown": "I agree with you. I tried using Resnet/Inception and got some results, but now I'm trying out my own architecture. In my thoughts this information coming from the filters of the camera are very sensitive to use very deep networks and does not make sense fine tuning with Imagenet weigths.\n\nBut I didn't get good results so maybe I'm totally wrong xD",
      "votes": null
    },
    {
      "id": "272378",
      "postDate": "01/23/2018 00:02:33",
      "content": "<p>Quote:</p>\n\n<blockquote>\n  <p>My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.</p>\n  \n  <p>Had it been so simple, don't you think people would already have done\n  that and there would have been no need to host this competition?</p>\n</blockquote>\n\n<p>Lets use reason and evidence. There's an absolute tonne of primary research on using CNNs on camera source identification and specifically, pre-trained deep networks like AlexNet, VGG etc..</p>\n\n<p>Simply CNNs will learn high-pass filters that can capture PRNU and other Fixed Pattern Noise (FPN). Your 'intuition' that deep pre-trained nets won't work well is simply wrong. Deep ImageNet pre-trained networks have been shown to work well on this problem (A non-exhaustive search):</p>\n\n<p><a href=\"http://ieeexplore.ieee.org/abstract/document/7823908/\">http://ieeexplore.ieee.org/abstract/document/7823908/</a></p>\n\n<p>^\nThey show DeepNets actually match if not outperform their custom CNN specific to the task of camera source identification (Image below) - So even if you can possibly design the perfect CNN for the task - you'll probably do just as well using a SOTA pre-trained net.</p>\n\n<p><img src=\"http://i67.tinypic.com/16730c4.png\" alt=\"From the paper above\"></p>\n\n<p><a href=\"http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN.pdf\">http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN.pdf</a></p>\n\n<p><a href=\"http://openaccess.thecvf.com/content_cvpr_2017_workshops/w28/papers/Delp_A_Counter-Forensic_Method_CVPR_2017_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2017_workshops/w28/papers/Delp_A_Counter-Forensic_Method_CVPR_2017_paper.pdf</a></p>\n\n<p><a href=\"https://arxiv.org/abs/1710.01257\">https://arxiv.org/abs/1710.01257</a></p>\n\n<p><a href=\"https://arxiv.org/pdf/1703.04856.pdf\">https://arxiv.org/pdf/1703.04856.pdf</a></p>",
      "rawMarkdown": "Quote:\n&gt; My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.\n\n&gt; Had it been so simple, don't you think people would already have done\n&gt; that and there would have been no need to host this competition?\n\nLets use reason and evidence. There's an absolute tonne of primary research on using CNNs on camera source identification and specifically, pre-trained deep networks like AlexNet, VGG etc..\n\nSimply CNNs will learn high-pass filters that can capture PRNU and other Fixed Pattern Noise (FPN). Your 'intuition' that deep pre-trained nets won't work well is simply wrong. Deep ImageNet pre-trained networks have been shown to work well on this problem (A non-exhaustive search):\n\nhttp://ieeexplore.ieee.org/abstract/document/7823908/\n\n^\nThey show DeepNets actually match if not outperform their custom CNN specific to the task of camera source identification (Image below) - So even if you can possibly design the perfect CNN for the task - you'll probably do just as well using a SOTA pre-trained net.\n\n![From the paper above][1]\n\n\nhttp://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN.pdf\n\nhttp://openaccess.thecvf.com/content_cvpr_2017_workshops/w28/papers/Delp_A_Counter-Forensic_Method_CVPR_2017_paper.pdf\n\nhttps://arxiv.org/abs/1710.01257\n\nhttps://arxiv.org/pdf/1703.04856.pdf\n\n\n  [1]: http://i67.tinypic.com/16730c4.png",
      "votes": null
    },
    {
      "id": "272698",
      "postDate": "01/23/2018 14:27:52",
      "content": "<p>Great summary. </p>\n\n<p>I implemented CA-CNN as per this paper: <a href=\"https://arxiv.org/pdf/1703.04856.pdf\">https://arxiv.org/pdf/1703.04856.pdf</a> here: <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/commit/d8643142bfdcb0f34e4811aff47af88d371f5c27#diff-e44f4a60e89e820dde9bb27afa634965R320\">https://github.com/antorsae/sp-society-camera-model-identification/commit/d8643142bfdcb0f34e4811aff47af88d371f5c27#diff-e44f4a60e89e820dde9bb27afa634965R320</a></p>\n\n<p>Two caveats:</p>\n\n<ul>\n<li>I also feed a 0,1. value as input to the FC layer depending on wether there's been a manipulation or not.</li>\n<li>The learned 3x3, 5x5 and 7x7 filters (Fig 3) are actually NxNx3x3 filters. I am not clear how to reuse a kernel (e.g. 3x3) for 3 channels (RGB) outputting 3 channels (RGB).</li>\n</ul>\n\n<p>But Im still at a loss as to why such discrepancy between val and test. </p>",
      "rawMarkdown": "Great summary. \n\nI implemented CA-CNN as per this paper: https://arxiv.org/pdf/1703.04856.pdf here: https://github.com/antorsae/sp-society-camera-model-identification/commit/d8643142bfdcb0f34e4811aff47af88d371f5c27#diff-e44f4a60e89e820dde9bb27afa634965R320\n\nTwo caveats:\n\n - I also feed a 0,1. value as input to the FC layer depending on wether there's been a manipulation or not.\n - The learned 3x3, 5x5 and 7x7 filters (Fig 3) are actually NxNx3x3 filters. I am not clear how to reuse a kernel (e.g. 3x3) for 3 channels (RGB) outputting 3 channels (RGB).\n\nBut Im still at a loss as to why such discrepancy between val and test.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 270756,
      "author_name": "hamyadlab",
      "author_url": "",
      "post_date": "01/18/2018 21:33:21",
      "content": "<p>So, what do you suggest for this task?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270888,
      "author_name": "datasec",
      "author_url": "",
      "post_date": "01/19/2018 06:48:03",
      "content": "<p>The way I like to think about this is and I might be totally wrong, we know that NNs are universal approximators and here we have a task to approximate whatever complex composition of functions that are happening inside the cameras include demosacing, denoising, compression etc. Now we could do this with a fully connected network but it would be difficult to train with SGD and related optimization algorithms. The only prior assumptions with using convolution layers are local relationships, translation equivariance for the functions learned (and some translation invariance due to pooling) which are presumably correct or do no harm for the task at hand.  Using convolutions and pooling are also lowering the memory footprint significantly and some computation and time cost so a win-win for all related grid based tasks including time series data, 1D conv for audio etc. The use of existing architecture is giving a proven arrangement of layers, some useful shared weights from imagenet (no idea how) and possibly a better than random initialisation if you proceed with training all the layers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 271070,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "01/19/2018 16:18:41",
          "content": "<p>I know that NN can approximate any mathematical function. The thing is you are trying to explain NNs to me which I know very well. Also, you are treating a CNN like a black box, more precisely a magical box. I understand convolution, pooling, etc, I am trying to understand how a CNN can learn things like compression, focal length, etc from a photo</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270958,
      "author_name": "shuhuagao",
      "author_url": "",
      "post_date": "01/19/2018 10:33:43",
      "content": "<p>However, it seems that CNN indeed worked. Please check this discussion <a href=\"https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/46688\">https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/46688</a>, where the author has ranked 3rd. </p>",
      "votes": null,
      "replies": [
        {
          "id": 271071,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "01/19/2018 16:20:47",
          "content": "<p>I am not interested in ranking. The only thing I want is a logical explanation about how a CNN can learn camera based stats from the images taken by it. Had it been so simple, don't you think people would already have done that and there would have been no need to host this competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271147,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "01/19/2018 19:24:50",
          "content": "<p>Each camera uses different optics, and those create very specific artifacts like shape of bokeh (google it up, it depends for example on aperture and focal length). Then you can have noise on different sensitivities (ISO). You may have different rolling shutter (how the image is captured when the camera is moving) - that you see the best for example when recording from inside the running train or the running train itself from the outside, normally straight lines get skewed. Also depending on the sensor each camera may have different color scale usually from 12 to 16 bits. Some of them may have build in stabilization and some not, that will also affect quality of picture. </p>\n\n<p>That's all about the theory, from the top of my head. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271149,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "01/19/2018 19:27:06",
          "content": "<p>I would have forgotten, also different cameras usually produce different quality jpegs, but also raws if they can output them. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271159,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "01/19/2018 19:51:02",
          "content": "<p>Exactly. Now, the question comes down to this \"Can a CNN learn such filters?\" If yes, then, what changes are required in a normal CNN architecture to capture this kind of information and how do these filters look like?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272378,
          "author_name": "craigglastonbury",
          "author_url": "",
          "post_date": "01/23/2018 00:02:33",
          "content": "<p>Quote:</p>\n\n<blockquote>\n  <p>My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.</p>\n  \n  <p>Had it been so simple, don't you think people would already have done\n  that and there would have been no need to host this competition?</p>\n</blockquote>\n\n<p>Lets use reason and evidence. There's an absolute tonne of primary research on using CNNs on camera source identification and specifically, pre-trained deep networks like AlexNet, VGG etc..</p>\n\n<p>Simply CNNs will learn high-pass filters that can capture PRNU and other Fixed Pattern Noise (FPN). Your 'intuition' that deep pre-trained nets won't work well is simply wrong. Deep ImageNet pre-trained networks have been shown to work well on this problem (A non-exhaustive search):</p>\n\n<p><a href=\"http://ieeexplore.ieee.org/abstract/document/7823908/\">http://ieeexplore.ieee.org/abstract/document/7823908/</a></p>\n\n<p>^\nThey show DeepNets actually match if not outperform their custom CNN specific to the task of camera source identification (Image below) - So even if you can possibly design the perfect CNN for the task - you'll probably do just as well using a SOTA pre-trained net.</p>\n\n<p><img src=\"http://i67.tinypic.com/16730c4.png\" alt=\"From the paper above\"></p>\n\n<p><a href=\"http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN.pdf\">http://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN.pdf</a></p>\n\n<p><a href=\"http://openaccess.thecvf.com/content_cvpr_2017_workshops/w28/papers/Delp_A_Counter-Forensic_Method_CVPR_2017_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2017_workshops/w28/papers/Delp_A_Counter-Forensic_Method_CVPR_2017_paper.pdf</a></p>\n\n<p><a href=\"https://arxiv.org/abs/1710.01257\">https://arxiv.org/abs/1710.01257</a></p>\n\n<p><a href=\"https://arxiv.org/pdf/1703.04856.pdf\">https://arxiv.org/pdf/1703.04856.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272698,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/23/2018 14:27:52",
          "content": "<p>Great summary. </p>\n\n<p>I implemented CA-CNN as per this paper: <a href=\"https://arxiv.org/pdf/1703.04856.pdf\">https://arxiv.org/pdf/1703.04856.pdf</a> here: <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/commit/d8643142bfdcb0f34e4811aff47af88d371f5c27#diff-e44f4a60e89e820dde9bb27afa634965R320\">https://github.com/antorsae/sp-society-camera-model-identification/commit/d8643142bfdcb0f34e4811aff47af88d371f5c27#diff-e44f4a60e89e820dde9bb27afa634965R320</a></p>\n\n<p>Two caveats:</p>\n\n<ul>\n<li>I also feed a 0,1. value as input to the FC layer depending on wether there's been a manipulation or not.</li>\n<li>The learned 3x3, 5x5 and 7x7 filters (Fig 3) are actually NxNx3x3 filters. I am not clear how to reuse a kernel (e.g. 3x3) for 3 channels (RGB) outputting 3 channels (RGB).</li>\n</ul>\n\n<p>But Im still at a loss as to why such discrepancy between val and test. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270971,
      "author_name": "mxdbld",
      "author_url": "",
      "post_date": "01/19/2018 10:50:58",
      "content": "<p>I have 97.2% accuracy on LB on unaltered pictures with one CNN. They do work but it is not the only possible approach. Abhishek does give an explanation of why they could perform</p>",
      "votes": null,
      "replies": [
        {
          "id": 271000,
          "author_name": "shuhuagao",
          "author_url": "",
          "post_date": "01/19/2018 12:29:58",
          "content": "<p>Hi, Max. You said \"I have 97.2% accuracy on LB on unaltered pictures\". In the test set, there are both unaltered and manipulated images, how did you know you accuracy only on unaltered ones?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271066,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "01/19/2018 16:14:26",
          "content": "<p>Wait for Private LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271772,
          "author_name": "",
          "author_url": "",
          "post_date": "01/21/2018 11:42:00",
          "content": "<p>To know accuracy of one type -- hardcode the other type to say 'iPhone-6' and submit .. since the dataset is scored 0.3x + 0.7y... and since the data is balanced your hardcoded values will fetch you 0.03 or 0.07 respectively (depending on which set you hardcoded) and the rest is your actual score on the other set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271793,
          "author_name": "mxdbld",
          "author_url": "",
          "post_date": "01/21/2018 13:35:01",
          "content": "<p>More precisely, you can know the score of each ones by submitting blanks for the other part.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 271169,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "01/19/2018 20:30:26",
      "content": "<p>Hi @NAIN, a partial answer to why a CNN can work here.</p>\n\n<p>I guess we can agree that a convnet can detect a pattern that a human can. And a human can easily spot differences in focal length, for example. Straigth lines, paralel lines, horizon, will be affected by focal length so there is at least one \"macroscopic\" footprint of a lense. </p>\n\n<p>I am not saying that this CNNs are actually using this, just , in the line of reasoning you suggest, putting it as an example of something that a CNN can possibly learn, sure there can be more subtle and better patterns explaining that performance, but I wouldn't say it is so weird that CNNs are working :-) </p>",
      "votes": null,
      "replies": [
        {
          "id": 271356,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "01/20/2018 08:42:25",
          "content": "<p>Hi @Miguel, \nI am not saying that CNN wouldn't work here. My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 272220,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/22/2018 15:06:05",
          "content": "<p>I agree with you. I tried using Resnet/Inception and got some results, but now I'm trying out my own architecture. In my thoughts this information coming from the filters of the camera are very sensitive to use very deep networks and does not make sense fine tuning with Imagenet weigths.</p>\n\n<p>But I didn't get good results so maybe I'm totally wrong xD</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "270628": "CNNs are great but I am not able to digest the fact that people are applying simple \nCNNs for making predictions. IMO, training architectures like VGG16 and ResNet50 are of very little use. Those architectures, as far as I know, cannot capture so dense information  about the image. They are good for object classification, object detection, etc but when you want to predict the camera model from an image, you are making too much of assumptions that the network can learn by itself.",
    "270756": "So, what do you suggest for this task?",
    "270888": "The way I like to think about this is and I might be totally wrong, we know that NNs are universal approximators and here we have a task to approximate whatever complex composition of functions that are happening inside the cameras include demosacing, denoising, compression etc. Now we could do this with a fully connected network but it would be difficult to train with SGD and related optimization algorithms. The only prior assumptions with using convolution layers are local relationships, translation equivariance for the functions learned (and some translation invariance due to pooling) which are presumably correct or do no harm for the task at hand.  Using convolutions and pooling are also lowering the memory footprint significantly and some computation and time cost so a win-win for all related grid based tasks including time series data, 1D conv for audio etc. The use of existing architecture is giving a proven arrangement of layers, some useful shared weights from imagenet (no idea how) and possibly a better than random initialisation if you proceed with training all the layers.",
    "270958": "However, it seems that CNN indeed worked. Please check this discussion https://www.kaggle.com/c/sp-society-camera-model-identification/discussion/46688, where the author has ranked 3rd.",
    "270971": "I have 97.2% accuracy on LB on unaltered pictures with one CNN. They do work but it is not the only possible approach. Abhishek does give an explanation of why they could perform",
    "271000": "Hi, Max. You said \"I have 97.2% accuracy on LB on unaltered pictures\". In the test set, there are both unaltered and manipulated images, how did you know you accuracy only on unaltered ones?",
    "271066": "Wait for Private LB",
    "271070": "I know that NN can approximate any mathematical function. The thing is you are trying to explain NNs to me which I know very well. Also, you are treating a CNN like a black box, more precisely a magical box. I understand convolution, pooling, etc, I am trying to understand how a CNN can learn things like compression, focal length, etc from a photo",
    "271071": "I am not interested in ranking. The only thing I want is a logical explanation about how a CNN can learn camera based stats from the images taken by it. Had it been so simple, don't you think people would already have done that and there would have been no need to host this competition?",
    "271147": "Each camera uses different optics, and those create very specific artifacts like shape of bokeh (google it up, it depends for example on aperture and focal length). Then you can have noise on different sensitivities (ISO). You may have different rolling shutter (how the image is captured when the camera is moving) - that you see the best for example when recording from inside the running train or the running train itself from the outside, normally straight lines get skewed. Also depending on the sensor each camera may have different color scale usually from 12 to 16 bits. Some of them may have build in stabilization and some not, that will also affect quality of picture. \n\nThat's all about the theory, from the top of my head. :)",
    "271149": "I would have forgotten, also different cameras usually produce different quality jpegs, but also raws if they can output them.",
    "271159": "Exactly. Now, the question comes down to this \"Can a CNN learn such filters?\" If yes, then, what changes are required in a normal CNN architecture to capture this kind of information and how do these filters look like?",
    "271169": "Hi @NAIN, a partial answer to why a CNN can work here.\n\n I guess we can agree that a convnet can detect a pattern that a human can. And a human can easily spot differences in focal length, for example. Straigth lines, paralel lines, horizon, will be affected by focal length so there is at least one \"macroscopic\" footprint of a lense. \n\nI am not saying that this CNNs are actually using this, just , in the line of reasoning you suggest, putting it as an example of something that a CNN can possibly learn, sure there can be more subtle and better patterns explaining that performance, but I wouldn't say it is so weird that CNNs are working :-)",
    "271356": "Hi @Miguel, \nI am not saying that CNN wouldn't work here. My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.",
    "271772": "To know accuracy of one type -- hardcode the other type to say 'iPhone-6' and submit .. since the dataset is scored 0.3x + 0.7y... and since the data is balanced your hardcoded values will fetch you 0.03 or 0.07 respectively (depending on which set you hardcoded) and the rest is your actual score on the other set.",
    "271793": "More precisely, you can know the score of each ones by submitting blanks for the other part.",
    "272220": "I agree with you. I tried using Resnet/Inception and got some results, but now I'm trying out my own architecture. In my thoughts this information coming from the filters of the camera are very sensitive to use very deep networks and does not make sense fine tuning with Imagenet weigths.\n\nBut I didn't get good results so maybe I'm totally wrong xD",
    "272378": "Quote:\n&gt; My point is that CNNs like VGG/ResNet are not the ones that are well suited for this task. In the end, the final model will be based on CNN for sure, no doubt in that, but we need to design a different type architecture for that which of course requires a lot of hypothesis as well as time.\n\n&gt; Had it been so simple, don't you think people would already have done\n&gt; that and there would have been no need to host this competition?\n\nLets use reason and evidence. There's an absolute tonne of primary research on using CNNs on camera source identification and specifically, pre-trained deep networks like AlexNet, VGG etc..\n\nSimply CNNs will learn high-pass filters that can capture PRNU and other Fixed Pattern Noise (FPN). Your 'intuition' that deep pre-trained nets won't work well is simply wrong. Deep ImageNet pre-trained networks have been shown to work well on this problem (A non-exhaustive search):\n\nhttp://ieeexplore.ieee.org/abstract/document/7823908/\n\n^\nThey show DeepNets actually match if not outperform their custom CNN specific to the task of camera source identification (Image below) - So even if you can possibly design the perfect CNN for the task - you'll probably do just as well using a SOTA pre-trained net.\n\n![From the paper above][1]\n\n\nhttp://www.lirmm.fr/~chaumont/publications/WIFS-2016_TUAMA_COMBY_CHAUMONT_Camera_Model_Identification_With_CNN.pdf\n\nhttp://openaccess.thecvf.com/content_cvpr_2017_workshops/w28/papers/Delp_A_Counter-Forensic_Method_CVPR_2017_paper.pdf\n\nhttps://arxiv.org/abs/1710.01257\n\nhttps://arxiv.org/pdf/1703.04856.pdf\n\n\n  [1]: http://i67.tinypic.com/16730c4.png",
    "272698": "Great summary. \n\nI implemented CA-CNN as per this paper: https://arxiv.org/pdf/1703.04856.pdf here: https://github.com/antorsae/sp-society-camera-model-identification/commit/d8643142bfdcb0f34e4811aff47af88d371f5c27#diff-e44f4a60e89e820dde9bb27afa634965R320\n\nTwo caveats:\n\n - I also feed a 0,1. value as input to the FC layer depending on wether there's been a manipulation or not.\n - The learned 3x3, 5x5 and 7x7 filters (Fig 3) are actually NxNx3x3 filters. I am not clear how to reuse a kernel (e.g. 3x3) for 3 channels (RGB) outputting 3 channels (RGB).\n\nBut Im still at a loss as to why such discrepancy between val and test."
  },
  "source": "meta"
}