{
  "id": 49343,
  "title": "How to deal with manipulated images/data augmentation with resize?",
  "url": "/competitions/sp-society-camera-model-identification/discussion/49343",
  "author_name": "",
  "post_date": "2018-02-09T16:50:52.124696600Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I realized that most of the top solutions used the same trained model to predict both manipulated and unaltered images. In my team's case, we trained the same ensemble model on four different input image scenarios: pristine images, jpeg recompressed (QFs 70 and 90), gamma corrected (0.8 and 1.2) and resize (0.5, 0.8, 1.5 and 2.0). Then, we used a CNN trained to detect these four operations, redirecting the testing sample to its corresponding ensemble, according to the detected manipulation. Although this strategy resulted in a good accuracy, we had a serious problem with very long training time. With flickr additional images shared here on Kaggle, a single resize-operation network took weeks to train (as there are 4 parameters for this operation) and, unfortunately, we didn't finish training our ensemble approach for the altered images on time, using for that just the best individual CNN.  So, my question is: did anybody else try this solution and could finish training ensembles on time? is this better than considering all the operations in the data augmentation task? I really hope we were not the only ones who thought about that solution (LOL).</p>\n\n<p>On the other hand, I see that teams used the data augmentation with the 8 possible manipulations, so the same model could be tested on both original and manipulated images. However, one of these operations is the resizing, which will CHANGE the training sample size. So, how is that possible to train a CNN on resized images? I mean, if my network uses 512x512 patches as input, how is possible to use patches of these different resized sizes after the data augmentation as input for a CNN training? </p>\n\n<p>Finally, one final question, I didn't make the data augmentation on the validation images, is this a good strategy? or depends on the problem? my early stopping prediction was reducing the validation loss.</p>\n\n<p>Thank you very much and I must say that I am feeling a winner because everything I am learning so far!</p>",
  "messages": [
    {
      "id": "280258",
      "postDate": "02/09/2018 16:50:52",
      "content": "<p>I realized that most of the top solutions used the same trained model to predict both manipulated and unaltered images. In my team's case, we trained the same ensemble model on four different input image scenarios: pristine images, jpeg recompressed (QFs 70 and 90), gamma corrected (0.8 and 1.2) and resize (0.5, 0.8, 1.5 and 2.0). Then, we used a CNN trained to detect these four operations, redirecting the testing sample to its corresponding ensemble, according to the detected manipulation. Although this strategy resulted in a good accuracy, we had a serious problem with very long training time. With flickr additional images shared here on Kaggle, a single resize-operation network took weeks to train (as there are 4 parameters for this operation) and, unfortunately, we didn't finish training our ensemble approach for the altered images on time, using for that just the best individual CNN.  So, my question is: did anybody else try this solution and could finish training ensembles on time? is this better than considering all the operations in the data augmentation task? I really hope we were not the only ones who thought about that solution (LOL).</p>\n\n<p>On the other hand, I see that teams used the data augmentation with the 8 possible manipulations, so the same model could be tested on both original and manipulated images. However, one of these operations is the resizing, which will CHANGE the training sample size. So, how is that possible to train a CNN on resized images? I mean, if my network uses 512x512 patches as input, how is possible to use patches of these different resized sizes after the data augmentation as input for a CNN training? </p>\n\n<p>Finally, one final question, I didn't make the data augmentation on the validation images, is this a good strategy? or depends on the problem? my early stopping prediction was reducing the validation loss.</p>\n\n<p>Thank you very much and I must say that I am feeling a winner because everything I am learning so far!</p>",
      "rawMarkdown": "I realized that most of the top solutions used the same trained model to predict both manipulated and unaltered images. In my team's case, we trained the same ensemble model on four different input image scenarios: pristine images, jpeg recompressed (QFs 70 and 90), gamma corrected (0.8 and 1.2) and resize (0.5, 0.8, 1.5 and 2.0). Then, we used a CNN trained to detect these four operations, redirecting the testing sample to its corresponding ensemble, according to the detected manipulation. Although this strategy resulted in a good accuracy, we had a serious problem with very long training time. With flickr additional images shared here on Kaggle, a single resize-operation network took weeks to train (as there are 4 parameters for this operation) and, unfortunately, we didn't finish training our ensemble approach for the altered images on time, using for that just the best individual CNN.  So, my question is: did anybody else try this solution and could finish training ensembles on time? is this better than considering all the operations in the data augmentation task? I really hope we were not the only ones who thought about that solution (LOL).\n\nOn the other hand, I see that teams used the data augmentation with the 8 possible manipulations, so the same model could be tested on both original and manipulated images. However, one of these operations is the resizing, which will CHANGE the training sample size. So, how is that possible to train a CNN on resized images? I mean, if my network uses 512x512 patches as input, how is possible to use patches of these different resized sizes after the data augmentation as input for a CNN training? \n\nFinally, one final question, I didn't make the data augmentation on the validation images, is this a good strategy? or depends on the problem? my early stopping prediction was reducing the validation loss.\n\nThank you very much and I must say that I am feeling a winner because everything I am learning so far!",
      "votes": null
    },
    {
      "id": "280560",
      "postDate": "02/10/2018 08:59:55",
      "content": "<p>Hi Anselmo,</p>\n\n<p>My best single-model used a two-output prediction: 1) camera 2) manipulation. The loss function I used was simply: x-entropy(camera_true, camera_pred) + 0.1 x-entropy(manipulation_true, manipulation_pred). Upon training, I feed images with all possible manipulations (including no manipulation) with same probability.</p>\n\n<p>There's two FC heads, and the output of the manipulation prediction FC head is concatenated to the input of the main camera FC head; so that the FC responsible for detecting camera type has information about the type of manipulation used. Validation accuracy for predicting manipulation was 0.98+.</p>\n\n<p>Upon inference/testing, I just ignore the output of the manipulation prediction.</p>\n\n<p>Re: data augmentation on validation images; well at least you should do manipulations, b/c in theory your validation images should be unaltered to start with.</p>\n\n<p>Re: feeding CNN with different input sizes... you just crop when the image is too big (the organization was doing that for bicubic scaling &gt;=1). More generally, conceptually is possible with a fully convolutional NN, not sure if doable in Keras, I'd go with dynamic graphs in pytorch if I had to do it.</p>",
      "rawMarkdown": "Hi Anselmo,\n\nMy best single-model used a two-output prediction: 1) camera 2) manipulation. The loss function I used was simply: x-entropy(camera_true, camera_pred) + 0.1 x-entropy(manipulation_true, manipulation_pred). Upon training, I feed images with all possible manipulations (including no manipulation) with same probability.\n\nThere's two FC heads, and the output of the manipulation prediction FC head is concatenated to the input of the main camera FC head; so that the FC responsible for detecting camera type has information about the type of manipulation used. Validation accuracy for predicting manipulation was 0.98+.\n\nUpon inference/testing, I just ignore the output of the manipulation prediction.\n\nRe: data augmentation on validation images; well at least you should do manipulations, b/c in theory your validation images should be unaltered to start with.\n\nRe: feeding CNN with different input sizes... you just crop when the image is too big (the organization was doing that for bicubic scaling &gt;=1). More generally, conceptually is possible with a fully convolutional NN, not sure if doable in Keras, I'd go with dynamic graphs in pytorch if I had to do it.",
      "votes": null
    },
    {
      "id": "280816",
      "postDate": "02/11/2018 05:26:12",
      "content": "<p>Hello Andres,</p>\n\n<p>let me ask you something, it seems that you are using just one trained model to predict either 'manip' and 'unalt' images. So, why predicting the manipulation? do the manipulation predictions in your training stage help adjusting your single network weights to better classify 'manip' images according to your loss function? Have you tried using the common loss function, no caring about the manipulation present on the training images? was it worse?</p>\n\n<p>Another question: did you try data augmentation on validation data for your solution? did you use any early stopping criteria on the validation data?</p>\n\n<p>Regarding the data augmentation, you are saying to me that we can simply apply the manipulations (8) in the full resolution training set, using the generated 8 times bigger dataset to extract patches, correct? when I asked the question I was thinking about not saving images on disk, loading them at training time.</p>\n\n<p>So, sorry for this stupid question, but is there any procedure which avoids me to save images on disk, doing data augmentation at training-time? I mean, for keras there is the ImageDataGenerator, but it could work only for gamma and jpeg, not for resize. How did you do it?</p>\n\n<p>Again, thank you so much.</p>",
      "rawMarkdown": "Hello Andres,\n\nlet me ask you something, it seems that you are using just one trained model to predict either 'manip' and 'unalt' images. So, why predicting the manipulation? do the manipulation predictions in your training stage help adjusting your single network weights to better classify 'manip' images according to your loss function? Have you tried using the common loss function, no caring about the manipulation present on the training images? was it worse?\n\nAnother question: did you try data augmentation on validation data for your solution? did you use any early stopping criteria on the validation data?\n\nRegarding the data augmentation, you are saying to me that we can simply apply the manipulations (8) in the full resolution training set, using the generated 8 times bigger dataset to extract patches, correct? when I asked the question I was thinking about not saving images on disk, loading them at training time.\n\nSo, sorry for this stupid question, but is there any procedure which avoids me to save images on disk, doing data augmentation at training-time? I mean, for keras there is the ImageDataGenerator, but it could work only for gamma and jpeg, not for resize. How did you do it?\n\nAgain, thank you so much.",
      "votes": null
    },
    {
      "id": "280880",
      "postDate": "02/11/2018 09:44:19",
      "content": "<p>Hi Anselmo,</p>\n\n<p>My best single model (not TTA) used the two heads. In general, when doing NN experiments I always try to force the network to learn secondary tasks which may help the main task. Take the case of bicubic interpolation by 2x (i.e. each pixel in the original image will now be ~2x2 pixels): the convolutional layers will supposedly learn PRNU (?) patterns for this case, but I thought they may pick up unrelated features from other conv filters that have converged for no bicubic (1x1), so I thought telling the main FC if there's a manipulation going on it may help it... hence another FC to predict manipulation type (9 in total) which is now an output to the model as well as an input to the main FC head.</p>\n\n<blockquote>\n  <p>Another question: did you try data augmentation on validation data for\n  your solution? did you use any early stopping criteria on the\n  validation data?</p>\n</blockquote>\n\n<p>Yes to both questions. You can check it in my code.</p>\n\n<blockquote>\n  <p>Regarding the data augmentation, you are saying to me that we can\n  simply apply the manipulations (8) in the full resolution training\n  set, using the generated 8 times bigger dataset to extract patches,\n  correct? when I asked the question I was thinking about not saving\n  images on disk, loading them at training time.</p>\n</blockquote>\n\n<p>You could do either. I didn't want to do preprocessing, so I do it in realtime. If you stick to Keras <code>ImageDataGenerator</code> you're going to be limited. Check my code, there's few or no comments ;-)</p>\n\n<p>Best - Andres</p>",
      "rawMarkdown": "Hi Anselmo,\n\nMy best single model (not TTA) used the two heads. In general, when doing NN experiments I always try to force the network to learn secondary tasks which may help the main task. Take the case of bicubic interpolation by 2x (i.e. each pixel in the original image will now be ~2x2 pixels): the convolutional layers will supposedly learn PRNU (?) patterns for this case, but I thought they may pick up unrelated features from other conv filters that have converged for no bicubic (1x1), so I thought telling the main FC if there's a manipulation going on it may help it... hence another FC to predict manipulation type (9 in total) which is now an output to the model as well as an input to the main FC head.\n\n&gt; Another question: did you try data augmentation on validation data for\n&gt; your solution? did you use any early stopping criteria on the\n&gt; validation data?\n\nYes to both questions. You can check it in my code.\n\n&gt; Regarding the data augmentation, you are saying to me that we can\n&gt; simply apply the manipulations (8) in the full resolution training\n&gt; set, using the generated 8 times bigger dataset to extract patches,\n&gt; correct? when I asked the question I was thinking about not saving\n&gt; images on disk, loading them at training time.\n\nYou could do either. I didn't want to do preprocessing, so I do it in realtime. If you stick to Keras `ImageDataGenerator` you're going to be limited. Check my code, there's few or no comments ;-)\n\nBest - Andres",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 280560,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/10/2018 08:59:55",
      "content": "<p>Hi Anselmo,</p>\n\n<p>My best single-model used a two-output prediction: 1) camera 2) manipulation. The loss function I used was simply: x-entropy(camera_true, camera_pred) + 0.1 x-entropy(manipulation_true, manipulation_pred). Upon training, I feed images with all possible manipulations (including no manipulation) with same probability.</p>\n\n<p>There's two FC heads, and the output of the manipulation prediction FC head is concatenated to the input of the main camera FC head; so that the FC responsible for detecting camera type has information about the type of manipulation used. Validation accuracy for predicting manipulation was 0.98+.</p>\n\n<p>Upon inference/testing, I just ignore the output of the manipulation prediction.</p>\n\n<p>Re: data augmentation on validation images; well at least you should do manipulations, b/c in theory your validation images should be unaltered to start with.</p>\n\n<p>Re: feeding CNN with different input sizes... you just crop when the image is too big (the organization was doing that for bicubic scaling &gt;=1). More generally, conceptually is possible with a fully convolutional NN, not sure if doable in Keras, I'd go with dynamic graphs in pytorch if I had to do it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 280816,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/11/2018 05:26:12",
          "content": "<p>Hello Andres,</p>\n\n<p>let me ask you something, it seems that you are using just one trained model to predict either 'manip' and 'unalt' images. So, why predicting the manipulation? do the manipulation predictions in your training stage help adjusting your single network weights to better classify 'manip' images according to your loss function? Have you tried using the common loss function, no caring about the manipulation present on the training images? was it worse?</p>\n\n<p>Another question: did you try data augmentation on validation data for your solution? did you use any early stopping criteria on the validation data?</p>\n\n<p>Regarding the data augmentation, you are saying to me that we can simply apply the manipulations (8) in the full resolution training set, using the generated 8 times bigger dataset to extract patches, correct? when I asked the question I was thinking about not saving images on disk, loading them at training time.</p>\n\n<p>So, sorry for this stupid question, but is there any procedure which avoids me to save images on disk, doing data augmentation at training-time? I mean, for keras there is the ImageDataGenerator, but it could work only for gamma and jpeg, not for resize. How did you do it?</p>\n\n<p>Again, thank you so much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280880,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/11/2018 09:44:19",
          "content": "<p>Hi Anselmo,</p>\n\n<p>My best single model (not TTA) used the two heads. In general, when doing NN experiments I always try to force the network to learn secondary tasks which may help the main task. Take the case of bicubic interpolation by 2x (i.e. each pixel in the original image will now be ~2x2 pixels): the convolutional layers will supposedly learn PRNU (?) patterns for this case, but I thought they may pick up unrelated features from other conv filters that have converged for no bicubic (1x1), so I thought telling the main FC if there's a manipulation going on it may help it... hence another FC to predict manipulation type (9 in total) which is now an output to the model as well as an input to the main FC head.</p>\n\n<blockquote>\n  <p>Another question: did you try data augmentation on validation data for\n  your solution? did you use any early stopping criteria on the\n  validation data?</p>\n</blockquote>\n\n<p>Yes to both questions. You can check it in my code.</p>\n\n<blockquote>\n  <p>Regarding the data augmentation, you are saying to me that we can\n  simply apply the manipulations (8) in the full resolution training\n  set, using the generated 8 times bigger dataset to extract patches,\n  correct? when I asked the question I was thinking about not saving\n  images on disk, loading them at training time.</p>\n</blockquote>\n\n<p>You could do either. I didn't want to do preprocessing, so I do it in realtime. If you stick to Keras <code>ImageDataGenerator</code> you're going to be limited. Check my code, there's few or no comments ;-)</p>\n\n<p>Best - Andres</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "280258": "I realized that most of the top solutions used the same trained model to predict both manipulated and unaltered images. In my team's case, we trained the same ensemble model on four different input image scenarios: pristine images, jpeg recompressed (QFs 70 and 90), gamma corrected (0.8 and 1.2) and resize (0.5, 0.8, 1.5 and 2.0). Then, we used a CNN trained to detect these four operations, redirecting the testing sample to its corresponding ensemble, according to the detected manipulation. Although this strategy resulted in a good accuracy, we had a serious problem with very long training time. With flickr additional images shared here on Kaggle, a single resize-operation network took weeks to train (as there are 4 parameters for this operation) and, unfortunately, we didn't finish training our ensemble approach for the altered images on time, using for that just the best individual CNN.  So, my question is: did anybody else try this solution and could finish training ensembles on time? is this better than considering all the operations in the data augmentation task? I really hope we were not the only ones who thought about that solution (LOL).\n\nOn the other hand, I see that teams used the data augmentation with the 8 possible manipulations, so the same model could be tested on both original and manipulated images. However, one of these operations is the resizing, which will CHANGE the training sample size. So, how is that possible to train a CNN on resized images? I mean, if my network uses 512x512 patches as input, how is possible to use patches of these different resized sizes after the data augmentation as input for a CNN training? \n\nFinally, one final question, I didn't make the data augmentation on the validation images, is this a good strategy? or depends on the problem? my early stopping prediction was reducing the validation loss.\n\nThank you very much and I must say that I am feeling a winner because everything I am learning so far!",
    "280560": "Hi Anselmo,\n\nMy best single-model used a two-output prediction: 1) camera 2) manipulation. The loss function I used was simply: x-entropy(camera_true, camera_pred) + 0.1 x-entropy(manipulation_true, manipulation_pred). Upon training, I feed images with all possible manipulations (including no manipulation) with same probability.\n\nThere's two FC heads, and the output of the manipulation prediction FC head is concatenated to the input of the main camera FC head; so that the FC responsible for detecting camera type has information about the type of manipulation used. Validation accuracy for predicting manipulation was 0.98+.\n\nUpon inference/testing, I just ignore the output of the manipulation prediction.\n\nRe: data augmentation on validation images; well at least you should do manipulations, b/c in theory your validation images should be unaltered to start with.\n\nRe: feeding CNN with different input sizes... you just crop when the image is too big (the organization was doing that for bicubic scaling &gt;=1). More generally, conceptually is possible with a fully convolutional NN, not sure if doable in Keras, I'd go with dynamic graphs in pytorch if I had to do it.",
    "280816": "Hello Andres,\n\nlet me ask you something, it seems that you are using just one trained model to predict either 'manip' and 'unalt' images. So, why predicting the manipulation? do the manipulation predictions in your training stage help adjusting your single network weights to better classify 'manip' images according to your loss function? Have you tried using the common loss function, no caring about the manipulation present on the training images? was it worse?\n\nAnother question: did you try data augmentation on validation data for your solution? did you use any early stopping criteria on the validation data?\n\nRegarding the data augmentation, you are saying to me that we can simply apply the manipulations (8) in the full resolution training set, using the generated 8 times bigger dataset to extract patches, correct? when I asked the question I was thinking about not saving images on disk, loading them at training time.\n\nSo, sorry for this stupid question, but is there any procedure which avoids me to save images on disk, doing data augmentation at training-time? I mean, for keras there is the ImageDataGenerator, but it could work only for gamma and jpeg, not for resize. How did you do it?\n\nAgain, thank you so much.",
    "280880": "Hi Anselmo,\n\nMy best single model (not TTA) used the two heads. In general, when doing NN experiments I always try to force the network to learn secondary tasks which may help the main task. Take the case of bicubic interpolation by 2x (i.e. each pixel in the original image will now be ~2x2 pixels): the convolutional layers will supposedly learn PRNU (?) patterns for this case, but I thought they may pick up unrelated features from other conv filters that have converged for no bicubic (1x1), so I thought telling the main FC if there's a manipulation going on it may help it... hence another FC to predict manipulation type (9 in total) which is now an output to the model as well as an input to the main FC head.\n\n&gt; Another question: did you try data augmentation on validation data for\n&gt; your solution? did you use any early stopping criteria on the\n&gt; validation data?\n\nYes to both questions. You can check it in my code.\n\n&gt; Regarding the data augmentation, you are saying to me that we can\n&gt; simply apply the manipulations (8) in the full resolution training\n&gt; set, using the generated 8 times bigger dataset to extract patches,\n&gt; correct? when I asked the question I was thinking about not saving\n&gt; images on disk, loading them at training time.\n\nYou could do either. I didn't want to do preprocessing, so I do it in realtime. If you stick to Keras `ImageDataGenerator` you're going to be limited. Check my code, there's few or no comments ;-)\n\nBest - Andres"
  },
  "source": "meta"
}