{
  "id": 49298,
  "title": "4th place solution",
  "url": "/competitions/sp-society-camera-model-identification/writeups/guanshuo-xu-4th-place-solution",
  "author_name": "",
  "post_date": "2018-02-11T14:02:01.943Z",
  "votes": 30,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Here is a brief summary of my steps.</p>\n\n<p>Edit: I only used the central 80% crop of the train data because the boundaries are often statistically very different from the test data. For example, if the original image size is 1000x1000, only the central 800x800 crop is used. This center-cropping applies to train data only, and it gave around 1% higher accuracy than training on the original size.</p>\n\n<ol>\n<li><p>Finetune a pretrained inception_v3 with random 480x480 crops. The provided training set and Gleb's data were used. My data augmentation include the eight possible manipulations but no transpose, rotation or flipping as I believe they should not help in theory. JPEG compression is always aligned (the 8x8 grid) as I bet re-compressions were done before cropping. This achieved Public LB 0.976 and Private LB 0.972. </p></li>\n<li><p>Predict the test set ('unalt' images only) using the finetuned model. Use the predicted probabilities as pseudo-labels for test data and merge the test data with the training set.  Continue tuning with the merged set. After the pseudo-labeling, the performance improved to Public LB 0.983 and Private LB 0.976.</p></li>\n<li><p>Group the 'unalt' images in test set by predicted labels and estimate the sensor noise patterns for each camera in test set (totally ten reference patterns). Then match each of the 'unalt' images with the ten reference patterns, and correct the predictions when the correlation between an image and a reference pattern is larger than a certain threshold. I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns. The last step gave the largest boost: Public LB 0.986 and Private LB 0.987.</p></li>\n</ol>\n\n<p>Thanks to Kaggle and IEEE SPS for hosting this interesting competition.\nThanks to everyone who generously shared their data and ideas.</p>",
  "messages": [
    {
      "id": "279967",
      "postDate": "02/09/2018 02:14:57",
      "content": "<p>Here is a brief summary of my steps.</p>\n\n<p>Edit: I only used the central 80% crop of the train data because the boundaries are often statistically very different from the test data. For example, if the original image size is 1000x1000, only the central 800x800 crop is used. This center-cropping applies to train data only, and it gave around 1% higher accuracy than training on the original size.</p>\n\n<ol>\n<li><p>Finetune a pretrained inception_v3 with random 480x480 crops. The provided training set and Gleb's data were used. My data augmentation include the eight possible manipulations but no transpose, rotation or flipping as I believe they should not help in theory. JPEG compression is always aligned (the 8x8 grid) as I bet re-compressions were done before cropping. This achieved Public LB 0.976 and Private LB 0.972. </p></li>\n<li><p>Predict the test set ('unalt' images only) using the finetuned model. Use the predicted probabilities as pseudo-labels for test data and merge the test data with the training set.  Continue tuning with the merged set. After the pseudo-labeling, the performance improved to Public LB 0.983 and Private LB 0.976.</p></li>\n<li><p>Group the 'unalt' images in test set by predicted labels and estimate the sensor noise patterns for each camera in test set (totally ten reference patterns). Then match each of the 'unalt' images with the ten reference patterns, and correct the predictions when the correlation between an image and a reference pattern is larger than a certain threshold. I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns. The last step gave the largest boost: Public LB 0.986 and Private LB 0.987.</p></li>\n</ol>\n\n<p>Thanks to Kaggle and IEEE SPS for hosting this interesting competition.\nThanks to everyone who generously shared their data and ideas.</p>",
      "rawMarkdown": "Here is a brief summary of my steps.\n\nEdit: I only used the central 80% crop of the train data because the boundaries are often statistically very different from the test data. For example, if the original image size is 1000x1000, only the central 800x800 crop is used. This center-cropping applies to train data only, and it gave around 1% higher accuracy than training on the original size.\n\n 1. Finetune a pretrained inception_v3 with random 480x480 crops. The provided training set and Gleb's data were used. My data augmentation include the eight possible manipulations but no transpose, rotation or flipping as I believe they should not help in theory. JPEG compression is always aligned (the 8x8 grid) as I bet re-compressions were done before cropping. This achieved Public LB 0.976 and Private LB 0.972. \n\n 2. Predict the test set ('unalt' images only) using the finetuned model. Use the predicted probabilities as pseudo-labels for test data and merge the test data with the training set.  Continue tuning with the merged set. After the pseudo-labeling, the performance improved to Public LB 0.983 and Private LB 0.976.\n\n 3. Group the 'unalt' images in test set by predicted labels and estimate the sensor noise patterns for each camera in test set (totally ten reference patterns). Then match each of the 'unalt' images with the ten reference patterns, and correct the predictions when the correlation between an image and a reference pattern is larger than a certain threshold. I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns. The last step gave the largest boost: Public LB 0.986 and Private LB 0.987.\n\nThanks to Kaggle and IEEE SPS for hosting this interesting competition.\nThanks to everyone who generously shared their data and ideas.",
      "votes": null
    },
    {
      "id": "279976",
      "postDate": "02/09/2018 02:38:15",
      "content": "<p>Hi Guanshuo, </p>\n\n<p>Thanks for sharing. Your methodology is excellent. I am happy to see those top solutions which do not rely too much on \"infinite\" datasets.</p>",
      "rawMarkdown": "Hi Guanshuo, \n\nThanks for sharing. Your methodology is excellent. I am happy to see those top solutions which do not rely too much on \"infinite\" datasets.",
      "votes": null
    },
    {
      "id": "279984",
      "postDate": "02/09/2018 02:57:56",
      "content": "<p>Really cool approach.</p>\n\n<blockquote>\n  <p>I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns.</p>\n</blockquote>\n\n<p>What do you mean by \"corrected the 'manip' part\"? Do you mean you psuedo-labeled the manip images? Or do you mean some manip images were wrongly assigned the 'manip' tag? Or do you mean you inferred which augmentations the organizers applied to each manip image by comparing each manip image's noise pattern to an iconic post-JPEG-compression noise pattern, post-resizing noise pattern, etc?</p>",
      "rawMarkdown": "Really cool approach.\n\n&gt; I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns.\n\nWhat do you mean by \"corrected the 'manip' part\"? Do you mean you psuedo-labeled the manip images? Or do you mean some manip images were wrongly assigned the 'manip' tag? Or do you mean you inferred which augmentations the organizers applied to each manip image by comparing each manip image's noise pattern to an iconic post-JPEG-compression noise pattern, post-resizing noise pattern, etc?",
      "votes": null
    },
    {
      "id": "279994",
      "postDate": "02/09/2018 03:37:59",
      "content": "<p>I gathered all the test data with disagreed labels by the two approaches (DL based and sensor noise based). The corrections are done by choosing the labels predicted by the sensor noise based method when the correlation values (between noise estimated by a test image and a reference pattern) is larger than a threshold, otherwise I keep using the DL produced label. To correct the 'manip' part, I first processed the 'unalt' test set by the eight manipulations. For each manipulation and for each camera, one reference pattern were estimated. So we have 10 classes x 8 manips reference patterns. Then, match each 'manip' image with the 80 ref patterns and choose the camera label with the largest correclation.</p>",
      "rawMarkdown": "I gathered all the test data with disagreed labels by the two approaches (DL based and sensor noise based). The corrections are done by choosing the labels predicted by the sensor noise based method when the correlation values (between noise estimated by a test image and a reference pattern) is larger than a threshold, otherwise I keep using the DL produced label. To correct the 'manip' part, I first processed the 'unalt' test set by the eight manipulations. For each manipulation and for each camera, one reference pattern were estimated. So we have 10 classes x 8 manips reference patterns. Then, match each 'manip' image with the 80 ref patterns and choose the camera label with the largest correclation.",
      "votes": null
    },
    {
      "id": "279999",
      "postDate": "02/09/2018 04:24:38",
      "content": "<p>Hi Guanshuo, </p>\n\n<p>I did not understand what do you mean by 'use of the predicted probabilities as pseudo-labels for test data and merge the test data with the training set. ' I mean, the training set has 10 possible labels and you are merging them with the testing which labels are found after test using 10 different probabilities (or maybe the same) as labels? so, this way, after trained again, will your inception_v3 predict 20 classes? what am I not understanding here?  </p>\n\n<p>Another thing, which approach did you use to estimate the sensor noise? did you use the mean sensor noise extracted from several images of each camera in the training set?</p>\n\n<p>Congratulations on your brilliant solution!</p>\n\n<p>BTW, if you replaced your Inception_v3 with Xception CNN you could probably win this challenge!</p>\n\n<p>Best wishes!</p>",
      "rawMarkdown": "Hi Guanshuo, \n\nI did not understand what do you mean by 'use of the predicted probabilities as pseudo-labels for test data and merge the test data with the training set. ' I mean, the training set has 10 possible labels and you are merging them with the testing which labels are found after test using 10 different probabilities (or maybe the same) as labels? so, this way, after trained again, will your inception_v3 predict 20 classes? what am I not understanding here?  \n\nAnother thing, which approach did you use to estimate the sensor noise? did you use the mean sensor noise extracted from several images of each camera in the training set?\n\nCongratulations on your brilliant solution!\n\nBTW, if you replaced your Inception_v3 with Xception CNN you could probably win this challenge!\n\nBest wishes!",
      "votes": null
    },
    {
      "id": "280035",
      "postDate": "02/09/2018 06:22:09",
      "content": "<p>Hi Guanshuo,</p>\n\n<p>Thanks for sharing that. I used more or less the same approach but could get as far as you. Especially I did not consider using the noise pattern for the altered images.</p>",
      "rawMarkdown": "Hi Guanshuo,\n\nThanks for sharing that. I used more or less the same approach but could get as far as you. Especially I did not consider using the noise pattern for the altered images.",
      "votes": null
    },
    {
      "id": "280037",
      "postDate": "02/09/2018 06:25:42",
      "content": "<p>What I understood from his description is that you train a network on the training set. You estimate the labels of the test set (eventually select those who have the largest response). You then assign a label to these test images and consider them as additional ground truth. </p>\n\n<p>The PRNU is estimated on the labeled test data, because the PRNU is specific to the particular camera used to take the images. </p>",
      "rawMarkdown": "What I understood from his description is that you train a network on the training set. You estimate the labels of the test set (eventually select those who have the largest response). You then assign a label to these test images and consider them as additional ground truth. \n\nThe PRNU is estimated on the labeled test data, because the PRNU is specific to the particular camera used to take the images.",
      "votes": null
    },
    {
      "id": "280041",
      "postDate": "02/09/2018 06:33:16",
      "content": "<p>Thanks for your answer!</p>\n\n<p>What I infer from the PRNU idea is: he has several images from 10 different cameras which label we know, then he estimates the reference pattern by extracting for several training images of the same camera their  PRNU patterns and averaging them to get 10 different PRNU reference patterns. For testing, he gets the PRNU pattern from the testing image and calculates the correlation with the 10 mean reference patterns he got in the training stage. The highest correlation identifies the source camera.</p>\n\n<p>Is this correct? </p>",
      "rawMarkdown": "Thanks for your answer!\n\nWhat I infer from the PRNU idea is: he has several images from 10 different cameras which label we know, then he estimates the reference pattern by extracting for several training images of the same camera their  PRNU patterns and averaging them to get 10 different PRNU reference patterns. For testing, he gets the PRNU pattern from the testing image and calculates the correlation with the 10 mean reference patterns he got in the training stage. The highest correlation identifies the source camera.\n\nIs this correct?",
      "votes": null
    },
    {
      "id": "280052",
      "postDate": "02/09/2018 07:27:49",
      "content": "<p>No, actually you estimate the PRNU on the test data labels. For instance you take the 100 most confident test images for each class. You estimate the PRNU for this test class and then you compute the correlation of the noise patttern with all test images again and reassign the images labels. </p>\n\n<p>As I said the PRNU is specific to a camera, so the prnu from the training set is not applicable to the test set. </p>",
      "rawMarkdown": "No, actually you estimate the PRNU on the test data labels. For instance you take the 100 most confident test images for each class. You estimate the PRNU for this test class and then you compute the correlation of the noise patttern with all test images again and reassign the images labels. \n\nAs I said the PRNU is specific to a camera, so the prnu from the training set is not applicable to the test set.",
      "votes": null
    },
    {
      "id": "280054",
      "postDate": "02/09/2018 07:31:01",
      "content": "<p>wow! great idea! However, it only works because here we know that only one another individual camera of the same brand and model of the used training camera was used in the testing. Correct? (LOL). I must confess that this idea was very far away from my mind.</p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "wow! great idea! However, it only works because here we know that only one another individual camera of the same brand and model of the used training camera was used in the testing. Correct? (LOL). I must confess that this idea was very far away from my mind.\n\nThank you!",
      "votes": null
    },
    {
      "id": "280161",
      "postDate": "02/09/2018 13:45:32",
      "content": "<p>Hi Guanshuo, \nI like your ideas. Your solution sounds more meaningful and more scientific than others.</p>",
      "rawMarkdown": "Hi Guanshuo, \nI like your ideas. Your solution sounds more meaningful and more scientific than others.",
      "votes": null
    },
    {
      "id": "280174",
      "postDate": "02/09/2018 14:11:26",
      "content": "<p>Congratulations on the solution. Really smart and novel 3) idea.</p>",
      "rawMarkdown": "Congratulations on the solution. Really smart and novel 3) idea.",
      "votes": null
    },
    {
      "id": "280217",
      "postDate": "02/09/2018 15:13:47",
      "content": "<p>Hi Anselmo,</p>\n\n<p>In the second step, I was actually doing training using the test data. Since we need the labels for optimization (training), what we can do is predicting the labels (for the test data) use the model obtained from the first step. In my case, I used the soft probabilities as labels. </p>\n\n<p>jeandebleau is accurate about the sensor noise part. Again I used the predicted labels for grouping the test data into ten clusters and estimate reference patterns. There is about 1% wrong labels in the prediction in my case for the 'unalt' test data, and fortunately this small 'labeling noise' does not have huge negative impact. Note a minor increase of the 'labeling noise' could totally collapse this method. One key point I would like to add is when matching the sensor noise of each test data with the ten reference patterns, there is an extra step to exclude the contribution of the given test data from the reference pattern, otherwise the method won't work well.</p>\n\n<p>You are right that the sensor noise based method might only work in this specific scenario. I don't believe it will bring improvement if the test data are obtained from a larger number of devices for each class,  and if we dont have enough data for each device, and also if the DL method is less accurate.</p>",
      "rawMarkdown": "Hi Anselmo,\n\n\nIn the second step, I was actually doing training using the test data. Since we need the labels for optimization (training), what we can do is predicting the labels (for the test data) use the model obtained from the first step. In my case, I used the soft probabilities as labels. \n\n\njeandebleau is accurate about the sensor noise part. Again I used the predicted labels for grouping the test data into ten clusters and estimate reference patterns. There is about 1% wrong labels in the prediction in my case for the 'unalt' test data, and fortunately this small 'labeling noise' does not have huge negative impact. Note a minor increase of the 'labeling noise' could totally collapse this method. One key point I would like to add is when matching the sensor noise of each test data with the ten reference patterns, there is an extra step to exclude the contribution of the given test data from the reference pattern, otherwise the method won't work well.\n\nYou are right that the sensor noise based method might only work in this specific scenario. I don't believe it will bring improvement if the test data are obtained from a larger number of devices for each class,  and if we dont have enough data for each device, and also if the DL method is less accurate.",
      "votes": null
    },
    {
      "id": "280246",
      "postDate": "02/09/2018 16:19:58",
      "content": "<p>Hi Guanshuo,</p>\n\n<p>Let me ask you further details about your approach, ok? I am learning a lot from this challenge so \nI must not waste this opportunity (LOL).</p>\n\n<p>This term pseudo-label is really novel to me. I've never seen this being used for camera attribution using DL, actually it sounds a little weird to me to use the testing data to retrain a network (LOL), so\ncan you answer these questions about your approach?</p>\n\n<p>1- basically, the pseudo-label approach means freezing your pre-trained weights and using the testing data and their predictions to retrain only the fully connected layers, correct?</p>\n\n<p>2- ok, suppose my CNN has two FC layers, one with 256 neurons and a last with 10 neurons, which one I retrain? the one with 256 neurons, right? what if I had three FC, with 256, 128 and 10, which ones do I retrain? sorry for the ignorance :-(</p>\n\n<p>3- You said you use the softmax probabilities as labels to retrain the network, is that correct? \nare they the same as the predicted labels from the network in the first test stage? maybe it is just a confusion I am having about terms :-)</p>\n\n<p>4- Did you use all testing data to do pseudo labeling or did you choose 100 of each class with highest predicted probabilities, for example?</p>\n\n<p>5- You said you excluded the tested data from the reference PRNU considered. So, can you correct if necessary my understanding of this approach below? </p>\n\n<p>5.1 you have all the predicted testing data and will create the reference PRNU trusting in their predicted labels. By clustering the testing data into 10 groups, for each group, you simply remove from the images their filtered version and average the results to have a reference for each camera. </p>\n\n<p>5.2 When you test, you calculate the correlation between the testing residual and the 10 references calculated above,  choosing the camera as the one with highest correlation. However, before that, you remove from the reference PRNU you created by trusting on CNN predictions the residual of that testing sample.</p>\n\n<p>You had a brilliant idea anyway. Congrats and thanks for sharing and helping!        </p>",
      "rawMarkdown": "Hi Guanshuo,\n\nLet me ask you further details about your approach, ok? I am learning a lot from this challenge so \nI must not waste this opportunity (LOL).\n\nThis term pseudo-label is really novel to me. I've never seen this being used for camera attribution using DL, actually it sounds a little weird to me to use the testing data to retrain a network (LOL), so\ncan you answer these questions about your approach?\n\n1- basically, the pseudo-label approach means freezing your pre-trained weights and using the testing data and their predictions to retrain only the fully connected layers, correct?\n\n2- ok, suppose my CNN has two FC layers, one with 256 neurons and a last with 10 neurons, which one I retrain? the one with 256 neurons, right? what if I had three FC, with 256, 128 and 10, which ones do I retrain? sorry for the ignorance :-(\n\n3- You said you use the softmax probabilities as labels to retrain the network, is that correct? \nare they the same as the predicted labels from the network in the first test stage? maybe it is just a confusion I am having about terms :-)\n\n4- Did you use all testing data to do pseudo labeling or did you choose 100 of each class with highest predicted probabilities, for example?\n\n5- You said you excluded the tested data from the reference PRNU considered. So, can you correct if necessary my understanding of this approach below? \n\n5.1 you have all the predicted testing data and will create the reference PRNU trusting in their predicted labels. By clustering the testing data into 10 groups, for each group, you simply remove from the images their filtered version and average the results to have a reference for each camera. \n\n5.2 When you test, you calculate the correlation between the testing residual and the 10 references calculated above,  choosing the camera as the one with highest correlation. However, before that, you remove from the reference PRNU you created by trusting on CNN predictions the residual of that testing sample.\n\nYou had a brilliant idea anyway. Congrats and thanks for sharing and helping!",
      "votes": null
    },
    {
      "id": "280264",
      "postDate": "02/09/2018 17:05:17",
      "content": "<p>1 and 2: In this competition, I always finetune all the layers of the CNN in step 1 and step 2, not just the FC layers.</p>\n\n<p>3: I used the softmax probabilities as the labels to NOT retrain but continue finetune the saved weights from step 1. And yes, the probabilities were obtained by the saved net from step 1.</p>\n\n<p>4: I used ALL the 'unalt' test set and I did not use the 'manip' set. The main intuition is to use the 'unalt' part to predict the 'manip' part of the test data because they share same noise patterns and I assume they have similar image content. What contradicts to my intuition is that the 'unalt' performance also improved a little bit and I don' understand why.</p>\n\n<p>5.1: That's pretty close. I actually used the MLE estimator proposed in \"Determining Image Origin and Integrity Using Sensor Noise\". But direct averaging should be okay too.</p>\n\n<p>5.2: Correct.</p>",
      "rawMarkdown": "1 and 2: In this competition, I always finetune all the layers of the CNN in step 1 and step 2, not just the FC layers.\n\n3: I used the softmax probabilities as the labels to NOT retrain but continue finetune the saved weights from step 1. And yes, the probabilities were obtained by the saved net from step 1.\n\n4: I used ALL the 'unalt' test set and I did not use the 'manip' set. The main intuition is to use the 'unalt' part to predict the 'manip' part of the test data because they share same noise patterns and I assume they have similar image content. What contradicts to my intuition is that the 'unalt' performance also improved a little bit and I don' understand why.\n\n5.1: That's pretty close. I actually used the MLE estimator proposed in \"Determining Image Origin and Integrity Using Sensor Noise\". But direct averaging should be okay too.\n\n5.2: Correct.",
      "votes": null
    },
    {
      "id": "280277",
      "postDate": "02/09/2018 17:16:31",
      "content": "<p>Thanks again Guanshuo,</p>\n\n<p>however I am still not understanding number 3: how you use the probabilities as labels to do finetune. I mean, do you use the probabilities to update the weights using probabilities as the ground truth of the testing samples? what I understand from CNN inputs is giving x as a matrix (sample) and y an integer (class) to do the training/fine tuning, You are giving x as sample and y as a double (probability). Is that correct? what is the reasoning behind this?</p>",
      "rawMarkdown": "Thanks again Guanshuo,\n\nhowever I am still not understanding number 3: how you use the probabilities as labels to do finetune. I mean, do you use the probabilities to update the weights using probabilities as the ground truth of the testing samples? what I understand from CNN inputs is giving x as a matrix (sample) and y an integer (class) to do the training/fine tuning, You are giving x as sample and y as a double (probability). Is that correct? what is the reasoning behind this?",
      "votes": null
    },
    {
      "id": "280296",
      "postDate": "02/09/2018 17:30:06",
      "content": "<p>In multi-class classification, labels should be (internally) one-hot encoded. Assume we have three classes, the input (data, label) pairs should be (data1, [0 1 0]), (data2, [1 0 0]), (data3, [1 0 0]) ... It's true that we can use single integers to represent the labels but they will be ultimately treated as one-hot encoded internally. So the pseudo-labeled test data, assume three classes, should look like (data1, [0.1 0.7 0.2]), (data2, [0.4 0.3 0.3]) ... You can also use 'hard' predictions instead of the probabilities. Honestly I don't know the best way to do the pseudo-labeling either.</p>",
      "rawMarkdown": "In multi-class classification, labels should be (internally) one-hot encoded. Assume we have three classes, the input (data, label) pairs should be (data1, [0 1 0]), (data2, [1 0 0]), (data3, [1 0 0]) ... It's true that we can use single integers to represent the labels but they will be ultimately treated as one-hot encoded internally. So the pseudo-labeled test data, assume three classes, should look like (data1, [0.1 0.7 0.2]), (data2, [0.4 0.3 0.3]) ... You can also use 'hard' predictions instead of the probabilities. Honestly I don't know the best way to do the pseudo-labeling either.",
      "votes": null
    },
    {
      "id": "280543",
      "postDate": "02/10/2018 07:52:55",
      "content": "<p>I just want to note that there was no mention of ensembling. If so, this is an even more impressive result. I wonder how Guanshuo Xu would have placed if they used ensembling.</p>",
      "rawMarkdown": "I just want to note that there was no mention of ensembling. If so, this is an even more impressive result. I wonder how Guanshuo Xu would have placed if they used ensembling.",
      "votes": null
    },
    {
      "id": "280550",
      "postDate": "02/10/2018 08:11:51",
      "content": "<p>A last question, what noise estimation method did you use ?</p>",
      "rawMarkdown": "A last question, what noise estimation method did you use ?",
      "votes": null
    },
    {
      "id": "280621",
      "postDate": "02/10/2018 13:50:18",
      "content": "<p>During competition, I did submit an ensemble result in which I averaged predictions using four inception models (inception_v3, inception_resnet_v2, inception_v4 xception) trained with various crop sizes. The LB result was public0.981/private0.981 (after step 2 of my solution). I don't know how much the improvement would transfer to after step 3. I feared that I would drop out of the 'gold' zone so I did not choose to continue with the ensemble result. The leading teams were just giving me too much pressure. </p>",
      "rawMarkdown": "During competition, I did submit an ensemble result in which I averaged predictions using four inception models (inception_v3, inception_resnet_v2, inception_v4 xception) trained with various crop sizes. The LB result was public0.981/private0.981 (after step 2 of my solution). I don't know how much the improvement would transfer to after step 3. I feared that I would drop out of the 'gold' zone so I did not choose to continue with the ensemble result. The leading teams were just giving me too much pressure.",
      "votes": null
    },
    {
      "id": "280668",
      "postDate": "02/10/2018 16:09:03",
      "content": "<p>I just ran step3 with the ensemble model - Private 0.988452 and Public 0.983541. Still cannot beat the 1st team. This confirms the importance of more data.</p>",
      "rawMarkdown": "I just ran step3 with the ensemble model - Private 0.988452 and Public 0.983541. Still cannot beat the 1st team. This confirms the importance of more data.",
      "votes": null
    },
    {
      "id": "280670",
      "postDate": "02/10/2018 16:10:36",
      "content": "<p>\"Determining Image Origin and Integrity Using Sensor Noise\"</p>",
      "rawMarkdown": "\"Determining Image Origin and Integrity Using Sensor Noise\"",
      "votes": null
    },
    {
      "id": "280682",
      "postDate": "02/10/2018 16:39:54",
      "content": "<p>So denoised image is estimated with a wavelet denoising filter. It might be worse trying something more accurate like bm3d.</p>",
      "rawMarkdown": "So denoised image is estimated with a wavelet denoising filter. It might be worse trying something more accurate like bm3d.",
      "votes": null
    },
    {
      "id": "280806",
      "postDate": "02/11/2018 04:49:54",
      "content": "<p>Thanks again Guanshuo.</p>\n\n<p>This pseudo-labeling you and the other people used is a kind of weird thing to me (LOL). I mean, if you use the testing data to retrain the network and after the network is trained you ensure that the same testing data will be better classified,  why not repeating the same process, say, 100 times? this way for each accuracy improvement using pseudo labeling, you have less noisy labels and this way you can get better accuracy at each pseudo-labeling iteration (LOL).</p>\n\n<p>As I said, this is a kind of strange solution to me, thanks for your patience!</p>",
      "rawMarkdown": "Thanks again Guanshuo.\n\nThis pseudo-labeling you and the other people used is a kind of weird thing to me (LOL). I mean, if you use the testing data to retrain the network and after the network is trained you ensure that the same testing data will be better classified,  why not repeating the same process, say, 100 times? this way for each accuracy improvement using pseudo labeling, you have less noisy labels and this way you can get better accuracy at each pseudo-labeling iteration (LOL).\n\nAs I said, this is a kind of strange solution to me, thanks for your patience!",
      "votes": null
    },
    {
      "id": "280818",
      "postDate": "02/11/2018 05:41:42",
      "content": "<p>Guanshuo, the ensemble you used is a little similar with what my team used (we didn't use inception v4), which crop sizes did you use? we used 299x299 only on areas with high texture.</p>\n\n<p>Another thing, did you consider the 8 image operations for the data augmentation of all your CNNs? </p>\n\n<p>Finally, which percentage of the data did you use for training and validation?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Guanshuo, the ensemble you used is a little similar with what my team used (we didn't use inception v4), which crop sizes did you use? we used 299x299 only on areas with high texture.\n\nAnother thing, did you consider the 8 image operations for the data augmentation of all your CNNs? \n\nFinally, which percentage of the data did you use for training and validation?\n\nThanks!",
      "votes": null
    },
    {
      "id": "280949",
      "postDate": "02/11/2018 14:15:04",
      "content": "<p>I forgot to mention that I only used the 80% central portion of the train data, which should be in agreement with your idea because the centers are generally more textured. I have updated the original post. Thank you for reminding me.</p>\n\n<p>The crop sizes: inception_v3 (480x480), inception_resnet_v2 (299x299), inception_v4 (352x352), xception (320x320).</p>\n\n<p>Yes, similar to others works, I augmented the random training crops by the 8 manips with a probability. In other words, one model for all the unalt and manip test data.</p>\n\n<p>In step 1, I used the original train data provided by the organizer plus Gleb's data (good jpgs) for training, and I used Gleb's validation data for validation. In step2, I used the original training data provided by the organizer plus Gleb's data (good jpgs) plus the 'unalt' part of the test data for training, and I still used Gleb's validation data for validation</p>",
      "rawMarkdown": "I forgot to mention that I only used the 80% central portion of the train data, which should be in agreement with your idea because the centers are generally more textured. I have updated the original post. Thank you for reminding me.\n\nThe crop sizes: inception_v3 (480x480), inception_resnet_v2 (299x299), inception_v4 (352x352), xception (320x320).\n\nYes, similar to others works, I augmented the random training crops by the 8 manips with a probability. In other words, one model for all the unalt and manip test data.\n\nIn step 1, I used the original train data provided by the organizer plus Gleb's data (good jpgs) for training, and I used Gleb's validation data for validation. In step2, I used the original training data provided by the organizer plus Gleb's data (good jpgs) plus the 'unalt' part of the test data for training, and I still used Gleb's validation data for validation",
      "votes": null
    },
    {
      "id": "280951",
      "postDate": "02/11/2018 14:18:06",
      "content": "<p>Right, I did iterate this process three times, but did not observe any concrete improvement.</p>",
      "rawMarkdown": "Right, I did iterate this process three times, but did not observe any concrete improvement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 279976,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "02/09/2018 02:38:15",
      "content": "<p>Hi Guanshuo, </p>\n\n<p>Thanks for sharing. Your methodology is excellent. I am happy to see those top solutions which do not rely too much on \"infinite\" datasets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 279984,
      "author_name": "kleinsmith",
      "author_url": "",
      "post_date": "02/09/2018 02:57:56",
      "content": "<p>Really cool approach.</p>\n\n<blockquote>\n  <p>I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns.</p>\n</blockquote>\n\n<p>What do you mean by \"corrected the 'manip' part\"? Do you mean you psuedo-labeled the manip images? Or do you mean some manip images were wrongly assigned the 'manip' tag? Or do you mean you inferred which augmentations the organizers applied to each manip image by comparing each manip image's noise pattern to an iconic post-JPEG-compression noise pattern, post-resizing noise pattern, etc?</p>",
      "votes": null,
      "replies": [
        {
          "id": 279994,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/09/2018 03:37:59",
          "content": "<p>I gathered all the test data with disagreed labels by the two approaches (DL based and sensor noise based). The corrections are done by choosing the labels predicted by the sensor noise based method when the correlation values (between noise estimated by a test image and a reference pattern) is larger than a threshold, otherwise I keep using the DL produced label. To correct the 'manip' part, I first processed the 'unalt' test set by the eight manipulations. For each manipulation and for each camera, one reference pattern were estimated. So we have 10 classes x 8 manips reference patterns. Then, match each 'manip' image with the 80 ref patterns and choose the camera label with the largest correclation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 279999,
      "author_name": "anselmoferreira35",
      "author_url": "",
      "post_date": "02/09/2018 04:24:38",
      "content": "<p>Hi Guanshuo, </p>\n\n<p>I did not understand what do you mean by 'use of the predicted probabilities as pseudo-labels for test data and merge the test data with the training set. ' I mean, the training set has 10 possible labels and you are merging them with the testing which labels are found after test using 10 different probabilities (or maybe the same) as labels? so, this way, after trained again, will your inception_v3 predict 20 classes? what am I not understanding here?  </p>\n\n<p>Another thing, which approach did you use to estimate the sensor noise? did you use the mean sensor noise extracted from several images of each camera in the training set?</p>\n\n<p>Congratulations on your brilliant solution!</p>\n\n<p>BTW, if you replaced your Inception_v3 with Xception CNN you could probably win this challenge!</p>\n\n<p>Best wishes!</p>",
      "votes": null,
      "replies": [
        {
          "id": 280037,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "02/09/2018 06:25:42",
          "content": "<p>What I understood from his description is that you train a network on the training set. You estimate the labels of the test set (eventually select those who have the largest response). You then assign a label to these test images and consider them as additional ground truth. </p>\n\n<p>The PRNU is estimated on the labeled test data, because the PRNU is specific to the particular camera used to take the images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280041,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/09/2018 06:33:16",
          "content": "<p>Thanks for your answer!</p>\n\n<p>What I infer from the PRNU idea is: he has several images from 10 different cameras which label we know, then he estimates the reference pattern by extracting for several training images of the same camera their  PRNU patterns and averaging them to get 10 different PRNU reference patterns. For testing, he gets the PRNU pattern from the testing image and calculates the correlation with the 10 mean reference patterns he got in the training stage. The highest correlation identifies the source camera.</p>\n\n<p>Is this correct? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280052,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "02/09/2018 07:27:49",
          "content": "<p>No, actually you estimate the PRNU on the test data labels. For instance you take the 100 most confident test images for each class. You estimate the PRNU for this test class and then you compute the correlation of the noise patttern with all test images again and reassign the images labels. </p>\n\n<p>As I said the PRNU is specific to a camera, so the prnu from the training set is not applicable to the test set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280054,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/09/2018 07:31:01",
          "content": "<p>wow! great idea! However, it only works because here we know that only one another individual camera of the same brand and model of the used training camera was used in the testing. Correct? (LOL). I must confess that this idea was very far away from my mind.</p>\n\n<p>Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280217,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/09/2018 15:13:47",
          "content": "<p>Hi Anselmo,</p>\n\n<p>In the second step, I was actually doing training using the test data. Since we need the labels for optimization (training), what we can do is predicting the labels (for the test data) use the model obtained from the first step. In my case, I used the soft probabilities as labels. </p>\n\n<p>jeandebleau is accurate about the sensor noise part. Again I used the predicted labels for grouping the test data into ten clusters and estimate reference patterns. There is about 1% wrong labels in the prediction in my case for the 'unalt' test data, and fortunately this small 'labeling noise' does not have huge negative impact. Note a minor increase of the 'labeling noise' could totally collapse this method. One key point I would like to add is when matching the sensor noise of each test data with the ten reference patterns, there is an extra step to exclude the contribution of the given test data from the reference pattern, otherwise the method won't work well.</p>\n\n<p>You are right that the sensor noise based method might only work in this specific scenario. I don't believe it will bring improvement if the test data are obtained from a larger number of devices for each class,  and if we dont have enough data for each device, and also if the DL method is less accurate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280246,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/09/2018 16:19:58",
          "content": "<p>Hi Guanshuo,</p>\n\n<p>Let me ask you further details about your approach, ok? I am learning a lot from this challenge so \nI must not waste this opportunity (LOL).</p>\n\n<p>This term pseudo-label is really novel to me. I've never seen this being used for camera attribution using DL, actually it sounds a little weird to me to use the testing data to retrain a network (LOL), so\ncan you answer these questions about your approach?</p>\n\n<p>1- basically, the pseudo-label approach means freezing your pre-trained weights and using the testing data and their predictions to retrain only the fully connected layers, correct?</p>\n\n<p>2- ok, suppose my CNN has two FC layers, one with 256 neurons and a last with 10 neurons, which one I retrain? the one with 256 neurons, right? what if I had three FC, with 256, 128 and 10, which ones do I retrain? sorry for the ignorance :-(</p>\n\n<p>3- You said you use the softmax probabilities as labels to retrain the network, is that correct? \nare they the same as the predicted labels from the network in the first test stage? maybe it is just a confusion I am having about terms :-)</p>\n\n<p>4- Did you use all testing data to do pseudo labeling or did you choose 100 of each class with highest predicted probabilities, for example?</p>\n\n<p>5- You said you excluded the tested data from the reference PRNU considered. So, can you correct if necessary my understanding of this approach below? </p>\n\n<p>5.1 you have all the predicted testing data and will create the reference PRNU trusting in their predicted labels. By clustering the testing data into 10 groups, for each group, you simply remove from the images their filtered version and average the results to have a reference for each camera. </p>\n\n<p>5.2 When you test, you calculate the correlation between the testing residual and the 10 references calculated above,  choosing the camera as the one with highest correlation. However, before that, you remove from the reference PRNU you created by trusting on CNN predictions the residual of that testing sample.</p>\n\n<p>You had a brilliant idea anyway. Congrats and thanks for sharing and helping!        </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280264,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/09/2018 17:05:17",
          "content": "<p>1 and 2: In this competition, I always finetune all the layers of the CNN in step 1 and step 2, not just the FC layers.</p>\n\n<p>3: I used the softmax probabilities as the labels to NOT retrain but continue finetune the saved weights from step 1. And yes, the probabilities were obtained by the saved net from step 1.</p>\n\n<p>4: I used ALL the 'unalt' test set and I did not use the 'manip' set. The main intuition is to use the 'unalt' part to predict the 'manip' part of the test data because they share same noise patterns and I assume they have similar image content. What contradicts to my intuition is that the 'unalt' performance also improved a little bit and I don' understand why.</p>\n\n<p>5.1: That's pretty close. I actually used the MLE estimator proposed in \"Determining Image Origin and Integrity Using Sensor Noise\". But direct averaging should be okay too.</p>\n\n<p>5.2: Correct.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280277,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/09/2018 17:16:31",
          "content": "<p>Thanks again Guanshuo,</p>\n\n<p>however I am still not understanding number 3: how you use the probabilities as labels to do finetune. I mean, do you use the probabilities to update the weights using probabilities as the ground truth of the testing samples? what I understand from CNN inputs is giving x as a matrix (sample) and y an integer (class) to do the training/fine tuning, You are giving x as sample and y as a double (probability). Is that correct? what is the reasoning behind this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280296,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/09/2018 17:30:06",
          "content": "<p>In multi-class classification, labels should be (internally) one-hot encoded. Assume we have three classes, the input (data, label) pairs should be (data1, [0 1 0]), (data2, [1 0 0]), (data3, [1 0 0]) ... It's true that we can use single integers to represent the labels but they will be ultimately treated as one-hot encoded internally. So the pseudo-labeled test data, assume three classes, should look like (data1, [0.1 0.7 0.2]), (data2, [0.4 0.3 0.3]) ... You can also use 'hard' predictions instead of the probabilities. Honestly I don't know the best way to do the pseudo-labeling either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280806,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/11/2018 04:49:54",
          "content": "<p>Thanks again Guanshuo.</p>\n\n<p>This pseudo-labeling you and the other people used is a kind of weird thing to me (LOL). I mean, if you use the testing data to retrain the network and after the network is trained you ensure that the same testing data will be better classified,  why not repeating the same process, say, 100 times? this way for each accuracy improvement using pseudo labeling, you have less noisy labels and this way you can get better accuracy at each pseudo-labeling iteration (LOL).</p>\n\n<p>As I said, this is a kind of strange solution to me, thanks for your patience!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280951,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/11/2018 14:18:06",
          "content": "<p>Right, I did iterate this process three times, but did not observe any concrete improvement.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280035,
      "author_name": "jeandebleau",
      "author_url": "",
      "post_date": "02/09/2018 06:22:09",
      "content": "<p>Hi Guanshuo,</p>\n\n<p>Thanks for sharing that. I used more or less the same approach but could get as far as you. Especially I did not consider using the noise pattern for the altered images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280161,
      "author_name": "jfkg16",
      "author_url": "",
      "post_date": "02/09/2018 13:45:32",
      "content": "<p>Hi Guanshuo, \nI like your ideas. Your solution sounds more meaningful and more scientific than others.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280174,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/09/2018 14:11:26",
      "content": "<p>Congratulations on the solution. Really smart and novel 3) idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280543,
      "author_name": "kleinsmith",
      "author_url": "",
      "post_date": "02/10/2018 07:52:55",
      "content": "<p>I just want to note that there was no mention of ensembling. If so, this is an even more impressive result. I wonder how Guanshuo Xu would have placed if they used ensembling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 280621,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/10/2018 13:50:18",
          "content": "<p>During competition, I did submit an ensemble result in which I averaged predictions using four inception models (inception_v3, inception_resnet_v2, inception_v4 xception) trained with various crop sizes. The LB result was public0.981/private0.981 (after step 2 of my solution). I don't know how much the improvement would transfer to after step 3. I feared that I would drop out of the 'gold' zone so I did not choose to continue with the ensemble result. The leading teams were just giving me too much pressure. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280668,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/10/2018 16:09:03",
          "content": "<p>I just ran step3 with the ensemble model - Private 0.988452 and Public 0.983541. Still cannot beat the 1st team. This confirms the importance of more data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280818,
          "author_name": "anselmoferreira35",
          "author_url": "",
          "post_date": "02/11/2018 05:41:42",
          "content": "<p>Guanshuo, the ensemble you used is a little similar with what my team used (we didn't use inception v4), which crop sizes did you use? we used 299x299 only on areas with high texture.</p>\n\n<p>Another thing, did you consider the 8 image operations for the data augmentation of all your CNNs? </p>\n\n<p>Finally, which percentage of the data did you use for training and validation?</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280949,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/11/2018 14:15:04",
          "content": "<p>I forgot to mention that I only used the 80% central portion of the train data, which should be in agreement with your idea because the centers are generally more textured. I have updated the original post. Thank you for reminding me.</p>\n\n<p>The crop sizes: inception_v3 (480x480), inception_resnet_v2 (299x299), inception_v4 (352x352), xception (320x320).</p>\n\n<p>Yes, similar to others works, I augmented the random training crops by the 8 manips with a probability. In other words, one model for all the unalt and manip test data.</p>\n\n<p>In step 1, I used the original train data provided by the organizer plus Gleb's data (good jpgs) for training, and I used Gleb's validation data for validation. In step2, I used the original training data provided by the organizer plus Gleb's data (good jpgs) plus the 'unalt' part of the test data for training, and I still used Gleb's validation data for validation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280550,
      "author_name": "jeandebleau",
      "author_url": "",
      "post_date": "02/10/2018 08:11:51",
      "content": "<p>A last question, what noise estimation method did you use ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 280670,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/10/2018 16:10:36",
          "content": "<p>\"Determining Image Origin and Integrity Using Sensor Noise\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280682,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "02/10/2018 16:39:54",
          "content": "<p>So denoised image is estimated with a wavelet denoising filter. It might be worse trying something more accurate like bm3d.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "279967": "Here is a brief summary of my steps.\n\nEdit: I only used the central 80% crop of the train data because the boundaries are often statistically very different from the test data. For example, if the original image size is 1000x1000, only the central 800x800 crop is used. This center-cropping applies to train data only, and it gave around 1% higher accuracy than training on the original size.\n\n 1. Finetune a pretrained inception_v3 with random 480x480 crops. The provided training set and Gleb's data were used. My data augmentation include the eight possible manipulations but no transpose, rotation or flipping as I believe they should not help in theory. JPEG compression is always aligned (the 8x8 grid) as I bet re-compressions were done before cropping. This achieved Public LB 0.976 and Private LB 0.972. \n\n 2. Predict the test set ('unalt' images only) using the finetuned model. Use the predicted probabilities as pseudo-labels for test data and merge the test data with the training set.  Continue tuning with the merged set. After the pseudo-labeling, the performance improved to Public LB 0.983 and Private LB 0.976.\n\n 3. Group the 'unalt' images in test set by predicted labels and estimate the sensor noise patterns for each camera in test set (totally ten reference patterns). Then match each of the 'unalt' images with the ten reference patterns, and correct the predictions when the correlation between an image and a reference pattern is larger than a certain threshold. I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns. The last step gave the largest boost: Public LB 0.986 and Private LB 0.987.\n\nThanks to Kaggle and IEEE SPS for hosting this interesting competition.\nThanks to everyone who generously shared their data and ideas.",
    "279976": "Hi Guanshuo, \n\nThanks for sharing. Your methodology is excellent. I am happy to see those top solutions which do not rely too much on \"infinite\" datasets.",
    "279984": "Really cool approach.\n\n&gt; I also corrected the 'manip' part by matching their sensor noises with the augmented (by the eight manipulations) reference patterns.\n\nWhat do you mean by \"corrected the 'manip' part\"? Do you mean you psuedo-labeled the manip images? Or do you mean some manip images were wrongly assigned the 'manip' tag? Or do you mean you inferred which augmentations the organizers applied to each manip image by comparing each manip image's noise pattern to an iconic post-JPEG-compression noise pattern, post-resizing noise pattern, etc?",
    "279994": "I gathered all the test data with disagreed labels by the two approaches (DL based and sensor noise based). The corrections are done by choosing the labels predicted by the sensor noise based method when the correlation values (between noise estimated by a test image and a reference pattern) is larger than a threshold, otherwise I keep using the DL produced label. To correct the 'manip' part, I first processed the 'unalt' test set by the eight manipulations. For each manipulation and for each camera, one reference pattern were estimated. So we have 10 classes x 8 manips reference patterns. Then, match each 'manip' image with the 80 ref patterns and choose the camera label with the largest correclation.",
    "279999": "Hi Guanshuo, \n\nI did not understand what do you mean by 'use of the predicted probabilities as pseudo-labels for test data and merge the test data with the training set. ' I mean, the training set has 10 possible labels and you are merging them with the testing which labels are found after test using 10 different probabilities (or maybe the same) as labels? so, this way, after trained again, will your inception_v3 predict 20 classes? what am I not understanding here?  \n\nAnother thing, which approach did you use to estimate the sensor noise? did you use the mean sensor noise extracted from several images of each camera in the training set?\n\nCongratulations on your brilliant solution!\n\nBTW, if you replaced your Inception_v3 with Xception CNN you could probably win this challenge!\n\nBest wishes!",
    "280035": "Hi Guanshuo,\n\nThanks for sharing that. I used more or less the same approach but could get as far as you. Especially I did not consider using the noise pattern for the altered images.",
    "280037": "What I understood from his description is that you train a network on the training set. You estimate the labels of the test set (eventually select those who have the largest response). You then assign a label to these test images and consider them as additional ground truth. \n\nThe PRNU is estimated on the labeled test data, because the PRNU is specific to the particular camera used to take the images.",
    "280041": "Thanks for your answer!\n\nWhat I infer from the PRNU idea is: he has several images from 10 different cameras which label we know, then he estimates the reference pattern by extracting for several training images of the same camera their  PRNU patterns and averaging them to get 10 different PRNU reference patterns. For testing, he gets the PRNU pattern from the testing image and calculates the correlation with the 10 mean reference patterns he got in the training stage. The highest correlation identifies the source camera.\n\nIs this correct?",
    "280052": "No, actually you estimate the PRNU on the test data labels. For instance you take the 100 most confident test images for each class. You estimate the PRNU for this test class and then you compute the correlation of the noise patttern with all test images again and reassign the images labels. \n\nAs I said the PRNU is specific to a camera, so the prnu from the training set is not applicable to the test set.",
    "280054": "wow! great idea! However, it only works because here we know that only one another individual camera of the same brand and model of the used training camera was used in the testing. Correct? (LOL). I must confess that this idea was very far away from my mind.\n\nThank you!",
    "280161": "Hi Guanshuo, \nI like your ideas. Your solution sounds more meaningful and more scientific than others.",
    "280174": "Congratulations on the solution. Really smart and novel 3) idea.",
    "280217": "Hi Anselmo,\n\n\nIn the second step, I was actually doing training using the test data. Since we need the labels for optimization (training), what we can do is predicting the labels (for the test data) use the model obtained from the first step. In my case, I used the soft probabilities as labels. \n\n\njeandebleau is accurate about the sensor noise part. Again I used the predicted labels for grouping the test data into ten clusters and estimate reference patterns. There is about 1% wrong labels in the prediction in my case for the 'unalt' test data, and fortunately this small 'labeling noise' does not have huge negative impact. Note a minor increase of the 'labeling noise' could totally collapse this method. One key point I would like to add is when matching the sensor noise of each test data with the ten reference patterns, there is an extra step to exclude the contribution of the given test data from the reference pattern, otherwise the method won't work well.\n\nYou are right that the sensor noise based method might only work in this specific scenario. I don't believe it will bring improvement if the test data are obtained from a larger number of devices for each class,  and if we dont have enough data for each device, and also if the DL method is less accurate.",
    "280246": "Hi Guanshuo,\n\nLet me ask you further details about your approach, ok? I am learning a lot from this challenge so \nI must not waste this opportunity (LOL).\n\nThis term pseudo-label is really novel to me. I've never seen this being used for camera attribution using DL, actually it sounds a little weird to me to use the testing data to retrain a network (LOL), so\ncan you answer these questions about your approach?\n\n1- basically, the pseudo-label approach means freezing your pre-trained weights and using the testing data and their predictions to retrain only the fully connected layers, correct?\n\n2- ok, suppose my CNN has two FC layers, one with 256 neurons and a last with 10 neurons, which one I retrain? the one with 256 neurons, right? what if I had three FC, with 256, 128 and 10, which ones do I retrain? sorry for the ignorance :-(\n\n3- You said you use the softmax probabilities as labels to retrain the network, is that correct? \nare they the same as the predicted labels from the network in the first test stage? maybe it is just a confusion I am having about terms :-)\n\n4- Did you use all testing data to do pseudo labeling or did you choose 100 of each class with highest predicted probabilities, for example?\n\n5- You said you excluded the tested data from the reference PRNU considered. So, can you correct if necessary my understanding of this approach below? \n\n5.1 you have all the predicted testing data and will create the reference PRNU trusting in their predicted labels. By clustering the testing data into 10 groups, for each group, you simply remove from the images their filtered version and average the results to have a reference for each camera. \n\n5.2 When you test, you calculate the correlation between the testing residual and the 10 references calculated above,  choosing the camera as the one with highest correlation. However, before that, you remove from the reference PRNU you created by trusting on CNN predictions the residual of that testing sample.\n\nYou had a brilliant idea anyway. Congrats and thanks for sharing and helping!",
    "280264": "1 and 2: In this competition, I always finetune all the layers of the CNN in step 1 and step 2, not just the FC layers.\n\n3: I used the softmax probabilities as the labels to NOT retrain but continue finetune the saved weights from step 1. And yes, the probabilities were obtained by the saved net from step 1.\n\n4: I used ALL the 'unalt' test set and I did not use the 'manip' set. The main intuition is to use the 'unalt' part to predict the 'manip' part of the test data because they share same noise patterns and I assume they have similar image content. What contradicts to my intuition is that the 'unalt' performance also improved a little bit and I don' understand why.\n\n5.1: That's pretty close. I actually used the MLE estimator proposed in \"Determining Image Origin and Integrity Using Sensor Noise\". But direct averaging should be okay too.\n\n5.2: Correct.",
    "280277": "Thanks again Guanshuo,\n\nhowever I am still not understanding number 3: how you use the probabilities as labels to do finetune. I mean, do you use the probabilities to update the weights using probabilities as the ground truth of the testing samples? what I understand from CNN inputs is giving x as a matrix (sample) and y an integer (class) to do the training/fine tuning, You are giving x as sample and y as a double (probability). Is that correct? what is the reasoning behind this?",
    "280296": "In multi-class classification, labels should be (internally) one-hot encoded. Assume we have three classes, the input (data, label) pairs should be (data1, [0 1 0]), (data2, [1 0 0]), (data3, [1 0 0]) ... It's true that we can use single integers to represent the labels but they will be ultimately treated as one-hot encoded internally. So the pseudo-labeled test data, assume three classes, should look like (data1, [0.1 0.7 0.2]), (data2, [0.4 0.3 0.3]) ... You can also use 'hard' predictions instead of the probabilities. Honestly I don't know the best way to do the pseudo-labeling either.",
    "280543": "I just want to note that there was no mention of ensembling. If so, this is an even more impressive result. I wonder how Guanshuo Xu would have placed if they used ensembling.",
    "280550": "A last question, what noise estimation method did you use ?",
    "280621": "During competition, I did submit an ensemble result in which I averaged predictions using four inception models (inception_v3, inception_resnet_v2, inception_v4 xception) trained with various crop sizes. The LB result was public0.981/private0.981 (after step 2 of my solution). I don't know how much the improvement would transfer to after step 3. I feared that I would drop out of the 'gold' zone so I did not choose to continue with the ensemble result. The leading teams were just giving me too much pressure.",
    "280668": "I just ran step3 with the ensemble model - Private 0.988452 and Public 0.983541. Still cannot beat the 1st team. This confirms the importance of more data.",
    "280670": "\"Determining Image Origin and Integrity Using Sensor Noise\"",
    "280682": "So denoised image is estimated with a wavelet denoising filter. It might be worse trying something more accurate like bm3d.",
    "280806": "Thanks again Guanshuo.\n\nThis pseudo-labeling you and the other people used is a kind of weird thing to me (LOL). I mean, if you use the testing data to retrain the network and after the network is trained you ensure that the same testing data will be better classified,  why not repeating the same process, say, 100 times? this way for each accuracy improvement using pseudo labeling, you have less noisy labels and this way you can get better accuracy at each pseudo-labeling iteration (LOL).\n\nAs I said, this is a kind of strange solution to me, thanks for your patience!",
    "280818": "Guanshuo, the ensemble you used is a little similar with what my team used (we didn't use inception v4), which crop sizes did you use? we used 299x299 only on areas with high texture.\n\nAnother thing, did you consider the 8 image operations for the data augmentation of all your CNNs? \n\nFinally, which percentage of the data did you use for training and validation?\n\nThanks!",
    "280949": "I forgot to mention that I only used the 80% central portion of the train data, which should be in agreement with your idea because the centers are generally more textured. I have updated the original post. Thank you for reminding me.\n\nThe crop sizes: inception_v3 (480x480), inception_resnet_v2 (299x299), inception_v4 (352x352), xception (320x320).\n\nYes, similar to others works, I augmented the random training crops by the 8 manips with a probability. In other words, one model for all the unalt and manip test data.\n\nIn step 1, I used the original train data provided by the organizer plus Gleb's data (good jpgs) for training, and I used Gleb's validation data for validation. In step2, I used the original training data provided by the organizer plus Gleb's data (good jpgs) plus the 'unalt' part of the test data for training, and I still used Gleb's validation data for validation",
    "280951": "Right, I did iterate this process three times, but did not observe any concrete improvement."
  },
  "source": "meta"
}