{
  "id": 49302,
  "title": "gold solution",
  "url": "/competitions/sp-society-camera-model-identification/writeups/master-gold-solution",
  "author_name": "",
  "post_date": "2018-12-19T10:54:36.497Z",
  "votes": 6,
  "comment_count": 14,
  "views": 0,
  "content": "<p>First, congratulations to all the winners and those that did well. Thanks to Kaggle, IEEE, and everyone who shared their ideas. This is my first gold medal.</p>\n\n<p>Convolutional neural networks were able to learn meaningful relationships, which is quite interesting. As a result of this competition, it is evident that deep learning is a powerful tool, generalizable to many tasks. The final model ensemble obtains a private LB score of 0.986 (Public: 0.983), using ImageNet models ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3.</p>\n\n<h2>Data Preparation:</h2>\n\n<ul>\n<li><p>Slice each image into 256 x 256 non-overlapping patches and discard any images with only one unique pixel, since they don't help with camera identification; Gleb's additional dataset was used. Each patch retains the label of the original image. My data augmentation included the eight possible manipulations.  </p>\n\n<ol><li>JPEG compression with quality factor = 70</li>\n<li>JPEG compression with quality factor = 90</li>\n<li>resizing (via bicubic interpolation) by a factor of 0.5</li>\n<li>resizing (via bicubic interpolation) by a factor of 0.8</li>\n<li>resizing (via bicubic interpolation) by a factor of 1.5</li>\n<li>resizing (via bicubic interpolation) by a factor of 2.0</li>\n<li>gamma correction using gamma = 0.8</li>\n<li>gamma correction using gamma = 1.2</li></ol></li>\n<li><p>Trained on unaltered patches and set each manipulation probability to 5%. No rotation or flipping was used, though in retrospect, those could have improved my score slightly.</p></li>\n<li>Validation patches are not from training images to ensure no possibility of data leakage.</li>\n</ul>\n\n<h2>Models:</h2>\n\n<ul>\n<li>ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3, all pretrained on ImageNet.</li>\n<li>All models use Global Average Pooling after the final convolution followed by a softmax layer. I experimented with fully-connected layers with dropout and regularization, but they didn't seem to improve the score.</li>\n</ul>\n\n<h2>Training:</h2>\n\n<ul>\n<li>Adam optimizer, learning rate: 0.0001    </li>\n<li>Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch; in most cases, the 3rd epoch actually lowered accuracy. I tried retraining with a learning rate of 0.00005 for another epoch, and managed to boost validation accuracies of most models to 0.975.  </li>\n</ul>\n\n<h2>Prediction:</h2>\n\n<ul>\n<li>Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. The validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.</li>\n</ul>\n\n<p>Adding some checkpoint models helped raise private LB score from 0.983 to 0.986. The additional models helped with corner cases.  </p>\n\n<p>The results could have been improved by adding more models like DenseNet201, NASNet, etc., and training for more epochs. But I guess at this point its for future competitions.</p>",
  "messages": [
    {
      "id": "279990",
      "postDate": "02/09/2018 03:28:52",
      "content": "<p>First, congratulations to all the winners and those that did well. Thanks to Kaggle, IEEE, and everyone who shared their ideas. This is my first gold medal.</p>\n\n<p>Convolutional neural networks were able to learn meaningful relationships, which is quite interesting. As a result of this competition, it is evident that deep learning is a powerful tool, generalizable to many tasks. The final model ensemble obtains a private LB score of 0.986 (Public: 0.983), using ImageNet models ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3.</p>\n\n<h2>Data Preparation:</h2>\n\n<ul>\n<li><p>Slice each image into 256 x 256 non-overlapping patches and discard any images with only one unique pixel, since they don't help with camera identification; Gleb's additional dataset was used. Each patch retains the label of the original image. My data augmentation included the eight possible manipulations.  </p>\n\n<ol><li>JPEG compression with quality factor = 70</li>\n<li>JPEG compression with quality factor = 90</li>\n<li>resizing (via bicubic interpolation) by a factor of 0.5</li>\n<li>resizing (via bicubic interpolation) by a factor of 0.8</li>\n<li>resizing (via bicubic interpolation) by a factor of 1.5</li>\n<li>resizing (via bicubic interpolation) by a factor of 2.0</li>\n<li>gamma correction using gamma = 0.8</li>\n<li>gamma correction using gamma = 1.2</li></ol></li>\n<li><p>Trained on unaltered patches and set each manipulation probability to 5%. No rotation or flipping was used, though in retrospect, those could have improved my score slightly.</p></li>\n<li>Validation patches are not from training images to ensure no possibility of data leakage.</li>\n</ul>\n\n<h2>Models:</h2>\n\n<ul>\n<li>ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3, all pretrained on ImageNet.</li>\n<li>All models use Global Average Pooling after the final convolution followed by a softmax layer. I experimented with fully-connected layers with dropout and regularization, but they didn't seem to improve the score.</li>\n</ul>\n\n<h2>Training:</h2>\n\n<ul>\n<li>Adam optimizer, learning rate: 0.0001    </li>\n<li>Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch; in most cases, the 3rd epoch actually lowered accuracy. I tried retraining with a learning rate of 0.00005 for another epoch, and managed to boost validation accuracies of most models to 0.975.  </li>\n</ul>\n\n<h2>Prediction:</h2>\n\n<ul>\n<li>Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. The validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.</li>\n</ul>\n\n<p>Adding some checkpoint models helped raise private LB score from 0.983 to 0.986. The additional models helped with corner cases.  </p>\n\n<p>The results could have been improved by adding more models like DenseNet201, NASNet, etc., and training for more epochs. But I guess at this point its for future competitions.</p>",
      "rawMarkdown": "First, congratulations to all the winners and those that did well. Thanks to Kaggle, IEEE, and everyone who shared their ideas. This is my first gold medal.\n\nConvolutional neural networks were able to learn meaningful relationships, which is quite interesting. As a result of this competition, it is evident that deep learning is a powerful tool, generalizable to many tasks. The final model ensemble obtains a private LB score of 0.986 (Public: 0.983), using ImageNet models ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3.\n\n\n##Data Preparation:\n* Slice each image into 256 x 256 non-overlapping patches and discard any images with only one unique pixel, since they don't help with camera identification; Gleb's additional dataset was used. Each patch retains the label of the original image. My data augmentation included the eight possible manipulations.  \n    1. JPEG compression with quality factor = 70\n    2. JPEG compression with quality factor = 90\n    3. resizing (via bicubic interpolation) by a factor of 0.5\n    4. resizing (via bicubic interpolation) by a factor of 0.8\n    5. resizing (via bicubic interpolation) by a factor of 1.5\n    6. resizing (via bicubic interpolation) by a factor of 2.0\n    7. gamma correction using gamma = 0.8\n    8. gamma correction using gamma = 1.2\n\n* Trained on unaltered patches and set each manipulation probability to 5%. No rotation or flipping was used, though in retrospect, those could have improved my score slightly.\n* Validation patches are not from training images to ensure no possibility of data leakage.\n\n\n##Models:\n* ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3, all pretrained on ImageNet.\n* All models use Global Average Pooling after the final convolution followed by a softmax layer. I experimented with fully-connected layers with dropout and regularization, but they didn't seem to improve the score.\n\n##Training:\n* Adam optimizer, learning rate: 0.0001    \n* Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch; in most cases, the 3rd epoch actually lowered accuracy. I tried retraining with a learning rate of 0.00005 for another epoch, and managed to boost validation accuracies of most models to 0.975.  \n\n\n##Prediction:\n*  Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. The validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.\n\n\nAdding some checkpoint models helped raise private LB score from 0.983 to 0.986. The additional models helped with corner cases.  \n\nThe results could have been improved by adding more models like DenseNet201, NASNet, etc., and training for more epochs. But I guess at this point its for future competitions.",
      "votes": null
    },
    {
      "id": "279998",
      "postDate": "02/09/2018 04:21:16",
      "content": "<p>Thank you for your post.</p>\n\n<blockquote>\n  <p>Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch</p>\n</blockquote>\n\n<p>That's very interesting. Knowing that this was possible I probably would have changed approaches earlier. I wonder if there's a lesson to be learned here. On ResNet50 I was getting about 0.75 validation accuracy after the second epoch. I center cropped the training images and validation images to 512x512, after applying manipulations, to match the test set. Also, I used Ivan's version of ResNet50 (the one with the binary manip feature included in the first FC layer). However, after looking at the code again I notice that there are no non-linearities after the the FC layers: <a href=\"https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/custom_models.py#L18\">https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/custom_models.py#L18</a></p>",
      "rawMarkdown": "Thank you for your post.\n\n&gt;Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch\n\nThat's very interesting. Knowing that this was possible I probably would have changed approaches earlier. I wonder if there's a lesson to be learned here. On ResNet50 I was getting about 0.75 validation accuracy after the second epoch. I center cropped the training images and validation images to 512x512, after applying manipulations, to match the test set. Also, I used Ivan's version of ResNet50 (the one with the binary manip feature included in the first FC layer). However, after looking at the code again I notice that there are no non-linearities after the the FC layers: https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/custom_models.py#L18",
      "votes": null
    },
    {
      "id": "280003",
      "postDate": "02/09/2018 04:32:34",
      "content": "<p>congrats to your gold model. Could I ask two questions? One is how do you ensemble your model, voting or average your probability? Another is how do you set your validation set? Thanks.</p>",
      "rawMarkdown": "congrats to your gold model. Could I ask two questions? One is how do you ensemble your model, voting or average your probability? Another is how do you set your validation set? Thanks.",
      "votes": null
    },
    {
      "id": "280009",
      "postDate": "02/09/2018 04:44:01",
      "content": "<p>My resizing was more like slicing an image into 256 x 256 patches, as opposed to merely using a center crop; this significantly increases the training data. I also initialized models from keras applications with imagenet weights, so that could have made a difference.</p>",
      "rawMarkdown": "My resizing was more like slicing an image into 256 x 256 patches, as opposed to merely using a center crop; this significantly increases the training data. I also initialized models from keras applications with imagenet weights, so that could have made a difference.",
      "votes": null
    },
    {
      "id": "280010",
      "postDate": "02/09/2018 04:51:01",
      "content": "<p>I also used pretrained imagenet weights, but I used PyTorch instead of Keras and used a slightly different model (shown in link above). I'll give the 256 x 256 slices a try tomorrow and post the results here.</p>\n\n<p>Did you default to slices or did you deliberately choose them over random patches?</p>",
      "rawMarkdown": "I also used pretrained imagenet weights, but I used PyTorch instead of Keras and used a slightly different model (shown in link above). I'll give the 256 x 256 slices a try tomorrow and post the results here.\n\nDid you default to slices or did you deliberately choose them over random patches?",
      "votes": null
    },
    {
      "id": "280012",
      "postDate": "02/09/2018 04:58:59",
      "content": "<p>Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. \nThe validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.</p>",
      "rawMarkdown": "Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. \nThe validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.",
      "votes": null
    },
    {
      "id": "280021",
      "postDate": "02/09/2018 05:20:03",
      "content": "<p>For some reason, random patches didn't yield any accuracy improvement for me, so I went for slicing instead. It also prevents overlapping, streamlining the training process.</p>",
      "rawMarkdown": "For some reason, random patches didn't yield any accuracy improvement for me, so I went for slicing instead. It also prevents overlapping, streamlining the training process.",
      "votes": null
    },
    {
      "id": "280045",
      "postDate": "02/09/2018 06:43:42",
      "content": "<p>Since you slice an image to 256x256 patches, how do you handle data augmentation of resizing by 0.5</p>",
      "rawMarkdown": "Since you slice an image to 256x256 patches, how do you handle data augmentation of resizing by 0.5",
      "votes": null
    },
    {
      "id": "280079",
      "postDate": "02/09/2018 09:15:35",
      "content": "<p>Thanks for your reply. </p>",
      "rawMarkdown": "Thanks for your reply.",
      "votes": null
    },
    {
      "id": "280433",
      "postDate": "02/10/2018 00:50:04",
      "content": "<p>Thanks for sharing. I would like to try to reproduce some of the results of the solutions that have been shared. \nI tried keras implementations, starting with imagenet weights, of Resnet50, InceptionV3, Densenet201 and MobileNet with patch sizes from 128x128 to 512x512. I wasn't able to hit those kind of accuracies.</p>\n\n<p>What kind of batch sizes where you using?</p>",
      "rawMarkdown": "Thanks for sharing. I would like to try to reproduce some of the results of the solutions that have been shared. \nI tried keras implementations, starting with imagenet weights, of Resnet50, InceptionV3, Densenet201 and MobileNet with patch sizes from 128x128 to 512x512. I wasn't able to hit those kind of accuracies.\n\nWhat kind of batch sizes where you using?",
      "votes": null
    },
    {
      "id": "280440",
      "postDate": "02/10/2018 01:33:10",
      "content": "<p>My guess: he resized the images before slicing them.</p>",
      "rawMarkdown": "My guess: he resized the images before slicing them.",
      "votes": null
    },
    {
      "id": "280445",
      "postDate": "02/10/2018 01:50:56",
      "content": "<p>64</p>",
      "rawMarkdown": "64",
      "votes": null
    },
    {
      "id": "280811",
      "postDate": "02/11/2018 05:02:40",
      "content": "<p>Congratulations! Thanks for sharing! Very interesting!</p>\n\n<p>I am especially interested in the validation accuracy score. \"Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch\". I tried many different models but never reach a valication accuracy score over 0.93. :(\nCould you share more about how you choose the training and validation set?</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing! Very interesting!\n\nI am especially interested in the validation accuracy score. \"Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch\". I tried many different models but never reach a valication accuracy score over 0.93. :(\nCould you share more about how you choose the training and validation set?",
      "votes": null
    },
    {
      "id": "282485",
      "postDate": "02/14/2018 03:40:52",
      "content": "<p>I'm going to try to reproduce your two-epoch results.</p>\n\n<blockquote>\n  <p>I only trained on 75% of unaltered patches and set each manipulation probability to 5%.</p>\n</blockquote>\n\n<p>Did you apply the manipulations to the patches directly? Or did you apply them to the image first and then extract the patches? This will affect the pixels on the borders of the patches.</p>\n\n<blockquote>\n  <p>I did not have enough time to utilize all the data. Thus, I only trained on 75% of unaltered patches</p>\n</blockquote>\n\n<p>Why did you choose to use less data rather than to train less? I would think seeing new information once would be better than seeing old information twice, but maybe it's problem dependent or close enough in benefit.</p>",
      "rawMarkdown": "I'm going to try to reproduce your two-epoch results.\n\n&gt; I only trained on 75% of unaltered patches and set each manipulation probability to 5%.\n\nDid you apply the manipulations to the patches directly? Or did you apply them to the image first and then extract the patches? This will affect the pixels on the borders of the patches.\n\n&gt; I did not have enough time to utilize all the data. Thus, I only trained on 75% of unaltered patches\n\nWhy did you choose to use less data rather than to train less? I would think seeing new information once would be better than seeing old information twice, but maybe it's problem dependent or close enough in benefit.",
      "votes": null
    },
    {
      "id": "283406",
      "postDate": "02/15/2018 13:30:48",
      "content": "<p>Hey Matt, let me know if you would like to collaborate on reproducing some of the published solution results. I want to see if I can get within 0.5% of Andres results and some of the other posted ones.</p>",
      "rawMarkdown": "Hey Matt, let me know if you would like to collaborate on reproducing some of the published solution results. I want to see if I can get within 0.5% of Andres results and some of the other posted ones.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 279998,
      "author_name": "kleinsmith",
      "author_url": "",
      "post_date": "02/09/2018 04:21:16",
      "content": "<p>Thank you for your post.</p>\n\n<blockquote>\n  <p>Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch</p>\n</blockquote>\n\n<p>That's very interesting. Knowing that this was possible I probably would have changed approaches earlier. I wonder if there's a lesson to be learned here. On ResNet50 I was getting about 0.75 validation accuracy after the second epoch. I center cropped the training images and validation images to 512x512, after applying manipulations, to match the test set. Also, I used Ivan's version of ResNet50 (the one with the binary manip feature included in the first FC layer). However, after looking at the code again I notice that there are no non-linearities after the the FC layers: <a href=\"https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/custom_models.py#L18\">https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/custom_models.py#L18</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 280009,
          "author_name": "nzabcd",
          "author_url": "",
          "post_date": "02/09/2018 04:44:01",
          "content": "<p>My resizing was more like slicing an image into 256 x 256 patches, as opposed to merely using a center crop; this significantly increases the training data. I also initialized models from keras applications with imagenet weights, so that could have made a difference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280010,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "02/09/2018 04:51:01",
          "content": "<p>I also used pretrained imagenet weights, but I used PyTorch instead of Keras and used a slightly different model (shown in link above). I'll give the 256 x 256 slices a try tomorrow and post the results here.</p>\n\n<p>Did you default to slices or did you deliberately choose them over random patches?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280021,
          "author_name": "nzabcd",
          "author_url": "",
          "post_date": "02/09/2018 05:20:03",
          "content": "<p>For some reason, random patches didn't yield any accuracy improvement for me, so I went for slicing instead. It also prevents overlapping, streamlining the training process.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280003,
      "author_name": "wuzuping",
      "author_url": "",
      "post_date": "02/09/2018 04:32:34",
      "content": "<p>congrats to your gold model. Could I ask two questions? One is how do you ensemble your model, voting or average your probability? Another is how do you set your validation set? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 280012,
          "author_name": "nzabcd",
          "author_url": "",
          "post_date": "02/09/2018 04:58:59",
          "content": "<p>Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. \nThe validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280079,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "02/09/2018 09:15:35",
          "content": "<p>Thanks for your reply. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280045,
      "author_name": "jiqiujia",
      "author_url": "",
      "post_date": "02/09/2018 06:43:42",
      "content": "<p>Since you slice an image to 256x256 patches, how do you handle data augmentation of resizing by 0.5</p>",
      "votes": null,
      "replies": [
        {
          "id": 280440,
          "author_name": "kleinsmith",
          "author_url": "",
          "post_date": "02/10/2018 01:33:10",
          "content": "<p>My guess: he resized the images before slicing them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280433,
      "author_name": "albertoa",
      "author_url": "",
      "post_date": "02/10/2018 00:50:04",
      "content": "<p>Thanks for sharing. I would like to try to reproduce some of the results of the solutions that have been shared. \nI tried keras implementations, starting with imagenet weights, of Resnet50, InceptionV3, Densenet201 and MobileNet with patch sizes from 128x128 to 512x512. I wasn't able to hit those kind of accuracies.</p>\n\n<p>What kind of batch sizes where you using?</p>",
      "votes": null,
      "replies": [
        {
          "id": 280445,
          "author_name": "nzabcd",
          "author_url": "",
          "post_date": "02/10/2018 01:50:56",
          "content": "<p>64</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280811,
      "author_name": "zhaoyangma",
      "author_url": "",
      "post_date": "02/11/2018 05:02:40",
      "content": "<p>Congratulations! Thanks for sharing! Very interesting!</p>\n\n<p>I am especially interested in the validation accuracy score. \"Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch\". I tried many different models but never reach a valication accuracy score over 0.93. :(\nCould you share more about how you choose the training and validation set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 282485,
      "author_name": "kleinsmith",
      "author_url": "",
      "post_date": "02/14/2018 03:40:52",
      "content": "<p>I'm going to try to reproduce your two-epoch results.</p>\n\n<blockquote>\n  <p>I only trained on 75% of unaltered patches and set each manipulation probability to 5%.</p>\n</blockquote>\n\n<p>Did you apply the manipulations to the patches directly? Or did you apply them to the image first and then extract the patches? This will affect the pixels on the borders of the patches.</p>\n\n<blockquote>\n  <p>I did not have enough time to utilize all the data. Thus, I only trained on 75% of unaltered patches</p>\n</blockquote>\n\n<p>Why did you choose to use less data rather than to train less? I would think seeing new information once would be better than seeing old information twice, but maybe it's problem dependent or close enough in benefit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 283406,
          "author_name": "albertoa",
          "author_url": "",
          "post_date": "02/15/2018 13:30:48",
          "content": "<p>Hey Matt, let me know if you would like to collaborate on reproducing some of the published solution results. I want to see if I can get within 0.5% of Andres results and some of the other posted ones.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "279990": "First, congratulations to all the winners and those that did well. Thanks to Kaggle, IEEE, and everyone who shared their ideas. This is my first gold medal.\n\nConvolutional neural networks were able to learn meaningful relationships, which is quite interesting. As a result of this competition, it is evident that deep learning is a powerful tool, generalizable to many tasks. The final model ensemble obtains a private LB score of 0.986 (Public: 0.983), using ImageNet models ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3.\n\n\n##Data Preparation:\n* Slice each image into 256 x 256 non-overlapping patches and discard any images with only one unique pixel, since they don't help with camera identification; Gleb's additional dataset was used. Each patch retains the label of the original image. My data augmentation included the eight possible manipulations.  \n    1. JPEG compression with quality factor = 70\n    2. JPEG compression with quality factor = 90\n    3. resizing (via bicubic interpolation) by a factor of 0.5\n    4. resizing (via bicubic interpolation) by a factor of 0.8\n    5. resizing (via bicubic interpolation) by a factor of 1.5\n    6. resizing (via bicubic interpolation) by a factor of 2.0\n    7. gamma correction using gamma = 0.8\n    8. gamma correction using gamma = 1.2\n\n* Trained on unaltered patches and set each manipulation probability to 5%. No rotation or flipping was used, though in retrospect, those could have improved my score slightly.\n* Validation patches are not from training images to ensure no possibility of data leakage.\n\n\n##Models:\n* ResNet50, DenseNet121, DenseNet169, Xception, and InceptionV3, all pretrained on ImageNet.\n* All models use Global Average Pooling after the final convolution followed by a softmax layer. I experimented with fully-connected layers with dropout and regularization, but they didn't seem to improve the score.\n\n##Training:\n* Adam optimizer, learning rate: 0.0001    \n* Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch; in most cases, the 3rd epoch actually lowered accuracy. I tried retraining with a learning rate of 0.00005 for another epoch, and managed to boost validation accuracies of most models to 0.975.  \n\n\n##Prediction:\n*  Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. The validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.\n\n\nAdding some checkpoint models helped raise private LB score from 0.983 to 0.986. The additional models helped with corner cases.  \n\nThe results could have been improved by adding more models like DenseNet201, NASNet, etc., and training for more epochs. But I guess at this point its for future competitions.",
    "279998": "Thank you for your post.\n\n&gt;Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch\n\nThat's very interesting. Knowing that this was possible I probably would have changed approaches earlier. I wonder if there's a lesson to be learned here. On ResNet50 I was getting about 0.75 validation accuracy after the second epoch. I center cropped the training images and validation images to 512x512, after applying manipulations, to match the test set. Also, I used Ivan's version of ResNet50 (the one with the binary manip feature included in the first FC layer). However, after looking at the code again I notice that there are no non-linearities after the the FC layers: https://github.com/irrmnv/pytorch-ieee-cmi/blob/master/custom_models.py#L18",
    "280003": "congrats to your gold model. Could I ask two questions? One is how do you ensemble your model, voting or average your probability? Another is how do you set your validation set? Thanks.",
    "280009": "My resizing was more like slicing an image into 256 x 256 patches, as opposed to merely using a center crop; this significantly increases the training data. I also initialized models from keras applications with imagenet weights, so that could have made a difference.",
    "280010": "I also used pretrained imagenet weights, but I used PyTorch instead of Keras and used a slightly different model (shown in link above). I'll give the 256 x 256 slices a try tomorrow and post the results here.\n\nDid you default to slices or did you deliberately choose them over random patches?",
    "280012": "Each test image is sliced into 4 256 x 256 images and augmented into 12 (vertical + horizontal flip). A model outputs raw probabilities for each image slice, which are then averaged together with those from other models. Then, the class for each slice is determined with np.argmax, and the mode is taken as the final prediction. \nThe validation set is just like the training set - with manipulations - just with fewer images. The original images are randomly assigned to training and validation sets so that validation crops do not come from the same image as training crops.",
    "280021": "For some reason, random patches didn't yield any accuracy improvement for me, so I went for slicing instead. It also prevents overlapping, streamlining the training process.",
    "280045": "Since you slice an image to 256x256 patches, how do you handle data augmentation of resizing by 0.5",
    "280079": "Thanks for your reply.",
    "280433": "Thanks for sharing. I would like to try to reproduce some of the results of the solutions that have been shared. \nI tried keras implementations, starting with imagenet weights, of Resnet50, InceptionV3, Densenet201 and MobileNet with patch sizes from 128x128 to 512x512. I wasn't able to hit those kind of accuracies.\n\nWhat kind of batch sizes where you using?",
    "280440": "My guess: he resized the images before slicing them.",
    "280445": "64",
    "280811": "Congratulations! Thanks for sharing! Very interesting!\n\nI am especially interested in the validation accuracy score. \"Each model achieves between 0.955 and 0.975 validation accuracy after the 2nd epoch\". I tried many different models but never reach a valication accuracy score over 0.93. :(\nCould you share more about how you choose the training and validation set?",
    "282485": "I'm going to try to reproduce your two-epoch results.\n\n&gt; I only trained on 75% of unaltered patches and set each manipulation probability to 5%.\n\nDid you apply the manipulations to the patches directly? Or did you apply them to the image first and then extract the patches? This will affect the pixels on the borders of the patches.\n\n&gt; I did not have enough time to utilize all the data. Thus, I only trained on 75% of unaltered patches\n\nWhy did you choose to use less data rather than to train less? I would think seeing new information once would be better than seeing old information twice, but maybe it's problem dependent or close enough in benefit.",
    "283406": "Hey Matt, let me know if you would like to collaborate on reproducing some of the published solution results. I want to see if I can get within 0.5% of Andres results and some of the other posted ones."
  },
  "source": "meta"
}