{
  "id": 49314,
  "title": "My solution: Single model with a few quirks.",
  "url": "/competitions/sp-society-camera-model-identification/writeups/andres-torrubia-my-solution-single-model-with-a-fe",
  "author_name": "",
  "post_date": "2018-02-09T09:56:26.200473600Z",
  "votes": 70,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi all! I had a lot of fun participating in this competition and in retrospect I've learned a few lessons which I'll share. Although I didn't score as high as I intended (my personal goal was to get 0.98 single model, gold medal) I am happy of having helped a lot of teams, both with the code, ideas and dataset.</p>\n\n<p>My latest code is @ <a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">Github repo</a> and you'll see that while I added a BUNCH of features, I didn't have time to even test them all (as most of us I'm limited by GPUs).</p>\n\n<h1>Solution overview</h1>\n\n<ul>\n<li>Single model (no ensemble)</li>\n<li>Inputs: 512x512 crop size, whether the image is manipulated or not,\nand relative location of crop wrt canonical (i.e. vertically\noriented) sensor resolution (e.g. center is 0,0; top left -1,1,\nbottom right 1,1). Random crops are displaced in 2x multiples since most CFA arrays are 2x2 so it is aligned.</li>\n<li>Outputs: Class and type of manipulation performed, manipulation prediction is fed to main (class) FC head, with the expectation that knowing the manipulation will help the main classifier.</li>\n<li>Feature extractor: DenseNet201 using avg pooling</li>\n<li>Two fully connected layers 512, and 256; with 0.3 dropout each for main </li>\n<li>Class-aware sampling</li>\n<li>Test-time augmentation using 28 total augmentations (for <code>unalt</code> images, less so for <code>manip ones</code>)</li>\n<li>Loss function: x-entropy of class + 0.1 x-entropy of manipulation</li>\n<li>Test probabilities equalization (instead of performing old-school\ninference)</li>\n<li>Used 2 x 1080 Tis and 1 x 1070 Ti</li>\n<li>Flickr CC dataset + organization dataset (I wasn't 100% positive Gleb's dataset was clean re: organization rules, so I didn't risk it).</li>\n</ul>\n\n<h1>What worked</h1>\n\n<ul>\n<li>Class-aware sampling: it got slightly better results than weighted class weights.</li>\n<li>Fine-tuning FC classes (both architecture and weights) freezing classifier.</li>\n<li>Flickr CC dataset :-)</li>\n<li>Equalizing test prediction probabilities: instead of just taking predicted classes, use probabilities to </li>\n</ul>\n\n<h1>What kind-of-worked</h1>\n\n<ul>\n<li>Feeding relative crop location to the main FC head. Motivation was that without information of where the crop is the net:  may not learn lens characteristics that easily (focal length, chromatic aberration, etc.) </li>\n<li>Validation set: my <code>val_acc</code> oscillated 0.01 points during epochs and then I saw a detrimental shakeup on private LB. I had a poor validation set.</li>\n</ul>\n\n<h1>What didn't work</h1>\n\n<ul>\n<li>Mixup: I had a lot of faith in mixup, however mixup didn't improve training accuracy beyond 0.97, reason is I was feeding random crops taken at different locations with different manipulations; and the dual-output net was expected to classify the mixed-up class PLUS the mixed-up augmentation type. I new it was far fetched, and I should have (for a given mixup pair) used the same manipulation type so the net would have focused much more on learning the main objetive.</li>\n<li>Experimental loss-function to \"force\" a flat target distribution. I thought this was so cool: adding a regularization term to the loss function which is basically the mean of the variance of target predictions within a batch. I needed to fine-tune the weight of the regularization term but didn't have time to test it properly. I could do this b/c freezing the classifier allowed me to focus on the FC layers which had ~10% of the net expressive power and do big batch sizes (128+) so the variance calculation was meaningful. Anyway, not sure if it even makes sense.</li>\n<li>Ensembles: At the end I had 6+ models with ~0.97ish LB. I only had &gt;ONE&lt; submission left and I coded a too far-fetched ensemble code which attempted to reject outliers (softprobs below a certain threshold) and at the same time equalizing the target distribution. I must have introduced a bug b/c the 6-model ensemble only got 0.95 LB... so I didn't use and I didn't have any extra submissions left.</li>\n</ul>\n\n<h1>Lessons learnt</h1>\n\n<ul>\n<li>Creating a good validation set is critical.</li>\n<li>Adding extra samples works.</li>\n<li>Should have used faux-labels</li>\n<li>Don't get too excited if you can add more features that the resources/time you have to test and compare them.</li>\n<li>Yes, I know, ensembles rule the LB. ;-)</li>\n<li>Pytorch is faster than Keras. </li>\n</ul>\n\n<p>Best and congratulations to all teams, it's been a great learning experience!</p>\n\n<p>-Andres</p>",
  "messages": [
    {
      "id": "280083",
      "postDate": "02/09/2018 09:56:26",
      "content": "<p>Hi all! I had a lot of fun participating in this competition and in retrospect I've learned a few lessons which I'll share. Although I didn't score as high as I intended (my personal goal was to get 0.98 single model, gold medal) I am happy of having helped a lot of teams, both with the code, ideas and dataset.</p>\n\n<p>My latest code is @ <a href=\"https://github.com/antorsae/sp-society-camera-model-identification\">Github repo</a> and you'll see that while I added a BUNCH of features, I didn't have time to even test them all (as most of us I'm limited by GPUs).</p>\n\n<h1>Solution overview</h1>\n\n<ul>\n<li>Single model (no ensemble)</li>\n<li>Inputs: 512x512 crop size, whether the image is manipulated or not,\nand relative location of crop wrt canonical (i.e. vertically\noriented) sensor resolution (e.g. center is 0,0; top left -1,1,\nbottom right 1,1). Random crops are displaced in 2x multiples since most CFA arrays are 2x2 so it is aligned.</li>\n<li>Outputs: Class and type of manipulation performed, manipulation prediction is fed to main (class) FC head, with the expectation that knowing the manipulation will help the main classifier.</li>\n<li>Feature extractor: DenseNet201 using avg pooling</li>\n<li>Two fully connected layers 512, and 256; with 0.3 dropout each for main </li>\n<li>Class-aware sampling</li>\n<li>Test-time augmentation using 28 total augmentations (for <code>unalt</code> images, less so for <code>manip ones</code>)</li>\n<li>Loss function: x-entropy of class + 0.1 x-entropy of manipulation</li>\n<li>Test probabilities equalization (instead of performing old-school\ninference)</li>\n<li>Used 2 x 1080 Tis and 1 x 1070 Ti</li>\n<li>Flickr CC dataset + organization dataset (I wasn't 100% positive Gleb's dataset was clean re: organization rules, so I didn't risk it).</li>\n</ul>\n\n<h1>What worked</h1>\n\n<ul>\n<li>Class-aware sampling: it got slightly better results than weighted class weights.</li>\n<li>Fine-tuning FC classes (both architecture and weights) freezing classifier.</li>\n<li>Flickr CC dataset :-)</li>\n<li>Equalizing test prediction probabilities: instead of just taking predicted classes, use probabilities to </li>\n</ul>\n\n<h1>What kind-of-worked</h1>\n\n<ul>\n<li>Feeding relative crop location to the main FC head. Motivation was that without information of where the crop is the net:  may not learn lens characteristics that easily (focal length, chromatic aberration, etc.) </li>\n<li>Validation set: my <code>val_acc</code> oscillated 0.01 points during epochs and then I saw a detrimental shakeup on private LB. I had a poor validation set.</li>\n</ul>\n\n<h1>What didn't work</h1>\n\n<ul>\n<li>Mixup: I had a lot of faith in mixup, however mixup didn't improve training accuracy beyond 0.97, reason is I was feeding random crops taken at different locations with different manipulations; and the dual-output net was expected to classify the mixed-up class PLUS the mixed-up augmentation type. I new it was far fetched, and I should have (for a given mixup pair) used the same manipulation type so the net would have focused much more on learning the main objetive.</li>\n<li>Experimental loss-function to \"force\" a flat target distribution. I thought this was so cool: adding a regularization term to the loss function which is basically the mean of the variance of target predictions within a batch. I needed to fine-tune the weight of the regularization term but didn't have time to test it properly. I could do this b/c freezing the classifier allowed me to focus on the FC layers which had ~10% of the net expressive power and do big batch sizes (128+) so the variance calculation was meaningful. Anyway, not sure if it even makes sense.</li>\n<li>Ensembles: At the end I had 6+ models with ~0.97ish LB. I only had &gt;ONE&lt; submission left and I coded a too far-fetched ensemble code which attempted to reject outliers (softprobs below a certain threshold) and at the same time equalizing the target distribution. I must have introduced a bug b/c the 6-model ensemble only got 0.95 LB... so I didn't use and I didn't have any extra submissions left.</li>\n</ul>\n\n<h1>Lessons learnt</h1>\n\n<ul>\n<li>Creating a good validation set is critical.</li>\n<li>Adding extra samples works.</li>\n<li>Should have used faux-labels</li>\n<li>Don't get too excited if you can add more features that the resources/time you have to test and compare them.</li>\n<li>Yes, I know, ensembles rule the LB. ;-)</li>\n<li>Pytorch is faster than Keras. </li>\n</ul>\n\n<p>Best and congratulations to all teams, it's been a great learning experience!</p>\n\n<p>-Andres</p>",
      "rawMarkdown": "Hi all! I had a lot of fun participating in this competition and in retrospect I've learned a few lessons which I'll share. Although I didn't score as high as I intended (my personal goal was to get 0.98 single model, gold medal) I am happy of having helped a lot of teams, both with the code, ideas and dataset.\n\nMy latest code is @ [Github repo][1] and you'll see that while I added a BUNCH of features, I didn't have time to even test them all (as most of us I'm limited by GPUs).\n\n# Solution overview #\n\n - Single model (no ensemble)\n - Inputs: 512x512 crop size, whether the image is manipulated or not,\n   and relative location of crop wrt canonical (i.e. vertically\n   oriented) sensor resolution (e.g. center is 0,0; top left -1,1,\n   bottom right 1,1). Random crops are displaced in 2x multiples since most CFA arrays are 2x2 so it is aligned.\n - Outputs: Class and type of manipulation performed, manipulation prediction is fed to main (class) FC head, with the expectation that knowing the manipulation will help the main classifier.\n - Feature extractor: DenseNet201 using avg pooling\n - Two fully connected layers 512, and 256; with 0.3 dropout each for main \n - Class-aware sampling\n - Test-time augmentation using 28 total augmentations (for `unalt` images, less so for `manip ones`)\n - Loss function: x-entropy of class + 0.1 x-entropy of manipulation\n - Test probabilities equalization (instead of performing old-school\n   inference)\n - Used 2 x 1080 Tis and 1 x 1070 Ti\n - Flickr CC dataset + organization dataset (I wasn't 100% positive Gleb's dataset was clean re: organization rules, so I didn't risk it).\n\n# What worked\n\n - Class-aware sampling: it got slightly better results than weighted class weights.\n - Fine-tuning FC classes (both architecture and weights) freezing classifier.\n - Flickr CC dataset :-)\n - Equalizing test prediction probabilities: instead of just taking predicted classes, use probabilities to \n\n# What kind-of-worked\n\n - Feeding relative crop location to the main FC head. Motivation was that without information of where the crop is the net:  may not learn lens characteristics that easily (focal length, chromatic aberration, etc.) \n - Validation set: my `val_acc` oscillated 0.01 points during epochs and then I saw a detrimental shakeup on private LB. I had a poor validation set.\n\n# What didn't work\n\n - Mixup: I had a lot of faith in mixup, however mixup didn't improve training accuracy beyond 0.97, reason is I was feeding random crops taken at different locations with different manipulations; and the dual-output net was expected to classify the mixed-up class PLUS the mixed-up augmentation type. I new it was far fetched, and I should have (for a given mixup pair) used the same manipulation type so the net would have focused much more on learning the main objetive.\n - Experimental loss-function to \"force\" a flat target distribution. I thought this was so cool: adding a regularization term to the loss function which is basically the mean of the variance of target predictions within a batch. I needed to fine-tune the weight of the regularization term but didn't have time to test it properly. I could do this b/c freezing the classifier allowed me to focus on the FC layers which had ~10% of the net expressive power and do big batch sizes (128+) so the variance calculation was meaningful. Anyway, not sure if it even makes sense.\n - Ensembles: At the end I had 6+ models with ~0.97ish LB. I only had &gt;ONE&lt; submission left and I coded a too far-fetched ensemble code which attempted to reject outliers (softprobs below a certain threshold) and at the same time equalizing the target distribution. I must have introduced a bug b/c the 6-model ensemble only got 0.95 LB... so I didn't use and I didn't have any extra submissions left.\n\n# Lessons learnt\n\n - Creating a good validation set is critical.\n - Adding extra samples works.\n - Should have used faux-labels\n - Don't get too excited if you can add more features that the resources/time you have to test and compare them.\n - Yes, I know, ensembles rule the LB. ;-)\n - Pytorch is faster than Keras. \n\nBest and congratulations to all teams, it's been a great learning experience!\n\n-Andres\n\n  [1]: https://github.com/antorsae/sp-society-camera-model-identification",
      "votes": null
    },
    {
      "id": "280106",
      "postDate": "02/09/2018 11:33:15",
      "content": "<p>Thanks for sharing your code and thoughts during this competition. Your impact is enormous and your work is definitely one of reasons of success of many teams. </p>",
      "rawMarkdown": "Thanks for sharing your code and thoughts during this competition. Your impact is enormous and your work is definitely one of reasons of success of many teams.",
      "votes": null
    },
    {
      "id": "280153",
      "postDate": "02/09/2018 13:36:59",
      "content": "<p>Huge respect for sharing your code! Without it, there would not be such amazing results in terms of final accuracy. I understand that it may be sad that others have used your code. I did same in other competitions (DSTL, DSB). But from sharing code everyone won: I found a bug in my code, new people came with new ideas. And as a result, it has moved the community to a new level of understanding of semantic segmentation. Also here, with the help of your solution, after a number of improvements, one can get a single model that shows very high accuracy without manual features and low-level sensor processing. And this contradicts the large number of publications.</p>",
      "rawMarkdown": "Huge respect for sharing your code! Without it, there would not be such amazing results in terms of final accuracy. I understand that it may be sad that others have used your code. I did same in other competitions (DSTL, DSB). But from sharing code everyone won: I found a bug in my code, new people came with new ideas. And as a result, it has moved the community to a new level of understanding of semantic segmentation. Also here, with the help of your solution, after a number of improvements, one can get a single model that shows very high accuracy without manual features and low-level sensor processing. And this contradicts the large number of publications.",
      "votes": null
    },
    {
      "id": "280164",
      "postDate": "02/09/2018 13:53:17",
      "content": "<p>Update: I submit all my CSVs to see what model was really the best... it turns out:</p>\n\n<p><code>submission_DenseNet201_cs512_fc512,512,512_doc0.0_do0.5_dol0.3_avg_xx_cas-epoch057-val_acc0.978009.csv</code></p>\n\n<p>It's a single-model, 512crop DenseNet201 w/ 2 heads of 3 FC connected layers, class-aware sampling and Flickr CC dataset... but importantly <em>no TTA</em> and <em>no test probability distribution equalization</em>. </p>\n\n<p>This model gets 0.977 in the private LB and 0.966 in the public one... so I think my biggest mistake was to have a poor validation set.</p>",
      "rawMarkdown": "Update: I submit all my CSVs to see what model was really the best... it turns out:\n\n`submission_DenseNet201_cs512_fc512,512,512_doc0.0_do0.5_dol0.3_avg_xx_cas-epoch057-val_acc0.978009.csv`\n\nIt's a single-model, 512crop DenseNet201 w/ 2 heads of 3 FC connected layers, class-aware sampling and Flickr CC dataset... but importantly *no TTA* and *no test probability distribution equalization*. \n\nThis model gets 0.977 in the private LB and 0.966 in the public one... so I think my biggest mistake was to have a poor validation set.",
      "votes": null
    },
    {
      "id": "280169",
      "postDate": "02/09/2018 13:57:56",
      "content": "<p>Thanks. Not sad at all about sharing the code, as you say everybody wins.</p>\n\n<p>I'm too positively surprised that an almost vanilla NN can distinguish camera sensors even after manipulations (and my net could identify <em>which</em> manipulation was performed with 0.97+ validation accuracy) ... it kind of blows your mind... :-)</p>",
      "rawMarkdown": "Thanks. Not sad at all about sharing the code, as you say everybody wins.\n\nI'm too positively surprised that an almost vanilla NN can distinguish camera sensors even after manipulations (and my net could identify *which* manipulation was performed with 0.97+ validation accuracy) ... it kind of blows your mind... :-)",
      "votes": null
    },
    {
      "id": "280171",
      "postDate": "02/09/2018 14:01:26",
      "content": "<p>Thank you Andres, I learned a lot from your ideas. Particularly, I was trying to get better results with my own approach without reading your code, but only with the discussions I increase my knowledge.</p>\n\n<p>You may not have achieved the gold medal but earned the respect of all teams!</p>",
      "rawMarkdown": "Thank you Andres, I learned a lot from your ideas. Particularly, I was trying to get better results with my own approach without reading your code, but only with the discussions I increase my knowledge.\n\nYou may not have achieved the gold medal but earned the respect of all teams!",
      "votes": null
    },
    {
      "id": "280193",
      "postDate": "02/09/2018 14:50:09",
      "content": "<p>Totally agreed about mind blowing performance of vanilla imagenet-like CNN. Because all articles said that such architecture can't find low level features that corresponding to camera sensor and processing. But imagenet-like architecture seems almost ultimate weapon. All our tricks is about training and ensemble, except augmentation flag input. </p>",
      "rawMarkdown": "Totally agreed about mind blowing performance of vanilla imagenet-like CNN. Because all articles said that such architecture can't find low level features that corresponding to camera sensor and processing. But imagenet-like architecture seems almost ultimate weapon. All our tricks is about training and ensemble, except augmentation flag input.",
      "votes": null
    },
    {
      "id": "280314",
      "postDate": "02/09/2018 17:51:10",
      "content": "<p>From the getgo I thought this would be a very different kind of image competition, as the organizers even indicated. Had no idea that the vanilla CNNs could be competitive, let alone perform so well. Which is why it took me a while to even start competing seriously here. </p>",
      "rawMarkdown": "From the getgo I thought this would be a very different kind of image competition, as the organizers even indicated. Had no idea that the vanilla CNNs could be competitive, let alone perform so well. Which is why it took me a while to even start competing seriously here.",
      "votes": null
    },
    {
      "id": "280317",
      "postDate": "02/09/2018 17:52:51",
      "content": "<p>Thank you for sharing your code, and for all these additional insights. Kaggle should really have a separate prize for people like you who are this helpful in a competition. </p>",
      "rawMarkdown": "Thank you for sharing your code, and for all these additional insights. Kaggle should really have a separate prize for people like you who are this helpful in a competition.",
      "votes": null
    },
    {
      "id": "280331",
      "postDate": "02/09/2018 18:14:42",
      "content": "<p>Thank you very much for your sharing, I learn a lot thanks to you (and I'm clearly not the only one!). I agree with @Bojan, you totally deserve a special prize. </p>",
      "rawMarkdown": "Thank you very much for your sharing, I learn a lot thanks to you (and I'm clearly not the only one!). I agree with @Bojan, you totally deserve a special prize.",
      "votes": null
    },
    {
      "id": "280390",
      "postDate": "02/09/2018 21:45:55",
      "content": "<p>I did not use Andres code, nor did I do as well as many here.  I got to ~80% on plain vanilla CNNs.  What I did enjoy was the community and discussion wrt Andres work and the spirit he brought to the competition.  Well done!</p>",
      "rawMarkdown": "I did not use Andres code, nor did I do as well as many here.  I got to ~80% on plain vanilla CNNs.  What I did enjoy was the community and discussion wrt Andres work and the spirit he brought to the competition.  Well done!",
      "votes": null
    },
    {
      "id": "280605",
      "postDate": "02/10/2018 12:37:12",
      "content": "<p>Hello and thank you so much for sharing the code before the deadline of the competition. It was my first participation in Kaggle and having available the pipeline of someone with  your experience taught me so much in so little time. I fully agree that you should also get some kind of prize :). Good luck! </p>",
      "rawMarkdown": "Hello and thank you so much for sharing the code before the deadline of the competition. It was my first participation in Kaggle and having available the pipeline of someone with  your experience taught me so much in so little time. I fully agree that you should also get some kind of prize :). Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 280106,
      "author_name": "opanichev",
      "author_url": "",
      "post_date": "02/09/2018 11:33:15",
      "content": "<p>Thanks for sharing your code and thoughts during this competition. Your impact is enormous and your work is definitely one of reasons of success of many teams. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280153,
      "author_name": "drn01z3",
      "author_url": "",
      "post_date": "02/09/2018 13:36:59",
      "content": "<p>Huge respect for sharing your code! Without it, there would not be such amazing results in terms of final accuracy. I understand that it may be sad that others have used your code. I did same in other competitions (DSTL, DSB). But from sharing code everyone won: I found a bug in my code, new people came with new ideas. And as a result, it has moved the community to a new level of understanding of semantic segmentation. Also here, with the help of your solution, after a number of improvements, one can get a single model that shows very high accuracy without manual features and low-level sensor processing. And this contradicts the large number of publications.</p>",
      "votes": null,
      "replies": [
        {
          "id": 280169,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/09/2018 13:57:56",
          "content": "<p>Thanks. Not sad at all about sharing the code, as you say everybody wins.</p>\n\n<p>I'm too positively surprised that an almost vanilla NN can distinguish camera sensors even after manipulations (and my net could identify <em>which</em> manipulation was performed with 0.97+ validation accuracy) ... it kind of blows your mind... :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280193,
          "author_name": "drn01z3",
          "author_url": "",
          "post_date": "02/09/2018 14:50:09",
          "content": "<p>Totally agreed about mind blowing performance of vanilla imagenet-like CNN. Because all articles said that such architecture can't find low level features that corresponding to camera sensor and processing. But imagenet-like architecture seems almost ultimate weapon. All our tricks is about training and ensemble, except augmentation flag input. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280314,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "02/09/2018 17:51:10",
          "content": "<p>From the getgo I thought this would be a very different kind of image competition, as the organizers even indicated. Had no idea that the vanilla CNNs could be competitive, let alone perform so well. Which is why it took me a while to even start competing seriously here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 280390,
          "author_name": "jbfarrar",
          "author_url": "",
          "post_date": "02/09/2018 21:45:55",
          "content": "<p>I did not use Andres code, nor did I do as well as many here.  I got to ~80% on plain vanilla CNNs.  What I did enjoy was the community and discussion wrt Andres work and the spirit he brought to the competition.  Well done!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280164,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/09/2018 13:53:17",
      "content": "<p>Update: I submit all my CSVs to see what model was really the best... it turns out:</p>\n\n<p><code>submission_DenseNet201_cs512_fc512,512,512_doc0.0_do0.5_dol0.3_avg_xx_cas-epoch057-val_acc0.978009.csv</code></p>\n\n<p>It's a single-model, 512crop DenseNet201 w/ 2 heads of 3 FC connected layers, class-aware sampling and Flickr CC dataset... but importantly <em>no TTA</em> and <em>no test probability distribution equalization</em>. </p>\n\n<p>This model gets 0.977 in the private LB and 0.966 in the public one... so I think my biggest mistake was to have a poor validation set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280171,
      "author_name": "igormunizims",
      "author_url": "",
      "post_date": "02/09/2018 14:01:26",
      "content": "<p>Thank you Andres, I learned a lot from your ideas. Particularly, I was trying to get better results with my own approach without reading your code, but only with the discussions I increase my knowledge.</p>\n\n<p>You may not have achieved the gold medal but earned the respect of all teams!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280317,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "02/09/2018 17:52:51",
      "content": "<p>Thank you for sharing your code, and for all these additional insights. Kaggle should really have a separate prize for people like you who are this helpful in a competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280331,
      "author_name": "arthurllau",
      "author_url": "",
      "post_date": "02/09/2018 18:14:42",
      "content": "<p>Thank you very much for your sharing, I learn a lot thanks to you (and I'm clearly not the only one!). I agree with @Bojan, you totally deserve a special prize. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280605,
      "author_name": "vmilias",
      "author_url": "",
      "post_date": "02/10/2018 12:37:12",
      "content": "<p>Hello and thank you so much for sharing the code before the deadline of the competition. It was my first participation in Kaggle and having available the pipeline of someone with  your experience taught me so much in so little time. I fully agree that you should also get some kind of prize :). Good luck! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "280083": "Hi all! I had a lot of fun participating in this competition and in retrospect I've learned a few lessons which I'll share. Although I didn't score as high as I intended (my personal goal was to get 0.98 single model, gold medal) I am happy of having helped a lot of teams, both with the code, ideas and dataset.\n\nMy latest code is @ [Github repo][1] and you'll see that while I added a BUNCH of features, I didn't have time to even test them all (as most of us I'm limited by GPUs).\n\n# Solution overview #\n\n - Single model (no ensemble)\n - Inputs: 512x512 crop size, whether the image is manipulated or not,\n   and relative location of crop wrt canonical (i.e. vertically\n   oriented) sensor resolution (e.g. center is 0,0; top left -1,1,\n   bottom right 1,1). Random crops are displaced in 2x multiples since most CFA arrays are 2x2 so it is aligned.\n - Outputs: Class and type of manipulation performed, manipulation prediction is fed to main (class) FC head, with the expectation that knowing the manipulation will help the main classifier.\n - Feature extractor: DenseNet201 using avg pooling\n - Two fully connected layers 512, and 256; with 0.3 dropout each for main \n - Class-aware sampling\n - Test-time augmentation using 28 total augmentations (for `unalt` images, less so for `manip ones`)\n - Loss function: x-entropy of class + 0.1 x-entropy of manipulation\n - Test probabilities equalization (instead of performing old-school\n   inference)\n - Used 2 x 1080 Tis and 1 x 1070 Ti\n - Flickr CC dataset + organization dataset (I wasn't 100% positive Gleb's dataset was clean re: organization rules, so I didn't risk it).\n\n# What worked\n\n - Class-aware sampling: it got slightly better results than weighted class weights.\n - Fine-tuning FC classes (both architecture and weights) freezing classifier.\n - Flickr CC dataset :-)\n - Equalizing test prediction probabilities: instead of just taking predicted classes, use probabilities to \n\n# What kind-of-worked\n\n - Feeding relative crop location to the main FC head. Motivation was that without information of where the crop is the net:  may not learn lens characteristics that easily (focal length, chromatic aberration, etc.) \n - Validation set: my `val_acc` oscillated 0.01 points during epochs and then I saw a detrimental shakeup on private LB. I had a poor validation set.\n\n# What didn't work\n\n - Mixup: I had a lot of faith in mixup, however mixup didn't improve training accuracy beyond 0.97, reason is I was feeding random crops taken at different locations with different manipulations; and the dual-output net was expected to classify the mixed-up class PLUS the mixed-up augmentation type. I new it was far fetched, and I should have (for a given mixup pair) used the same manipulation type so the net would have focused much more on learning the main objetive.\n - Experimental loss-function to \"force\" a flat target distribution. I thought this was so cool: adding a regularization term to the loss function which is basically the mean of the variance of target predictions within a batch. I needed to fine-tune the weight of the regularization term but didn't have time to test it properly. I could do this b/c freezing the classifier allowed me to focus on the FC layers which had ~10% of the net expressive power and do big batch sizes (128+) so the variance calculation was meaningful. Anyway, not sure if it even makes sense.\n - Ensembles: At the end I had 6+ models with ~0.97ish LB. I only had &gt;ONE&lt; submission left and I coded a too far-fetched ensemble code which attempted to reject outliers (softprobs below a certain threshold) and at the same time equalizing the target distribution. I must have introduced a bug b/c the 6-model ensemble only got 0.95 LB... so I didn't use and I didn't have any extra submissions left.\n\n# Lessons learnt\n\n - Creating a good validation set is critical.\n - Adding extra samples works.\n - Should have used faux-labels\n - Don't get too excited if you can add more features that the resources/time you have to test and compare them.\n - Yes, I know, ensembles rule the LB. ;-)\n - Pytorch is faster than Keras. \n\nBest and congratulations to all teams, it's been a great learning experience!\n\n-Andres\n\n  [1]: https://github.com/antorsae/sp-society-camera-model-identification",
    "280106": "Thanks for sharing your code and thoughts during this competition. Your impact is enormous and your work is definitely one of reasons of success of many teams.",
    "280153": "Huge respect for sharing your code! Without it, there would not be such amazing results in terms of final accuracy. I understand that it may be sad that others have used your code. I did same in other competitions (DSTL, DSB). But from sharing code everyone won: I found a bug in my code, new people came with new ideas. And as a result, it has moved the community to a new level of understanding of semantic segmentation. Also here, with the help of your solution, after a number of improvements, one can get a single model that shows very high accuracy without manual features and low-level sensor processing. And this contradicts the large number of publications.",
    "280164": "Update: I submit all my CSVs to see what model was really the best... it turns out:\n\n`submission_DenseNet201_cs512_fc512,512,512_doc0.0_do0.5_dol0.3_avg_xx_cas-epoch057-val_acc0.978009.csv`\n\nIt's a single-model, 512crop DenseNet201 w/ 2 heads of 3 FC connected layers, class-aware sampling and Flickr CC dataset... but importantly *no TTA* and *no test probability distribution equalization*. \n\nThis model gets 0.977 in the private LB and 0.966 in the public one... so I think my biggest mistake was to have a poor validation set.",
    "280169": "Thanks. Not sad at all about sharing the code, as you say everybody wins.\n\nI'm too positively surprised that an almost vanilla NN can distinguish camera sensors even after manipulations (and my net could identify *which* manipulation was performed with 0.97+ validation accuracy) ... it kind of blows your mind... :-)",
    "280171": "Thank you Andres, I learned a lot from your ideas. Particularly, I was trying to get better results with my own approach without reading your code, but only with the discussions I increase my knowledge.\n\nYou may not have achieved the gold medal but earned the respect of all teams!",
    "280193": "Totally agreed about mind blowing performance of vanilla imagenet-like CNN. Because all articles said that such architecture can't find low level features that corresponding to camera sensor and processing. But imagenet-like architecture seems almost ultimate weapon. All our tricks is about training and ensemble, except augmentation flag input.",
    "280314": "From the getgo I thought this would be a very different kind of image competition, as the organizers even indicated. Had no idea that the vanilla CNNs could be competitive, let alone perform so well. Which is why it took me a while to even start competing seriously here.",
    "280317": "Thank you for sharing your code, and for all these additional insights. Kaggle should really have a separate prize for people like you who are this helpful in a competition.",
    "280331": "Thank you very much for your sharing, I learn a lot thanks to you (and I'm clearly not the only one!). I agree with @Bojan, you totally deserve a special prize.",
    "280390": "I did not use Andres code, nor did I do as well as many here.  I got to ~80% on plain vanilla CNNs.  What I did enjoy was the community and discussion wrt Andres work and the spirit he brought to the competition.  Well done!",
    "280605": "Hello and thank you so much for sharing the code before the deadline of the competition. It was my first participation in Kaggle and having available the pipeline of someone with  your experience taught me so much in so little time. I fully agree that you should also get some kind of prize :). Good luck!"
  },
  "source": "meta"
}