{
  "id": 49319,
  "title": "9th Place Solution",
  "url": "/competitions/sp-society-camera-model-identification/writeups/ods-ai-svm-punks-9th-place-solution",
  "author_name": "",
  "post_date": "2018-02-09T12:48:14.573Z",
  "votes": 30,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, congratulations to all of my teammates: <a href=\"https://www.kaggle.com/iglovikov\">ternaus</a> for getting Grand Master title; <a href=\"https://www.kaggle.com/arsenyinfo\">arsenyinfo</a> and <a href=\"https://www.kaggle.com/nesterov\">mephistopheies</a> for Master titles and <a href=\"https://www.kaggle.com/cortwave\">cortwave</a> for the first gold medal. Well-deserved guys!</p>\n\n<p>Here is a brief overview of our solution:</p>\n\n<ol>\n<li>Each of us created a solid single model earlier in the competition with the initial data and Gleb's data. There were MobileNets, VGGs, custom ResNets, DenseNets (about 6-7 models) with different crop sizes, augmentations and TTAs. We constructed majority voting out of these models and put test images with the maximum votes to pseudolabels.</li>\n<li>Got more data from flickr: about 40k images. They have been filtered by camera model, resolution, quality and any software changes. Constructed validation from Gleb's and new flickr images.</li>\n<li>The major idea of our models was to make training process iterative. In the first stage we've used train+flickr data (Gleb's and ours). Adding pseudolabels and resetting LR in the second stage. In particular, my own approach included: in the first stage, training VGG-16 on initial images + some flickr (overall 7k images). In stage 2, keep only pseudolabels and highly overfit to them reaching 100% accuracy on the train set. Such model gave 0.985 Private LB.</li>\n<li>Occasionally, we noticed that distribution of test photos is quite uniform in the test set and decided to force it to be exactly uniform. <a href=\"https://www.kaggle.com/nesterov\">mephistopheies</a> made some Analysis magic and gave a formula for such a normalization (it's better to ask him directly what he's done :) )</li>\n<li>Our final submission was a blend of 10 models (with and without stages) and subsequent classes balancing.</li>\n</ol>\n\n<p>Repo of our solution:\n<a href=\"https://github.com/cortwave/camera-model-identification\">https://github.com/cortwave/camera-model-identification</a></p>\n\n<p>TL; DR:</p>\n\n<p>Did Work:</p>\n\n<ul>\n<li>Pseudolabels</li>\n<li>External data and data cleaning</li>\n<li>Different crop sizes </li>\n<li>Picking the argmax of probabilities during TTA</li>\n<li>Balancing helped in the Public LB, but now we see that is has been overfitting</li>\n</ul>\n\n<p>Did Not Work:</p>\n\n<ul>\n<li>Averaging model checkpoints</li>\n<li>Training on denoised images</li>\n<li>Non-standard loss functions like hinge loss</li>\n<li>KNN on pseudolabel embeddings</li>\n<li>GAN and Siamese architectures</li>\n</ul>\n\n<p>P.S. Late submission of the blend on top-3 models scored 0.989 Private LB. Unfortunately, we haven't tried this one due to the lack of submissions..</p>",
  "messages": [
    {
      "id": "280104",
      "postDate": "02/09/2018 11:30:38",
      "content": "<p>First of all, congratulations to all of my teammates: <a href=\"https://www.kaggle.com/iglovikov\">ternaus</a> for getting Grand Master title; <a href=\"https://www.kaggle.com/arsenyinfo\">arsenyinfo</a> and <a href=\"https://www.kaggle.com/nesterov\">mephistopheies</a> for Master titles and <a href=\"https://www.kaggle.com/cortwave\">cortwave</a> for the first gold medal. Well-deserved guys!</p>\n\n<p>Here is a brief overview of our solution:</p>\n\n<ol>\n<li>Each of us created a solid single model earlier in the competition with the initial data and Gleb's data. There were MobileNets, VGGs, custom ResNets, DenseNets (about 6-7 models) with different crop sizes, augmentations and TTAs. We constructed majority voting out of these models and put test images with the maximum votes to pseudolabels.</li>\n<li>Got more data from flickr: about 40k images. They have been filtered by camera model, resolution, quality and any software changes. Constructed validation from Gleb's and new flickr images.</li>\n<li>The major idea of our models was to make training process iterative. In the first stage we've used train+flickr data (Gleb's and ours). Adding pseudolabels and resetting LR in the second stage. In particular, my own approach included: in the first stage, training VGG-16 on initial images + some flickr (overall 7k images). In stage 2, keep only pseudolabels and highly overfit to them reaching 100% accuracy on the train set. Such model gave 0.985 Private LB.</li>\n<li>Occasionally, we noticed that distribution of test photos is quite uniform in the test set and decided to force it to be exactly uniform. <a href=\"https://www.kaggle.com/nesterov\">mephistopheies</a> made some Analysis magic and gave a formula for such a normalization (it's better to ask him directly what he's done :) )</li>\n<li>Our final submission was a blend of 10 models (with and without stages) and subsequent classes balancing.</li>\n</ol>\n\n<p>Repo of our solution:\n<a href=\"https://github.com/cortwave/camera-model-identification\">https://github.com/cortwave/camera-model-identification</a></p>\n\n<p>TL; DR:</p>\n\n<p>Did Work:</p>\n\n<ul>\n<li>Pseudolabels</li>\n<li>External data and data cleaning</li>\n<li>Different crop sizes </li>\n<li>Picking the argmax of probabilities during TTA</li>\n<li>Balancing helped in the Public LB, but now we see that is has been overfitting</li>\n</ul>\n\n<p>Did Not Work:</p>\n\n<ul>\n<li>Averaging model checkpoints</li>\n<li>Training on denoised images</li>\n<li>Non-standard loss functions like hinge loss</li>\n<li>KNN on pseudolabel embeddings</li>\n<li>GAN and Siamese architectures</li>\n</ul>\n\n<p>P.S. Late submission of the blend on top-3 models scored 0.989 Private LB. Unfortunately, we haven't tried this one due to the lack of submissions..</p>",
      "rawMarkdown": "First of all, congratulations to all of my teammates: [ternaus][1] for getting Grand Master title; [arsenyinfo][2] and [mephistopheies][3] for Master titles and [cortwave][4] for the first gold medal. Well-deserved guys!\n\nHere is a brief overview of our solution:\n\n 1. Each of us created a solid single model earlier in the competition with the initial data and Gleb's data. There were MobileNets, VGGs, custom ResNets, DenseNets (about 6-7 models) with different crop sizes, augmentations and TTAs. We constructed majority voting out of these models and put test images with the maximum votes to pseudolabels.\n 2. Got more data from flickr: about 40k images. They have been filtered by camera model, resolution, quality and any software changes. Constructed validation from Gleb's and new flickr images.\n 3. The major idea of our models was to make training process iterative. In the first stage we've used train+flickr data (Gleb's and ours). Adding pseudolabels and resetting LR in the second stage. In particular, my own approach included: in the first stage, training VGG-16 on initial images + some flickr (overall 7k images). In stage 2, keep only pseudolabels and highly overfit to them reaching 100% accuracy on the train set. Such model gave 0.985 Private LB.\n 4. Occasionally, we noticed that distribution of test photos is quite uniform in the test set and decided to force it to be exactly uniform. [mephistopheies][5] made some Analysis magic and gave a formula for such a normalization (it's better to ask him directly what he's done :) )\n 5. Our final submission was a blend of 10 models (with and without stages) and subsequent classes balancing.\n\nRepo of our solution:\nhttps://github.com/cortwave/camera-model-identification\n\nTL; DR:\n\nDid Work:\n\n - Pseudolabels\n - External data and data cleaning\n - Different crop sizes \n - Picking the argmax of probabilities during TTA\n - Balancing helped in the Public LB, but now we see that is has been overfitting\n\nDid Not Work:\n\n - Averaging model checkpoints\n - Training on denoised images\n - Non-standard loss functions like hinge loss\n - KNN on pseudolabel embeddings\n - GAN and Siamese architectures\n\nP.S. Late submission of the blend on top-3 models scored 0.989 Private LB. Unfortunately, we haven't tried this one due to the lack of submissions..\n\n\n  [1]: https://www.kaggle.com/iglovikov\n  [2]: https://www.kaggle.com/arsenyinfo\n  [3]: https://www.kaggle.com/nesterov\n  [4]: https://www.kaggle.com/cortwave\n  [5]: https://www.kaggle.com/nesterov",
      "votes": null
    },
    {
      "id": "280434",
      "postDate": "02/10/2018 00:57:10",
      "content": "<p>Thanks for sharing. I'm going to try to reproduce the results as laid out in your step 3. So that I can plan for time, how many epochs (approx.) did you train the VGG-16 in each stage?</p>",
      "rawMarkdown": "Thanks for sharing. I'm going to try to reproduce the results as laid out in your step 3. So that I can plan for time, how many epochs (approx.) did you train the VGG-16 in each stage?",
      "votes": null
    },
    {
      "id": "280585",
      "postDate": "02/10/2018 11:03:13",
      "content": "<p>One epoch was 4 full passes through the training set. It took about 50 such epochs at the first stage and about 10 at the second.</p>",
      "rawMarkdown": "One epoch was 4 full passes through the training set. It took about 50 such epochs at the first stage and about 10 at the second.",
      "votes": null
    },
    {
      "id": "280622",
      "postDate": "02/10/2018 13:50:45",
      "content": "<p>Brilliant! Thanks so much for the detailed overview!</p>\n\n<p>Just a question. In stage 2, you say: </p>\n\n<blockquote>\n  <p>keep only pseudolabels and highly overfit to them reaching 100%\n  accuracy on the train set.</p>\n</blockquote>\n\n<p>If you reach 100% accuracy on your pseudo-labeled test set, doesn't it imply, that predictions of your model equal exactly the pseudo-labels? But if that is the case, then your model just replicates pseudo-labels. I suppose the idea is to improve on the pseudo-labels. </p>",
      "rawMarkdown": "Brilliant! Thanks so much for the detailed overview!\n\nJust a question. In stage 2, you say: \n\n&gt; keep only pseudolabels and highly overfit to them reaching 100%\n&gt; accuracy on the train set.\n\nIf you reach 100% accuracy on your pseudo-labeled test set, doesn't it imply, that predictions of your model equal exactly the pseudo-labels? But if that is the case, then your model just replicates pseudo-labels. I suppose the idea is to improve on the pseudo-labels.",
      "votes": null
    },
    {
      "id": "280691",
      "postDate": "02/10/2018 17:00:56",
      "content": "<p>As pseudo-labels we've been using only confident predictions from all the models. So, there were about 2400 test images. I was highly overfitting to these images to make my model replicate confident pseudolabels. And hopefully such model will learn how to classify the rest 240 test images.</p>",
      "rawMarkdown": "As pseudo-labels we've been using only confident predictions from all the models. So, there were about 2400 test images. I was highly overfitting to these images to make my model replicate confident pseudolabels. And hopefully such model will learn how to classify the rest 240 test images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 280434,
      "author_name": "albertoa",
      "author_url": "",
      "post_date": "02/10/2018 00:57:10",
      "content": "<p>Thanks for sharing. I'm going to try to reproduce the results as laid out in your step 3. So that I can plan for time, how many epochs (approx.) did you train the VGG-16 in each stage?</p>",
      "votes": null,
      "replies": [
        {
          "id": 280585,
          "author_name": "ybabakhin",
          "author_url": "",
          "post_date": "02/10/2018 11:03:13",
          "content": "<p>One epoch was 4 full passes through the training set. It took about 50 such epochs at the first stage and about 10 at the second.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280622,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "02/10/2018 13:50:45",
      "content": "<p>Brilliant! Thanks so much for the detailed overview!</p>\n\n<p>Just a question. In stage 2, you say: </p>\n\n<blockquote>\n  <p>keep only pseudolabels and highly overfit to them reaching 100%\n  accuracy on the train set.</p>\n</blockquote>\n\n<p>If you reach 100% accuracy on your pseudo-labeled test set, doesn't it imply, that predictions of your model equal exactly the pseudo-labels? But if that is the case, then your model just replicates pseudo-labels. I suppose the idea is to improve on the pseudo-labels. </p>",
      "votes": null,
      "replies": [
        {
          "id": 280691,
          "author_name": "ybabakhin",
          "author_url": "",
          "post_date": "02/10/2018 17:00:56",
          "content": "<p>As pseudo-labels we've been using only confident predictions from all the models. So, there were about 2400 test images. I was highly overfitting to these images to make my model replicate confident pseudolabels. And hopefully such model will learn how to classify the rest 240 test images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "280104": "First of all, congratulations to all of my teammates: [ternaus][1] for getting Grand Master title; [arsenyinfo][2] and [mephistopheies][3] for Master titles and [cortwave][4] for the first gold medal. Well-deserved guys!\n\nHere is a brief overview of our solution:\n\n 1. Each of us created a solid single model earlier in the competition with the initial data and Gleb's data. There were MobileNets, VGGs, custom ResNets, DenseNets (about 6-7 models) with different crop sizes, augmentations and TTAs. We constructed majority voting out of these models and put test images with the maximum votes to pseudolabels.\n 2. Got more data from flickr: about 40k images. They have been filtered by camera model, resolution, quality and any software changes. Constructed validation from Gleb's and new flickr images.\n 3. The major idea of our models was to make training process iterative. In the first stage we've used train+flickr data (Gleb's and ours). Adding pseudolabels and resetting LR in the second stage. In particular, my own approach included: in the first stage, training VGG-16 on initial images + some flickr (overall 7k images). In stage 2, keep only pseudolabels and highly overfit to them reaching 100% accuracy on the train set. Such model gave 0.985 Private LB.\n 4. Occasionally, we noticed that distribution of test photos is quite uniform in the test set and decided to force it to be exactly uniform. [mephistopheies][5] made some Analysis magic and gave a formula for such a normalization (it's better to ask him directly what he's done :) )\n 5. Our final submission was a blend of 10 models (with and without stages) and subsequent classes balancing.\n\nRepo of our solution:\nhttps://github.com/cortwave/camera-model-identification\n\nTL; DR:\n\nDid Work:\n\n - Pseudolabels\n - External data and data cleaning\n - Different crop sizes \n - Picking the argmax of probabilities during TTA\n - Balancing helped in the Public LB, but now we see that is has been overfitting\n\nDid Not Work:\n\n - Averaging model checkpoints\n - Training on denoised images\n - Non-standard loss functions like hinge loss\n - KNN on pseudolabel embeddings\n - GAN and Siamese architectures\n\nP.S. Late submission of the blend on top-3 models scored 0.989 Private LB. Unfortunately, we haven't tried this one due to the lack of submissions..\n\n\n  [1]: https://www.kaggle.com/iglovikov\n  [2]: https://www.kaggle.com/arsenyinfo\n  [3]: https://www.kaggle.com/nesterov\n  [4]: https://www.kaggle.com/cortwave\n  [5]: https://www.kaggle.com/nesterov",
    "280434": "Thanks for sharing. I'm going to try to reproduce the results as laid out in your step 3. So that I can plan for time, how many epochs (approx.) did you train the VGG-16 in each stage?",
    "280585": "One epoch was 4 full passes through the training set. It took about 50 such epochs at the first stage and about 10 at the second.",
    "280622": "Brilliant! Thanks so much for the detailed overview!\n\nJust a question. In stage 2, you say: \n\n&gt; keep only pseudolabels and highly overfit to them reaching 100%\n&gt; accuracy on the train set.\n\nIf you reach 100% accuracy on your pseudo-labeled test set, doesn't it imply, that predictions of your model equal exactly the pseudo-labels? But if that is the case, then your model just replicates pseudo-labels. I suppose the idea is to improve on the pseudo-labels.",
    "280691": "As pseudo-labels we've been using only confident predictions from all the models. So, there were about 2400 test images. I was highly overfitting to these images to make my model replicate confident pseudolabels. And hopefully such model will learn how to classify the rest 240 test images."
  },
  "source": "meta"
}