{
  "id": 45798,
  "title": "12th place solution",
  "url": "/competitions/cdiscount-image-classification-challenge/writeups/vladimir-iglovikov-12th-place-solution",
  "author_name": "",
  "post_date": "2017-12-16T00:04:09.124183600Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p><img src=\"https://habrastorage.org/webt/0f/q7/q9/0fq7q9l_r53hhzvjgoxtknwsm2g.jpeg\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>(each bar represents an image)</p>\n\n<p>There are duplicate images in train and test and a set of images is shared between them (data leak). The easiest way to find them is to calculate md5 hash for each image. </p>\n\n<p><a href=\"https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_train.py\">https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_train.py</a>\n<a href=\"https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_test.py\">https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_test.py</a></p>\n\n<p>This allows:</p>\n\n<ol>\n<li>Significantly decrease the size of the train and test speeding up training and inference: 12m =&gt; 7.5m, 3m =&gt; 2m</li>\n<li>Get class labels from train for images in test</li>\n</ol>\n\n<p>After this, I trained Resnet 50, 101 and 152, on 160x160 crops, dropping learning rate on the plateau.\nAt the first epoch, all layers except last are frozen.</p>\n\n<p>For each model test time augmentation + geometric mean=&gt; 0.75 on LB</p>\n\n<p>Geometric mean of the previous step =&gt; 0.77 on LB</p>",
  "messages": [
    {
      "id": "258330",
      "postDate": "12/16/2017 00:04:09",
      "content": "<p><img src=\"https://habrastorage.org/webt/0f/q7/q9/0fq7q9l_r53hhzvjgoxtknwsm2g.jpeg\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>(each bar represents an image)</p>\n\n<p>There are duplicate images in train and test and a set of images is shared between them (data leak). The easiest way to find them is to calculate md5 hash for each image. </p>\n\n<p><a href=\"https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_train.py\">https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_train.py</a>\n<a href=\"https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_test.py\">https://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_test.py</a></p>\n\n<p>This allows:</p>\n\n<ol>\n<li>Significantly decrease the size of the train and test speeding up training and inference: 12m =&gt; 7.5m, 3m =&gt; 2m</li>\n<li>Get class labels from train for images in test</li>\n</ol>\n\n<p>After this, I trained Resnet 50, 101 and 152, on 160x160 crops, dropping learning rate on the plateau.\nAt the first epoch, all layers except last are frozen.</p>\n\n<p>For each model test time augmentation + geometric mean=&gt; 0.75 on LB</p>\n\n<p>Geometric mean of the previous step =&gt; 0.77 on LB</p>",
      "rawMarkdown": "![enter image description here][1]\n\n(each bar represents an image)\n\nThere are duplicate images in train and test and a set of images is shared between them (data leak). The easiest way to find them is to calculate md5 hash for each image. \n\nhttps://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_train.py\nhttps://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_test.py\n\nThis allows:\n\n 1. Significantly decrease the size of the train and test speeding up training and inference: 12m =&gt; 7.5m, 3m =&gt; 2m\n 2. Get class labels from train for images in test\n\nAfter this, I trained Resnet 50, 101 and 152, on 160x160 crops, dropping learning rate on the plateau.\nAt the first epoch, all layers except last are frozen.\n\nFor each model test time augmentation + geometric mean=&gt; 0.75 on LB\n\nGeometric mean of the previous step =&gt; 0.77 on LB\n\n  [1]: https://habrastorage.org/webt/0f/q7/q9/0fq7q9l_r53hhzvjgoxtknwsm2g.jpeg",
      "votes": null
    },
    {
      "id": "258336",
      "postDate": "12/16/2017 00:12:50",
      "content": "<p>Good job finding useful leaks! I'm too lazy to even look at the images...</p>\n\n<p>Do you have any results without using the leaks?</p>",
      "rawMarkdown": "Good job finding useful leaks! I'm too lazy to even look at the images...\n\nDo you have any results without using the leaks?",
      "votes": null
    },
    {
      "id": "258342",
      "postDate": "12/16/2017 00:23:30",
      "content": "<ol>\n<li><p>The leak is not really useful here for improving accuracy but speeds up training and inference. 0.74 at the LB, that I reported in one of the threads, was achieved before I figured out that this leak exists. </p></li>\n<li><p>Training Resnet 101 with dropped duplicates =&gt; 0.75 at the LB. Traning Resnet without dropped duplicates =&gt; same 0.75 at the LB.</p></li>\n<li><p>If admins will clean up test and remove those images that could be found in test, score at the LB for everyone will drop by 10%.</p></li>\n</ol>",
      "rawMarkdown": "1. The leak is not really useful here for improving accuracy but speeds up training and inference. 0.74 at the LB, that I reported in one of the threads, was achieved before I figured out that this leak exists. \n\n2. Training Resnet 101 with dropped duplicates =&gt; 0.75 at the LB. Traning Resnet without dropped duplicates =&gt; same 0.75 at the LB.\n\n3. If admins will clean up test and remove those images that could be found in test, score at the LB for everyone will drop by 10%.",
      "votes": null
    },
    {
      "id": "258349",
      "postDate": "12/16/2017 00:37:45",
      "content": "<p>Now I know why the score was so good.</p>\n\n<p>Until now I thought that there was some kind of watermark on some classes that the model was learning. It was clear that, whatever it was, the score was too good to be true and a model that performs well on this dataset  doesn't necessarily generalize.</p>",
      "rawMarkdown": "Now I know why the score was so good.\n\nUntil now I thought that there was some kind of watermark on some classes that the model was learning. It was clear that, whatever it was, the score was too good to be true and a model that performs well on this dataset  doesn't necessarily generalize.",
      "votes": null
    },
    {
      "id": "258509",
      "postDate": "12/16/2017 09:48:06",
      "content": "<p>Congrats Vladimir, Extremely elegant solution with impressive score. This is a beauty.</p>",
      "rawMarkdown": "Congrats Vladimir, Extremely elegant solution with impressive score. This is a beauty.",
      "votes": null
    },
    {
      "id": "258636",
      "postDate": "12/16/2017 17:09:50",
      "content": "<p>Thanks for sharing your solution. </p>\n\n<p>There were lots of duplicate images that exist in multiple classes such as placeholder. If you found that, how did you attack this problem? FYI, we replace their predictions with their most frequent class in train data.\nAnd why did you use geometric mean rather than arithmetic?</p>",
      "rawMarkdown": "Thanks for sharing your solution. \n\nThere were lots of duplicate images that exist in multiple classes such as placeholder. If you found that, how did you attack this problem? FYI, we replace their predictions with their most frequent class in train data.\nAnd why did you use geometric mean rather than arithmetic?",
      "votes": null
    },
    {
      "id": "258640",
      "postDate": "12/16/2017 17:13:46",
      "content": "<p>You can find one example here: </p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45850\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45850</a></p>",
      "rawMarkdown": "You can find one example here: \n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45850",
      "votes": null
    },
    {
      "id": "258642",
      "postDate": "12/16/2017 17:23:28",
      "content": "<p>I did not really invest time into this problem (first months in a new job, conferences, wrapping up Carvana problem (releasing code at github and writing a blog post for Kaggle blog), etc..), so my approach to this problem was very straightforward. If some image presented in classes </p>\n\n<ul>\n<li>A - 10 times </li>\n<li>B - 20 times </li>\n<li>C - 30 times</li>\n</ul>\n\n<p>I just created sparse vector with</p>\n\n<ul>\n<li>10 / (10 + 20 + 30)</li>\n<li>20 / (10 + 20 + 30) </li>\n<li>30 / (10 + 20 + 30) </li>\n</ul>\n\n<p>in corresponding spots.</p>\n\n<p>I believe that better solution would be either to remove images that presented in different classes in the train or predict them using networks in a hope that networks will generate better embedding.</p>\n\n<p>I was thinking about dropping placeholders, but there were several classes that contained only placeholder images.</p>\n\n<p>Typically, when I work with probabilities geometric mean shows better score. For one of the models I compared arithmetic and geometric means, results were very similar but geometric was still a bit better, so I stick with it.</p>",
      "rawMarkdown": "I did not really invest time into this problem (first months in a new job, conferences, wrapping up Carvana problem (releasing code at github and writing a blog post for Kaggle blog), etc..), so my approach to this problem was very straightforward. If some image presented in classes \n \n* A - 10 times \n* B - 20 times \n* C - 30 times\n\nI just created sparse vector with\n\n * 10 / (10 + 20 + 30)\n * 20 / (10 + 20 + 30) \n * 30 / (10 + 20 + 30) \n\nin corresponding spots.\n\nI believe that better solution would be either to remove images that presented in different classes in the train or predict them using networks in a hope that networks will generate better embedding.\n\nI was thinking about dropping placeholders, but there were several classes that contained only placeholder images.\n\nTypically, when I work with probabilities geometric mean shows better score. For one of the models I compared arithmetic and geometric means, results were very similar but geometric was still a bit better, so I stick with it.",
      "votes": null
    },
    {
      "id": "258656",
      "postDate": "12/16/2017 18:01:46",
      "content": "<p>Oh, you wrote the solution of Carvana contest! I will read it definitely.</p>\n\n<p>Interestingly, the predicted probabilities of placeholder by my nets were in proportion to their frequency of each class (like your way).</p>\n\n<blockquote>\n  <p>there were several classes that contained only placeholder images.</p>\n</blockquote>\n\n<p>I didn't notice that. But I'm not surprised this matter since there were amounts of placeholder images.</p>",
      "rawMarkdown": "Oh, you wrote the solution of Carvana contest! I will read it definitely.\n\nInterestingly, the predicted probabilities of placeholder by my nets were in proportion to their frequency of each class (like your way).\n\n&gt; there were several classes that contained only placeholder images.\n\nI didn't notice that. But I'm not surprised this matter since there were amounts of placeholder images.",
      "votes": null
    },
    {
      "id": "258766",
      "postDate": "12/17/2017 01:35:33",
      "content": "<p>Looks great!!!!!</p>",
      "rawMarkdown": "Looks great!!!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 258336,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "12/16/2017 00:12:50",
      "content": "<p>Good job finding useful leaks! I'm too lazy to even look at the images...</p>\n\n<p>Do you have any results without using the leaks?</p>",
      "votes": null,
      "replies": [
        {
          "id": 258342,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "12/16/2017 00:23:30",
          "content": "<ol>\n<li><p>The leak is not really useful here for improving accuracy but speeds up training and inference. 0.74 at the LB, that I reported in one of the threads, was achieved before I figured out that this leak exists. </p></li>\n<li><p>Training Resnet 101 with dropped duplicates =&gt; 0.75 at the LB. Traning Resnet without dropped duplicates =&gt; same 0.75 at the LB.</p></li>\n<li><p>If admins will clean up test and remove those images that could be found in test, score at the LB for everyone will drop by 10%.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258349,
          "author_name": "radustoicescu",
          "author_url": "",
          "post_date": "12/16/2017 00:37:45",
          "content": "<p>Now I know why the score was so good.</p>\n\n<p>Until now I thought that there was some kind of watermark on some classes that the model was learning. It was clear that, whatever it was, the score was too good to be true and a model that performs well on this dataset  doesn't necessarily generalize.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258509,
      "author_name": "andreanobile",
      "author_url": "",
      "post_date": "12/16/2017 09:48:06",
      "content": "<p>Congrats Vladimir, Extremely elegant solution with impressive score. This is a beauty.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258636,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "12/16/2017 17:09:50",
      "content": "<p>Thanks for sharing your solution. </p>\n\n<p>There were lots of duplicate images that exist in multiple classes such as placeholder. If you found that, how did you attack this problem? FYI, we replace their predictions with their most frequent class in train data.\nAnd why did you use geometric mean rather than arithmetic?</p>",
      "votes": null,
      "replies": [
        {
          "id": 258640,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "12/16/2017 17:13:46",
          "content": "<p>You can find one example here: </p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45850\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45850</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258642,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "12/16/2017 17:23:28",
          "content": "<p>I did not really invest time into this problem (first months in a new job, conferences, wrapping up Carvana problem (releasing code at github and writing a blog post for Kaggle blog), etc..), so my approach to this problem was very straightforward. If some image presented in classes </p>\n\n<ul>\n<li>A - 10 times </li>\n<li>B - 20 times </li>\n<li>C - 30 times</li>\n</ul>\n\n<p>I just created sparse vector with</p>\n\n<ul>\n<li>10 / (10 + 20 + 30)</li>\n<li>20 / (10 + 20 + 30) </li>\n<li>30 / (10 + 20 + 30) </li>\n</ul>\n\n<p>in corresponding spots.</p>\n\n<p>I believe that better solution would be either to remove images that presented in different classes in the train or predict them using networks in a hope that networks will generate better embedding.</p>\n\n<p>I was thinking about dropping placeholders, but there were several classes that contained only placeholder images.</p>\n\n<p>Typically, when I work with probabilities geometric mean shows better score. For one of the models I compared arithmetic and geometric means, results were very similar but geometric was still a bit better, so I stick with it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258656,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "12/16/2017 18:01:46",
          "content": "<p>Oh, you wrote the solution of Carvana contest! I will read it definitely.</p>\n\n<p>Interestingly, the predicted probabilities of placeholder by my nets were in proportion to their frequency of each class (like your way).</p>\n\n<blockquote>\n  <p>there were several classes that contained only placeholder images.</p>\n</blockquote>\n\n<p>I didn't notice that. But I'm not surprised this matter since there were amounts of placeholder images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258766,
      "author_name": "raqueldeneige",
      "author_url": "",
      "post_date": "12/17/2017 01:35:33",
      "content": "<p>Looks great!!!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "258330": "![enter image description here][1]\n\n(each bar represents an image)\n\nThere are duplicate images in train and test and a set of images is shared between them (data leak). The easiest way to find them is to calculate md5 hash for each image. \n\nhttps://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_train.py\nhttps://github.com/ternaus/kaggle_cdiscount/blob/master/find_md5_test.py\n\nThis allows:\n\n 1. Significantly decrease the size of the train and test speeding up training and inference: 12m =&gt; 7.5m, 3m =&gt; 2m\n 2. Get class labels from train for images in test\n\nAfter this, I trained Resnet 50, 101 and 152, on 160x160 crops, dropping learning rate on the plateau.\nAt the first epoch, all layers except last are frozen.\n\nFor each model test time augmentation + geometric mean=&gt; 0.75 on LB\n\nGeometric mean of the previous step =&gt; 0.77 on LB\n\n  [1]: https://habrastorage.org/webt/0f/q7/q9/0fq7q9l_r53hhzvjgoxtknwsm2g.jpeg",
    "258336": "Good job finding useful leaks! I'm too lazy to even look at the images...\n\nDo you have any results without using the leaks?",
    "258342": "1. The leak is not really useful here for improving accuracy but speeds up training and inference. 0.74 at the LB, that I reported in one of the threads, was achieved before I figured out that this leak exists. \n\n2. Training Resnet 101 with dropped duplicates =&gt; 0.75 at the LB. Traning Resnet without dropped duplicates =&gt; same 0.75 at the LB.\n\n3. If admins will clean up test and remove those images that could be found in test, score at the LB for everyone will drop by 10%.",
    "258349": "Now I know why the score was so good.\n\nUntil now I thought that there was some kind of watermark on some classes that the model was learning. It was clear that, whatever it was, the score was too good to be true and a model that performs well on this dataset  doesn't necessarily generalize.",
    "258509": "Congrats Vladimir, Extremely elegant solution with impressive score. This is a beauty.",
    "258636": "Thanks for sharing your solution. \n\nThere were lots of duplicate images that exist in multiple classes such as placeholder. If you found that, how did you attack this problem? FYI, we replace their predictions with their most frequent class in train data.\nAnd why did you use geometric mean rather than arithmetic?",
    "258640": "You can find one example here: \n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45850",
    "258642": "I did not really invest time into this problem (first months in a new job, conferences, wrapping up Carvana problem (releasing code at github and writing a blog post for Kaggle blog), etc..), so my approach to this problem was very straightforward. If some image presented in classes \n \n* A - 10 times \n* B - 20 times \n* C - 30 times\n\nI just created sparse vector with\n\n * 10 / (10 + 20 + 30)\n * 20 / (10 + 20 + 30) \n * 30 / (10 + 20 + 30) \n\nin corresponding spots.\n\nI believe that better solution would be either to remove images that presented in different classes in the train or predict them using networks in a hope that networks will generate better embedding.\n\nI was thinking about dropping placeholders, but there were several classes that contained only placeholder images.\n\nTypically, when I work with probabilities geometric mean shows better score. For one of the models I compared arithmetic and geometric means, results were very similar but geometric was still a bit better, so I stick with it.",
    "258656": "Oh, you wrote the solution of Carvana contest! I will read it definitely.\n\nInterestingly, the predicted probabilities of placeholder by my nets were in proportion to their frequency of each class (like your way).\n\n&gt; there were several classes that contained only placeholder images.\n\nI didn't notice that. But I'm not surprised this matter since there were amounts of placeholder images.",
    "258766": "Looks great!!!!!"
  },
  "source": "meta"
}