{
  "id": 45774,
  "title": "GTX 1050 Ti solution [43rd place]",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/45774",
  "author_name": "NighTurs",
  "post_date": "2017-12-15T18:14:42.840000",
  "votes": 23,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I was crazy enough to enter this competition with my single GTX 1050 ti and 8 GB of RAM. Soon I knew that it is not realistic to train even one epoch of full CNN with this setup. Instead I took ImageNet pre-trained ResNet50 and VGG16 models and saved features from layer before FC to disk. 6 days of pure running time only to do that. From that point I was training models of FC layers on top of these features. And I trained a lot of them experimenting with different FC setup, batch sizes, LR etc. Overall all of them were quite weak, best of them should be around 0.68. </p>\n\n<p>Some findings:</p>\n\n<ol>\n<li>Adding number of images in product and image index to features gave me 1% boost. </li>\n<li>Combining single image predictions by multiplication of probabilities is 0.0025 better then averaging.</li>\n<li>Although majority of my models predict single image, my best model (looking at ensemble weights) predicts product by maxpooling CNN features for each image (inspired by <a href=\"https://arxiv.org/abs/1505.00880\">Multi-view Convolutional Neural Networks for 3D Shape Recognition</a>)</li>\n</ol>\n\n<p>For each such model I predicted top ten categories, and then linearly ensembled them. Overall my final submission consists of 44 such weak models. Together they should score around 0.725 on LB (haven't checked it). Then I also took available <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">Heng CherKeng models</a> (Thanks a lot Heng!), my ensemble of them gives 0.71054 private LB. And again mixed top 10 predictions from both ensembles with weights 0.6 0.4.</p>\n\n<p>I had to struggle a lot with the size of this dataset, lack of RAM, lack of SSD space. It made me extra careful with how I organize my pipeline, and how I can restore anything I computed so far (in a case of unlucky rm). I have a huge make file with targets for every submission. Because every change is under git, and I avoided manual tweaks it should be possible (in theory) to reproduce everything I did. That is one of the most valuable things I got away from this competition. Uploaded code to git <a href=\"https://github.com/NighTurs/kaggle-cdiscount-image-classification\">https://github.com/NighTurs/kaggle-cdiscount-image-classification</a></p>",
  "messages": [
    {
      "id": 258213,
      "postDate": "2017-12-15T18:14:42.840Z",
      "content": "<p>I was crazy enough to enter this competition with my single GTX 1050 ti and 8 GB of RAM. Soon I knew that it is not realistic to train even one epoch of full CNN with this setup. Instead I took ImageNet pre-trained ResNet50 and VGG16 models and saved features from layer before FC to disk. 6 days of pure running time only to do that. From that point I was training models of FC layers on top of these features. And I trained a lot of them experimenting with different FC setup, batch sizes, LR etc. Overall all of them were quite weak, best of them should be around 0.68. </p>\n\n<p>Some findings:</p>\n\n<ol>\n<li>Adding number of images in product and image index to features gave me 1% boost. </li>\n<li>Combining single image predictions by multiplication of probabilities is 0.0025 better then averaging.</li>\n<li>Although majority of my models predict single image, my best model (looking at ensemble weights) predicts product by maxpooling CNN features for each image (inspired by <a href=\"https://arxiv.org/abs/1505.00880\">Multi-view Convolutional Neural Networks for 3D Shape Recognition</a>)</li>\n</ol>\n\n<p>For each such model I predicted top ten categories, and then linearly ensembled them. Overall my final submission consists of 44 such weak models. Together they should score around 0.725 on LB (haven't checked it). Then I also took available <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">Heng CherKeng models</a> (Thanks a lot Heng!), my ensemble of them gives 0.71054 private LB. And again mixed top 10 predictions from both ensembles with weights 0.6 0.4.</p>\n\n<p>I had to struggle a lot with the size of this dataset, lack of RAM, lack of SSD space. It made me extra careful with how I organize my pipeline, and how I can restore anything I computed so far (in a case of unlucky rm). I have a huge make file with targets for every submission. Because every change is under git, and I avoided manual tweaks it should be possible (in theory) to reproduce everything I did. That is one of the most valuable things I got away from this competition. Uploaded code to git <a href=\"https://github.com/NighTurs/kaggle-cdiscount-image-classification\">https://github.com/NighTurs/kaggle-cdiscount-image-classification</a></p>",
      "rawMarkdown": "I was crazy enough to enter this competition with my single GTX 1050 ti and 8 GB of RAM. Soon I knew that it is not realistic to train even one epoch of full CNN with this setup. Instead I took ImageNet pre-trained ResNet50 and VGG16 models and saved features from layer before FC to disk. 6 days of pure running time only to do that. From that point I was training models of FC layers on top of these features. And I trained a lot of them experimenting with different FC setup, batch sizes, LR etc. Overall all of them were quite weak, best of them should be around 0.68. \n\nSome findings:\n\n1. Adding number of images in product and image index to features gave me 1% boost. \n2. Combining single image predictions by multiplication of probabilities is 0.0025 better then averaging.\n3. Although majority of my models predict single image, my best model (looking at ensemble weights) predicts product by maxpooling CNN features for each image (inspired by [Multi-view Convolutional Neural Networks for 3D Shape Recognition](https://arxiv.org/abs/1505.00880))\n\nFor each such model I predicted top ten categories, and then linearly ensembled them. Overall my final submission consists of 44 such weak models. Together they should score around 0.725 on LB (haven't checked it). Then I also took available [Heng CherKeng models](https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021) (Thanks a lot Heng!), my ensemble of them gives 0.71054 private LB. And again mixed top 10 predictions from both ensembles with weights 0.6 0.4.\n\nI had to struggle a lot with the size of this dataset, lack of RAM, lack of SSD space. It made me extra careful with how I organize my pipeline, and how I can restore anything I computed so far (in a case of unlucky rm). I have a huge make file with targets for every submission. Because every change is under git, and I avoided manual tweaks it should be possible (in theory) to reproduce everything I did. That is one of the most valuable things I got away from this competition. Uploaded code to git https://github.com/NighTurs/kaggle-cdiscount-image-classification",
      "votes": 23
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "258213": "I was crazy enough to enter this competition with my single GTX 1050 ti and 8 GB of RAM. Soon I knew that it is not realistic to train even one epoch of full CNN with this setup. Instead I took ImageNet pre-trained ResNet50 and VGG16 models and saved features from layer before FC to disk. 6 days of pure running time only to do that. From that point I was training models of FC layers on top of these features. And I trained a lot of them experimenting with different FC setup, batch sizes, LR etc. Overall all of them were quite weak, best of them should be around 0.68. \n\nSome findings:\n\n1. Adding number of images in product and image index to features gave me 1% boost. \n2. Combining single image predictions by multiplication of probabilities is 0.0025 better then averaging.\n3. Although majority of my models predict single image, my best model (looking at ensemble weights) predicts product by maxpooling CNN features for each image (inspired by [Multi-view Convolutional Neural Networks for 3D Shape Recognition](https://arxiv.org/abs/1505.00880))\n\nFor each such model I predicted top ten categories, and then linearly ensembled them. Overall my final submission consists of 44 such weak models. Together they should score around 0.725 on LB (haven't checked it). Then I also took available [Heng CherKeng models](https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021) (Thanks a lot Heng!), my ensemble of them gives 0.71054 private LB. And again mixed top 10 predictions from both ensembles with weights 0.6 0.4.\n\nI had to struggle a lot with the size of this dataset, lack of RAM, lack of SSD space. It made me extra careful with how I organize my pipeline, and how I can restore anything I computed so far (in a case of unlucky rm). I have a huge make file with targets for every submission. Because every change is under git, and I avoided manual tweaks it should be possible (in theory) to reproduce everything I did. That is one of the most valuable things I got away from this competition. Uploaded code to git https://github.com/NighTurs/kaggle-cdiscount-image-classification"
  }
}