{
  "id": 45850,
  "title": "[0.776] Single model solution",
  "url": "/competitions/cdiscount-image-classification-challenge/writeups/azat-davletshin-0-776-single-model-solution",
  "author_name": "",
  "post_date": "2017-12-16T14:58:40.077Z",
  "votes": 29,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I managed to get 0.77603 on Private Set with single model (my final result is 0.77883 - it's an ensemble of two models). Here's what I did:</p>\n\n<p>1) First I finetuned the pre-trained SE-ResNet-50 for about 7 epochs: <br>\n- I took the weights from here: <a href=\"https://github.com/hujie-frank/SENet\">https://github.com/hujie-frank/SENet</a> (I've used Caffe) <br>\n- Finetuned the network for 7 epochs using Nesterov Momemtum SGD (momentum = 0.9, batch_size = 256) with the following training schedule: <br>\n- epochs 1..5: lr = 0.001 <br>\n - epoch 6: lr = 0.0001 <br>\n - epoch 7: lr = 0.00001 <br>\n- The input image size was 161x161. I've used random crops + random horizontal flips for augmentation. <br>\n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.7133 <br>\n- I tried to finetune network with the exact same protocol, but without random crops (resized images to 161x161 and used only random horizontal flips). It turned out almost 0.01 better than with crops: 0.721. So I decided that I would not use augmentation, except random horizontal flip. <br>\n- I decided to try KNN on top of the features of SE-ResNet-50: <br>\n- I've generated features for the train, test and validation (formula for averaging flip with noflip:   features = l2_normalize(l2_normalize(flip_features) + l2_normalize(noflip_features))). The matrix for the training set weighed about 96 GB and I used numpy.memmap, because all this did not fit in RAM. I saved the Matrix with features on the SSD. <br>\n- For each picture from the test, I found 5 closest (by cosine) pictures from the train for each class. I stored the search result on the HDD via memmap (Matrix size 5270 x N_TEST_IMAGES x 5 - weighed about 300 GB for the test.). <br>\n- Next, I calculated the score of each class for each picture: scores[picture, label] = numpy.power((1 + search_scores[label, picture, :]) / 2, 35).sum() (the sum of the scores from 5 closest pictures). The formula was tuned on validation. <br>\n- For each image, I stored scores from the top10 classes (ie equated the scores of other classes to 0). Score of the product was the sum of scores of pictures. After these manipulations I've got 0.752 on LB. Thus KNN worked at 0.03 better than softmax. <br>\n2) I decided to use SE-ResNet-101 for the final model (I had to convert weights to PyTorch, because Caffe consumes too much memory). <br>\n- Finetuned the network for 15 epochs using Nesterov Momemtum (momentum = 0.9, batch_size = 512) with the following training schedule: <br>\n- epochs 1..9: lr = 0.01 <br>\n- epochs 10..15: lr = 0.001 <br>\n- The size of the input is 161x161. Of the augmentations, only Random Horizontal Flip was used. In addition, I added dropout with p = 0.2 after global average pooling. <br>\n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.746 <br>\n- KNN (according to the scheme described above) gave 0.765 <br>\n- I tried to average the classifier predictions with KNN. To do this, I multiply the KNN scores by 300 and took the softmax of the top10 classes and summed it with the classifier predictions. This gave 0.776 on Private Set. <br>\n3) For the final ensemble, I recalculated KNN combining the training set with the validation set and averaged the classifier predictions with the SE-ResNet-50. This gave 0.77883. <br>\nP.S My hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD</p>",
  "messages": [
    {
      "id": "258594",
      "postDate": "12/16/2017 14:42:43",
      "content": "<p>I managed to get 0.77603 on Private Set with single model (my final result is 0.77883 - it's an ensemble of two models). Here's what I did:</p>\n\n<p>1) First I finetuned the pre-trained SE-ResNet-50 for about 7 epochs: <br>\n- I took the weights from here: <a href=\"https://github.com/hujie-frank/SENet\">https://github.com/hujie-frank/SENet</a> (I've used Caffe) <br>\n- Finetuned the network for 7 epochs using Nesterov Momemtum SGD (momentum = 0.9, batch_size = 256) with the following training schedule: <br>\n- epochs 1..5: lr = 0.001 <br>\n - epoch 6: lr = 0.0001 <br>\n - epoch 7: lr = 0.00001 <br>\n- The input image size was 161x161. I've used random crops + random horizontal flips for augmentation. <br>\n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.7133 <br>\n- I tried to finetune network with the exact same protocol, but without random crops (resized images to 161x161 and used only random horizontal flips). It turned out almost 0.01 better than with crops: 0.721. So I decided that I would not use augmentation, except random horizontal flip. <br>\n- I decided to try KNN on top of the features of SE-ResNet-50: <br>\n- I've generated features for the train, test and validation (formula for averaging flip with noflip:   features = l2_normalize(l2_normalize(flip_features) + l2_normalize(noflip_features))). The matrix for the training set weighed about 96 GB and I used numpy.memmap, because all this did not fit in RAM. I saved the Matrix with features on the SSD. <br>\n- For each picture from the test, I found 5 closest (by cosine) pictures from the train for each class. I stored the search result on the HDD via memmap (Matrix size 5270 x N_TEST_IMAGES x 5 - weighed about 300 GB for the test.). <br>\n- Next, I calculated the score of each class for each picture: scores[picture, label] = numpy.power((1 + search_scores[label, picture, :]) / 2, 35).sum() (the sum of the scores from 5 closest pictures). The formula was tuned on validation. <br>\n- For each image, I stored scores from the top10 classes (ie equated the scores of other classes to 0). Score of the product was the sum of scores of pictures. After these manipulations I've got 0.752 on LB. Thus KNN worked at 0.03 better than softmax. <br>\n2) I decided to use SE-ResNet-101 for the final model (I had to convert weights to PyTorch, because Caffe consumes too much memory). <br>\n- Finetuned the network for 15 epochs using Nesterov Momemtum (momentum = 0.9, batch_size = 512) with the following training schedule: <br>\n- epochs 1..9: lr = 0.01 <br>\n- epochs 10..15: lr = 0.001 <br>\n- The size of the input is 161x161. Of the augmentations, only Random Horizontal Flip was used. In addition, I added dropout with p = 0.2 after global average pooling. <br>\n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.746 <br>\n- KNN (according to the scheme described above) gave 0.765 <br>\n- I tried to average the classifier predictions with KNN. To do this, I multiply the KNN scores by 300 and took the softmax of the top10 classes and summed it with the classifier predictions. This gave 0.776 on Private Set. <br>\n3) For the final ensemble, I recalculated KNN combining the training set with the validation set and averaged the classifier predictions with the SE-ResNet-50. This gave 0.77883. <br>\nP.S My hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD</p>",
      "rawMarkdown": "I managed to get 0.77603 on Private Set with single model (my final result is 0.77883 - it's an ensemble of two models). Here's what I did:\n\n1) First I finetuned the pre-trained SE-ResNet-50 for about 7 epochs:  \n- I took the weights from here: https://github.com/hujie-frank/SENet (I've used Caffe)  \n- Finetuned the network for 7 epochs using Nesterov Momemtum SGD (momentum = 0.9, batch_size = 256) with the following training schedule:  \n- epochs 1..5: lr = 0.001  \n - epoch 6: lr = 0.0001  \n - epoch 7: lr = 0.00001  \n- The input image size was 161x161. I've used random crops + random horizontal flips for augmentation.  \n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.7133  \n- I tried to finetune network with the exact same protocol, but without random crops (resized images to 161x161 and used only random horizontal flips). It turned out almost 0.01 better than with crops: 0.721. So I decided that I would not use augmentation, except random horizontal flip.  \n- I decided to try KNN on top of the features of SE-ResNet-50:  \n- I've generated features for the train, test and validation (formula for averaging flip with noflip:   features = l2_normalize(l2_normalize(flip_features) + l2_normalize(noflip_features))). The matrix for the training set weighed about 96 GB and I used numpy.memmap, because all this did not fit in RAM. I saved the Matrix with features on the SSD.  \n- For each picture from the test, I found 5 closest (by cosine) pictures from the train for each class. I stored the search result on the HDD via memmap (Matrix size 5270 x N_TEST_IMAGES x 5 - weighed about 300 GB for the test.).  \n- Next, I calculated the score of each class for each picture: scores[picture, label] = numpy.power((1 + search_scores[label, picture, :]) / 2, 35).sum() (the sum of the scores from 5 closest pictures). The formula was tuned on validation.  \n- For each image, I stored scores from the top10 classes (ie equated the scores of other classes to 0). Score of the product was the sum of scores of pictures. After these manipulations I've got 0.752 on LB. Thus KNN worked at 0.03 better than softmax.  \n2) I decided to use SE-ResNet-101 for the final model (I had to convert weights to PyTorch, because Caffe consumes too much memory).  \n- Finetuned the network for 15 epochs using Nesterov Momemtum (momentum = 0.9, batch_size = 512) with the following training schedule:  \n- epochs 1..9: lr = 0.01  \n- epochs 10..15: lr = 0.001  \n- The size of the input is 161x161. Of the augmentations, only Random Horizontal Flip was used. In addition, I added dropout with p = 0.2 after global average pooling.  \n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.746  \n- KNN (according to the scheme described above) gave 0.765  \n- I tried to average the classifier predictions with KNN. To do this, I multiply the KNN scores by 300 and took the softmax of the top10 classes and summed it with the classifier predictions. This gave 0.776 on Private Set.  \n3) For the final ensemble, I recalculated KNN combining the training set with the validation set and averaged the classifier predictions with the SE-ResNet-50. This gave 0.77883.  \nP.S My hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD",
      "votes": null
    },
    {
      "id": "258639",
      "postDate": "12/16/2017 17:12:47",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": null
    },
    {
      "id": "258667",
      "postDate": "12/16/2017 18:48:26",
      "content": "<p>Impressive! Thanks for sharing!</p>",
      "rawMarkdown": "Impressive! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "301513",
      "postDate": "03/22/2018 21:20:04",
      "content": "<p>Congratulations and thanks for sharing. Your solution is very interesting to me and I am trying to reproduce what you did but I have some difficulties. Would you please share your code?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing. Your solution is very interesting to me and I am trying to reproduce what you did but I have some difficulties. Would you please share your code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 258639,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "12/16/2017 17:12:47",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258667,
      "author_name": "andreanobile",
      "author_url": "",
      "post_date": "12/16/2017 18:48:26",
      "content": "<p>Impressive! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 301513,
      "author_name": "ehsanfathi",
      "author_url": "",
      "post_date": "03/22/2018 21:20:04",
      "content": "<p>Congratulations and thanks for sharing. Your solution is very interesting to me and I am trying to reproduce what you did but I have some difficulties. Would you please share your code?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "258594": "I managed to get 0.77603 on Private Set with single model (my final result is 0.77883 - it's an ensemble of two models). Here's what I did:\n\n1) First I finetuned the pre-trained SE-ResNet-50 for about 7 epochs:  \n- I took the weights from here: https://github.com/hujie-frank/SENet (I've used Caffe)  \n- Finetuned the network for 7 epochs using Nesterov Momemtum SGD (momentum = 0.9, batch_size = 256) with the following training schedule:  \n- epochs 1..5: lr = 0.001  \n - epoch 6: lr = 0.0001  \n - epoch 7: lr = 0.00001  \n- The input image size was 161x161. I've used random crops + random horizontal flips for augmentation.  \n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.7133  \n- I tried to finetune network with the exact same protocol, but without random crops (resized images to 161x161 and used only random horizontal flips). It turned out almost 0.01 better than with crops: 0.721. So I decided that I would not use augmentation, except random horizontal flip.  \n- I decided to try KNN on top of the features of SE-ResNet-50:  \n- I've generated features for the train, test and validation (formula for averaging flip with noflip:   features = l2_normalize(l2_normalize(flip_features) + l2_normalize(noflip_features))). The matrix for the training set weighed about 96 GB and I used numpy.memmap, because all this did not fit in RAM. I saved the Matrix with features on the SSD.  \n- For each picture from the test, I found 5 closest (by cosine) pictures from the train for each class. I stored the search result on the HDD via memmap (Matrix size 5270 x N_TEST_IMAGES x 5 - weighed about 300 GB for the test.).  \n- Next, I calculated the score of each class for each picture: scores[picture, label] = numpy.power((1 + search_scores[label, picture, :]) / 2, 35).sum() (the sum of the scores from 5 closest pictures). The formula was tuned on validation.  \n- For each image, I stored scores from the top10 classes (ie equated the scores of other classes to 0). Score of the product was the sum of scores of pictures. After these manipulations I've got 0.752 on LB. Thus KNN worked at 0.03 better than softmax.  \n2) I decided to use SE-ResNet-101 for the final model (I had to convert weights to PyTorch, because Caffe consumes too much memory).  \n- Finetuned the network for 15 epochs using Nesterov Momemtum (momentum = 0.9, batch_size = 512) with the following training schedule:  \n- epochs 1..9: lr = 0.01  \n- epochs 10..15: lr = 0.001  \n- The size of the input is 161x161. Of the augmentations, only Random Horizontal Flip was used. In addition, I added dropout with p = 0.2 after global average pooling.  \n- I've used horizontal flips for test time augmentation. After averaging the pictures of the products I've got 0.746  \n- KNN (according to the scheme described above) gave 0.765  \n- I tried to average the classifier predictions with KNN. To do this, I multiply the KNN scores by 300 and took the softmax of the top10 classes and summed it with the classifier predictions. This gave 0.776 on Private Set.  \n3) For the final ensemble, I recalculated KNN combining the training set with the validation set and averaged the classifier predictions with the SE-ResNet-50. This gave 0.77883.  \nP.S My hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD",
    "258639": "Congratulations and thanks for sharing!",
    "258667": "Impressive! Thanks for sharing!",
    "301513": "Congratulations and thanks for sharing. Your solution is very interesting to me and I am trying to reproduce what you did but I have some difficulties. Would you please share your code?"
  },
  "source": "meta"
}