{
  "id": 45733,
  "title": "Our solution [5th place]",
  "url": "/competitions/cdiscount-image-classification-challenge/writeups/10011000-our-solution-5th-place",
  "author_name": "",
  "post_date": "2017-12-15T21:02:07.147Z",
  "votes": 60,
  "comment_count": 17,
  "views": 0,
  "content": "<p>The key idea was to train several different convolutional neural networks and then build an ensemble.</p>\n\n<p>We trained many different CNNs: Resnet50, Resnet101, Resnet152, InceptionResnetV2, DenseNet161. Each one gives 74-75 LB score except of DenseNet161 which gives about 73. Using simple averaging we got 77.5. However, it was not enough and we decided to train a second-layer model.</p>\n\n<p>We had only one CNN which was trained using KFolds split. It was Resnet-50. For each product and for each image of this product we extracted TOP-5 probabilities and 5 relevant classes. So, there were 5 categorical and 5 float features for each image. We also noticed that one image could be present in the data multiple times, thus, we computed MD5 hash for each picture and matched it with the most common class IDs. So, we built another 10 features for each image: 5 TOP most common class IDs with such hash and the number of images with such hash.</p>\n\n<p>After this procedure, we had 80 raw features (40 numerical and 40 categorical). We preprocessed categorical features using different approaches. Since it is impossible to train xgboost on 5270 classes we added a new feature \"possible_class_id\" and changed the multiclassification problem to binary classification. Next, we were training a model which predicts 'is it true that this product has class_id which is equal to \"possible_class_id\"'. Therefore, a new dataset consisted of samples (product_features, possible_class_id, binary_target). We only used samples with possible_class_id which was one of the TOP5 predictions of CNN or one of the possible classes based on hash trick.</p>\n\n<p>Using the new dataset it was possible to train a second-layer model. We tried different models: XGBoost, LightGBM, CatBoost, Random Forest, Extra Trees. They gave similiar results and averaging of their predictions improved our score. Since we didn't want to spend too much time to finding the best coefficient for each model in our ensemble, we decided to train a third-layer model.</p>\n\n<p>For third layer we used:</p>\n\n<ol>\n<li>The features which were using for train second-layer models</li>\n<li>Predictions of second-layer models. We also computed several statistics of predictions on a set of possible class ids.</li>\n<li>Priors for categorical features. For each class_id we computed a frequency of this class in the whole dataset and a frequency of this class among the closest products based on their IDs (50k closest products). It was helpful because the dataset has a slight leak.</li>\n</ol>\n\n<p>Our third-layer models are three different LightGBM models and XGBoost one. Then, we used weighted arifmetic mean of these models to compute predictions for the test set. It seems that we could improve our score by building a fourth layer model :)</p>\n\n<p>For test predictions we used two different approaches:</p>\n\n<ol>\n<li>First, average probabilities extracted from CNNs, then build features and use them in second and third layer models.</li>\n<li>Extract predictions of each CNN, build features, use it for predictions on second and third layers and only then merge predictions of different CNNs.</li>\n</ol>\n\n<p>These two approaches gave almost the same score. We averaged their predictions and got a slight improvement.</p>\n\n<p>Some notes:</p>\n\n<ul>\n<li><p>I also tried to use predictions of the test set for training CNN. It improved the score of my single Resnet101 from 0.749 to 0.754. I stopped training after 1.5 epochs because such model showed worse results when I tried to put the predictions to XGBoost. However, I think it was possible to get 76+ from this single model using this approach. I also tried to apply this trick to second-layer models, the score also improved but it decreased the score of the third-layer model. I do believe that we could get better results applying this trick to our third-layer models. However, we didn't have extra time for doing it :(</p></li>\n<li><p>For training CNN we (Alexey Kharlamov and me) used such approach: first, initalize a model using pretrained weights from Imagenet, there were not frozen layers, the augmentation was disabled. Then, as soon as the validation score stopped growing, we added augmentation and doubled batch size. After this procedure, the score began to grow sharply. But, after a while the growth stopped, so we repeated this procedure again and again.</p></li>\n</ul>\n\n<p>Congratulations to the winners and thanks for all the participants!</p>",
  "messages": [
    {
      "id": "257907",
      "postDate": "12/15/2017 04:10:58",
      "content": "<p>The key idea was to train several different convolutional neural networks and then build an ensemble.</p>\n\n<p>We trained many different CNNs: Resnet50, Resnet101, Resnet152, InceptionResnetV2, DenseNet161. Each one gives 74-75 LB score except of DenseNet161 which gives about 73. Using simple averaging we got 77.5. However, it was not enough and we decided to train a second-layer model.</p>\n\n<p>We had only one CNN which was trained using KFolds split. It was Resnet-50. For each product and for each image of this product we extracted TOP-5 probabilities and 5 relevant classes. So, there were 5 categorical and 5 float features for each image. We also noticed that one image could be present in the data multiple times, thus, we computed MD5 hash for each picture and matched it with the most common class IDs. So, we built another 10 features for each image: 5 TOP most common class IDs with such hash and the number of images with such hash.</p>\n\n<p>After this procedure, we had 80 raw features (40 numerical and 40 categorical). We preprocessed categorical features using different approaches. Since it is impossible to train xgboost on 5270 classes we added a new feature \"possible_class_id\" and changed the multiclassification problem to binary classification. Next, we were training a model which predicts 'is it true that this product has class_id which is equal to \"possible_class_id\"'. Therefore, a new dataset consisted of samples (product_features, possible_class_id, binary_target). We only used samples with possible_class_id which was one of the TOP5 predictions of CNN or one of the possible classes based on hash trick.</p>\n\n<p>Using the new dataset it was possible to train a second-layer model. We tried different models: XGBoost, LightGBM, CatBoost, Random Forest, Extra Trees. They gave similiar results and averaging of their predictions improved our score. Since we didn't want to spend too much time to finding the best coefficient for each model in our ensemble, we decided to train a third-layer model.</p>\n\n<p>For third layer we used:</p>\n\n<ol>\n<li>The features which were using for train second-layer models</li>\n<li>Predictions of second-layer models. We also computed several statistics of predictions on a set of possible class ids.</li>\n<li>Priors for categorical features. For each class_id we computed a frequency of this class in the whole dataset and a frequency of this class among the closest products based on their IDs (50k closest products). It was helpful because the dataset has a slight leak.</li>\n</ol>\n\n<p>Our third-layer models are three different LightGBM models and XGBoost one. Then, we used weighted arifmetic mean of these models to compute predictions for the test set. It seems that we could improve our score by building a fourth layer model :)</p>\n\n<p>For test predictions we used two different approaches:</p>\n\n<ol>\n<li>First, average probabilities extracted from CNNs, then build features and use them in second and third layer models.</li>\n<li>Extract predictions of each CNN, build features, use it for predictions on second and third layers and only then merge predictions of different CNNs.</li>\n</ol>\n\n<p>These two approaches gave almost the same score. We averaged their predictions and got a slight improvement.</p>\n\n<p>Some notes:</p>\n\n<ul>\n<li><p>I also tried to use predictions of the test set for training CNN. It improved the score of my single Resnet101 from 0.749 to 0.754. I stopped training after 1.5 epochs because such model showed worse results when I tried to put the predictions to XGBoost. However, I think it was possible to get 76+ from this single model using this approach. I also tried to apply this trick to second-layer models, the score also improved but it decreased the score of the third-layer model. I do believe that we could get better results applying this trick to our third-layer models. However, we didn't have extra time for doing it :(</p></li>\n<li><p>For training CNN we (Alexey Kharlamov and me) used such approach: first, initalize a model using pretrained weights from Imagenet, there were not frozen layers, the augmentation was disabled. Then, as soon as the validation score stopped growing, we added augmentation and doubled batch size. After this procedure, the score began to grow sharply. But, after a while the growth stopped, so we repeated this procedure again and again.</p></li>\n</ul>\n\n<p>Congratulations to the winners and thanks for all the participants!</p>",
      "rawMarkdown": "The key idea was to train several different convolutional neural networks and then build an ensemble.\n\nWe trained many different CNNs: Resnet50, Resnet101, Resnet152, InceptionResnetV2, DenseNet161. Each one gives 74-75 LB score except of DenseNet161 which gives about 73. Using simple averaging we got 77.5. However, it was not enough and we decided to train a second-layer model.\n\nWe had only one CNN which was trained using KFolds split. It was Resnet-50. For each product and for each image of this product we extracted TOP-5 probabilities and 5 relevant classes. So, there were 5 categorical and 5 float features for each image. We also noticed that one image could be present in the data multiple times, thus, we computed MD5 hash for each picture and matched it with the most common class IDs. So, we built another 10 features for each image: 5 TOP most common class IDs with such hash and the number of images with such hash.\n\nAfter this procedure, we had 80 raw features (40 numerical and 40 categorical). We preprocessed categorical features using different approaches. Since it is impossible to train xgboost on 5270 classes we added a new feature \"possible_class_id\" and changed the multiclassification problem to binary classification. Next, we were training a model which predicts 'is it true that this product has class_id which is equal to \"possible_class_id\"'. Therefore, a new dataset consisted of samples (product_features, possible_class_id, binary_target). We only used samples with possible_class_id which was one of the TOP5 predictions of CNN or one of the possible classes based on hash trick.\n\nUsing the new dataset it was possible to train a second-layer model. We tried different models: XGBoost, LightGBM, CatBoost, Random Forest, Extra Trees. They gave similiar results and averaging of their predictions improved our score. Since we didn't want to spend too much time to finding the best coefficient for each model in our ensemble, we decided to train a third-layer model.\n\nFor third layer we used:\n\n1. The features which were using for train second-layer models\n2. Predictions of second-layer models. We also computed several statistics of predictions on a set of possible class ids.\n3. Priors for categorical features. For each class_id we computed a frequency of this class in the whole dataset and a frequency of this class among the closest products based on their IDs (50k closest products). It was helpful because the dataset has a slight leak.\n\nOur third-layer models are three different LightGBM models and XGBoost one. Then, we used weighted arifmetic mean of these models to compute predictions for the test set. It seems that we could improve our score by building a fourth layer model :)\n\nFor test predictions we used two different approaches:\n\n1. First, average probabilities extracted from CNNs, then build features and use them in second and third layer models.\n2. Extract predictions of each CNN, build features, use it for predictions on second and third layers and only then merge predictions of different CNNs.\n\nThese two approaches gave almost the same score. We averaged their predictions and got a slight improvement.\n\n\nSome notes:\n\n- I also tried to use predictions of the test set for training CNN. It improved the score of my single Resnet101 from 0.749 to 0.754. I stopped training after 1.5 epochs because such model showed worse results when I tried to put the predictions to XGBoost. However, I think it was possible to get 76+ from this single model using this approach. I also tried to apply this trick to second-layer models, the score also improved but it decreased the score of the third-layer model. I do believe that we could get better results applying this trick to our third-layer models. However, we didn't have extra time for doing it :(\n\n- For training CNN we (Alexey Kharlamov and me) used such approach: first, initalize a model using pretrained weights from Imagenet, there were not frozen layers, the augmentation was disabled. Then, as soon as the validation score stopped growing, we added augmentation and doubled batch size. After this procedure, the score began to grow sharply. But, after a while the growth stopped, so we repeated this procedure again and again.\n\n\nCongratulations to the winners and thanks for all the participants!",
      "votes": null
    },
    {
      "id": "257938",
      "postDate": "12/15/2017 05:27:10",
      "content": "<p>Thank you for sharing your results. </p>\n\n<p>Your results are impressive and the incremental augmentation for training CNN is smart. My base CNN models are only on 72-73 range and i still cannot find a way to boost it to 75 range. I will try your method. Thanks!</p>\n\n<p>I have always want to make a comparsion between traditional xgboost tree and deep neural trees. You may want to try them for post submission.</p>\n\n<p><a href=\"https://github.com/chrischoy/fully-differentiable-deep-ndf-tf\">https://github.com/chrischoy/fully-differentiable-deep-ndf-tf</a></p>\n\n<p><a href=\"https://discuss.pytorch.org/t/problems-on-implementation-of-deep-neural-decision-forest/837\">https://discuss.pytorch.org/t/problems-on-implementation-of-deep-neural-decision-forest/837</a></p>",
      "rawMarkdown": "Thank you for sharing your results. \n\nYour results are impressive and the incremental augmentation for training CNN is smart. My base CNN models are only on 72-73 range and i still cannot find a way to boost it to 75 range. I will try your method. Thanks!\n\nI have always want to make a comparsion between traditional xgboost tree and deep neural trees. You may want to try them for post submission.\n\nhttps://github.com/chrischoy/fully-differentiable-deep-ndf-tf\n\nhttps://discuss.pytorch.org/t/problems-on-implementation-of-deep-neural-decision-forest/837",
      "votes": null
    },
    {
      "id": "257953",
      "postDate": "12/15/2017 06:30:00",
      "content": "<p>great sharing !</p>",
      "rawMarkdown": "great sharing !",
      "votes": null
    },
    {
      "id": "258064",
      "postDate": "12/15/2017 11:23:52",
      "content": "<p>Thanks for sharing!</p>\n\n<p>What exactly do you mean with \"average probabilities extracted from CNN\"?\nI'm not sure what the probabilities are here.</p>",
      "rawMarkdown": "Thanks for sharing!\n\nWhat exactly do you mean with \"average probabilities extracted from CNN\"?\nI'm not sure what the probabilities are here.",
      "votes": null
    },
    {
      "id": "258070",
      "postDate": "12/15/2017 11:38:23",
      "content": "<p>Thanks for sharing! Impressive results!</p>",
      "rawMarkdown": "Thanks for sharing! Impressive results!",
      "votes": null
    },
    {
      "id": "258284",
      "postDate": "12/15/2017 22:04:18",
      "content": "<p>Since the output of each CNN is 5270-dimensional vector where the sum of elements is equal to 1, it is interpreted that i-th element of this vector is a probability that an image, which was passing to CNN, has class i.</p>",
      "rawMarkdown": "Since the output of each CNN is 5270-dimensional vector where the sum of elements is equal to 1, it is interpreted that i-th element of this vector is a probability that an image, which was passing to CNN, has class i.",
      "votes": null
    },
    {
      "id": "258505",
      "postDate": "12/16/2017 09:37:33",
      "content": "<p>Congrats for your result! Very interesting techniques. How have you got a r101 to get such high score is amazing! How have you modified it? How have you done the training? Which data augmentation? </p>",
      "rawMarkdown": "Congrats for your result! Very interesting techniques. How have you got a r101 to get such high score is amazing! How have you modified it? How have you done the training? Which data augmentation?",
      "votes": null
    },
    {
      "id": "258507",
      "postDate": "12/16/2017 09:45:42",
      "content": "<p>Ah yes obviously, I guess the \"extracting\" got me confused.</p>",
      "rawMarkdown": "Ah yes obviously, I guess the \"extracting\" got me confused.",
      "votes": null
    },
    {
      "id": "258730",
      "postDate": "12/16/2017 22:59:38",
      "content": "<p>We didn't modified CNNs so much, we only changed pooling type to a global average pooling and the size of the latest dense from 1000 to 5270. We used 160x160 central crops, random shifts, rotations, scaling. We used different data augmentation for each CNN because of different convergence. As I mentioned above, we were adding augmentation several times until the network converged. For example, the final angle of rotation for Resnet101 is 80 degrees and only 25 degrees for DenseNet161.</p>",
      "rawMarkdown": "We didn't modified CNNs so much, we only changed pooling type to a global average pooling and the size of the latest dense from 1000 to 5270. We used 160x160 central crops, random shifts, rotations, scaling. We used different data augmentation for each CNN because of different convergence. As I mentioned above, we were adding augmentation several times until the network converged. For example, the final angle of rotation for Resnet101 is 80 degrees and only 25 degrees for DenseNet161.",
      "votes": null
    },
    {
      "id": "258822",
      "postDate": "12/17/2017 05:01:55",
      "content": "<p>Very Nice..\nThanks for sharing...</p>",
      "rawMarkdown": "Very Nice..\nThanks for sharing...",
      "votes": null
    },
    {
      "id": "261939",
      "postDate": "12/24/2017 16:11:43",
      "content": "<p>Can you share code for training a CNN? </p>",
      "rawMarkdown": "Can you share code for training a CNN?",
      "votes": null
    },
    {
      "id": "285074",
      "postDate": "02/19/2018 08:35:26",
      "content": "<p>Pavel Ostyakov and Alexey Kharlamov have told their solution on Yandex ML trainings meet up in Moscow. We have made an English subtitles for this video.</p>\n\n<p><a href=\"https://youtu.be/Mw2vdYv4ups\">https://youtu.be/Mw2vdYv4ups</a></p>",
      "rawMarkdown": "Pavel Ostyakov and Alexey Kharlamov have told their solution on Yandex ML trainings meet up in Moscow. We have made an English subtitles for this video.\n\nhttps://youtu.be/Mw2vdYv4ups",
      "votes": null
    },
    {
      "id": "327937",
      "postDate": "05/13/2018 01:05:08",
      "content": "<p>Hi, Pavel. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~</p>",
      "rawMarkdown": "Hi, Pavel. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~",
      "votes": null
    },
    {
      "id": "363594",
      "postDate": "07/29/2018 16:22:02",
      "content": "<p>How do i pass the output of CNN's to xgboost in order to get the final predictions</p>",
      "rawMarkdown": "How do i pass the output of CNN's to xgboost in order to get the final predictions",
      "votes": null
    },
    {
      "id": "435949",
      "postDate": "12/09/2018 06:10:58",
      "content": "<p>This may help  <a href=\"https://alan.do/deep-gradient-boosted-learning-4e33adaf2969\">https://alan.do/deep-gradient-boosted-learning-4e33adaf2969</a> </p>",
      "rawMarkdown": "This may help  https://alan.do/deep-gradient-boosted-learning-4e33adaf2969",
      "votes": null
    },
    {
      "id": "438856",
      "postDate": "12/14/2018 09:44:37",
      "content": "<p>Thank you for the assist</p>",
      "rawMarkdown": "Thank you for the assist",
      "votes": null
    },
    {
      "id": "467058",
      "postDate": "02/06/2019 11:52:38",
      "content": "<p>For training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?</p>",
      "rawMarkdown": "For training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?",
      "votes": null
    },
    {
      "id": "467059",
      "postDate": "02/06/2019 11:53:21",
      "content": "<p><a href=\"/pavelost\">@pavelost</a>\nFor training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?</p>",
      "rawMarkdown": "pavelost\nFor training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 257938,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/15/2017 05:27:10",
      "content": "<p>Thank you for sharing your results. </p>\n\n<p>Your results are impressive and the incremental augmentation for training CNN is smart. My base CNN models are only on 72-73 range and i still cannot find a way to boost it to 75 range. I will try your method. Thanks!</p>\n\n<p>I have always want to make a comparsion between traditional xgboost tree and deep neural trees. You may want to try them for post submission.</p>\n\n<p><a href=\"https://github.com/chrischoy/fully-differentiable-deep-ndf-tf\">https://github.com/chrischoy/fully-differentiable-deep-ndf-tf</a></p>\n\n<p><a href=\"https://discuss.pytorch.org/t/problems-on-implementation-of-deep-neural-decision-forest/837\">https://discuss.pytorch.org/t/problems-on-implementation-of-deep-neural-decision-forest/837</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 257953,
      "author_name": "yanchao727",
      "author_url": "",
      "post_date": "12/15/2017 06:30:00",
      "content": "<p>great sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258064,
      "author_name": "antonvanmoere",
      "author_url": "",
      "post_date": "12/15/2017 11:23:52",
      "content": "<p>Thanks for sharing!</p>\n\n<p>What exactly do you mean with \"average probabilities extracted from CNN\"?\nI'm not sure what the probabilities are here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 258284,
          "author_name": "pavelost",
          "author_url": "",
          "post_date": "12/15/2017 22:04:18",
          "content": "<p>Since the output of each CNN is 5270-dimensional vector where the sum of elements is equal to 1, it is interpreted that i-th element of this vector is a probability that an image, which was passing to CNN, has class i.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258507,
          "author_name": "antonvanmoere",
          "author_url": "",
          "post_date": "12/16/2017 09:45:42",
          "content": "<p>Ah yes obviously, I guess the \"extracting\" got me confused.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 467058,
          "author_name": "tonmoyj",
          "author_url": "",
          "post_date": "02/06/2019 11:52:38",
          "content": "<p>For training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258070,
      "author_name": "andreaslup",
      "author_url": "",
      "post_date": "12/15/2017 11:38:23",
      "content": "<p>Thanks for sharing! Impressive results!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258505,
      "author_name": "andreanobile",
      "author_url": "",
      "post_date": "12/16/2017 09:37:33",
      "content": "<p>Congrats for your result! Very interesting techniques. How have you got a r101 to get such high score is amazing! How have you modified it? How have you done the training? Which data augmentation? </p>",
      "votes": null,
      "replies": [
        {
          "id": 258730,
          "author_name": "pavelost",
          "author_url": "",
          "post_date": "12/16/2017 22:59:38",
          "content": "<p>We didn't modified CNNs so much, we only changed pooling type to a global average pooling and the size of the latest dense from 1000 to 5270. We used 160x160 central crops, random shifts, rotations, scaling. We used different data augmentation for each CNN because of different convergence. As I mentioned above, we were adding augmentation several times until the network converged. For example, the final angle of rotation for Resnet101 is 80 degrees and only 25 degrees for DenseNet161.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258822,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "12/17/2017 05:01:55",
      "content": "<p>Very Nice..\nThanks for sharing...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 261939,
      "author_name": "rebirth",
      "author_url": "",
      "post_date": "12/24/2017 16:11:43",
      "content": "<p>Can you share code for training a CNN? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 285074,
      "author_name": "emilkayumov",
      "author_url": "",
      "post_date": "02/19/2018 08:35:26",
      "content": "<p>Pavel Ostyakov and Alexey Kharlamov have told their solution on Yandex ML trainings meet up in Moscow. We have made an English subtitles for this video.</p>\n\n<p><a href=\"https://youtu.be/Mw2vdYv4ups\">https://youtu.be/Mw2vdYv4ups</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327937,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "05/13/2018 01:05:08",
      "content": "<p>Hi, Pavel. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 363594,
      "author_name": "ayushpatel567",
      "author_url": "",
      "post_date": "07/29/2018 16:22:02",
      "content": "<p>How do i pass the output of CNN's to xgboost in order to get the final predictions</p>",
      "votes": null,
      "replies": [
        {
          "id": 435949,
          "author_name": "rahulpathak",
          "author_url": "",
          "post_date": "12/09/2018 06:10:58",
          "content": "<p>This may help  <a href=\"https://alan.do/deep-gradient-boosted-learning-4e33adaf2969\">https://alan.do/deep-gradient-boosted-learning-4e33adaf2969</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438856,
          "author_name": "ayushpatel567",
          "author_url": "",
          "post_date": "12/14/2018 09:44:37",
          "content": "<p>Thank you for the assist</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 467059,
      "author_name": "tonmoyj",
      "author_url": "",
      "post_date": "02/06/2019 11:53:21",
      "content": "<p><a href=\"/pavelost\">@pavelost</a>\nFor training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "257907": "The key idea was to train several different convolutional neural networks and then build an ensemble.\n\nWe trained many different CNNs: Resnet50, Resnet101, Resnet152, InceptionResnetV2, DenseNet161. Each one gives 74-75 LB score except of DenseNet161 which gives about 73. Using simple averaging we got 77.5. However, it was not enough and we decided to train a second-layer model.\n\nWe had only one CNN which was trained using KFolds split. It was Resnet-50. For each product and for each image of this product we extracted TOP-5 probabilities and 5 relevant classes. So, there were 5 categorical and 5 float features for each image. We also noticed that one image could be present in the data multiple times, thus, we computed MD5 hash for each picture and matched it with the most common class IDs. So, we built another 10 features for each image: 5 TOP most common class IDs with such hash and the number of images with such hash.\n\nAfter this procedure, we had 80 raw features (40 numerical and 40 categorical). We preprocessed categorical features using different approaches. Since it is impossible to train xgboost on 5270 classes we added a new feature \"possible_class_id\" and changed the multiclassification problem to binary classification. Next, we were training a model which predicts 'is it true that this product has class_id which is equal to \"possible_class_id\"'. Therefore, a new dataset consisted of samples (product_features, possible_class_id, binary_target). We only used samples with possible_class_id which was one of the TOP5 predictions of CNN or one of the possible classes based on hash trick.\n\nUsing the new dataset it was possible to train a second-layer model. We tried different models: XGBoost, LightGBM, CatBoost, Random Forest, Extra Trees. They gave similiar results and averaging of their predictions improved our score. Since we didn't want to spend too much time to finding the best coefficient for each model in our ensemble, we decided to train a third-layer model.\n\nFor third layer we used:\n\n1. The features which were using for train second-layer models\n2. Predictions of second-layer models. We also computed several statistics of predictions on a set of possible class ids.\n3. Priors for categorical features. For each class_id we computed a frequency of this class in the whole dataset and a frequency of this class among the closest products based on their IDs (50k closest products). It was helpful because the dataset has a slight leak.\n\nOur third-layer models are three different LightGBM models and XGBoost one. Then, we used weighted arifmetic mean of these models to compute predictions for the test set. It seems that we could improve our score by building a fourth layer model :)\n\nFor test predictions we used two different approaches:\n\n1. First, average probabilities extracted from CNNs, then build features and use them in second and third layer models.\n2. Extract predictions of each CNN, build features, use it for predictions on second and third layers and only then merge predictions of different CNNs.\n\nThese two approaches gave almost the same score. We averaged their predictions and got a slight improvement.\n\n\nSome notes:\n\n- I also tried to use predictions of the test set for training CNN. It improved the score of my single Resnet101 from 0.749 to 0.754. I stopped training after 1.5 epochs because such model showed worse results when I tried to put the predictions to XGBoost. However, I think it was possible to get 76+ from this single model using this approach. I also tried to apply this trick to second-layer models, the score also improved but it decreased the score of the third-layer model. I do believe that we could get better results applying this trick to our third-layer models. However, we didn't have extra time for doing it :(\n\n- For training CNN we (Alexey Kharlamov and me) used such approach: first, initalize a model using pretrained weights from Imagenet, there were not frozen layers, the augmentation was disabled. Then, as soon as the validation score stopped growing, we added augmentation and doubled batch size. After this procedure, the score began to grow sharply. But, after a while the growth stopped, so we repeated this procedure again and again.\n\n\nCongratulations to the winners and thanks for all the participants!",
    "257938": "Thank you for sharing your results. \n\nYour results are impressive and the incremental augmentation for training CNN is smart. My base CNN models are only on 72-73 range and i still cannot find a way to boost it to 75 range. I will try your method. Thanks!\n\nI have always want to make a comparsion between traditional xgboost tree and deep neural trees. You may want to try them for post submission.\n\nhttps://github.com/chrischoy/fully-differentiable-deep-ndf-tf\n\nhttps://discuss.pytorch.org/t/problems-on-implementation-of-deep-neural-decision-forest/837",
    "257953": "great sharing !",
    "258064": "Thanks for sharing!\n\nWhat exactly do you mean with \"average probabilities extracted from CNN\"?\nI'm not sure what the probabilities are here.",
    "258070": "Thanks for sharing! Impressive results!",
    "258284": "Since the output of each CNN is 5270-dimensional vector where the sum of elements is equal to 1, it is interpreted that i-th element of this vector is a probability that an image, which was passing to CNN, has class i.",
    "258505": "Congrats for your result! Very interesting techniques. How have you got a r101 to get such high score is amazing! How have you modified it? How have you done the training? Which data augmentation?",
    "258507": "Ah yes obviously, I guess the \"extracting\" got me confused.",
    "258730": "We didn't modified CNNs so much, we only changed pooling type to a global average pooling and the size of the latest dense from 1000 to 5270. We used 160x160 central crops, random shifts, rotations, scaling. We used different data augmentation for each CNN because of different convergence. As I mentioned above, we were adding augmentation several times until the network converged. For example, the final angle of rotation for Resnet101 is 80 degrees and only 25 degrees for DenseNet161.",
    "258822": "Very Nice..\nThanks for sharing...",
    "261939": "Can you share code for training a CNN?",
    "285074": "Pavel Ostyakov and Alexey Kharlamov have told their solution on Yandex ML trainings meet up in Moscow. We have made an English subtitles for this video.\n\nhttps://youtu.be/Mw2vdYv4ups",
    "327937": "Hi, Pavel. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~",
    "363594": "How do i pass the output of CNN's to xgboost in order to get the final predictions",
    "435949": "This may help  https://alan.do/deep-gradient-boosted-learning-4e33adaf2969",
    "438856": "Thank you for the assist",
    "467058": "For training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?",
    "467059": "pavelost\nFor training the CNN when you increase the batch size and augmentation, do you increase the learning rate or change it in anyway?"
  },
  "source": "meta"
}