{
  "id": 45737,
  "title": "Our 7th place solution",
  "url": "/competitions/cdiscount-image-classification-challenge/writeups/deeptortoise-our-7th-place-solution",
  "author_name": "",
  "post_date": "2017-12-15T06:11:16.052214500Z",
  "votes": 25,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First many thanks to Kaggle and CDiscount to host this very interesting competition. </p>\n\n<p>Congratulations to the prize winners, your performances are impressive and I look forward to read your solution. </p>\n\n<p>A special kudos to Heng CherKeng for all your sharings, your comments. You are definitely #1 in this competition in style!</p>\n\n<p>Thanks to my team: Kyle, alup and voglinio, it was a pleasure to work and share together. Special kudo to Kyle whose ensembling speed was very impressive.</p>\n\n<p>Here a a quick summary of our solution. It is a combination of 3 things:</p>\n\n<ul>\n<li><p>Baseline CNN models: InceptionResnetv2, Resnet101, SE-InceptionV3, Xception. These models got performance from .69+ to .72+</p></li>\n<li><p>We extracted the bottleneck features from these models, group them by product (+padding) to shape (bs, 4, dim_features) and train different models to combine the images and classify the object. I shared my top level models here <a href=\"https://www.kaggle.com/lamdang/models-to-combine-images-and-predict-items\">https://www.kaggle.com/lamdang/models-to-combine-images-and-predict-items</a></p>\n\n<ul><li>RNN: LSTM and GRU</li>\n<li>flat model: just flatten the inputs and put dense layers on top</li>\n<li>NetVlad: implementation based on <a href=\"https://arxiv.org/abs/1706.06905\">https://arxiv.org/abs/1706.06905</a> and <a href=\"https://arxiv.org/abs/1511.07247\">https://arxiv.org/abs/1511.07247</a>. Also from <a href=\"https://github.com/antoine77340/LOUPE\">https://github.com/antoine77340/LOUPE</a>. Kudos to Antoine Miech and Inria researchers on this great work.\nThese models individually get around .73-.74, when averaged get around .75</li>\n<li>Doing this previous steps and ensembling across baseline model get around .77</li></ul></li>\n<li><p>We extracted text box with EAST <a href=\"https://github.com/argman/EAST\">https://github.com/argman/EAST</a> and did OCR with <a href=\"https://github.com/meijieru/crnn.pytorch\">https://github.com/meijieru/crnn.pytorch</a>. We did not do language modelling correction after on the output. This boost the score further by ~0.01. We only take the detection and ocr model out of the box and did not retrain, further tuning of these models on this dataset may improve the performance and give extraboost to the global model.\nAlso, we could only manage to use the text in bag of char ngram fashion. Other way of exploiting it could also improve the result.</p></li>\n</ul>\n\n<p>So that is, globally it is quite simple. I am really happy to learn about some really nice aggregation techniques like NetVlad and text detection/OCR models.</p>",
  "messages": [
    {
      "id": "257948",
      "postDate": "12/15/2017 06:11:16",
      "content": "<p>First many thanks to Kaggle and CDiscount to host this very interesting competition. </p>\n\n<p>Congratulations to the prize winners, your performances are impressive and I look forward to read your solution. </p>\n\n<p>A special kudos to Heng CherKeng for all your sharings, your comments. You are definitely #1 in this competition in style!</p>\n\n<p>Thanks to my team: Kyle, alup and voglinio, it was a pleasure to work and share together. Special kudo to Kyle whose ensembling speed was very impressive.</p>\n\n<p>Here a a quick summary of our solution. It is a combination of 3 things:</p>\n\n<ul>\n<li><p>Baseline CNN models: InceptionResnetv2, Resnet101, SE-InceptionV3, Xception. These models got performance from .69+ to .72+</p></li>\n<li><p>We extracted the bottleneck features from these models, group them by product (+padding) to shape (bs, 4, dim_features) and train different models to combine the images and classify the object. I shared my top level models here <a href=\"https://www.kaggle.com/lamdang/models-to-combine-images-and-predict-items\">https://www.kaggle.com/lamdang/models-to-combine-images-and-predict-items</a></p>\n\n<ul><li>RNN: LSTM and GRU</li>\n<li>flat model: just flatten the inputs and put dense layers on top</li>\n<li>NetVlad: implementation based on <a href=\"https://arxiv.org/abs/1706.06905\">https://arxiv.org/abs/1706.06905</a> and <a href=\"https://arxiv.org/abs/1511.07247\">https://arxiv.org/abs/1511.07247</a>. Also from <a href=\"https://github.com/antoine77340/LOUPE\">https://github.com/antoine77340/LOUPE</a>. Kudos to Antoine Miech and Inria researchers on this great work.\nThese models individually get around .73-.74, when averaged get around .75</li>\n<li>Doing this previous steps and ensembling across baseline model get around .77</li></ul></li>\n<li><p>We extracted text box with EAST <a href=\"https://github.com/argman/EAST\">https://github.com/argman/EAST</a> and did OCR with <a href=\"https://github.com/meijieru/crnn.pytorch\">https://github.com/meijieru/crnn.pytorch</a>. We did not do language modelling correction after on the output. This boost the score further by ~0.01. We only take the detection and ocr model out of the box and did not retrain, further tuning of these models on this dataset may improve the performance and give extraboost to the global model.\nAlso, we could only manage to use the text in bag of char ngram fashion. Other way of exploiting it could also improve the result.</p></li>\n</ul>\n\n<p>So that is, globally it is quite simple. I am really happy to learn about some really nice aggregation techniques like NetVlad and text detection/OCR models.</p>",
      "rawMarkdown": "First many thanks to Kaggle and CDiscount to host this very interesting competition. \n\nCongratulations to the prize winners, your performances are impressive and I look forward to read your solution. \n\nA special kudos to Heng CherKeng for all your sharings, your comments. You are definitely #1 in this competition in style!\n\nThanks to my team: Kyle, alup and voglinio, it was a pleasure to work and share together. Special kudo to Kyle whose ensembling speed was very impressive.\n\nHere a a quick summary of our solution. It is a combination of 3 things:\n\n* Baseline CNN models: InceptionResnetv2, Resnet101, SE-InceptionV3, Xception. These models got performance from .69+ to .72+\n\n* We extracted the bottleneck features from these models, group them by product (+padding) to shape (bs, 4, dim_features) and train different models to combine the images and classify the object. I shared my top level models here https://www.kaggle.com/lamdang/models-to-combine-images-and-predict-items\n\t* RNN: LSTM and GRU\n\t* flat model: just flatten the inputs and put dense layers on top\n\t* NetVlad: implementation based on https://arxiv.org/abs/1706.06905 and https://arxiv.org/abs/1511.07247. Also from https://github.com/antoine77340/LOUPE. Kudos to Antoine Miech and Inria researchers on this great work.\nThese models individually get around .73-.74, when averaged get around .75\n    * Doing this previous steps and ensembling across baseline model get around .77\n\n* We extracted text box with EAST https://github.com/argman/EAST and did OCR with https://github.com/meijieru/crnn.pytorch. We did not do language modelling correction after on the output. This boost the score further by ~0.01. We only take the detection and ocr model out of the box and did not retrain, further tuning of these models on this dataset may improve the performance and give extraboost to the global model.\nAlso, we could only manage to use the text in bag of char ngram fashion. Other way of exploiting it could also improve the result.\n\nSo that is, globally it is quite simple. I am really happy to learn about some really nice aggregation techniques like NetVlad and text detection/OCR models.",
      "votes": null
    },
    {
      "id": "258013",
      "postDate": "12/15/2017 08:40:04",
      "content": "<p>Thank you for sharing knowledge!</p>",
      "rawMarkdown": "Thank you for sharing knowledge!",
      "votes": null
    },
    {
      "id": "258380",
      "postDate": "12/16/2017 02:29:48",
      "content": "<p>Thanks @Lam Dang and team for sharing your solution.</p>",
      "rawMarkdown": "Thanks @Lam Dang and team for sharing your solution.",
      "votes": null
    },
    {
      "id": "327941",
      "postDate": "05/13/2018 01:32:34",
      "content": "<p>Hi, Heng. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~</p>",
      "rawMarkdown": "Hi, Heng. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 258013,
      "author_name": "sondaoduy",
      "author_url": "",
      "post_date": "12/15/2017 08:40:04",
      "content": "<p>Thank you for sharing knowledge!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258380,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "12/16/2017 02:29:48",
      "content": "<p>Thanks @Lam Dang and team for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327941,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "05/13/2018 01:32:34",
      "content": "<p>Hi, Heng. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "257948": "First many thanks to Kaggle and CDiscount to host this very interesting competition. \n\nCongratulations to the prize winners, your performances are impressive and I look forward to read your solution. \n\nA special kudos to Heng CherKeng for all your sharings, your comments. You are definitely #1 in this competition in style!\n\nThanks to my team: Kyle, alup and voglinio, it was a pleasure to work and share together. Special kudo to Kyle whose ensembling speed was very impressive.\n\nHere a a quick summary of our solution. It is a combination of 3 things:\n\n* Baseline CNN models: InceptionResnetv2, Resnet101, SE-InceptionV3, Xception. These models got performance from .69+ to .72+\n\n* We extracted the bottleneck features from these models, group them by product (+padding) to shape (bs, 4, dim_features) and train different models to combine the images and classify the object. I shared my top level models here https://www.kaggle.com/lamdang/models-to-combine-images-and-predict-items\n\t* RNN: LSTM and GRU\n\t* flat model: just flatten the inputs and put dense layers on top\n\t* NetVlad: implementation based on https://arxiv.org/abs/1706.06905 and https://arxiv.org/abs/1511.07247. Also from https://github.com/antoine77340/LOUPE. Kudos to Antoine Miech and Inria researchers on this great work.\nThese models individually get around .73-.74, when averaged get around .75\n    * Doing this previous steps and ensembling across baseline model get around .77\n\n* We extracted text box with EAST https://github.com/argman/EAST and did OCR with https://github.com/meijieru/crnn.pytorch. We did not do language modelling correction after on the output. This boost the score further by ~0.01. We only take the detection and ocr model out of the box and did not retrain, further tuning of these models on this dataset may improve the performance and give extraboost to the global model.\nAlso, we could only manage to use the text in bag of char ngram fashion. Other way of exploiting it could also improve the result.\n\nSo that is, globally it is quite simple. I am really happy to learn about some really nice aggregation techniques like NetVlad and text detection/OCR models.",
    "258013": "Thank you for sharing knowledge!",
    "258380": "Thanks @Lam Dang and team for sharing your solution.",
    "327941": "Hi, Heng. Did you save the data(.bson or .jpg)? The data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~"
  },
  "source": "meta"
}