{
  "id": 61307,
  "title": "Sharing my solution",
  "url": "/competitions/avito-demand-prediction/discussion/61307",
  "author_name": "Adel Valiullin",
  "post_date": "2018-07-17T14:27:13.597000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, thanks to Avito and Kaggle for organizing such a wonderful competition! It was interesting all-in-one problem with text features, structured data and images. Really very big front of ideas and techniques to use during this competition. Data size was not so big and accessible to everyone and also it was leak free. </p>\n\n<p>But there was some interesting nature of target values (lots of zeros and discrete values that differed for different categories). I couldn't find usage of this fact and I don't know could someone use it for boosting his score.</p>\n\n<p>Last week of competition was really challenging. It was harder to climb on the leaderboard than in previous days. Thanks for all competitors for these last sleepless nights!</p>\n\n<p>Here is the summary of my approach:</p>\n\n<p><strong>Validation</strong> <br>\nI used 5 CV Folds as validation strategy. <br>\nIt differs with Public LB ~0.0034-0.004 and that different was stable</p>\n\n<p><strong>Feature Engineering</strong> <br>\nFeatures were computed on the concatenation of train and test sets.</p>\n\n<p><strong>Images</strong> <br>\nI used three Neural Network models (<strong>ResNet50, InceptionV3, Xception</strong>) for predictiong first top3 labels and scores for images.\nOther image features:</p>\n\n<ul>\n<li>image size</li>\n<li>height, width</li>\n<li>image blurness/lightness</li>\n<li>color channels</li>\n<li>rgb average color</li>\n<li>other image features (Average Pixel Width, Dominant Color etc) (<a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">from here</a>)</li>\n</ul>\n\n<p><strong>Text features</strong></p>\n\n<ul>\n<li>TF-IDF of bigrams for description</li>\n<li>TF-IDF of 1gram for the params and title</li>\n<li>word2vec features on words</li>\n<li>pretrained fasttext features on words</li>\n<li>pymorphy2 text normalization</li>\n<li>svd title/description</li>\n</ul>\n\n<p><strong>Other hand-made features</strong></p>\n\n<ul>\n<li>count of symbols in title/desc</li>\n<li>count of words in title/desc</li>\n<li>count of digits in title/desc</li>\n<li>number of uppercase/special symbols/punctuation</li>\n<li>number of stopwords</li>\n<li>missing price/image/descr flag</li>\n<li>log1n price</li>\n<li>price aggregations: min/max, median, var, mean of the groups of cat features</li>\n<li>other group statistics (city counts, groups counts/means/vars)</li>\n</ul>\n\n<p><strong>Models</strong> <br>\nI used <strong>10 LightGBM</strong> (lb 0.2202), <strong>10 XGBoost</strong> (lb 0.2228) and <strong>10 CatBoost</strong> (lb 0.2240) models. <br>\nAnd pretrained <strong>Neural Networks: ResNet50, InceptionV3, Xception</strong> for label prediction of the images.</p>\n\n<p><strong>Stacking</strong> <br>\nIt was 2 layers stack. <br>\nFirst layer included models of LightGBM, XGBoost, CatBoost with metafeatures. <br>\nSecond layer was LightGBM based on features generated with previous layer. <br>\nStacking helped to improve the score to lb <strong>0.2192</strong> (public) and <strong>0.2230</strong> (private)</p>",
  "messages": [
    {
      "id": 358096,
      "postDate": "2018-07-17T14:27:13.597Z",
      "content": "<p>First of all, thanks to Avito and Kaggle for organizing such a wonderful competition! It was interesting all-in-one problem with text features, structured data and images. Really very big front of ideas and techniques to use during this competition. Data size was not so big and accessible to everyone and also it was leak free. </p>\n\n<p>But there was some interesting nature of target values (lots of zeros and discrete values that differed for different categories). I couldn't find usage of this fact and I don't know could someone use it for boosting his score.</p>\n\n<p>Last week of competition was really challenging. It was harder to climb on the leaderboard than in previous days. Thanks for all competitors for these last sleepless nights!</p>\n\n<p>Here is the summary of my approach:</p>\n\n<p><strong>Validation</strong> <br>\nI used 5 CV Folds as validation strategy. <br>\nIt differs with Public LB ~0.0034-0.004 and that different was stable</p>\n\n<p><strong>Feature Engineering</strong> <br>\nFeatures were computed on the concatenation of train and test sets.</p>\n\n<p><strong>Images</strong> <br>\nI used three Neural Network models (<strong>ResNet50, InceptionV3, Xception</strong>) for predictiong first top3 labels and scores for images.\nOther image features:</p>\n\n<ul>\n<li>image size</li>\n<li>height, width</li>\n<li>image blurness/lightness</li>\n<li>color channels</li>\n<li>rgb average color</li>\n<li>other image features (Average Pixel Width, Dominant Color etc) (<a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">from here</a>)</li>\n</ul>\n\n<p><strong>Text features</strong></p>\n\n<ul>\n<li>TF-IDF of bigrams for description</li>\n<li>TF-IDF of 1gram for the params and title</li>\n<li>word2vec features on words</li>\n<li>pretrained fasttext features on words</li>\n<li>pymorphy2 text normalization</li>\n<li>svd title/description</li>\n</ul>\n\n<p><strong>Other hand-made features</strong></p>\n\n<ul>\n<li>count of symbols in title/desc</li>\n<li>count of words in title/desc</li>\n<li>count of digits in title/desc</li>\n<li>number of uppercase/special symbols/punctuation</li>\n<li>number of stopwords</li>\n<li>missing price/image/descr flag</li>\n<li>log1n price</li>\n<li>price aggregations: min/max, median, var, mean of the groups of cat features</li>\n<li>other group statistics (city counts, groups counts/means/vars)</li>\n</ul>\n\n<p><strong>Models</strong> <br>\nI used <strong>10 LightGBM</strong> (lb 0.2202), <strong>10 XGBoost</strong> (lb 0.2228) and <strong>10 CatBoost</strong> (lb 0.2240) models. <br>\nAnd pretrained <strong>Neural Networks: ResNet50, InceptionV3, Xception</strong> for label prediction of the images.</p>\n\n<p><strong>Stacking</strong> <br>\nIt was 2 layers stack. <br>\nFirst layer included models of LightGBM, XGBoost, CatBoost with metafeatures. <br>\nSecond layer was LightGBM based on features generated with previous layer. <br>\nStacking helped to improve the score to lb <strong>0.2192</strong> (public) and <strong>0.2230</strong> (private)</p>",
      "rawMarkdown": "First of all, thanks to Avito and Kaggle for organizing such a wonderful competition! It was interesting all-in-one problem with text features, structured data and images. Really very big front of ideas and techniques to use during this competition. Data size was not so big and accessible to everyone and also it was leak free. \n\nBut there was some interesting nature of target values (lots of zeros and discrete values that differed for different categories). I couldn't find usage of this fact and I don't know could someone use it for boosting his score.\n\nLast week of competition was really challenging. It was harder to climb on the leaderboard than in previous days. Thanks for all competitors for these last sleepless nights!\n\nHere is the summary of my approach:\n\n**Validation**  \nI used 5 CV Folds as validation strategy.  \nIt differs with Public LB ~0.0034-0.004 and that different was stable\n\n**Feature Engineering**  \nFeatures were computed on the concatenation of train and test sets.\n\n**Images**  \nI used three Neural Network models (**ResNet50, InceptionV3, Xception**) for predictiong first top3 labels and scores for images.\nOther image features:\n\n- image size\n- height, width\n- image blurness/lightness\n- color channels\n- rgb average color\n- other image features (Average Pixel Width, Dominant Color etc) ([from here][1])\n\n**Text features**\n\n- TF-IDF of bigrams for description\n- TF-IDF of 1gram for the params and title\n- word2vec features on words\n- pretrained fasttext features on words\n- pymorphy2 text normalization\n- svd title/description\n\n**Other hand-made features**\n\n- count of symbols in title/desc\n- count of words in title/desc\n- count of digits in title/desc\n- number of uppercase/special symbols/punctuation\n- number of stopwords\n- missing price/image/descr flag\n- log1n price\n- price aggregations: min/max, median, var, mean of the groups of cat features\n- other group statistics (city counts, groups counts/means/vars)\n\n**Models**  \nI used **10 LightGBM** (lb 0.2202), **10 XGBoost** (lb 0.2228) and **10 CatBoost** (lb 0.2240) models.  \nAnd pretrained **Neural Networks: ResNet50, InceptionV3, Xception** for label prediction of the images.\n\n**Stacking**  \nIt was 2 layers stack.  \nFirst layer included models of LightGBM, XGBoost, CatBoost with metafeatures.  \nSecond layer was LightGBM based on features generated with previous layer.  \nStacking helped to improve the score to lb **0.2192** (public) and **0.2230** (private)\n\n  [1]: https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "358096": "First of all, thanks to Avito and Kaggle for organizing such a wonderful competition! It was interesting all-in-one problem with text features, structured data and images. Really very big front of ideas and techniques to use during this competition. Data size was not so big and accessible to everyone and also it was leak free. \n\nBut there was some interesting nature of target values (lots of zeros and discrete values that differed for different categories). I couldn't find usage of this fact and I don't know could someone use it for boosting his score.\n\nLast week of competition was really challenging. It was harder to climb on the leaderboard than in previous days. Thanks for all competitors for these last sleepless nights!\n\nHere is the summary of my approach:\n\n**Validation**  \nI used 5 CV Folds as validation strategy.  \nIt differs with Public LB ~0.0034-0.004 and that different was stable\n\n**Feature Engineering**  \nFeatures were computed on the concatenation of train and test sets.\n\n**Images**  \nI used three Neural Network models (**ResNet50, InceptionV3, Xception**) for predictiong first top3 labels and scores for images.\nOther image features:\n\n- image size\n- height, width\n- image blurness/lightness\n- color channels\n- rgb average color\n- other image features (Average Pixel Width, Dominant Color etc) ([from here][1])\n\n**Text features**\n\n- TF-IDF of bigrams for description\n- TF-IDF of 1gram for the params and title\n- word2vec features on words\n- pretrained fasttext features on words\n- pymorphy2 text normalization\n- svd title/description\n\n**Other hand-made features**\n\n- count of symbols in title/desc\n- count of words in title/desc\n- count of digits in title/desc\n- number of uppercase/special symbols/punctuation\n- number of stopwords\n- missing price/image/descr flag\n- log1n price\n- price aggregations: min/max, median, var, mean of the groups of cat features\n- other group statistics (city counts, groups counts/means/vars)\n\n**Models**  \nI used **10 LightGBM** (lb 0.2202), **10 XGBoost** (lb 0.2228) and **10 CatBoost** (lb 0.2240) models.  \nAnd pretrained **Neural Networks: ResNet50, InceptionV3, Xception** for label prediction of the images.\n\n**Stacking**  \nIt was 2 layers stack.  \nFirst layer included models of LightGBM, XGBoost, CatBoost with metafeatures.  \nSecond layer was LightGBM based on features generated with previous layer.  \nStacking helped to improve the score to lb **0.2192** (public) and **0.2230** (private)\n\n  [1]: https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality"
  }
}