{
  "id": 60026,
  "title": "7th place solution",
  "url": "/competitions/avito-demand-prediction/writeups/light-in-june-7th-place-solution",
  "author_name": "",
  "post_date": "2018-06-29T13:40:08.427Z",
  "votes": 33,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, dear kagglers.\nFirst of all, thanks to Kaggle and Avito for holding such a wonderful competition. And congratulations to the winners and all Kagglers.</p>\n\n<p>Our final solution was large stacking that involved 64 models in total. It’s quite similar to the post that was posted by KazAnova on the <a href=\"http://blog.kaggle.com/2017/06/15/stacking-made-easy-an-introduction-to-stacknet-by-competitions-grandmaster-marios-michailidis-kazanova/\">kaggle blog</a> before. And after we completed building models, We grouped all the level-0 model into 4 group in order to gain some diversity :</p>\n\n<ul>\n<li>All models</li>\n<li>Branden and Takuoko-Angus models</li>\n<li>Takuoko-Angus and Yusaku models</li>\n<li>Yusaku and Branden models</li>\n</ul>\n\n<p>We used the following models for our stacking (level-1 &amp; level-2) : {XGB gblinear, LGBM,  Ridge, Extra Trees, simple NN}. Finally we got 20 level-1 models &amp; 5 level-2 model in consequence. than we just use a simple ridge for level-3 stacking (but excluded the meta-features which had negative weight in ridge, than rerun, until there is no negative weight appeared). It was one of our last submission, it scored 0.2148 on public LB (private 0.2186 / CV 0.208749).</p>\n\n<p>In another submission, I just applied the optim() function in R with BFGS solver on level-1 meta-features to find the best weight to blend, with following formula : x1*model1 + x2*model2 + ...... +x20*model20 + x21. It worked almost same well as our level-3 stacking. It scored 0.2148 on public LB(private 0.2186 / CV 0.2087724). Actually two sub are quite same.</p>\n\n<h1>Takuoko and Angus’s solution</h1>\n\n<p>Takuoko and I merged on the 10 days before merger deadline.</p>\n\n<h3>Feature Enginnering</h3>\n\n<ul>\n<li><p>Applying (min, max, mean, var) to numeric features that was already grouped by some categorical features (e.g. groupby(by=“region”)[“price”].mean())</p></li>\n<li><p>nunique feature 2-way interactions</p></li>\n<li><p>count encoding on categorical features</p></li>\n<li><p><a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">aggregated feature of categorical features</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">image features : mean, whiteness, dullness of RGB</a></p></li>\n<li><p>Applying PCA / TSVD to OHE categorical features</p></li>\n<li><p><a href=\"https://www.kaggle.com/sudalairajkumar/simple-feature-engg-notebook-spooky-author\">text stats features</a></p></li>\n<li><p>different word n-gram and char n-gram </p></li>\n<li><p>applying TF-IDF or not</p></li>\n<li><p>Some features from Angus’s FE (see the attachment)</p></li>\n</ul>\n\n<h3>Features not work</h3>\n\n<ul>\n<li><p>nunique feature 3-way interactions</p></li>\n<li><p>Applying tsvd on tf-idf text feature or VGG16 features</p></li>\n<li><p>Ohe-hot encoding on categorical features</p></li>\n<li><p>Applying GaussianRandomProjection / FastICA / LDA / SparseRandomProjection on OHE  </p></li>\n<li><p><a href=\"https://github.com/seatgeek/fuzzywuzzy\">fuzzywuzzy features</a></p></li>\n</ul>\n\n<h3>Modeling</h3>\n\n<p>-- tree based model : LGBM, XGB --\nBasically we used Bayesian for our tuning. And there is something special in our setting, In LGBM, Takuoko set 0.1 to feature_fraction and low colsample_byleve in XGB. We achieve great success by performing bagging with different seed. We can get about 0.2184 on public LB with single(5 seed avg.) LGBM by this approach.</p>\n\n<p>We also used objective=poisson. It needs more time to converge and its performance is not so good, but it gave us some extra boost when we doing stacking.</p>\n\n<p>-- NN --\nWe constructed a lot of different NN like GRU, Conv1D, Conv2D based on this <a href=\"https://www.kaggle.com/shanth84/rnn-detailed-explanation-0-2246\">public kernel</a>.\nWe also made a MLP since its score is pretty bad but good for stacking. Until end, We still can’t let our NN model beat the 0.2210 on public. even with BN and fine-tuned Dropout. And that’s pretty frustrated for both of us.</p>\n\n<p>-- Ridge, drop0 model --\nRidge and <a href=\"https://www.kaggle.com/c/allstate-claims-severity/discussion/26416\">drop0 model</a> also were not good solo, but both of them improved our stacking score.</p>\n\n<p>Drop 0 model was inspired by the link above. We just dropped target=0 and train a LGBM.</p>\n\n<h2>Branden and Yusaku Solution</h2>\n\n<p>Branden and Yusaku were working together before merging with us. Branden and Yusaku used the same 5-folds based on a random split to train their level 0 models.</p>\n\n<h2>Yusaku’s solution</h2>\n\n<p>I trained 10 level-0 LGBM models that utilized image features.\nMy models were largely based on the kernel <a href=\"https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code\">https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code</a> (v14) with some tweaks and addition of CNN image features.\nTo add CNN features, the images were first resized to 224x224 and activations from the layer right before classification layers were extracted and average pooled.  The models were initialized with pre-trained weights from ImageNet.\nThe follow image features were used in the final models, though other CNN models were experimented with and did not work as well or were not usable due to memory constraints:</p>\n\n<ul>\n<li><p>VGG16 (512x7x7)  ||  Average pool (512x1x1) - Improved 0.0008 LB</p></li>\n<li><p>Densenet121 (1024x7x7) || Average pool (1024x1x1) - Improved 0.0012 LB</p></li>\n</ul>\n\n<p>Both average and max pool were tried, but average pool gave better CV/LB, so I ended up just using average pooled features in the final models.\nIt would have been interesting to try lower level image features extracted from the CNN models (to better capture qualities such as image focus, fuzziness, graininess), but did not get around to it.</p>\n\n<p>To encourage diversity and reduce correlations for the trained models, some were trained by removing features that were deemed important by LGBM, such as “image_top_1”, “city”, etc.  </p>\n\n<p>For text features, TFIDF of bigrams were used, unmodified from the original base kernel.  I also experimented with adding text features based on gensim’s doc2vec but did not help.  I did not try FastText, but I probably should have based on favorable results reported by other teams.</p>\n\n<p>My models’ predictions ended up having relatively low correlations with Branden’s models despite both of us using LGBM; combining my models with Branden’s gave a pretty big boost.</p>\n\n<h2>Branden’s solution</h2>\n\n<p>He will release it in the comment field after few days!!!</p>",
  "messages": [
    {
      "id": "350264",
      "postDate": "06/29/2018 13:23:37",
      "content": "<p>Hi, dear kagglers.\nFirst of all, thanks to Kaggle and Avito for holding such a wonderful competition. And congratulations to the winners and all Kagglers.</p>\n\n<p>Our final solution was large stacking that involved 64 models in total. It’s quite similar to the post that was posted by KazAnova on the <a href=\"http://blog.kaggle.com/2017/06/15/stacking-made-easy-an-introduction-to-stacknet-by-competitions-grandmaster-marios-michailidis-kazanova/\">kaggle blog</a> before. And after we completed building models, We grouped all the level-0 model into 4 group in order to gain some diversity :</p>\n\n<ul>\n<li>All models</li>\n<li>Branden and Takuoko-Angus models</li>\n<li>Takuoko-Angus and Yusaku models</li>\n<li>Yusaku and Branden models</li>\n</ul>\n\n<p>We used the following models for our stacking (level-1 &amp; level-2) : {XGB gblinear, LGBM,  Ridge, Extra Trees, simple NN}. Finally we got 20 level-1 models &amp; 5 level-2 model in consequence. than we just use a simple ridge for level-3 stacking (but excluded the meta-features which had negative weight in ridge, than rerun, until there is no negative weight appeared). It was one of our last submission, it scored 0.2148 on public LB (private 0.2186 / CV 0.208749).</p>\n\n<p>In another submission, I just applied the optim() function in R with BFGS solver on level-1 meta-features to find the best weight to blend, with following formula : x1*model1 + x2*model2 + ...... +x20*model20 + x21. It worked almost same well as our level-3 stacking. It scored 0.2148 on public LB(private 0.2186 / CV 0.2087724). Actually two sub are quite same.</p>\n\n<h1>Takuoko and Angus’s solution</h1>\n\n<p>Takuoko and I merged on the 10 days before merger deadline.</p>\n\n<h3>Feature Enginnering</h3>\n\n<ul>\n<li><p>Applying (min, max, mean, var) to numeric features that was already grouped by some categorical features (e.g. groupby(by=“region”)[“price”].mean())</p></li>\n<li><p>nunique feature 2-way interactions</p></li>\n<li><p>count encoding on categorical features</p></li>\n<li><p><a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">aggregated feature of categorical features</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">image features : mean, whiteness, dullness of RGB</a></p></li>\n<li><p>Applying PCA / TSVD to OHE categorical features</p></li>\n<li><p><a href=\"https://www.kaggle.com/sudalairajkumar/simple-feature-engg-notebook-spooky-author\">text stats features</a></p></li>\n<li><p>different word n-gram and char n-gram </p></li>\n<li><p>applying TF-IDF or not</p></li>\n<li><p>Some features from Angus’s FE (see the attachment)</p></li>\n</ul>\n\n<h3>Features not work</h3>\n\n<ul>\n<li><p>nunique feature 3-way interactions</p></li>\n<li><p>Applying tsvd on tf-idf text feature or VGG16 features</p></li>\n<li><p>Ohe-hot encoding on categorical features</p></li>\n<li><p>Applying GaussianRandomProjection / FastICA / LDA / SparseRandomProjection on OHE  </p></li>\n<li><p><a href=\"https://github.com/seatgeek/fuzzywuzzy\">fuzzywuzzy features</a></p></li>\n</ul>\n\n<h3>Modeling</h3>\n\n<p>-- tree based model : LGBM, XGB --\nBasically we used Bayesian for our tuning. And there is something special in our setting, In LGBM, Takuoko set 0.1 to feature_fraction and low colsample_byleve in XGB. We achieve great success by performing bagging with different seed. We can get about 0.2184 on public LB with single(5 seed avg.) LGBM by this approach.</p>\n\n<p>We also used objective=poisson. It needs more time to converge and its performance is not so good, but it gave us some extra boost when we doing stacking.</p>\n\n<p>-- NN --\nWe constructed a lot of different NN like GRU, Conv1D, Conv2D based on this <a href=\"https://www.kaggle.com/shanth84/rnn-detailed-explanation-0-2246\">public kernel</a>.\nWe also made a MLP since its score is pretty bad but good for stacking. Until end, We still can’t let our NN model beat the 0.2210 on public. even with BN and fine-tuned Dropout. And that’s pretty frustrated for both of us.</p>\n\n<p>-- Ridge, drop0 model --\nRidge and <a href=\"https://www.kaggle.com/c/allstate-claims-severity/discussion/26416\">drop0 model</a> also were not good solo, but both of them improved our stacking score.</p>\n\n<p>Drop 0 model was inspired by the link above. We just dropped target=0 and train a LGBM.</p>\n\n<h2>Branden and Yusaku Solution</h2>\n\n<p>Branden and Yusaku were working together before merging with us. Branden and Yusaku used the same 5-folds based on a random split to train their level 0 models.</p>\n\n<h2>Yusaku’s solution</h2>\n\n<p>I trained 10 level-0 LGBM models that utilized image features.\nMy models were largely based on the kernel <a href=\"https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code\">https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code</a> (v14) with some tweaks and addition of CNN image features.\nTo add CNN features, the images were first resized to 224x224 and activations from the layer right before classification layers were extracted and average pooled.  The models were initialized with pre-trained weights from ImageNet.\nThe follow image features were used in the final models, though other CNN models were experimented with and did not work as well or were not usable due to memory constraints:</p>\n\n<ul>\n<li><p>VGG16 (512x7x7)  ||  Average pool (512x1x1) - Improved 0.0008 LB</p></li>\n<li><p>Densenet121 (1024x7x7) || Average pool (1024x1x1) - Improved 0.0012 LB</p></li>\n</ul>\n\n<p>Both average and max pool were tried, but average pool gave better CV/LB, so I ended up just using average pooled features in the final models.\nIt would have been interesting to try lower level image features extracted from the CNN models (to better capture qualities such as image focus, fuzziness, graininess), but did not get around to it.</p>\n\n<p>To encourage diversity and reduce correlations for the trained models, some were trained by removing features that were deemed important by LGBM, such as “image_top_1”, “city”, etc.  </p>\n\n<p>For text features, TFIDF of bigrams were used, unmodified from the original base kernel.  I also experimented with adding text features based on gensim’s doc2vec but did not help.  I did not try FastText, but I probably should have based on favorable results reported by other teams.</p>\n\n<p>My models’ predictions ended up having relatively low correlations with Branden’s models despite both of us using LGBM; combining my models with Branden’s gave a pretty big boost.</p>\n\n<h2>Branden’s solution</h2>\n\n<p>He will release it in the comment field after few days!!!</p>",
      "rawMarkdown": "Hi, dear kagglers.\nFirst of all, thanks to Kaggle and Avito for holding such a wonderful competition. And congratulations to the winners and all Kagglers.\n\nOur final solution was large stacking that involved 64 models in total. It’s quite similar to the post that was posted by KazAnova on the [kaggle blog](http://blog.kaggle.com/2017/06/15/stacking-made-easy-an-introduction-to-stacknet-by-competitions-grandmaster-marios-michailidis-kazanova/) before. And after we completed building models, We grouped all the level-0 model into 4 group in order to gain some diversity :\n\n+ All models\n+ Branden and Takuoko-Angus models\n+ Takuoko-Angus and Yusaku models\n+ Yusaku and Branden models\n\nWe used the following models for our stacking (level-1 &amp; level-2) : {XGB gblinear, LGBM,  Ridge, Extra Trees, simple NN}. Finally we got 20 level-1 models &amp; 5 level-2 model in consequence. than we just use a simple ridge for level-3 stacking (but excluded the meta-features which had negative weight in ridge, than rerun, until there is no negative weight appeared). It was one of our last submission, it scored 0.2148 on public LB (private 0.2186 / CV 0.208749).\n\nIn another submission, I just applied the optim() function in R with BFGS solver on level-1 meta-features to find the best weight to blend, with following formula : x1*model1 + x2*model2 + ...... +x20*model20 + x21. It worked almost same well as our level-3 stacking. It scored 0.2148 on public LB(private 0.2186 / CV 0.2087724). Actually two sub are quite same.\n\n\n# Takuoko and Angus’s solution \n\nTakuoko and I merged on the 10 days before merger deadline.\n\n### Feature Enginnering \n+ Applying (min, max, mean, var) to numeric features that was already grouped by some categorical features (e.g. groupby(by=“region”)[“price”].mean())\n\n+ nunique feature 2-way interactions\n\n+ count encoding on categorical features\n\n+ [aggregated feature of categorical features](https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm )\n\n+ [image features : mean, whiteness, dullness of RGB](https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality )\n\n+ Applying PCA / TSVD to OHE categorical features\n\n+ [text stats features](https://www.kaggle.com/sudalairajkumar/simple-feature-engg-notebook-spooky-author)\n\n+ different word n-gram and char n-gram \n\n+ applying TF-IDF or not\n\n+ Some features from Angus’s FE (see the attachment)\n\n\n### Features not work   \n\n+ nunique feature 3-way interactions\n\n+ Applying tsvd on tf-idf text feature or VGG16 features\n\n+ Ohe-hot encoding on categorical features\n\n+ Applying GaussianRandomProjection / FastICA / LDA / SparseRandomProjection on OHE  \n\n+ [fuzzywuzzy features](https://github.com/seatgeek/fuzzywuzzy )\n\n\n### Modeling\n\n-- tree based model : LGBM, XGB --\nBasically we used Bayesian for our tuning. And there is something special in our setting, In LGBM, Takuoko set 0.1 to feature_fraction and low colsample_byleve in XGB. We achieve great success by performing bagging with different seed. We can get about 0.2184 on public LB with single(5 seed avg.) LGBM by this approach.\n \nWe also used objective=poisson. It needs more time to converge and its performance is not so good, but it gave us some extra boost when we doing stacking.\n\n-- NN --\nWe constructed a lot of different NN like GRU, Conv1D, Conv2D based on this [public kernel](https://www.kaggle.com/shanth84/rnn-detailed-explanation-0-2246).\nWe also made a MLP since its score is pretty bad but good for stacking. Until end, We still can’t let our NN model beat the 0.2210 on public. even with BN and fine-tuned Dropout. And that’s pretty frustrated for both of us.\n\n-- Ridge, drop0 model --\nRidge and [drop0 model](https://www.kaggle.com/c/allstate-claims-severity/discussion/26416) also were not good solo, but both of them improved our stacking score.\n\nDrop 0 model was inspired by the link above. We just dropped target=0 and train a LGBM.\n\n## Branden and Yusaku Solution\n\nBranden and Yusaku were working together before merging with us. Branden and Yusaku used the same 5-folds based on a random split to train their level 0 models.\n\n## Yusaku’s solution\nI trained 10 level-0 LGBM models that utilized image features.\nMy models were largely based on the kernel https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code (v14) with some tweaks and addition of CNN image features.\nTo add CNN features, the images were first resized to 224x224 and activations from the layer right before classification layers were extracted and average pooled.  The models were initialized with pre-trained weights from ImageNet.\nThe follow image features were used in the final models, though other CNN models were experimented with and did not work as well or were not usable due to memory constraints:\n\n+ VGG16 (512x7x7)  ||  Average pool (512x1x1) - Improved 0.0008 LB\n\n+ Densenet121 (1024x7x7) || Average pool (1024x1x1) - Improved 0.0012 LB\n\nBoth average and max pool were tried, but average pool gave better CV/LB, so I ended up just using average pooled features in the final models.\nIt would have been interesting to try lower level image features extracted from the CNN models (to better capture qualities such as image focus, fuzziness, graininess), but did not get around to it.\n\nTo encourage diversity and reduce correlations for the trained models, some were trained by removing features that were deemed important by LGBM, such as “image_top_1”, “city”, etc.  \n\nFor text features, TFIDF of bigrams were used, unmodified from the original base kernel.  I also experimented with adding text features based on gensim’s doc2vec but did not help.  I did not try FastText, but I probably should have based on favorable results reported by other teams.\n\nMy models’ predictions ended up having relatively low correlations with Branden’s models despite both of us using LGBM; combining my models with Branden’s gave a pretty big boost.\n \n## Branden’s solution\nHe will release it in the comment field after few days!!!",
      "votes": null
    },
    {
      "id": "350278",
      "postDate": "06/29/2018 13:54:34",
      "content": "<p>awesome,Thank you for sharing!</p>",
      "rawMarkdown": "awesome,Thank you for sharing!",
      "votes": null
    },
    {
      "id": "350315",
      "postDate": "06/29/2018 14:45:54",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "350361",
      "postDate": "06/29/2018 16:01:37",
      "content": "<p>Your team's final boost was amazing, can't even see your back, lol, congrats!</p>",
      "rawMarkdown": "Your team's final boost was amazing, can't even see your back, lol, congrats!",
      "votes": null
    },
    {
      "id": "350619",
      "postDate": "06/30/2018 04:27:08",
      "content": "<p>Congratulations @Shuo Jen, Chang and team. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations @Shuo Jen, Chang and team. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "350631",
      "postDate": "06/30/2018 05:27:02",
      "content": "<p>Thank you for your kind words!</p>",
      "rawMarkdown": "Thank you for your kind words!",
      "votes": null
    },
    {
      "id": "350632",
      "postDate": "06/30/2018 05:27:19",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 350278,
      "author_name": "jetouxu",
      "author_url": "",
      "post_date": "06/29/2018 13:54:34",
      "content": "<p>awesome,Thank you for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 350315,
          "author_name": "andrew60909",
          "author_url": "",
          "post_date": "06/29/2018 14:45:54",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350361,
      "author_name": "peterzheng",
      "author_url": "",
      "post_date": "06/29/2018 16:01:37",
      "content": "<p>Your team's final boost was amazing, can't even see your back, lol, congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 350631,
          "author_name": "andrew60909",
          "author_url": "",
          "post_date": "06/30/2018 05:27:02",
          "content": "<p>Thank you for your kind words!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350619,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/30/2018 04:27:08",
      "content": "<p>Congratulations @Shuo Jen, Chang and team. Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 350632,
          "author_name": "andrew60909",
          "author_url": "",
          "post_date": "06/30/2018 05:27:19",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "350264": "Hi, dear kagglers.\nFirst of all, thanks to Kaggle and Avito for holding such a wonderful competition. And congratulations to the winners and all Kagglers.\n\nOur final solution was large stacking that involved 64 models in total. It’s quite similar to the post that was posted by KazAnova on the [kaggle blog](http://blog.kaggle.com/2017/06/15/stacking-made-easy-an-introduction-to-stacknet-by-competitions-grandmaster-marios-michailidis-kazanova/) before. And after we completed building models, We grouped all the level-0 model into 4 group in order to gain some diversity :\n\n+ All models\n+ Branden and Takuoko-Angus models\n+ Takuoko-Angus and Yusaku models\n+ Yusaku and Branden models\n\nWe used the following models for our stacking (level-1 &amp; level-2) : {XGB gblinear, LGBM,  Ridge, Extra Trees, simple NN}. Finally we got 20 level-1 models &amp; 5 level-2 model in consequence. than we just use a simple ridge for level-3 stacking (but excluded the meta-features which had negative weight in ridge, than rerun, until there is no negative weight appeared). It was one of our last submission, it scored 0.2148 on public LB (private 0.2186 / CV 0.208749).\n\nIn another submission, I just applied the optim() function in R with BFGS solver on level-1 meta-features to find the best weight to blend, with following formula : x1*model1 + x2*model2 + ...... +x20*model20 + x21. It worked almost same well as our level-3 stacking. It scored 0.2148 on public LB(private 0.2186 / CV 0.2087724). Actually two sub are quite same.\n\n\n# Takuoko and Angus’s solution \n\nTakuoko and I merged on the 10 days before merger deadline.\n\n### Feature Enginnering \n+ Applying (min, max, mean, var) to numeric features that was already grouped by some categorical features (e.g. groupby(by=“region”)[“price”].mean())\n\n+ nunique feature 2-way interactions\n\n+ count encoding on categorical features\n\n+ [aggregated feature of categorical features](https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm )\n\n+ [image features : mean, whiteness, dullness of RGB](https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality )\n\n+ Applying PCA / TSVD to OHE categorical features\n\n+ [text stats features](https://www.kaggle.com/sudalairajkumar/simple-feature-engg-notebook-spooky-author)\n\n+ different word n-gram and char n-gram \n\n+ applying TF-IDF or not\n\n+ Some features from Angus’s FE (see the attachment)\n\n\n### Features not work   \n\n+ nunique feature 3-way interactions\n\n+ Applying tsvd on tf-idf text feature or VGG16 features\n\n+ Ohe-hot encoding on categorical features\n\n+ Applying GaussianRandomProjection / FastICA / LDA / SparseRandomProjection on OHE  \n\n+ [fuzzywuzzy features](https://github.com/seatgeek/fuzzywuzzy )\n\n\n### Modeling\n\n-- tree based model : LGBM, XGB --\nBasically we used Bayesian for our tuning. And there is something special in our setting, In LGBM, Takuoko set 0.1 to feature_fraction and low colsample_byleve in XGB. We achieve great success by performing bagging with different seed. We can get about 0.2184 on public LB with single(5 seed avg.) LGBM by this approach.\n \nWe also used objective=poisson. It needs more time to converge and its performance is not so good, but it gave us some extra boost when we doing stacking.\n\n-- NN --\nWe constructed a lot of different NN like GRU, Conv1D, Conv2D based on this [public kernel](https://www.kaggle.com/shanth84/rnn-detailed-explanation-0-2246).\nWe also made a MLP since its score is pretty bad but good for stacking. Until end, We still can’t let our NN model beat the 0.2210 on public. even with BN and fine-tuned Dropout. And that’s pretty frustrated for both of us.\n\n-- Ridge, drop0 model --\nRidge and [drop0 model](https://www.kaggle.com/c/allstate-claims-severity/discussion/26416) also were not good solo, but both of them improved our stacking score.\n\nDrop 0 model was inspired by the link above. We just dropped target=0 and train a LGBM.\n\n## Branden and Yusaku Solution\n\nBranden and Yusaku were working together before merging with us. Branden and Yusaku used the same 5-folds based on a random split to train their level 0 models.\n\n## Yusaku’s solution\nI trained 10 level-0 LGBM models that utilized image features.\nMy models were largely based on the kernel https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code (v14) with some tweaks and addition of CNN image features.\nTo add CNN features, the images were first resized to 224x224 and activations from the layer right before classification layers were extracted and average pooled.  The models were initialized with pre-trained weights from ImageNet.\nThe follow image features were used in the final models, though other CNN models were experimented with and did not work as well or were not usable due to memory constraints:\n\n+ VGG16 (512x7x7)  ||  Average pool (512x1x1) - Improved 0.0008 LB\n\n+ Densenet121 (1024x7x7) || Average pool (1024x1x1) - Improved 0.0012 LB\n\nBoth average and max pool were tried, but average pool gave better CV/LB, so I ended up just using average pooled features in the final models.\nIt would have been interesting to try lower level image features extracted from the CNN models (to better capture qualities such as image focus, fuzziness, graininess), but did not get around to it.\n\nTo encourage diversity and reduce correlations for the trained models, some were trained by removing features that were deemed important by LGBM, such as “image_top_1”, “city”, etc.  \n\nFor text features, TFIDF of bigrams were used, unmodified from the original base kernel.  I also experimented with adding text features based on gensim’s doc2vec but did not help.  I did not try FastText, but I probably should have based on favorable results reported by other teams.\n\nMy models’ predictions ended up having relatively low correlations with Branden’s models despite both of us using LGBM; combining my models with Branden’s gave a pretty big boost.\n \n## Branden’s solution\nHe will release it in the comment field after few days!!!",
    "350278": "awesome,Thank you for sharing!",
    "350315": "Thanks!",
    "350361": "Your team's final boost was amazing, can't even see your back, lol, congrats!",
    "350619": "Congratulations @Shuo Jen, Chang and team. Thanks for sharing.",
    "350631": "Thank you for your kind words!",
    "350632": "Thanks!"
  },
  "source": "meta"
}