{
  "id": 60154,
  "title": "11th Solution Overview ",
  "url": "/competitions/avito-demand-prediction/writeups/11th-solution-overview",
  "author_name": "",
  "post_date": "2018-07-01T00:12:46.213Z",
  "votes": 38,
  "comment_count": 6,
  "views": 0,
  "content": "<p>What a competition this has been - it was such a rare opportunity to be able to tackle image, text,  tabular and time series competition in one single competition.   I personally found that the ability to iterate and execute ideas as quickly as possible is really critical in a competition like this. and it is for the same reason that I really admire all the solo participants who did well here, none more so than <a href=\"/dmitrylarko\">@dmitrylarko</a>, <a href=\"/kelexu\">@kelexu</a> and <a href=\"/wenbozhao\">@wenbozhao</a>   Despite managing to finish in the gold medal zone, it is obvious to me that the distance between us and the top teams is very big, so it is somehow humbling to see there is still so much room to improve.  We will definitely keep going to get better </p>\n\n<p>Our team names during this competition more or less reflect our journey in the last 7 weeks when we were active in the competition. First of all, 每天进步一点点means making small progress every day, while 挣扎到最后一刻 means fighting til the bitter end.  I don't think we have come across any single magical features that improve our score by more than 0.001,  while we were working frantically till the last seconds of the competition to make sure we can stay in the gold medal zone. The fight for the last gold medal positions was probably the most intensive that I have witnessed will multiple teams making sudden big jumps, so we are really happy that we manage to stay with them. </p>\n\n<p>I started the competition with <a href=\"/cxl923cc\">@cxl923cc</a>, <a href=\"/zhiqiangzhong\">@zhiqiangzhong</a> and <a href=\"/yl1202\">@yl1202</a>, we are close friends and colleagues with each other and wanted to use this exercise to form a group that can share the burden of kaggle competitions (seems it worked out well). Towards the merger deadline, we formed with <a href=\"/mzr2017\">@mzr2017</a> and <a href=\"/oyxuan\">@oyxuan</a> - because their modelling approach was quite different from ours, and the two groups complemented each other well. </p>\n\n<h2>Feature Engineering</h2>\n\n<p>I will describe the features we have been working on within our original team and let our teammate add theirs in following threads.  Most of our features are already covered by other top solutions, but anyway I will provide a list here. The following are shared by all models:</p>\n\n<ul>\n<li>text feature with TFIDF vectorizer for title, description and params,\nwe  played around with many different combination of TFIDF\nparameters, and it all added to our model diversity.</li>\n<li>SVD of TFIDF vectorizer features </li>\n<li>text statistical features such as length of text, number of number</li>\n<li>text features on ngrams, text distance features </li>\n<li>various groupby statistics between different categoricals </li>\n<li>aggregated features like the one shared in the kernels, </li>\n<li>LDA features </li>\n<li>Image features like the one that is extacted from pretrained models, as well as dullness, brightness</li>\n<li>rolling statistics such as number of ad in the same category in different time windows</li>\n</ul>\n\n<p>For non-NN models we specifically also created:</p>\n\n<ul>\n<li>rnn extracted features： rnn features extracted from rnn architecture (Bidirectional-GRU, attention, global max, global avg), and feed into non-nn model as tabular features. </li>\n<li>sentence2vec features\nboth group of features played important roles in our models. </li>\n</ul>\n\n<p>For NN model, we made heavy use of pre-trained word embedding models - three variant of FastText models shared in external data thread, and the self-train embedding model trained on all data. </p>\n\n<p>Aside from the above, we have also added the following features trained on five-fold oof manner:</p>\n\n<ul>\n<li>\"zero deal probability rate\" - five-fold oof probability prediction on if the deal probability is zero. </li>\n<li>target encoding features - again five-fold oof target encoding features for each original categorical features apart from user-id </li>\n</ul>\n\n<p>All these oof features are used directly on level 1 of the stacking, along side with other lv1 meta features (oof train and test)</p>\n\n<h2> Modelling Approach </h2>\n\n<p>We have used a wide range of models/algorithms in this competition. LightGBM with different parameters, XGB, Wordbatch FTRL_FM, Ridge, LinearSVR, RNN with Keras, Fully-connected NN with keras.</p>\n\n<p>In the last week of the competition, <a href=\"/mzr2017\">@mzr2017</a> found out that using xentropy in LightGBM, and building model by different parent categories helps to improve scores and give more diversity - and we started to retrained our better models with the corresponding settings.  We achieved our best LB model at public LB 0.2186 with xentropy training on all data - it was a model with 400+ features and some dense rnn &amp; sentence2vec features, and trained for more than 30 hours.  meanwhile Our best RNN model was around 0.2195 in public LB, courtesy of our teammates.</p>\n\n<h2> Stacking Approach </h2>\n\n<p>Stacking turned out to be very effective in this competition, and we were able to generate stacking score that is almost 0.004 better than our best l1 models. </p>\n\n<p>With the range of model/algorithms mentioned above, we generate 154 level 1 models. we found that almost all models add some diversity to stacking, and all efforts to trim the selection resulted in worse CV/LB, so in the end we went with including all level1 models. </p>\n\n<p>The stacking approach we used were very similar to the one that I have described in my <a href=\"https://www.kaggle.com/yifanxie/porto-seguro-tutorial-end-to-end-ensemble\">ensemble kernel</a> in the Porto competition. Except from the fact that we went for weight averaging for level 3 instead of stacking. I found that weight averaging always performed better on this level with a combination of lgb and keras level 2 oof. Didn't attempt to perform stacking beyond level 3, which we could have done - but this way we would have to start stacking a bit more earlier into the competition. </p>\n\n<p>We found that training with alternative objective functions on level 1, and retraining level 1 model by different combinations of category helped to increase the diversity.  Having discovered this we focused solely on retraining our existing models in the final day of the competition and managed to generate more than 40 models to add to our mix. This contributed a lot to our stacking effort with more 0.0005 gained in the last day. </p>",
  "messages": [
    {
      "id": "350976",
      "postDate": "06/30/2018 23:53:58",
      "content": "<p>What a competition this has been - it was such a rare opportunity to be able to tackle image, text,  tabular and time series competition in one single competition.   I personally found that the ability to iterate and execute ideas as quickly as possible is really critical in a competition like this. and it is for the same reason that I really admire all the solo participants who did well here, none more so than <a href=\"/dmitrylarko\">@dmitrylarko</a>, <a href=\"/kelexu\">@kelexu</a> and <a href=\"/wenbozhao\">@wenbozhao</a>   Despite managing to finish in the gold medal zone, it is obvious to me that the distance between us and the top teams is very big, so it is somehow humbling to see there is still so much room to improve.  We will definitely keep going to get better </p>\n\n<p>Our team names during this competition more or less reflect our journey in the last 7 weeks when we were active in the competition. First of all, 每天进步一点点means making small progress every day, while 挣扎到最后一刻 means fighting til the bitter end.  I don't think we have come across any single magical features that improve our score by more than 0.001,  while we were working frantically till the last seconds of the competition to make sure we can stay in the gold medal zone. The fight for the last gold medal positions was probably the most intensive that I have witnessed will multiple teams making sudden big jumps, so we are really happy that we manage to stay with them. </p>\n\n<p>I started the competition with <a href=\"/cxl923cc\">@cxl923cc</a>, <a href=\"/zhiqiangzhong\">@zhiqiangzhong</a> and <a href=\"/yl1202\">@yl1202</a>, we are close friends and colleagues with each other and wanted to use this exercise to form a group that can share the burden of kaggle competitions (seems it worked out well). Towards the merger deadline, we formed with <a href=\"/mzr2017\">@mzr2017</a> and <a href=\"/oyxuan\">@oyxuan</a> - because their modelling approach was quite different from ours, and the two groups complemented each other well. </p>\n\n<h2>Feature Engineering</h2>\n\n<p>I will describe the features we have been working on within our original team and let our teammate add theirs in following threads.  Most of our features are already covered by other top solutions, but anyway I will provide a list here. The following are shared by all models:</p>\n\n<ul>\n<li>text feature with TFIDF vectorizer for title, description and params,\nwe  played around with many different combination of TFIDF\nparameters, and it all added to our model diversity.</li>\n<li>SVD of TFIDF vectorizer features </li>\n<li>text statistical features such as length of text, number of number</li>\n<li>text features on ngrams, text distance features </li>\n<li>various groupby statistics between different categoricals </li>\n<li>aggregated features like the one shared in the kernels, </li>\n<li>LDA features </li>\n<li>Image features like the one that is extacted from pretrained models, as well as dullness, brightness</li>\n<li>rolling statistics such as number of ad in the same category in different time windows</li>\n</ul>\n\n<p>For non-NN models we specifically also created:</p>\n\n<ul>\n<li>rnn extracted features： rnn features extracted from rnn architecture (Bidirectional-GRU, attention, global max, global avg), and feed into non-nn model as tabular features. </li>\n<li>sentence2vec features\nboth group of features played important roles in our models. </li>\n</ul>\n\n<p>For NN model, we made heavy use of pre-trained word embedding models - three variant of FastText models shared in external data thread, and the self-train embedding model trained on all data. </p>\n\n<p>Aside from the above, we have also added the following features trained on five-fold oof manner:</p>\n\n<ul>\n<li>\"zero deal probability rate\" - five-fold oof probability prediction on if the deal probability is zero. </li>\n<li>target encoding features - again five-fold oof target encoding features for each original categorical features apart from user-id </li>\n</ul>\n\n<p>All these oof features are used directly on level 1 of the stacking, along side with other lv1 meta features (oof train and test)</p>\n\n<h2> Modelling Approach </h2>\n\n<p>We have used a wide range of models/algorithms in this competition. LightGBM with different parameters, XGB, Wordbatch FTRL_FM, Ridge, LinearSVR, RNN with Keras, Fully-connected NN with keras.</p>\n\n<p>In the last week of the competition, <a href=\"/mzr2017\">@mzr2017</a> found out that using xentropy in LightGBM, and building model by different parent categories helps to improve scores and give more diversity - and we started to retrained our better models with the corresponding settings.  We achieved our best LB model at public LB 0.2186 with xentropy training on all data - it was a model with 400+ features and some dense rnn &amp; sentence2vec features, and trained for more than 30 hours.  meanwhile Our best RNN model was around 0.2195 in public LB, courtesy of our teammates.</p>\n\n<h2> Stacking Approach </h2>\n\n<p>Stacking turned out to be very effective in this competition, and we were able to generate stacking score that is almost 0.004 better than our best l1 models. </p>\n\n<p>With the range of model/algorithms mentioned above, we generate 154 level 1 models. we found that almost all models add some diversity to stacking, and all efforts to trim the selection resulted in worse CV/LB, so in the end we went with including all level1 models. </p>\n\n<p>The stacking approach we used were very similar to the one that I have described in my <a href=\"https://www.kaggle.com/yifanxie/porto-seguro-tutorial-end-to-end-ensemble\">ensemble kernel</a> in the Porto competition. Except from the fact that we went for weight averaging for level 3 instead of stacking. I found that weight averaging always performed better on this level with a combination of lgb and keras level 2 oof. Didn't attempt to perform stacking beyond level 3, which we could have done - but this way we would have to start stacking a bit more earlier into the competition. </p>\n\n<p>We found that training with alternative objective functions on level 1, and retraining level 1 model by different combinations of category helped to increase the diversity.  Having discovered this we focused solely on retraining our existing models in the final day of the competition and managed to generate more than 40 models to add to our mix. This contributed a lot to our stacking effort with more 0.0005 gained in the last day. </p>",
      "rawMarkdown": "What a competition this has been - it was such a rare opportunity to be able to tackle image, text,  tabular and time series competition in one single competition.   I personally found that the ability to iterate and execute ideas as quickly as possible is really critical in a competition like this. and it is for the same reason that I really admire all the solo participants who did well here, none more so than @dmitrylarko, @kelexu and @wenbozhao   Despite managing to finish in the gold medal zone, it is obvious to me that the distance between us and the top teams is very big, so it is somehow humbling to see there is still so much room to improve.  We will definitely keep going to get better \n\nOur team names during this competition more or less reflect our journey in the last 7 weeks when we were active in the competition. First of all, 每天进步一点点means making small progress every day, while 挣扎到最后一刻 means fighting til the bitter end.  I don't think we have come across any single magical features that improve our score by more than 0.001,  while we were working frantically till the last seconds of the competition to make sure we can stay in the gold medal zone. The fight for the last gold medal positions was probably the most intensive that I have witnessed will multiple teams making sudden big jumps, so we are really happy that we manage to stay with them. \n\nI started the competition with @cxl923cc, @zhiqiangzhong and @yl1202, we are close friends and colleagues with each other and wanted to use this exercise to form a group that can share the burden of kaggle competitions (seems it worked out well). Towards the merger deadline, we formed with @mzr2017 and @oyxuan - because their modelling approach was quite different from ours, and the two groups complemented each other well. \n\n<h2>Feature Engineering</h2>\nI will describe the features we have been working on within our original team and let our teammate add theirs in following threads.  Most of our features are already covered by other top solutions, but anyway I will provide a list here. The following are shared by all models:\n\n - text feature with TFIDF vectorizer for title, description and params,\n   we  played around with many different combination of TFIDF\n   parameters, and it all added to our model diversity.\n - SVD of TFIDF vectorizer features \n - text statistical features such as length of text, number of number\n - text features on ngrams, text distance features \n -  various groupby statistics between different categoricals \n - aggregated features like the one shared in the kernels, \n - LDA features \n - Image features like the one that is extacted from pretrained models, as well as dullness, brightness\n - rolling statistics such as number of ad in the same category in different time windows\n\nFor non-NN models we specifically also created:\n\n - rnn extracted features： rnn features extracted from rnn architecture (Bidirectional-GRU, attention, global max, global avg), and feed into non-nn model as tabular features. \n - sentence2vec features\nboth group of features played important roles in our models. \n\nFor NN model, we made heavy use of pre-trained word embedding models - three variant of FastText models shared in external data thread, and the self-train embedding model trained on all data. \n\nAside from the above, we have also added the following features trained on five-fold oof manner:\n\n -  \"zero deal probability rate\" - five-fold oof probability prediction on if the deal probability is zero. \n - target encoding features - again five-fold oof target encoding features for each original categorical features apart from user-id \n\nAll these oof features are used directly on level 1 of the stacking, along side with other lv1 meta features (oof train and test)\n\n<h2> Modelling Approach </h2>\n\nWe have used a wide range of models/algorithms in this competition. LightGBM with different parameters, XGB, Wordbatch FTRL_FM, Ridge, LinearSVR, RNN with Keras, Fully-connected NN with keras.\n\nIn the last week of the competition, @mzr2017 found out that using xentropy in LightGBM, and building model by different parent categories helps to improve scores and give more diversity - and we started to retrained our better models with the corresponding settings.  We achieved our best LB model at public LB 0.2186 with xentropy training on all data - it was a model with 400+ features and some dense rnn &amp; sentence2vec features, and trained for more than 30 hours.  meanwhile Our best RNN model was around 0.2195 in public LB, courtesy of our teammates.\n\n<h2> Stacking Approach </h2>\nStacking turned out to be very effective in this competition, and we were able to generate stacking score that is almost 0.004 better than our best l1 models. \n\nWith the range of model/algorithms mentioned above, we generate 154 level 1 models. we found that almost all models add some diversity to stacking, and all efforts to trim the selection resulted in worse CV/LB, so in the end we went with including all level1 models. \n\nThe stacking approach we used were very similar to the one that I have described in my [ensemble kernel][1] in the Porto competition. Except from the fact that we went for weight averaging for level 3 instead of stacking. I found that weight averaging always performed better on this level with a combination of lgb and keras level 2 oof. Didn't attempt to perform stacking beyond level 3, which we could have done - but this way we would have to start stacking a bit more earlier into the competition. \n\nWe found that training with alternative objective functions on level 1, and retraining level 1 model by different combinations of category helped to increase the diversity.  Having discovered this we focused solely on retraining our existing models in the final day of the competition and managed to generate more than 40 models to add to our mix. This contributed a lot to our stacking effort with more 0.0005 gained in the last day. \n\n\n  [1]: https://www.kaggle.com/yifanxie/porto-seguro-tutorial-end-to-end-ensemble",
      "votes": null
    },
    {
      "id": "351022",
      "postDate": "07/01/2018 04:38:29",
      "content": "<p>Just simply astonishing and don't forget the team effort.</p>",
      "rawMarkdown": "Just simply astonishing and don't forget the team effort.",
      "votes": null
    },
    {
      "id": "351249",
      "postDate": "07/01/2018 17:49:31",
      "content": "<p>Congratulations to you and your team ;) Your progression has been steady and really amazing from what I could see. Well done and happy you got a well deserved gold medal !</p>",
      "rawMarkdown": "Congratulations to you and your team ;) Your progression has been steady and really amazing from what I could see. Well done and happy you got a well deserved gold medal !",
      "votes": null
    },
    {
      "id": "351302",
      "postDate": "07/01/2018 21:39:34",
      "content": "<p><a href=\"/ogrellier\">@ogrellier</a> thank you for your kind words, a large part of the success was built on our previous teaming experience, so some of the credits must also go to you. Hopefully in the not so distant future we find some common space &amp; time to work on another competition again.</p>",
      "rawMarkdown": "ogrellier thank you for your kind words, a large part of the success was built on our previous teaming experience, so some of the credits must also go to you. Hopefully in the not so distant future we find some common space &amp; time to work on another competition again.",
      "votes": null
    },
    {
      "id": "351308",
      "postDate": "07/01/2018 22:00:08",
      "content": "<p>Congratulations !!!. <br>\nI was not clear the stacking until I see your awesome work (Thanks for your link!). <br>\nI wonder why we need so many stacking layers as other team did and how many layers is the best choice?  </p>",
      "rawMarkdown": "Congratulations !!!.  \nI was not clear the stacking until I see your awesome work (Thanks for your link!).  \nI wonder why we need so many stacking layers as other team did and how many layers is the best choice?",
      "votes": null
    },
    {
      "id": "351520",
      "postDate": "07/02/2018 12:09:55",
      "content": "<p><a href=\"/backaggle\">@backaggle</a> thanks and glad that my kernel helps.</p>\n\n<p>In theory, you can stack as many layers as you can, but personally I never tried to go beyond level 3. As you see the 2nd team did a 6-layer stacking which is quite an eye opener form me :) </p>\n\n<p>From my point of view, if you want to do multilayers stacking,  you would need to spend time and work out a good stacking workflow for you - so that you can generate oof predictions on different layers &amp; different algorithms, but at the same time manage the large amount of files generated </p>",
      "rawMarkdown": "backaggle thanks and glad that my kernel helps.\n\nIn theory, you can stack as many layers as you can, but personally I never tried to go beyond level 3. As you see the 2nd team did a 6-layer stacking which is quite an eye opener form me :) \n\nFrom my point of view, if you want to do multilayers stacking,  you would need to spend time and work out a good stacking workflow for you - so that you can generate oof predictions on different layers &amp; different algorithms, but at the same time manage the large amount of files generated",
      "votes": null
    },
    {
      "id": "352843",
      "postDate": "07/05/2018 08:42:47",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 351022,
      "author_name": "piyush1912",
      "author_url": "",
      "post_date": "07/01/2018 04:38:29",
      "content": "<p>Just simply astonishing and don't forget the team effort.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 351249,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "07/01/2018 17:49:31",
      "content": "<p>Congratulations to you and your team ;) Your progression has been steady and really amazing from what I could see. Well done and happy you got a well deserved gold medal !</p>",
      "votes": null,
      "replies": [
        {
          "id": 351302,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/01/2018 21:39:34",
          "content": "<p><a href=\"/ogrellier\">@ogrellier</a> thank you for your kind words, a large part of the success was built on our previous teaming experience, so some of the credits must also go to you. Hopefully in the not so distant future we find some common space &amp; time to work on another competition again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 351308,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "07/01/2018 22:00:08",
      "content": "<p>Congratulations !!!. <br>\nI was not clear the stacking until I see your awesome work (Thanks for your link!). <br>\nI wonder why we need so many stacking layers as other team did and how many layers is the best choice?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 351520,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/02/2018 12:09:55",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> thanks and glad that my kernel helps.</p>\n\n<p>In theory, you can stack as many layers as you can, but personally I never tried to go beyond level 3. As you see the 2nd team did a 6-layer stacking which is quite an eye opener form me :) </p>\n\n<p>From my point of view, if you want to do multilayers stacking,  you would need to spend time and work out a good stacking workflow for you - so that you can generate oof predictions on different layers &amp; different algorithms, but at the same time manage the large amount of files generated </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 352843,
      "author_name": "akshitarora9211",
      "author_url": "",
      "post_date": "07/05/2018 08:42:47",
      "content": "<p>Congratulations!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "350976": "What a competition this has been - it was such a rare opportunity to be able to tackle image, text,  tabular and time series competition in one single competition.   I personally found that the ability to iterate and execute ideas as quickly as possible is really critical in a competition like this. and it is for the same reason that I really admire all the solo participants who did well here, none more so than @dmitrylarko, @kelexu and @wenbozhao   Despite managing to finish in the gold medal zone, it is obvious to me that the distance between us and the top teams is very big, so it is somehow humbling to see there is still so much room to improve.  We will definitely keep going to get better \n\nOur team names during this competition more or less reflect our journey in the last 7 weeks when we were active in the competition. First of all, 每天进步一点点means making small progress every day, while 挣扎到最后一刻 means fighting til the bitter end.  I don't think we have come across any single magical features that improve our score by more than 0.001,  while we were working frantically till the last seconds of the competition to make sure we can stay in the gold medal zone. The fight for the last gold medal positions was probably the most intensive that I have witnessed will multiple teams making sudden big jumps, so we are really happy that we manage to stay with them. \n\nI started the competition with @cxl923cc, @zhiqiangzhong and @yl1202, we are close friends and colleagues with each other and wanted to use this exercise to form a group that can share the burden of kaggle competitions (seems it worked out well). Towards the merger deadline, we formed with @mzr2017 and @oyxuan - because their modelling approach was quite different from ours, and the two groups complemented each other well. \n\n<h2>Feature Engineering</h2>\nI will describe the features we have been working on within our original team and let our teammate add theirs in following threads.  Most of our features are already covered by other top solutions, but anyway I will provide a list here. The following are shared by all models:\n\n - text feature with TFIDF vectorizer for title, description and params,\n   we  played around with many different combination of TFIDF\n   parameters, and it all added to our model diversity.\n - SVD of TFIDF vectorizer features \n - text statistical features such as length of text, number of number\n - text features on ngrams, text distance features \n -  various groupby statistics between different categoricals \n - aggregated features like the one shared in the kernels, \n - LDA features \n - Image features like the one that is extacted from pretrained models, as well as dullness, brightness\n - rolling statistics such as number of ad in the same category in different time windows\n\nFor non-NN models we specifically also created:\n\n - rnn extracted features： rnn features extracted from rnn architecture (Bidirectional-GRU, attention, global max, global avg), and feed into non-nn model as tabular features. \n - sentence2vec features\nboth group of features played important roles in our models. \n\nFor NN model, we made heavy use of pre-trained word embedding models - three variant of FastText models shared in external data thread, and the self-train embedding model trained on all data. \n\nAside from the above, we have also added the following features trained on five-fold oof manner:\n\n -  \"zero deal probability rate\" - five-fold oof probability prediction on if the deal probability is zero. \n - target encoding features - again five-fold oof target encoding features for each original categorical features apart from user-id \n\nAll these oof features are used directly on level 1 of the stacking, along side with other lv1 meta features (oof train and test)\n\n<h2> Modelling Approach </h2>\n\nWe have used a wide range of models/algorithms in this competition. LightGBM with different parameters, XGB, Wordbatch FTRL_FM, Ridge, LinearSVR, RNN with Keras, Fully-connected NN with keras.\n\nIn the last week of the competition, @mzr2017 found out that using xentropy in LightGBM, and building model by different parent categories helps to improve scores and give more diversity - and we started to retrained our better models with the corresponding settings.  We achieved our best LB model at public LB 0.2186 with xentropy training on all data - it was a model with 400+ features and some dense rnn &amp; sentence2vec features, and trained for more than 30 hours.  meanwhile Our best RNN model was around 0.2195 in public LB, courtesy of our teammates.\n\n<h2> Stacking Approach </h2>\nStacking turned out to be very effective in this competition, and we were able to generate stacking score that is almost 0.004 better than our best l1 models. \n\nWith the range of model/algorithms mentioned above, we generate 154 level 1 models. we found that almost all models add some diversity to stacking, and all efforts to trim the selection resulted in worse CV/LB, so in the end we went with including all level1 models. \n\nThe stacking approach we used were very similar to the one that I have described in my [ensemble kernel][1] in the Porto competition. Except from the fact that we went for weight averaging for level 3 instead of stacking. I found that weight averaging always performed better on this level with a combination of lgb and keras level 2 oof. Didn't attempt to perform stacking beyond level 3, which we could have done - but this way we would have to start stacking a bit more earlier into the competition. \n\nWe found that training with alternative objective functions on level 1, and retraining level 1 model by different combinations of category helped to increase the diversity.  Having discovered this we focused solely on retraining our existing models in the final day of the competition and managed to generate more than 40 models to add to our mix. This contributed a lot to our stacking effort with more 0.0005 gained in the last day. \n\n\n  [1]: https://www.kaggle.com/yifanxie/porto-seguro-tutorial-end-to-end-ensemble",
    "351022": "Just simply astonishing and don't forget the team effort.",
    "351249": "Congratulations to you and your team ;) Your progression has been steady and really amazing from what I could see. Well done and happy you got a well deserved gold medal !",
    "351302": "ogrellier thank you for your kind words, a large part of the success was built on our previous teaming experience, so some of the credits must also go to you. Hopefully in the not so distant future we find some common space &amp; time to work on another competition again.",
    "351308": "Congratulations !!!.  \nI was not clear the stacking until I see your awesome work (Thanks for your link!).  \nI wonder why we need so many stacking layers as other team did and how many layers is the best choice?",
    "351520": "backaggle thanks and glad that my kernel helps.\n\nIn theory, you can stack as many layers as you can, but personally I never tried to go beyond level 3. As you see the 2nd team did a 6-layer stacking which is quite an eye opener form me :) \n\nFrom my point of view, if you want to do multilayers stacking,  you would need to spend time and work out a good stacking workflow for you - so that you can generate oof predictions on different layers &amp; different algorithms, but at the same time manage the large amount of files generated",
    "352843": "Congratulations!!"
  },
  "source": "meta"
}