{
  "id": 60059,
  "title": "14th Place Solution: The Almost Golden Defenders",
  "url": "/competitions/avito-demand-prediction/writeups/the-almost-golden-defenders-14th-place-solution-th",
  "author_name": "",
  "post_date": "2018-07-01T02:19:09.357Z",
  "votes": 41,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This competition was very overwhelming. I’ve never seen a Kaggle competition before where we’ve had to intelligently deal with images, text, geospatial information, unlabeled additional data, numeric data, categorical data, and a low-key time series problem to boot! Not to mention that we still don’t even know what the dependent variable is and a few of the independent variables remain a mystery as well.</p>\n\n<p>Congratulations to everyone who got a gold medal -- especially the soloists. In my opinion, anyone who got a solo gold in this competition deserves two gold medals. I have no idea how you’d even manage something like that given everything this competition has to offer.</p>\n\n<p>We’re really excited that we pushed way farther than we ever thought would be possible in the final weeks, but are really astonished by the caliber of the competition that we faced. I’ve never seen such a competitive final week. ...If only there were a few more gold medal slots, as there certainly were quite a few deserving teams.</p>\n\n<p>Naturally, spending most of the competition in the gold medal zone only to lose out in the end by a smidge is a humbling experience. It leads us to second guess everything we did and wonder where the extra few points could have come from that could have made the difference.</p>\n\n<p>Anyways, I’ll stop complaining and try to make up for our second guessing with an instructive write-up.</p>\n\n<h2>The Stacker</h2>\n\n<p>I’m too lazy to make a pretty diagram, but Sijun was not too lazy, <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/60059#350998\">so you can read more detail here</a>. The original plan was to stop at Level 3 with just a single LightGBM, but the competitiveness of the competition forced us to keep going to level 4. Here's some brief discussion, though see Sijun's additional analysis.</p>\n\n<p><strong>Level 4</strong></p>\n\n<p>We created a four-level model, using a special Lasso model to aggregate together submodels. The two key features of this Lasso model was that the weights were constrained to be only positive and the model was trained separately for each <code>parent_category_name</code>. We found that <code>parent_category_name</code> introduced a good degree of variability in the errors and capturing this by building models separately was good for a very small (we’re talking ~0.00002) boost in score compared to training just one Lasso overall. Creating this insane Level 4 stacking system was key to boosting us from 0.2149 LB to 0.2148 LB on the public leaderboard.</p>\n\n<p><strong>Level 3</strong></p>\n\n<p>On the Level 3, we had 3x LightGBMs trained using (1) the L2 models and a few key features (parent_category_name, price, 20-dimensional SVD of text embeddings and a few image features), (2) all the L2 models and all the features used in L1, and (3) the same as the first model but trained with a Poisson objective. We also trained an MLP, a Lasso, and the aforementioned grouped Lasso by <code>parent_category_name</code> on all the L2 models (but no other features). Each of these models were also lazily bagged with some previous versions of the model that did not have some of our final models, for a total of 13 L3 models (7 LightGBM, 4 MLP, 2 Lasso). The LightGBM stackers each got ~0.2149 on the public LB and the Lasso ones each got ~0.2153. This was a large jump from Level 2 and Level 1.</p>\n\n<p><strong>Level 2</strong></p>\n\n<p>There was a small bump between Level 1 and Level 3 where we added all of our Ridge models to our LightGBM combined with all of our features. We trained this with both normal (regression) and Poisson objectives. We also lazily bagged this model for a total of four copies (3 regression,1 Poisson) by adding in previous copies of the model that didn’t have some of our final features. Our best Level 2 model scored 0.2173 on the public LB.</p>\n\n<p><strong>Level 1</strong></p>\n\n<p>We trained a ton of features into LightGBMs and varied the kind of encoding for the categoricals (either one hot encoding, LightGBM’s built-in target encoding, or Bayesian target encoding), varied the objective (regression and Poisson), and did some lazy bagging (using prior copies of the model that didn’t have all the features) for a total of 19 LightGBMs. Our best LightGBM scored 0.2183 on the public LB.</p>\n\n<p>On the NN front, we trained three CNNs with FastText and two RNNs with Attention and pooling. Our best NN we 0.2181 on the public LB. <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59917\">You can see how our best NN looked here.</a> We also diversified our Level 3 blend by adding two NNs trained with multiclass and one on binary objectives, which provided a good boost.</p>\n\n<p>We also trained a large number of Ridge models that were varied on training just on TFIDF for the title, just on TFIDF for the description, just on character-level TFIDF for the description, and on all text (concat of all TFIDFs). We then trained these models individually for each parent_category_name, cat_bin (described below), and parent_category_name - region interaction. We also trained several Ridges on all the data and included interaction terms in one. All told, there were 17 Ridge models. All of these Ridges (but no other L1 models) formed the basis of Level 2.</p>\n\n<p>Lastly, we trained three FM models, one using Wordbatch and two using TFFM. These didn’t end up outperforming our best Ridge models.</p>\n\n<h2>Features</h2>\n\n<p>Our feature engineering work was very exhaustive and tended toward a kitchen sink. In the beginning, we would diligently add each feature into the model one-by-one and keep it if it improved the first fold. Toward the end of the competition, we ran out of time for that, so we would just dump tons of features in at once. The one-by-one approach is nice because frequently there were features that looked good but ended up making the model worse.</p>\n\n<p>We added a lot of features, so we’ll be posting separately sometime next week about everything we did.</p>\n\n<h2>Code Sharing</h2>\n\n<p>We’re still cleaning the code and getting everything together, and hope to fully share all our code sometime next week.</p>\n\n<h2>Until Next Time (Very Soon)</h2>\n\n<p>See you guys more next week. I look forward to continuing to digest all the lessons learned and plotting my steps toward a real gold medal. We have some unfinished business. :)</p>\n\n<p><img src=\"https://i.imgur.com/Qj8kqVY.jpg\" alt=\"missed it by that much\"></p>",
  "messages": [
    {
      "id": "350426",
      "postDate": "06/29/2018 18:30:31",
      "content": "<p>This competition was very overwhelming. I’ve never seen a Kaggle competition before where we’ve had to intelligently deal with images, text, geospatial information, unlabeled additional data, numeric data, categorical data, and a low-key time series problem to boot! Not to mention that we still don’t even know what the dependent variable is and a few of the independent variables remain a mystery as well.</p>\n\n<p>Congratulations to everyone who got a gold medal -- especially the soloists. In my opinion, anyone who got a solo gold in this competition deserves two gold medals. I have no idea how you’d even manage something like that given everything this competition has to offer.</p>\n\n<p>We’re really excited that we pushed way farther than we ever thought would be possible in the final weeks, but are really astonished by the caliber of the competition that we faced. I’ve never seen such a competitive final week. ...If only there were a few more gold medal slots, as there certainly were quite a few deserving teams.</p>\n\n<p>Naturally, spending most of the competition in the gold medal zone only to lose out in the end by a smidge is a humbling experience. It leads us to second guess everything we did and wonder where the extra few points could have come from that could have made the difference.</p>\n\n<p>Anyways, I’ll stop complaining and try to make up for our second guessing with an instructive write-up.</p>\n\n<h2>The Stacker</h2>\n\n<p>I’m too lazy to make a pretty diagram, but Sijun was not too lazy, <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/60059#350998\">so you can read more detail here</a>. The original plan was to stop at Level 3 with just a single LightGBM, but the competitiveness of the competition forced us to keep going to level 4. Here's some brief discussion, though see Sijun's additional analysis.</p>\n\n<p><strong>Level 4</strong></p>\n\n<p>We created a four-level model, using a special Lasso model to aggregate together submodels. The two key features of this Lasso model was that the weights were constrained to be only positive and the model was trained separately for each <code>parent_category_name</code>. We found that <code>parent_category_name</code> introduced a good degree of variability in the errors and capturing this by building models separately was good for a very small (we’re talking ~0.00002) boost in score compared to training just one Lasso overall. Creating this insane Level 4 stacking system was key to boosting us from 0.2149 LB to 0.2148 LB on the public leaderboard.</p>\n\n<p><strong>Level 3</strong></p>\n\n<p>On the Level 3, we had 3x LightGBMs trained using (1) the L2 models and a few key features (parent_category_name, price, 20-dimensional SVD of text embeddings and a few image features), (2) all the L2 models and all the features used in L1, and (3) the same as the first model but trained with a Poisson objective. We also trained an MLP, a Lasso, and the aforementioned grouped Lasso by <code>parent_category_name</code> on all the L2 models (but no other features). Each of these models were also lazily bagged with some previous versions of the model that did not have some of our final models, for a total of 13 L3 models (7 LightGBM, 4 MLP, 2 Lasso). The LightGBM stackers each got ~0.2149 on the public LB and the Lasso ones each got ~0.2153. This was a large jump from Level 2 and Level 1.</p>\n\n<p><strong>Level 2</strong></p>\n\n<p>There was a small bump between Level 1 and Level 3 where we added all of our Ridge models to our LightGBM combined with all of our features. We trained this with both normal (regression) and Poisson objectives. We also lazily bagged this model for a total of four copies (3 regression,1 Poisson) by adding in previous copies of the model that didn’t have some of our final features. Our best Level 2 model scored 0.2173 on the public LB.</p>\n\n<p><strong>Level 1</strong></p>\n\n<p>We trained a ton of features into LightGBMs and varied the kind of encoding for the categoricals (either one hot encoding, LightGBM’s built-in target encoding, or Bayesian target encoding), varied the objective (regression and Poisson), and did some lazy bagging (using prior copies of the model that didn’t have all the features) for a total of 19 LightGBMs. Our best LightGBM scored 0.2183 on the public LB.</p>\n\n<p>On the NN front, we trained three CNNs with FastText and two RNNs with Attention and pooling. Our best NN we 0.2181 on the public LB. <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59917\">You can see how our best NN looked here.</a> We also diversified our Level 3 blend by adding two NNs trained with multiclass and one on binary objectives, which provided a good boost.</p>\n\n<p>We also trained a large number of Ridge models that were varied on training just on TFIDF for the title, just on TFIDF for the description, just on character-level TFIDF for the description, and on all text (concat of all TFIDFs). We then trained these models individually for each parent_category_name, cat_bin (described below), and parent_category_name - region interaction. We also trained several Ridges on all the data and included interaction terms in one. All told, there were 17 Ridge models. All of these Ridges (but no other L1 models) formed the basis of Level 2.</p>\n\n<p>Lastly, we trained three FM models, one using Wordbatch and two using TFFM. These didn’t end up outperforming our best Ridge models.</p>\n\n<h2>Features</h2>\n\n<p>Our feature engineering work was very exhaustive and tended toward a kitchen sink. In the beginning, we would diligently add each feature into the model one-by-one and keep it if it improved the first fold. Toward the end of the competition, we ran out of time for that, so we would just dump tons of features in at once. The one-by-one approach is nice because frequently there were features that looked good but ended up making the model worse.</p>\n\n<p>We added a lot of features, so we’ll be posting separately sometime next week about everything we did.</p>\n\n<h2>Code Sharing</h2>\n\n<p>We’re still cleaning the code and getting everything together, and hope to fully share all our code sometime next week.</p>\n\n<h2>Until Next Time (Very Soon)</h2>\n\n<p>See you guys more next week. I look forward to continuing to digest all the lessons learned and plotting my steps toward a real gold medal. We have some unfinished business. :)</p>\n\n<p><img src=\"https://i.imgur.com/Qj8kqVY.jpg\" alt=\"missed it by that much\"></p>",
      "rawMarkdown": "This competition was very overwhelming. I’ve never seen a Kaggle competition before where we’ve had to intelligently deal with images, text, geospatial information, unlabeled additional data, numeric data, categorical data, and a low-key time series problem to boot! Not to mention that we still don’t even know what the dependent variable is and a few of the independent variables remain a mystery as well.\n\nCongratulations to everyone who got a gold medal -- especially the soloists. In my opinion, anyone who got a solo gold in this competition deserves two gold medals. I have no idea how you’d even manage something like that given everything this competition has to offer.\n\nWe’re really excited that we pushed way farther than we ever thought would be possible in the final weeks, but are really astonished by the caliber of the competition that we faced. I’ve never seen such a competitive final week. ...If only there were a few more gold medal slots, as there certainly were quite a few deserving teams.\n\nNaturally, spending most of the competition in the gold medal zone only to lose out in the end by a smidge is a humbling experience. It leads us to second guess everything we did and wonder where the extra few points could have come from that could have made the difference.\n \nAnyways, I’ll stop complaining and try to make up for our second guessing with an instructive write-up.\n\n\nThe Stacker\n-----------\n\nI’m too lazy to make a pretty diagram, but Sijun was not too lazy, [so you can read more detail here](https://www.kaggle.com/c/avito-demand-prediction/discussion/60059#350998). The original plan was to stop at Level 3 with just a single LightGBM, but the competitiveness of the competition forced us to keep going to level 4. Here's some brief discussion, though see Sijun's additional analysis.\n\n\n**Level 4**\n\nWe created a four-level model, using a special Lasso model to aggregate together submodels. The two key features of this Lasso model was that the weights were constrained to be only positive and the model was trained separately for each `parent_category_name`. We found that `parent_category_name` introduced a good degree of variability in the errors and capturing this by building models separately was good for a very small (we’re talking ~0.00002) boost in score compared to training just one Lasso overall. Creating this insane Level 4 stacking system was key to boosting us from 0.2149 LB to 0.2148 LB on the public leaderboard.\n\n\n**Level 3**\n\nOn the Level 3, we had 3x LightGBMs trained using (1) the L2 models and a few key features (parent_category_name, price, 20-dimensional SVD of text embeddings and a few image features), (2) all the L2 models and all the features used in L1, and (3) the same as the first model but trained with a Poisson objective. We also trained an MLP, a Lasso, and the aforementioned grouped Lasso by `parent_category_name` on all the L2 models (but no other features). Each of these models were also lazily bagged with some previous versions of the model that did not have some of our final models, for a total of 13 L3 models (7 LightGBM, 4 MLP, 2 Lasso). The LightGBM stackers each got ~0.2149 on the public LB and the Lasso ones each got ~0.2153. This was a large jump from Level 2 and Level 1.\n\n\n**Level 2**\n\nThere was a small bump between Level 1 and Level 3 where we added all of our Ridge models to our LightGBM combined with all of our features. We trained this with both normal (regression) and Poisson objectives. We also lazily bagged this model for a total of four copies (3 regression,1 Poisson) by adding in previous copies of the model that didn’t have some of our final features. Our best Level 2 model scored 0.2173 on the public LB.\n\n\n**Level 1**\n\nWe trained a ton of features into LightGBMs and varied the kind of encoding for the categoricals (either one hot encoding, LightGBM’s built-in target encoding, or Bayesian target encoding), varied the objective (regression and Poisson), and did some lazy bagging (using prior copies of the model that didn’t have all the features) for a total of 19 LightGBMs. Our best LightGBM scored 0.2183 on the public LB.\n\nOn the NN front, we trained three CNNs with FastText and two RNNs with Attention and pooling. Our best NN we 0.2181 on the public LB. [You can see how our best NN looked here.](https://www.kaggle.com/c/avito-demand-prediction/discussion/59917) We also diversified our Level 3 blend by adding two NNs trained with multiclass and one on binary objectives, which provided a good boost.\n\nWe also trained a large number of Ridge models that were varied on training just on TFIDF for the title, just on TFIDF for the description, just on character-level TFIDF for the description, and on all text (concat of all TFIDFs). We then trained these models individually for each parent_category_name, cat_bin (described below), and parent_category_name - region interaction. We also trained several Ridges on all the data and included interaction terms in one. All told, there were 17 Ridge models. All of these Ridges (but no other L1 models) formed the basis of Level 2.\n\nLastly, we trained three FM models, one using Wordbatch and two using TFFM. These didn’t end up outperforming our best Ridge models.\n\n\nFeatures\n-----------\n\nOur feature engineering work was very exhaustive and tended toward a kitchen sink. In the beginning, we would diligently add each feature into the model one-by-one and keep it if it improved the first fold. Toward the end of the competition, we ran out of time for that, so we would just dump tons of features in at once. The one-by-one approach is nice because frequently there were features that looked good but ended up making the model worse.\n\nWe added a lot of features, so we’ll be posting separately sometime next week about everything we did.\n\n\nCode Sharing\n----------------\n\nWe’re still cleaning the code and getting everything together, and hope to fully share all our code sometime next week.\n\n\nUntil Next Time (Very Soon)\n---------------------------------\n\nSee you guys more next week. I look forward to continuing to digest all the lessons learned and plotting my steps toward a real gold medal. We have some unfinished business. :)\n\n![missed it by that much][1]\n\n\n  [1]: https://i.imgur.com/Qj8kqVY.jpg",
      "votes": null
    },
    {
      "id": "350438",
      "postDate": "06/29/2018 18:57:01",
      "content": "<p>Thanks for your sharing!</p>",
      "rawMarkdown": "Thanks for your sharing!",
      "votes": null
    },
    {
      "id": "350503",
      "postDate": "06/29/2018 21:31:05",
      "content": "<p><a href=\"/peterhurford\">@peterhurford</a> it was very tight, very very tight finishing. We also spent the last hours frantically pushing our score to stay within the gold medal zone.  Hopefully I (or my team mates) will write something up in the next few days - was brain dead for two days and now trying to catch up on too many things in work and life!</p>\n\n<p>Definitely look forward to seeing more of you around the LB, thanks for all the sharing! </p>",
      "rawMarkdown": "peterhurford it was very tight, very very tight finishing. We also spent the last hours frantically pushing our score to stay within the gold medal zone.  Hopefully I (or my team mates) will write something up in the next few days - was brain dead for two days and now trying to catch up on too many things in work and life!\n\nDefinitely look forward to seeing more of you around the LB, thanks for all the sharing!",
      "votes": null
    },
    {
      "id": "350533",
      "postDate": "06/29/2018 22:32:19",
      "content": "<p>Thanks. We'll share even more soon!</p>",
      "rawMarkdown": "Thanks. We'll share even more soon!",
      "votes": null
    },
    {
      "id": "350656",
      "postDate": "06/30/2018 07:26:07",
      "content": "<p>@Peter, very sorry you missed it :( Thanks for sharing this amazing stacking tower ! You'll surely become a competition master very soon !</p>",
      "rawMarkdown": "Peter, very sorry you missed it :( Thanks for sharing this amazing stacking tower ! You'll surely become a competition master very soon !",
      "votes": null
    },
    {
      "id": "350684",
      "postDate": "06/30/2018 08:24:32",
      "content": "<p>Thanks. Stacking towers are my favorite part of Kaggle. I look forward to stacking toward a gold after a few months off.</p>",
      "rawMarkdown": "Thanks. Stacking towers are my favorite part of Kaggle. I look forward to stacking toward a gold after a few months off.",
      "votes": null
    },
    {
      "id": "350998",
      "postDate": "07/01/2018 01:33:44",
      "content": "<p>More details about our blender tower. \n<img src=\"https://s3-us-west-1.amazonaws.com/sijunhe-blog/plots/post17/stacking_tower.png\" alt=\"stacking\"></p>\n\n<p>Until the last 4 days, our plan has always been building/improving base (L1) models and stacking them with a single L2 lgb model. This got us to public LB 0.2155 but we weren’t able to push beyond that. The level of competition around us was incredible and teams kept surpassing us. We then thought of the <a href=\"https://www.slideshare.net/jeongyoonlee/winning-data-science-competitions-74391113\">winning solution from KDD Cup 2015</a>, where Jeong-yoon Lee came up with a 3-level stacking tower. We built more L2 models and a few Lasso for L3 and pushed our score to 0.2148. Unfortunately we missed the gold zone, but we did much better than we thought we could.</p>\n\n<p>There are a few characteristics of this competition that gave us the confidence to do this crazy stacking:</p>\n\n<ul>\n<li>CV reflects LB very well. For us, a decrease in CV ALWAYS lead to a decrease in LB</li>\n<li>The CV-LB gap is also fairly consistent for us, though the gap differs a bit for different models (gap for LGB is 0.004x, NN is 0.002x and some of the blends are 0.005x). </li>\n<li>Other than target encoding, overfitting never seems to be a issue. In most of our models, regularization doesn’t play a big part.</li>\n</ul>",
      "rawMarkdown": "More details about our blender tower. \n![stacking](https://s3-us-west-1.amazonaws.com/sijunhe-blog/plots/post17/stacking_tower.png)\n\nUntil the last 4 days, our plan has always been building/improving base (L1) models and stacking them with a single L2 lgb model. This got us to public LB 0.2155 but we weren’t able to push beyond that. The level of competition around us was incredible and teams kept surpassing us. We then thought of the [winning solution from KDD Cup 2015](https://www.slideshare.net/jeongyoonlee/winning-data-science-competitions-74391113), where Jeong-yoon Lee came up with a 3-level stacking tower. We built more L2 models and a few Lasso for L3 and pushed our score to 0.2148. Unfortunately we missed the gold zone, but we did much better than we thought we could.\n\nThere are a few characteristics of this competition that gave us the confidence to do this crazy stacking:\n\n- CV reflects LB very well. For us, a decrease in CV ALWAYS lead to a decrease in LB\n- The CV-LB gap is also fairly consistent for us, though the gap differs a bit for different models (gap for LGB is 0.004x, NN is 0.002x and some of the blends are 0.005x). \n- Other than target encoding, overfitting never seems to be a issue. In most of our models, regularization doesn’t play a big part.",
      "votes": null
    },
    {
      "id": "353086",
      "postDate": "07/05/2018 21:42:52",
      "content": "<p>One of my main contributions to the team was Bayesian target encoding.  The idea is to use bayesian statistics to encode the categorical variables.   The cool thing is that we can encode not only the target mean, but other statistics like the median, mode, variance, skewness, and kurtosis using the same framework.   We found that this style of target encoding outperforms the built-in LightGBM categorical encoding. </p>\n\n<p><img src=\"https://image.ibb.co/jcvTRd/Screenshot_from_2018_07_05_11_29_33.png\" alt=\"enter image description here\"></p>\n\n<p>Here are some links with more information. </p>\n\n<ul>\n<li><a href=\"https://mattmotoki.github.io/beta-target-encoding.html\">A more detailed write-up here</a>. </li>\n<li><a href=\"https://www.kaggle.com/mmotoki/avito-target-encoding\">Code for implementing bayesian target encoding and getting the results for reproducing the graph</a>.</li>\n</ul>",
      "rawMarkdown": "One of my main contributions to the team was Bayesian target encoding.  The idea is to use bayesian statistics to encode the categorical variables.   The cool thing is that we can encode not only the target mean, but other statistics like the median, mode, variance, skewness, and kurtosis using the same framework.   We found that this style of target encoding outperforms the built-in LightGBM categorical encoding. \n\n![enter image description here][1]\n\nHere are some links with more information. \n\n* [A more detailed write-up here][2]. \n* [Code for implementing bayesian target encoding and getting the results for reproducing the graph][3].\n\n\n\n  [1]: https://image.ibb.co/jcvTRd/Screenshot_from_2018_07_05_11_29_33.png\n  [2]: https://mattmotoki.github.io/beta-target-encoding.html\n  [3]: https://www.kaggle.com/mmotoki/avito-target-encoding",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 350438,
      "author_name": "andrew60909",
      "author_url": "",
      "post_date": "06/29/2018 18:57:01",
      "content": "<p>Thanks for your sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 350503,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "06/29/2018 21:31:05",
      "content": "<p><a href=\"/peterhurford\">@peterhurford</a> it was very tight, very very tight finishing. We also spent the last hours frantically pushing our score to stay within the gold medal zone.  Hopefully I (or my team mates) will write something up in the next few days - was brain dead for two days and now trying to catch up on too many things in work and life!</p>\n\n<p>Definitely look forward to seeing more of you around the LB, thanks for all the sharing! </p>",
      "votes": null,
      "replies": [
        {
          "id": 350533,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "06/29/2018 22:32:19",
          "content": "<p>Thanks. We'll share even more soon!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350656,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "06/30/2018 07:26:07",
      "content": "<p>@Peter, very sorry you missed it :( Thanks for sharing this amazing stacking tower ! You'll surely become a competition master very soon !</p>",
      "votes": null,
      "replies": [
        {
          "id": 350684,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "06/30/2018 08:24:32",
          "content": "<p>Thanks. Stacking towers are my favorite part of Kaggle. I look forward to stacking toward a gold after a few months off.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350998,
      "author_name": "sijunhe9248",
      "author_url": "",
      "post_date": "07/01/2018 01:33:44",
      "content": "<p>More details about our blender tower. \n<img src=\"https://s3-us-west-1.amazonaws.com/sijunhe-blog/plots/post17/stacking_tower.png\" alt=\"stacking\"></p>\n\n<p>Until the last 4 days, our plan has always been building/improving base (L1) models and stacking them with a single L2 lgb model. This got us to public LB 0.2155 but we weren’t able to push beyond that. The level of competition around us was incredible and teams kept surpassing us. We then thought of the <a href=\"https://www.slideshare.net/jeongyoonlee/winning-data-science-competitions-74391113\">winning solution from KDD Cup 2015</a>, where Jeong-yoon Lee came up with a 3-level stacking tower. We built more L2 models and a few Lasso for L3 and pushed our score to 0.2148. Unfortunately we missed the gold zone, but we did much better than we thought we could.</p>\n\n<p>There are a few characteristics of this competition that gave us the confidence to do this crazy stacking:</p>\n\n<ul>\n<li>CV reflects LB very well. For us, a decrease in CV ALWAYS lead to a decrease in LB</li>\n<li>The CV-LB gap is also fairly consistent for us, though the gap differs a bit for different models (gap for LGB is 0.004x, NN is 0.002x and some of the blends are 0.005x). </li>\n<li>Other than target encoding, overfitting never seems to be a issue. In most of our models, regularization doesn’t play a big part.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 353086,
      "author_name": "mmotoki",
      "author_url": "",
      "post_date": "07/05/2018 21:42:52",
      "content": "<p>One of my main contributions to the team was Bayesian target encoding.  The idea is to use bayesian statistics to encode the categorical variables.   The cool thing is that we can encode not only the target mean, but other statistics like the median, mode, variance, skewness, and kurtosis using the same framework.   We found that this style of target encoding outperforms the built-in LightGBM categorical encoding. </p>\n\n<p><img src=\"https://image.ibb.co/jcvTRd/Screenshot_from_2018_07_05_11_29_33.png\" alt=\"enter image description here\"></p>\n\n<p>Here are some links with more information. </p>\n\n<ul>\n<li><a href=\"https://mattmotoki.github.io/beta-target-encoding.html\">A more detailed write-up here</a>. </li>\n<li><a href=\"https://www.kaggle.com/mmotoki/avito-target-encoding\">Code for implementing bayesian target encoding and getting the results for reproducing the graph</a>.</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "350426": "This competition was very overwhelming. I’ve never seen a Kaggle competition before where we’ve had to intelligently deal with images, text, geospatial information, unlabeled additional data, numeric data, categorical data, and a low-key time series problem to boot! Not to mention that we still don’t even know what the dependent variable is and a few of the independent variables remain a mystery as well.\n\nCongratulations to everyone who got a gold medal -- especially the soloists. In my opinion, anyone who got a solo gold in this competition deserves two gold medals. I have no idea how you’d even manage something like that given everything this competition has to offer.\n\nWe’re really excited that we pushed way farther than we ever thought would be possible in the final weeks, but are really astonished by the caliber of the competition that we faced. I’ve never seen such a competitive final week. ...If only there were a few more gold medal slots, as there certainly were quite a few deserving teams.\n\nNaturally, spending most of the competition in the gold medal zone only to lose out in the end by a smidge is a humbling experience. It leads us to second guess everything we did and wonder where the extra few points could have come from that could have made the difference.\n \nAnyways, I’ll stop complaining and try to make up for our second guessing with an instructive write-up.\n\n\nThe Stacker\n-----------\n\nI’m too lazy to make a pretty diagram, but Sijun was not too lazy, [so you can read more detail here](https://www.kaggle.com/c/avito-demand-prediction/discussion/60059#350998). The original plan was to stop at Level 3 with just a single LightGBM, but the competitiveness of the competition forced us to keep going to level 4. Here's some brief discussion, though see Sijun's additional analysis.\n\n\n**Level 4**\n\nWe created a four-level model, using a special Lasso model to aggregate together submodels. The two key features of this Lasso model was that the weights were constrained to be only positive and the model was trained separately for each `parent_category_name`. We found that `parent_category_name` introduced a good degree of variability in the errors and capturing this by building models separately was good for a very small (we’re talking ~0.00002) boost in score compared to training just one Lasso overall. Creating this insane Level 4 stacking system was key to boosting us from 0.2149 LB to 0.2148 LB on the public leaderboard.\n\n\n**Level 3**\n\nOn the Level 3, we had 3x LightGBMs trained using (1) the L2 models and a few key features (parent_category_name, price, 20-dimensional SVD of text embeddings and a few image features), (2) all the L2 models and all the features used in L1, and (3) the same as the first model but trained with a Poisson objective. We also trained an MLP, a Lasso, and the aforementioned grouped Lasso by `parent_category_name` on all the L2 models (but no other features). Each of these models were also lazily bagged with some previous versions of the model that did not have some of our final models, for a total of 13 L3 models (7 LightGBM, 4 MLP, 2 Lasso). The LightGBM stackers each got ~0.2149 on the public LB and the Lasso ones each got ~0.2153. This was a large jump from Level 2 and Level 1.\n\n\n**Level 2**\n\nThere was a small bump between Level 1 and Level 3 where we added all of our Ridge models to our LightGBM combined with all of our features. We trained this with both normal (regression) and Poisson objectives. We also lazily bagged this model for a total of four copies (3 regression,1 Poisson) by adding in previous copies of the model that didn’t have some of our final features. Our best Level 2 model scored 0.2173 on the public LB.\n\n\n**Level 1**\n\nWe trained a ton of features into LightGBMs and varied the kind of encoding for the categoricals (either one hot encoding, LightGBM’s built-in target encoding, or Bayesian target encoding), varied the objective (regression and Poisson), and did some lazy bagging (using prior copies of the model that didn’t have all the features) for a total of 19 LightGBMs. Our best LightGBM scored 0.2183 on the public LB.\n\nOn the NN front, we trained three CNNs with FastText and two RNNs with Attention and pooling. Our best NN we 0.2181 on the public LB. [You can see how our best NN looked here.](https://www.kaggle.com/c/avito-demand-prediction/discussion/59917) We also diversified our Level 3 blend by adding two NNs trained with multiclass and one on binary objectives, which provided a good boost.\n\nWe also trained a large number of Ridge models that were varied on training just on TFIDF for the title, just on TFIDF for the description, just on character-level TFIDF for the description, and on all text (concat of all TFIDFs). We then trained these models individually for each parent_category_name, cat_bin (described below), and parent_category_name - region interaction. We also trained several Ridges on all the data and included interaction terms in one. All told, there were 17 Ridge models. All of these Ridges (but no other L1 models) formed the basis of Level 2.\n\nLastly, we trained three FM models, one using Wordbatch and two using TFFM. These didn’t end up outperforming our best Ridge models.\n\n\nFeatures\n-----------\n\nOur feature engineering work was very exhaustive and tended toward a kitchen sink. In the beginning, we would diligently add each feature into the model one-by-one and keep it if it improved the first fold. Toward the end of the competition, we ran out of time for that, so we would just dump tons of features in at once. The one-by-one approach is nice because frequently there were features that looked good but ended up making the model worse.\n\nWe added a lot of features, so we’ll be posting separately sometime next week about everything we did.\n\n\nCode Sharing\n----------------\n\nWe’re still cleaning the code and getting everything together, and hope to fully share all our code sometime next week.\n\n\nUntil Next Time (Very Soon)\n---------------------------------\n\nSee you guys more next week. I look forward to continuing to digest all the lessons learned and plotting my steps toward a real gold medal. We have some unfinished business. :)\n\n![missed it by that much][1]\n\n\n  [1]: https://i.imgur.com/Qj8kqVY.jpg",
    "350438": "Thanks for your sharing!",
    "350503": "peterhurford it was very tight, very very tight finishing. We also spent the last hours frantically pushing our score to stay within the gold medal zone.  Hopefully I (or my team mates) will write something up in the next few days - was brain dead for two days and now trying to catch up on too many things in work and life!\n\nDefinitely look forward to seeing more of you around the LB, thanks for all the sharing!",
    "350533": "Thanks. We'll share even more soon!",
    "350656": "Peter, very sorry you missed it :( Thanks for sharing this amazing stacking tower ! You'll surely become a competition master very soon !",
    "350684": "Thanks. Stacking towers are my favorite part of Kaggle. I look forward to stacking toward a gold after a few months off.",
    "350998": "More details about our blender tower. \n![stacking](https://s3-us-west-1.amazonaws.com/sijunhe-blog/plots/post17/stacking_tower.png)\n\nUntil the last 4 days, our plan has always been building/improving base (L1) models and stacking them with a single L2 lgb model. This got us to public LB 0.2155 but we weren’t able to push beyond that. The level of competition around us was incredible and teams kept surpassing us. We then thought of the [winning solution from KDD Cup 2015](https://www.slideshare.net/jeongyoonlee/winning-data-science-competitions-74391113), where Jeong-yoon Lee came up with a 3-level stacking tower. We built more L2 models and a few Lasso for L3 and pushed our score to 0.2148. Unfortunately we missed the gold zone, but we did much better than we thought we could.\n\nThere are a few characteristics of this competition that gave us the confidence to do this crazy stacking:\n\n- CV reflects LB very well. For us, a decrease in CV ALWAYS lead to a decrease in LB\n- The CV-LB gap is also fairly consistent for us, though the gap differs a bit for different models (gap for LGB is 0.004x, NN is 0.002x and some of the blends are 0.005x). \n- Other than target encoding, overfitting never seems to be a issue. In most of our models, regularization doesn’t play a big part.",
    "353086": "One of my main contributions to the team was Bayesian target encoding.  The idea is to use bayesian statistics to encode the categorical variables.   The cool thing is that we can encode not only the target mean, but other statistics like the median, mode, variance, skewness, and kurtosis using the same framework.   We found that this style of target encoding outperforms the built-in LightGBM categorical encoding. \n\n![enter image description here][1]\n\nHere are some links with more information. \n\n* [A more detailed write-up here][2]. \n* [Code for implementing bayesian target encoding and getting the results for reproducing the graph][3].\n\n\n\n  [1]: https://image.ibb.co/jcvTRd/Screenshot_from_2018_07_05_11_29_33.png\n  [2]: https://mattmotoki.github.io/beta-target-encoding.html\n  [3]: https://www.kaggle.com/mmotoki/avito-target-encoding"
  },
  "source": "meta"
}