{
  "id": 59885,
  "title": "3 place solution",
  "url": "/competitions/avito-demand-prediction/writeups/superanova-3-place-solution",
  "author_name": "",
  "post_date": "2018-06-28T02:08:02.716718600Z",
  "votes": 123,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Congratulations to the winners, especially <strong>Dance with Ensemble</strong> your score was amazing from start to finish. Also, congrats to the solo gold medallists, and especially to the solo kaggler who claimed it was <strong>too hard to get a solo gold</strong> and eventually snuck in at #13– we have been secretly rooting for you.</p>\n\n<p>Here we will outline a brief explanation of our solution.</p>\n\n<h2>Ensemble/approach</h2>\n\n<p>Our solution is the average of 2 ensembles. The first ensemble was trained using a standard 5-fold validation schema.</p>\n\n<p>The second ensemble used a time-series schema. The test data begins a number of days past the final day included in the training data. We tried to replicate this in our internal validation. To do so, we trained on the first six days of training and used days ten through thirteen as validation. This meant days six through nine were excluded, mimicking the gap between train and test.Then, to generate predictions for the test data, we trained on the entire training set.  </p>\n\n<p>In order to generate likelihoods of categorical features for this approach, we always applied a gap of 4 days. For example to estimate likelihoods for day four of the training data, we use the average of target for day zero. To estimate likelihoods for day five, we used likelihoods of (day0+day1)/2. We decided on a gap of four days for stability, as it gave similar CV and LB performance. </p>\n\n<p>Our best single models came from this approach. Our best single lgb clocked in at 0.2163, and our best single nn scored .2180.</p>\n\n<h2>LightGBM features</h2>\n\n<ul>\n<li>Tf-idf on words  (2 grams on description, 1gram for the params and\ntitle)</li>\n<li>Tf-idf on chars (5 grams)</li>\n<li>Word2vec-based features on words  (this worked better for us than\nfastext)</li>\n<li>Pretrained fastext features on words</li>\n<li>Image quality features (from <a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">here</a>)</li>\n<li>Vgg16 feature (from <a>here</a> \nand <a>here</a> )</li>\n<li>Vgg19 feature (similar to vgg16)</li>\n<li>Resnet prediction (of object top 3)  features</li>\n<li>Inception prediction (of object  top 3)  features</li>\n<li>Xception prediction (of object  top 3)  features</li>\n<li>Some binnings of numerical features (like price)</li>\n<li>Some group-by user type of features (like average number of words,\naverage number of days of Displaying ads)</li>\n<li>Some text based counts (like upper counts, punctuation, counts of\nemojis)</li>\n<li>Some location features based on latitude and  longtitude</li>\n<li>Likelihood and counts on almost all 3-way interactions of categorical\nfeatures (for kfold approach we excluded user_id interactions as they\nover fitted. For the time series approach, user_id likelihoods gave a\ngood boost)</li>\n</ul>\n\n<p>We primarily used <strong>xentropy</strong> for our objective, as probabilities are constrained between zero and one. </p>\n\n<h2>Neural Networks</h2>\n\n<p>Our best nn used two stacked, bidirectional GRUs on the text with concatenated embeddings of fastext and those trained on our own with word2vec, along with some of the numerical features used in LightGBM. All categorical features (minus user id) had embedding layers of size 100. Our objective was binary cross-entropy with sigmoid output. </p>\n\n<p>A few small additions/key points:</p>\n\n<ul>\n<li>There was a dense input for the vgg16 and vgg19 features</li>\n<li>Text was stemmed based using nltk.</li>\n<li>For text we combined description,title and params into one field,\nseparated by a dividing character.</li>\n</ul>\n\n<p>In addition to the above network, we ran a NN for each feature channel (image, text, character) with minimal categorical embeddings and likelihood features. When stacking with just these models, they scored ~.216 at the second level, which was comparatively weak to lightGBM. </p>\n\n<p>For these models, we primarily used each feature channel independently, than concatenated them with basic likelihood features and categorical embeddings before feeding them to dense layers for the output. While these models were individually weaker, they stacked well, and maximized the 2nd-level information from each feature channel.</p>\n\n<h2>Stacking</h2>\n\n<p>Both approaches include mostly lightGBM and neural nets. In addition, simple ridge models were used to improve performance for some lightGBM models, in a similar manner to a couple of shared kernels.</p>\n\n<p>To stack our 5-fold approach, we used the time schema as explained above (0-5, 10-14) to validate and optimize hyperparameters. We also did a bit of re-stacking, adding some counts and likelihoods spanning user-based interactions of categorical features and price. Our second level stack combined lightGBM, neural nets using 2 hidden layers, linear output, and mse objective, and sklearn’s ExtratTeesRegressor. This approach scored 0.2145.</p>\n\n<p>For the time series approach, we found best parameters using days [10,11] for training and the remaining for validation. Restacking did not help here and we used lgb and nn of equal weight .  This scored 0.2140.</p>\n\n<p>Finally, a blend (35% of time stack + 65% of 5-fold stack) gave our best public (0.2136) and private score.</p>",
  "messages": [
    {
      "id": "349306",
      "postDate": "06/28/2018 02:08:02",
      "content": "<p>Congratulations to the winners, especially <strong>Dance with Ensemble</strong> your score was amazing from start to finish. Also, congrats to the solo gold medallists, and especially to the solo kaggler who claimed it was <strong>too hard to get a solo gold</strong> and eventually snuck in at #13– we have been secretly rooting for you.</p>\n\n<p>Here we will outline a brief explanation of our solution.</p>\n\n<h2>Ensemble/approach</h2>\n\n<p>Our solution is the average of 2 ensembles. The first ensemble was trained using a standard 5-fold validation schema.</p>\n\n<p>The second ensemble used a time-series schema. The test data begins a number of days past the final day included in the training data. We tried to replicate this in our internal validation. To do so, we trained on the first six days of training and used days ten through thirteen as validation. This meant days six through nine were excluded, mimicking the gap between train and test.Then, to generate predictions for the test data, we trained on the entire training set.  </p>\n\n<p>In order to generate likelihoods of categorical features for this approach, we always applied a gap of 4 days. For example to estimate likelihoods for day four of the training data, we use the average of target for day zero. To estimate likelihoods for day five, we used likelihoods of (day0+day1)/2. We decided on a gap of four days for stability, as it gave similar CV and LB performance. </p>\n\n<p>Our best single models came from this approach. Our best single lgb clocked in at 0.2163, and our best single nn scored .2180.</p>\n\n<h2>LightGBM features</h2>\n\n<ul>\n<li>Tf-idf on words  (2 grams on description, 1gram for the params and\ntitle)</li>\n<li>Tf-idf on chars (5 grams)</li>\n<li>Word2vec-based features on words  (this worked better for us than\nfastext)</li>\n<li>Pretrained fastext features on words</li>\n<li>Image quality features (from <a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">here</a>)</li>\n<li>Vgg16 feature (from <a>here</a> \nand <a>here</a> )</li>\n<li>Vgg19 feature (similar to vgg16)</li>\n<li>Resnet prediction (of object top 3)  features</li>\n<li>Inception prediction (of object  top 3)  features</li>\n<li>Xception prediction (of object  top 3)  features</li>\n<li>Some binnings of numerical features (like price)</li>\n<li>Some group-by user type of features (like average number of words,\naverage number of days of Displaying ads)</li>\n<li>Some text based counts (like upper counts, punctuation, counts of\nemojis)</li>\n<li>Some location features based on latitude and  longtitude</li>\n<li>Likelihood and counts on almost all 3-way interactions of categorical\nfeatures (for kfold approach we excluded user_id interactions as they\nover fitted. For the time series approach, user_id likelihoods gave a\ngood boost)</li>\n</ul>\n\n<p>We primarily used <strong>xentropy</strong> for our objective, as probabilities are constrained between zero and one. </p>\n\n<h2>Neural Networks</h2>\n\n<p>Our best nn used two stacked, bidirectional GRUs on the text with concatenated embeddings of fastext and those trained on our own with word2vec, along with some of the numerical features used in LightGBM. All categorical features (minus user id) had embedding layers of size 100. Our objective was binary cross-entropy with sigmoid output. </p>\n\n<p>A few small additions/key points:</p>\n\n<ul>\n<li>There was a dense input for the vgg16 and vgg19 features</li>\n<li>Text was stemmed based using nltk.</li>\n<li>For text we combined description,title and params into one field,\nseparated by a dividing character.</li>\n</ul>\n\n<p>In addition to the above network, we ran a NN for each feature channel (image, text, character) with minimal categorical embeddings and likelihood features. When stacking with just these models, they scored ~.216 at the second level, which was comparatively weak to lightGBM. </p>\n\n<p>For these models, we primarily used each feature channel independently, than concatenated them with basic likelihood features and categorical embeddings before feeding them to dense layers for the output. While these models were individually weaker, they stacked well, and maximized the 2nd-level information from each feature channel.</p>\n\n<h2>Stacking</h2>\n\n<p>Both approaches include mostly lightGBM and neural nets. In addition, simple ridge models were used to improve performance for some lightGBM models, in a similar manner to a couple of shared kernels.</p>\n\n<p>To stack our 5-fold approach, we used the time schema as explained above (0-5, 10-14) to validate and optimize hyperparameters. We also did a bit of re-stacking, adding some counts and likelihoods spanning user-based interactions of categorical features and price. Our second level stack combined lightGBM, neural nets using 2 hidden layers, linear output, and mse objective, and sklearn’s ExtratTeesRegressor. This approach scored 0.2145.</p>\n\n<p>For the time series approach, we found best parameters using days [10,11] for training and the remaining for validation. Restacking did not help here and we used lgb and nn of equal weight .  This scored 0.2140.</p>\n\n<p>Finally, a blend (35% of time stack + 65% of 5-fold stack) gave our best public (0.2136) and private score.</p>",
      "rawMarkdown": "Congratulations to the winners, especially **Dance with Ensemble** your score was amazing from start to finish. Also, congrats to the solo gold medallists, and especially to the solo kaggler who claimed it was **too hard to get a solo gold** and eventually snuck in at #13– we have been secretly rooting for you.\n\nHere we will outline a brief explanation of our solution.\n\nEnsemble/approach\n-----------------\n\nOur solution is the average of 2 ensembles. The first ensemble was trained using a standard 5-fold validation schema.\n\nThe second ensemble used a time-series schema. The test data begins a number of days past the final day included in the training data. We tried to replicate this in our internal validation. To do so, we trained on the first six days of training and used days ten through thirteen as validation. This meant days six through nine were excluded, mimicking the gap between train and test.Then, to generate predictions for the test data, we trained on the entire training set.  \n\nIn order to generate likelihoods of categorical features for this approach, we always applied a gap of 4 days. For example to estimate likelihoods for day four of the training data, we use the average of target for day zero. To estimate likelihoods for day five, we used likelihoods of (day0+day1)/2. We decided on a gap of four days for stability, as it gave similar CV and LB performance. \n\nOur best single models came from this approach. Our best single lgb clocked in at 0.2163, and our best single nn scored .2180.\n\n\nLightGBM features\n-----------------\n\n - Tf-idf on words  (2 grams on description, 1gram for the params and\n   title)\n - Tf-idf on chars (5 grams)\n - Word2vec-based features on words  (this worked better for us than\n   fastext)\n - Pretrained fastext features on words\n - Image quality features (from [here][1])\n - Vgg16 feature (from [here][2] \n   and [here][3] )\n - Vgg19 feature (similar to vgg16)\n - Resnet prediction (of object top 3)  features\n - Inception prediction (of object  top 3)  features\n - Xception prediction (of object  top 3)  features\n - Some binnings of numerical features (like price)\n - Some group-by user type of features (like average number of words,\n   average number of days of Displaying ads)\n - Some text based counts (like upper counts, punctuation, counts of\n   emojis)\n - Some location features based on latitude and  longtitude\n - Likelihood and counts on almost all 3-way interactions of categorical\n   features (for kfold approach we excluded user_id interactions as they\n   over fitted. For the time series approach, user_id likelihoods gave a\n   good boost)\n\nWe primarily used **xentropy** for our objective, as probabilities are constrained between zero and one. \n\n\nNeural Networks\n---------------\n\nOur best nn used two stacked, bidirectional GRUs on the text with concatenated embeddings of fastext and those trained on our own with word2vec, along with some of the numerical features used in LightGBM. All categorical features (minus user id) had embedding layers of size 100. Our objective was binary cross-entropy with sigmoid output. \n\nA few small additions/key points:\n\n - There was a dense input for the vgg16 and vgg19 features\n - Text was stemmed based using nltk.\n - For text we combined description,title and params into one field,\n   separated by a dividing character.\n\n\nIn addition to the above network, we ran a NN for each feature channel (image, text, character) with minimal categorical embeddings and likelihood features. When stacking with just these models, they scored ~.216 at the second level, which was comparatively weak to lightGBM. \n\nFor these models, we primarily used each feature channel independently, than concatenated them with basic likelihood features and categorical embeddings before feeding them to dense layers for the output. While these models were individually weaker, they stacked well, and maximized the 2nd-level information from each feature channel.\n\n\nStacking\n--------\n\nBoth approaches include mostly lightGBM and neural nets. In addition, simple ridge models were used to improve performance for some lightGBM models, in a similar manner to a couple of shared kernels.\n\nTo stack our 5-fold approach, we used the time schema as explained above (0-5, 10-14) to validate and optimize hyperparameters. We also did a bit of re-stacking, adding some counts and likelihoods spanning user-based interactions of categorical features and price. Our second level stack combined lightGBM, neural nets using 2 hidden layers, linear output, and mse objective, and sklearn’s ExtratTeesRegressor. This approach scored 0.2145.\n\nFor the time series approach, we found best parameters using days [10,11] for training and the remaining for validation. Restacking did not help here and we used lgb and nn of equal weight .  This scored 0.2140.\n\nFinally, a blend (35% of time stack + 65% of 5-fold stack) gave our best public (0.2136) and private score.\n\n\n  [1]: https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\n  [2]: http://%20https://www.kaggle.com/bguberfain/vgg16-train-features\n  [3]: http://%20https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16",
      "votes": null
    },
    {
      "id": "349326",
      "postDate": "06/28/2018 02:30:29",
      "content": "<p>Well done Marios &amp; team. Thanks for sharing the appraoch.! </p>",
      "rawMarkdown": "Well done Marios &amp; team. Thanks for sharing the appraoch.!",
      "votes": null
    },
    {
      "id": "349329",
      "postDate": "06/28/2018 02:32:33",
      "content": "<p>Congratulations.....</p>",
      "rawMarkdown": "Congratulations.....",
      "votes": null
    },
    {
      "id": "349344",
      "postDate": "06/28/2018 02:47:20",
      "content": "<p>Congrats KazAnova and your team  and thanks for sharing !</p>",
      "rawMarkdown": "Congrats KazAnova and your team  and thanks for sharing !",
      "votes": null
    },
    {
      "id": "349405",
      "postDate": "06/28/2018 04:39:27",
      "content": "<p>Well done and thanks for sharing.</p>",
      "rawMarkdown": "Well done and thanks for sharing.",
      "votes": null
    },
    {
      "id": "349412",
      "postDate": "06/28/2018 04:46:51",
      "content": "<p>Congratulations. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "349420",
      "postDate": "06/28/2018 04:56:53",
      "content": "<p>Congratulations @KazAnova and team. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations @KazAnova and team. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "349450",
      "postDate": "06/28/2018 06:10:25",
      "content": "<p>Congrats guys and thx for the nice writeup!</p>",
      "rawMarkdown": "Congrats guys and thx for the nice writeup!",
      "votes": null
    },
    {
      "id": "349599",
      "postDate": "06/28/2018 10:45:57",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": null
    },
    {
      "id": "349620",
      "postDate": "06/28/2018 11:38:07",
      "content": "<p>Well done </p>",
      "rawMarkdown": "Well done",
      "votes": null
    },
    {
      "id": "349713",
      "postDate": "06/28/2018 14:27:16",
      "content": "<p>Congratulation on your 3rd place Kazanova! <br>\nMaybe not your favorite position(4th), but I can only imagine that winning prizes feels really good :) <br>\nYour solution is really interesting, as I could not come up of a good way to set up a time-series model. <br></p>\n\n<p>I have a few questions about stacking (your team has brilliant stackers)<br>\n1) What was your reason behind adding some counts and likelihoods to your stackers?<br>\n   Or, as a general question, what features should we add to the metamodels (level2 models)? Is it try and error?<br>\n2) About your 5fold models. When you merge teams, you often have different folds between teammates. <br>\n   Are we supposed to re-run all our models in the same folds? Or is it okay to use different folds and stack them?</p>",
      "rawMarkdown": "Congratulation on your 3rd place Kazanova! <br>\nMaybe not your favorite position(4th), but I can only imagine that winning prizes feels really good :) <br>\nYour solution is really interesting, as I could not come up of a good way to set up a time-series model. <br>\n\nI have a few questions about stacking (your team has brilliant stackers)<br>\n1) What was your reason behind adding some counts and likelihoods to your stackers?<br>\n   Or, as a general question, what features should we add to the metamodels (level2 models)? Is it try and error?<br>\n2) About your 5fold models. When you merge teams, you often have different folds between teammates. <br>\n   Are we supposed to re-run all our models in the same folds? Or is it okay to use different folds and stack them?",
      "votes": null
    },
    {
      "id": "349800",
      "postDate": "06/28/2018 16:53:12",
      "content": "<p>Thank you </p>\n\n<p>Haha, I guess I can live with a 3rd place :)</p>\n\n<p>1) Experimentally this often works . I guess this is because a model has a chance to look at this data from a slightly difference perspective when you already have powerful models in . In other words it tries to explore information it may have missed before.  It is mostly trial and error.</p>\n\n<p>2) No. you can see it as a form of bagging if you trained on different folds - to my experience it does not change the results. People can use different folds</p>",
      "rawMarkdown": "Thank you \n\nHaha, I guess I can live with a 3rd place :)\n\n1) Experimentally this often works . I guess this is because a model has a chance to look at this data from a slightly difference perspective when you already have powerful models in . In other words it tries to explore information it may have missed before.  It is mostly trial and error.\n\n2) No. you can see it as a form of bagging if you trained on different folds - to my experience it does not change the results. People can use different folds",
      "votes": null
    },
    {
      "id": "349812",
      "postDate": "06/28/2018 17:18:24",
      "content": "<p>Congrats Kaza, Pavel and Raymond :-)\nGood work!</p>",
      "rawMarkdown": "Congrats Kaza, Pavel and Raymond :-)\nGood work!",
      "votes": null
    },
    {
      "id": "349942",
      "postDate": "06/28/2018 23:09:55",
      "content": "<p>congrats Kazanova!</p>",
      "rawMarkdown": "congrats Kazanova!",
      "votes": null
    },
    {
      "id": "349983",
      "postDate": "06/29/2018 01:28:34",
      "content": "<p>Congrats, Marios!!! </p>",
      "rawMarkdown": "Congrats, Marios!!!",
      "votes": null
    },
    {
      "id": "350214",
      "postDate": "06/29/2018 12:03:46",
      "content": "<p>Thanks you Kaza. I have a question. In the last part of \"LGBM features\" , Is the \"Likelihood\" means something like \"target encode\" ?</p>",
      "rawMarkdown": "Thanks you Kaza. I have a question. In the last part of \"LGBM features\" , Is the \"Likelihood\" means something like \"target encode\" ?",
      "votes": null
    },
    {
      "id": "350226",
      "postDate": "06/29/2018 12:39:08",
      "content": "<p>Congratulations , Could you please tell me what are '' Some location features based on latitude and longtitude'', Thank you.</p>",
      "rawMarkdown": "Congratulations , Could you please tell me what are '' Some location features based on latitude and longtitude'', Thank you.",
      "votes": null
    },
    {
      "id": "350243",
      "postDate": "06/29/2018 12:57:27",
      "content": "<p>Thank you for the reply!<br>\n2) is really interesting, because I heard some other Grand Master say the same thing.<br> \nHowever when I asked @Kohei, he told me that he definitely had experiences (from different competitions) where different folds lead to leakages (especially when there were highly predictive features)<br>\nPerhaps this is a controversial issue...?</p>",
      "rawMarkdown": "Thank you for the reply!<br>\n2) is really interesting, because I heard some other Grand Master say the same thing.<br> \nHowever when I asked @Kohei, he told me that he definitely had experiences (from different competitions) where different folds lead to leakages (especially when there were highly predictive features)<br>\nPerhaps this is a controversial issue...?",
      "votes": null
    },
    {
      "id": "350254",
      "postDate": "06/29/2018 13:03:30",
      "content": "<p><a href=\"/rhgrossm\">@rhgrossm</a> (To train them is my cause)\nCongratulation to you too! <br>\nCan I ask one question to you too? <br>\nI remember from the toxic comment competition, that you advocated trying pseudo-labeling in any competition. <br>\nThis is my second competition (after talkingdata), that I tried it, and didn't work... <br>\nDoes pseudo-labeling not work well in general for GDBT? Or is it that we don't have a good enough prediction? (but in talkingdata we had a +0.97AUC, and it still didn't work)<br>\nI am really lost, and I hope you can help me :)</p>",
      "rawMarkdown": "rhgrossm (To train them is my cause)\nCongratulation to you too! <br>\nCan I ask one question to you too? <br>\nI remember from the toxic comment competition, that you advocated trying pseudo-labeling in any competition. <br>\nThis is my second competition (after talkingdata), that I tried it, and didn't work... <br>\nDoes pseudo-labeling not work well in general for GDBT? Or is it that we don't have a good enough prediction? (but in talkingdata we had a +0.97AUC, and it still didn't work)<br>\nI am really lost, and I hope you can help me :)",
      "votes": null
    },
    {
      "id": "350292",
      "postDate": "06/29/2018 14:07:03",
      "content": "<p>Yes, it is exactly that.</p>",
      "rawMarkdown": "Yes, it is exactly that.",
      "votes": null
    },
    {
      "id": "350295",
      "postDate": "06/29/2018 14:09:02",
      "content": "<p>Very similar to <a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">these</a> </p>\n\n<p>We created some additional clusters based on kmeans. </p>",
      "rawMarkdown": "Very similar to [these][1] \n\nWe created some additional clusters based on kmeans. \n\n  [1]: https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters",
      "votes": null
    },
    {
      "id": "350297",
      "postDate": "06/29/2018 14:10:59",
      "content": "<blockquote>\n  <p>I remember from the toxic comment competition, that you advocated\n  trying  pseudo-labeling in any competition.</p>\n</blockquote>\n\n<p>I tried that here and it did not work. </p>\n\n<p>I tried with both lgbt and nn</p>\n\n<p>I could see performance of  pseudo-labeling was worse with lgb than nns. </p>",
      "rawMarkdown": "&gt; I remember from the toxic comment competition, that you advocated\n&gt; trying  pseudo-labeling in any competition.\n\nI tried that here and it did not work. \n\nI tried with both lgbt and nn\n\nI could see performance of  pseudo-labeling was worse with lgb than nns.",
      "votes": null
    },
    {
      "id": "350369",
      "postDate": "06/29/2018 16:11:57",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks.",
      "votes": null
    },
    {
      "id": "350383",
      "postDate": "06/29/2018 16:58:10",
      "content": "<p>There are some competitions (especially when the data is not very big) where you can introduce leakage, no matter what cv you are doing. Sometimes it is good to ensure everyone is cving the same way to avoid any mistakes. Other than that, I have not seen huge differences, but I trust kohei if he says he has.  </p>",
      "rawMarkdown": "There are some competitions (especially when the data is not very big) where you can introduce leakage, no matter what cv you are doing. Sometimes it is good to ensure everyone is cving the same way to avoid any mistakes. Other than that, I have not seen huge differences, but I trust kohei if he says he has.",
      "votes": null
    },
    {
      "id": "350544",
      "postDate": "06/29/2018 23:26:19",
      "content": "<p>Thanks again for all the replies!<br>\nHope to see you in future competitions again :)</p>",
      "rawMarkdown": "Thanks again for all the replies!<br>\nHope to see you in future competitions again :)",
      "votes": null
    },
    {
      "id": "350562",
      "postDate": "06/30/2018 00:28:04",
      "content": "<p>I am sure you will :)</p>",
      "rawMarkdown": "I am sure you will :)",
      "votes": null
    },
    {
      "id": "351103",
      "postDate": "07/01/2018 10:25:45",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "357779",
      "postDate": "07/16/2018 20:47:58",
      "content": "<p>Sorry I didn't respond! For pseudolabeling, you have to be careful about the amount of noise that you are introducing into the training and test set. With a problem like this, with a lot of high dimensional categoricals, introducing predictions is liable to introduce a high level of label noise. I actually found that using test predictions as targets led to a similar score as with zero pseudolabeling, which means we weren't gaining a lot of extra information at risk of overfitting. </p>\n\n<p>I didn't spend a ton of time on TD, but PL worked well there. How did you deal with adding your test targets?</p>\n\n<p>And, in general, GBDT can't handle pseudolabeling because it is not continuous. Hope that helps!</p>",
      "rawMarkdown": "Sorry I didn't respond! For pseudolabeling, you have to be careful about the amount of noise that you are introducing into the training and test set. With a problem like this, with a lot of high dimensional categoricals, introducing predictions is liable to introduce a high level of label noise. I actually found that using test predictions as targets led to a similar score as with zero pseudolabeling, which means we weren't gaining a lot of extra information at risk of overfitting. \n\nI didn't spend a ton of time on TD, but PL worked well there. How did you deal with adding your test targets?\n\nAnd, in general, GBDT can't handle pseudolabeling because it is not continuous. Hope that helps!",
      "votes": null
    },
    {
      "id": "364403",
      "postDate": "07/31/2018 13:24:39",
      "content": "<p>Thanks for the reply :) <br>\nAll my effort into PL was using GBDT.... You should have told me earlier that it doesn't work :( hahaha <br>\nIt is encoraging that PL worked in TD though. We didn't try it with NN, so that was the issue.</p>",
      "rawMarkdown": "Thanks for the reply :) <br>\nAll my effort into PL was using GBDT.... You should have told me earlier that it doesn't work :( hahaha <br>\nIt is encoraging that PL worked in TD though. We didn't try it with NN, so that was the issue.",
      "votes": null
    },
    {
      "id": "2180792",
      "postDate": "03/14/2023 05:18:30",
      "content": "<p>Where can i find the code behind this submission? Almost every solution on the leaderboard has no notebook :(</p>",
      "rawMarkdown": "Where can i find the code behind this submission? Almost every solution on the leaderboard has no notebook :(",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2180792,
      "author_name": "tripmanas",
      "author_url": "",
      "post_date": "03/14/2023 05:18:30",
      "content": "<p>Where can i find the code behind this submission? Almost every solution on the leaderboard has no notebook :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349326,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "06/28/2018 02:30:29",
      "content": "<p>Well done Marios &amp; team. Thanks for sharing the appraoch.! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349329,
      "author_name": "samratp",
      "author_url": "",
      "post_date": "06/28/2018 02:32:33",
      "content": "<p>Congratulations.....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349344,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "06/28/2018 02:47:20",
      "content": "<p>Congrats KazAnova and your team  and thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349405,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "06/28/2018 04:39:27",
      "content": "<p>Well done and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349412,
      "author_name": "a45632",
      "author_url": "",
      "post_date": "06/28/2018 04:46:51",
      "content": "<p>Congratulations. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349420,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/28/2018 04:56:53",
      "content": "<p>Congratulations @KazAnova and team. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349450,
      "author_name": "apgeorg",
      "author_url": "",
      "post_date": "06/28/2018 06:10:25",
      "content": "<p>Congrats guys and thx for the nice writeup!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349599,
      "author_name": "rakibilly",
      "author_url": "",
      "post_date": "06/28/2018 10:45:57",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349620,
      "author_name": "classtag",
      "author_url": "",
      "post_date": "06/28/2018 11:38:07",
      "content": "<p>Well done </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349713,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "06/28/2018 14:27:16",
      "content": "<p>Congratulation on your 3rd place Kazanova! <br>\nMaybe not your favorite position(4th), but I can only imagine that winning prizes feels really good :) <br>\nYour solution is really interesting, as I could not come up of a good way to set up a time-series model. <br></p>\n\n<p>I have a few questions about stacking (your team has brilliant stackers)<br>\n1) What was your reason behind adding some counts and likelihoods to your stackers?<br>\n   Or, as a general question, what features should we add to the metamodels (level2 models)? Is it try and error?<br>\n2) About your 5fold models. When you merge teams, you often have different folds between teammates. <br>\n   Are we supposed to re-run all our models in the same folds? Or is it okay to use different folds and stack them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 349800,
          "author_name": "kazanova",
          "author_url": "",
          "post_date": "06/28/2018 16:53:12",
          "content": "<p>Thank you </p>\n\n<p>Haha, I guess I can live with a 3rd place :)</p>\n\n<p>1) Experimentally this often works . I guess this is because a model has a chance to look at this data from a slightly difference perspective when you already have powerful models in . In other words it tries to explore information it may have missed before.  It is mostly trial and error.</p>\n\n<p>2) No. you can see it as a form of bagging if you trained on different folds - to my experience it does not change the results. People can use different folds</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350243,
          "author_name": "pocketsuteado",
          "author_url": "",
          "post_date": "06/29/2018 12:57:27",
          "content": "<p>Thank you for the reply!<br>\n2) is really interesting, because I heard some other Grand Master say the same thing.<br> \nHowever when I asked @Kohei, he told me that he definitely had experiences (from different competitions) where different folds lead to leakages (especially when there were highly predictive features)<br>\nPerhaps this is a controversial issue...?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350383,
          "author_name": "kazanova",
          "author_url": "",
          "post_date": "06/29/2018 16:58:10",
          "content": "<p>There are some competitions (especially when the data is not very big) where you can introduce leakage, no matter what cv you are doing. Sometimes it is good to ensure everyone is cving the same way to avoid any mistakes. Other than that, I have not seen huge differences, but I trust kohei if he says he has.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350544,
          "author_name": "pocketsuteado",
          "author_url": "",
          "post_date": "06/29/2018 23:26:19",
          "content": "<p>Thanks again for all the replies!<br>\nHope to see you in future competitions again :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350562,
          "author_name": "kazanova",
          "author_url": "",
          "post_date": "06/30/2018 00:28:04",
          "content": "<p>I am sure you will :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349812,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/28/2018 17:18:24",
      "content": "<p>Congrats Kaza, Pavel and Raymond :-)\nGood work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349942,
      "author_name": "stexyz",
      "author_url": "",
      "post_date": "06/28/2018 23:09:55",
      "content": "<p>congrats Kazanova!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349983,
      "author_name": "u39kun",
      "author_url": "",
      "post_date": "06/29/2018 01:28:34",
      "content": "<p>Congrats, Marios!!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 350214,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "06/29/2018 12:03:46",
      "content": "<p>Thanks you Kaza. I have a question. In the last part of \"LGBM features\" , Is the \"Likelihood\" means something like \"target encode\" ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 350292,
          "author_name": "kazanova",
          "author_url": "",
          "post_date": "06/29/2018 14:07:03",
          "content": "<p>Yes, it is exactly that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350369,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "06/29/2018 16:11:57",
          "content": "<p>Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350226,
      "author_name": "zahedi",
      "author_url": "",
      "post_date": "06/29/2018 12:39:08",
      "content": "<p>Congratulations , Could you please tell me what are '' Some location features based on latitude and longtitude'', Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 350295,
          "author_name": "kazanova",
          "author_url": "",
          "post_date": "06/29/2018 14:09:02",
          "content": "<p>Very similar to <a href=\"https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters\">these</a> </p>\n\n<p>We created some additional clusters based on kmeans. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350254,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "06/29/2018 13:03:30",
      "content": "<p><a href=\"/rhgrossm\">@rhgrossm</a> (To train them is my cause)\nCongratulation to you too! <br>\nCan I ask one question to you too? <br>\nI remember from the toxic comment competition, that you advocated trying pseudo-labeling in any competition. <br>\nThis is my second competition (after talkingdata), that I tried it, and didn't work... <br>\nDoes pseudo-labeling not work well in general for GDBT? Or is it that we don't have a good enough prediction? (but in talkingdata we had a +0.97AUC, and it still didn't work)<br>\nI am really lost, and I hope you can help me :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 350297,
          "author_name": "kazanova",
          "author_url": "",
          "post_date": "06/29/2018 14:10:59",
          "content": "<blockquote>\n  <p>I remember from the toxic comment competition, that you advocated\n  trying  pseudo-labeling in any competition.</p>\n</blockquote>\n\n<p>I tried that here and it did not work. </p>\n\n<p>I tried with both lgbt and nn</p>\n\n<p>I could see performance of  pseudo-labeling was worse with lgb than nns. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 357779,
          "author_name": "rhgrossm",
          "author_url": "",
          "post_date": "07/16/2018 20:47:58",
          "content": "<p>Sorry I didn't respond! For pseudolabeling, you have to be careful about the amount of noise that you are introducing into the training and test set. With a problem like this, with a lot of high dimensional categoricals, introducing predictions is liable to introduce a high level of label noise. I actually found that using test predictions as targets led to a similar score as with zero pseudolabeling, which means we weren't gaining a lot of extra information at risk of overfitting. </p>\n\n<p>I didn't spend a ton of time on TD, but PL worked well there. How did you deal with adding your test targets?</p>\n\n<p>And, in general, GBDT can't handle pseudolabeling because it is not continuous. Hope that helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 364403,
          "author_name": "pocketsuteado",
          "author_url": "",
          "post_date": "07/31/2018 13:24:39",
          "content": "<p>Thanks for the reply :) <br>\nAll my effort into PL was using GBDT.... You should have told me earlier that it doesn't work :( hahaha <br>\nIt is encoraging that PL worked in TD though. We didn't try it with NN, so that was the issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 351103,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "07/01/2018 10:25:45",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349306": "Congratulations to the winners, especially **Dance with Ensemble** your score was amazing from start to finish. Also, congrats to the solo gold medallists, and especially to the solo kaggler who claimed it was **too hard to get a solo gold** and eventually snuck in at #13– we have been secretly rooting for you.\n\nHere we will outline a brief explanation of our solution.\n\nEnsemble/approach\n-----------------\n\nOur solution is the average of 2 ensembles. The first ensemble was trained using a standard 5-fold validation schema.\n\nThe second ensemble used a time-series schema. The test data begins a number of days past the final day included in the training data. We tried to replicate this in our internal validation. To do so, we trained on the first six days of training and used days ten through thirteen as validation. This meant days six through nine were excluded, mimicking the gap between train and test.Then, to generate predictions for the test data, we trained on the entire training set.  \n\nIn order to generate likelihoods of categorical features for this approach, we always applied a gap of 4 days. For example to estimate likelihoods for day four of the training data, we use the average of target for day zero. To estimate likelihoods for day five, we used likelihoods of (day0+day1)/2. We decided on a gap of four days for stability, as it gave similar CV and LB performance. \n\nOur best single models came from this approach. Our best single lgb clocked in at 0.2163, and our best single nn scored .2180.\n\n\nLightGBM features\n-----------------\n\n - Tf-idf on words  (2 grams on description, 1gram for the params and\n   title)\n - Tf-idf on chars (5 grams)\n - Word2vec-based features on words  (this worked better for us than\n   fastext)\n - Pretrained fastext features on words\n - Image quality features (from [here][1])\n - Vgg16 feature (from [here][2] \n   and [here][3] )\n - Vgg19 feature (similar to vgg16)\n - Resnet prediction (of object top 3)  features\n - Inception prediction (of object  top 3)  features\n - Xception prediction (of object  top 3)  features\n - Some binnings of numerical features (like price)\n - Some group-by user type of features (like average number of words,\n   average number of days of Displaying ads)\n - Some text based counts (like upper counts, punctuation, counts of\n   emojis)\n - Some location features based on latitude and  longtitude\n - Likelihood and counts on almost all 3-way interactions of categorical\n   features (for kfold approach we excluded user_id interactions as they\n   over fitted. For the time series approach, user_id likelihoods gave a\n   good boost)\n\nWe primarily used **xentropy** for our objective, as probabilities are constrained between zero and one. \n\n\nNeural Networks\n---------------\n\nOur best nn used two stacked, bidirectional GRUs on the text with concatenated embeddings of fastext and those trained on our own with word2vec, along with some of the numerical features used in LightGBM. All categorical features (minus user id) had embedding layers of size 100. Our objective was binary cross-entropy with sigmoid output. \n\nA few small additions/key points:\n\n - There was a dense input for the vgg16 and vgg19 features\n - Text was stemmed based using nltk.\n - For text we combined description,title and params into one field,\n   separated by a dividing character.\n\n\nIn addition to the above network, we ran a NN for each feature channel (image, text, character) with minimal categorical embeddings and likelihood features. When stacking with just these models, they scored ~.216 at the second level, which was comparatively weak to lightGBM. \n\nFor these models, we primarily used each feature channel independently, than concatenated them with basic likelihood features and categorical embeddings before feeding them to dense layers for the output. While these models were individually weaker, they stacked well, and maximized the 2nd-level information from each feature channel.\n\n\nStacking\n--------\n\nBoth approaches include mostly lightGBM and neural nets. In addition, simple ridge models were used to improve performance for some lightGBM models, in a similar manner to a couple of shared kernels.\n\nTo stack our 5-fold approach, we used the time schema as explained above (0-5, 10-14) to validate and optimize hyperparameters. We also did a bit of re-stacking, adding some counts and likelihoods spanning user-based interactions of categorical features and price. Our second level stack combined lightGBM, neural nets using 2 hidden layers, linear output, and mse objective, and sklearn’s ExtratTeesRegressor. This approach scored 0.2145.\n\nFor the time series approach, we found best parameters using days [10,11] for training and the remaining for validation. Restacking did not help here and we used lgb and nn of equal weight .  This scored 0.2140.\n\nFinally, a blend (35% of time stack + 65% of 5-fold stack) gave our best public (0.2136) and private score.\n\n\n  [1]: https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\n  [2]: http://%20https://www.kaggle.com/bguberfain/vgg16-train-features\n  [3]: http://%20https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16",
    "349326": "Well done Marios &amp; team. Thanks for sharing the appraoch.!",
    "349329": "Congratulations.....",
    "349344": "Congrats KazAnova and your team  and thanks for sharing !",
    "349405": "Well done and thanks for sharing.",
    "349412": "Congratulations. Thanks for sharing.",
    "349420": "Congratulations @KazAnova and team. Thanks for sharing.",
    "349450": "Congrats guys and thx for the nice writeup!",
    "349599": "Congratulations and thanks for sharing!",
    "349620": "Well done",
    "349713": "Congratulation on your 3rd place Kazanova! <br>\nMaybe not your favorite position(4th), but I can only imagine that winning prizes feels really good :) <br>\nYour solution is really interesting, as I could not come up of a good way to set up a time-series model. <br>\n\nI have a few questions about stacking (your team has brilliant stackers)<br>\n1) What was your reason behind adding some counts and likelihoods to your stackers?<br>\n   Or, as a general question, what features should we add to the metamodels (level2 models)? Is it try and error?<br>\n2) About your 5fold models. When you merge teams, you often have different folds between teammates. <br>\n   Are we supposed to re-run all our models in the same folds? Or is it okay to use different folds and stack them?",
    "349800": "Thank you \n\nHaha, I guess I can live with a 3rd place :)\n\n1) Experimentally this often works . I guess this is because a model has a chance to look at this data from a slightly difference perspective when you already have powerful models in . In other words it tries to explore information it may have missed before.  It is mostly trial and error.\n\n2) No. you can see it as a form of bagging if you trained on different folds - to my experience it does not change the results. People can use different folds",
    "349812": "Congrats Kaza, Pavel and Raymond :-)\nGood work!",
    "349942": "congrats Kazanova!",
    "349983": "Congrats, Marios!!!",
    "350214": "Thanks you Kaza. I have a question. In the last part of \"LGBM features\" , Is the \"Likelihood\" means something like \"target encode\" ?",
    "350226": "Congratulations , Could you please tell me what are '' Some location features based on latitude and longtitude'', Thank you.",
    "350243": "Thank you for the reply!<br>\n2) is really interesting, because I heard some other Grand Master say the same thing.<br> \nHowever when I asked @Kohei, he told me that he definitely had experiences (from different competitions) where different folds lead to leakages (especially when there were highly predictive features)<br>\nPerhaps this is a controversial issue...?",
    "350254": "rhgrossm (To train them is my cause)\nCongratulation to you too! <br>\nCan I ask one question to you too? <br>\nI remember from the toxic comment competition, that you advocated trying pseudo-labeling in any competition. <br>\nThis is my second competition (after talkingdata), that I tried it, and didn't work... <br>\nDoes pseudo-labeling not work well in general for GDBT? Or is it that we don't have a good enough prediction? (but in talkingdata we had a +0.97AUC, and it still didn't work)<br>\nI am really lost, and I hope you can help me :)",
    "350292": "Yes, it is exactly that.",
    "350295": "Very similar to [these][1] \n\nWe created some additional clusters based on kmeans. \n\n  [1]: https://www.kaggle.com/frankherfert/region-and-city-details-with-lat-lon-and-clusters",
    "350297": "&gt; I remember from the toxic comment competition, that you advocated\n&gt; trying  pseudo-labeling in any competition.\n\nI tried that here and it did not work. \n\nI tried with both lgbt and nn\n\nI could see performance of  pseudo-labeling was worse with lgb than nns.",
    "350369": "Thanks.",
    "350383": "There are some competitions (especially when the data is not very big) where you can introduce leakage, no matter what cv you are doing. Sometimes it is good to ensure everyone is cving the same way to avoid any mistakes. Other than that, I have not seen huge differences, but I trust kohei if he says he has.",
    "350544": "Thanks again for all the replies!<br>\nHope to see you in future competitions again :)",
    "350562": "I am sure you will :)",
    "351103": "Congrats!",
    "357779": "Sorry I didn't respond! For pseudolabeling, you have to be careful about the amount of noise that you are introducing into the training and test set. With a problem like this, with a lot of high dimensional categoricals, introducing predictions is liable to introduce a high level of label noise. I actually found that using test predictions as targets led to a similar score as with zero pseudolabeling, which means we weren't gaining a lot of extra information at risk of overfitting. \n\nI didn't spend a ton of time on TD, but PL worked well there. How did you deal with adding your test targets?\n\nAnd, in general, GBDT can't handle pseudolabeling because it is not continuous. Hope that helps!",
    "364403": "Thanks for the reply :) <br>\nAll my effort into PL was using GBDT.... You should have told me earlier that it doesn't work :( hahaha <br>\nIt is encoraging that PL worked in TD though. We didn't try it with NN, so that was the issue.",
    "2180792": "Where can i find the code behind this submission? Almost every solution on the leaderboard has no notebook :("
  },
  "source": "meta"
}