{
  "id": 59914,
  "title": "5th place solution",
  "url": "/competitions/avito-demand-prediction/discussion/59914",
  "author_name": "Darragh",
  "post_date": "2018-06-28T09:15:15.418000",
  "votes": 71,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Congrats to top3 winners, and all top teams – its funny to see the solutions – I was pretty sure there was something massive we missed on images, but after a first read of other solutions it looks like there was no big thing missed there. <br>\nWhen we got into the competition we would have been pretty happy with a top50, so are psyched to end up with 5th. Our team are all working together at Optum Health (hence the name :) ) so it was great to benchmark some of our techniques here. <br>\nOur solution involved a LGBM stack of 4 different types of models – lgbm, RNN, MLP and ridge. The best of each scored on public LB approx <code>0.216, 0.2185, 0.2215 and 0.222</code> – but it was really in the stack where diversity between these helped. This is our repo <a href=\"https://github.com/darraghdog/avito-demand/\">linky</a>. We tracked progression for different models/changes in there on the front page. <br>\nFor teams starting in Kaggle or Data Science I cannot underestimate the importance of getting a good local validation that tracks to the Leaderboard and tracking improvement on val and lb. Initially we used a small validation set of a few of the final days of train; when we had a few good models we set up a stack and we used 5 fold with timesplit.    </p>\n\n<p><strong>MLP</strong> <br>\nIn general we leaned a lot on the Mercari solution’s it was a very similar problem. The winning MLP there scored 0.2215 here with very little changes ( <a href=\"https://github.com/darraghdog/avito-demand/blob/master/nnet/mlp_1705.py\">MLP code</a> ). Tried a few other things to improve, but it did not help in the stack. All credits to Konstantin and Pawel who developed this and for sharing a simple 75 line version <a href=\"https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\">linky</a>      </p>\n\n<p><strong>RNN</strong> <br>\nTim set up the RNN on a macbook GPU which was pretty impressive . We also introduced pymorphy2 which gave good improvement here and on lgb for tokenization. A lot of work was done on regularization for tuning and we concatenated on the penultimate layer of a densenet feature map of the images.  Example <a href=\"https://github.com/darraghdog/avito-demand/blob/master/nnet/rnntmp/nnetdh5CV_2705A.py\">RNN code</a> <br>\nWe tried adding numerical features and pretrained embeddings; while it helped at L1, it did not add much in the stack for the first few tries. One of the main challenges here was hardware – on 5CV it took about 36 hours to run on an AWS P2, as we bagged 2 times and used 256 wide embedding layer – so we gave up here and concentrated a bit more on LGB.   </p>\n\n<p><strong>LGB</strong> <br>\nFeatures engineering in LGB was the happy tree that did not stop giving. On the repo front page you can see the progression. <br>\nOne of the strongest was relative price. We did a kind of Bayesian mean of item price vs price of the group – <code>((item_price/mean_price_grp)*ct_grp + (prior))/(ct_grp+prior)</code>. This allowed the ratio be weighted on how many items were in the group - which is pretty important, if an item is the cheapest of 2 similar items its a lot less significant than being the cheapest of 100 similar items.  Just doing this over lots of different groups – title, params categories, clusters etc. added close to 0.002. <br>\nThis is an <a href=\"https://github.com/darraghdog/avito-demand/blob/master/features/code/pratioFestivitiesR1206.R\">example</a> of price ratio. <br>\nImage features in the public kernels helped a little – dullness, channel intensity etc. All credits to the author. <br>\nBayesian mean and counts over different groups helped maybe 0.001 also. We used a few combinations of tfidf for diversity. Entropy helped also. Dropping categoricals in their raw encoded form helped; and letting the model learn their representation through the FE mentioned. Also moving up to 1000 leaves helped some – 2000 or 5000 leaves probably would have helped more, but took too long to run.   </p>\n\n<p><strong>Ridge</strong> <br>\nLate in the game we set up a few ridge models; they were very fast. Although a lot weaker, they gave about 0.0005 on the stack. We had a separate model running on each parent category; models on image feature maps (vgg19 and densenet) – and on different types of count vectorizer and tfidf of text features.   </p>\n\n<p><strong>Stack</strong> <br>\nThe final stack just combined about 30 different models – also added all two way combinations of sums and differences of each model. We got an MLP stack scoring approximately the same but correlation to the lgb was very high so averaging did not help.   </p>\n\n<p><strong>Final note</strong> <br>\nOk, I was wondering this morning how teams have such long write-ups, so looking back I can see why. There was so much opportunity to try things – thanks to Avito for hosting this and I hope the Avito team get some benefit from the solutions and we’ll see you back on Kaggle with another competition soon! </p>",
  "messages": [
    {
      "id": 349552,
      "postDate": "2018-06-28T09:15:15.420Z",
      "content": "<p>Congrats to top3 winners, and all top teams – its funny to see the solutions – I was pretty sure there was something massive we missed on images, but after a first read of other solutions it looks like there was no big thing missed there. <br>\nWhen we got into the competition we would have been pretty happy with a top50, so are psyched to end up with 5th. Our team are all working together at Optum Health (hence the name :) ) so it was great to benchmark some of our techniques here. <br>\nOur solution involved a LGBM stack of 4 different types of models – lgbm, RNN, MLP and ridge. The best of each scored on public LB approx <code>0.216, 0.2185, 0.2215 and 0.222</code> – but it was really in the stack where diversity between these helped. This is our repo <a href=\"https://github.com/darraghdog/avito-demand/\">linky</a>. We tracked progression for different models/changes in there on the front page. <br>\nFor teams starting in Kaggle or Data Science I cannot underestimate the importance of getting a good local validation that tracks to the Leaderboard and tracking improvement on val and lb. Initially we used a small validation set of a few of the final days of train; when we had a few good models we set up a stack and we used 5 fold with timesplit.    </p>\n\n<p><strong>MLP</strong> <br>\nIn general we leaned a lot on the Mercari solution’s it was a very similar problem. The winning MLP there scored 0.2215 here with very little changes ( <a href=\"https://github.com/darraghdog/avito-demand/blob/master/nnet/mlp_1705.py\">MLP code</a> ). Tried a few other things to improve, but it did not help in the stack. All credits to Konstantin and Pawel who developed this and for sharing a simple 75 line version <a href=\"https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\">linky</a>      </p>\n\n<p><strong>RNN</strong> <br>\nTim set up the RNN on a macbook GPU which was pretty impressive . We also introduced pymorphy2 which gave good improvement here and on lgb for tokenization. A lot of work was done on regularization for tuning and we concatenated on the penultimate layer of a densenet feature map of the images.  Example <a href=\"https://github.com/darraghdog/avito-demand/blob/master/nnet/rnntmp/nnetdh5CV_2705A.py\">RNN code</a> <br>\nWe tried adding numerical features and pretrained embeddings; while it helped at L1, it did not add much in the stack for the first few tries. One of the main challenges here was hardware – on 5CV it took about 36 hours to run on an AWS P2, as we bagged 2 times and used 256 wide embedding layer – so we gave up here and concentrated a bit more on LGB.   </p>\n\n<p><strong>LGB</strong> <br>\nFeatures engineering in LGB was the happy tree that did not stop giving. On the repo front page you can see the progression. <br>\nOne of the strongest was relative price. We did a kind of Bayesian mean of item price vs price of the group – <code>((item_price/mean_price_grp)*ct_grp + (prior))/(ct_grp+prior)</code>. This allowed the ratio be weighted on how many items were in the group - which is pretty important, if an item is the cheapest of 2 similar items its a lot less significant than being the cheapest of 100 similar items.  Just doing this over lots of different groups – title, params categories, clusters etc. added close to 0.002. <br>\nThis is an <a href=\"https://github.com/darraghdog/avito-demand/blob/master/features/code/pratioFestivitiesR1206.R\">example</a> of price ratio. <br>\nImage features in the public kernels helped a little – dullness, channel intensity etc. All credits to the author. <br>\nBayesian mean and counts over different groups helped maybe 0.001 also. We used a few combinations of tfidf for diversity. Entropy helped also. Dropping categoricals in their raw encoded form helped; and letting the model learn their representation through the FE mentioned. Also moving up to 1000 leaves helped some – 2000 or 5000 leaves probably would have helped more, but took too long to run.   </p>\n\n<p><strong>Ridge</strong> <br>\nLate in the game we set up a few ridge models; they were very fast. Although a lot weaker, they gave about 0.0005 on the stack. We had a separate model running on each parent category; models on image feature maps (vgg19 and densenet) – and on different types of count vectorizer and tfidf of text features.   </p>\n\n<p><strong>Stack</strong> <br>\nThe final stack just combined about 30 different models – also added all two way combinations of sums and differences of each model. We got an MLP stack scoring approximately the same but correlation to the lgb was very high so averaging did not help.   </p>\n\n<p><strong>Final note</strong> <br>\nOk, I was wondering this morning how teams have such long write-ups, so looking back I can see why. There was so much opportunity to try things – thanks to Avito for hosting this and I hope the Avito team get some benefit from the solutions and we’ll see you back on Kaggle with another competition soon! </p>",
      "rawMarkdown": "Congrats to top3 winners, and all top teams – its funny to see the solutions – I was pretty sure there was something massive we missed on images, but after a first read of other solutions it looks like there was no big thing missed there.   \nWhen we got into the competition we would have been pretty happy with a top50, so are psyched to end up with 5th. Our team are all working together at Optum Health (hence the name :) ) so it was great to benchmark some of our techniques here.    \nOur solution involved a LGBM stack of 4 different types of models – lgbm, RNN, MLP and ridge. The best of each scored on public LB approx `0.216, 0.2185, 0.2215 and 0.222` – but it was really in the stack where diversity between these helped. This is our repo [linky][1]. We tracked progression for different models/changes in there on the front page.    \nFor teams starting in Kaggle or Data Science I cannot underestimate the importance of getting a good local validation that tracks to the Leaderboard and tracking improvement on val and lb. Initially we used a small validation set of a few of the final days of train; when we had a few good models we set up a stack and we used 5 fold with timesplit.    \n \n**MLP**    \nIn general we leaned a lot on the Mercari solution’s it was a very similar problem. The winning MLP there scored 0.2215 here with very little changes ( [MLP code][2] ). Tried a few other things to improve, but it did not help in the stack. All credits to Konstantin and Pawel who developed this and for sharing a simple 75 line version [linky][3]      \n\n**RNN**    \nTim set up the RNN on a macbook GPU which was pretty impressive . We also introduced pymorphy2 which gave good improvement here and on lgb for tokenization. A lot of work was done on regularization for tuning and we concatenated on the penultimate layer of a densenet feature map of the images.  Example [RNN code][4]    \nWe tried adding numerical features and pretrained embeddings; while it helped at L1, it did not add much in the stack for the first few tries. One of the main challenges here was hardware – on 5CV it took about 36 hours to run on an AWS P2, as we bagged 2 times and used 256 wide embedding layer – so we gave up here and concentrated a bit more on LGB.   \n\n**LGB**   \nFeatures engineering in LGB was the happy tree that did not stop giving. On the repo front page you can see the progression.    \nOne of the strongest was relative price. We did a kind of Bayesian mean of item price vs price of the group – `((item_price/mean_price_grp)*ct_grp + (prior))/(ct_grp+prior)`. This allowed the ratio be weighted on how many items were in the group - which is pretty important, if an item is the cheapest of 2 similar items its a lot less significant than being the cheapest of 100 similar items.  Just doing this over lots of different groups – title, params categories, clusters etc. added close to 0.002.    \nThis is an [example][5] of price ratio.   \nImage features in the public kernels helped a little – dullness, channel intensity etc. All credits to the author.     \nBayesian mean and counts over different groups helped maybe 0.001 also. We used a few combinations of tfidf for diversity. Entropy helped also. Dropping categoricals in their raw encoded form helped; and letting the model learn their representation through the FE mentioned. Also moving up to 1000 leaves helped some – 2000 or 5000 leaves probably would have helped more, but took too long to run.   \n\n**Ridge**    \nLate in the game we set up a few ridge models; they were very fast. Although a lot weaker, they gave about 0.0005 on the stack. We had a separate model running on each parent category; models on image feature maps (vgg19 and densenet) – and on different types of count vectorizer and tfidf of text features.   \n\n**Stack**   \nThe final stack just combined about 30 different models – also added all two way combinations of sums and differences of each model. We got an MLP stack scoring approximately the same but correlation to the lgb was very high so averaging did not help.   \n\n**Final note**  \nOk, I was wondering this morning how teams have such long write-ups, so looking back I can see why. There was so much opportunity to try things – thanks to Avito for hosting this and I hope the Avito team get some benefit from the solutions and we’ll see you back on Kaggle with another competition soon! \n\n\n  [1]: https://github.com/darraghdog/avito-demand/\n  [2]: https://github.com/darraghdog/avito-demand/blob/master/nnet/mlp_1705.py\n  [3]: https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\n  [4]: https://github.com/darraghdog/avito-demand/blob/master/nnet/rnntmp/nnetdh5CV_2705A.py\n  [5]: https://github.com/darraghdog/avito-demand/blob/master/features/code/pratioFestivitiesR1206.R",
      "votes": 71
    },
    {
      "id": 351560,
      "postDate": "2018-07-02T13:56:14.213Z",
      "content": "<p>Thank you for the code, it was very helpful!</p>",
      "rawMarkdown": "Thank you for the code, it was very helpful!",
      "votes": 1
    },
    {
      "id": 350221,
      "postDate": "2018-06-29T12:32:22.483Z",
      "content": "<p>Thank you for explanation and code. Great job!</p>\n\n<p>Could you clarify, what group size have used for parameter <code>prior</code> in Bayesian formula?</p>\n\n<p>Unfortunately, provided example link is not working.</p>",
      "rawMarkdown": "Thank you for explanation and code. Great job!\n\nCould you clarify, what group size have used for parameter `prior` in Bayesian formula?\n\nUnfortunately, provided example link is not working.",
      "votes": 1,
      "replies": [
        {
          "id": 350300,
          "postDate": "2018-06-29T14:15:58.990Z",
          "content": "<p>hey, sorry, I cleaned up the repo yesterday, and moved the link - but its corrected now... depending on the groups I pick different priors... or very granular groups, like where they all have the same <code>user_id</code> &amp; <code>title</code>, I would pick a smaller prior, maybe 5 or 10; and on larger groups; like <code>category_name</code>, a larger one of 100 may be used.  </p>",
          "rawMarkdown": "hey, sorry, I cleaned up the repo yesterday, and moved the link - but its corrected now... depending on the groups I pick different priors... or very granular groups, like where they all have the same `user_id` &amp; `title`, I would pick a smaller prior, maybe 5 or 10; and on larger groups; like `category_name`, a larger one of 100 may be used.  ",
          "votes": 1
        },
        {
          "id": 350461,
          "postDate": "2018-06-29T19:44:15.107Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 350003,
      "postDate": "2018-06-29T02:33:34.753Z",
      "content": "<p>WOW! \nAwesome work!\nCongrats! Thanks for sharing code.</p>",
      "rawMarkdown": "WOW! \nAwesome work!\nCongrats! Thanks for sharing code.",
      "votes": 1
    },
    {
      "id": 349992,
      "postDate": "2018-06-29T01:49:29.430Z",
      "content": "<p>Thank you very much for the description and links. Very big congratulations to your team ! </p>",
      "rawMarkdown": "Thank you very much for the description and links. Very big congratulations to your team ! ",
      "votes": 1
    },
    {
      "id": 349854,
      "postDate": "2018-06-28T18:51:24.520Z",
      "content": "<p>For the \"Bayesian mean of item price vs price of the group\" how did you get 20 for the prior? Thank you</p>",
      "rawMarkdown": "For the \"Bayesian mean of item price vs price of the group\" how did you get 20 for the prior? Thank you",
      "votes": 1,
      "replies": [
        {
          "id": 349857,
          "postDate": "2018-06-28T19:03:33.970Z",
          "content": "<p>it's a bit of a guess.... think of it like any groups below the prior will be sort of dampened; and the more lower than the prior the group count is the more it will be dampened. To really see it; its good to set up an excel sheet with a couple of groups with different prices and run the formula over it to calculate the bayes mean for each item depending on the group size. You'll see how moving up and down the prior affects the high count and low count group values. </p>",
          "rawMarkdown": "it's a bit of a guess.... think of it like any groups below the prior will be sort of dampened; and the more lower than the prior the group count is the more it will be dampened. To really see it; its good to set up an excel sheet with a couple of groups with different prices and run the formula over it to calculate the bayes mean for each item depending on the group size. You'll see how moving up and down the prior affects the high count and low count group values. ",
          "votes": 2
        },
        {
          "id": 349865,
          "postDate": "2018-06-28T19:21:44.317Z",
          "content": "<p>Ok, thank you, I want to try to experiment with doing this for the Home Credit Default Risk challenge...also nice profile picture!!!! =)</p>",
          "rawMarkdown": "Ok, thank you, I want to try to experiment with doing this for the Home Credit Default Risk challenge...also nice profile picture!!!! =)",
          "votes": 1
        }
      ]
    },
    {
      "id": 349802,
      "postDate": "2018-06-28T16:59:38.303Z",
      "content": "<p><code>df[cols] = df[cols].astype(str).fillna('nicapotato') # FILL NA</code> <br>\nFlattered to have sneaked into this LGBM feature engineering :D!! Thank you so much for sharing, so much to learn here.</p>",
      "rawMarkdown": "`df[cols] = df[cols].astype(str).fillna('nicapotato') # FILL NA` <br>\nFlattered to have sneaked into this LGBM feature engineering :D!! Thank you so much for sharing, so much to learn here.",
      "votes": 1,
      "replies": [
        {
          "id": 349805,
          "postDate": "2018-06-28T17:03:38.460Z",
          "content": "<p>Thank you very much Nick for publishing the script; we used it as basis for LGBM. </p>",
          "rawMarkdown": "Thank you very much Nick for publishing the script; we used it as basis for LGBM. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 349714,
      "postDate": "2018-06-28T14:27:56.147Z",
      "content": "<p>Congratulations @Darragh and team. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations @Darragh and team. Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 349634,
      "postDate": "2018-06-28T12:15:12.367Z",
      "content": "<p>Could you describe how you processed the images?</p>",
      "rawMarkdown": "Could you describe how you processed the images?",
      "votes": 1,
      "replies": [
        {
          "id": 349671,
          "postDate": "2018-06-28T13:19:12.023Z",
          "content": "<p>Code here to extract densenet featuremap : <a href=\"https://github.com/darraghdog/avito-demand/blob/master/imgfeatures/densenet_extractor.py\">https://github.com/darraghdog/avito-demand/blob/master/imgfeatures/densenet_extractor.py</a>\nJust load pretrained model; chop off the end, like in the code line 47; and then run all the images through and save the output. Each image will give an matrix of shape for example <code>32x32</code> but use numpy to flatten it and you get (1*1024) - do the same for the 130K images and you get a matrix 130K*1024 dimensions; then concat that to your RNN of whatever. We filled blank images with 0. </p>",
          "rawMarkdown": "Code here to extract densenet featuremap : https://github.com/darraghdog/avito-demand/blob/master/imgfeatures/densenet_extractor.py\nJust load pretrained model; chop off the end, like in the code line 47; and then run all the images through and save the output. Each image will give an matrix of shape for example `32x32` but use numpy to flatten it and you get (1*1024) - do the same for the 130K images and you get a matrix 130K*1024 dimensions; then concat that to your RNN of whatever. We filled blank images with 0. ",
          "votes": 1
        },
        {
          "id": 349797,
          "postDate": "2018-06-28T16:48:38.987Z",
          "content": "<p>Thank you for sharing. Congrats on your success. </p>",
          "rawMarkdown": "Thank you for sharing. Congrats on your success. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 349633,
      "postDate": "2018-06-28T12:14:08.427Z",
      "content": "<p>Thank you for the detailed write-up. Amazing work as always ! \nI only have a small request, could you tell us more about the hardware used throughout all the phases of the competition ? </p>",
      "rawMarkdown": "Thank you for the detailed write-up. Amazing work as always ! \nI only have a small request, could you tell us more about the hardware used throughout all the phases of the competition ? ",
      "votes": 1,
      "replies": [
        {
          "id": 349648,
          "postDate": "2018-06-28T12:44:33.777Z",
          "content": "<p>Darragh can touch on the AWS instance we occasionally shared but I ran most of my computation from my laptop. Nothing notable about it except the GTX 970M video card which was helpful to run the RNNs on. 12 GB of RAM was hardly enough to run the scripts so I constantly ran into memory problems.</p>",
          "rawMarkdown": "Darragh can touch on the AWS instance we occasionally shared but I ran most of my computation from my laptop. Nothing notable about it except the GTX 970M video card which was helpful to run the RNNs on. 12 GB of RAM was hardly enough to run the scripts so I constantly ran into memory problems.",
          "votes": 1
        },
        {
          "id": 349665,
          "postDate": "2018-06-28T13:09:37.427Z",
          "content": "<p>I got an 3.4xlarge instance, I believe - 16 core and 128GB RAM; about 28c per hour on spot pricing to run jobs when they were ready for 5CV. but got the jobs ready locally on the laptop and then kicked off the AWS instance. This was for the Lightgbm. Sometimes you can get vouchers for AWS, which helps on cost. </p>",
          "rawMarkdown": "I got an 3.4xlarge instance, I believe - 16 core and 128GB RAM; about 28c per hour on spot pricing to run jobs when they were ready for 5CV. but got the jobs ready locally on the laptop and then kicked off the AWS instance. This was for the Lightgbm. Sometimes you can get vouchers for AWS, which helps on cost. "
        }
      ]
    },
    {
      "id": 349566,
      "postDate": "2018-06-28T09:26:08.660Z",
      "content": "<p>thanks for the write-up. the ridge models are amazing. did you use the same features for each ridge model on different parent cateogry? </p>",
      "rawMarkdown": "thanks for the write-up. the ridge models are amazing. did you use the same features for each ridge model on different parent cateogry? ",
      "votes": 1,
      "replies": [
        {
          "id": 349568,
          "postDate": "2018-06-28T09:29:37.580Z",
          "content": "<p>Yeah I think we only ran one set of features on parent categories - given more time, more could have been helpful- I thnk this was set up Monday night :) . One of the parent categories took about 60% of the data; so we split that up to its subcategories and ran a second model - I think it was clothes, where we split to kids clothes and adults clothes. </p>",
          "rawMarkdown": "Yeah I think we only ran one set of features on parent categories - given more time, more could have been helpful- I thnk this was set up Monday night :) . One of the parent categories took about 60% of the data; so we split that up to its subcategories and ran a second model - I think it was clothes, where we split to kids clothes and adults clothes. ",
          "votes": 1
        },
        {
          "id": 349572,
          "postDate": "2018-06-28T09:44:07.603Z",
          "content": "<p>I observed the same thing 2-3 5rdtays ago and tried to group parent category with less observation but I used lgb models. obviously, it takes too long to finish for all categories so I drop this idea. your ridge add 0.0005 to the stack is really good! </p>",
          "rawMarkdown": "I observed the same thing 2-3 5rdtays ago and tried to group parent category with less observation but I used lgb models. obviously, it takes too long to finish for all categories so I drop this idea. your ridge add 0.0005 to the stack is really good! ",
          "votes": 1
        },
        {
          "id": 349578,
          "postDate": "2018-06-28T09:54:51.937Z",
          "content": "<p>Yeah, we actually ran it with lightgbm also, same features but ran it on parent parent category only. I was thinking to try on all categories - ridge probably would have been fast enough to do it. </p>",
          "rawMarkdown": "Yeah, we actually ran it with lightgbm also, same features but ran it on parent parent category only. I was thinking to try on all categories - ridge probably would have been fast enough to do it. "
        }
      ]
    },
    {
      "id": 349825,
      "postDate": "2018-06-28T17:28:46.160Z",
      "content": "<p>Nice solution Darragh. Congrats!!!</p>",
      "rawMarkdown": "Nice solution Darragh. Congrats!!!",
      "votes": 2
    },
    {
      "id": 364052,
      "postDate": "2018-07-30T15:58:11.350Z",
      "content": "<p>I see most of the competition winners use Python for coding. Most of the things I see here, like text mining, modeling etc., can also be done in R. Could you tell me why you chose Python and not R? </p>",
      "rawMarkdown": "I see most of the competition winners use Python for coding. Most of the things I see here, like text mining, modeling etc., can also be done in R. Could you tell me why you chose Python and not R? ",
      "replies": [
        {
          "id": 364058,
          "postDate": "2018-07-30T16:07:54.533Z",
          "content": "<p>I personally have more experience with Python and to my taste the packages we use have much better APIs through Python than R (for instance Keras). It is also not uncommon for R packages to be several versions behind their Python counterparts. </p>\n\n<p>Overall I use R for quick and dirty analysis/ fitting simple models  and switch to Python for heavier lifting </p>",
          "rawMarkdown": "I personally have more experience with Python and to my taste the packages we use have much better APIs through Python than R (for instance Keras). It is also not uncommon for R packages to be several versions behind their Python counterparts. \n\nOverall I use R for quick and dirty analysis/ fitting simple models  and switch to Python for heavier lifting "
        },
        {
          "id": 364087,
          "postDate": "2018-07-30T17:43:59.480Z",
          "content": "<p>Got it. Thanks! :)</p>",
          "rawMarkdown": "Got it. Thanks! :)"
        }
      ]
    },
    {
      "id": 350205,
      "postDate": "2018-06-29T11:56:55.007Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 349909,
      "postDate": "2018-06-28T20:31:29.073Z",
      "content": "<p>Congrats! Thanks for posting code:)</p>",
      "rawMarkdown": "Congrats! Thanks for posting code:)",
      "votes": 1
    },
    {
      "id": 349621,
      "postDate": "2018-06-28T11:41:08.430Z",
      "content": "<p>Thank you for very detailed solution.</p>",
      "rawMarkdown": "Thank you for very detailed solution.\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 351560,
      "author_name": "Jeffery Cao",
      "author_url": "",
      "post_date": "2018-07-02T13:56:14.213000",
      "content": "<p>Thank you for the code, it was very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 350221,
      "author_name": "Sasha Turutin",
      "author_url": "",
      "post_date": "2018-06-29T12:32:22.483000",
      "content": "<p>Thank you for explanation and code. Great job!</p>\n\n<p>Could you clarify, what group size have used for parameter <code>prior</code> in Bayesian formula?</p>\n\n<p>Unfortunately, provided example link is not working.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 350300,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-29T14:15:58.990000",
          "content": "<p>hey, sorry, I cleaned up the repo yesterday, and moved the link - but its corrected now... depending on the groups I pick different priors... or very granular groups, like where they all have the same <code>user_id</code> &amp; <code>title</code>, I would pick a smaller prior, maybe 5 or 10; and on larger groups; like <code>category_name</code>, a larger one of 100 may be used.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 350461,
          "author_name": "Sasha Turutin",
          "author_url": "",
          "post_date": "2018-06-29T19:44:15.107000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 350003,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "2018-06-29T02:33:34.753000",
      "content": "<p>WOW! \nAwesome work!\nCongrats! Thanks for sharing code.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349992,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2018-06-29T01:49:29.430000",
      "content": "<p>Thank you very much for the description and links. Very big congratulations to your team ! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349854,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2018-06-28T18:51:24.520000",
      "content": "<p>For the \"Bayesian mean of item price vs price of the group\" how did you get 20 for the prior? Thank you</p>",
      "votes": 1,
      "replies": [
        {
          "id": 349857,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T19:03:33.970000",
          "content": "<p>it's a bit of a guess.... think of it like any groups below the prior will be sort of dampened; and the more lower than the prior the group count is the more it will be dampened. To really see it; its good to set up an excel sheet with a couple of groups with different prices and run the formula over it to calculate the bayes mean for each item depending on the group size. You'll see how moving up and down the prior affects the high count and low count group values. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 349865,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2018-06-28T19:21:44.317000",
          "content": "<p>Ok, thank you, I want to try to experiment with doing this for the Home Credit Default Risk challenge...also nice profile picture!!!! =)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 349802,
      "author_name": "nicapotato",
      "author_url": "",
      "post_date": "2018-06-28T16:59:38.303000",
      "content": "<p><code>df[cols] = df[cols].astype(str).fillna('nicapotato') # FILL NA</code> <br>\nFlattered to have sneaked into this LGBM feature engineering :D!! Thank you so much for sharing, so much to learn here.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 349805,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T17:03:38.460000",
          "content": "<p>Thank you very much Nick for publishing the script; we used it as basis for LGBM. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 349714,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-06-28T14:27:56.147000",
      "content": "<p>Congratulations @Darragh and team. Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349634,
      "author_name": "Tapioca",
      "author_url": "",
      "post_date": "2018-06-28T12:15:12.367000",
      "content": "<p>Could you describe how you processed the images?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 349671,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T13:19:12.023000",
          "content": "<p>Code here to extract densenet featuremap : <a href=\"https://github.com/darraghdog/avito-demand/blob/master/imgfeatures/densenet_extractor.py\">https://github.com/darraghdog/avito-demand/blob/master/imgfeatures/densenet_extractor.py</a>\nJust load pretrained model; chop off the end, like in the code line 47; and then run all the images through and save the output. Each image will give an matrix of shape for example <code>32x32</code> but use numpy to flatten it and you get (1*1024) - do the same for the 130K images and you get a matrix 130K*1024 dimensions; then concat that to your RNN of whatever. We filled blank images with 0. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349797,
          "author_name": "Tapioca",
          "author_url": "",
          "post_date": "2018-06-28T16:48:38.987000",
          "content": "<p>Thank you for sharing. Congrats on your success. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 349633,
      "author_name": "Zakaria EL Mesaoudi",
      "author_url": "",
      "post_date": "2018-06-28T12:14:08.427000",
      "content": "<p>Thank you for the detailed write-up. Amazing work as always ! \nI only have a small request, could you tell us more about the hardware used throughout all the phases of the competition ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 349648,
          "author_name": "Tim Rosenflanz",
          "author_url": "",
          "post_date": "2018-06-28T12:44:33.777000",
          "content": "<p>Darragh can touch on the AWS instance we occasionally shared but I ran most of my computation from my laptop. Nothing notable about it except the GTX 970M video card which was helpful to run the RNNs on. 12 GB of RAM was hardly enough to run the scripts so I constantly ran into memory problems.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349665,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T13:09:37.427000",
          "content": "<p>I got an 3.4xlarge instance, I believe - 16 core and 128GB RAM; about 28c per hour on spot pricing to run jobs when they were ready for 5CV. but got the jobs ready locally on the laptop and then kicked off the AWS instance. This was for the Lightgbm. Sometimes you can get vouchers for AWS, which helps on cost. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349566,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "2018-06-28T09:26:08.660000",
      "content": "<p>thanks for the write-up. the ridge models are amazing. did you use the same features for each ridge model on different parent cateogry? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 349568,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T09:29:37.580000",
          "content": "<p>Yeah I think we only ran one set of features on parent categories - given more time, more could have been helpful- I thnk this was set up Monday night :) . One of the parent categories took about 60% of the data; so we split that up to its subcategories and ran a second model - I think it was clothes, where we split to kids clothes and adults clothes. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349572,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2018-06-28T09:44:07.603000",
          "content": "<p>I observed the same thing 2-3 5rdtays ago and tried to group parent category with less observation but I used lgb models. obviously, it takes too long to finish for all categories so I drop this idea. your ridge add 0.0005 to the stack is really good! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349578,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T09:54:51.937000",
          "content": "<p>Yeah, we actually ran it with lightgbm also, same features but ran it on parent parent category only. I was thinking to try on all categories - ridge probably would have been fast enough to do it. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349825,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-06-28T17:28:46.160000",
      "content": "<p>Nice solution Darragh. Congrats!!!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 364052,
      "author_name": "KaivalyaKandukuri",
      "author_url": "",
      "post_date": "2018-07-30T15:58:11.350000",
      "content": "<p>I see most of the competition winners use Python for coding. Most of the things I see here, like text mining, modeling etc., can also be done in R. Could you tell me why you chose Python and not R? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 364058,
          "author_name": "Tim Rosenflanz",
          "author_url": "",
          "post_date": "2018-07-30T16:07:54.533000",
          "content": "<p>I personally have more experience with Python and to my taste the packages we use have much better APIs through Python than R (for instance Keras). It is also not uncommon for R packages to be several versions behind their Python counterparts. </p>\n\n<p>Overall I use R for quick and dirty analysis/ fitting simple models  and switch to Python for heavier lifting </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364087,
          "author_name": "KaivalyaKandukuri",
          "author_url": "",
          "post_date": "2018-07-30T17:43:59.480000",
          "content": "<p>Got it. Thanks! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 350205,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-29T11:56:55.007000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349909,
      "author_name": "Devin Anzelmo",
      "author_url": "",
      "post_date": "2018-06-28T20:31:29.073000",
      "content": "<p>Congrats! Thanks for posting code:)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349621,
      "author_name": "Pooh",
      "author_url": "",
      "post_date": "2018-06-28T11:41:08.430000",
      "content": "<p>Thank you for very detailed solution.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349552": "Congrats to top3 winners, and all top teams – its funny to see the solutions – I was pretty sure there was something massive we missed on images, but after a first read of other solutions it looks like there was no big thing missed there.   \nWhen we got into the competition we would have been pretty happy with a top50, so are psyched to end up with 5th. Our team are all working together at Optum Health (hence the name :) ) so it was great to benchmark some of our techniques here.    \nOur solution involved a LGBM stack of 4 different types of models – lgbm, RNN, MLP and ridge. The best of each scored on public LB approx `0.216, 0.2185, 0.2215 and 0.222` – but it was really in the stack where diversity between these helped. This is our repo [linky][1]. We tracked progression for different models/changes in there on the front page.    \nFor teams starting in Kaggle or Data Science I cannot underestimate the importance of getting a good local validation that tracks to the Leaderboard and tracking improvement on val and lb. Initially we used a small validation set of a few of the final days of train; when we had a few good models we set up a stack and we used 5 fold with timesplit.    \n \n**MLP**    \nIn general we leaned a lot on the Mercari solution’s it was a very similar problem. The winning MLP there scored 0.2215 here with very little changes ( [MLP code][2] ). Tried a few other things to improve, but it did not help in the stack. All credits to Konstantin and Pawel who developed this and for sharing a simple 75 line version [linky][3]      \n\n**RNN**    \nTim set up the RNN on a macbook GPU which was pretty impressive . We also introduced pymorphy2 which gave good improvement here and on lgb for tokenization. A lot of work was done on regularization for tuning and we concatenated on the penultimate layer of a densenet feature map of the images.  Example [RNN code][4]    \nWe tried adding numerical features and pretrained embeddings; while it helped at L1, it did not add much in the stack for the first few tries. One of the main challenges here was hardware – on 5CV it took about 36 hours to run on an AWS P2, as we bagged 2 times and used 256 wide embedding layer – so we gave up here and concentrated a bit more on LGB.   \n\n**LGB**   \nFeatures engineering in LGB was the happy tree that did not stop giving. On the repo front page you can see the progression.    \nOne of the strongest was relative price. We did a kind of Bayesian mean of item price vs price of the group – `((item_price/mean_price_grp)*ct_grp + (prior))/(ct_grp+prior)`. This allowed the ratio be weighted on how many items were in the group - which is pretty important, if an item is the cheapest of 2 similar items its a lot less significant than being the cheapest of 100 similar items.  Just doing this over lots of different groups – title, params categories, clusters etc. added close to 0.002.    \nThis is an [example][5] of price ratio.   \nImage features in the public kernels helped a little – dullness, channel intensity etc. All credits to the author.     \nBayesian mean and counts over different groups helped maybe 0.001 also. We used a few combinations of tfidf for diversity. Entropy helped also. Dropping categoricals in their raw encoded form helped; and letting the model learn their representation through the FE mentioned. Also moving up to 1000 leaves helped some – 2000 or 5000 leaves probably would have helped more, but took too long to run.   \n\n**Ridge**    \nLate in the game we set up a few ridge models; they were very fast. Although a lot weaker, they gave about 0.0005 on the stack. We had a separate model running on each parent category; models on image feature maps (vgg19 and densenet) – and on different types of count vectorizer and tfidf of text features.   \n\n**Stack**   \nThe final stack just combined about 30 different models – also added all two way combinations of sums and differences of each model. We got an MLP stack scoring approximately the same but correlation to the lgb was very high so averaging did not help.   \n\n**Final note**  \nOk, I was wondering this morning how teams have such long write-ups, so looking back I can see why. There was so much opportunity to try things – thanks to Avito for hosting this and I hope the Avito team get some benefit from the solutions and we’ll see you back on Kaggle with another competition soon! \n\n\n  [1]: https://github.com/darraghdog/avito-demand/\n  [2]: https://github.com/darraghdog/avito-demand/blob/master/nnet/mlp_1705.py\n  [3]: https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\n  [4]: https://github.com/darraghdog/avito-demand/blob/master/nnet/rnntmp/nnetdh5CV_2705A.py\n  [5]: https://github.com/darraghdog/avito-demand/blob/master/features/code/pratioFestivitiesR1206.R",
    "351560": "Thank you for the code, it was very helpful!",
    "350221": "Thank you for explanation and code. Great job!\n\nCould you clarify, what group size have used for parameter `prior` in Bayesian formula?\n\nUnfortunately, provided example link is not working.",
    "350003": "WOW! \nAwesome work!\nCongrats! Thanks for sharing code.",
    "349992": "Thank you very much for the description and links. Very big congratulations to your team ! ",
    "349854": "For the \"Bayesian mean of item price vs price of the group\" how did you get 20 for the prior? Thank you",
    "349802": "`df[cols] = df[cols].astype(str).fillna('nicapotato') # FILL NA` <br>\nFlattered to have sneaked into this LGBM feature engineering :D!! Thank you so much for sharing, so much to learn here.",
    "349714": "Congratulations @Darragh and team. Thanks for sharing.",
    "349634": "Could you describe how you processed the images?",
    "349633": "Thank you for the detailed write-up. Amazing work as always ! \nI only have a small request, could you tell us more about the hardware used throughout all the phases of the competition ? ",
    "349566": "thanks for the write-up. the ridge models are amazing. did you use the same features for each ridge model on different parent cateogry? ",
    "349825": "Nice solution Darragh. Congrats!!!",
    "364052": "I see most of the competition winners use Python for coding. Most of the things I see here, like text mining, modeling etc., can also be done in R. Could you tell me why you chose Python and not R? ",
    "350205": "",
    "349909": "Congrats! Thanks for posting code:)",
    "349621": "Thank you for very detailed solution.\n"
  }
}