{
  "id": 59936,
  "title": "20th place solution",
  "url": "/competitions/avito-demand-prediction/writeups/ram-20th-place-solution",
  "author_name": "",
  "post_date": "2018-06-28T13:09:55.213989700Z",
  "votes": 27,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Firstly thank you to Avito for hosting such an awesome competition and thank you to both my teammates for all their hard work. I'll try to keep this brief.</p>\n\n<p><strong>Price model from active data</strong></p>\n\n<p>I actually wrote about this <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57410\">here</a> and am surprised that many others didn't do this - it was perhaps our single biggest success. Whilst a price model on active data was quite poor in error metric terms it made a large contribution to our score. We used NNs on all train and test active data and Rob trained a per category xgb price model after extracting lots of text features (numerical fields for things like cars/houses etc...) which helped a lot.</p>\n\n<p><strong>NNs and LGBM/XGB</strong></p>\n\n<p>Like others we had good success with NNs and trained a variety on different text embeddings (fasttext cc and wiki, self-trained w2v with and without stemming). We didn't add many other features to the NNs but did use a variety of architectures with and without attention. Modelling raw log1p(price) as a categorical and then embedding it worked well too, as did lower batch sizes and averaging predictions over several epochs. It was my first time using DL in a competition so I was very happy to be reaching ~0.2195 with a NN without many new features.</p>\n\n<p><strong>CV/stacking</strong></p>\n\n<p>Everything was done on a 5 fold basis and using a lgbm stacker. We got a lot of success from this but perhaps could have been more careful in trying to construct diverse base models and using a more rigorous feed forward selection. Antoine in particular did a lot of work ensuring we were robust in what we finally included.</p>\n\n<p><strong>Other thoughts</strong></p>\n\n<ul>\n<li>We did try to train models to predict period from the active data to no avail.</li>\n<li>We didn't use images beyond basic features</li>\n<li>We tried translating all the text, probably too late, but google's api gave up on us. The hope was the translation, whilst average, would help with the cases in the Russian language - as well as letting us finally understand what was being sold!</li>\n<li>We trained models on a few different targets (e.g a NN to predict if an item was a 0 or not, treating target as categorical) which then went into the stack. The pdf entropy from these classifications was a good feature also.</li>\n<li>We did train a user_id model (only using users that were in test and train) but ended up excluding any models which used user_id as were still finding too unstable CV-LB relationship. Be interested to hear if anyone included it successfully.</li>\n<li>I second some of the other comments about workflow: we saved all our transformations down and then simple read in .pkl files and concatenated. Some of the transformations were expensive and this allowed us to get going quickly.</li>\n<li><a href=\"https://radimrehurek.com/gensim/\">Gensim</a> is a cool package, though slightly steeper learning curve than sklearn.</li>\n<li>A note on lgbm: we actually found better results by sometimes excluding strong categorical features, particularly from the stack. My intuition for this is that given lgb fits in a leaf-wise (best first) manner it perhaps over-uses categorical features which always offer easy splits but ultimately might not be the best choice/don't allow slower exploration of weaker features or more subtle relationships. Keen to hear what others think of this.</li>\n</ul>\n\n<p>I'll let Rob and Antoine add anything if they see fit. Thanks again for a fun competition!</p>\n\n<p>Mark</p>",
  "messages": [
    {
      "id": "349666",
      "postDate": "06/28/2018 13:09:55",
      "content": "<p>Firstly thank you to Avito for hosting such an awesome competition and thank you to both my teammates for all their hard work. I'll try to keep this brief.</p>\n\n<p><strong>Price model from active data</strong></p>\n\n<p>I actually wrote about this <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57410\">here</a> and am surprised that many others didn't do this - it was perhaps our single biggest success. Whilst a price model on active data was quite poor in error metric terms it made a large contribution to our score. We used NNs on all train and test active data and Rob trained a per category xgb price model after extracting lots of text features (numerical fields for things like cars/houses etc...) which helped a lot.</p>\n\n<p><strong>NNs and LGBM/XGB</strong></p>\n\n<p>Like others we had good success with NNs and trained a variety on different text embeddings (fasttext cc and wiki, self-trained w2v with and without stemming). We didn't add many other features to the NNs but did use a variety of architectures with and without attention. Modelling raw log1p(price) as a categorical and then embedding it worked well too, as did lower batch sizes and averaging predictions over several epochs. It was my first time using DL in a competition so I was very happy to be reaching ~0.2195 with a NN without many new features.</p>\n\n<p><strong>CV/stacking</strong></p>\n\n<p>Everything was done on a 5 fold basis and using a lgbm stacker. We got a lot of success from this but perhaps could have been more careful in trying to construct diverse base models and using a more rigorous feed forward selection. Antoine in particular did a lot of work ensuring we were robust in what we finally included.</p>\n\n<p><strong>Other thoughts</strong></p>\n\n<ul>\n<li>We did try to train models to predict period from the active data to no avail.</li>\n<li>We didn't use images beyond basic features</li>\n<li>We tried translating all the text, probably too late, but google's api gave up on us. The hope was the translation, whilst average, would help with the cases in the Russian language - as well as letting us finally understand what was being sold!</li>\n<li>We trained models on a few different targets (e.g a NN to predict if an item was a 0 or not, treating target as categorical) which then went into the stack. The pdf entropy from these classifications was a good feature also.</li>\n<li>We did train a user_id model (only using users that were in test and train) but ended up excluding any models which used user_id as were still finding too unstable CV-LB relationship. Be interested to hear if anyone included it successfully.</li>\n<li>I second some of the other comments about workflow: we saved all our transformations down and then simple read in .pkl files and concatenated. Some of the transformations were expensive and this allowed us to get going quickly.</li>\n<li><a href=\"https://radimrehurek.com/gensim/\">Gensim</a> is a cool package, though slightly steeper learning curve than sklearn.</li>\n<li>A note on lgbm: we actually found better results by sometimes excluding strong categorical features, particularly from the stack. My intuition for this is that given lgb fits in a leaf-wise (best first) manner it perhaps over-uses categorical features which always offer easy splits but ultimately might not be the best choice/don't allow slower exploration of weaker features or more subtle relationships. Keen to hear what others think of this.</li>\n</ul>\n\n<p>I'll let Rob and Antoine add anything if they see fit. Thanks again for a fun competition!</p>\n\n<p>Mark</p>",
      "rawMarkdown": "Firstly thank you to Avito for hosting such an awesome competition and thank you to both my teammates for all their hard work. I'll try to keep this brief.\n\n**Price model from active data**\n\nI actually wrote about this [here][1] and am surprised that many others didn't do this - it was perhaps our single biggest success. Whilst a price model on active data was quite poor in error metric terms it made a large contribution to our score. We used NNs on all train and test active data and Rob trained a per category xgb price model after extracting lots of text features (numerical fields for things like cars/houses etc...) which helped a lot.\n\n**NNs and LGBM/XGB**\n\nLike others we had good success with NNs and trained a variety on different text embeddings (fasttext cc and wiki, self-trained w2v with and without stemming). We didn't add many other features to the NNs but did use a variety of architectures with and without attention. Modelling raw log1p(price) as a categorical and then embedding it worked well too, as did lower batch sizes and averaging predictions over several epochs. It was my first time using DL in a competition so I was very happy to be reaching ~0.2195 with a NN without many new features.\n\n**CV/stacking**\n\nEverything was done on a 5 fold basis and using a lgbm stacker. We got a lot of success from this but perhaps could have been more careful in trying to construct diverse base models and using a more rigorous feed forward selection. Antoine in particular did a lot of work ensuring we were robust in what we finally included.\n\n**Other thoughts**\n\n - We did try to train models to predict period from the active data to no avail.\n - We didn't use images beyond basic features\n - We tried translating all the text, probably too late, but google's api gave up on us. The hope was the translation, whilst average, would help with the cases in the Russian language - as well as letting us finally understand what was being sold!\n - We trained models on a few different targets (e.g a NN to predict if an item was a 0 or not, treating target as categorical) which then went into the stack. The pdf entropy from these classifications was a good feature also.\n - We did train a user_id model (only using users that were in test and train) but ended up excluding any models which used user_id as were still finding too unstable CV-LB relationship. Be interested to hear if anyone included it successfully.\n - I second some of the other comments about workflow: we saved all our transformations down and then simple read in .pkl files and concatenated. Some of the transformations were expensive and this allowed us to get going quickly.\n - [Gensim][2] is a cool package, though slightly steeper learning curve than sklearn.\n - A note on lgbm: we actually found better results by sometimes excluding strong categorical features, particularly from the stack. My intuition for this is that given lgb fits in a leaf-wise (best first) manner it perhaps over-uses categorical features which always offer easy splits but ultimately might not be the best choice/don't allow slower exploration of weaker features or more subtle relationships. Keen to hear what others think of this.\n\nI'll let Rob and Antoine add anything if they see fit. Thanks again for a fun competition!\n\nMark\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/57410\n  [2]: https://radimrehurek.com/gensim/",
      "votes": null
    },
    {
      "id": "350018",
      "postDate": "06/29/2018 03:07:03",
      "content": "<p>Congratulations @Maw501 and team on a very strong finish. Thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congratulations @Maw501 and team on a very strong finish. Thanks for sharing your solution overview.",
      "votes": null
    },
    {
      "id": "350045",
      "postDate": "06/29/2018 04:06:28",
      "content": "<p>Congratulations @Maw501,thanks for sharing your price idea.However it's pity that many people like me,have not see it</p>",
      "rawMarkdown": "Congratulations @Maw501,thanks for sharing your price idea.However it's pity that many people like me,have not see it",
      "votes": null
    },
    {
      "id": "350046",
      "postDate": "06/29/2018 04:08:36",
      "content": "<p>But I use a model to predict image_top_1 to fillna, but improve a little little score.</p>",
      "rawMarkdown": "But I use a model to predict image_top_1 to fillna, but improve a little little score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 350018,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/29/2018 03:07:03",
      "content": "<p>Congratulations @Maw501 and team on a very strong finish. Thanks for sharing your solution overview.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 350045,
      "author_name": "liuhdsgoal",
      "author_url": "",
      "post_date": "06/29/2018 04:06:28",
      "content": "<p>Congratulations @Maw501,thanks for sharing your price idea.However it's pity that many people like me,have not see it</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 350046,
      "author_name": "liuhdsgoal",
      "author_url": "",
      "post_date": "06/29/2018 04:08:36",
      "content": "<p>But I use a model to predict image_top_1 to fillna, but improve a little little score.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349666": "Firstly thank you to Avito for hosting such an awesome competition and thank you to both my teammates for all their hard work. I'll try to keep this brief.\n\n**Price model from active data**\n\nI actually wrote about this [here][1] and am surprised that many others didn't do this - it was perhaps our single biggest success. Whilst a price model on active data was quite poor in error metric terms it made a large contribution to our score. We used NNs on all train and test active data and Rob trained a per category xgb price model after extracting lots of text features (numerical fields for things like cars/houses etc...) which helped a lot.\n\n**NNs and LGBM/XGB**\n\nLike others we had good success with NNs and trained a variety on different text embeddings (fasttext cc and wiki, self-trained w2v with and without stemming). We didn't add many other features to the NNs but did use a variety of architectures with and without attention. Modelling raw log1p(price) as a categorical and then embedding it worked well too, as did lower batch sizes and averaging predictions over several epochs. It was my first time using DL in a competition so I was very happy to be reaching ~0.2195 with a NN without many new features.\n\n**CV/stacking**\n\nEverything was done on a 5 fold basis and using a lgbm stacker. We got a lot of success from this but perhaps could have been more careful in trying to construct diverse base models and using a more rigorous feed forward selection. Antoine in particular did a lot of work ensuring we were robust in what we finally included.\n\n**Other thoughts**\n\n - We did try to train models to predict period from the active data to no avail.\n - We didn't use images beyond basic features\n - We tried translating all the text, probably too late, but google's api gave up on us. The hope was the translation, whilst average, would help with the cases in the Russian language - as well as letting us finally understand what was being sold!\n - We trained models on a few different targets (e.g a NN to predict if an item was a 0 or not, treating target as categorical) which then went into the stack. The pdf entropy from these classifications was a good feature also.\n - We did train a user_id model (only using users that were in test and train) but ended up excluding any models which used user_id as were still finding too unstable CV-LB relationship. Be interested to hear if anyone included it successfully.\n - I second some of the other comments about workflow: we saved all our transformations down and then simple read in .pkl files and concatenated. Some of the transformations were expensive and this allowed us to get going quickly.\n - [Gensim][2] is a cool package, though slightly steeper learning curve than sklearn.\n - A note on lgbm: we actually found better results by sometimes excluding strong categorical features, particularly from the stack. My intuition for this is that given lgb fits in a leaf-wise (best first) manner it perhaps over-uses categorical features which always offer easy splits but ultimately might not be the best choice/don't allow slower exploration of weaker features or more subtle relationships. Keen to hear what others think of this.\n\nI'll let Rob and Antoine add anything if they see fit. Thanks again for a fun competition!\n\nMark\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/57410\n  [2]: https://radimrehurek.com/gensim/",
    "350018": "Congratulations @Maw501 and team on a very strong finish. Thanks for sharing your solution overview.",
    "350045": "Congratulations @Maw501,thanks for sharing your price idea.However it's pity that many people like me,have not see it",
    "350046": "But I use a model to predict image_top_1 to fillna, but improve a little little score."
  },
  "source": "meta"
}