{
  "id": 59881,
  "title": "4th Place Solution",
  "url": "/competitions/avito-demand-prediction/discussion/59881",
  "author_name": "Joe Eddy",
  "post_date": "2018-06-28T01:51:00.161000",
  "votes": 115,
  "comment_count": 31,
  "views": 0,
  "content": "<p><strong>Edit/update</strong>: please also see <a href=\"https://www.kaggle.com/jtrotman\"></a><a href=\"/jtrotman\">@jtrotman</a>'s writeup in the responses below! :) </p>\n\n<p>Thank you @ Kaggle and Avito - this competition was just awesome. There were so many interesting facets of the data to explore, and this was a rare competition where I wish I had joined earlier and spent even more time with it instead of getting sick of the data by the end. And a huge, huge thank you to my team <a href=\"https://www.kaggle.com/neongen\"></a><a href=\"/neongen\">@neongen</a>, <a href=\"https://www.kaggle.com/gphilippis\"></a><a href=\"/gphilippis\">@gphilippis</a>, and <a href=\"https://www.kaggle.com/jtrotman\"></a><a href=\"/jtrotman\">@jtrotman</a>. This was a great and inspiring team that I feel very lucky to have been a part of.</p>\n\n<p>Our solution is a lgbm stacker trained on a bunch of good base models and a chunk of our strongest features. Nonlinear stacking and including features for stacking definitely helped. My teammates can share more about that framework and all of the different models that went into the stack (lgbs, nns with lstm, sparse nns, weaker models like ridge), but my part of this thread will focus most on my biggest contribution to the result - a single lgbm model that scores .2175 public /.2213 private and the features that go into this model.</p>\n\n<p><strong>Hyperparameters</strong></p>\n\n<p>Hyperparameter tuning was not my main focus, but I found that very deep trees with a low learning rate worked quite well. Here are the final parameters:</p>\n\n<pre><code>lgb_params = {\n            'boosting_type': 'gbdt',\n            'objective': 'regression',\n            'metric': 'rmse',\n            'learning_rate': 0.01,\n            'num_leaves': 400, \n            'colsample_bytree': .45\n          }\n</code></pre>\n\n<p><strong>Features</strong></p>\n\n<p>Now to get to the good part and my focus in this competition - features. I ended up with about 800 tabular features (original + engineered), along with 100k tf-idf features for description and title + param_1. Here are the tf-idf settings:</p>\n\n<pre><code>TfidfVectorizer(stop_words=stopwords.words('russian'), \n                         lowercase=True, ngram_range=(1, 3),\n                         max_features=50000,\n                         sublinear_tf=True)\n</code></pre>\n\n<p>I'd summarize my tabular features by saying that they try to extract as much information as possible using price and the text fields. At first I only used very simple image features like color channels and size, and the image features I added at the end only gave about a .0003 boost.</p>\n\n<p>My core feature ideas:</p>\n\n<p><strong>Price statistics</strong>: </p>\n\n<p>It seemed clear early on that price information should have a very import impact on the target. So I wanted to extract price distribution information / statistics on pretty much every basis of aggregation I could come up with (e.g. by category, by category and city, by user_id, by image_top_1, etc.). For a bunch of different aggregates, I took these stats across all records  (with duplicate item_ids dropped):</p>\n\n<ul>\n<li>20th percentile</li>\n<li>median</li>\n<li>max</li>\n<li>standard deviation</li>\n<li>skew</li>\n</ul>\n\n<p>Then computed some relative price numbers for each post record with respect to the corresponding aggregate -- this measures how expensive or cheap the item is relative to the grouping\n- row / 20th percentile, row / median, row / max </p>\n\n<p>These computations are at the core of my FE, including for the ideas just below. </p>\n\n<p><strong>Granular textual aggregates on price</strong>: </p>\n\n<p>I believe my most important discovery in this competition was that applying my price statistics features to very specific text groupings worked extremely well, and these features are at the core of my model. Taking price / median price for each post's title alone gave me a big boost. I hypothesize that this may be a way to partially reverse engineer keyword search rankings (e.g. search sorted by price) and that search rankings heavily impact views and therefore deal_probability. Image that you're an avito user looking browsing for items - you search for a few specific keywords (e.g. \"red bicycle\"), sort by price, and start looking at many of the cheapest options. I've extended the idea to add more value in a few different ways that try to reduce the number of distinct groupings while still being granular enough to capture this level of signal -- I run price statistics on all of the following:</p>\n\n<ul>\n<li>title_noun aggregates: I extract all nouns from each title, normalized them, removed duplicates and sorted them alphabetically, and then used the resulting key as a basis for aggregation</li>\n<li>title_noun_adjs aggregates: same as above, but adding adjectives as well </li>\n<li>title_cluster: using title tf-idf features, I run SVD with 500 components and form 30,000 k-means clusters on these components. k-means is very slow, so I used mini-batch k-means and limited it to 500 components / 30k clusters even though I think more of both may have performed better.</li>\n<li>text_cluster: same as above, but based on a concatenation of title, description, and the param fields. </li>\n</ul>\n\n<p><strong>User semantics</strong>:</p>\n\n<p>When you build ML models, it's easy to get stuck in a row-based world and lose sight of the many interesting relationships between the different rows. Normal aggregate feature engineering like column statistics help with this, but I got to a point beyond that where I realized there was still more to be done with the text fields to help capture user characteristics.</p>\n\n<p>Coming off of talkingdata where it saw some success, I had the idea to use matrix factorization for this. So for users in train/test, I concatenated their text fields across all rows (title, desc, params), ran tf-idf -&gt; 300 component SVD, adding these components as features. A performance tip here is to use hashingvectorizer like below rather than straight tf-idf, since it is significantly more memory efficient (at the cost of a bit of accuracy). Set n_features higher than the number of features you really want, since there will be collisions.</p>\n\n<pre><code>print('Applying hash vectorizer then tf-idf to text')\n\nprint('Hash Vectorizer')\nhv = HashingVectorizer(stop_words=stopwords.words('russian'), \n                       lowercase=True, ngram_range=(1, 3),\n                       n_features=200000)\nhv_feats = hv.fit_transform(user_text_df['all_text'])\n\nprint('TF-idf transformer')\ntfidf_user_text = TfidfTransformer() \ntfidf_user_text_feats = tfidf_user_text.fit_transform(hv_feats)\n</code></pre>\n\n<p>Another nice way to capture user-based information from text is with aggregations on meta-textual features like percentage of caps in title, etc. I took a bunch of meta-textual features and computed the same aggregate statistics that I did for price, across each user_id. For example, one of my top 30 features was the rather unfortunately named (by my conventions) </p>\n\n<pre><code>\"pct_caps_description_pct_user_id_median:pct_caps_description\"\n</code></pre>\n\n<p>This means taking the specific row's percentage of caps in description and dividing it by that user's median percentage of caps in the description. Maybe I should have called it \"user shoutiness relative to their norm\" instead.  </p>\n\n<p><strong>Image Features</strong></p>\n\n<p>For a big part of the competition I thought the images were mostly a red herring. My worldview was very focused on exploiting the possibilities of textual information and thought its relationship with search rankings had a much more dominant effect than the images would. It was very interesting to hear that people were seeing significant gains from good image features, and I think this is the area where I would try to improve my model more if given more time. I used a few very simple features (color channels, size dimensions), and some nice features prepared by my teammates (they could explain in more detail) that gave about a .0003 boost to my model as the final feature addition --</p>\n\n<ul>\n<li>Image blurness</li>\n<li>Color histograms - SVD components and some statistical features</li>\n<li>NIMA - activations, score, stds, and score + std</li>\n</ul>\n\n<p>Aside from features mentioned, I used a bunch of other miscellaneous features similar to those seen in kernels - simple meta-textual features, lat/long + lat/long clusters, avg days a user's post is active and similar stats from the periods data.  I did clean and lemmatize title and description before applying tf-idf and some of the other text processing mentioned above, but I believe it made a pretty minor difference. </p>\n\n<p>In addition to lgb I used these features in a few other models, e.g. a neural net that scores .2197 public. Architecture is very similar to what you can see in kernels - biGRU applied to title and description, category embeddings, numerics concatenated and passed through 2 dense layers. It was difficult to close the gap between lgb and NN with this feature set, I think largely because there were many null values as a result of aggregate stats on unique occurrences, and these aggregate stats were among the key features to the lgb. I tried some different imputation strategies, but nothing worked all that well (although I will say that imputing price in a smart way, whether by a groupby-average on some category combinations or by a dedicated price prediction model, improved the NN score by about .001). </p>\n\n<p><strong>Some Comments on Workflow</strong></p>\n\n<p>I want to end on a few comments about workflow that I hope may be useful. What works well in one competition might not be ideal for another, but maybe it's helpful :) I've already written about some of these approaches here: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/56986#330023\">https://www.kaggle.com/c/avito-demand-prediction/discussion/56986#330023</a></p>\n\n<p>If you do a lot of feature engineering (especially with categorical aggregates), you are likely creating what amounts to a miniature relational database. You don't have to go as heavyweight as creating an actual database, but I found it helpful to create a bunch of feather files for storing aggregate features then merge all the ones I wanted back into the main data to build models. For example I had many feature files like \"global_city_features.ftr\" that I would merge back into the training dataframe using city as a key. Storing features in a relational manner saves you processing time (run FE once, not every time you model), saves you disk space, and gives you a lot of flexibility for choosing which features to include in a model run (just don't merge them in if you don't want that set).</p>\n\n<p>I used greedy forward feature selection - engineer a new set of related features, add them to my current best model, see if validation result improves. If it doesn't improve, leave them out of the model. This is a fast and fairly reliable (if not optimal) way of selecting for good features to include in a model.</p>\n\n<p>I noticed early on that there was very little variance in RMSE between validation folds (&lt;.0002), so I felt pretty safe testing for feature contributions to my model by evaluating on only one fold at a time. This helped speed up iteration on feature sets.</p>\n\n<p>I've already written a lot and my teammates will have additions to make as well so I'll stop there. If you're not sick of my blabbering and are interested, I'll likely write a blog post about my ideas / approach and share a link here once it's up. I hope this was a useful read, and happy kaggling!</p>",
  "messages": [
    {
      "id": 349294,
      "postDate": "2018-06-28T01:51:00.163Z",
      "content": "<p><strong>Edit/update</strong>: please also see <a href=\"https://www.kaggle.com/jtrotman\"></a><a href=\"/jtrotman\">@jtrotman</a>'s writeup in the responses below! :) </p>\n\n<p>Thank you @ Kaggle and Avito - this competition was just awesome. There were so many interesting facets of the data to explore, and this was a rare competition where I wish I had joined earlier and spent even more time with it instead of getting sick of the data by the end. And a huge, huge thank you to my team <a href=\"https://www.kaggle.com/neongen\"></a><a href=\"/neongen\">@neongen</a>, <a href=\"https://www.kaggle.com/gphilippis\"></a><a href=\"/gphilippis\">@gphilippis</a>, and <a href=\"https://www.kaggle.com/jtrotman\"></a><a href=\"/jtrotman\">@jtrotman</a>. This was a great and inspiring team that I feel very lucky to have been a part of.</p>\n\n<p>Our solution is a lgbm stacker trained on a bunch of good base models and a chunk of our strongest features. Nonlinear stacking and including features for stacking definitely helped. My teammates can share more about that framework and all of the different models that went into the stack (lgbs, nns with lstm, sparse nns, weaker models like ridge), but my part of this thread will focus most on my biggest contribution to the result - a single lgbm model that scores .2175 public /.2213 private and the features that go into this model.</p>\n\n<p><strong>Hyperparameters</strong></p>\n\n<p>Hyperparameter tuning was not my main focus, but I found that very deep trees with a low learning rate worked quite well. Here are the final parameters:</p>\n\n<pre><code>lgb_params = {\n            'boosting_type': 'gbdt',\n            'objective': 'regression',\n            'metric': 'rmse',\n            'learning_rate': 0.01,\n            'num_leaves': 400, \n            'colsample_bytree': .45\n          }\n</code></pre>\n\n<p><strong>Features</strong></p>\n\n<p>Now to get to the good part and my focus in this competition - features. I ended up with about 800 tabular features (original + engineered), along with 100k tf-idf features for description and title + param_1. Here are the tf-idf settings:</p>\n\n<pre><code>TfidfVectorizer(stop_words=stopwords.words('russian'), \n                         lowercase=True, ngram_range=(1, 3),\n                         max_features=50000,\n                         sublinear_tf=True)\n</code></pre>\n\n<p>I'd summarize my tabular features by saying that they try to extract as much information as possible using price and the text fields. At first I only used very simple image features like color channels and size, and the image features I added at the end only gave about a .0003 boost.</p>\n\n<p>My core feature ideas:</p>\n\n<p><strong>Price statistics</strong>: </p>\n\n<p>It seemed clear early on that price information should have a very import impact on the target. So I wanted to extract price distribution information / statistics on pretty much every basis of aggregation I could come up with (e.g. by category, by category and city, by user_id, by image_top_1, etc.). For a bunch of different aggregates, I took these stats across all records  (with duplicate item_ids dropped):</p>\n\n<ul>\n<li>20th percentile</li>\n<li>median</li>\n<li>max</li>\n<li>standard deviation</li>\n<li>skew</li>\n</ul>\n\n<p>Then computed some relative price numbers for each post record with respect to the corresponding aggregate -- this measures how expensive or cheap the item is relative to the grouping\n- row / 20th percentile, row / median, row / max </p>\n\n<p>These computations are at the core of my FE, including for the ideas just below. </p>\n\n<p><strong>Granular textual aggregates on price</strong>: </p>\n\n<p>I believe my most important discovery in this competition was that applying my price statistics features to very specific text groupings worked extremely well, and these features are at the core of my model. Taking price / median price for each post's title alone gave me a big boost. I hypothesize that this may be a way to partially reverse engineer keyword search rankings (e.g. search sorted by price) and that search rankings heavily impact views and therefore deal_probability. Image that you're an avito user looking browsing for items - you search for a few specific keywords (e.g. \"red bicycle\"), sort by price, and start looking at many of the cheapest options. I've extended the idea to add more value in a few different ways that try to reduce the number of distinct groupings while still being granular enough to capture this level of signal -- I run price statistics on all of the following:</p>\n\n<ul>\n<li>title_noun aggregates: I extract all nouns from each title, normalized them, removed duplicates and sorted them alphabetically, and then used the resulting key as a basis for aggregation</li>\n<li>title_noun_adjs aggregates: same as above, but adding adjectives as well </li>\n<li>title_cluster: using title tf-idf features, I run SVD with 500 components and form 30,000 k-means clusters on these components. k-means is very slow, so I used mini-batch k-means and limited it to 500 components / 30k clusters even though I think more of both may have performed better.</li>\n<li>text_cluster: same as above, but based on a concatenation of title, description, and the param fields. </li>\n</ul>\n\n<p><strong>User semantics</strong>:</p>\n\n<p>When you build ML models, it's easy to get stuck in a row-based world and lose sight of the many interesting relationships between the different rows. Normal aggregate feature engineering like column statistics help with this, but I got to a point beyond that where I realized there was still more to be done with the text fields to help capture user characteristics.</p>\n\n<p>Coming off of talkingdata where it saw some success, I had the idea to use matrix factorization for this. So for users in train/test, I concatenated their text fields across all rows (title, desc, params), ran tf-idf -&gt; 300 component SVD, adding these components as features. A performance tip here is to use hashingvectorizer like below rather than straight tf-idf, since it is significantly more memory efficient (at the cost of a bit of accuracy). Set n_features higher than the number of features you really want, since there will be collisions.</p>\n\n<pre><code>print('Applying hash vectorizer then tf-idf to text')\n\nprint('Hash Vectorizer')\nhv = HashingVectorizer(stop_words=stopwords.words('russian'), \n                       lowercase=True, ngram_range=(1, 3),\n                       n_features=200000)\nhv_feats = hv.fit_transform(user_text_df['all_text'])\n\nprint('TF-idf transformer')\ntfidf_user_text = TfidfTransformer() \ntfidf_user_text_feats = tfidf_user_text.fit_transform(hv_feats)\n</code></pre>\n\n<p>Another nice way to capture user-based information from text is with aggregations on meta-textual features like percentage of caps in title, etc. I took a bunch of meta-textual features and computed the same aggregate statistics that I did for price, across each user_id. For example, one of my top 30 features was the rather unfortunately named (by my conventions) </p>\n\n<pre><code>\"pct_caps_description_pct_user_id_median:pct_caps_description\"\n</code></pre>\n\n<p>This means taking the specific row's percentage of caps in description and dividing it by that user's median percentage of caps in the description. Maybe I should have called it \"user shoutiness relative to their norm\" instead.  </p>\n\n<p><strong>Image Features</strong></p>\n\n<p>For a big part of the competition I thought the images were mostly a red herring. My worldview was very focused on exploiting the possibilities of textual information and thought its relationship with search rankings had a much more dominant effect than the images would. It was very interesting to hear that people were seeing significant gains from good image features, and I think this is the area where I would try to improve my model more if given more time. I used a few very simple features (color channels, size dimensions), and some nice features prepared by my teammates (they could explain in more detail) that gave about a .0003 boost to my model as the final feature addition --</p>\n\n<ul>\n<li>Image blurness</li>\n<li>Color histograms - SVD components and some statistical features</li>\n<li>NIMA - activations, score, stds, and score + std</li>\n</ul>\n\n<p>Aside from features mentioned, I used a bunch of other miscellaneous features similar to those seen in kernels - simple meta-textual features, lat/long + lat/long clusters, avg days a user's post is active and similar stats from the periods data.  I did clean and lemmatize title and description before applying tf-idf and some of the other text processing mentioned above, but I believe it made a pretty minor difference. </p>\n\n<p>In addition to lgb I used these features in a few other models, e.g. a neural net that scores .2197 public. Architecture is very similar to what you can see in kernels - biGRU applied to title and description, category embeddings, numerics concatenated and passed through 2 dense layers. It was difficult to close the gap between lgb and NN with this feature set, I think largely because there were many null values as a result of aggregate stats on unique occurrences, and these aggregate stats were among the key features to the lgb. I tried some different imputation strategies, but nothing worked all that well (although I will say that imputing price in a smart way, whether by a groupby-average on some category combinations or by a dedicated price prediction model, improved the NN score by about .001). </p>\n\n<p><strong>Some Comments on Workflow</strong></p>\n\n<p>I want to end on a few comments about workflow that I hope may be useful. What works well in one competition might not be ideal for another, but maybe it's helpful :) I've already written about some of these approaches here: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/56986#330023\">https://www.kaggle.com/c/avito-demand-prediction/discussion/56986#330023</a></p>\n\n<p>If you do a lot of feature engineering (especially with categorical aggregates), you are likely creating what amounts to a miniature relational database. You don't have to go as heavyweight as creating an actual database, but I found it helpful to create a bunch of feather files for storing aggregate features then merge all the ones I wanted back into the main data to build models. For example I had many feature files like \"global_city_features.ftr\" that I would merge back into the training dataframe using city as a key. Storing features in a relational manner saves you processing time (run FE once, not every time you model), saves you disk space, and gives you a lot of flexibility for choosing which features to include in a model run (just don't merge them in if you don't want that set).</p>\n\n<p>I used greedy forward feature selection - engineer a new set of related features, add them to my current best model, see if validation result improves. If it doesn't improve, leave them out of the model. This is a fast and fairly reliable (if not optimal) way of selecting for good features to include in a model.</p>\n\n<p>I noticed early on that there was very little variance in RMSE between validation folds (&lt;.0002), so I felt pretty safe testing for feature contributions to my model by evaluating on only one fold at a time. This helped speed up iteration on feature sets.</p>\n\n<p>I've already written a lot and my teammates will have additions to make as well so I'll stop there. If you're not sick of my blabbering and are interested, I'll likely write a blog post about my ideas / approach and share a link here once it's up. I hope this was a useful read, and happy kaggling!</p>",
      "rawMarkdown": "**Edit/update**: please also see [@jtrotman][1]'s writeup in the responses below! :) \n\nThank you @ Kaggle and Avito - this competition was just awesome. There were so many interesting facets of the data to explore, and this was a rare competition where I wish I had joined earlier and spent even more time with it instead of getting sick of the data by the end. And a huge, huge thank you to my team [@neongen][2], [@gphilippis][3], and [@jtrotman][4]. This was a great and inspiring team that I feel very lucky to have been a part of.\n\nOur solution is a lgbm stacker trained on a bunch of good base models and a chunk of our strongest features. Nonlinear stacking and including features for stacking definitely helped. My teammates can share more about that framework and all of the different models that went into the stack (lgbs, nns with lstm, sparse nns, weaker models like ridge), but my part of this thread will focus most on my biggest contribution to the result - a single lgbm model that scores .2175 public /.2213 private and the features that go into this model.\n\n**Hyperparameters**\n\nHyperparameter tuning was not my main focus, but I found that very deep trees with a low learning rate worked quite well. Here are the final parameters:\n\n    lgb_params = {\n                'boosting_type': 'gbdt',\n                'objective': 'regression',\n                'metric': 'rmse',\n                'learning_rate': 0.01,\n                'num_leaves': 400, \n                'colsample_bytree': .45\n              }\n\n**Features**\n\nNow to get to the good part and my focus in this competition - features. I ended up with about 800 tabular features (original + engineered), along with 100k tf-idf features for description and title + param_1. Here are the tf-idf settings:\n\n    TfidfVectorizer(stop_words=stopwords.words('russian'), \n                             lowercase=True, ngram_range=(1, 3),\n                             max_features=50000,\n                             sublinear_tf=True)\n\nI'd summarize my tabular features by saying that they try to extract as much information as possible using price and the text fields. At first I only used very simple image features like color channels and size, and the image features I added at the end only gave about a .0003 boost.\n\nMy core feature ideas:\n\n**Price statistics**: \n\nIt seemed clear early on that price information should have a very import impact on the target. So I wanted to extract price distribution information / statistics on pretty much every basis of aggregation I could come up with (e.g. by category, by category and city, by user_id, by image_top_1, etc.). For a bunch of different aggregates, I took these stats across all records  (with duplicate item_ids dropped):\n\n - 20th percentile\n - median\n - max\n - standard deviation\n - skew\n\nThen computed some relative price numbers for each post record with respect to the corresponding aggregate -- this measures how expensive or cheap the item is relative to the grouping\n- row / 20th percentile, row / median, row / max \n\nThese computations are at the core of my FE, including for the ideas just below. \n\n**Granular textual aggregates on price**: \n\nI believe my most important discovery in this competition was that applying my price statistics features to very specific text groupings worked extremely well, and these features are at the core of my model. Taking price / median price for each post's title alone gave me a big boost. I hypothesize that this may be a way to partially reverse engineer keyword search rankings (e.g. search sorted by price) and that search rankings heavily impact views and therefore deal_probability. Image that you're an avito user looking browsing for items - you search for a few specific keywords (e.g. \"red bicycle\"), sort by price, and start looking at many of the cheapest options. I've extended the idea to add more value in a few different ways that try to reduce the number of distinct groupings while still being granular enough to capture this level of signal -- I run price statistics on all of the following:\n\n - title_noun aggregates: I extract all nouns from each title, normalized them, removed duplicates and sorted them alphabetically, and then used the resulting key as a basis for aggregation\n - title_noun_adjs aggregates: same as above, but adding adjectives as well \n - title_cluster: using title tf-idf features, I run SVD with 500 components and form 30,000 k-means clusters on these components. k-means is very slow, so I used mini-batch k-means and limited it to 500 components / 30k clusters even though I think more of both may have performed better.\n - text_cluster: same as above, but based on a concatenation of title, description, and the param fields. \n\n**User semantics**:\n\nWhen you build ML models, it's easy to get stuck in a row-based world and lose sight of the many interesting relationships between the different rows. Normal aggregate feature engineering like column statistics help with this, but I got to a point beyond that where I realized there was still more to be done with the text fields to help capture user characteristics.\n    \nComing off of talkingdata where it saw some success, I had the idea to use matrix factorization for this. So for users in train/test, I concatenated their text fields across all rows (title, desc, params), ran tf-idf -&gt; 300 component SVD, adding these components as features. A performance tip here is to use hashingvectorizer like below rather than straight tf-idf, since it is significantly more memory efficient (at the cost of a bit of accuracy). Set n_features higher than the number of features you really want, since there will be collisions.\n\n    print('Applying hash vectorizer then tf-idf to text')\n    \n    print('Hash Vectorizer')\n    hv = HashingVectorizer(stop_words=stopwords.words('russian'), \n                           lowercase=True, ngram_range=(1, 3),\n                           n_features=200000)\n    hv_feats = hv.fit_transform(user_text_df['all_text'])\n    \n    print('TF-idf transformer')\n    tfidf_user_text = TfidfTransformer() \n    tfidf_user_text_feats = tfidf_user_text.fit_transform(hv_feats)\n\nAnother nice way to capture user-based information from text is with aggregations on meta-textual features like percentage of caps in title, etc. I took a bunch of meta-textual features and computed the same aggregate statistics that I did for price, across each user_id. For example, one of my top 30 features was the rather unfortunately named (by my conventions) \n\n    \"pct_caps_description_pct_user_id_median:pct_caps_description\"\n\nThis means taking the specific row's percentage of caps in description and dividing it by that user's median percentage of caps in the description. Maybe I should have called it \"user shoutiness relative to their norm\" instead.  \n\n**Image Features**\n\nFor a big part of the competition I thought the images were mostly a red herring. My worldview was very focused on exploiting the possibilities of textual information and thought its relationship with search rankings had a much more dominant effect than the images would. It was very interesting to hear that people were seeing significant gains from good image features, and I think this is the area where I would try to improve my model more if given more time. I used a few very simple features (color channels, size dimensions), and some nice features prepared by my teammates (they could explain in more detail) that gave about a .0003 boost to my model as the final feature addition --\n\n - Image blurness\n - Color histograms - SVD components and some statistical features\n - NIMA - activations, score, stds, and score + std\n\nAside from features mentioned, I used a bunch of other miscellaneous features similar to those seen in kernels - simple meta-textual features, lat/long + lat/long clusters, avg days a user's post is active and similar stats from the periods data.  I did clean and lemmatize title and description before applying tf-idf and some of the other text processing mentioned above, but I believe it made a pretty minor difference. \n \nIn addition to lgb I used these features in a few other models, e.g. a neural net that scores .2197 public. Architecture is very similar to what you can see in kernels - biGRU applied to title and description, category embeddings, numerics concatenated and passed through 2 dense layers. It was difficult to close the gap between lgb and NN with this feature set, I think largely because there were many null values as a result of aggregate stats on unique occurrences, and these aggregate stats were among the key features to the lgb. I tried some different imputation strategies, but nothing worked all that well (although I will say that imputing price in a smart way, whether by a groupby-average on some category combinations or by a dedicated price prediction model, improved the NN score by about .001). \n\n**Some Comments on Workflow**\n\nI want to end on a few comments about workflow that I hope may be useful. What works well in one competition might not be ideal for another, but maybe it's helpful :) I've already written about some of these approaches here: https://www.kaggle.com/c/avito-demand-prediction/discussion/56986#330023\n\nIf you do a lot of feature engineering (especially with categorical aggregates), you are likely creating what amounts to a miniature relational database. You don't have to go as heavyweight as creating an actual database, but I found it helpful to create a bunch of feather files for storing aggregate features then merge all the ones I wanted back into the main data to build models. For example I had many feature files like \"global_city_features.ftr\" that I would merge back into the training dataframe using city as a key. Storing features in a relational manner saves you processing time (run FE once, not every time you model), saves you disk space, and gives you a lot of flexibility for choosing which features to include in a model run (just don't merge them in if you don't want that set).\n\nI used greedy forward feature selection - engineer a new set of related features, add them to my current best model, see if validation result improves. If it doesn't improve, leave them out of the model. This is a fast and fairly reliable (if not optimal) way of selecting for good features to include in a model.\n\nI noticed early on that there was very little variance in RMSE between validation folds (&lt;.0002), so I felt pretty safe testing for feature contributions to my model by evaluating on only one fold at a time. This helped speed up iteration on feature sets.\n\nI've already written a lot and my teammates will have additions to make as well so I'll stop there. If you're not sick of my blabbering and are interested, I'll likely write a blog post about my ideas / approach and share a link here once it's up. I hope this was a useful read, and happy kaggling!\n  \n\n\n  [1]: https://www.kaggle.com/jtrotman\n  [2]: https://www.kaggle.com/neongen\n  [3]: https://www.kaggle.com/gphilippis\n  [4]: https://www.kaggle.com/jtrotman",
      "votes": 115
    },
    {
      "id": 349682,
      "postDate": "2018-06-28T13:29:16.427Z",
      "content": "<p>Thanks Joe – here’s a quick write up of what I did and our path to teaming up. I’ll avoid mentioning what <a href=\"/neongen\">@neongen</a> and <a href=\"/gphilippis\">@gphilippis</a> did so these two posts together still only describe half the ensemble!</p>\n\n<p>My general approach was to start stacking straight away, using a LightGBM stacker, testing new features in there along with a growing set of level 1 models. Just a plain 5 fold CV lead to the same 0.004 public LB delta others had mentioned, with very reliable correlation.</p>\n\n<p>I started off with sparse nets on word/char trigrams with custom tokenisation (using pymorphy2) and interactions, inspired by the <a href=\"https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\">top</a> <a href=\"https://www.kaggle.com/mchahhou/mercari-second-place-solution\">two</a> Mercari solutions, mainly because they are much faster to train and I thought they’d make a good start, these reached 0.2200 in CV. I also tried one BiGRU &amp; categorical embedding network based on the fasttext-russian-2m shared vectors, but did not work on it much and it stayed at 0.2245. Occasionally I’d build the LightGBM model as a standalone (with deeper trees), removing it’s L1 model prediction features, adding some tf-idf / SVD features, and adding that in as a new L1 model helped.</p>\n\n<p>I generated hundreds of features, here are a selection...</p>\n\n<h1>Photos</h1>\n\n<p>Inspired by this great share of <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">VGG16 features</a> I decided to use Kaggle Kernels to do all the image feature engineering, just a simple loop over the jpg files and dump the generated features into Python pickles for download, 6 hours runtime and 1Gb of space (and a GPU if you want!) is more than enough.</p>\n\n<h3>VGG16 features</h3>\n\n<p>I tried adding these (raw and log scaled) to sparse nets as extra inputs and got slight improvements. Also tried predicting the category using softmax (cars &amp; clothing were very predictable), and summarised those predictions, thinking it might pick up on mistakes, e.g. a confident picture of a car in a clothing listing or many other combinations.</p>\n\n<h3>Color Histograms</h3>\n\n<p>(just the RGB histograms with 32 bins per channel).</p>\n\n<ul>\n<li>SVD on normalised histograms down to 12 columns</li>\n<li>mean bin for each of R/G/B</li>\n<li>KL divergence between all of R/B, R/G etc.</li>\n</ul>\n\n<p>This surprised me – it weighed in at about 0.0004 with the feature set I had at the time.</p>\n\n<h3>ImageHash Hashes</h3>\n\n<p>ahash, dhash, phash, used for counts of images with same hash. Then aggregating over users, i.e. has a user listed something with a widely copied photo? Or are all their images unique?</p>\n\n<h3>Basic Info</h3>\n\n<p>Simply open the file and get the [width, height] dimensions, then the file size info from the ZipFile API – from this you can get to a pixel count and bytes_per_pixel which is a rough measure of image quality. (The CRC from the ZipFile API is also useful to use as another ‘hash’ like above.)</p>\n\n<p>After teaming, we went looking for more features and I adapted my Kernel code to generate these :</p>\n\n<h3>NIMA</h3>\n\n<p><a href=\"https://github.com/titu1994/neural-image-assessment\">The model</a> that <a href=\"https://www.kaggle.com/christofhenkel\">Dieter</a> kindly shared in the external data thread was small enough to run in a Kernel, scoring ~70 images per second single threaded with a GPU. The best of those features was the lower bound on the aesthetic score i.e. score_mean-score_stdv.</p>\n\n<h3>Image keypoints</h3>\n\n<p><a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59414#347781\">Shared here</a> by <a href=\"https://www.kaggle.com/nuhsikander\">Nooh</a> – I used a range of thresholds, 5 and 10 worked much better than higher ones.</p>\n\n<h1>Train / Test Active</h1>\n\n<p>The shared <code>avg_days_up_user</code> and <code>avg_times_up_user</code> features from <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">aggregated-features-lightgbm</a> proved very strong, I thought they were acting as a safe approximate form of user_id target coding. There were many deal_probability==0 training rows and this probably comes from held-back evidence Avito has that the item was at some point re-listed. If a user re-lists things a lot they’re probably not getting many deals...</p>\n\n<p>To expand on that I built a few models on train_active and test_active combined (only unique items), and scored them on train &amp; test to make new features. I chunked the files based on parent_category_name and created sparse feature sets for LightGBM, running similar loops:</p>\n\n<h3>Predict Re-list Probability:</h3>\n\n<ul>\n<li>Predict if item_id appears more than once, using categories &amp; tf-idf on titles with different tokenisations.</li>\n</ul>\n\n<h3>Predict Log Price:</h3>\n\n<ul>\n<li>Similar features: predict the log of the price, remove 0 / NA prices from training, but scoring for all train/test rows.</li>\n</ul>\n\n<p>These features helped a lot, the pred_relist_probability feature has 0.32 correlation to deal_probability, and user_id_max_pred_relist_prob was a top-3 feature in one base model.</p>\n\n<p>Adding pred_log_price gained about 0.0002 in CV and adding the interaction <code>log1p(price)-pred_log_price</code> as well as aggregations like user_id_mean_pred_log_price_diff added the same again.</p>\n\n<h1>Text Stats</h1>\n\n<p>I had old code from the <a href=\"https://www.kaggle.com/c/two-sigma-connect-rental-listing-inquiries/discussion/32146#178428\">Renthop competition</a> I could use here, I tried looking for things tf-idf / tokenisation type processing might miss, like turning descriptions into a kind of Morse code, e.g. replacing non exclamation marks with ‘a’ and summarizing the length of the string, and also using it as a new category, but this was a flop. It’s possible there were better things to find here, but it didn’t make the cut… Simpler punctuation counts and stats were enough.</p>\n\n<p>Simple title token statistics, word appearance counts transformed into probabilities and summed to give a naive title appearance probability helped to give some diversity, with a resulting spectrum of [ common .. rare ] words that overlaps in the middle. e.g. is there one rare word in the title or a few? The probabilities were from global counts in all of train_active, test_active, train and test combined, and also per category, e.g. p(see_word | category==Transport).</p>\n\n<p>Another view was: reduce the titles per category with tf-idf / SVD and compute centroids, then cosine distance to the category centroid: is it a ‘normal’ (common) listing for the category?</p>\n\n<p>One last trick: groupby user_id and zip their concatenated item descriptions, then take the length of the zipped encoding (trying both lowercase / original case). Then divide that by a user's listing count to give a compression ratio of their descriptions: are they repetitive or not? This might add a little something extra, but I liked Joe’s user semantics idea much more.</p>\n\n<h1>Matrix Factorizations</h1>\n\n<p>Some categorical combinations were small enough to do a full SVD and use the first N eigenvectors as features for tree models, for example [user_id, category_name_param_1 ], reducing listing counts (log1p scaled) with <code>np.linalg.svd</code> and using the first 12 vectors of each. The user_id vectors were the intended use: a ‘category’ space for each user describing what they list, but the category_name_param_1 space over user_id’s was used much more heavily by LGB, perhaps capturing category/param_1 combinations that are mainly ‘company’ listings versus ‘private’ listings in a better way.</p>\n\n<h1>Team Up</h1>\n\n<p>Before we teamed I guessed from my prospective team-mates past forum posts (e.g. <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\">by neongen</a> and <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44579\">by Joe</a> ) that we’d used quite different approaches and we merged as neighbouring top 10 entries on the leaderboard.</p>\n\n<p>A blend of our best solutions at 0.2161 and 0.2163 gave 0.2151, and combining all our L1 models (~40 by the end) into an LGB stacker with 240 features overall reached 0.2139. This is a great demonstration of diverse approaches combining well. We reached that peak after a few days and got stuck  - perhaps converging a bit – finding improvements proved very hard.</p>\n\n<p>This was my first time teaming up here and my teammates <a href=\"/aquatic\">@aquatic</a>, <a href=\"/neongen\">@neongen</a> and <a href=\"/gphilippis\">@gphilippis</a> are shining examples of exactly what you'd want from teammates, a relentless source of ideas to discuss and a constant flow of new L1 models! Thank you all guys, this was a great experience.</p>\n\n<p>Finally, thanks to Kaggle and Avito for hosting this competition – there were so many sources of variance to explore and surprises along the way, it’s both fun and great practice at honing ML intuition.</p>",
      "rawMarkdown": "Thanks Joe – here’s a quick write up of what I did and our path to teaming up. I’ll avoid mentioning what @neongen and @gphilippis did so these two posts together still only describe half the ensemble!\n\nMy general approach was to start stacking straight away, using a LightGBM stacker, testing new features in there along with a growing set of level 1 models. Just a plain 5 fold CV lead to the same 0.004 public LB delta others had mentioned, with very reliable correlation.\n\nI started off with sparse nets on word/char trigrams with custom tokenisation (using pymorphy2) and interactions, inspired by the [top][1] [two][2] Mercari solutions, mainly because they are much faster to train and I thought they’d make a good start, these reached 0.2200 in CV. I also tried one BiGRU &amp; categorical embedding network based on the fasttext-russian-2m shared vectors, but did not work on it much and it stayed at 0.2245. Occasionally I’d build the LightGBM model as a standalone (with deeper trees), removing it’s L1 model prediction features, adding some tf-idf / SVD features, and adding that in as a new L1 model helped.\n\nI generated hundreds of features, here are a selection...\n\n# Photos\n\nInspired by this great share of [VGG16 features][3] I decided to use Kaggle Kernels to do all the image feature engineering, just a simple loop over the jpg files and dump the generated features into Python pickles for download, 6 hours runtime and 1Gb of space (and a GPU if you want!) is more than enough.\n\n### VGG16 features\nI tried adding these (raw and log scaled) to sparse nets as extra inputs and got slight improvements. Also tried predicting the category using softmax (cars &amp; clothing were very predictable), and summarised those predictions, thinking it might pick up on mistakes, e.g. a confident picture of a car in a clothing listing or many other combinations.\n\n### Color Histograms\n (just the RGB histograms with 32 bins per channel).\n\n- SVD on normalised histograms down to 12 columns\n- mean bin for each of R/G/B\n- KL divergence between all of R/B, R/G etc.\n\nThis surprised me – it weighed in at about 0.0004 with the feature set I had at the time.\n\n### ImageHash Hashes\nahash, dhash, phash, used for counts of images with same hash. Then aggregating over users, i.e. has a user listed something with a widely copied photo? Or are all their images unique?\n\n### Basic Info\nSimply open the file and get the [width, height] dimensions, then the file size info from the ZipFile API – from this you can get to a pixel count and bytes_per_pixel which is a rough measure of image quality. (The CRC from the ZipFile API is also useful to use as another ‘hash’ like above.)\n\nAfter teaming, we went looking for more features and I adapted my Kernel code to generate these :\n\n### NIMA\n[The model][4] that [Dieter][5] kindly shared in the external data thread was small enough to run in a Kernel, scoring ~70 images per second single threaded with a GPU. The best of those features was the lower bound on the aesthetic score i.e. score_mean-score_stdv.\n\n### Image keypoints\n[Shared here][6] by [Nooh][7] – I used a range of thresholds, 5 and 10 worked much better than higher ones.\n\n\n# Train / Test Active\n\nThe shared `avg_days_up_user` and `avg_times_up_user` features from [aggregated-features-lightgbm][8] proved very strong, I thought they were acting as a safe approximate form of user_id target coding. There were many deal_probability==0 training rows and this probably comes from held-back evidence Avito has that the item was at some point re-listed. If a user re-lists things a lot they’re probably not getting many deals...\n\nTo expand on that I built a few models on train_active and test_active combined (only unique items), and scored them on train &amp; test to make new features. I chunked the files based on parent_category_name and created sparse feature sets for LightGBM, running similar loops:\n\n### Predict Re-list Probability:\n\n - Predict if item_id appears more than once, using categories &amp; tf-idf on titles with different tokenisations.\n\n### Predict Log Price:\n\n - Similar features: predict the log of the price, remove 0 / NA prices from training, but scoring for all train/test rows.\n\nThese features helped a lot, the pred_relist_probability feature has 0.32 correlation to deal_probability, and user_id_max_pred_relist_prob was a top-3 feature in one base model.\n\nAdding pred_log_price gained about 0.0002 in CV and adding the interaction `log1p(price)-pred_log_price` as well as aggregations like user_id_mean_pred_log_price_diff added the same again.\n\n# Text Stats\n\nI had old code from the [Renthop competition][9] I could use here, I tried looking for things tf-idf / tokenisation type processing might miss, like turning descriptions into a kind of Morse code, e.g. replacing non exclamation marks with ‘a’ and summarizing the length of the string, and also using it as a new category, but this was a flop. It’s possible there were better things to find here, but it didn’t make the cut… Simpler punctuation counts and stats were enough.\n\nSimple title token statistics, word appearance counts transformed into probabilities and summed to give a naive title appearance probability helped to give some diversity, with a resulting spectrum of [ common .. rare ] words that overlaps in the middle. e.g. is there one rare word in the title or a few? The probabilities were from global counts in all of train_active, test_active, train and test combined, and also per category, e.g. p(see_word | category==Transport).\n\nAnother view was: reduce the titles per category with tf-idf / SVD and compute centroids, then cosine distance to the category centroid: is it a ‘normal’ (common) listing for the category?\n\nOne last trick: groupby user_id and zip their concatenated item descriptions, then take the length of the zipped encoding (trying both lowercase / original case). Then divide that by a user's listing count to give a compression ratio of their descriptions: are they repetitive or not? This might add a little something extra, but I liked Joe’s user semantics idea much more.\n\n# Matrix Factorizations\n\nSome categorical combinations were small enough to do a full SVD and use the first N eigenvectors as features for tree models, for example [user_id, category_name_param_1 ], reducing listing counts (log1p scaled) with `np.linalg.svd` and using the first 12 vectors of each. The user_id vectors were the intended use: a ‘category’ space for each user describing what they list, but the category_name_param_1 space over user_id’s was used much more heavily by LGB, perhaps capturing category/param_1 combinations that are mainly ‘company’ listings versus ‘private’ listings in a better way.\n\n# Team Up\n\nBefore we teamed I guessed from my prospective team-mates past forum posts (e.g. [by neongen][10] and [by Joe][11] ) that we’d used quite different approaches and we merged as neighbouring top 10 entries on the leaderboard.\n\nA blend of our best solutions at 0.2161 and 0.2163 gave 0.2151, and combining all our L1 models (~40 by the end) into an LGB stacker with 240 features overall reached 0.2139. This is a great demonstration of diverse approaches combining well. We reached that peak after a few days and got stuck  - perhaps converging a bit – finding improvements proved very hard.\n\nThis was my first time teaming up here and my teammates @aquatic, @neongen and @gphilippis are shining examples of exactly what you'd want from teammates, a relentless source of ideas to discuss and a constant flow of new L1 models! Thank you all guys, this was a great experience.\n\nFinally, thanks to Kaggle and Avito for hosting this competition – there were so many sources of variance to explore and surprises along the way, it’s both fun and great practice at honing ML intuition.\n\n\n  [1]: https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\n  [2]: https://www.kaggle.com/mchahhou/mercari-second-place-solution\n  [3]: https://www.kaggle.com/bguberfain/vgg16-train-features\n  [4]: https://github.com/titu1994/neural-image-assessment\n  [5]: https://www.kaggle.com/christofhenkel\n  [6]: https://www.kaggle.com/c/avito-demand-prediction/discussion/59414#347781\n  [7]: https://www.kaggle.com/nuhsikander\n  [8]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\n  [9]: https://www.kaggle.com/c/two-sigma-connect-rental-listing-inquiries/discussion/32146#178428\n  [10]: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\n  [11]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44579\n",
      "votes": 13,
      "replies": [
        {
          "id": 351946,
          "postDate": "2018-07-03T10:29:30.397Z",
          "content": "<p>Congrats!!!</p>",
          "rawMarkdown": "Congrats!!!"
        }
      ]
    },
    {
      "id": 349305,
      "postDate": "2018-06-28T02:05:53.370Z",
      "content": "<p>Congrats on your 4th place finish. Thanks for posting such a detailed solution. I used a strategy quite similar to your's in terms of determining price competitiveness in a peer group. Obviously you went quite far with that approach, and I now see what else I could have done (i.e. lots of things) Yeah, k-means at these data sizes is indeed slow. I discovered <a href=\"https://github.com/src-d/kmcuda\">https://github.com/src-d/kmcuda</a> along the way and that was blazingly fast on a GPU. I split the data by category_name before forming clusters so that further reduced my computation time.</p>",
      "rawMarkdown": "Congrats on your 4th place finish. Thanks for posting such a detailed solution. I used a strategy quite similar to your's in terms of determining price competitiveness in a peer group. Obviously you went quite far with that approach, and I now see what else I could have done (i.e. lots of things) Yeah, k-means at these data sizes is indeed slow. I discovered https://github.com/src-d/kmcuda along the way and that was blazingly fast on a GPU. I split the data by category_name before forming clusters so that further reduced my computation time.\n",
      "votes": 5,
      "replies": [
        {
          "id": 349894,
          "postDate": "2018-06-28T20:19:35.793Z",
          "content": "<p>Thanks! Your approach is very nice, thanks for sharing it and the link - I'll have to check this out, it looks great.</p>",
          "rawMarkdown": "Thanks! Your approach is very nice, thanks for sharing it and the link - I'll have to check this out, it looks great."
        },
        {
          "id": 349975,
          "postDate": "2018-06-29T01:14:45.353Z",
          "content": "<p>Thanks <a href=\"/eigenvector\">@eigenvector</a> for that fast k-means link. I will definitely use it in the future.</p>",
          "rawMarkdown": "Thanks @eigenvector for that fast k-means link. I will definitely use it in the future.",
          "votes": 1
        }
      ]
    },
    {
      "id": 349323,
      "postDate": "2018-06-28T02:28:29.703Z",
      "content": "<p>Congrats Joe and thanks for the detailed solution... BTW, congratulations for becoming a Kaggle Master....</p>",
      "rawMarkdown": "Congrats Joe and thanks for the detailed solution... BTW, congratulations for becoming a Kaggle Master....",
      "votes": 3,
      "replies": [
        {
          "id": 349899,
          "postDate": "2018-06-28T20:21:33.023Z",
          "content": "<p>Thanks Samrat!</p>",
          "rawMarkdown": "Thanks Samrat!"
        }
      ]
    },
    {
      "id": 349826,
      "postDate": "2018-06-28T17:29:12.567Z",
      "content": "<p>Great solution. Congrats!!!</p>",
      "rawMarkdown": "Great solution. Congrats!!!",
      "votes": 4
    },
    {
      "id": 350150,
      "postDate": "2018-06-29T09:30:38.387Z",
      "content": "<blockquote>\n  <p>Price statistics \n   - row / 20th percentile, row / median, row / max</p>\n</blockquote>\n\n<p>Thanks for sharing. In this part, do you mean that you use row divide 20th percentile</p>",
      "rawMarkdown": "  \n\n&gt;  Price statistics \n - row / 20th percentile, row / median, row / max\n\nThanks for sharing. In this part, do you mean that you use row divide 20th percentile",
      "votes": 1,
      "replies": [
        {
          "id": 350456,
          "postDate": "2018-06-29T19:42:13.603Z",
          "content": "<p>The idea was to get features that capture price characteristics relative to a grouping. So for example, I want to know how expensive an item is vs. other items that share the same image_top_1 category. To get a feature like that, I first compute aggregate statistics for the category - its 20th percentile, median etc. Then for each individual item I merge in the aggregate statistic as a new column, and divide the item's price by the value of that column (that's what I mean by row / 20th percentile). Then it would look like df['price_pct_image_top_1_20thp:price'] = df['price'] / df['image_top_1_20thp:price'].</p>\n\n<p>Hope that's clear :)  </p>",
          "rawMarkdown": "The idea was to get features that capture price characteristics relative to a grouping. So for example, I want to know how expensive an item is vs. other items that share the same image_top_1 category. To get a feature like that, I first compute aggregate statistics for the category - its 20th percentile, median etc. Then for each individual item I merge in the aggregate statistic as a new column, and divide the item's price by the value of that column (that's what I mean by row / 20th percentile). Then it would look like df['price_pct_image_top_1_20thp:price'] = df['price'] / df['image_top_1_20thp:price'].\n\nHope that's clear :)  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 349506,
      "postDate": "2018-06-28T07:48:28.947Z",
      "content": "<p>Congrats! Really impressive!</p>",
      "rawMarkdown": "Congrats! Really impressive!",
      "votes": 1
    },
    {
      "id": 349449,
      "postDate": "2018-06-28T06:08:32.123Z",
      "content": "<p>Congrats and thx for the nice writeup :)</p>",
      "rawMarkdown": "Congrats and thx for the nice writeup :)",
      "votes": 1
    },
    {
      "id": 349448,
      "postDate": "2018-06-28T05:59:10.973Z",
      "content": "<p>Great solution. Congrats!</p>",
      "rawMarkdown": "Great solution. Congrats!",
      "votes": 1
    },
    {
      "id": 349379,
      "postDate": "2018-06-28T03:40:09.687Z",
      "content": "<p>Granular textual aggregates idea is GREAT! thanks for share your thoughts about that ! </p>",
      "rawMarkdown": "Granular textual aggregates idea is GREAT! thanks for share your thoughts about that ! ",
      "votes": 1
    },
    {
      "id": 349372,
      "postDate": "2018-06-28T03:27:31.347Z",
      "content": "<p>Fantastic.. Any chance on the code going public? I am particularly interested in your detailed feature engineering :D</p>",
      "rawMarkdown": "Fantastic.. Any chance on the code going public? I am particularly interested in your detailed feature engineering :D",
      "votes": 1
    },
    {
      "id": 350306,
      "postDate": "2018-06-29T14:25:08.357Z",
      "content": "<p>Thanks! Cool solution and great write-up!</p>",
      "rawMarkdown": "Thanks! Cool solution and great write-up!",
      "votes": 2
    },
    {
      "id": 349775,
      "postDate": "2018-06-28T16:31:38.583Z",
      "content": "<p>Congratulations Joe! This is a fantastic achievement. I see that you have LGB num_leaves as 400. A couple of other successful solutions also went deep into depth/num_leaves while my conventional wisdom says that it will result in overfitting. In fact, most of the public kernels were around 5 depths and 2**5-1 (num_leaves=31). Any thoughts on this? Any decision criteria you suggest for this hyper parameter?</p>",
      "rawMarkdown": "Congratulations Joe! This is a fantastic achievement. I see that you have LGB num_leaves as 400. A couple of other successful solutions also went deep into depth/num_leaves while my conventional wisdom says that it will result in overfitting. In fact, most of the public kernels were around 5 depths and 2**5-1 (num_leaves=31). Any thoughts on this? Any decision criteria you suggest for this hyper parameter?",
      "votes": 2,
      "replies": [
        {
          "id": 349914,
          "postDate": "2018-06-28T20:57:48.963Z",
          "content": "<p>Thanks Prasun! I was surprised by this too - I've usually seen shallow trees (3-8 max depth) work best for boosting. <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">Thousandvoices</a> went all the way to 1024 leaves! Tree depth is definitely very problem dependent, and one way to think about it intuitively is that it controls the level of feature interactivity that's allowed. With a massive number of low signal sparse features (from something like tf-idf, in my case 100k sparse features), it's sensible to me that trees need to go deep to properly capture the multitude of very specific interactions like word usages. Also, I think there's a lot of granularity in this prediction problem in general - specific target values are associated with multi-level category combinations as well as the very particular items that fall within those combinations and the very specific traits of those items relative to their peers (e.g. why granular price aggregate features work well). Another thing worth mentioning is that LGB uses a leaf-wise tree growing algorithm, which allows it to capitalize on going extremely deep if there's better information gain to be had that way.</p>\n\n<p>I think the reasons I've outlined can be used to an extent as an intuitive decision criterion, but at the end of the day your best answer will always be from proper validation. When you start building a gradient boosting model, try pushing the max depth up until you stop seeing validation improvement, and narrow it down to a reasonable choice in something like a binary-style search. E.g. try depth 3, 8, 15, discover 8 works better than 3 and 15, check around 8 until you find the right balance.    </p>",
          "rawMarkdown": "Thanks Prasun! I was surprised by this too - I've usually seen shallow trees (3-8 max depth) work best for boosting. [Thousandvoices][1] went all the way to 1024 leaves! Tree depth is definitely very problem dependent, and one way to think about it intuitively is that it controls the level of feature interactivity that's allowed. With a massive number of low signal sparse features (from something like tf-idf, in my case 100k sparse features), it's sensible to me that trees need to go deep to properly capture the multitude of very specific interactions like word usages. Also, I think there's a lot of granularity in this prediction problem in general - specific target values are associated with multi-level category combinations as well as the very particular items that fall within those combinations and the very specific traits of those items relative to their peers (e.g. why granular price aggregate features work well). Another thing worth mentioning is that LGB uses a leaf-wise tree growing algorithm, which allows it to capitalize on going extremely deep if there's better information gain to be had that way.\n\nI think the reasons I've outlined can be used to an extent as an intuitive decision criterion, but at the end of the day your best answer will always be from proper validation. When you start building a gradient boosting model, try pushing the max depth up until you stop seeing validation improvement, and narrow it down to a reasonable choice in something like a binary-style search. E.g. try depth 3, 8, 15, discover 8 works better than 3 and 15, check around 8 until you find the right balance.    \n\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/59880",
          "votes": 7
        }
      ]
    },
    {
      "id": 349739,
      "postDate": "2018-06-28T14:59:51.740Z",
      "content": "<p>Hi Joe and James, thank you for your detailed writeup, and congratulation on your 4th place!<br>\nTwo question to Joe, <br>\n1) Why did you choose tf-idf and svd as the matrix factorization method of categorical data? <br>\nComing from talkingdata, my first instinct was to use LDA(talkingdata 1st place solution) on label-encoded categories.<br>\n2) You chose 500 for SVD, 30000 for kmeans, 50000 for tfidf. How did you tune the numbers? <br>\nWas it trial and error on you local CV? Or did those numbers come intuitively?</p>",
      "rawMarkdown": "Hi Joe and James, thank you for your detailed writeup, and congratulation on your 4th place!<br>\nTwo question to Joe, <br>\n1) Why did you choose tf-idf and svd as the matrix factorization method of categorical data? <br>\nComing from talkingdata, my first instinct was to use LDA(talkingdata 1st place solution) on label-encoded categories.<br>\n2) You chose 500 for SVD, 30000 for kmeans, 50000 for tfidf. How did you tune the numbers? <br>\nWas it trial and error on you local CV? Or did those numbers come intuitively?\n",
      "votes": 2,
      "replies": [
        {
          "id": 349910,
          "postDate": "2018-06-28T20:41:48.663Z",
          "content": "<p>Thanks pocket!</p>\n\n<p>1) My approach only used matrix factorization of text data, not categorical (applied the tf-idf to the concatenation of title, description, params1-3). I tried some user-level categorical factorization too but it didn't seem to add anything (but see James' post, he saw success with flipping it to category-level factorization across users). I chose tf-idf -&gt; SVD because it's the classic approach used for latent semantic analysis on text / the easiest one for me to implement and tune, but it's very possible that LDA would have worked well for this too. LDA would have been a nice thing to try with more time. Also, it's possible that just sparse counts/tf-idf without factorization is more optimal (should have tried this). <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">Thousandvoices</a> used raw countvectorizer on user-concatenated titles, for example.</p>\n\n<p>2) I used a combination of intuition, model feedback, and time/space tradeoff decisions. With tf-idf I found that going up to 3-grams and 50k features gave a validation boost over simpler settings, but I think going higher than that had diminishing returns. For SVD components, values in the hundreds are a common heuristic for semantic models, but for the clustering process k-means run time scales with number of features and number of clusters, so too many of either makes things very slow. I tried fewer clusters at first and found it to not work well, and from my work with aggregating on title nouns I had additional evidence that very specific text groupings worked well. So I chose 30k as a tradeoff between granularity and excessive run time for the k-means algorithm, and found it to work better. For what it's worth, the text clustering didn't add much over the noun and noun+adjective aggregates, and I think it's possible that even much more than 30k clusters would have been preferable.       </p>",
          "rawMarkdown": "Thanks pocket!\n\n1) My approach only used matrix factorization of text data, not categorical (applied the tf-idf to the concatenation of title, description, params1-3). I tried some user-level categorical factorization too but it didn't seem to add anything (but see James' post, he saw success with flipping it to category-level factorization across users). I chose tf-idf -&gt; SVD because it's the classic approach used for latent semantic analysis on text / the easiest one for me to implement and tune, but it's very possible that LDA would have worked well for this too. LDA would have been a nice thing to try with more time. Also, it's possible that just sparse counts/tf-idf without factorization is more optimal (should have tried this). [Thousandvoices][1] used raw countvectorizer on user-concatenated titles, for example.\n\n2) I used a combination of intuition, model feedback, and time/space tradeoff decisions. With tf-idf I found that going up to 3-grams and 50k features gave a validation boost over simpler settings, but I think going higher than that had diminishing returns. For SVD components, values in the hundreds are a common heuristic for semantic models, but for the clustering process k-means run time scales with number of features and number of clusters, so too many of either makes things very slow. I tried fewer clusters at first and found it to not work well, and from my work with aggregating on title nouns I had additional evidence that very specific text groupings worked well. So I chose 30k as a tradeoff between granularity and excessive run time for the k-means algorithm, and found it to work better. For what it's worth, the text clustering didn't add much over the noun and noun+adjective aggregates, and I think it's possible that even much more than 30k clusters would have been preferable.       \n\n   \n\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/59880",
          "votes": 3
        },
        {
          "id": 350232,
          "postDate": "2018-06-29T12:46:56.723Z",
          "content": "<p>Thank you for the detailed response!<br>\nIt is always a pleasure to read your posts. <br>\nHope you will keep contributing to the community in the future as well :) </p>",
          "rawMarkdown": "Thank you for the detailed response!<br>\nIt is always a pleasure to read your posts. <br>\nHope you will keep contributing to the community in the future as well :) ",
          "votes": 1
        },
        {
          "id": 350453,
          "postDate": "2018-06-29T19:35:26.143Z",
          "content": "<p>Thanks for the kind words pocket! I'm very glad if my posts are useful.</p>",
          "rawMarkdown": "Thanks for the kind words pocket! I'm very glad if my posts are useful."
        }
      ]
    },
    {
      "id": 349706,
      "postDate": "2018-06-28T14:14:59.967Z",
      "content": "<p>Congratulations and thanks for the post!</p>",
      "rawMarkdown": "Congratulations and thanks for the post!\n",
      "votes": 2
    },
    {
      "id": 349421,
      "postDate": "2018-06-28T04:58:19.690Z",
      "content": "<p>Congratulations @Joe Eddy and team. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations @Joe Eddy and team. Thanks for sharing.",
      "votes": 2
    },
    {
      "id": 349338,
      "postDate": "2018-06-28T02:38:55.403Z",
      "content": "<p>Really impressive and thorough features engeneering.   Congrats  Joe Eddy and your team !</p>",
      "rawMarkdown": "Really impressive and thorough features engeneering.   Congrats  Joe Eddy and your team !",
      "votes": 2,
      "replies": [
        {
          "id": 349901,
          "postDate": "2018-06-28T20:21:56.180Z",
          "content": "<p>Thanks Serigne!</p>",
          "rawMarkdown": "Thanks Serigne!"
        }
      ]
    },
    {
      "id": 349321,
      "postDate": "2018-06-28T02:27:22.507Z",
      "content": "<p>Very interesting and impressive feature engineering work. I'd love to read that blog post.</p>",
      "rawMarkdown": "Very interesting and impressive feature engineering work. I'd love to read that blog post.",
      "votes": 2,
      "replies": [
        {
          "id": 349898,
          "postDate": "2018-06-28T20:21:16.400Z",
          "content": "<p>Thanks so much Peter! And congrats, looking forward to reading your team's approach if you share :)</p>",
          "rawMarkdown": "Thanks so much Peter! And congrats, looking forward to reading your team's approach if you share :)"
        }
      ]
    },
    {
      "id": 351104,
      "postDate": "2018-07-01T10:26:01.227Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 350965,
      "postDate": "2018-06-30T22:53:14.453Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 349413,
      "postDate": "2018-06-28T04:47:41.870Z",
      "content": "<p>Congratulations. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 349682,
      "author_name": "James Trotman",
      "author_url": "",
      "post_date": "2018-06-28T13:29:16.427000",
      "content": "<p>Thanks Joe – here’s a quick write up of what I did and our path to teaming up. I’ll avoid mentioning what <a href=\"/neongen\">@neongen</a> and <a href=\"/gphilippis\">@gphilippis</a> did so these two posts together still only describe half the ensemble!</p>\n\n<p>My general approach was to start stacking straight away, using a LightGBM stacker, testing new features in there along with a growing set of level 1 models. Just a plain 5 fold CV lead to the same 0.004 public LB delta others had mentioned, with very reliable correlation.</p>\n\n<p>I started off with sparse nets on word/char trigrams with custom tokenisation (using pymorphy2) and interactions, inspired by the <a href=\"https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\">top</a> <a href=\"https://www.kaggle.com/mchahhou/mercari-second-place-solution\">two</a> Mercari solutions, mainly because they are much faster to train and I thought they’d make a good start, these reached 0.2200 in CV. I also tried one BiGRU &amp; categorical embedding network based on the fasttext-russian-2m shared vectors, but did not work on it much and it stayed at 0.2245. Occasionally I’d build the LightGBM model as a standalone (with deeper trees), removing it’s L1 model prediction features, adding some tf-idf / SVD features, and adding that in as a new L1 model helped.</p>\n\n<p>I generated hundreds of features, here are a selection...</p>\n\n<h1>Photos</h1>\n\n<p>Inspired by this great share of <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">VGG16 features</a> I decided to use Kaggle Kernels to do all the image feature engineering, just a simple loop over the jpg files and dump the generated features into Python pickles for download, 6 hours runtime and 1Gb of space (and a GPU if you want!) is more than enough.</p>\n\n<h3>VGG16 features</h3>\n\n<p>I tried adding these (raw and log scaled) to sparse nets as extra inputs and got slight improvements. Also tried predicting the category using softmax (cars &amp; clothing were very predictable), and summarised those predictions, thinking it might pick up on mistakes, e.g. a confident picture of a car in a clothing listing or many other combinations.</p>\n\n<h3>Color Histograms</h3>\n\n<p>(just the RGB histograms with 32 bins per channel).</p>\n\n<ul>\n<li>SVD on normalised histograms down to 12 columns</li>\n<li>mean bin for each of R/G/B</li>\n<li>KL divergence between all of R/B, R/G etc.</li>\n</ul>\n\n<p>This surprised me – it weighed in at about 0.0004 with the feature set I had at the time.</p>\n\n<h3>ImageHash Hashes</h3>\n\n<p>ahash, dhash, phash, used for counts of images with same hash. Then aggregating over users, i.e. has a user listed something with a widely copied photo? Or are all their images unique?</p>\n\n<h3>Basic Info</h3>\n\n<p>Simply open the file and get the [width, height] dimensions, then the file size info from the ZipFile API – from this you can get to a pixel count and bytes_per_pixel which is a rough measure of image quality. (The CRC from the ZipFile API is also useful to use as another ‘hash’ like above.)</p>\n\n<p>After teaming, we went looking for more features and I adapted my Kernel code to generate these :</p>\n\n<h3>NIMA</h3>\n\n<p><a href=\"https://github.com/titu1994/neural-image-assessment\">The model</a> that <a href=\"https://www.kaggle.com/christofhenkel\">Dieter</a> kindly shared in the external data thread was small enough to run in a Kernel, scoring ~70 images per second single threaded with a GPU. The best of those features was the lower bound on the aesthetic score i.e. score_mean-score_stdv.</p>\n\n<h3>Image keypoints</h3>\n\n<p><a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59414#347781\">Shared here</a> by <a href=\"https://www.kaggle.com/nuhsikander\">Nooh</a> – I used a range of thresholds, 5 and 10 worked much better than higher ones.</p>\n\n<h1>Train / Test Active</h1>\n\n<p>The shared <code>avg_days_up_user</code> and <code>avg_times_up_user</code> features from <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">aggregated-features-lightgbm</a> proved very strong, I thought they were acting as a safe approximate form of user_id target coding. There were many deal_probability==0 training rows and this probably comes from held-back evidence Avito has that the item was at some point re-listed. If a user re-lists things a lot they’re probably not getting many deals...</p>\n\n<p>To expand on that I built a few models on train_active and test_active combined (only unique items), and scored them on train &amp; test to make new features. I chunked the files based on parent_category_name and created sparse feature sets for LightGBM, running similar loops:</p>\n\n<h3>Predict Re-list Probability:</h3>\n\n<ul>\n<li>Predict if item_id appears more than once, using categories &amp; tf-idf on titles with different tokenisations.</li>\n</ul>\n\n<h3>Predict Log Price:</h3>\n\n<ul>\n<li>Similar features: predict the log of the price, remove 0 / NA prices from training, but scoring for all train/test rows.</li>\n</ul>\n\n<p>These features helped a lot, the pred_relist_probability feature has 0.32 correlation to deal_probability, and user_id_max_pred_relist_prob was a top-3 feature in one base model.</p>\n\n<p>Adding pred_log_price gained about 0.0002 in CV and adding the interaction <code>log1p(price)-pred_log_price</code> as well as aggregations like user_id_mean_pred_log_price_diff added the same again.</p>\n\n<h1>Text Stats</h1>\n\n<p>I had old code from the <a href=\"https://www.kaggle.com/c/two-sigma-connect-rental-listing-inquiries/discussion/32146#178428\">Renthop competition</a> I could use here, I tried looking for things tf-idf / tokenisation type processing might miss, like turning descriptions into a kind of Morse code, e.g. replacing non exclamation marks with ‘a’ and summarizing the length of the string, and also using it as a new category, but this was a flop. It’s possible there were better things to find here, but it didn’t make the cut… Simpler punctuation counts and stats were enough.</p>\n\n<p>Simple title token statistics, word appearance counts transformed into probabilities and summed to give a naive title appearance probability helped to give some diversity, with a resulting spectrum of [ common .. rare ] words that overlaps in the middle. e.g. is there one rare word in the title or a few? The probabilities were from global counts in all of train_active, test_active, train and test combined, and also per category, e.g. p(see_word | category==Transport).</p>\n\n<p>Another view was: reduce the titles per category with tf-idf / SVD and compute centroids, then cosine distance to the category centroid: is it a ‘normal’ (common) listing for the category?</p>\n\n<p>One last trick: groupby user_id and zip their concatenated item descriptions, then take the length of the zipped encoding (trying both lowercase / original case). Then divide that by a user's listing count to give a compression ratio of their descriptions: are they repetitive or not? This might add a little something extra, but I liked Joe’s user semantics idea much more.</p>\n\n<h1>Matrix Factorizations</h1>\n\n<p>Some categorical combinations were small enough to do a full SVD and use the first N eigenvectors as features for tree models, for example [user_id, category_name_param_1 ], reducing listing counts (log1p scaled) with <code>np.linalg.svd</code> and using the first 12 vectors of each. The user_id vectors were the intended use: a ‘category’ space for each user describing what they list, but the category_name_param_1 space over user_id’s was used much more heavily by LGB, perhaps capturing category/param_1 combinations that are mainly ‘company’ listings versus ‘private’ listings in a better way.</p>\n\n<h1>Team Up</h1>\n\n<p>Before we teamed I guessed from my prospective team-mates past forum posts (e.g. <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\">by neongen</a> and <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44579\">by Joe</a> ) that we’d used quite different approaches and we merged as neighbouring top 10 entries on the leaderboard.</p>\n\n<p>A blend of our best solutions at 0.2161 and 0.2163 gave 0.2151, and combining all our L1 models (~40 by the end) into an LGB stacker with 240 features overall reached 0.2139. This is a great demonstration of diverse approaches combining well. We reached that peak after a few days and got stuck  - perhaps converging a bit – finding improvements proved very hard.</p>\n\n<p>This was my first time teaming up here and my teammates <a href=\"/aquatic\">@aquatic</a>, <a href=\"/neongen\">@neongen</a> and <a href=\"/gphilippis\">@gphilippis</a> are shining examples of exactly what you'd want from teammates, a relentless source of ideas to discuss and a constant flow of new L1 models! Thank you all guys, this was a great experience.</p>\n\n<p>Finally, thanks to Kaggle and Avito for hosting this competition – there were so many sources of variance to explore and surprises along the way, it’s both fun and great practice at honing ML intuition.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 351946,
          "author_name": "KishoreNeduri",
          "author_url": "",
          "post_date": "2018-07-03T10:29:30.397000",
          "content": "<p>Congrats!!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349305,
      "author_name": "EV",
      "author_url": "",
      "post_date": "2018-06-28T02:05:53.370000",
      "content": "<p>Congrats on your 4th place finish. Thanks for posting such a detailed solution. I used a strategy quite similar to your's in terms of determining price competitiveness in a peer group. Obviously you went quite far with that approach, and I now see what else I could have done (i.e. lots of things) Yeah, k-means at these data sizes is indeed slow. I discovered <a href=\"https://github.com/src-d/kmcuda\">https://github.com/src-d/kmcuda</a> along the way and that was blazingly fast on a GPU. I split the data by category_name before forming clusters so that further reduced my computation time.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 349894,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-28T20:19:35.793000",
          "content": "<p>Thanks! Your approach is very nice, thanks for sharing it and the link - I'll have to check this out, it looks great.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349975,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-29T01:14:45.353000",
          "content": "<p>Thanks <a href=\"/eigenvector\">@eigenvector</a> for that fast k-means link. I will definitely use it in the future.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 349323,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-06-28T02:28:29.703000",
      "content": "<p>Congrats Joe and thanks for the detailed solution... BTW, congratulations for becoming a Kaggle Master....</p>",
      "votes": 3,
      "replies": [
        {
          "id": 349899,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-28T20:21:33.023000",
          "content": "<p>Thanks Samrat!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349826,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-06-28T17:29:12.567000",
      "content": "<p>Great solution. Congrats!!!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 350150,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-06-29T09:30:38.387000",
      "content": "<blockquote>\n  <p>Price statistics \n   - row / 20th percentile, row / median, row / max</p>\n</blockquote>\n\n<p>Thanks for sharing. In this part, do you mean that you use row divide 20th percentile</p>",
      "votes": 1,
      "replies": [
        {
          "id": 350456,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-29T19:42:13.603000",
          "content": "<p>The idea was to get features that capture price characteristics relative to a grouping. So for example, I want to know how expensive an item is vs. other items that share the same image_top_1 category. To get a feature like that, I first compute aggregate statistics for the category - its 20th percentile, median etc. Then for each individual item I merge in the aggregate statistic as a new column, and divide the item's price by the value of that column (that's what I mean by row / 20th percentile). Then it would look like df['price_pct_image_top_1_20thp:price'] = df['price'] / df['image_top_1_20thp:price'].</p>\n\n<p>Hope that's clear :)  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 349506,
      "author_name": "alup",
      "author_url": "",
      "post_date": "2018-06-28T07:48:28.947000",
      "content": "<p>Congrats! Really impressive!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349449,
      "author_name": "apgeorg",
      "author_url": "",
      "post_date": "2018-06-28T06:08:32.123000",
      "content": "<p>Congrats and thx for the nice writeup :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349448,
      "author_name": "den3b",
      "author_url": "",
      "post_date": "2018-06-28T05:59:10.973000",
      "content": "<p>Great solution. Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349379,
      "author_name": "gufra",
      "author_url": "",
      "post_date": "2018-06-28T03:40:09.687000",
      "content": "<p>Granular textual aggregates idea is GREAT! thanks for share your thoughts about that ! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349372,
      "author_name": "nicapotato",
      "author_url": "",
      "post_date": "2018-06-28T03:27:31.347000",
      "content": "<p>Fantastic.. Any chance on the code going public? I am particularly interested in your detailed feature engineering :D</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 350306,
      "author_name": "Vadym",
      "author_url": "",
      "post_date": "2018-06-29T14:25:08.357000",
      "content": "<p>Thanks! Cool solution and great write-up!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 349775,
      "author_name": "PrasunMishra",
      "author_url": "",
      "post_date": "2018-06-28T16:31:38.583000",
      "content": "<p>Congratulations Joe! This is a fantastic achievement. I see that you have LGB num_leaves as 400. A couple of other successful solutions also went deep into depth/num_leaves while my conventional wisdom says that it will result in overfitting. In fact, most of the public kernels were around 5 depths and 2**5-1 (num_leaves=31). Any thoughts on this? Any decision criteria you suggest for this hyper parameter?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 349914,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-28T20:57:48.963000",
          "content": "<p>Thanks Prasun! I was surprised by this too - I've usually seen shallow trees (3-8 max depth) work best for boosting. <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">Thousandvoices</a> went all the way to 1024 leaves! Tree depth is definitely very problem dependent, and one way to think about it intuitively is that it controls the level of feature interactivity that's allowed. With a massive number of low signal sparse features (from something like tf-idf, in my case 100k sparse features), it's sensible to me that trees need to go deep to properly capture the multitude of very specific interactions like word usages. Also, I think there's a lot of granularity in this prediction problem in general - specific target values are associated with multi-level category combinations as well as the very particular items that fall within those combinations and the very specific traits of those items relative to their peers (e.g. why granular price aggregate features work well). Another thing worth mentioning is that LGB uses a leaf-wise tree growing algorithm, which allows it to capitalize on going extremely deep if there's better information gain to be had that way.</p>\n\n<p>I think the reasons I've outlined can be used to an extent as an intuitive decision criterion, but at the end of the day your best answer will always be from proper validation. When you start building a gradient boosting model, try pushing the max depth up until you stop seeing validation improvement, and narrow it down to a reasonable choice in something like a binary-style search. E.g. try depth 3, 8, 15, discover 8 works better than 3 and 15, check around 8 until you find the right balance.    </p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 349739,
      "author_name": "pocket",
      "author_url": "",
      "post_date": "2018-06-28T14:59:51.740000",
      "content": "<p>Hi Joe and James, thank you for your detailed writeup, and congratulation on your 4th place!<br>\nTwo question to Joe, <br>\n1) Why did you choose tf-idf and svd as the matrix factorization method of categorical data? <br>\nComing from talkingdata, my first instinct was to use LDA(talkingdata 1st place solution) on label-encoded categories.<br>\n2) You chose 500 for SVD, 30000 for kmeans, 50000 for tfidf. How did you tune the numbers? <br>\nWas it trial and error on you local CV? Or did those numbers come intuitively?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 349910,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-28T20:41:48.663000",
          "content": "<p>Thanks pocket!</p>\n\n<p>1) My approach only used matrix factorization of text data, not categorical (applied the tf-idf to the concatenation of title, description, params1-3). I tried some user-level categorical factorization too but it didn't seem to add anything (but see James' post, he saw success with flipping it to category-level factorization across users). I chose tf-idf -&gt; SVD because it's the classic approach used for latent semantic analysis on text / the easiest one for me to implement and tune, but it's very possible that LDA would have worked well for this too. LDA would have been a nice thing to try with more time. Also, it's possible that just sparse counts/tf-idf without factorization is more optimal (should have tried this). <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">Thousandvoices</a> used raw countvectorizer on user-concatenated titles, for example.</p>\n\n<p>2) I used a combination of intuition, model feedback, and time/space tradeoff decisions. With tf-idf I found that going up to 3-grams and 50k features gave a validation boost over simpler settings, but I think going higher than that had diminishing returns. For SVD components, values in the hundreds are a common heuristic for semantic models, but for the clustering process k-means run time scales with number of features and number of clusters, so too many of either makes things very slow. I tried fewer clusters at first and found it to not work well, and from my work with aggregating on title nouns I had additional evidence that very specific text groupings worked well. So I chose 30k as a tradeoff between granularity and excessive run time for the k-means algorithm, and found it to work better. For what it's worth, the text clustering didn't add much over the noun and noun+adjective aggregates, and I think it's possible that even much more than 30k clusters would have been preferable.       </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 350232,
          "author_name": "pocket",
          "author_url": "",
          "post_date": "2018-06-29T12:46:56.723000",
          "content": "<p>Thank you for the detailed response!<br>\nIt is always a pleasure to read your posts. <br>\nHope you will keep contributing to the community in the future as well :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 350453,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-29T19:35:26.143000",
          "content": "<p>Thanks for the kind words pocket! I'm very glad if my posts are useful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349706,
      "author_name": "btk1",
      "author_url": "",
      "post_date": "2018-06-28T14:14:59.967000",
      "content": "<p>Congratulations and thanks for the post!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 349421,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-06-28T04:58:19.690000",
      "content": "<p>Congratulations @Joe Eddy and team. Thanks for sharing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 349338,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-06-28T02:38:55.403000",
      "content": "<p>Really impressive and thorough features engeneering.   Congrats  Joe Eddy and your team !</p>",
      "votes": 2,
      "replies": [
        {
          "id": 349901,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-28T20:21:56.180000",
          "content": "<p>Thanks Serigne!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349321,
      "author_name": "Peter Hurford",
      "author_url": "",
      "post_date": "2018-06-28T02:27:22.507000",
      "content": "<p>Very interesting and impressive feature engineering work. I'd love to read that blog post.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 349898,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-28T20:21:16.400000",
          "content": "<p>Thanks so much Peter! And congrats, looking forward to reading your team's approach if you share :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 351104,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2018-07-01T10:26:01.227000",
      "content": "<p>Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 350965,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-30T22:53:14.453000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349413,
      "author_name": "Eric Vos",
      "author_url": "",
      "post_date": "2018-06-28T04:47:41.870000",
      "content": "<p>Congratulations. Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349294": "**Edit/update**: please also see [@jtrotman][1]'s writeup in the responses below! :) \n\nThank you @ Kaggle and Avito - this competition was just awesome. There were so many interesting facets of the data to explore, and this was a rare competition where I wish I had joined earlier and spent even more time with it instead of getting sick of the data by the end. And a huge, huge thank you to my team [@neongen][2], [@gphilippis][3], and [@jtrotman][4]. This was a great and inspiring team that I feel very lucky to have been a part of.\n\nOur solution is a lgbm stacker trained on a bunch of good base models and a chunk of our strongest features. Nonlinear stacking and including features for stacking definitely helped. My teammates can share more about that framework and all of the different models that went into the stack (lgbs, nns with lstm, sparse nns, weaker models like ridge), but my part of this thread will focus most on my biggest contribution to the result - a single lgbm model that scores .2175 public /.2213 private and the features that go into this model.\n\n**Hyperparameters**\n\nHyperparameter tuning was not my main focus, but I found that very deep trees with a low learning rate worked quite well. Here are the final parameters:\n\n    lgb_params = {\n                'boosting_type': 'gbdt',\n                'objective': 'regression',\n                'metric': 'rmse',\n                'learning_rate': 0.01,\n                'num_leaves': 400, \n                'colsample_bytree': .45\n              }\n\n**Features**\n\nNow to get to the good part and my focus in this competition - features. I ended up with about 800 tabular features (original + engineered), along with 100k tf-idf features for description and title + param_1. Here are the tf-idf settings:\n\n    TfidfVectorizer(stop_words=stopwords.words('russian'), \n                             lowercase=True, ngram_range=(1, 3),\n                             max_features=50000,\n                             sublinear_tf=True)\n\nI'd summarize my tabular features by saying that they try to extract as much information as possible using price and the text fields. At first I only used very simple image features like color channels and size, and the image features I added at the end only gave about a .0003 boost.\n\nMy core feature ideas:\n\n**Price statistics**: \n\nIt seemed clear early on that price information should have a very import impact on the target. So I wanted to extract price distribution information / statistics on pretty much every basis of aggregation I could come up with (e.g. by category, by category and city, by user_id, by image_top_1, etc.). For a bunch of different aggregates, I took these stats across all records  (with duplicate item_ids dropped):\n\n - 20th percentile\n - median\n - max\n - standard deviation\n - skew\n\nThen computed some relative price numbers for each post record with respect to the corresponding aggregate -- this measures how expensive or cheap the item is relative to the grouping\n- row / 20th percentile, row / median, row / max \n\nThese computations are at the core of my FE, including for the ideas just below. \n\n**Granular textual aggregates on price**: \n\nI believe my most important discovery in this competition was that applying my price statistics features to very specific text groupings worked extremely well, and these features are at the core of my model. Taking price / median price for each post's title alone gave me a big boost. I hypothesize that this may be a way to partially reverse engineer keyword search rankings (e.g. search sorted by price) and that search rankings heavily impact views and therefore deal_probability. Image that you're an avito user looking browsing for items - you search for a few specific keywords (e.g. \"red bicycle\"), sort by price, and start looking at many of the cheapest options. I've extended the idea to add more value in a few different ways that try to reduce the number of distinct groupings while still being granular enough to capture this level of signal -- I run price statistics on all of the following:\n\n - title_noun aggregates: I extract all nouns from each title, normalized them, removed duplicates and sorted them alphabetically, and then used the resulting key as a basis for aggregation\n - title_noun_adjs aggregates: same as above, but adding adjectives as well \n - title_cluster: using title tf-idf features, I run SVD with 500 components and form 30,000 k-means clusters on these components. k-means is very slow, so I used mini-batch k-means and limited it to 500 components / 30k clusters even though I think more of both may have performed better.\n - text_cluster: same as above, but based on a concatenation of title, description, and the param fields. \n\n**User semantics**:\n\nWhen you build ML models, it's easy to get stuck in a row-based world and lose sight of the many interesting relationships between the different rows. Normal aggregate feature engineering like column statistics help with this, but I got to a point beyond that where I realized there was still more to be done with the text fields to help capture user characteristics.\n    \nComing off of talkingdata where it saw some success, I had the idea to use matrix factorization for this. So for users in train/test, I concatenated their text fields across all rows (title, desc, params), ran tf-idf -&gt; 300 component SVD, adding these components as features. A performance tip here is to use hashingvectorizer like below rather than straight tf-idf, since it is significantly more memory efficient (at the cost of a bit of accuracy). Set n_features higher than the number of features you really want, since there will be collisions.\n\n    print('Applying hash vectorizer then tf-idf to text')\n    \n    print('Hash Vectorizer')\n    hv = HashingVectorizer(stop_words=stopwords.words('russian'), \n                           lowercase=True, ngram_range=(1, 3),\n                           n_features=200000)\n    hv_feats = hv.fit_transform(user_text_df['all_text'])\n    \n    print('TF-idf transformer')\n    tfidf_user_text = TfidfTransformer() \n    tfidf_user_text_feats = tfidf_user_text.fit_transform(hv_feats)\n\nAnother nice way to capture user-based information from text is with aggregations on meta-textual features like percentage of caps in title, etc. I took a bunch of meta-textual features and computed the same aggregate statistics that I did for price, across each user_id. For example, one of my top 30 features was the rather unfortunately named (by my conventions) \n\n    \"pct_caps_description_pct_user_id_median:pct_caps_description\"\n\nThis means taking the specific row's percentage of caps in description and dividing it by that user's median percentage of caps in the description. Maybe I should have called it \"user shoutiness relative to their norm\" instead.  \n\n**Image Features**\n\nFor a big part of the competition I thought the images were mostly a red herring. My worldview was very focused on exploiting the possibilities of textual information and thought its relationship with search rankings had a much more dominant effect than the images would. It was very interesting to hear that people were seeing significant gains from good image features, and I think this is the area where I would try to improve my model more if given more time. I used a few very simple features (color channels, size dimensions), and some nice features prepared by my teammates (they could explain in more detail) that gave about a .0003 boost to my model as the final feature addition --\n\n - Image blurness\n - Color histograms - SVD components and some statistical features\n - NIMA - activations, score, stds, and score + std\n\nAside from features mentioned, I used a bunch of other miscellaneous features similar to those seen in kernels - simple meta-textual features, lat/long + lat/long clusters, avg days a user's post is active and similar stats from the periods data.  I did clean and lemmatize title and description before applying tf-idf and some of the other text processing mentioned above, but I believe it made a pretty minor difference. \n \nIn addition to lgb I used these features in a few other models, e.g. a neural net that scores .2197 public. Architecture is very similar to what you can see in kernels - biGRU applied to title and description, category embeddings, numerics concatenated and passed through 2 dense layers. It was difficult to close the gap between lgb and NN with this feature set, I think largely because there were many null values as a result of aggregate stats on unique occurrences, and these aggregate stats were among the key features to the lgb. I tried some different imputation strategies, but nothing worked all that well (although I will say that imputing price in a smart way, whether by a groupby-average on some category combinations or by a dedicated price prediction model, improved the NN score by about .001). \n\n**Some Comments on Workflow**\n\nI want to end on a few comments about workflow that I hope may be useful. What works well in one competition might not be ideal for another, but maybe it's helpful :) I've already written about some of these approaches here: https://www.kaggle.com/c/avito-demand-prediction/discussion/56986#330023\n\nIf you do a lot of feature engineering (especially with categorical aggregates), you are likely creating what amounts to a miniature relational database. You don't have to go as heavyweight as creating an actual database, but I found it helpful to create a bunch of feather files for storing aggregate features then merge all the ones I wanted back into the main data to build models. For example I had many feature files like \"global_city_features.ftr\" that I would merge back into the training dataframe using city as a key. Storing features in a relational manner saves you processing time (run FE once, not every time you model), saves you disk space, and gives you a lot of flexibility for choosing which features to include in a model run (just don't merge them in if you don't want that set).\n\nI used greedy forward feature selection - engineer a new set of related features, add them to my current best model, see if validation result improves. If it doesn't improve, leave them out of the model. This is a fast and fairly reliable (if not optimal) way of selecting for good features to include in a model.\n\nI noticed early on that there was very little variance in RMSE between validation folds (&lt;.0002), so I felt pretty safe testing for feature contributions to my model by evaluating on only one fold at a time. This helped speed up iteration on feature sets.\n\nI've already written a lot and my teammates will have additions to make as well so I'll stop there. If you're not sick of my blabbering and are interested, I'll likely write a blog post about my ideas / approach and share a link here once it's up. I hope this was a useful read, and happy kaggling!\n  \n\n\n  [1]: https://www.kaggle.com/jtrotman\n  [2]: https://www.kaggle.com/neongen\n  [3]: https://www.kaggle.com/gphilippis\n  [4]: https://www.kaggle.com/jtrotman",
    "349682": "Thanks Joe – here’s a quick write up of what I did and our path to teaming up. I’ll avoid mentioning what @neongen and @gphilippis did so these two posts together still only describe half the ensemble!\n\nMy general approach was to start stacking straight away, using a LightGBM stacker, testing new features in there along with a growing set of level 1 models. Just a plain 5 fold CV lead to the same 0.004 public LB delta others had mentioned, with very reliable correlation.\n\nI started off with sparse nets on word/char trigrams with custom tokenisation (using pymorphy2) and interactions, inspired by the [top][1] [two][2] Mercari solutions, mainly because they are much faster to train and I thought they’d make a good start, these reached 0.2200 in CV. I also tried one BiGRU &amp; categorical embedding network based on the fasttext-russian-2m shared vectors, but did not work on it much and it stayed at 0.2245. Occasionally I’d build the LightGBM model as a standalone (with deeper trees), removing it’s L1 model prediction features, adding some tf-idf / SVD features, and adding that in as a new L1 model helped.\n\nI generated hundreds of features, here are a selection...\n\n# Photos\n\nInspired by this great share of [VGG16 features][3] I decided to use Kaggle Kernels to do all the image feature engineering, just a simple loop over the jpg files and dump the generated features into Python pickles for download, 6 hours runtime and 1Gb of space (and a GPU if you want!) is more than enough.\n\n### VGG16 features\nI tried adding these (raw and log scaled) to sparse nets as extra inputs and got slight improvements. Also tried predicting the category using softmax (cars &amp; clothing were very predictable), and summarised those predictions, thinking it might pick up on mistakes, e.g. a confident picture of a car in a clothing listing or many other combinations.\n\n### Color Histograms\n (just the RGB histograms with 32 bins per channel).\n\n- SVD on normalised histograms down to 12 columns\n- mean bin for each of R/G/B\n- KL divergence between all of R/B, R/G etc.\n\nThis surprised me – it weighed in at about 0.0004 with the feature set I had at the time.\n\n### ImageHash Hashes\nahash, dhash, phash, used for counts of images with same hash. Then aggregating over users, i.e. has a user listed something with a widely copied photo? Or are all their images unique?\n\n### Basic Info\nSimply open the file and get the [width, height] dimensions, then the file size info from the ZipFile API – from this you can get to a pixel count and bytes_per_pixel which is a rough measure of image quality. (The CRC from the ZipFile API is also useful to use as another ‘hash’ like above.)\n\nAfter teaming, we went looking for more features and I adapted my Kernel code to generate these :\n\n### NIMA\n[The model][4] that [Dieter][5] kindly shared in the external data thread was small enough to run in a Kernel, scoring ~70 images per second single threaded with a GPU. The best of those features was the lower bound on the aesthetic score i.e. score_mean-score_stdv.\n\n### Image keypoints\n[Shared here][6] by [Nooh][7] – I used a range of thresholds, 5 and 10 worked much better than higher ones.\n\n\n# Train / Test Active\n\nThe shared `avg_days_up_user` and `avg_times_up_user` features from [aggregated-features-lightgbm][8] proved very strong, I thought they were acting as a safe approximate form of user_id target coding. There were many deal_probability==0 training rows and this probably comes from held-back evidence Avito has that the item was at some point re-listed. If a user re-lists things a lot they’re probably not getting many deals...\n\nTo expand on that I built a few models on train_active and test_active combined (only unique items), and scored them on train &amp; test to make new features. I chunked the files based on parent_category_name and created sparse feature sets for LightGBM, running similar loops:\n\n### Predict Re-list Probability:\n\n - Predict if item_id appears more than once, using categories &amp; tf-idf on titles with different tokenisations.\n\n### Predict Log Price:\n\n - Similar features: predict the log of the price, remove 0 / NA prices from training, but scoring for all train/test rows.\n\nThese features helped a lot, the pred_relist_probability feature has 0.32 correlation to deal_probability, and user_id_max_pred_relist_prob was a top-3 feature in one base model.\n\nAdding pred_log_price gained about 0.0002 in CV and adding the interaction `log1p(price)-pred_log_price` as well as aggregations like user_id_mean_pred_log_price_diff added the same again.\n\n# Text Stats\n\nI had old code from the [Renthop competition][9] I could use here, I tried looking for things tf-idf / tokenisation type processing might miss, like turning descriptions into a kind of Morse code, e.g. replacing non exclamation marks with ‘a’ and summarizing the length of the string, and also using it as a new category, but this was a flop. It’s possible there were better things to find here, but it didn’t make the cut… Simpler punctuation counts and stats were enough.\n\nSimple title token statistics, word appearance counts transformed into probabilities and summed to give a naive title appearance probability helped to give some diversity, with a resulting spectrum of [ common .. rare ] words that overlaps in the middle. e.g. is there one rare word in the title or a few? The probabilities were from global counts in all of train_active, test_active, train and test combined, and also per category, e.g. p(see_word | category==Transport).\n\nAnother view was: reduce the titles per category with tf-idf / SVD and compute centroids, then cosine distance to the category centroid: is it a ‘normal’ (common) listing for the category?\n\nOne last trick: groupby user_id and zip their concatenated item descriptions, then take the length of the zipped encoding (trying both lowercase / original case). Then divide that by a user's listing count to give a compression ratio of their descriptions: are they repetitive or not? This might add a little something extra, but I liked Joe’s user semantics idea much more.\n\n# Matrix Factorizations\n\nSome categorical combinations were small enough to do a full SVD and use the first N eigenvectors as features for tree models, for example [user_id, category_name_param_1 ], reducing listing counts (log1p scaled) with `np.linalg.svd` and using the first 12 vectors of each. The user_id vectors were the intended use: a ‘category’ space for each user describing what they list, but the category_name_param_1 space over user_id’s was used much more heavily by LGB, perhaps capturing category/param_1 combinations that are mainly ‘company’ listings versus ‘private’ listings in a better way.\n\n# Team Up\n\nBefore we teamed I guessed from my prospective team-mates past forum posts (e.g. [by neongen][10] and [by Joe][11] ) that we’d used quite different approaches and we merged as neighbouring top 10 entries on the leaderboard.\n\nA blend of our best solutions at 0.2161 and 0.2163 gave 0.2151, and combining all our L1 models (~40 by the end) into an LGB stacker with 240 features overall reached 0.2139. This is a great demonstration of diverse approaches combining well. We reached that peak after a few days and got stuck  - perhaps converging a bit – finding improvements proved very hard.\n\nThis was my first time teaming up here and my teammates @aquatic, @neongen and @gphilippis are shining examples of exactly what you'd want from teammates, a relentless source of ideas to discuss and a constant flow of new L1 models! Thank you all guys, this was a great experience.\n\nFinally, thanks to Kaggle and Avito for hosting this competition – there were so many sources of variance to explore and surprises along the way, it’s both fun and great practice at honing ML intuition.\n\n\n  [1]: https://www.kaggle.com/lopuhin/mercari-golf-0-3875-cv-in-75-loc-1900-s\n  [2]: https://www.kaggle.com/mchahhou/mercari-second-place-solution\n  [3]: https://www.kaggle.com/bguberfain/vgg16-train-features\n  [4]: https://github.com/titu1994/neural-image-assessment\n  [5]: https://www.kaggle.com/christofhenkel\n  [6]: https://www.kaggle.com/c/avito-demand-prediction/discussion/59414#347781\n  [7]: https://www.kaggle.com/nuhsikander\n  [8]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\n  [9]: https://www.kaggle.com/c/two-sigma-connect-rental-listing-inquiries/discussion/32146#178428\n  [10]: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\n  [11]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44579\n",
    "349305": "Congrats on your 4th place finish. Thanks for posting such a detailed solution. I used a strategy quite similar to your's in terms of determining price competitiveness in a peer group. Obviously you went quite far with that approach, and I now see what else I could have done (i.e. lots of things) Yeah, k-means at these data sizes is indeed slow. I discovered https://github.com/src-d/kmcuda along the way and that was blazingly fast on a GPU. I split the data by category_name before forming clusters so that further reduced my computation time.\n",
    "349323": "Congrats Joe and thanks for the detailed solution... BTW, congratulations for becoming a Kaggle Master....",
    "349826": "Great solution. Congrats!!!",
    "350150": "  \n\n&gt;  Price statistics \n - row / 20th percentile, row / median, row / max\n\nThanks for sharing. In this part, do you mean that you use row divide 20th percentile",
    "349506": "Congrats! Really impressive!",
    "349449": "Congrats and thx for the nice writeup :)",
    "349448": "Great solution. Congrats!",
    "349379": "Granular textual aggregates idea is GREAT! thanks for share your thoughts about that ! ",
    "349372": "Fantastic.. Any chance on the code going public? I am particularly interested in your detailed feature engineering :D",
    "350306": "Thanks! Cool solution and great write-up!",
    "349775": "Congratulations Joe! This is a fantastic achievement. I see that you have LGB num_leaves as 400. A couple of other successful solutions also went deep into depth/num_leaves while my conventional wisdom says that it will result in overfitting. In fact, most of the public kernels were around 5 depths and 2**5-1 (num_leaves=31). Any thoughts on this? Any decision criteria you suggest for this hyper parameter?",
    "349739": "Hi Joe and James, thank you for your detailed writeup, and congratulation on your 4th place!<br>\nTwo question to Joe, <br>\n1) Why did you choose tf-idf and svd as the matrix factorization method of categorical data? <br>\nComing from talkingdata, my first instinct was to use LDA(talkingdata 1st place solution) on label-encoded categories.<br>\n2) You chose 500 for SVD, 30000 for kmeans, 50000 for tfidf. How did you tune the numbers? <br>\nWas it trial and error on you local CV? Or did those numbers come intuitively?\n",
    "349706": "Congratulations and thanks for the post!\n",
    "349421": "Congratulations @Joe Eddy and team. Thanks for sharing.",
    "349338": "Really impressive and thorough features engeneering.   Congrats  Joe Eddy and your team !",
    "349321": "Very interesting and impressive feature engineering work. I'd love to read that blog post.",
    "351104": "Congrats!",
    "350965": "",
    "349413": "Congratulations. Thanks for sharing."
  }
}