{
  "id": 59899,
  "title": "37th place solution",
  "url": "/competitions/avito-demand-prediction/writeups/attentionheads-37th-place-solution",
  "author_name": "",
  "post_date": "2018-06-28T06:59:50.876713900Z",
  "votes": 36,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, congratulations to all the winners, especially solo gold winners! </p>\n\n<p>I would also like to thank my team members @cutlass90 and @yaroshevskiy for your invaluable contribution to our solution!</p>\n\n<p>Now I will briefly describe our approach. We didn't use sophisticated feature engineering nor complicated neural network based models. Our main model is a fairly simple lgbm with a handful of hand-engineered features and some metafeatures coming from the other models.</p>\n\n<p>Hand-engineered features:</p>\n\n<ul>\n<li>Price aggregations: mean, median, var, min/max over the groups of categorical features (approximately 100 features in total). After joining these aggregations to the main dataframe we divided all these features on the price of the item to get some sort of relative price compared to the average price on the market (responsible for approx. 0.0004 improvement).</li>\n<li>Number of unique users in each city, region + category_name</li>\n<li>Text features: length, word count, number of uppercase/special symbols/punctuation vs length, number of stopwords vs word count</li>\n<li>n/a features: separate boolean columns for price/image missing (gave 0.0001 improvement)</li>\n<li>Average days active per category/city/param_1 calculated from periods data</li>\n</ul>\n\n<p>Metafeatures:</p>\n\n<ul>\n<li>RNNs (BiGRU + pool/attention) over title/description with trainable embeddings</li>\n<li>Same networks with pretrained fasttext embeddings (Wikipedia). It seems that training own fasttext model on active data could be a good idea, we didn't do it though.</li>\n<li>Sparse feedforward NN on top of Tfidf features + categorical/numerical features (responsible for approx. 0.0008 improvement). The model is very simple: sparse matrix multiplication on tfidf matrix -&gt; concatenate with categorical embeddings + numerical features -&gt; 2-3 dense layers with PReLU and Dropout -&gt; sigmoid (we observed cross-entropy works slightly better than RMSE as an objective)</li>\n<li>Finetuned resnet34 with ImageNet weights. We used predicts from two checkpoints (2nd and 3rd epoch). The former was a bit underfitted and the latter was a bit overfitted. Together they formed two good metafeatures responsible for approx. 0.001 improvement.</li>\n<li>Simple feedforward neural network on top of the hidden state of a denoising autoencoder. The autoencoder was trained on train/test active data in the similar fashion described in <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">this</a> fantastic thread. We didn't have enough time to optimize the parameters of the autoencoder, but finally we got around 0.0002 improvement from it.</li>\n<li>CNN for images + RNN for texts model. It wasn't properly optimized but gave a nice improvement of 0.0002.</li>\n</ul>\n\n<p>We also trained a bunch of other models that didn't work well such as FM-like neural network and Transformer-like model (from Attention Is All You Need paper).</p>\n\n<p>We trained a simple neural network on active data to estimate price and number of active days (gave approx. 0.0003 improvement).</p>\n\n<p>The final model is a simple blend of several trained boosted trees on slightly different feature sets.</p>\n\n<p>Once again, thank all the competitors for your work and fruitful discussions! We hope to learn a lot exploring the winning solutions. See you next time!</p>",
  "messages": [
    {
      "id": "349473",
      "postDate": "06/28/2018 06:59:50",
      "content": "<p>First of all, congratulations to all the winners, especially solo gold winners! </p>\n\n<p>I would also like to thank my team members @cutlass90 and @yaroshevskiy for your invaluable contribution to our solution!</p>\n\n<p>Now I will briefly describe our approach. We didn't use sophisticated feature engineering nor complicated neural network based models. Our main model is a fairly simple lgbm with a handful of hand-engineered features and some metafeatures coming from the other models.</p>\n\n<p>Hand-engineered features:</p>\n\n<ul>\n<li>Price aggregations: mean, median, var, min/max over the groups of categorical features (approximately 100 features in total). After joining these aggregations to the main dataframe we divided all these features on the price of the item to get some sort of relative price compared to the average price on the market (responsible for approx. 0.0004 improvement).</li>\n<li>Number of unique users in each city, region + category_name</li>\n<li>Text features: length, word count, number of uppercase/special symbols/punctuation vs length, number of stopwords vs word count</li>\n<li>n/a features: separate boolean columns for price/image missing (gave 0.0001 improvement)</li>\n<li>Average days active per category/city/param_1 calculated from periods data</li>\n</ul>\n\n<p>Metafeatures:</p>\n\n<ul>\n<li>RNNs (BiGRU + pool/attention) over title/description with trainable embeddings</li>\n<li>Same networks with pretrained fasttext embeddings (Wikipedia). It seems that training own fasttext model on active data could be a good idea, we didn't do it though.</li>\n<li>Sparse feedforward NN on top of Tfidf features + categorical/numerical features (responsible for approx. 0.0008 improvement). The model is very simple: sparse matrix multiplication on tfidf matrix -&gt; concatenate with categorical embeddings + numerical features -&gt; 2-3 dense layers with PReLU and Dropout -&gt; sigmoid (we observed cross-entropy works slightly better than RMSE as an objective)</li>\n<li>Finetuned resnet34 with ImageNet weights. We used predicts from two checkpoints (2nd and 3rd epoch). The former was a bit underfitted and the latter was a bit overfitted. Together they formed two good metafeatures responsible for approx. 0.001 improvement.</li>\n<li>Simple feedforward neural network on top of the hidden state of a denoising autoencoder. The autoencoder was trained on train/test active data in the similar fashion described in <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">this</a> fantastic thread. We didn't have enough time to optimize the parameters of the autoencoder, but finally we got around 0.0002 improvement from it.</li>\n<li>CNN for images + RNN for texts model. It wasn't properly optimized but gave a nice improvement of 0.0002.</li>\n</ul>\n\n<p>We also trained a bunch of other models that didn't work well such as FM-like neural network and Transformer-like model (from Attention Is All You Need paper).</p>\n\n<p>We trained a simple neural network on active data to estimate price and number of active days (gave approx. 0.0003 improvement).</p>\n\n<p>The final model is a simple blend of several trained boosted trees on slightly different feature sets.</p>\n\n<p>Once again, thank all the competitors for your work and fruitful discussions! We hope to learn a lot exploring the winning solutions. See you next time!</p>",
      "rawMarkdown": "First of all, congratulations to all the winners, especially solo gold winners! \n\nI would also like to thank my team members @cutlass90 and @yaroshevskiy for your invaluable contribution to our solution!\n\nNow I will briefly describe our approach. We didn't use sophisticated feature engineering nor complicated neural network based models. Our main model is a fairly simple lgbm with a handful of hand-engineered features and some metafeatures coming from the other models.\n\nHand-engineered features:\n\n - Price aggregations: mean, median, var, min/max over the groups of categorical features (approximately 100 features in total). After joining these aggregations to the main dataframe we divided all these features on the price of the item to get some sort of relative price compared to the average price on the market (responsible for approx. 0.0004 improvement).\n - Number of unique users in each city, region + category_name\n - Text features: length, word count, number of uppercase/special symbols/punctuation vs length, number of stopwords vs word count\n - n/a features: separate boolean columns for price/image missing (gave 0.0001 improvement)\n - Average days active per category/city/param_1 calculated from periods data\n\nMetafeatures:\n\n- RNNs (BiGRU + pool/attention) over title/description with trainable embeddings\n- Same networks with pretrained fasttext embeddings (Wikipedia). It seems that training own fasttext model on active data could be a good idea, we didn't do it though.\n- Sparse feedforward NN on top of Tfidf features + categorical/numerical features (responsible for approx. 0.0008 improvement). The model is very simple: sparse matrix multiplication on tfidf matrix -&gt; concatenate with categorical embeddings + numerical features -&gt; 2-3 dense layers with PReLU and Dropout -&gt; sigmoid (we observed cross-entropy works slightly better than RMSE as an objective)\n- Finetuned resnet34 with ImageNet weights. We used predicts from two checkpoints (2nd and 3rd epoch). The former was a bit underfitted and the latter was a bit overfitted. Together they formed two good metafeatures responsible for approx. 0.001 improvement.\n- Simple feedforward neural network on top of the hidden state of a denoising autoencoder. The autoencoder was trained on train/test active data in the similar fashion described in [this][1] fantastic thread. We didn't have enough time to optimize the parameters of the autoencoder, but finally we got around 0.0002 improvement from it.\n- CNN for images + RNN for texts model. It wasn't properly optimized but gave a nice improvement of 0.0002.\n\nWe also trained a bunch of other models that didn't work well such as FM-like neural network and Transformer-like model (from Attention Is All You Need paper).\n\nWe trained a simple neural network on active data to estimate price and number of active days (gave approx. 0.0003 improvement).\n\nThe final model is a simple blend of several trained boosted trees on slightly different feature sets.\n\nOnce again, thank all the competitors for your work and fruitful discussions! We hope to learn a lot exploring the winning solutions. See you next time!\n\n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629",
      "votes": null
    },
    {
      "id": "350014",
      "postDate": "06/29/2018 03:03:56",
      "content": "<p>Congratulations @Dmitriy and team on a strong finish. Thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congratulations @Dmitriy and team on a strong finish. Thanks for sharing your solution overview.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 350014,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/29/2018 03:03:56",
      "content": "<p>Congratulations @Dmitriy and team on a strong finish. Thanks for sharing your solution overview.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349473": "First of all, congratulations to all the winners, especially solo gold winners! \n\nI would also like to thank my team members @cutlass90 and @yaroshevskiy for your invaluable contribution to our solution!\n\nNow I will briefly describe our approach. We didn't use sophisticated feature engineering nor complicated neural network based models. Our main model is a fairly simple lgbm with a handful of hand-engineered features and some metafeatures coming from the other models.\n\nHand-engineered features:\n\n - Price aggregations: mean, median, var, min/max over the groups of categorical features (approximately 100 features in total). After joining these aggregations to the main dataframe we divided all these features on the price of the item to get some sort of relative price compared to the average price on the market (responsible for approx. 0.0004 improvement).\n - Number of unique users in each city, region + category_name\n - Text features: length, word count, number of uppercase/special symbols/punctuation vs length, number of stopwords vs word count\n - n/a features: separate boolean columns for price/image missing (gave 0.0001 improvement)\n - Average days active per category/city/param_1 calculated from periods data\n\nMetafeatures:\n\n- RNNs (BiGRU + pool/attention) over title/description with trainable embeddings\n- Same networks with pretrained fasttext embeddings (Wikipedia). It seems that training own fasttext model on active data could be a good idea, we didn't do it though.\n- Sparse feedforward NN on top of Tfidf features + categorical/numerical features (responsible for approx. 0.0008 improvement). The model is very simple: sparse matrix multiplication on tfidf matrix -&gt; concatenate with categorical embeddings + numerical features -&gt; 2-3 dense layers with PReLU and Dropout -&gt; sigmoid (we observed cross-entropy works slightly better than RMSE as an objective)\n- Finetuned resnet34 with ImageNet weights. We used predicts from two checkpoints (2nd and 3rd epoch). The former was a bit underfitted and the latter was a bit overfitted. Together they formed two good metafeatures responsible for approx. 0.001 improvement.\n- Simple feedforward neural network on top of the hidden state of a denoising autoencoder. The autoencoder was trained on train/test active data in the similar fashion described in [this][1] fantastic thread. We didn't have enough time to optimize the parameters of the autoencoder, but finally we got around 0.0002 improvement from it.\n- CNN for images + RNN for texts model. It wasn't properly optimized but gave a nice improvement of 0.0002.\n\nWe also trained a bunch of other models that didn't work well such as FM-like neural network and Transformer-like model (from Attention Is All You Need paper).\n\nWe trained a simple neural network on active data to estimate price and number of active days (gave approx. 0.0003 improvement).\n\nThe final model is a simple blend of several trained boosted trees on slightly different feature sets.\n\nOnce again, thank all the competitors for your work and fruitful discussions! We hope to learn a lot exploring the winning solutions. See you next time!\n\n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629",
    "350014": "Congratulations @Dmitriy and team on a strong finish. Thanks for sharing your solution overview."
  },
  "source": "meta"
}