{
  "id": 60006,
  "title": "Our 30th Solution: In which our heroes tried Quantum Gravity, Adaptive Noise and other cool stuff…",
  "url": "/competitions/avito-demand-prediction/writeups/spain-belarus-singapore-our-30th-solution-in-which",
  "author_name": "",
  "post_date": "2018-07-04T09:01:34.130Z",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This was a great competition, with the most variety of data I have seen in kaggle, maybe a time series competition with the same kind of data present here would be great next 😄 Now to our solution</p>\n\n<p><strong>Features</strong></p>\n\n<p><em>Image Features</em> : Extracted simple features(all from public kernels), Imagenet pretrained VGG16 to SVD, NIMA(Neural Image assessment) avg and std from Mobilenet and NASnet, SVD on <a href=\"https://www.kaggle.com/the1owl/natural-growth-patterns-fractals-of-nature\">https://www.kaggle.com/the1owl/natural-growth-patterns-fractals-of-nature</a>.</p>\n\n<p><em>Text features</em>: wordbatch ngrams, text stats from public kernels, english text stats, tfidf translated english and russian , Several different embeddings(Fasttext self-trained on description, wordtovec self-trained on description, 3 embeddings from <a href=\"https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md\">https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md</a> preprocessed text according to how the embeddings was trained, english GLOVE wiki vectors, russian fasttext wiki vectors), SVD and LDA on description and title, Naive Bayes Tfidf, Sentiment </p>\n\n<p><em>Tabular features</em>: All features on public kernels, normalized log price by groupby param_2 train+test+test &amp; train_active with the idea that relative price difference between items in the same category matter, binning of 'norm_price' by different total number of discrete values, count of 'norm_price' bins by param_2 with test &amp; train_active, test &amp; train_active average item count per day of 10+ categorical features by using test &amp; train period using from &gt; to time, Target Encoding of all categoricals with smoothing and noise, almost 60 different aggregate counts of items groupby a set of categoricals and unique counts of categoricals grouped by different sets of categoricals including test and train active, impute by prediction with NN image top 1</p>\n\n<p><strong>Models</strong></p>\n\n<p>Our final submision was based on stacking and weighted averaging, it provides the boost we needed from our seemingly poor performing models, even the best ones.</p>\n\n<p>19 Level one models that consist of pure RNN , LGBM, Ridge, Lasso, Elastic Net, Random Forest Regressor, XGB , FM-FTRL, RNN + categoricals, RNN + categoricals + continuous, RNN + categoricals + continuous + images.</p>\n\n<p>10 Level two models build on top isotonic regression k-fold transformed level one predictions, consist of Linear Regression, LGBM, XGB, RNN + categoricals + continuous.</p>\n\n<p>3rd Level model is just a Linear Regression with isotonic regression</p>\n\n<p>Final submission a weighted average with best RNN + categoricals + continuous + images model and 3rd Model.</p>\n\n<p>Best models is LGBM- two of them with very uncorrelated predictions, both at 0.2203 Public LB trained with entirely different subset of features listed above, NNs with all type of feature 0.2207 Public LB</p>\n\n<p><strong>Notes</strong></p>\n\n<p>We had two distinct RNNs. One implemented by Chin which is just the best architecture adopted from toxic comment classification and the other implemented by @Pavel and @Andres. The later can be found here: <a href=\"https://github.com/antorsae/avito-demand-prediction\">https://github.com/antorsae/avito-demand-prediction</a> and didn't use as many features as described above, and it has a few different features.</p>\n\n<p><strong>Things that didn't work</strong></p>\n\n<p>Discretizing predictions aka \"Quantum Gravity\": As discovered by one competitor, predictions are very discrete, so we built a list of predictions grouped by category and added a post-processing layer to convert them to discrete values based on proximity and strength of discrete probability. We dubbed this approach <em>quantum gravity</em> but although the name was cool it didn't work.</p>\n\n<p>Noise: We added this the last day of the competition so our findings were inconclusive. We implemented \"swap noise\" and \"smart noise\", the first just swaps a fraction of columns by values of the same column in different samples, whereas the second picks samples to swap columns from whose prediction is similar to the current sample.\nAdaptive noise: We saw that controlling the rate of noise was very delicate and if set too low (e.g. <code>-fnr 0.1</code>)  the network would eventually overfit, and setting it too high (e.g. <code>-fnr 0.3</code>) would make the network converge very slowly or not converge at all; so we added a callback to adjust noise rate based in a target rmse. We didn't have time to test it properly.</p>\n\n<p>Image pixels: We implemented computing image features and optimizing them in multiple networks (all controlled by command line), and with support for freezing layers; while our initial tests showed promise, we did not have time/GPUs to run it at the end.</p>",
  "messages": [
    {
      "id": "350175",
      "postDate": "06/29/2018 10:35:14",
      "content": "<p>This was a great competition, with the most variety of data I have seen in kaggle, maybe a time series competition with the same kind of data present here would be great next 😄 Now to our solution</p>\n\n<p><strong>Features</strong></p>\n\n<p><em>Image Features</em> : Extracted simple features(all from public kernels), Imagenet pretrained VGG16 to SVD, NIMA(Neural Image assessment) avg and std from Mobilenet and NASnet, SVD on <a href=\"https://www.kaggle.com/the1owl/natural-growth-patterns-fractals-of-nature\">https://www.kaggle.com/the1owl/natural-growth-patterns-fractals-of-nature</a>.</p>\n\n<p><em>Text features</em>: wordbatch ngrams, text stats from public kernels, english text stats, tfidf translated english and russian , Several different embeddings(Fasttext self-trained on description, wordtovec self-trained on description, 3 embeddings from <a href=\"https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md\">https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md</a> preprocessed text according to how the embeddings was trained, english GLOVE wiki vectors, russian fasttext wiki vectors), SVD and LDA on description and title, Naive Bayes Tfidf, Sentiment </p>\n\n<p><em>Tabular features</em>: All features on public kernels, normalized log price by groupby param_2 train+test+test &amp; train_active with the idea that relative price difference between items in the same category matter, binning of 'norm_price' by different total number of discrete values, count of 'norm_price' bins by param_2 with test &amp; train_active, test &amp; train_active average item count per day of 10+ categorical features by using test &amp; train period using from &gt; to time, Target Encoding of all categoricals with smoothing and noise, almost 60 different aggregate counts of items groupby a set of categoricals and unique counts of categoricals grouped by different sets of categoricals including test and train active, impute by prediction with NN image top 1</p>\n\n<p><strong>Models</strong></p>\n\n<p>Our final submision was based on stacking and weighted averaging, it provides the boost we needed from our seemingly poor performing models, even the best ones.</p>\n\n<p>19 Level one models that consist of pure RNN , LGBM, Ridge, Lasso, Elastic Net, Random Forest Regressor, XGB , FM-FTRL, RNN + categoricals, RNN + categoricals + continuous, RNN + categoricals + continuous + images.</p>\n\n<p>10 Level two models build on top isotonic regression k-fold transformed level one predictions, consist of Linear Regression, LGBM, XGB, RNN + categoricals + continuous.</p>\n\n<p>3rd Level model is just a Linear Regression with isotonic regression</p>\n\n<p>Final submission a weighted average with best RNN + categoricals + continuous + images model and 3rd Model.</p>\n\n<p>Best models is LGBM- two of them with very uncorrelated predictions, both at 0.2203 Public LB trained with entirely different subset of features listed above, NNs with all type of feature 0.2207 Public LB</p>\n\n<p><strong>Notes</strong></p>\n\n<p>We had two distinct RNNs. One implemented by Chin which is just the best architecture adopted from toxic comment classification and the other implemented by @Pavel and @Andres. The later can be found here: <a href=\"https://github.com/antorsae/avito-demand-prediction\">https://github.com/antorsae/avito-demand-prediction</a> and didn't use as many features as described above, and it has a few different features.</p>\n\n<p><strong>Things that didn't work</strong></p>\n\n<p>Discretizing predictions aka \"Quantum Gravity\": As discovered by one competitor, predictions are very discrete, so we built a list of predictions grouped by category and added a post-processing layer to convert them to discrete values based on proximity and strength of discrete probability. We dubbed this approach <em>quantum gravity</em> but although the name was cool it didn't work.</p>\n\n<p>Noise: We added this the last day of the competition so our findings were inconclusive. We implemented \"swap noise\" and \"smart noise\", the first just swaps a fraction of columns by values of the same column in different samples, whereas the second picks samples to swap columns from whose prediction is similar to the current sample.\nAdaptive noise: We saw that controlling the rate of noise was very delicate and if set too low (e.g. <code>-fnr 0.1</code>)  the network would eventually overfit, and setting it too high (e.g. <code>-fnr 0.3</code>) would make the network converge very slowly or not converge at all; so we added a callback to adjust noise rate based in a target rmse. We didn't have time to test it properly.</p>\n\n<p>Image pixels: We implemented computing image features and optimizing them in multiple networks (all controlled by command line), and with support for freezing layers; while our initial tests showed promise, we did not have time/GPUs to run it at the end.</p>",
      "rawMarkdown": "This was a great competition, with the most variety of data I have seen in kaggle, maybe a time series competition with the same kind of data present here would be great next 😄 Now to our solution\n\n**Features**\n\n*Image Features* : Extracted simple features(all from public kernels), Imagenet pretrained VGG16 to SVD, NIMA(Neural Image assessment) avg and std from Mobilenet and NASnet, SVD on https://www.kaggle.com/the1owl/natural-growth-patterns-fractals-of-nature.\n\n*Text features*: wordbatch ngrams, text stats from public kernels, english text stats, tfidf translated english and russian , Several different embeddings(Fasttext self-trained on description, wordtovec self-trained on description, 3 embeddings from https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md preprocessed text according to how the embeddings was trained, english GLOVE wiki vectors, russian fasttext wiki vectors), SVD and LDA on description and title, Naive Bayes Tfidf, Sentiment \n\n*Tabular features*: All features on public kernels, normalized log price by groupby param_2 train+test+test &amp; train_active with the idea that relative price difference between items in the same category matter, binning of 'norm_price' by different total number of discrete values, count of 'norm_price' bins by param_2 with test &amp; train_active, test &amp; train_active average item count per day of 10+ categorical features by using test &amp; train period using from &gt; to time, Target Encoding of all categoricals with smoothing and noise, almost 60 different aggregate counts of items groupby a set of categoricals and unique counts of categoricals grouped by different sets of categoricals including test and train active, impute by prediction with NN image top 1\n\n**Models**\n\nOur final submision was based on stacking and weighted averaging, it provides the boost we needed from our seemingly poor performing models, even the best ones.\n\n19 Level one models that consist of pure RNN , LGBM, Ridge, Lasso, Elastic Net, Random Forest Regressor, XGB , FM-FTRL, RNN + categoricals, RNN + categoricals + continuous, RNN + categoricals + continuous + images.\n\n10 Level two models build on top isotonic regression k-fold transformed level one predictions, consist of Linear Regression, LGBM, XGB, RNN + categoricals + continuous.\n\n3rd Level model is just a Linear Regression with isotonic regression\n\nFinal submission a weighted average with best RNN + categoricals + continuous + images model and 3rd Model.\n\nBest models is LGBM- two of them with very uncorrelated predictions, both at 0.2203 Public LB trained with entirely different subset of features listed above, NNs with all type of feature 0.2207 Public LB\n\n**Notes**\n\nWe had two distinct RNNs. One implemented by Chin which is just the best architecture adopted from toxic comment classification and the other implemented by @Pavel and @Andres. The later can be found here: https://github.com/antorsae/avito-demand-prediction and didn't use as many features as described above, and it has a few different features.\n\n**Things that didn't work**\n\nDiscretizing predictions aka \"Quantum Gravity\": As discovered by one competitor, predictions are very discrete, so we built a list of predictions grouped by category and added a post-processing layer to convert them to discrete values based on proximity and strength of discrete probability. We dubbed this approach _quantum gravity_ but although the name was cool it didn't work.\n\nNoise: We added this the last day of the competition so our findings were inconclusive. We implemented \"swap noise\" and \"smart noise\", the first just swaps a fraction of columns by values of the same column in different samples, whereas the second picks samples to swap columns from whose prediction is similar to the current sample.\nAdaptive noise: We saw that controlling the rate of noise was very delicate and if set too low (e.g. `-fnr 0.1`)  the network would eventually overfit, and setting it too high (e.g. `-fnr 0.3`) would make the network converge very slowly or not converge at all; so we added a callback to adjust noise rate based in a target rmse. We didn't have time to test it properly.\n\nImage pixels: We implemented computing image features and optimizing them in multiple networks (all controlled by command line), and with support for freezing layers; while our initial tests showed promise, we did not have time/GPUs to run it at the end.",
      "votes": null
    },
    {
      "id": "350289",
      "postDate": "06/29/2018 14:05:39",
      "content": "<p>Forgot to mention that the NN had three heads which tried to predict whether the <code>deal_probability</code> was zero or not (binary classification), and <code>imgtop1</code> which would condition the main FC head. This worked, as often is the case with joint learning (even if you are only interested in one of the tasks). I think of it as a tiny autoencoder whispering hits to the main FC head.</p>",
      "rawMarkdown": "Forgot to mention that the NN had three heads which tried to predict whether the `deal_probability` was zero or not (binary classification), and `imgtop1` which would condition the main FC head. This worked, as often is the case with joint learning (even if you are only interested in one of the tasks). I think of it as a tiny autoencoder whispering hits to the main FC head.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 350289,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "06/29/2018 14:05:39",
      "content": "<p>Forgot to mention that the NN had three heads which tried to predict whether the <code>deal_probability</code> was zero or not (binary classification), and <code>imgtop1</code> which would condition the main FC head. This worked, as often is the case with joint learning (even if you are only interested in one of the tasks). I think of it as a tiny autoencoder whispering hits to the main FC head.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "350175": "This was a great competition, with the most variety of data I have seen in kaggle, maybe a time series competition with the same kind of data present here would be great next 😄 Now to our solution\n\n**Features**\n\n*Image Features* : Extracted simple features(all from public kernels), Imagenet pretrained VGG16 to SVD, NIMA(Neural Image assessment) avg and std from Mobilenet and NASnet, SVD on https://www.kaggle.com/the1owl/natural-growth-patterns-fractals-of-nature.\n\n*Text features*: wordbatch ngrams, text stats from public kernels, english text stats, tfidf translated english and russian , Several different embeddings(Fasttext self-trained on description, wordtovec self-trained on description, 3 embeddings from https://github.com/deepmipt/DeepPavlov/blob/master/pretrained-vectors.md preprocessed text according to how the embeddings was trained, english GLOVE wiki vectors, russian fasttext wiki vectors), SVD and LDA on description and title, Naive Bayes Tfidf, Sentiment \n\n*Tabular features*: All features on public kernels, normalized log price by groupby param_2 train+test+test &amp; train_active with the idea that relative price difference between items in the same category matter, binning of 'norm_price' by different total number of discrete values, count of 'norm_price' bins by param_2 with test &amp; train_active, test &amp; train_active average item count per day of 10+ categorical features by using test &amp; train period using from &gt; to time, Target Encoding of all categoricals with smoothing and noise, almost 60 different aggregate counts of items groupby a set of categoricals and unique counts of categoricals grouped by different sets of categoricals including test and train active, impute by prediction with NN image top 1\n\n**Models**\n\nOur final submision was based on stacking and weighted averaging, it provides the boost we needed from our seemingly poor performing models, even the best ones.\n\n19 Level one models that consist of pure RNN , LGBM, Ridge, Lasso, Elastic Net, Random Forest Regressor, XGB , FM-FTRL, RNN + categoricals, RNN + categoricals + continuous, RNN + categoricals + continuous + images.\n\n10 Level two models build on top isotonic regression k-fold transformed level one predictions, consist of Linear Regression, LGBM, XGB, RNN + categoricals + continuous.\n\n3rd Level model is just a Linear Regression with isotonic regression\n\nFinal submission a weighted average with best RNN + categoricals + continuous + images model and 3rd Model.\n\nBest models is LGBM- two of them with very uncorrelated predictions, both at 0.2203 Public LB trained with entirely different subset of features listed above, NNs with all type of feature 0.2207 Public LB\n\n**Notes**\n\nWe had two distinct RNNs. One implemented by Chin which is just the best architecture adopted from toxic comment classification and the other implemented by @Pavel and @Andres. The later can be found here: https://github.com/antorsae/avito-demand-prediction and didn't use as many features as described above, and it has a few different features.\n\n**Things that didn't work**\n\nDiscretizing predictions aka \"Quantum Gravity\": As discovered by one competitor, predictions are very discrete, so we built a list of predictions grouped by category and added a post-processing layer to convert them to discrete values based on proximity and strength of discrete probability. We dubbed this approach _quantum gravity_ but although the name was cool it didn't work.\n\nNoise: We added this the last day of the competition so our findings were inconclusive. We implemented \"swap noise\" and \"smart noise\", the first just swaps a fraction of columns by values of the same column in different samples, whereas the second picks samples to swap columns from whose prediction is similar to the current sample.\nAdaptive noise: We saw that controlling the rate of noise was very delicate and if set too low (e.g. `-fnr 0.1`)  the network would eventually overfit, and setting it too high (e.g. `-fnr 0.3`) would make the network converge very slowly or not converge at all; so we added a callback to adjust noise rate based in a target rmse. We didn't have time to test it properly.\n\nImage pixels: We implemented computing image features and optimizing them in multiple networks (all controlled by command line), and with support for freezing layers; while our initial tests showed promise, we did not have time/GPUs to run it at the end.",
    "350289": "Forgot to mention that the NN had three heads which tried to predict whether the `deal_probability` was zero or not (binary classification), and `imgtop1` which would condition the main FC head. This worked, as often is the case with joint learning (even if you are only interested in one of the tasks). I think of it as a tiny autoencoder whispering hits to the main FC head."
  },
  "source": "meta"
}