{
  "id": 60005,
  "title": "Place 40, tried to get something from train_active and test_active with DAE",
  "url": "/competitions/avito-demand-prediction/discussion/60005",
  "author_name": "abzaliev",
  "post_date": "2018-06-29T10:29:34.206000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey guys, would like to share what I learned, will try to highlight only the things that may be interesting,</p>\n\n<p>First of all, thanks for the great competition and for learning - was awesome.</p>\n\n<p>From the beginning I thought that semi-supervised/unsupervised learning from active data may play a big role in this competition, and I concentrated on this. Spent waay to many time trying to reimplement Michael's Denoising Autoencoders from <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">here</a> in keras, there are many small details(like proper learning rate and proper standartization) that slowed me down. The idea was to take only categorical and numerical data (ignoring text and images), that avaliable for both train-test and train_active-test_active and train denoising autoencoder on the whole dataset. Afterwards replace these features in the supervised learning model with the output from the autoencoder. I assumed that this will allow me to skip feature generation and generalize better - was a partially wrong assumption =( I see now from your solution that feature generation could really have helped me. </p>\n\n<p>Overall, my solution is weighted average of 4 NN models  (based on public leaderboard, I manually have chosen the weights) , mainly those from <a href=\"https://www.kaggle.com/shadowwarrior\">Harlan</a> (<a href=\"https://www.kaggle.com/shadowwarrior/1st-place-solution/notebook\">here</a>),  with some twists (for instance going deep didn't work for me, wide LSTMs was better) and ResNet last layer image features mentioned by others. All predictions are 10 fold averages. I didn't use any gradient boosting models - was too lazy, and I wanted to test how much can I get from the deep learning alone. </p>\n\n<p>It was a big pleasure to learn from you and I see now how much room for improvement for me is there. Happy kaggling!</p>",
  "messages": [
    {
      "id": 350173,
      "postDate": "2018-06-29T10:29:34.207Z",
      "content": "<p>Hey guys, would like to share what I learned, will try to highlight only the things that may be interesting,</p>\n\n<p>First of all, thanks for the great competition and for learning - was awesome.</p>\n\n<p>From the beginning I thought that semi-supervised/unsupervised learning from active data may play a big role in this competition, and I concentrated on this. Spent waay to many time trying to reimplement Michael's Denoising Autoencoders from <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">here</a> in keras, there are many small details(like proper learning rate and proper standartization) that slowed me down. The idea was to take only categorical and numerical data (ignoring text and images), that avaliable for both train-test and train_active-test_active and train denoising autoencoder on the whole dataset. Afterwards replace these features in the supervised learning model with the output from the autoencoder. I assumed that this will allow me to skip feature generation and generalize better - was a partially wrong assumption =( I see now from your solution that feature generation could really have helped me. </p>\n\n<p>Overall, my solution is weighted average of 4 NN models  (based on public leaderboard, I manually have chosen the weights) , mainly those from <a href=\"https://www.kaggle.com/shadowwarrior\">Harlan</a> (<a href=\"https://www.kaggle.com/shadowwarrior/1st-place-solution/notebook\">here</a>),  with some twists (for instance going deep didn't work for me, wide LSTMs was better) and ResNet last layer image features mentioned by others. All predictions are 10 fold averages. I didn't use any gradient boosting models - was too lazy, and I wanted to test how much can I get from the deep learning alone. </p>\n\n<p>It was a big pleasure to learn from you and I see now how much room for improvement for me is there. Happy kaggling!</p>",
      "rawMarkdown": "Hey guys, would like to share what I learned, will try to highlight only the things that may be interesting,\n\nFirst of all, thanks for the great competition and for learning - was awesome.\n\nFrom the beginning I thought that semi-supervised/unsupervised learning from active data may play a big role in this competition, and I concentrated on this. Spent waay to many time trying to reimplement Michael's Denoising Autoencoders from [here][1] in keras, there are many small details(like proper learning rate and proper standartization) that slowed me down. The idea was to take only categorical and numerical data (ignoring text and images), that avaliable for both train-test and train_active-test_active and train denoising autoencoder on the whole dataset. Afterwards replace these features in the supervised learning model with the output from the autoencoder. I assumed that this will allow me to skip feature generation and generalize better - was a partially wrong assumption =( I see now from your solution that feature generation could really have helped me. \n\nOverall, my solution is weighted average of 4 NN models  (based on public leaderboard, I manually have chosen the weights) , mainly those from [Harlan][2] ([here][3]),  with some twists (for instance going deep didn't work for me, wide LSTMs was better) and ResNet last layer image features mentioned by others. All predictions are 10 fold averages. I didn't use any gradient boosting models - was too lazy, and I wanted to test how much can I get from the deep learning alone. \n\nIt was a big pleasure to learn from you and I see now how much room for improvement for me is there. Happy kaggling!\n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\n  [2]: https://www.kaggle.com/shadowwarrior\n  [3]: https://www.kaggle.com/shadowwarrior/1st-place-solution/notebook",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "350173": "Hey guys, would like to share what I learned, will try to highlight only the things that may be interesting,\n\nFirst of all, thanks for the great competition and for learning - was awesome.\n\nFrom the beginning I thought that semi-supervised/unsupervised learning from active data may play a big role in this competition, and I concentrated on this. Spent waay to many time trying to reimplement Michael's Denoising Autoencoders from [here][1] in keras, there are many small details(like proper learning rate and proper standartization) that slowed me down. The idea was to take only categorical and numerical data (ignoring text and images), that avaliable for both train-test and train_active-test_active and train denoising autoencoder on the whole dataset. Afterwards replace these features in the supervised learning model with the output from the autoencoder. I assumed that this will allow me to skip feature generation and generalize better - was a partially wrong assumption =( I see now from your solution that feature generation could really have helped me. \n\nOverall, my solution is weighted average of 4 NN models  (based on public leaderboard, I manually have chosen the weights) , mainly those from [Harlan][2] ([here][3]),  with some twists (for instance going deep didn't work for me, wide LSTMs was better) and ResNet last layer image features mentioned by others. All predictions are 10 fold averages. I didn't use any gradient boosting models - was too lazy, and I wanted to test how much can I get from the deep learning alone. \n\nIt was a big pleasure to learn from you and I see now how much room for improvement for me is there. Happy kaggling!\n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\n  [2]: https://www.kaggle.com/shadowwarrior\n  [3]: https://www.kaggle.com/shadowwarrior/1st-place-solution/notebook"
  }
}