{
  "id": 59959,
  "title": "A silver solution (31st place)",
  "url": "/competitions/avito-demand-prediction/discussion/59959",
  "author_name": "MPWARE",
  "post_date": "2018-06-28T20:12:31.061000",
  "votes": 27,
  "comment_count": 9,
  "views": 0,
  "content": "<p>It's time to share about solutions after this great challenge. Here is an insight of my models:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/349888/9726/solution.png\" alt=\"enter image description here\"></p>\n\n<p>A few different models followed by simple stacking with a CatBoost. My driver was diversity. \nEach model tuned with stratified CV5 and used to compute OOF for final stacking. Stratification is based on day of week + bin of deal probability.\nA few words about important items of each model:</p>\n\n<ul>\n<li>LGB: Based on pymorphy2 text normalization + TF-IDF + categories, images features (NIMA scores, INet scores and basic statistics).\nVGG16 features not used. SVD not used. Public LB 0.2200</li>\n<li>A set of NN with RNN (BiGRU, BiLSTM), embeddings and dense layers. \nRNN are based on fixed W2V and/or FastText (trained from all title + description available).\nThe one relying on Capsules provided best results (LB 0.2198). Other with Attention Context scored LB 0.2207.\nLast one with Conv2D scored 0.2212. One key point for NN was to take care of categories embedding dimensions. </li>\n<li>Stacked XGB with mix of target encoding, TF-IDF and WordBatch (seen in kernels). It scored LB 0.2201.</li>\n</ul>\n\n<p>Finally, main features + each model prediction were used as inputs of a CatBoost model that scored public LB 0.2171 (private 0.2210) and reached 32nd place.\nFrom what I already read in published solutions I should have spent more time on target encoding with different levels as it was key.\nI also tried a second layer of stacking (not in the picture) but results were comparable, so I published the simplest one.</p>\n\n<p>I had a lot of fun with this competition, I was my first steps with RNN and NLP, I learned a lot from kernels and discussions here. Thanks to all competitors!</p>",
  "messages": [
    {
      "id": 349888,
      "postDate": "2018-06-28T20:12:31.060Z",
      "content": "<p>It's time to share about solutions after this great challenge. Here is an insight of my models:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/349888/9726/solution.png\" alt=\"enter image description here\"></p>\n\n<p>A few different models followed by simple stacking with a CatBoost. My driver was diversity. \nEach model tuned with stratified CV5 and used to compute OOF for final stacking. Stratification is based on day of week + bin of deal probability.\nA few words about important items of each model:</p>\n\n<ul>\n<li>LGB: Based on pymorphy2 text normalization + TF-IDF + categories, images features (NIMA scores, INet scores and basic statistics).\nVGG16 features not used. SVD not used. Public LB 0.2200</li>\n<li>A set of NN with RNN (BiGRU, BiLSTM), embeddings and dense layers. \nRNN are based on fixed W2V and/or FastText (trained from all title + description available).\nThe one relying on Capsules provided best results (LB 0.2198). Other with Attention Context scored LB 0.2207.\nLast one with Conv2D scored 0.2212. One key point for NN was to take care of categories embedding dimensions. </li>\n<li>Stacked XGB with mix of target encoding, TF-IDF and WordBatch (seen in kernels). It scored LB 0.2201.</li>\n</ul>\n\n<p>Finally, main features + each model prediction were used as inputs of a CatBoost model that scored public LB 0.2171 (private 0.2210) and reached 32nd place.\nFrom what I already read in published solutions I should have spent more time on target encoding with different levels as it was key.\nI also tried a second layer of stacking (not in the picture) but results were comparable, so I published the simplest one.</p>\n\n<p>I had a lot of fun with this competition, I was my first steps with RNN and NLP, I learned a lot from kernels and discussions here. Thanks to all competitors!</p>",
      "rawMarkdown": "It's time to share about solutions after this great challenge. Here is an insight of my models:\n\n![enter image description here][1]\n\nA few different models followed by simple stacking with a CatBoost. My driver was diversity. \nEach model tuned with stratified CV5 and used to compute OOF for final stacking. Stratification is based on day of week + bin of deal probability.\nA few words about important items of each model:\n\n- LGB: Based on pymorphy2 text normalization + TF-IDF + categories, images features (NIMA scores, INet scores and basic statistics).\n  VGG16 features not used. SVD not used. Public LB 0.2200\n- A set of NN with RNN (BiGRU, BiLSTM), embeddings and dense layers. \n  RNN are based on fixed W2V and/or FastText (trained from all title + description available).\n  The one relying on Capsules provided best results (LB 0.2198). Other with Attention Context scored LB 0.2207.\n  Last one with Conv2D scored 0.2212. One key point for NN was to take care of categories embedding dimensions. \n- Stacked XGB with mix of target encoding, TF-IDF and WordBatch (seen in kernels). It scored LB 0.2201.\n\nFinally, main features + each model prediction were used as inputs of a CatBoost model that scored public LB 0.2171 (private 0.2210) and reached 32nd place.\nFrom what I already read in published solutions I should have spent more time on target encoding with different levels as it was key.\nI also tried a second layer of stacking (not in the picture) but results were comparable, so I published the simplest one.\n\nI had a lot of fun with this competition, I was my first steps with RNN and NLP, I learned a lot from kernels and discussions here. Thanks to all competitors!\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/349888/9726/solution.png",
      "votes": 27
    },
    {
      "id": 350049,
      "postDate": "2018-06-29T04:21:47.790Z",
      "content": "<blockquote>\n  <p>One key point for NN was to take care of categories embedding dimensions.</p>\n</blockquote>\n\n<p>Thanks for sharing. And I saw most embedding dimensions is size of categories,how to choose right dimensions?</p>",
      "rawMarkdown": "&gt; One key point for NN was to take care of categories embedding dimensions.\n\n\nThanks for sharing. And I saw most embedding dimensions is size of categories,how to choose right dimensions?",
      "replies": [
        {
          "id": 350096,
          "postDate": "2018-06-29T07:04:49.867Z",
          "content": "<p>I did not use size of categories and I limited it for high cardinality.   I made different evaluations with 64, 50, 32, 24, 12 and 10 and convergence and RMSE was quite better with small one. I used the following formula found in many NN articles about categories embedding: embed_size = min(MAXSIZE, (category_size + 1) // 2)</p>\n\n<pre><code> MAXSIZE = 12\n for category, category_size in categorical_features.items():\n      embed_size = min(MAXSIZE, (category_size + 1) // 2)\n      inp = Input(shape = (1, ), name=category)\n      emb = Embedding(category_size, embed_size, input_length=1)(inp)\n      ....\n</code></pre>\n\n<p>Also, I dropped user_id from category embedding because it was really too high.</p>\n\n<p>Here is one link about category size selection: <a href=\"https://towardsdatascience.com/deep-learning-structured-data-8d6a278f3088\">https://towardsdatascience.com/deep-learning-structured-data-8d6a278f3088</a></p>\n\n<p>See 'Choosing the embedding size\" section.</p>",
          "rawMarkdown": "I did not use size of categories and I limited it for high cardinality.   I made different evaluations with 64, 50, 32, 24, 12 and 10 and convergence and RMSE was quite better with small one. I used the following formula found in many NN articles about categories embedding: embed_size = min(MAXSIZE, (category_size + 1) // 2)\n\n\n     MAXSIZE = 12\n     for category, category_size in categorical_features.items():\n          embed_size = min(MAXSIZE, (category_size + 1) // 2)\n          inp = Input(shape = (1, ), name=category)\n          emb = Embedding(category_size, embed_size, input_length=1)(inp)\n          ....\n\nAlso, I dropped user_id from category embedding because it was really too high.\n\nHere is one link about category size selection: https://towardsdatascience.com/deep-learning-structured-data-8d6a278f3088\n\nSee 'Choosing the embedding size\" section.",
          "votes": 1
        },
        {
          "id": 350149,
          "postDate": "2018-06-29T09:27:58.283Z",
          "content": "<p>Thanks,in this competition I am very afraid to use NN model,it's so complicated and hard to tune, next time I will try to use it.</p>",
          "rawMarkdown": "Thanks,in this competition I am very afraid to use NN model,it's so complicated and hard to tune, next time I will try to use it."
        }
      ]
    },
    {
      "id": 349930,
      "postDate": "2018-06-28T21:48:47.233Z",
      "content": "<p>The idea comes from a kernel available Toxic Comment challenge. </p>\n\n<p>Here is the link: <a href=\"https://www.kaggle.com/yekenot/textcnn-2d-convolution\">https://www.kaggle.com/yekenot/textcnn-2d-convolution</a></p>",
      "rawMarkdown": "The idea comes from a kernel available Toxic Comment challenge. \n\nHere is the link: https://www.kaggle.com/yekenot/textcnn-2d-convolution"
    },
    {
      "id": 349928,
      "postDate": "2018-06-28T21:42:38.747Z",
      "content": "<p>Thanks for the sharing. How did you do conv2D with text? Thanks.</p>",
      "rawMarkdown": "Thanks for the sharing. How did you do conv2D with text? Thanks."
    },
    {
      "id": 350184,
      "postDate": "2018-06-29T10:57:07.913Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 350399,
          "postDate": "2018-06-29T17:20:19.230Z",
          "content": "<p>No package. Handmade powerpoint just with important items.</p>",
          "rawMarkdown": "No package. Handmade powerpoint just with important items."
        },
        {
          "id": 350610,
          "postDate": "2018-06-30T04:00:16.353Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 350666,
          "postDate": "2018-06-30T07:45:05.623Z",
          "content": "<p>I had another layer of models stacking but results were quite similar (it was my second submission) so I did not include it in this chart.</p>",
          "rawMarkdown": "I had another layer of models stacking but results were quite similar (it was my second submission) so I did not include it in this chart.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 350049,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-06-29T04:21:47.790000",
      "content": "<blockquote>\n  <p>One key point for NN was to take care of categories embedding dimensions.</p>\n</blockquote>\n\n<p>Thanks for sharing. And I saw most embedding dimensions is size of categories,how to choose right dimensions?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 350096,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2018-06-29T07:04:49.867000",
          "content": "<p>I did not use size of categories and I limited it for high cardinality.   I made different evaluations with 64, 50, 32, 24, 12 and 10 and convergence and RMSE was quite better with small one. I used the following formula found in many NN articles about categories embedding: embed_size = min(MAXSIZE, (category_size + 1) // 2)</p>\n\n<pre><code> MAXSIZE = 12\n for category, category_size in categorical_features.items():\n      embed_size = min(MAXSIZE, (category_size + 1) // 2)\n      inp = Input(shape = (1, ), name=category)\n      emb = Embedding(category_size, embed_size, input_length=1)(inp)\n      ....\n</code></pre>\n\n<p>Also, I dropped user_id from category embedding because it was really too high.</p>\n\n<p>Here is one link about category size selection: <a href=\"https://towardsdatascience.com/deep-learning-structured-data-8d6a278f3088\">https://towardsdatascience.com/deep-learning-structured-data-8d6a278f3088</a></p>\n\n<p>See 'Choosing the embedding size\" section.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 350149,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-29T09:27:58.283000",
          "content": "<p>Thanks,in this competition I am very afraid to use NN model,it's so complicated and hard to tune, next time I will try to use it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349930,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2018-06-28T21:48:47.233000",
      "content": "<p>The idea comes from a kernel available Toxic Comment challenge. </p>\n\n<p>Here is the link: <a href=\"https://www.kaggle.com/yekenot/textcnn-2d-convolution\">https://www.kaggle.com/yekenot/textcnn-2d-convolution</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349928,
      "author_name": "Jialun He",
      "author_url": "",
      "post_date": "2018-06-28T21:42:38.747000",
      "content": "<p>Thanks for the sharing. How did you do conv2D with text? Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 350184,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-29T10:57:07.913000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 350399,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2018-06-29T17:20:19.230000",
          "content": "<p>No package. Handmade powerpoint just with important items.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350610,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-30T04:00:16.353000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350666,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2018-06-30T07:45:05.623000",
          "content": "<p>I had another layer of models stacking but results were quite similar (it was my second submission) so I did not include it in this chart.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "349888": "It's time to share about solutions after this great challenge. Here is an insight of my models:\n\n![enter image description here][1]\n\nA few different models followed by simple stacking with a CatBoost. My driver was diversity. \nEach model tuned with stratified CV5 and used to compute OOF for final stacking. Stratification is based on day of week + bin of deal probability.\nA few words about important items of each model:\n\n- LGB: Based on pymorphy2 text normalization + TF-IDF + categories, images features (NIMA scores, INet scores and basic statistics).\n  VGG16 features not used. SVD not used. Public LB 0.2200\n- A set of NN with RNN (BiGRU, BiLSTM), embeddings and dense layers. \n  RNN are based on fixed W2V and/or FastText (trained from all title + description available).\n  The one relying on Capsules provided best results (LB 0.2198). Other with Attention Context scored LB 0.2207.\n  Last one with Conv2D scored 0.2212. One key point for NN was to take care of categories embedding dimensions. \n- Stacked XGB with mix of target encoding, TF-IDF and WordBatch (seen in kernels). It scored LB 0.2201.\n\nFinally, main features + each model prediction were used as inputs of a CatBoost model that scored public LB 0.2171 (private 0.2210) and reached 32nd place.\nFrom what I already read in published solutions I should have spent more time on target encoding with different levels as it was key.\nI also tried a second layer of stacking (not in the picture) but results were comparable, so I published the simplest one.\n\nI had a lot of fun with this competition, I was my first steps with RNN and NLP, I learned a lot from kernels and discussions here. Thanks to all competitors!\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/349888/9726/solution.png",
    "350049": "&gt; One key point for NN was to take care of categories embedding dimensions.\n\n\nThanks for sharing. And I saw most embedding dimensions is size of categories,how to choose right dimensions?",
    "349930": "The idea comes from a kernel available Toxic Comment challenge. \n\nHere is the link: https://www.kaggle.com/yekenot/textcnn-2d-convolution",
    "349928": "Thanks for the sharing. How did you do conv2D with text? Thanks.",
    "350184": ""
  }
}