{
  "id": 59902,
  "title": "25th place solution",
  "url": "/competitions/avito-demand-prediction/writeups/addimewe-25th-place-solution",
  "author_name": "",
  "post_date": "2018-06-28T07:12:54.114411300Z",
  "votes": 64,
  "comment_count": 20,
  "views": 0,
  "content": "<p>First of all, thanks to kaggle as well as avito for this challenging competition. </p>\n\n<p>Let me quickly give an overview of my part of our solution.</p>\n\n<p><strong>Summary</strong></p>\n\n<p>Until last day we used a single layer stacking of different NN and LGB models. For the final submission we also trained our stack with xgb and blended the results with our LGB stacker 1:1.\nThe best NN scored 0.2194 the best LGB 0.2196. Some notes on what helped and did not help.</p>\n\n<ul>\n<li>I spent a lot of time trying different factorization methods like LDA, libfm etc. but none improved my score</li>\n<li>What helped a lot where features derived in public kernels like city population from wikipedia etc. </li>\n<li>I made a test on pseudo-labelling and found out it helped, but we had no submission and time left to incorporate</li>\n<li>Features from pretrained image models (DenseNet, ILGNet, Ava) </li>\n<li>w2v embeddings trained from train active as described in one of my kernels</li>\n<li>a higher k-folder significantly improved the score, even more when bagging same models with different seeds</li>\n</ul>\n\n<p><strong>Feature engineering</strong></p>\n\n<p>We tried/ used most of the features from public kernels, however two interesting features we found were:</p>\n\n<ul>\n<li>number of duplicates per user. While Benjamin drops duplicates in his kernel we considered the number of duplicates important and this helped</li>\n<li>my teammate had a great idea seeing item_seq_number as a sequence. (sometimes the obvious is hard to see) and came up with the idea to use the price difference of consecutive items as feature, which improved LB a lot </li>\n</ul>\n\n<p><strong>NN structcture</strong></p>\n\n<p>I used the following NN structure, where I did TFIDF on character level on param_3  and w2v embeddings on text. I fed all numerical features we had (including a ridge model trained with price as target)</p>\n\n<pre><code>def build_model():\nsparse_params = Input(shape=[X_train['sparse_params'].shape[1]], dtype='float32', sparse=True, name='sparse_params')\n\ncategorical_inputs = []\nfor cat in cat_cols:\n    categorical_inputs.append(Input(shape=[1], name=cat))\n\ncategorical_embeddings = []\nfor i, cat in enumerate(cat_cols):\n    categorical_embeddings.append(\n        Embedding(embed_sizes[i], 10, embeddings_regularizer=l2(0.00001))(categorical_inputs[i]))\n\ncategorical_logits = Concatenate()([Flatten()(cat_emb) for cat_emb in categorical_embeddings])\ncategorical_logits = prense(categorical_logits, 256)\ncategorical_logits = prense(categorical_logits, 128)\n\nnumerical_inputs = Input(shape=[len(num_cols)], name='numerical')\n\nnumerical_logits = numerical_inputs\nnumerical_logits = BatchNormalization()(numerical_logits)\nnumerical_logits = prense(numerical_logits, 256)\nnumerical_logits = prense(numerical_logits, 128)\n\nparams_logits = prense(sparse_params, 64)\nparams_logits = prense(params_logits, 32)\n\ndesc_inp = Input(shape=[max_len_desc], name='desc')\ntitle_inp = Input(shape=[max_len_title], name='title')\nembedding = Embedding(nb_words, embed_size, weights=[embedding_matrix], trainable=False)  # nb_words\nemb_desc = embedding(desc_inp)\nemb_title = embedding(title_inp)\nemb_text = Concatenate(axis=1)([emb_desc,emb_title])\n\ntext_logits = SpatialDropout1D(0.2)(emb_text)\ntext_logits = Bidirectional(CuDNNLSTM(128, return_sequences=True))(text_logits)\ntext_logits = Conv1D(64, kernel_size=3, padding=\"valid\", kernel_initializer=\"glorot_uniform\")(text_logits)\navg_pool = GlobalAveragePooling1D()(text_logits)\nmax_pool = GlobalMaxPool1D()(text_logits)\ntext_logits = Concatenate()([avg_pool, max_pool])\nx = Dropout(0.2)(text_logits)\nx = Concatenate()([categorical_logits, text_logits])\nx = BatchNormalization()(x)\nx = Concatenate()([x, params_logits, numerical_logits])\nx = Dense(512, kernel_initializer=he_uniform(seed=0))(x)\nx = PReLU()(x)\nx = Dense(256, kernel_initializer=he_uniform(seed=0))(x)\nx = PReLU()(x)\nx = Dense(128, kernel_initializer=he_uniform(seed=0))(x)\nx = PReLU()(x)\nx = LayerNorm1D()(x)\nout = Dense(1, activation='sigmoid')(x)\n\nmodel = Model(inputs=[desc_inp] + [title_inp] + [sparse_params] + categorical_inputs + [numerical_inputs],\n              outputs=out)\n\nmodel.compile(optimizer=Adam(lr=0.0005, clipnorm=0.5), loss='mean_squared_error',\n              metrics=[root_mean_squared_error])\nreturn model\n</code></pre>",
  "messages": [
    {
      "id": "349481",
      "postDate": "06/28/2018 07:12:54",
      "content": "<p>First of all, thanks to kaggle as well as avito for this challenging competition. </p>\n\n<p>Let me quickly give an overview of my part of our solution.</p>\n\n<p><strong>Summary</strong></p>\n\n<p>Until last day we used a single layer stacking of different NN and LGB models. For the final submission we also trained our stack with xgb and blended the results with our LGB stacker 1:1.\nThe best NN scored 0.2194 the best LGB 0.2196. Some notes on what helped and did not help.</p>\n\n<ul>\n<li>I spent a lot of time trying different factorization methods like LDA, libfm etc. but none improved my score</li>\n<li>What helped a lot where features derived in public kernels like city population from wikipedia etc. </li>\n<li>I made a test on pseudo-labelling and found out it helped, but we had no submission and time left to incorporate</li>\n<li>Features from pretrained image models (DenseNet, ILGNet, Ava) </li>\n<li>w2v embeddings trained from train active as described in one of my kernels</li>\n<li>a higher k-folder significantly improved the score, even more when bagging same models with different seeds</li>\n</ul>\n\n<p><strong>Feature engineering</strong></p>\n\n<p>We tried/ used most of the features from public kernels, however two interesting features we found were:</p>\n\n<ul>\n<li>number of duplicates per user. While Benjamin drops duplicates in his kernel we considered the number of duplicates important and this helped</li>\n<li>my teammate had a great idea seeing item_seq_number as a sequence. (sometimes the obvious is hard to see) and came up with the idea to use the price difference of consecutive items as feature, which improved LB a lot </li>\n</ul>\n\n<p><strong>NN structcture</strong></p>\n\n<p>I used the following NN structure, where I did TFIDF on character level on param_3  and w2v embeddings on text. I fed all numerical features we had (including a ridge model trained with price as target)</p>\n\n<pre><code>def build_model():\nsparse_params = Input(shape=[X_train['sparse_params'].shape[1]], dtype='float32', sparse=True, name='sparse_params')\n\ncategorical_inputs = []\nfor cat in cat_cols:\n    categorical_inputs.append(Input(shape=[1], name=cat))\n\ncategorical_embeddings = []\nfor i, cat in enumerate(cat_cols):\n    categorical_embeddings.append(\n        Embedding(embed_sizes[i], 10, embeddings_regularizer=l2(0.00001))(categorical_inputs[i]))\n\ncategorical_logits = Concatenate()([Flatten()(cat_emb) for cat_emb in categorical_embeddings])\ncategorical_logits = prense(categorical_logits, 256)\ncategorical_logits = prense(categorical_logits, 128)\n\nnumerical_inputs = Input(shape=[len(num_cols)], name='numerical')\n\nnumerical_logits = numerical_inputs\nnumerical_logits = BatchNormalization()(numerical_logits)\nnumerical_logits = prense(numerical_logits, 256)\nnumerical_logits = prense(numerical_logits, 128)\n\nparams_logits = prense(sparse_params, 64)\nparams_logits = prense(params_logits, 32)\n\ndesc_inp = Input(shape=[max_len_desc], name='desc')\ntitle_inp = Input(shape=[max_len_title], name='title')\nembedding = Embedding(nb_words, embed_size, weights=[embedding_matrix], trainable=False)  # nb_words\nemb_desc = embedding(desc_inp)\nemb_title = embedding(title_inp)\nemb_text = Concatenate(axis=1)([emb_desc,emb_title])\n\ntext_logits = SpatialDropout1D(0.2)(emb_text)\ntext_logits = Bidirectional(CuDNNLSTM(128, return_sequences=True))(text_logits)\ntext_logits = Conv1D(64, kernel_size=3, padding=\"valid\", kernel_initializer=\"glorot_uniform\")(text_logits)\navg_pool = GlobalAveragePooling1D()(text_logits)\nmax_pool = GlobalMaxPool1D()(text_logits)\ntext_logits = Concatenate()([avg_pool, max_pool])\nx = Dropout(0.2)(text_logits)\nx = Concatenate()([categorical_logits, text_logits])\nx = BatchNormalization()(x)\nx = Concatenate()([x, params_logits, numerical_logits])\nx = Dense(512, kernel_initializer=he_uniform(seed=0))(x)\nx = PReLU()(x)\nx = Dense(256, kernel_initializer=he_uniform(seed=0))(x)\nx = PReLU()(x)\nx = Dense(128, kernel_initializer=he_uniform(seed=0))(x)\nx = PReLU()(x)\nx = LayerNorm1D()(x)\nout = Dense(1, activation='sigmoid')(x)\n\nmodel = Model(inputs=[desc_inp] + [title_inp] + [sparse_params] + categorical_inputs + [numerical_inputs],\n              outputs=out)\n\nmodel.compile(optimizer=Adam(lr=0.0005, clipnorm=0.5), loss='mean_squared_error',\n              metrics=[root_mean_squared_error])\nreturn model\n</code></pre>",
      "rawMarkdown": "First of all, thanks to kaggle as well as avito for this challenging competition. \n\nLet me quickly give an overview of my part of our solution.\n\n**Summary**\n\nUntil last day we used a single layer stacking of different NN and LGB models. For the final submission we also trained our stack with xgb and blended the results with our LGB stacker 1:1.\nThe best NN scored 0.2194 the best LGB 0.2196. Some notes on what helped and did not help.\n\n - I spent a lot of time trying different factorization methods like LDA, libfm etc. but none improved my score\n - What helped a lot where features derived in public kernels like city population from wikipedia etc. \n - I made a test on pseudo-labelling and found out it helped, but we had no submission and time left to incorporate\n - Features from pretrained image models (DenseNet, ILGNet, Ava) \n - w2v embeddings trained from train active as described in one of my kernels\n - a higher k-folder significantly improved the score, even more when bagging same models with different seeds\n \n**Feature engineering**\n\nWe tried/ used most of the features from public kernels, however two interesting features we found were:\n\n - number of duplicates per user. While Benjamin drops duplicates in his kernel we considered the number of duplicates important and this helped\n - my teammate had a great idea seeing item_seq_number as a sequence. (sometimes the obvious is hard to see) and came up with the idea to use the price difference of consecutive items as feature, which improved LB a lot \n\n\n\n**NN structcture**\n\nI used the following NN structure, where I did TFIDF on character level on param_3  and w2v embeddings on text. I fed all numerical features we had (including a ridge model trained with price as target)\n\n\n    def build_model():\n    sparse_params = Input(shape=[X_train['sparse_params'].shape[1]], dtype='float32', sparse=True, name='sparse_params')\n\n    categorical_inputs = []\n    for cat in cat_cols:\n        categorical_inputs.append(Input(shape=[1], name=cat))\n\n    categorical_embeddings = []\n    for i, cat in enumerate(cat_cols):\n        categorical_embeddings.append(\n            Embedding(embed_sizes[i], 10, embeddings_regularizer=l2(0.00001))(categorical_inputs[i]))\n\n    categorical_logits = Concatenate()([Flatten()(cat_emb) for cat_emb in categorical_embeddings])\n    categorical_logits = prense(categorical_logits, 256)\n    categorical_logits = prense(categorical_logits, 128)\n\n    numerical_inputs = Input(shape=[len(num_cols)], name='numerical')\n\n    numerical_logits = numerical_inputs\n    numerical_logits = BatchNormalization()(numerical_logits)\n    numerical_logits = prense(numerical_logits, 256)\n    numerical_logits = prense(numerical_logits, 128)\n\n    params_logits = prense(sparse_params, 64)\n    params_logits = prense(params_logits, 32)\n\n    desc_inp = Input(shape=[max_len_desc], name='desc')\n    title_inp = Input(shape=[max_len_title], name='title')\n    embedding = Embedding(nb_words, embed_size, weights=[embedding_matrix], trainable=False)  # nb_words\n    emb_desc = embedding(desc_inp)\n    emb_title = embedding(title_inp)\n    emb_text = Concatenate(axis=1)([emb_desc,emb_title])\n\n    text_logits = SpatialDropout1D(0.2)(emb_text)\n    text_logits = Bidirectional(CuDNNLSTM(128, return_sequences=True))(text_logits)\n    text_logits = Conv1D(64, kernel_size=3, padding=\"valid\", kernel_initializer=\"glorot_uniform\")(text_logits)\n    avg_pool = GlobalAveragePooling1D()(text_logits)\n    max_pool = GlobalMaxPool1D()(text_logits)\n    text_logits = Concatenate()([avg_pool, max_pool])\n    x = Dropout(0.2)(text_logits)\n    x = Concatenate()([categorical_logits, text_logits])\n    x = BatchNormalization()(x)\n    x = Concatenate()([x, params_logits, numerical_logits])\n    x = Dense(512, kernel_initializer=he_uniform(seed=0))(x)\n    x = PReLU()(x)\n    x = Dense(256, kernel_initializer=he_uniform(seed=0))(x)\n    x = PReLU()(x)\n    x = Dense(128, kernel_initializer=he_uniform(seed=0))(x)\n    x = PReLU()(x)\n    x = LayerNorm1D()(x)\n    out = Dense(1, activation='sigmoid')(x)\n\n    model = Model(inputs=[desc_inp] + [title_inp] + [sparse_params] + categorical_inputs + [numerical_inputs],\n                  outputs=out)\n\n    model.compile(optimizer=Adam(lr=0.0005, clipnorm=0.5), loss='mean_squared_error',\n                  metrics=[root_mean_squared_error])\n    return model",
      "votes": null
    },
    {
      "id": "349518",
      "postDate": "06/28/2018 08:12:29",
      "content": "<p>Congrats on the 25th place! Also I want to thank you a lot for your contributions to kernels and discussions during the competition. I think they helped a lot to all of us.</p>\n\n<p>One thing from your solution I didn't get well... Can you be more specific on </p>\n\n<p>&gt; seeing item_seq_number as a sequence.</p>\n\n<p>please?</p>",
      "rawMarkdown": "Congrats on the 25th place! Also I want to thank you a lot for your contributions to kernels and discussions during the competition. I think they helped a lot to all of us.\n\nOne thing from your solution I didn't get well... Can you be more specific on \n\n&gt; seeing item_seq_number as a sequence.\n\nplease?",
      "votes": null
    },
    {
      "id": "349530",
      "postDate": "06/28/2018 08:31:52",
      "content": "<p>If you sort the data for each user_id by item_seq_number you get a sequence for the price. E.g.</p>\n\n<p>user_id =5, item_seq_number = 1, price = 10 </p>\n\n<p>user_id =5, item_seq_number = 5, price = 20 </p>\n\n<p>user_id =5, item_seq_number = 6, price = 100</p>\n\n<p>would result in the sequence [10,20,100]. We then derived delta features between price of previous item and actual item (we actually used previous 2 and following 2, but for illustration I only show delta +- 1):</p>\n\n<p>user_id =5, item_seq_number = 1, price = 10, delta_-1 = N/A, delta_1 = 10</p>\n\n<p>user_id =5, item_seq_number = 5, price = 20, delta_-1 = -10, delta_1 = 80 </p>\n\n<p>user_id =5, item_seq_number = 6, price = 100, delta_-80 = N/A, delta_1 = N/A</p>\n\n<p>If we had this idea earlier I would have tried to actually use RNN on the whole sequence within the NN architecture :)</p>\n\n<p>Thank you for valuing my contribution. That means a lot and encourages to share my ideas in further competitions. </p>",
      "rawMarkdown": "If you sort the data for each user_id by item_seq_number you get a sequence for the price. E.g.\n\nuser_id =5, item_seq_number = 1, price = 10 \n\nuser_id =5, item_seq_number = 5, price = 20 \n\nuser_id =5, item_seq_number = 6, price = 100\n\nwould result in the sequence [10,20,100]. We then derived delta features between price of previous item and actual item (we actually used previous 2 and following 2, but for illustration I only show delta +- 1):\n\nuser_id =5, item_seq_number = 1, price = 10, delta_-1 = N/A, delta_1 = 10\n\nuser_id =5, item_seq_number = 5, price = 20, delta_-1 = -10, delta_1 = 80 \n\nuser_id =5, item_seq_number = 6, price = 100, delta_-80 = N/A, delta_1 = N/A\n\nIf we had this idea earlier I would have tried to actually use RNN on the whole sequence within the NN architecture :)\n\nThank you for valuing my contribution. That means a lot and encourages to share my ideas in further competitions.",
      "votes": null
    },
    {
      "id": "349532",
      "postDate": "06/28/2018 08:34:53",
      "content": "<p>Thanks. Now it's clear :)</p>",
      "rawMarkdown": "Thanks. Now it's clear :)",
      "votes": null
    },
    {
      "id": "349533",
      "postDate": "06/28/2018 08:35:24",
      "content": "<p>brilliant. </p>",
      "rawMarkdown": "brilliant.",
      "votes": null
    },
    {
      "id": "349555",
      "postDate": "06/28/2018 09:18:00",
      "content": "<p>Awesome,I like your code</p>",
      "rawMarkdown": "Awesome,I like your code",
      "votes": null
    },
    {
      "id": "349556",
      "postDate": "06/28/2018 09:18:35",
      "content": "<p>How long to train this model</p>",
      "rawMarkdown": "How long to train this model",
      "votes": null
    },
    {
      "id": "349597",
      "postDate": "06/28/2018 10:37:45",
      "content": "<p>Thank you for all your commentary and contributions! Congratulations on your results!</p>",
      "rawMarkdown": "Thank you for all your commentary and contributions! Congratulations on your results!",
      "votes": null
    },
    {
      "id": "349618",
      "postDate": "06/28/2018 11:36:17",
      "content": "<p>Consgrats and thanks very much. During this competition, your kernel helped me a lot. </p>",
      "rawMarkdown": "Consgrats and thanks very much. During this competition, your kernel helped me a lot.",
      "votes": null
    },
    {
      "id": "349632",
      "postDate": "06/28/2018 12:09:10",
      "content": "<p>Thank you for sharing. Learnt a lot of NN from your awesome kernels!</p>",
      "rawMarkdown": "Thank you for sharing. Learnt a lot of NN from your awesome kernels!",
      "votes": null
    },
    {
      "id": "349685",
      "postDate": "06/28/2018 13:32:19",
      "content": "<p>Congatulations ! What is in prense() function ?</p>",
      "rawMarkdown": "Congatulations ! What is in prense() function ?",
      "votes": null
    },
    {
      "id": "349730",
      "postDate": "06/28/2018 14:47:14",
      "content": "<p>Congrats ! And thanks for sharing your NN !</p>",
      "rawMarkdown": "Congrats ! And thanks for sharing your NN !",
      "votes": null
    },
    {
      "id": "349742",
      "postDate": "06/28/2018 15:07:25",
      "content": "<p>Congrats! I really appreciate your kernels at Kaggle. That really helped me build up my own first NN model. I have a question on your NN solution:</p>\n\n<p>what is prense and what is that for? Is it like Dense layer?</p>\n\n<p>Thank you again!</p>",
      "rawMarkdown": "Congrats! I really appreciate your kernels at Kaggle. That really helped me build up my own first NN model. I have a question on your NN solution:\n\nwhat is prense and what is that for? Is it like Dense layer?\n\nThank you again!",
      "votes": null
    },
    {
      "id": "349778",
      "postDate": "06/28/2018 16:33:16",
      "content": "<p>Sorry that I did not put it in code. prense() is just a shortcut I used.</p>\n\n<pre><code>def prense(x, units):\n    x = Dense(units)(x)\n    x = PReLU()(x)\nreturn x\n</code></pre>",
      "rawMarkdown": "Sorry that I did not put it in code. prense() is just a shortcut I used.\n\n    def prense(x, units):\n        x = Dense(units)(x)\n        x = PReLU()(x)\n    return x",
      "votes": null
    },
    {
      "id": "349782",
      "postDate": "06/28/2018 16:34:48",
      "content": "<p>10 fold cv around 6h on Ti 1070. Average number of training epoch before early stopping is around 6</p>",
      "rawMarkdown": "10 fold cv around 6h on Ti 1070. Average number of training epoch before early stopping is around 6",
      "votes": null
    },
    {
      "id": "349960",
      "postDate": "06/29/2018 00:24:44",
      "content": "<p>Congratulations @Dieter and team for a very strong finish. Thanks for sharing your solution overview and your interesting NN architecture. What is your HW setup and how long did it take to run you NN model?</p>\n\n<p>Thanks so much @Dieter for all your contributions in kernels as well as discussions. This is what makes Kaggle unique.</p>",
      "rawMarkdown": "Congratulations @Dieter and team for a very strong finish. Thanks for sharing your solution overview and your interesting NN architecture. What is your HW setup and how long did it take to run you NN model?\n\nThanks so much @Dieter for all your contributions in kernels as well as discussions. This is what makes Kaggle unique.",
      "votes": null
    },
    {
      "id": "403755",
      "postDate": "10/14/2018 14:39:43",
      "content": "<p>I am refactoring into a GitHub repository. Since I did a picture of the architecture, I'd like to share here</p>\n\n<p><img src=\"http://github.com/khumbuai/kaggle_avito_demand/raw/master/avito_NN_overview.jpeg\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "I am refactoring into a GitHub repository. Since I did a picture of the architecture, I'd like to share here\n\n![enter image description here][1]\n\n\n  [1]: http://github.com/khumbuai/kaggle_avito_demand/raw/master/avito_NN_overview.jpeg",
      "votes": null
    },
    {
      "id": "403924",
      "postDate": "10/14/2018 22:58:39",
      "content": "<p>Thanks for the update @Dieter.</p>",
      "rawMarkdown": "Thanks for the update @Dieter.",
      "votes": null
    },
    {
      "id": "405150",
      "postDate": "10/17/2018 01:01:04",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> This is a beautiful diagram!! I presume this is done with <code>graphviz</code>? Can you share the <code>dot</code> file? I'd love to learn how to align things like this.</p>",
      "rawMarkdown": "christofhenkel This is a beautiful diagram!! I presume this is done with `graphviz`? Can you share the `dot` file? I'd love to learn how to align things like this.",
      "votes": null
    },
    {
      "id": "405215",
      "postDate": "10/17/2018 04:24:10",
      "content": "<p>I used keynote. The file is in the git repo: <a href=\"https://github.com/khumbuai/kaggle_avito_demand\">https://github.com/khumbuai/kaggle_avito_demand</a></p>",
      "rawMarkdown": "I used keynote. The file is in the git repo: [https://github.com/khumbuai/kaggle_avito_demand][1]\n\n\n  [1]: https://github.com/khumbuai/kaggle_avito_demand",
      "votes": null
    },
    {
      "id": "405242",
      "postDate": "10/17/2018 05:23:28",
      "content": "<p>Newbie question, where you got inspiration/ide ? i am still strugle to create Good NN architecture. maybe i need to buy more powerfull GPU to speedup the experiment</p>\n\n<p>Thank you very much for the code, </p>",
      "rawMarkdown": "Newbie question, where you got inspiration/ide ? i am still strugle to create Good NN architecture. maybe i need to buy more powerfull GPU to speedup the experiment\n\nThank you very much for the code,",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 349518,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "06/28/2018 08:12:29",
      "content": "<p>Congrats on the 25th place! Also I want to thank you a lot for your contributions to kernels and discussions during the competition. I think they helped a lot to all of us.</p>\n\n<p>One thing from your solution I didn't get well... Can you be more specific on </p>\n\n<p>&gt; seeing item_seq_number as a sequence.</p>\n\n<p>please?</p>",
      "votes": null,
      "replies": [
        {
          "id": 349530,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/28/2018 08:31:52",
          "content": "<p>If you sort the data for each user_id by item_seq_number you get a sequence for the price. E.g.</p>\n\n<p>user_id =5, item_seq_number = 1, price = 10 </p>\n\n<p>user_id =5, item_seq_number = 5, price = 20 </p>\n\n<p>user_id =5, item_seq_number = 6, price = 100</p>\n\n<p>would result in the sequence [10,20,100]. We then derived delta features between price of previous item and actual item (we actually used previous 2 and following 2, but for illustration I only show delta +- 1):</p>\n\n<p>user_id =5, item_seq_number = 1, price = 10, delta_-1 = N/A, delta_1 = 10</p>\n\n<p>user_id =5, item_seq_number = 5, price = 20, delta_-1 = -10, delta_1 = 80 </p>\n\n<p>user_id =5, item_seq_number = 6, price = 100, delta_-80 = N/A, delta_1 = N/A</p>\n\n<p>If we had this idea earlier I would have tried to actually use RNN on the whole sequence within the NN architecture :)</p>\n\n<p>Thank you for valuing my contribution. That means a lot and encourages to share my ideas in further competitions. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349532,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "06/28/2018 08:34:53",
          "content": "<p>Thanks. Now it's clear :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349533,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "06/28/2018 08:35:24",
          "content": "<p>brilliant. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349555,
      "author_name": "liuhdsgoal",
      "author_url": "",
      "post_date": "06/28/2018 09:18:00",
      "content": "<p>Awesome,I like your code</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349556,
      "author_name": "liuhdsgoal",
      "author_url": "",
      "post_date": "06/28/2018 09:18:35",
      "content": "<p>How long to train this model</p>",
      "votes": null,
      "replies": [
        {
          "id": 349782,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/28/2018 16:34:48",
          "content": "<p>10 fold cv around 6h on Ti 1070. Average number of training epoch before early stopping is around 6</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349597,
      "author_name": "rakibilly",
      "author_url": "",
      "post_date": "06/28/2018 10:37:45",
      "content": "<p>Thank you for all your commentary and contributions! Congratulations on your results!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349618,
      "author_name": "poohtls",
      "author_url": "",
      "post_date": "06/28/2018 11:36:17",
      "content": "<p>Consgrats and thanks very much. During this competition, your kernel helped me a lot. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349632,
      "author_name": "wangking",
      "author_url": "",
      "post_date": "06/28/2018 12:09:10",
      "content": "<p>Thank you for sharing. Learnt a lot of NN from your awesome kernels!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349685,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "06/28/2018 13:32:19",
      "content": "<p>Congatulations ! What is in prense() function ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 349778,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/28/2018 16:33:16",
          "content": "<p>Sorry that I did not put it in code. prense() is just a shortcut I used.</p>\n\n<pre><code>def prense(x, units):\n    x = Dense(units)(x)\n    x = PReLU()(x)\nreturn x\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349730,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "06/28/2018 14:47:14",
      "content": "<p>Congrats ! And thanks for sharing your NN !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349742,
      "author_name": "haotianwu",
      "author_url": "",
      "post_date": "06/28/2018 15:07:25",
      "content": "<p>Congrats! I really appreciate your kernels at Kaggle. That really helped me build up my own first NN model. I have a question on your NN solution:</p>\n\n<p>what is prense and what is that for? Is it like Dense layer?</p>\n\n<p>Thank you again!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349960,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/29/2018 00:24:44",
      "content": "<p>Congratulations @Dieter and team for a very strong finish. Thanks for sharing your solution overview and your interesting NN architecture. What is your HW setup and how long did it take to run you NN model?</p>\n\n<p>Thanks so much @Dieter for all your contributions in kernels as well as discussions. This is what makes Kaggle unique.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 403755,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "10/14/2018 14:39:43",
      "content": "<p>I am refactoring into a GitHub repository. Since I did a picture of the architecture, I'd like to share here</p>\n\n<p><img src=\"http://github.com/khumbuai/kaggle_avito_demand/raw/master/avito_NN_overview.jpeg\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 403924,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "10/14/2018 22:58:39",
          "content": "<p>Thanks for the update @Dieter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 405150,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "10/17/2018 01:01:04",
          "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> This is a beautiful diagram!! I presume this is done with <code>graphviz</code>? Can you share the <code>dot</code> file? I'd love to learn how to align things like this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 405215,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "10/17/2018 04:24:10",
          "content": "<p>I used keynote. The file is in the git repo: <a href=\"https://github.com/khumbuai/kaggle_avito_demand\">https://github.com/khumbuai/kaggle_avito_demand</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 405242,
          "author_name": "hendraherviawan",
          "author_url": "",
          "post_date": "10/17/2018 05:23:28",
          "content": "<p>Newbie question, where you got inspiration/ide ? i am still strugle to create Good NN architecture. maybe i need to buy more powerfull GPU to speedup the experiment</p>\n\n<p>Thank you very much for the code, </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "349481": "First of all, thanks to kaggle as well as avito for this challenging competition. \n\nLet me quickly give an overview of my part of our solution.\n\n**Summary**\n\nUntil last day we used a single layer stacking of different NN and LGB models. For the final submission we also trained our stack with xgb and blended the results with our LGB stacker 1:1.\nThe best NN scored 0.2194 the best LGB 0.2196. Some notes on what helped and did not help.\n\n - I spent a lot of time trying different factorization methods like LDA, libfm etc. but none improved my score\n - What helped a lot where features derived in public kernels like city population from wikipedia etc. \n - I made a test on pseudo-labelling and found out it helped, but we had no submission and time left to incorporate\n - Features from pretrained image models (DenseNet, ILGNet, Ava) \n - w2v embeddings trained from train active as described in one of my kernels\n - a higher k-folder significantly improved the score, even more when bagging same models with different seeds\n \n**Feature engineering**\n\nWe tried/ used most of the features from public kernels, however two interesting features we found were:\n\n - number of duplicates per user. While Benjamin drops duplicates in his kernel we considered the number of duplicates important and this helped\n - my teammate had a great idea seeing item_seq_number as a sequence. (sometimes the obvious is hard to see) and came up with the idea to use the price difference of consecutive items as feature, which improved LB a lot \n\n\n\n**NN structcture**\n\nI used the following NN structure, where I did TFIDF on character level on param_3  and w2v embeddings on text. I fed all numerical features we had (including a ridge model trained with price as target)\n\n\n    def build_model():\n    sparse_params = Input(shape=[X_train['sparse_params'].shape[1]], dtype='float32', sparse=True, name='sparse_params')\n\n    categorical_inputs = []\n    for cat in cat_cols:\n        categorical_inputs.append(Input(shape=[1], name=cat))\n\n    categorical_embeddings = []\n    for i, cat in enumerate(cat_cols):\n        categorical_embeddings.append(\n            Embedding(embed_sizes[i], 10, embeddings_regularizer=l2(0.00001))(categorical_inputs[i]))\n\n    categorical_logits = Concatenate()([Flatten()(cat_emb) for cat_emb in categorical_embeddings])\n    categorical_logits = prense(categorical_logits, 256)\n    categorical_logits = prense(categorical_logits, 128)\n\n    numerical_inputs = Input(shape=[len(num_cols)], name='numerical')\n\n    numerical_logits = numerical_inputs\n    numerical_logits = BatchNormalization()(numerical_logits)\n    numerical_logits = prense(numerical_logits, 256)\n    numerical_logits = prense(numerical_logits, 128)\n\n    params_logits = prense(sparse_params, 64)\n    params_logits = prense(params_logits, 32)\n\n    desc_inp = Input(shape=[max_len_desc], name='desc')\n    title_inp = Input(shape=[max_len_title], name='title')\n    embedding = Embedding(nb_words, embed_size, weights=[embedding_matrix], trainable=False)  # nb_words\n    emb_desc = embedding(desc_inp)\n    emb_title = embedding(title_inp)\n    emb_text = Concatenate(axis=1)([emb_desc,emb_title])\n\n    text_logits = SpatialDropout1D(0.2)(emb_text)\n    text_logits = Bidirectional(CuDNNLSTM(128, return_sequences=True))(text_logits)\n    text_logits = Conv1D(64, kernel_size=3, padding=\"valid\", kernel_initializer=\"glorot_uniform\")(text_logits)\n    avg_pool = GlobalAveragePooling1D()(text_logits)\n    max_pool = GlobalMaxPool1D()(text_logits)\n    text_logits = Concatenate()([avg_pool, max_pool])\n    x = Dropout(0.2)(text_logits)\n    x = Concatenate()([categorical_logits, text_logits])\n    x = BatchNormalization()(x)\n    x = Concatenate()([x, params_logits, numerical_logits])\n    x = Dense(512, kernel_initializer=he_uniform(seed=0))(x)\n    x = PReLU()(x)\n    x = Dense(256, kernel_initializer=he_uniform(seed=0))(x)\n    x = PReLU()(x)\n    x = Dense(128, kernel_initializer=he_uniform(seed=0))(x)\n    x = PReLU()(x)\n    x = LayerNorm1D()(x)\n    out = Dense(1, activation='sigmoid')(x)\n\n    model = Model(inputs=[desc_inp] + [title_inp] + [sparse_params] + categorical_inputs + [numerical_inputs],\n                  outputs=out)\n\n    model.compile(optimizer=Adam(lr=0.0005, clipnorm=0.5), loss='mean_squared_error',\n                  metrics=[root_mean_squared_error])\n    return model",
    "349518": "Congrats on the 25th place! Also I want to thank you a lot for your contributions to kernels and discussions during the competition. I think they helped a lot to all of us.\n\nOne thing from your solution I didn't get well... Can you be more specific on \n\n&gt; seeing item_seq_number as a sequence.\n\nplease?",
    "349530": "If you sort the data for each user_id by item_seq_number you get a sequence for the price. E.g.\n\nuser_id =5, item_seq_number = 1, price = 10 \n\nuser_id =5, item_seq_number = 5, price = 20 \n\nuser_id =5, item_seq_number = 6, price = 100\n\nwould result in the sequence [10,20,100]. We then derived delta features between price of previous item and actual item (we actually used previous 2 and following 2, but for illustration I only show delta +- 1):\n\nuser_id =5, item_seq_number = 1, price = 10, delta_-1 = N/A, delta_1 = 10\n\nuser_id =5, item_seq_number = 5, price = 20, delta_-1 = -10, delta_1 = 80 \n\nuser_id =5, item_seq_number = 6, price = 100, delta_-80 = N/A, delta_1 = N/A\n\nIf we had this idea earlier I would have tried to actually use RNN on the whole sequence within the NN architecture :)\n\nThank you for valuing my contribution. That means a lot and encourages to share my ideas in further competitions.",
    "349532": "Thanks. Now it's clear :)",
    "349533": "brilliant.",
    "349555": "Awesome,I like your code",
    "349556": "How long to train this model",
    "349597": "Thank you for all your commentary and contributions! Congratulations on your results!",
    "349618": "Consgrats and thanks very much. During this competition, your kernel helped me a lot.",
    "349632": "Thank you for sharing. Learnt a lot of NN from your awesome kernels!",
    "349685": "Congatulations ! What is in prense() function ?",
    "349730": "Congrats ! And thanks for sharing your NN !",
    "349742": "Congrats! I really appreciate your kernels at Kaggle. That really helped me build up my own first NN model. I have a question on your NN solution:\n\nwhat is prense and what is that for? Is it like Dense layer?\n\nThank you again!",
    "349778": "Sorry that I did not put it in code. prense() is just a shortcut I used.\n\n    def prense(x, units):\n        x = Dense(units)(x)\n        x = PReLU()(x)\n    return x",
    "349782": "10 fold cv around 6h on Ti 1070. Average number of training epoch before early stopping is around 6",
    "349960": "Congratulations @Dieter and team for a very strong finish. Thanks for sharing your solution overview and your interesting NN architecture. What is your HW setup and how long did it take to run you NN model?\n\nThanks so much @Dieter for all your contributions in kernels as well as discussions. This is what makes Kaggle unique.",
    "403755": "I am refactoring into a GitHub repository. Since I did a picture of the architecture, I'd like to share here\n\n![enter image description here][1]\n\n\n  [1]: http://github.com/khumbuai/kaggle_avito_demand/raw/master/avito_NN_overview.jpeg",
    "403924": "Thanks for the update @Dieter.",
    "405150": "christofhenkel This is a beautiful diagram!! I presume this is done with `graphviz`? Can you share the `dot` file? I'd love to learn how to align things like this.",
    "405215": "I used keynote. The file is in the git repo: [https://github.com/khumbuai/kaggle_avito_demand][1]\n\n\n  [1]: https://github.com/khumbuai/kaggle_avito_demand",
    "405242": "Newbie question, where you got inspiration/ide ? i am still strugle to create Good NN architecture. maybe i need to buy more powerfull GPU to speedup the experiment\n\nThank you very much for the code,"
  },
  "source": "meta"
}