{
  "id": 57085,
  "title": "Best Neural Network Model and structure",
  "url": "/competitions/avito-demand-prediction/discussion/57085",
  "author_name": "Strideradu",
  "post_date": "2018-05-18T23:42:49.197000",
  "votes": 25,
  "comment_count": 61,
  "views": 0,
  "content": "<p>@Dieter have a very good discussion on the best single model. However, most results are for the lgbm model. Since I start working on the neural network model, I would like to discuss the best model for neural network and also the network structure ( if you like to share).</p>\n\n<p>Right now my 2 layer GRU model with global max pooling for title and description is about 0.23, I just copy the model I used for the toxic comment challenge.</p>\n\n<p>Haven't try an NN model just like a classifier</p>",
  "messages": [
    {
      "id": 330488,
      "postDate": "2018-05-18T23:42:49.197Z",
      "content": "<p>@Dieter have a very good discussion on the best single model. However, most results are for the lgbm model. Since I start working on the neural network model, I would like to discuss the best model for neural network and also the network structure ( if you like to share).</p>\n\n<p>Right now my 2 layer GRU model with global max pooling for title and description is about 0.23, I just copy the model I used for the toxic comment challenge.</p>\n\n<p>Haven't try an NN model just like a classifier</p>",
      "rawMarkdown": "@Dieter have a very good discussion on the best single model. However, most results are for the lgbm model. Since I start working on the neural network model, I would like to discuss the best model for neural network and also the network structure ( if you like to share).\n\nRight now my 2 layer GRU model with global max pooling for title and description is about 0.23, I just copy the model I used for the toxic comment challenge.\n\nHaven't try an NN model just like a classifier",
      "votes": 25
    },
    {
      "id": 330989,
      "postDate": "2018-05-20T05:47:26.070Z",
      "content": "<p>My single NN model is now scoring 0.2217 on LB</p>\n\n<p>It's RNN and Description and title processed with self_trained embeddings...</p>",
      "rawMarkdown": "My single NN model is now scoring 0.2217 on LB\n\nIt's RNN and Description and title processed with self_trained embeddings...\n",
      "votes": 13,
      "replies": [
        {
          "id": 331101,
          "postDate": "2018-05-20T11:11:19.130Z",
          "content": "<p>Nice result! Is it only text features or do you include structured features (price, categorical features, ...) as well?</p>",
          "rawMarkdown": "Nice result! Is it only text features or do you include structured features (price, categorical features, ...) as well?"
        },
        {
          "id": 331105,
          "postDate": "2018-05-20T11:27:12.023Z",
          "content": "<p>I use almost all the features (unless ids and image)....</p>\n\n<p>Item_seq_number was very noisy at the beginning but I finally manged to include  it.</p>",
          "rawMarkdown": "I use almost all the features (unless ids and image)....\n\nItem_seq_number was very noisy at the beginning but I finally manged to include  it.",
          "votes": 4
        },
        {
          "id": 331111,
          "postDate": "2018-05-20T11:39:15Z",
          "content": "<p>Thanks for the hint. I'm usually not a big fan of neural nets for structured data, but I want to try it for this competition. Seems like it's actually worth a try.</p>",
          "rawMarkdown": "Thanks for the hint. I'm usually not a big fan of neural nets for structured data, but I want to try it for this competition. Seems like it's actually worth a try."
        },
        {
          "id": 331190,
          "postDate": "2018-05-20T15:26:57.313Z",
          "content": "<blockquote>\n  <p>I'm usually not a big fan of neural nets for structured data</p>\n</blockquote>\n\n<p>Why?</p>",
          "rawMarkdown": "&gt; I'm usually not a big fan of neural nets for structured data\n\nWhy?",
          "votes": 1
        },
        {
          "id": 331332,
          "postDate": "2018-05-21T02:55:38.113Z",
          "content": "<p>Serigne, I will start this competition with your instructions. Thanks.</p>",
          "rawMarkdown": "Serigne, I will start this competition with your instructions. Thanks.\n",
          "votes": 2
        },
        {
          "id": 331352,
          "postDate": "2018-05-21T04:33:04.757Z",
          "content": "<p>@ Maximilian </p>\n\n<p>Any reason why you would not use NN for structured data? And what would be your definition for structured data in this context?</p>",
          "rawMarkdown": "@ Maximilian \n\nAny reason why you would not use NN for structured data? And what would be your definition for structured data in this context?"
        },
        {
          "id": 331588,
          "postDate": "2018-05-21T14:28:56.880Z",
          "content": "<p>Neural networks being relatively less powerful on structured data is a pretty common heuristic, with good reason. They impose the functional form assumption that the target can be well-modeled by layering of hierarchical feature representations, so the data should support that hierarchical representation for the model to be ideal. That's certainly the case with data like images or text, where processing tasks involve building up higher order concepts from simpler ones - the best visual example is how image classification CNN filters learn generic edge and shape patterns in early layers and build out to more specific object representations in late layers.</p>\n\n<p>It's harder to see how hierarchical layering would be as naturally useful for most tabular tasks. For example, something like \"price\" is already a very well-described feature that doesn't need to be intricately manipulated to have a lot of predictive value, whereas individual pixels in an image need a lot of work to be leveraged as signals. This is why you'll often see that shallow neural networks outperform deeper ones on tabular data. And to capture explicit, direct feature interactions, it's more natural to use tree based methods (especially gradient-boosted trees) that allow for repeated interactive partitioning that derives from meaningful features as is instead of from more complex representations that need to be learned. </p>\n\n<p>So I usually expect to see gradient boosting outperform neural networks on tabular data. That doesn't mean they can't be competitive, especially with careful preprocessing and tuning. </p>\n\n<p>But this competition isn't tabular data, it's mixed. And in addition to their natural application to text and image data, one of the great benefits of neural networks is that their architecture is flexible - it's possible to incorporate as many different raw data types as you want into a single model. This obviously comes with its own challenges, but it is one possible solution to handling this dataset.  </p>",
          "rawMarkdown": "Neural networks being relatively less powerful on structured data is a pretty common heuristic, with good reason. They impose the functional form assumption that the target can be well-modeled by layering of hierarchical feature representations, so the data should support that hierarchical representation for the model to be ideal. That's certainly the case with data like images or text, where processing tasks involve building up higher order concepts from simpler ones - the best visual example is how image classification CNN filters learn generic edge and shape patterns in early layers and build out to more specific object representations in late layers.\n\nIt's harder to see how hierarchical layering would be as naturally useful for most tabular tasks. For example, something like \"price\" is already a very well-described feature that doesn't need to be intricately manipulated to have a lot of predictive value, whereas individual pixels in an image need a lot of work to be leveraged as signals. This is why you'll often see that shallow neural networks outperform deeper ones on tabular data. And to capture explicit, direct feature interactions, it's more natural to use tree based methods (especially gradient-boosted trees) that allow for repeated interactive partitioning that derives from meaningful features as is instead of from more complex representations that need to be learned. \n\nSo I usually expect to see gradient boosting outperform neural networks on tabular data. That doesn't mean they can't be competitive, especially with careful preprocessing and tuning. \n\nBut this competition isn't tabular data, it's mixed. And in addition to their natural application to text and image data, one of the great benefits of neural networks is that their architecture is flexible - it's possible to incorporate as many different raw data types as you want into a single model. This obviously comes with its own challenges, but it is one possible solution to handling this dataset.  ",
          "votes": 33
        },
        {
          "id": 331626,
          "postDate": "2018-05-21T15:47:04.230Z",
          "content": "<p>Just what Joe Eddy said. If you look at all past Kaggle competitions not involving text or images, you will almost always find a GBM model as the best single model. There are some exceptions from this rule, but I think if you had one shot, you should always go for a GBM model in these cases. One counter example is the competition mentioned in this article: <a href=\"https://towardsdatascience.com/structured-deep-learning-b8ca4138b848\">https://towardsdatascience.com/structured-deep-learning-b8ca4138b848</a></p>\n\n<p>But as you can see, people had to come up with new ideas to actually build good NN models for structured date.</p>\n\n<p>P.S.: By \"structured\" I mean tabular data with categorical and numerical variables.</p>",
          "rawMarkdown": "Just what Joe Eddy said. If you look at all past Kaggle competitions not involving text or images, you will almost always find a GBM model as the best single model. There are some exceptions from this rule, but I think if you had one shot, you should always go for a GBM model in these cases. One counter example is the competition mentioned in this article: https://towardsdatascience.com/structured-deep-learning-b8ca4138b848\n\nBut as you can see, people had to come up with new ideas to actually build good NN models for structured date.\n\nP.S.: By \"structured\" I mean tabular data with categorical and numerical variables.",
          "votes": 3
        },
        {
          "id": 331859,
          "postDate": "2018-05-22T01:53:12.217Z",
          "content": "<p>Another very cool exception that I think is worth mentioning is Porto Seguro - Michael Jahrer won with a solution that was primarily based around <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">denoising autoencoders</a>. It seems the reason this worked so well was that since the raw representations in that dataset were so messy (many missing values in some of the most important features), learned representations offered significant predictive gain. </p>",
          "rawMarkdown": "Another very cool exception that I think is worth mentioning is Porto Seguro - Michael Jahrer won with a solution that was primarily based around [denoising autoencoders][1]. It seems the reason this worked so well was that since the raw representations in that dataset were so messy (many missing values in some of the most important features), learned representations offered significant predictive gain. \n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629",
          "votes": 2
        },
        {
          "id": 334803,
          "postDate": "2018-05-28T12:47:54.410Z",
          "content": "<p>Interesting. That would go in the direction proposed in another discussion (by Dieter ?) to have price and image models in order to have cleaner data before training.</p>",
          "rawMarkdown": "Interesting. That would go in the direction proposed in another discussion (by Dieter ?) to have price and image models in order to have cleaner data before training."
        },
        {
          "id": 335317,
          "postDate": "2018-05-29T15:16:13.193Z",
          "content": "<p>Hi Serigine, I am wondering if you use both train_active and test_active to train your model? And how many epoches? It seens one epoch took long time to finish</p>",
          "rawMarkdown": "Hi Serigine, I am wondering if you use both train_active and test_active to train your model? And how many epoches? It seens one epoch took long time to finish"
        },
        {
          "id": 336964,
          "postDate": "2018-06-01T16:19:04.097Z",
          "content": "<p>Hi Strideradu </p>\n\n<p>Sorry for the late reply ...</p>\n\n<p>No I don't use them yet for training. </p>\n\n<p>As for the number of epoches, I don't specify it. I just use Callback with early_stopping criterion to let the model decide the number of epoches (to avoid overfitting) </p>",
          "rawMarkdown": "Hi Strideradu \n\nSorry for the late reply ...\n\nNo I don't use them yet for training. \n\nAs for the number of epoches, I don't specify it. I just use Callback with early_stopping criterion to let the model decide the number of epoches (to avoid overfitting) "
        },
        {
          "id": 336970,
          "postDate": "2018-06-01T16:30:43.200Z",
          "content": "<p>Hi Serigine,</p>\n\n<p>Thanks for your answer. I am actually asking how many epochs you were using to train you embeddings, it seems no good way to decide the epochs?</p>",
          "rawMarkdown": "Hi Serigine,\n\nThanks for your answer. I am actually asking how many epochs you were using to train you embeddings, it seems no good way to decide the epochs?"
        },
        {
          "id": 336987,
          "postDate": "2018-06-01T17:29:47.880Z",
          "content": "<p>Ah ok ! I read \"train your model\"  and thought you're  talking about the RNN framework. </p>\n\n<p>I use 3 epochs for the embeddings.</p>",
          "rawMarkdown": "Ah ok ! I read \"train your model\"  and thought you're  talking about the RNN framework. \n\nI use 3 epochs for the embeddings."
        },
        {
          "id": 337099,
          "postDate": "2018-06-01T23:30:06.977Z",
          "content": "<p>@Serigne, are you able to get your NN model better than LB=0.2217? My best lgb model is at 0.2217 LB as well right now and am going to start modeling an NN solution. I am just wondering what the best NN score is before I start spending money on my cloud account.</p>",
          "rawMarkdown": "@Serigne, are you able to get your NN model better than LB=0.2217? My best lgb model is at 0.2217 LB as well right now and am going to start modeling an NN solution. I am just wondering what the best NN score is before I start spending money on my cloud account."
        },
        {
          "id": 337382,
          "postDate": "2018-06-02T17:48:31.743Z",
          "content": "<p>Hi ...I haven't touch it since it achieved this score.  ...I am working now with another model. </p>\n\n<p>However it can surely be improved with active and image features. </p>",
          "rawMarkdown": "Hi ...I haven't touch it since it achieved this score.  ...I am working now with another model. \n\nHowever it can surely be improved with active and image features. "
        },
        {
          "id": 337505,
          "postDate": "2018-06-03T02:55:31.213Z",
          "content": "<p>Thanks @Serigne. I forgot to ask, your NN model ... 0.2217 LB score is with n-Folds right? Perhaps 5-Fold CV?</p>\n\n<p>I have my 1st attempt NN at LB = 0.2245 and the best lgb mentioned is for a single run. Hence, I am working on 5-Fold CV for my lgbm and other models before going back to improving my NN model. Hopefully, I can make a good jump with the 5-Fold :-)</p>",
          "rawMarkdown": "Thanks @Serigne. I forgot to ask, your NN model ... 0.2217 LB score is with n-Folds right? Perhaps 5-Fold CV?\n\nI have my 1st attempt NN at LB = 0.2245 and the best lgb mentioned is for a single run. Hence, I am working on 5-Fold CV for my lgbm and other models before going back to improving my NN model. Hopefully, I can make a good jump with the 5-Fold :-)"
        },
        {
          "id": 337871,
          "postDate": "2018-06-04T00:01:00.330Z",
          "content": "<blockquote>\n  <p><strong>YaGana Sheriff-Hussaini wrote</strong></p>\n  \n  <blockquote>\n    <p>Thanks @Serigne. I forgot to ask, your NN model ... 0.2217 LB score is with n-Folds right? Perhaps 5-Fold CV?</p>\n  </blockquote>\n  \n  <p>I have my 1st attempt NN at LB = 0.2245 and the best lgb mentioned is for a single run. Hence, I am working on 5-Fold CV for my lgbm and other models before going back to improving my NN model. Hopefully, I can make a good jump with the 5-Fold :-)</p>\n</blockquote>\n\n<p>I am using 5 fold bagging predictions but I don't think it can make huge difference compared to single run , given that train and test seem to have pretty much the same distribution.</p>\n\n<p>I use it mainly to have coherent validation strategy for all my experiments and also to have OOF predictions for further blending/stacking.</p>",
          "rawMarkdown": "\n&gt; **YaGana Sheriff-Hussaini wrote**\n&gt; \n&gt; &gt; Thanks @Serigne. I forgot to ask, your NN model ... 0.2217 LB score is with n-Folds right? Perhaps 5-Fold CV?\n&gt; \n&gt; I have my 1st attempt NN at LB = 0.2245 and the best lgb mentioned is for a single run. Hence, I am working on 5-Fold CV for my lgbm and other models before going back to improving my NN model. Hopefully, I can make a good jump with the 5-Fold :-)\n\nI am using 5 fold bagging predictions but I don't think it can make huge difference compared to single run , given that train and test seem to have pretty much the same distribution.\n\nI use it mainly to have coherent validation strategy for all my experiments and also to have OOF predictions for further blending/stacking.",
          "votes": 2
        },
        {
          "id": 337947,
          "postDate": "2018-06-04T05:56:16.253Z",
          "content": "<p>Thanks @Serigne.  </p>\n\n<p>I have finally got my 5-Fold predictions working. My best model is now scoring 0.2215 on LB. Although hoping to get improvement with 5-fold, I am also doing it to get the OOF predictions for blending/bagging. Fingers crossed :-)</p>",
          "rawMarkdown": "Thanks @Serigne.  \n\nI have finally got my 5-Fold predictions working. My best model is now scoring 0.2215 on LB. Although hoping to get improvement with 5-fold, I am also doing it to get the OOF predictions for blending/bagging. Fingers crossed :-)"
        },
        {
          "id": 338870,
          "postDate": "2018-06-05T22:01:42.410Z",
          "content": "<p>@Serigne may I know how you deal with item_seq_number?  I just simply removed it to reduce noise and this did help a bit.</p>",
          "rawMarkdown": "@Serigne may I know how you deal with item_seq_number?  I just simply removed it to reduce noise and this did help a bit."
        }
      ]
    },
    {
      "id": 342809,
      "postDate": "2018-06-14T06:41:17.957Z",
      "content": "<p>Below my current best-ish model (0.2210 in 1 run local validation using 20% of samples, I have not submitted yet). </p>\n\n<p>I use pre-trained 300d vectors for language and i only take the 50,000 most common tokens which get fine-tuned (<code>emb_desc</code> layer). </p>\n\n<p>There's embeddings for a lot of other thing and <code>inp_price</code> is just  log(price+eps). I am not using image-features below (but I have one which uses it extracting ResNet50 last layer and finetuning the whole thing, but takes a while to train so I'll leave it until the end).</p>\n\n<p>Things in my list:\n -  Add features like items per user\n - Make sense of item_seq\n - Automatic feature engineering a la DAE\n - Conv after RNN (GRU)</p>\n\n<p>Ideas welcome.</p>\n\n<pre><code>__________________________________________________________________________________________________\nLayer (type)                    Output Shape         Param #     Connected to\n==================================================================================================\ninp_region (InputLayer)         (None, 1)            0\n__________________________________________________________________________________________________\ninp_parent_category_name (Input (None, 1)            0\n__________________________________________________________________________________________________\ninp_category_name (InputLayer)  (None, 1)            0\n__________________________________________________________________________________________________\ninp_user_type (InputLayer)      (None, 1)            0\n__________________________________________________________________________________________________\ninp_city (InputLayer)           (None, 1)            0\n__________________________________________________________________________________________________\ninp_week (InputLayer)           (None, 1)            0\n__________________________________________________________________________________________________\ninp_imgt1 (InputLayer)          (None, 1)            0\n__________________________________________________________________________________________________\ninp_p1 (InputLayer)             (None, 1)            0\n__________________________________________________________________________________________________\ninp_p2 (InputLayer)             (None, 1)            0\n__________________________________________________________________________________________________\ninp_p3 (InputLayer)             (None, 1)            0\n__________________________________________________________________________________________________\nemb_region (Embedding)          (None, 1, 14)        392         inp_region[0][0]\n__________________________________________________________________________________________________\nemb_parent_category_name (Embed (None, 1, 5)         45          inp_parent_category_name[0][0]\n__________________________________________________________________________________________________\nemb_category_name (Embedding)   (None, 1, 24)        1128        inp_category_name[0][0]\n__________________________________________________________________________________________________\nemb_user_type (Embedding)       (None, 1, 2)         6           inp_user_type[0][0]\n__________________________________________________________________________________________________\nemb_city (Embedding)            (None, 1, 64)        112128      inp_city[0][0]\n__________________________________________________________________________________________________\nemb_week (Embedding)            (None, 1, 4)         28          inp_week[0][0]\n__________________________________________________________________________________________________\nemb_imgt1 (Embedding)           (None, 1, 64)        196352      inp_imgt1[0][0]\n__________________________________________________________________________________________________\nemb_p1 (Embedding)              (None, 1, 64)        23808       inp_p1[0][0]\n__________________________________________________________________________________________________\nemb_p2 (Embedding)              (None, 1, 64)        17792       inp_p2[0][0]\n__________________________________________________________________________________________________\nemb_p3 (Embedding)              (None, 1, 64)        81728       inp_p3[0][0]\n__________________________________________________________________________________________________\nconcat_categorical_vars (Concat (None, 1, 369)       0           emb_region[0][0]\n                                                                 emb_parent_category_name[0][0]\n                                                                 emb_category_name[0][0]\n                                                                 emb_user_type[0][0]\n                                                                 emb_city[0][0]\n                                                                 emb_week[0][0]\n                                                                 emb_imgt1[0][0]\n                                                                 emb_p1[0][0]\n                                                                 emb_p2[0][0]\n                                                                 emb_p3[0][0]\n__________________________________________________________________________________________________\ninp_itemseq (InputLayer)        (None, 1)            0\n__________________________________________________________________________________________________\nflatten_1 (Flatten)             (None, 369)          0           concat_categorical_vars[0][0]\n__________________________________________________________________________________________________\ninp_price (InputLayer)          (None, 1)            0\n__________________________________________________________________________________________________\nemb_itemseq (Dense)             (None, 16)           32          inp_itemseq[0][0]\n__________________________________________________________________________________________________\ninp_avg_days_up_user (InputLaye (None, 1)            0\n__________________________________________________________________________________________________\navg_times_up_user (InputLayer)  (None, 1)            0\n__________________________________________________________________________________________________\nn_user_items (InputLayer)       (None, 1)            0\n__________________________________________________________________________________________________\ninp_filesize (InputLayer)       (None, 1)            0\n__________________________________________________________________________________________________\nconcatenate_1 (Concatenate)     (None, 390)          0           flatten_1[0][0]\n                                                                 inp_price[0][0]\n                                                                 emb_itemseq[0][0]\n                                                                 inp_avg_days_up_user[0][0]\n                                                                 avg_times_up_user[0][0]\n                                                                 n_user_items[0][0]\n                                                                 inp_filesize[0][0]\n__________________________________________________________________________________________________\ndense_1 (Dense)                 (None, 512)          200192      concatenate_1[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_1 (BatchNor (None, 512)          2048        dense_1[0][0]\n__________________________________________________________________________________________________\np_re_lu_1 (PReLU)               (None, 512)          512         batch_normalization_1[0][0]\n__________________________________________________________________________________________________\ndense_2 (Dense)                 (None, 512)          262656      p_re_lu_1[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_2 (BatchNor (None, 512)          2048        dense_2[0][0]\n__________________________________________________________________________________________________\np_re_lu_2 (PReLU)               (None, 512)          512         batch_normalization_2[0][0]\n__________________________________________________________________________________________________\ninp_desc (InputLayer)           (None, 256)          0\n__________________________________________________________________________________________________\ninp_title (InputLayer)          (None, 16)           0\n__________________________________________________________________________________________________\ndense_3 (Dense)                 (None, 512)          262656      p_re_lu_2[0][0]\n__________________________________________________________________________________________________\nemb_desc (Embedding)            multiple             15000300    inp_desc[0][0]\n                                                                 inp_title[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_3 (BatchNor (None, 512)          2048        dense_3[0][0]\n__________________________________________________________________________________________________\nbidirectional_1 (Bidirectional) (None, 256, 600)     1083600     emb_desc[0][0]\n__________________________________________________________________________________________________\nbidirectional_3 (Bidirectional) (None, 16, 600)      1083600     emb_desc[1][0]\n__________________________________________________________________________________________________\np_re_lu_3 (PReLU)               (None, 512)          512         batch_normalization_3[0][0]\n__________________________________________________________________________________________________\nbidirectional_2 (Bidirectional) (None, 600)          1623600     bidirectional_1[0][0]\n__________________________________________________________________________________________________\nbidirectional_4 (Bidirectional) (None, 600)          1623600     bidirectional_3[0][0]\n__________________________________________________________________________________________________\nconcatenate_2 (Concatenate)     (None, 1712)         0           p_re_lu_3[0][0]\n                                                                 bidirectional_2[0][0]\n                                                                 bidirectional_4[0][0]\n__________________________________________________________________________________________________\ndense_4 (Dense)                 (None, 4096)         7016448     concatenate_2[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_4 (BatchNor (None, 4096)         16384       dense_4[0][0]\n__________________________________________________________________________________________________\np_re_lu_4 (PReLU)               (None, 4096)         4096        batch_normalization_4[0][0]\n__________________________________________________________________________________________________\ndense_5 (Dense)                 (None, 4096)         16781312    p_re_lu_4[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_5 (BatchNor (None, 4096)         16384       dense_5[0][0]\n__________________________________________________________________________________________________\np_re_lu_5 (PReLU)               (None, 4096)         4096        batch_normalization_5[0][0]\n__________________________________________________________________________________________________\noutput (Dense)                  (None, 1)            4097        p_re_lu_5[0][0]\n==================================================================================================\nTotal params: 45,424,140\nTrainable params: 45,404,684\nNon-trainable params: 19,456\n__________________________________________________________________________________________________\n</code></pre>",
      "rawMarkdown": "Below my current best-ish model (0.2210 in 1 run local validation using 20% of samples, I have not submitted yet). \n\nI use pre-trained 300d vectors for language and i only take the 50,000 most common tokens which get fine-tuned (`emb_desc` layer). \n\nThere's embeddings for a lot of other thing and `inp_price` is just  log(price+eps). I am not using image-features below (but I have one which uses it extracting ResNet50 last layer and finetuning the whole thing, but takes a while to train so I'll leave it until the end).\n\nThings in my list:\n -  Add features like items per user\n - Make sense of item_seq\n - Automatic feature engineering a la DAE\n - Conv after RNN (GRU)\n\nIdeas welcome.\n\n    __________________________________________________________________________________________________\n    Layer (type)                    Output Shape         Param #     Connected to\n    ==================================================================================================\n    inp_region (InputLayer)         (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_parent_category_name (Input (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_category_name (InputLayer)  (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_user_type (InputLayer)      (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_city (InputLayer)           (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_week (InputLayer)           (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_imgt1 (InputLayer)          (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_p1 (InputLayer)             (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_p2 (InputLayer)             (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_p3 (InputLayer)             (None, 1)            0\n    __________________________________________________________________________________________________\n    emb_region (Embedding)          (None, 1, 14)        392         inp_region[0][0]\n    __________________________________________________________________________________________________\n    emb_parent_category_name (Embed (None, 1, 5)         45          inp_parent_category_name[0][0]\n    __________________________________________________________________________________________________\n    emb_category_name (Embedding)   (None, 1, 24)        1128        inp_category_name[0][0]\n    __________________________________________________________________________________________________\n    emb_user_type (Embedding)       (None, 1, 2)         6           inp_user_type[0][0]\n    __________________________________________________________________________________________________\n    emb_city (Embedding)            (None, 1, 64)        112128      inp_city[0][0]\n    __________________________________________________________________________________________________\n    emb_week (Embedding)            (None, 1, 4)         28          inp_week[0][0]\n    __________________________________________________________________________________________________\n    emb_imgt1 (Embedding)           (None, 1, 64)        196352      inp_imgt1[0][0]\n    __________________________________________________________________________________________________\n    emb_p1 (Embedding)              (None, 1, 64)        23808       inp_p1[0][0]\n    __________________________________________________________________________________________________\n    emb_p2 (Embedding)              (None, 1, 64)        17792       inp_p2[0][0]\n    __________________________________________________________________________________________________\n    emb_p3 (Embedding)              (None, 1, 64)        81728       inp_p3[0][0]\n    __________________________________________________________________________________________________\n    concat_categorical_vars (Concat (None, 1, 369)       0           emb_region[0][0]\n                                                                     emb_parent_category_name[0][0]\n                                                                     emb_category_name[0][0]\n                                                                     emb_user_type[0][0]\n                                                                     emb_city[0][0]\n                                                                     emb_week[0][0]\n                                                                     emb_imgt1[0][0]\n                                                                     emb_p1[0][0]\n                                                                     emb_p2[0][0]\n                                                                     emb_p3[0][0]\n    __________________________________________________________________________________________________\n    inp_itemseq (InputLayer)        (None, 1)            0\n    __________________________________________________________________________________________________\n    flatten_1 (Flatten)             (None, 369)          0           concat_categorical_vars[0][0]\n    __________________________________________________________________________________________________\n    inp_price (InputLayer)          (None, 1)            0\n    __________________________________________________________________________________________________\n    emb_itemseq (Dense)             (None, 16)           32          inp_itemseq[0][0]\n    __________________________________________________________________________________________________\n    inp_avg_days_up_user (InputLaye (None, 1)            0\n    __________________________________________________________________________________________________\n    avg_times_up_user (InputLayer)  (None, 1)            0\n    __________________________________________________________________________________________________\n    n_user_items (InputLayer)       (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_filesize (InputLayer)       (None, 1)            0\n    __________________________________________________________________________________________________\n    concatenate_1 (Concatenate)     (None, 390)          0           flatten_1[0][0]\n                                                                     inp_price[0][0]\n                                                                     emb_itemseq[0][0]\n                                                                     inp_avg_days_up_user[0][0]\n                                                                     avg_times_up_user[0][0]\n                                                                     n_user_items[0][0]\n                                                                     inp_filesize[0][0]\n    __________________________________________________________________________________________________\n    dense_1 (Dense)                 (None, 512)          200192      concatenate_1[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_1 (BatchNor (None, 512)          2048        dense_1[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_1 (PReLU)               (None, 512)          512         batch_normalization_1[0][0]\n    __________________________________________________________________________________________________\n    dense_2 (Dense)                 (None, 512)          262656      p_re_lu_1[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_2 (BatchNor (None, 512)          2048        dense_2[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_2 (PReLU)               (None, 512)          512         batch_normalization_2[0][0]\n    __________________________________________________________________________________________________\n    inp_desc (InputLayer)           (None, 256)          0\n    __________________________________________________________________________________________________\n    inp_title (InputLayer)          (None, 16)           0\n    __________________________________________________________________________________________________\n    dense_3 (Dense)                 (None, 512)          262656      p_re_lu_2[0][0]\n    __________________________________________________________________________________________________\n    emb_desc (Embedding)            multiple             15000300    inp_desc[0][0]\n                                                                     inp_title[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_3 (BatchNor (None, 512)          2048        dense_3[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_1 (Bidirectional) (None, 256, 600)     1083600     emb_desc[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_3 (Bidirectional) (None, 16, 600)      1083600     emb_desc[1][0]\n    __________________________________________________________________________________________________\n    p_re_lu_3 (PReLU)               (None, 512)          512         batch_normalization_3[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_2 (Bidirectional) (None, 600)          1623600     bidirectional_1[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_4 (Bidirectional) (None, 600)          1623600     bidirectional_3[0][0]\n    __________________________________________________________________________________________________\n    concatenate_2 (Concatenate)     (None, 1712)         0           p_re_lu_3[0][0]\n                                                                     bidirectional_2[0][0]\n                                                                     bidirectional_4[0][0]\n    __________________________________________________________________________________________________\n    dense_4 (Dense)                 (None, 4096)         7016448     concatenate_2[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_4 (BatchNor (None, 4096)         16384       dense_4[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_4 (PReLU)               (None, 4096)         4096        batch_normalization_4[0][0]\n    __________________________________________________________________________________________________\n    dense_5 (Dense)                 (None, 4096)         16781312    p_re_lu_4[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_5 (BatchNor (None, 4096)         16384       dense_5[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_5 (PReLU)               (None, 4096)         4096        batch_normalization_5[0][0]\n    __________________________________________________________________________________________________\n    output (Dense)                  (None, 1)            4097        p_re_lu_5[0][0]\n    ==================================================================================================\n    Total params: 45,424,140\n    Trainable params: 45,404,684\n    Non-trainable params: 19,456\n    __________________________________________________________________________________________________",
      "votes": 6,
      "replies": [
        {
          "id": 343260,
          "postDate": "2018-06-15T01:16:39.727Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 343332,
          "postDate": "2018-06-15T06:53:59.867Z",
          "content": "<p>IMO you have way too many parameters. Setting word embeddings aside your model has 30M parameters (for comparison: our model (scoring 0.2196 LB) has around 1.5M trainable parameters). Often less is more. I think especially your RNN layers have too many units which you then mitigate with batchnorm layers. Furthermore I am curious how did you set the embedding sizes of your categories? I currently use a \"one for all\" size.</p>",
          "rawMarkdown": "IMO you have way too many parameters. Setting word embeddings aside your model has 30M parameters (for comparison: our model (scoring 0.2196 LB) has around 1.5M trainable parameters). Often less is more. I think especially your RNN layers have too many units which you then mitigate with batchnorm layers. Furthermore I am curious how did you set the embedding sizes of your categories? I currently use a \"one for all\" size.",
          "votes": 2
        },
        {
          "id": 343389,
          "postDate": "2018-06-15T09:07:45.540Z",
          "content": "<p>Thanks for the insights. I also have a version with non-trainable text embeddings, so instead of picking top 50,000 tokens, preload them with fastText 300d embeddings from Facebook:</p>\n\n<pre><code>emb_desc (Embedding)            multiple             15000300    inp_desc[0][0] inp_title[0][0]\n</code></pre>\n\n<p>I use all tokens (~900,000 different) and set them to not trainable:</p>\n\n<pre><code>emb_desc (Embedding)            multiple             269846400   inp_desc[0][0] inp_title[0][0]\n</code></pre>\n\n<p>But all things being equal it still get bit worse CV (just 1-fold,  ˜10 epochs). Need to dig more.</p>\n\n<p>The  embedding size for the other is just:</p>\n\n<pre><code>## categorical\nmax_emb = 64\nconfig.emb_reg   = min(max_emb,(config.len_reg   + 1)//2)\nconfig.emb_pcn   = min(max_emb,(config.len_pcn   + 1)//2)\nconfig.emb_cn    = min(max_emb,(config.len_cn    + 1)//2)\nconfig.emb_ut    = min(max_emb,(config.len_ut    + 1)//2)\nconfig.emb_city  = min(max_emb,(config.len_city  + 1)//2)\nconfig.emb_week  = min(max_emb,(config.len_week  + 1)//2)\nconfig.emb_imgt1 = min(max_emb,(config.len_imgt1 + 1)//2)\nconfig.emb_p1    = min(max_emb,(config.len_p1    + 1)//2)\nconfig.emb_p2    = min(max_emb,(config.len_p2    + 1)//2)\nconfig.emb_p3    = min(max_emb,(config.len_p3    + 1)//2)\n</code></pre>\n\n<p>Which is roughly following Jeremy Howard (fast.ai...) guidelines for setting embedding size here: <a href=\"https://youtu.be/gbceqO8PpBg?t=3372\">https://youtu.be/gbceqO8PpBg?t=3372</a></p>\n\n<p>I think Im hitting a hard block getting lower than 0.221x I think I need more feature engineering but I really don't like going that route. Probably going to try DAE next. Even if it doesn't work it's going to be exciting.</p>",
          "rawMarkdown": "Thanks for the insights. I also have a version with non-trainable text embeddings, so instead of picking top 50,000 tokens, preload them with fastText 300d embeddings from Facebook:\n\n    emb_desc (Embedding)            multiple             15000300    inp_desc[0][0] inp_title[0][0]\n\nI use all tokens (~900,000 different) and set them to not trainable:\n\n    emb_desc (Embedding)            multiple             269846400   inp_desc[0][0] inp_title[0][0]\n\nBut all things being equal it still get bit worse CV (just 1-fold,  ˜10 epochs). Need to dig more.\n\nThe  embedding size for the other is just:\n\n    ## categorical\n    max_emb = 64\n    config.emb_reg   = min(max_emb,(config.len_reg   + 1)//2)\n    config.emb_pcn   = min(max_emb,(config.len_pcn   + 1)//2)\n    config.emb_cn    = min(max_emb,(config.len_cn    + 1)//2)\n    config.emb_ut    = min(max_emb,(config.len_ut    + 1)//2)\n    config.emb_city  = min(max_emb,(config.len_city  + 1)//2)\n    config.emb_week  = min(max_emb,(config.len_week  + 1)//2)\n    config.emb_imgt1 = min(max_emb,(config.len_imgt1 + 1)//2)\n    config.emb_p1    = min(max_emb,(config.len_p1    + 1)//2)\n    config.emb_p2    = min(max_emb,(config.len_p2    + 1)//2)\n    config.emb_p3    = min(max_emb,(config.len_p3    + 1)//2)\n\nWhich is roughly following Jeremy Howard (fast.ai...) guidelines for setting embedding size here: https://youtu.be/gbceqO8PpBg?t=3372\n\nI think Im hitting a hard block getting lower than 0.221x I think I need more feature engineering but I really don't like going that route. Probably going to try DAE next. Even if it doesn't work it's going to be exciting.",
          "votes": 1
        },
        {
          "id": 343395,
          "postDate": "2018-06-15T09:17:16.263Z",
          "content": "<p>What is your batch size?</p>",
          "rawMarkdown": "What is your batch size?"
        },
        {
          "id": 343397,
          "postDate": "2018-06-15T09:22:21.797Z",
          "content": "<p>1024 or 2048 if not using image (pixels) features, 64 otherwise. I use 2 x 1080 Ti for training.</p>",
          "rawMarkdown": "1024 or 2048 if not using image (pixels) features, 64 otherwise. I use 2 x 1080 Ti for training."
        }
      ]
    },
    {
      "id": 333164,
      "postDate": "2018-05-24T14:23:27.167Z",
      "content": "<p>By using pre-trained fasttext embedding and aggregated features now I can get around 0.224 LB score. </p>",
      "rawMarkdown": "By using pre-trained fasttext embedding and aggregated features now I can get around 0.224 LB score. ",
      "votes": 3
    },
    {
      "id": 342205,
      "postDate": "2018-06-13T04:56:35.840Z",
      "content": "<p>Neural Network 2 bagging with 5kfold is 0.2206.\nText feature is applied by CNN</p>",
      "rawMarkdown": "Neural Network 2 bagging with 5kfold is 0.2206.\nText feature is applied by CNN",
      "votes": 1,
      "replies": [
        {
          "id": 343257,
          "postDate": "2018-06-15T01:04:43.010Z",
          "content": "<p>what dense features are u using？</p>",
          "rawMarkdown": "what dense features are u using？",
          "votes": 1
        }
      ]
    },
    {
      "id": 341695,
      "postDate": "2018-06-12T05:19:25.363Z",
      "content": "<p>With a first shot we are at 0.2204 with simple single convolution for text. We will try more complicated (and time consuming ) structures next.</p>\n\n<p>Update: with RNN we are at 0.2196 </p>",
      "rawMarkdown": "With a first shot we are at 0.2204 with simple single convolution for text. We will try more complicated (and time consuming ) structures next.\n\nUpdate: with RNN we are at 0.2196 ",
      "votes": 1,
      "replies": [
        {
          "id": 341699,
          "postDate": "2018-06-12T05:26:10.017Z",
          "content": "<p>Would you mind to tell us what's your kernel size for CNN?</p>",
          "rawMarkdown": "Would you mind to tell us what's your kernel size for CNN?"
        },
        {
          "id": 341700,
          "postDate": "2018-06-12T05:27:46.033Z",
          "content": "<p>kernel size is 3. I need to add that lgb with roughly same features scores 0.2200. So text is only a minor  part</p>",
          "rawMarkdown": "kernel size is 3. I need to add that lgb with roughly same features scores 0.2200. So text is only a minor  part",
          "votes": 1
        },
        {
          "id": 342132,
          "postDate": "2018-06-13T00:22:37.200Z",
          "content": "<p>So this means you feed all manual features to your network?</p>",
          "rawMarkdown": "So this means you feed all manual features to your network?"
        },
        {
          "id": 343345,
          "postDate": "2018-06-15T07:21:39.590Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes",
          "votes": 1
        },
        {
          "id": 345584,
          "postDate": "2018-06-20T04:37:04.167Z",
          "content": "<p>When I add the features from lgbm, I found nn model is seriously overfit. Have you also met this problem?</p>",
          "rawMarkdown": "When I add the features from lgbm, I found nn model is seriously overfit. Have you also met this problem?"
        },
        {
          "id": 345633,
          "postDate": "2018-06-20T07:07:22.213Z",
          "content": "<p>Simply manually feeding the numerical features straight to the NN did not work for you ?</p>\n\n<p>Beware when you use features derived from others models. If they have «seen&nbsp;»  the whole target inside the training fold. they will necessarily overfit</p>",
          "rawMarkdown": "Simply manually feeding the numerical features straight to the NN did not work for you ?\n\nBeware when you use features derived from others models. If they have «seen&nbsp;»  the whole target inside the training fold. they will necessarily overfit"
        },
        {
          "id": 346131,
          "postDate": "2018-06-21T06:27:44.050Z",
          "content": "<p>Are you using just CNN's or combining GRU &amp; CNN's ?</p>",
          "rawMarkdown": "Are you using just CNN's or combining GRU &amp; CNN's ?"
        }
      ]
    },
    {
      "id": 333230,
      "postDate": "2018-05-24T17:52:34.843Z",
      "content": "<p>Single NN model at 0.2245 LB </p>",
      "rawMarkdown": "Single NN model at 0.2245 LB ",
      "votes": 1,
      "replies": [
        {
          "id": 339249,
          "postDate": "2018-06-06T15:34:53.503Z",
          "content": "<p>How many layers and what other hyper-parameters?</p>",
          "rawMarkdown": "How many layers and what other hyper-parameters?"
        }
      ]
    },
    {
      "id": 341661,
      "postDate": "2018-06-12T03:40:38.850Z",
      "content": "<p>Able to touch around 0.2215 with more epoch training and adjust dropout rate with GRU model plus dense feature</p>",
      "rawMarkdown": "Able to touch around 0.2215 with more epoch training and adjust dropout rate with GRU model plus dense feature"
    },
    {
      "id": 334479,
      "postDate": "2018-05-27T14:47:36.693Z",
      "content": "<p>my single nn model is now scoring 0.2219 on lb</p>",
      "rawMarkdown": "my single nn model is now scoring 0.2219 on lb",
      "replies": [
        {
          "id": 336935,
          "postDate": "2018-06-01T15:11:15.270Z",
          "content": "<p>Can you share some tips, for me I am staying at 0.224 for a while</p>",
          "rawMarkdown": "Can you share some tips, for me I am staying at 0.224 for a while"
        },
        {
          "id": 337491,
          "postDate": "2018-06-03T01:36:22.853Z",
          "content": "<p>Now my score is 0.2213 on the leaderboard and I haven't used image features yet. Structure like this <a href=\"https://www.kaggle.com/eashish/bidirectional-gru-with-convolution\"> LSTM with Convolution</a>. 10fold</p>",
          "rawMarkdown": "Now my score is 0.2213 on the leaderboard and I haven't used image features yet. Structure like this [ LSTM with Convolution][1]. 10fold\n\n\n  [1]: https://www.kaggle.com/eashish/bidirectional-gru-with-convolution",
          "votes": 5
        },
        {
          "id": 337492,
          "postDate": "2018-06-03T01:41:36.853Z",
          "content": "<p>I've seen getting global average and global max as features out of a RNN layer multiple times. Does it really make a big difference over only one of them ?</p>",
          "rawMarkdown": "I've seen getting global average and global max as features out of a RNN layer multiple times. Does it really make a big difference over only one of them ?"
        },
        {
          "id": 347393,
          "postDate": "2018-06-24T06:40:59.030Z",
          "content": "<p>@Arnaud Roussel, please check Sec 3.3 on this paper: <a href=\"https://arxiv.org/abs/1801.06146\">https://arxiv.org/abs/1801.06146</a></p>",
          "rawMarkdown": "@Arnaud Roussel, please check Sec 3.3 on this paper: https://arxiv.org/abs/1801.06146",
          "votes": 3
        },
        {
          "id": 347581,
          "postDate": "2018-06-24T18:39:46.713Z",
          "content": "<p>@Shujian Liu, Thanks ! So they suggest concatenating both the last state and Max/Avg across all states. So to do that in Keras you'd have to use </p>\n\n<p><code>lstm, state_h, _ = GRU(... , return_sequences=True, return_state=True)(...)</code></p>\n\n<p>and then send lstm through Global avg/max and concatenate that to state_h ?</p>",
          "rawMarkdown": "@Shujian Liu, Thanks ! So they suggest concatenating both the last state and Max/Avg across all states. So to do that in Keras you'd have to use \n\n`lstm, state_h, _ = GRU(... , return_sequences=True, return_state=True)(...)`\n\nand then send lstm through Global avg/max and concatenate that to state_h ?",
          "votes": 1
        }
      ]
    },
    {
      "id": 330926,
      "postDate": "2018-05-20T03:15:17.517Z",
      "content": "<p>I am especially interested how to successfully process the description and title using neural networks</p>",
      "rawMarkdown": "I am especially interested how to successfully process the description and title using neural networks",
      "replies": [
        {
          "id": 330953,
          "postDate": "2018-05-20T04:02:18.667Z",
          "content": "<p>From my experience processing description and title is the easy part. (I also just reused the toxic comments stuff) However its hard to get numerical, categorical and text parts to converge at the same rate.</p>",
          "rawMarkdown": "From my experience processing description and title is the easy part. (I also just reused the toxic comments stuff) However its hard to get numerical, categorical and text parts to converge at the same rate.",
          "votes": 4
        },
        {
          "id": 331302,
          "postDate": "2018-05-20T22:48:36.103Z",
          "content": "<p>In this competition seems LGBM is more powerful than NN. I think the only thing LGBM may missing is how to process the text, so that's why I am interested in this part</p>",
          "rawMarkdown": "In this competition seems LGBM is more powerful than NN. I think the only thing LGBM may missing is how to process the text, so that's why I am interested in this part"
        },
        {
          "id": 331310,
          "postDate": "2018-05-20T23:56:30.543Z",
          "content": "<p>I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.</p>",
          "rawMarkdown": "I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer."
        },
        {
          "id": 331541,
          "postDate": "2018-05-21T13:20:59.383Z",
          "content": "<blockquote>\n  <p>I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.</p>\n</blockquote>\n\n<p>For me it doesn't seem to work :( !  Are you using Adam?</p>",
          "rawMarkdown": "\n&gt; I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.\n\nFor me it doesn't seem to work :( !  Are you using Adam?"
        },
        {
          "id": 331555,
          "postDate": "2018-05-21T13:38:05.610Z",
          "content": "<p>@ Mohsin </p>\n\n<p>How many epochs are you running? </p>",
          "rawMarkdown": "@ Mohsin \n\nHow many epochs are you running? "
        },
        {
          "id": 331562,
          "postDate": "2018-05-21T13:44:16.563Z",
          "content": "<blockquote>\n  <p><strong>Strideradu wrote</strong></p>\n  \n  <blockquote>\n    <p>In this competition seems LGBM is more powerful than NN.</p>\n  </blockquote>\n</blockquote>\n\n<p>I don't think so ... (consider text features ...and probably images)</p>",
          "rawMarkdown": "\n&gt; **Strideradu wrote**\n&gt;\n&gt; &gt; In this competition seems LGBM is more powerful than NN.\n\nI don't think so ... (consider text features ...and probably images)"
        },
        {
          "id": 331566,
          "postDate": "2018-05-21T13:48:45.793Z",
          "content": "<blockquote>\n  <p><strong>Mohsin hasan wrote</strong></p>\n  \n  <blockquote>\n    <p>&gt; I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.</p>\n  </blockquote>\n  \n  <p>For me it doesn't seem to work :( !  Are you using Adam?</p>\n</blockquote>\n\n<p>Just finished trying it ...Didn't work for me either.  ('didn't improve'...'didn't improve' ...for all the extra-epochs)</p>",
          "rawMarkdown": "\n&gt; **Mohsin hasan wrote**\n&gt; \n&gt; &gt; \n&gt; &gt; I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.\n&gt; \n&gt; For me it doesn't seem to work :( !  Are you using Adam?\n\nJust finished trying it ...Didn't work for me either.  ('didn't improve'...'didn't improve' ...for all the extra-epochs)",
          "votes": 1
        },
        {
          "id": 331594,
          "postDate": "2018-05-21T14:38:42.887Z",
          "content": "<p>I apologize, after trying without it I realize it isn't working for me either :(</p>",
          "rawMarkdown": "I apologize, after trying without it I realize it isn't working for me either :("
        },
        {
          "id": 331621,
          "postDate": "2018-05-21T15:34:15.183Z",
          "content": "<p>@Dieter: you say, \"However it's hard to get numerical, categorical and text parts to converge at the same rate.\"  How did you get this insight?  Do you have a way to measure how different branches of your NN are converging?</p>\n\n<p>Personally, I use a combination of intuition and hyperparameter tuning / optimization.  I experiment with different depths and widths in my \"numerical, categorical and text parts\" sub-architectures (before they are concatenated together) and search for what gives me the nicest overall score and convergence.  It's only through this indirect procedure that I get information about part-wise convergence.</p>",
          "rawMarkdown": "@Dieter: you say, \"However it's hard to get numerical, categorical and text parts to converge at the same rate.\"  How did you get this insight?  Do you have a way to measure how different branches of your NN are converging?\n\nPersonally, I use a combination of intuition and hyperparameter tuning / optimization.  I experiment with different depths and widths in my \"numerical, categorical and text parts\" sub-architectures (before they are concatenated together) and search for what gives me the nicest overall score and convergence.  It's only through this indirect procedure that I get information about part-wise convergence.",
          "votes": 4
        },
        {
          "id": 331778,
          "postDate": "2018-05-21T20:37:42.803Z",
          "content": "<p>@Harlan Seymour</p>\n\n<p>Yes, I'm intrigued by this comment as well as it appears to provide a good explanation of some of the difficulty I've been having. I've searched the NN hypothesis space with some pretty elaborate architectures (and many inputs) and am settling on this being the issue. </p>\n\n<p>On the plus side for humanity, it's nice to report from (several steps away from) the front line that deep learning isn't about to take over the world, far too fiddly to train for that. ;)</p>\n\n<p>P.S. Are you using tf/keras and if so, have you tried TensorBoard?</p>",
          "rawMarkdown": "@Harlan Seymour\n\nYes, I'm intrigued by this comment as well as it appears to provide a good explanation of some of the difficulty I've been having. I've searched the NN hypothesis space with some pretty elaborate architectures (and many inputs) and am settling on this being the issue. \n\nOn the plus side for humanity, it's nice to report from (several steps away from) the front line that deep learning isn't about to take over the world, far too fiddly to train for that. ;)\n\nP.S. Are you using tf/keras and if so, have you tried TensorBoard?",
          "votes": 1
        },
        {
          "id": 331817,
          "postDate": "2018-05-21T23:14:05.223Z",
          "content": "<p><a href=\"/maw501\">@maw501</a>:  </p>\n\n<p>I use Keras and plot training and validation RMSE in my code.   I've played with TensorBoard (TB) but don't use it.  I understand that TB can also give me graphs of my layer activations and weights.  If I have a pretty decent model, and the activations and weights look OK (not pathologically bad), then I  don't know how TB can help me tweak my model to improve it further.  </p>\n\n<p>DNN's, partnered with gradient boosted trees, will provide the top solutions to this competition.   It's true that the DNN's are more fiddly than GBT, but we need them nonetheless!</p>",
          "rawMarkdown": "@maw501:  \n\nI use Keras and plot training and validation RMSE in my code.   I've played with TensorBoard (TB) but don't use it.  I understand that TB can also give me graphs of my layer activations and weights.  If I have a pretty decent model, and the activations and weights look OK (not pathologically bad), then I  don't know how TB can help me tweak my model to improve it further.  \n\nDNN's, partnered with gradient boosted trees, will provide the top solutions to this competition.   It's true that the DNN's are more fiddly than GBT, but we need them nonetheless!\n\n",
          "votes": 2
        },
        {
          "id": 331832,
          "postDate": "2018-05-22T00:01:40.250Z",
          "content": "<blockquote>\n  <p><strong>Shanth wrote</strong></p>\n  \n  <blockquote>\n    <p>@ Mohsin </p>\n  </blockquote>\n  \n  <p>How many epochs are you running? </p>\n</blockquote>\n\n<p>1 epoch to initial fit and then tried 1-3 for retraining</p>",
          "rawMarkdown": "\n&gt; **Shanth wrote**\n&gt; \n&gt; &gt; @ Mohsin \n&gt; \n&gt; How many epochs are you running? \n\n1 epoch to initial fit and then tried 1-3 for retraining",
          "votes": 3
        },
        {
          "id": 341698,
          "postDate": "2018-06-12T05:24:25.667Z",
          "content": "<p>@Harlan Seymor</p>\n\n<blockquote>\n  <p>How did you get this insight? Do you have a way to measure how different branches of your NN are converging?</p>\n</blockquote>\n\n<p>We take the individual parts of the NN (text, numerical, categorical) and see how many epochs the network trains before overfitting. Aim is to have approximately the same number of epochs. Then we put the network together and optimize hyperparameters a bit</p>",
          "rawMarkdown": "@Harlan Seymor\n\n&gt; How did you get this insight? Do you have a way to measure how different branches of your NN are converging?\n\nWe take the individual parts of the NN (text, numerical, categorical) and see how many epochs the network trains before overfitting. Aim is to have approximately the same number of epochs. Then we put the network together and optimize hyperparameters a bit"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 330989,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-05-20T05:47:26.070000",
      "content": "<p>My single NN model is now scoring 0.2217 on LB</p>\n\n<p>It's RNN and Description and title processed with self_trained embeddings...</p>",
      "votes": 13,
      "replies": [
        {
          "id": 331101,
          "author_name": "Maximilian Hahn",
          "author_url": "",
          "post_date": "2018-05-20T11:11:19.130000",
          "content": "<p>Nice result! Is it only text features or do you include structured features (price, categorical features, ...) as well?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331105,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-20T11:27:12.023000",
          "content": "<p>I use almost all the features (unless ids and image)....</p>\n\n<p>Item_seq_number was very noisy at the beginning but I finally manged to include  it.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 331111,
          "author_name": "Maximilian Hahn",
          "author_url": "",
          "post_date": "2018-05-20T11:39:15",
          "content": "<p>Thanks for the hint. I'm usually not a big fan of neural nets for structured data, but I want to try it for this competition. Seems like it's actually worth a try.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331190,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-20T15:26:57.313000",
          "content": "<blockquote>\n  <p>I'm usually not a big fan of neural nets for structured data</p>\n</blockquote>\n\n<p>Why?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331332,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-05-21T02:55:38.113000",
          "content": "<p>Serigne, I will start this competition with your instructions. Thanks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 331352,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-05-21T04:33:04.757000",
          "content": "<p>@ Maximilian </p>\n\n<p>Any reason why you would not use NN for structured data? And what would be your definition for structured data in this context?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331588,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-05-21T14:28:56.880000",
          "content": "<p>Neural networks being relatively less powerful on structured data is a pretty common heuristic, with good reason. They impose the functional form assumption that the target can be well-modeled by layering of hierarchical feature representations, so the data should support that hierarchical representation for the model to be ideal. That's certainly the case with data like images or text, where processing tasks involve building up higher order concepts from simpler ones - the best visual example is how image classification CNN filters learn generic edge and shape patterns in early layers and build out to more specific object representations in late layers.</p>\n\n<p>It's harder to see how hierarchical layering would be as naturally useful for most tabular tasks. For example, something like \"price\" is already a very well-described feature that doesn't need to be intricately manipulated to have a lot of predictive value, whereas individual pixels in an image need a lot of work to be leveraged as signals. This is why you'll often see that shallow neural networks outperform deeper ones on tabular data. And to capture explicit, direct feature interactions, it's more natural to use tree based methods (especially gradient-boosted trees) that allow for repeated interactive partitioning that derives from meaningful features as is instead of from more complex representations that need to be learned. </p>\n\n<p>So I usually expect to see gradient boosting outperform neural networks on tabular data. That doesn't mean they can't be competitive, especially with careful preprocessing and tuning. </p>\n\n<p>But this competition isn't tabular data, it's mixed. And in addition to their natural application to text and image data, one of the great benefits of neural networks is that their architecture is flexible - it's possible to incorporate as many different raw data types as you want into a single model. This obviously comes with its own challenges, but it is one possible solution to handling this dataset.  </p>",
          "votes": 33,
          "replies": []
        },
        {
          "id": 331626,
          "author_name": "Maximilian Hahn",
          "author_url": "",
          "post_date": "2018-05-21T15:47:04.230000",
          "content": "<p>Just what Joe Eddy said. If you look at all past Kaggle competitions not involving text or images, you will almost always find a GBM model as the best single model. There are some exceptions from this rule, but I think if you had one shot, you should always go for a GBM model in these cases. One counter example is the competition mentioned in this article: <a href=\"https://towardsdatascience.com/structured-deep-learning-b8ca4138b848\">https://towardsdatascience.com/structured-deep-learning-b8ca4138b848</a></p>\n\n<p>But as you can see, people had to come up with new ideas to actually build good NN models for structured date.</p>\n\n<p>P.S.: By \"structured\" I mean tabular data with categorical and numerical variables.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 331859,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-05-22T01:53:12.217000",
          "content": "<p>Another very cool exception that I think is worth mentioning is Porto Seguro - Michael Jahrer won with a solution that was primarily based around <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">denoising autoencoders</a>. It seems the reason this worked so well was that since the raw representations in that dataset were so messy (many missing values in some of the most important features), learned representations offered significant predictive gain. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 334803,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2018-05-28T12:47:54.410000",
          "content": "<p>Interesting. That would go in the direction proposed in another discussion (by Dieter ?) to have price and image models in order to have cleaner data before training.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335317,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-29T15:16:13.193000",
          "content": "<p>Hi Serigine, I am wondering if you use both train_active and test_active to train your model? And how many epoches? It seens one epoch took long time to finish</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 336964,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-06-01T16:19:04.097000",
          "content": "<p>Hi Strideradu </p>\n\n<p>Sorry for the late reply ...</p>\n\n<p>No I don't use them yet for training. </p>\n\n<p>As for the number of epoches, I don't specify it. I just use Callback with early_stopping criterion to let the model decide the number of epoches (to avoid overfitting) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 336970,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-06-01T16:30:43.200000",
          "content": "<p>Hi Serigine,</p>\n\n<p>Thanks for your answer. I am actually asking how many epochs you were using to train you embeddings, it seems no good way to decide the epochs?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 336987,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-06-01T17:29:47.880000",
          "content": "<p>Ah ok ! I read \"train your model\"  and thought you're  talking about the RNN framework. </p>\n\n<p>I use 3 epochs for the embeddings.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337099,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-01T23:30:06.977000",
          "content": "<p>@Serigne, are you able to get your NN model better than LB=0.2217? My best lgb model is at 0.2217 LB as well right now and am going to start modeling an NN solution. I am just wondering what the best NN score is before I start spending money on my cloud account.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337382,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-06-02T17:48:31.743000",
          "content": "<p>Hi ...I haven't touch it since it achieved this score.  ...I am working now with another model. </p>\n\n<p>However it can surely be improved with active and image features. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337505,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-03T02:55:31.213000",
          "content": "<p>Thanks @Serigne. I forgot to ask, your NN model ... 0.2217 LB score is with n-Folds right? Perhaps 5-Fold CV?</p>\n\n<p>I have my 1st attempt NN at LB = 0.2245 and the best lgb mentioned is for a single run. Hence, I am working on 5-Fold CV for my lgbm and other models before going back to improving my NN model. Hopefully, I can make a good jump with the 5-Fold :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337871,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-06-04T00:01:00.330000",
          "content": "<blockquote>\n  <p><strong>YaGana Sheriff-Hussaini wrote</strong></p>\n  \n  <blockquote>\n    <p>Thanks @Serigne. I forgot to ask, your NN model ... 0.2217 LB score is with n-Folds right? Perhaps 5-Fold CV?</p>\n  </blockquote>\n  \n  <p>I have my 1st attempt NN at LB = 0.2245 and the best lgb mentioned is for a single run. Hence, I am working on 5-Fold CV for my lgbm and other models before going back to improving my NN model. Hopefully, I can make a good jump with the 5-Fold :-)</p>\n</blockquote>\n\n<p>I am using 5 fold bagging predictions but I don't think it can make huge difference compared to single run , given that train and test seem to have pretty much the same distribution.</p>\n\n<p>I use it mainly to have coherent validation strategy for all my experiments and also to have OOF predictions for further blending/stacking.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 337947,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-04T05:56:16.253000",
          "content": "<p>Thanks @Serigne.  </p>\n\n<p>I have finally got my 5-Fold predictions working. My best model is now scoring 0.2215 on LB. Although hoping to get improvement with 5-fold, I am also doing it to get the OOF predictions for blending/bagging. Fingers crossed :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338870,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-06-05T22:01:42.410000",
          "content": "<p>@Serigne may I know how you deal with item_seq_number?  I just simply removed it to reduce noise and this did help a bit.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 342809,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2018-06-14T06:41:17.957000",
      "content": "<p>Below my current best-ish model (0.2210 in 1 run local validation using 20% of samples, I have not submitted yet). </p>\n\n<p>I use pre-trained 300d vectors for language and i only take the 50,000 most common tokens which get fine-tuned (<code>emb_desc</code> layer). </p>\n\n<p>There's embeddings for a lot of other thing and <code>inp_price</code> is just  log(price+eps). I am not using image-features below (but I have one which uses it extracting ResNet50 last layer and finetuning the whole thing, but takes a while to train so I'll leave it until the end).</p>\n\n<p>Things in my list:\n -  Add features like items per user\n - Make sense of item_seq\n - Automatic feature engineering a la DAE\n - Conv after RNN (GRU)</p>\n\n<p>Ideas welcome.</p>\n\n<pre><code>__________________________________________________________________________________________________\nLayer (type)                    Output Shape         Param #     Connected to\n==================================================================================================\ninp_region (InputLayer)         (None, 1)            0\n__________________________________________________________________________________________________\ninp_parent_category_name (Input (None, 1)            0\n__________________________________________________________________________________________________\ninp_category_name (InputLayer)  (None, 1)            0\n__________________________________________________________________________________________________\ninp_user_type (InputLayer)      (None, 1)            0\n__________________________________________________________________________________________________\ninp_city (InputLayer)           (None, 1)            0\n__________________________________________________________________________________________________\ninp_week (InputLayer)           (None, 1)            0\n__________________________________________________________________________________________________\ninp_imgt1 (InputLayer)          (None, 1)            0\n__________________________________________________________________________________________________\ninp_p1 (InputLayer)             (None, 1)            0\n__________________________________________________________________________________________________\ninp_p2 (InputLayer)             (None, 1)            0\n__________________________________________________________________________________________________\ninp_p3 (InputLayer)             (None, 1)            0\n__________________________________________________________________________________________________\nemb_region (Embedding)          (None, 1, 14)        392         inp_region[0][0]\n__________________________________________________________________________________________________\nemb_parent_category_name (Embed (None, 1, 5)         45          inp_parent_category_name[0][0]\n__________________________________________________________________________________________________\nemb_category_name (Embedding)   (None, 1, 24)        1128        inp_category_name[0][0]\n__________________________________________________________________________________________________\nemb_user_type (Embedding)       (None, 1, 2)         6           inp_user_type[0][0]\n__________________________________________________________________________________________________\nemb_city (Embedding)            (None, 1, 64)        112128      inp_city[0][0]\n__________________________________________________________________________________________________\nemb_week (Embedding)            (None, 1, 4)         28          inp_week[0][0]\n__________________________________________________________________________________________________\nemb_imgt1 (Embedding)           (None, 1, 64)        196352      inp_imgt1[0][0]\n__________________________________________________________________________________________________\nemb_p1 (Embedding)              (None, 1, 64)        23808       inp_p1[0][0]\n__________________________________________________________________________________________________\nemb_p2 (Embedding)              (None, 1, 64)        17792       inp_p2[0][0]\n__________________________________________________________________________________________________\nemb_p3 (Embedding)              (None, 1, 64)        81728       inp_p3[0][0]\n__________________________________________________________________________________________________\nconcat_categorical_vars (Concat (None, 1, 369)       0           emb_region[0][0]\n                                                                 emb_parent_category_name[0][0]\n                                                                 emb_category_name[0][0]\n                                                                 emb_user_type[0][0]\n                                                                 emb_city[0][0]\n                                                                 emb_week[0][0]\n                                                                 emb_imgt1[0][0]\n                                                                 emb_p1[0][0]\n                                                                 emb_p2[0][0]\n                                                                 emb_p3[0][0]\n__________________________________________________________________________________________________\ninp_itemseq (InputLayer)        (None, 1)            0\n__________________________________________________________________________________________________\nflatten_1 (Flatten)             (None, 369)          0           concat_categorical_vars[0][0]\n__________________________________________________________________________________________________\ninp_price (InputLayer)          (None, 1)            0\n__________________________________________________________________________________________________\nemb_itemseq (Dense)             (None, 16)           32          inp_itemseq[0][0]\n__________________________________________________________________________________________________\ninp_avg_days_up_user (InputLaye (None, 1)            0\n__________________________________________________________________________________________________\navg_times_up_user (InputLayer)  (None, 1)            0\n__________________________________________________________________________________________________\nn_user_items (InputLayer)       (None, 1)            0\n__________________________________________________________________________________________________\ninp_filesize (InputLayer)       (None, 1)            0\n__________________________________________________________________________________________________\nconcatenate_1 (Concatenate)     (None, 390)          0           flatten_1[0][0]\n                                                                 inp_price[0][0]\n                                                                 emb_itemseq[0][0]\n                                                                 inp_avg_days_up_user[0][0]\n                                                                 avg_times_up_user[0][0]\n                                                                 n_user_items[0][0]\n                                                                 inp_filesize[0][0]\n__________________________________________________________________________________________________\ndense_1 (Dense)                 (None, 512)          200192      concatenate_1[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_1 (BatchNor (None, 512)          2048        dense_1[0][0]\n__________________________________________________________________________________________________\np_re_lu_1 (PReLU)               (None, 512)          512         batch_normalization_1[0][0]\n__________________________________________________________________________________________________\ndense_2 (Dense)                 (None, 512)          262656      p_re_lu_1[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_2 (BatchNor (None, 512)          2048        dense_2[0][0]\n__________________________________________________________________________________________________\np_re_lu_2 (PReLU)               (None, 512)          512         batch_normalization_2[0][0]\n__________________________________________________________________________________________________\ninp_desc (InputLayer)           (None, 256)          0\n__________________________________________________________________________________________________\ninp_title (InputLayer)          (None, 16)           0\n__________________________________________________________________________________________________\ndense_3 (Dense)                 (None, 512)          262656      p_re_lu_2[0][0]\n__________________________________________________________________________________________________\nemb_desc (Embedding)            multiple             15000300    inp_desc[0][0]\n                                                                 inp_title[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_3 (BatchNor (None, 512)          2048        dense_3[0][0]\n__________________________________________________________________________________________________\nbidirectional_1 (Bidirectional) (None, 256, 600)     1083600     emb_desc[0][0]\n__________________________________________________________________________________________________\nbidirectional_3 (Bidirectional) (None, 16, 600)      1083600     emb_desc[1][0]\n__________________________________________________________________________________________________\np_re_lu_3 (PReLU)               (None, 512)          512         batch_normalization_3[0][0]\n__________________________________________________________________________________________________\nbidirectional_2 (Bidirectional) (None, 600)          1623600     bidirectional_1[0][0]\n__________________________________________________________________________________________________\nbidirectional_4 (Bidirectional) (None, 600)          1623600     bidirectional_3[0][0]\n__________________________________________________________________________________________________\nconcatenate_2 (Concatenate)     (None, 1712)         0           p_re_lu_3[0][0]\n                                                                 bidirectional_2[0][0]\n                                                                 bidirectional_4[0][0]\n__________________________________________________________________________________________________\ndense_4 (Dense)                 (None, 4096)         7016448     concatenate_2[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_4 (BatchNor (None, 4096)         16384       dense_4[0][0]\n__________________________________________________________________________________________________\np_re_lu_4 (PReLU)               (None, 4096)         4096        batch_normalization_4[0][0]\n__________________________________________________________________________________________________\ndense_5 (Dense)                 (None, 4096)         16781312    p_re_lu_4[0][0]\n__________________________________________________________________________________________________\nbatch_normalization_5 (BatchNor (None, 4096)         16384       dense_5[0][0]\n__________________________________________________________________________________________________\np_re_lu_5 (PReLU)               (None, 4096)         4096        batch_normalization_5[0][0]\n__________________________________________________________________________________________________\noutput (Dense)                  (None, 1)            4097        p_re_lu_5[0][0]\n==================================================================================================\nTotal params: 45,424,140\nTrainable params: 45,404,684\nNon-trainable params: 19,456\n__________________________________________________________________________________________________\n</code></pre>",
      "votes": 6,
      "replies": [
        {
          "id": 343260,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-15T01:16:39.727000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343332,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-15T06:53:59.867000",
          "content": "<p>IMO you have way too many parameters. Setting word embeddings aside your model has 30M parameters (for comparison: our model (scoring 0.2196 LB) has around 1.5M trainable parameters). Often less is more. I think especially your RNN layers have too many units which you then mitigate with batchnorm layers. Furthermore I am curious how did you set the embedding sizes of your categories? I currently use a \"one for all\" size.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 343389,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-06-15T09:07:45.540000",
          "content": "<p>Thanks for the insights. I also have a version with non-trainable text embeddings, so instead of picking top 50,000 tokens, preload them with fastText 300d embeddings from Facebook:</p>\n\n<pre><code>emb_desc (Embedding)            multiple             15000300    inp_desc[0][0] inp_title[0][0]\n</code></pre>\n\n<p>I use all tokens (~900,000 different) and set them to not trainable:</p>\n\n<pre><code>emb_desc (Embedding)            multiple             269846400   inp_desc[0][0] inp_title[0][0]\n</code></pre>\n\n<p>But all things being equal it still get bit worse CV (just 1-fold,  ˜10 epochs). Need to dig more.</p>\n\n<p>The  embedding size for the other is just:</p>\n\n<pre><code>## categorical\nmax_emb = 64\nconfig.emb_reg   = min(max_emb,(config.len_reg   + 1)//2)\nconfig.emb_pcn   = min(max_emb,(config.len_pcn   + 1)//2)\nconfig.emb_cn    = min(max_emb,(config.len_cn    + 1)//2)\nconfig.emb_ut    = min(max_emb,(config.len_ut    + 1)//2)\nconfig.emb_city  = min(max_emb,(config.len_city  + 1)//2)\nconfig.emb_week  = min(max_emb,(config.len_week  + 1)//2)\nconfig.emb_imgt1 = min(max_emb,(config.len_imgt1 + 1)//2)\nconfig.emb_p1    = min(max_emb,(config.len_p1    + 1)//2)\nconfig.emb_p2    = min(max_emb,(config.len_p2    + 1)//2)\nconfig.emb_p3    = min(max_emb,(config.len_p3    + 1)//2)\n</code></pre>\n\n<p>Which is roughly following Jeremy Howard (fast.ai...) guidelines for setting embedding size here: <a href=\"https://youtu.be/gbceqO8PpBg?t=3372\">https://youtu.be/gbceqO8PpBg?t=3372</a></p>\n\n<p>I think Im hitting a hard block getting lower than 0.221x I think I need more feature engineering but I really don't like going that route. Probably going to try DAE next. Even if it doesn't work it's going to be exciting.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343395,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-06-15T09:17:16.263000",
          "content": "<p>What is your batch size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343397,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-06-15T09:22:21.797000",
          "content": "<p>1024 or 2048 if not using image (pixels) features, 64 otherwise. I use 2 x 1080 Ti for training.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 333164,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2018-05-24T14:23:27.167000",
      "content": "<p>By using pre-trained fasttext embedding and aggregated features now I can get around 0.224 LB score. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 342205,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "2018-06-13T04:56:35.840000",
      "content": "<p>Neural Network 2 bagging with 5kfold is 0.2206.\nText feature is applied by CNN</p>",
      "votes": 1,
      "replies": [
        {
          "id": 343257,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-06-15T01:04:43.010000",
          "content": "<p>what dense features are u using？</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 341695,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-06-12T05:19:25.363000",
      "content": "<p>With a first shot we are at 0.2204 with simple single convolution for text. We will try more complicated (and time consuming ) structures next.</p>\n\n<p>Update: with RNN we are at 0.2196 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 341699,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-06-12T05:26:10.017000",
          "content": "<p>Would you mind to tell us what's your kernel size for CNN?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341700,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-12T05:27:46.033000",
          "content": "<p>kernel size is 3. I need to add that lgb with roughly same features scores 0.2200. So text is only a minor  part</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 342132,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-06-13T00:22:37.200000",
          "content": "<p>So this means you feed all manual features to your network?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343345,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-15T07:21:39.590000",
          "content": "<p>yes</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345584,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-06-20T04:37:04.167000",
          "content": "<p>When I add the features from lgbm, I found nn model is seriously overfit. Have you also met this problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345633,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-06-20T07:07:22.213000",
          "content": "<p>Simply manually feeding the numerical features straight to the NN did not work for you ?</p>\n\n<p>Beware when you use features derived from others models. If they have «seen&nbsp;»  the whole target inside the training fold. they will necessarily overfit</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346131,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2018-06-21T06:27:44.050000",
          "content": "<p>Are you using just CNN's or combining GRU &amp; CNN's ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 333230,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-05-24T17:52:34.843000",
      "content": "<p>Single NN model at 0.2245 LB </p>",
      "votes": 1,
      "replies": [
        {
          "id": 339249,
          "author_name": "Temi Babs",
          "author_url": "",
          "post_date": "2018-06-06T15:34:53.503000",
          "content": "<p>How many layers and what other hyper-parameters?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 341661,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2018-06-12T03:40:38.850000",
      "content": "<p>Able to touch around 0.2215 with more epoch training and adjust dropout rate with GRU model plus dense feature</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 334479,
      "author_name": "jetou Xu",
      "author_url": "",
      "post_date": "2018-05-27T14:47:36.693000",
      "content": "<p>my single nn model is now scoring 0.2219 on lb</p>",
      "votes": 0,
      "replies": [
        {
          "id": 336935,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-06-01T15:11:15.270000",
          "content": "<p>Can you share some tips, for me I am staying at 0.224 for a while</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337491,
          "author_name": "jetou Xu",
          "author_url": "",
          "post_date": "2018-06-03T01:36:22.853000",
          "content": "<p>Now my score is 0.2213 on the leaderboard and I haven't used image features yet. Structure like this <a href=\"https://www.kaggle.com/eashish/bidirectional-gru-with-convolution\"> LSTM with Convolution</a>. 10fold</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 337492,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2018-06-03T01:41:36.853000",
          "content": "<p>I've seen getting global average and global max as features out of a RNN layer multiple times. Does it really make a big difference over only one of them ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 347393,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-06-24T06:40:59.030000",
          "content": "<p>@Arnaud Roussel, please check Sec 3.3 on this paper: <a href=\"https://arxiv.org/abs/1801.06146\">https://arxiv.org/abs/1801.06146</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 347581,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2018-06-24T18:39:46.713000",
          "content": "<p>@Shujian Liu, Thanks ! So they suggest concatenating both the last state and Max/Avg across all states. So to do that in Keras you'd have to use </p>\n\n<p><code>lstm, state_h, _ = GRU(... , return_sequences=True, return_state=True)(...)</code></p>\n\n<p>and then send lstm through Global avg/max and concatenate that to state_h ?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 330926,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2018-05-20T03:15:17.517000",
      "content": "<p>I am especially interested how to successfully process the description and title using neural networks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 330953,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-20T04:02:18.667000",
          "content": "<p>From my experience processing description and title is the easy part. (I also just reused the toxic comments stuff) However its hard to get numerical, categorical and text parts to converge at the same rate.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 331302,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-20T22:48:36.103000",
          "content": "<p>In this competition seems LGBM is more powerful than NN. I think the only thing LGBM may missing is how to process the text, so that's why I am interested in this part</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331310,
          "author_name": "Callum Gundlach",
          "author_url": "",
          "post_date": "2018-05-20T23:56:30.543000",
          "content": "<p>I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331541,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2018-05-21T13:20:59.383000",
          "content": "<blockquote>\n  <p>I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.</p>\n</blockquote>\n\n<p>For me it doesn't seem to work :( !  Are you using Adam?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331555,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-05-21T13:38:05.610000",
          "content": "<p>@ Mohsin </p>\n\n<p>How many epochs are you running? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331562,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-21T13:44:16.563000",
          "content": "<blockquote>\n  <p><strong>Strideradu wrote</strong></p>\n  \n  <blockquote>\n    <p>In this competition seems LGBM is more powerful than NN.</p>\n  </blockquote>\n</blockquote>\n\n<p>I don't think so ... (consider text features ...and probably images)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331566,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-21T13:48:45.793000",
          "content": "<blockquote>\n  <p><strong>Mohsin hasan wrote</strong></p>\n  \n  <blockquote>\n    <p>&gt; I've been disabling the numerical and categorical embeddings after so many epochs and letting the text processing layers train for longer.</p>\n  </blockquote>\n  \n  <p>For me it doesn't seem to work :( !  Are you using Adam?</p>\n</blockquote>\n\n<p>Just finished trying it ...Didn't work for me either.  ('didn't improve'...'didn't improve' ...for all the extra-epochs)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331594,
          "author_name": "Callum Gundlach",
          "author_url": "",
          "post_date": "2018-05-21T14:38:42.887000",
          "content": "<p>I apologize, after trying without it I realize it isn't working for me either :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331621,
          "author_name": "Harlan Seymour",
          "author_url": "",
          "post_date": "2018-05-21T15:34:15.183000",
          "content": "<p>@Dieter: you say, \"However it's hard to get numerical, categorical and text parts to converge at the same rate.\"  How did you get this insight?  Do you have a way to measure how different branches of your NN are converging?</p>\n\n<p>Personally, I use a combination of intuition and hyperparameter tuning / optimization.  I experiment with different depths and widths in my \"numerical, categorical and text parts\" sub-architectures (before they are concatenated together) and search for what gives me the nicest overall score and convergence.  It's only through this indirect procedure that I get information about part-wise convergence.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 331778,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-05-21T20:37:42.803000",
          "content": "<p>@Harlan Seymour</p>\n\n<p>Yes, I'm intrigued by this comment as well as it appears to provide a good explanation of some of the difficulty I've been having. I've searched the NN hypothesis space with some pretty elaborate architectures (and many inputs) and am settling on this being the issue. </p>\n\n<p>On the plus side for humanity, it's nice to report from (several steps away from) the front line that deep learning isn't about to take over the world, far too fiddly to train for that. ;)</p>\n\n<p>P.S. Are you using tf/keras and if so, have you tried TensorBoard?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331817,
          "author_name": "Harlan Seymour",
          "author_url": "",
          "post_date": "2018-05-21T23:14:05.223000",
          "content": "<p><a href=\"/maw501\">@maw501</a>:  </p>\n\n<p>I use Keras and plot training and validation RMSE in my code.   I've played with TensorBoard (TB) but don't use it.  I understand that TB can also give me graphs of my layer activations and weights.  If I have a pretty decent model, and the activations and weights look OK (not pathologically bad), then I  don't know how TB can help me tweak my model to improve it further.  </p>\n\n<p>DNN's, partnered with gradient boosted trees, will provide the top solutions to this competition.   It's true that the DNN's are more fiddly than GBT, but we need them nonetheless!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 331832,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2018-05-22T00:01:40.250000",
          "content": "<blockquote>\n  <p><strong>Shanth wrote</strong></p>\n  \n  <blockquote>\n    <p>@ Mohsin </p>\n  </blockquote>\n  \n  <p>How many epochs are you running? </p>\n</blockquote>\n\n<p>1 epoch to initial fit and then tried 1-3 for retraining</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 341698,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-12T05:24:25.667000",
          "content": "<p>@Harlan Seymor</p>\n\n<blockquote>\n  <p>How did you get this insight? Do you have a way to measure how different branches of your NN are converging?</p>\n</blockquote>\n\n<p>We take the individual parts of the NN (text, numerical, categorical) and see how many epochs the network trains before overfitting. Aim is to have approximately the same number of epochs. Then we put the network together and optimize hyperparameters a bit</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "330488": "@Dieter have a very good discussion on the best single model. However, most results are for the lgbm model. Since I start working on the neural network model, I would like to discuss the best model for neural network and also the network structure ( if you like to share).\n\nRight now my 2 layer GRU model with global max pooling for title and description is about 0.23, I just copy the model I used for the toxic comment challenge.\n\nHaven't try an NN model just like a classifier",
    "330989": "My single NN model is now scoring 0.2217 on LB\n\nIt's RNN and Description and title processed with self_trained embeddings...\n",
    "342809": "Below my current best-ish model (0.2210 in 1 run local validation using 20% of samples, I have not submitted yet). \n\nI use pre-trained 300d vectors for language and i only take the 50,000 most common tokens which get fine-tuned (`emb_desc` layer). \n\nThere's embeddings for a lot of other thing and `inp_price` is just  log(price+eps). I am not using image-features below (but I have one which uses it extracting ResNet50 last layer and finetuning the whole thing, but takes a while to train so I'll leave it until the end).\n\nThings in my list:\n -  Add features like items per user\n - Make sense of item_seq\n - Automatic feature engineering a la DAE\n - Conv after RNN (GRU)\n\nIdeas welcome.\n\n    __________________________________________________________________________________________________\n    Layer (type)                    Output Shape         Param #     Connected to\n    ==================================================================================================\n    inp_region (InputLayer)         (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_parent_category_name (Input (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_category_name (InputLayer)  (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_user_type (InputLayer)      (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_city (InputLayer)           (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_week (InputLayer)           (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_imgt1 (InputLayer)          (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_p1 (InputLayer)             (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_p2 (InputLayer)             (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_p3 (InputLayer)             (None, 1)            0\n    __________________________________________________________________________________________________\n    emb_region (Embedding)          (None, 1, 14)        392         inp_region[0][0]\n    __________________________________________________________________________________________________\n    emb_parent_category_name (Embed (None, 1, 5)         45          inp_parent_category_name[0][0]\n    __________________________________________________________________________________________________\n    emb_category_name (Embedding)   (None, 1, 24)        1128        inp_category_name[0][0]\n    __________________________________________________________________________________________________\n    emb_user_type (Embedding)       (None, 1, 2)         6           inp_user_type[0][0]\n    __________________________________________________________________________________________________\n    emb_city (Embedding)            (None, 1, 64)        112128      inp_city[0][0]\n    __________________________________________________________________________________________________\n    emb_week (Embedding)            (None, 1, 4)         28          inp_week[0][0]\n    __________________________________________________________________________________________________\n    emb_imgt1 (Embedding)           (None, 1, 64)        196352      inp_imgt1[0][0]\n    __________________________________________________________________________________________________\n    emb_p1 (Embedding)              (None, 1, 64)        23808       inp_p1[0][0]\n    __________________________________________________________________________________________________\n    emb_p2 (Embedding)              (None, 1, 64)        17792       inp_p2[0][0]\n    __________________________________________________________________________________________________\n    emb_p3 (Embedding)              (None, 1, 64)        81728       inp_p3[0][0]\n    __________________________________________________________________________________________________\n    concat_categorical_vars (Concat (None, 1, 369)       0           emb_region[0][0]\n                                                                     emb_parent_category_name[0][0]\n                                                                     emb_category_name[0][0]\n                                                                     emb_user_type[0][0]\n                                                                     emb_city[0][0]\n                                                                     emb_week[0][0]\n                                                                     emb_imgt1[0][0]\n                                                                     emb_p1[0][0]\n                                                                     emb_p2[0][0]\n                                                                     emb_p3[0][0]\n    __________________________________________________________________________________________________\n    inp_itemseq (InputLayer)        (None, 1)            0\n    __________________________________________________________________________________________________\n    flatten_1 (Flatten)             (None, 369)          0           concat_categorical_vars[0][0]\n    __________________________________________________________________________________________________\n    inp_price (InputLayer)          (None, 1)            0\n    __________________________________________________________________________________________________\n    emb_itemseq (Dense)             (None, 16)           32          inp_itemseq[0][0]\n    __________________________________________________________________________________________________\n    inp_avg_days_up_user (InputLaye (None, 1)            0\n    __________________________________________________________________________________________________\n    avg_times_up_user (InputLayer)  (None, 1)            0\n    __________________________________________________________________________________________________\n    n_user_items (InputLayer)       (None, 1)            0\n    __________________________________________________________________________________________________\n    inp_filesize (InputLayer)       (None, 1)            0\n    __________________________________________________________________________________________________\n    concatenate_1 (Concatenate)     (None, 390)          0           flatten_1[0][0]\n                                                                     inp_price[0][0]\n                                                                     emb_itemseq[0][0]\n                                                                     inp_avg_days_up_user[0][0]\n                                                                     avg_times_up_user[0][0]\n                                                                     n_user_items[0][0]\n                                                                     inp_filesize[0][0]\n    __________________________________________________________________________________________________\n    dense_1 (Dense)                 (None, 512)          200192      concatenate_1[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_1 (BatchNor (None, 512)          2048        dense_1[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_1 (PReLU)               (None, 512)          512         batch_normalization_1[0][0]\n    __________________________________________________________________________________________________\n    dense_2 (Dense)                 (None, 512)          262656      p_re_lu_1[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_2 (BatchNor (None, 512)          2048        dense_2[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_2 (PReLU)               (None, 512)          512         batch_normalization_2[0][0]\n    __________________________________________________________________________________________________\n    inp_desc (InputLayer)           (None, 256)          0\n    __________________________________________________________________________________________________\n    inp_title (InputLayer)          (None, 16)           0\n    __________________________________________________________________________________________________\n    dense_3 (Dense)                 (None, 512)          262656      p_re_lu_2[0][0]\n    __________________________________________________________________________________________________\n    emb_desc (Embedding)            multiple             15000300    inp_desc[0][0]\n                                                                     inp_title[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_3 (BatchNor (None, 512)          2048        dense_3[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_1 (Bidirectional) (None, 256, 600)     1083600     emb_desc[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_3 (Bidirectional) (None, 16, 600)      1083600     emb_desc[1][0]\n    __________________________________________________________________________________________________\n    p_re_lu_3 (PReLU)               (None, 512)          512         batch_normalization_3[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_2 (Bidirectional) (None, 600)          1623600     bidirectional_1[0][0]\n    __________________________________________________________________________________________________\n    bidirectional_4 (Bidirectional) (None, 600)          1623600     bidirectional_3[0][0]\n    __________________________________________________________________________________________________\n    concatenate_2 (Concatenate)     (None, 1712)         0           p_re_lu_3[0][0]\n                                                                     bidirectional_2[0][0]\n                                                                     bidirectional_4[0][0]\n    __________________________________________________________________________________________________\n    dense_4 (Dense)                 (None, 4096)         7016448     concatenate_2[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_4 (BatchNor (None, 4096)         16384       dense_4[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_4 (PReLU)               (None, 4096)         4096        batch_normalization_4[0][0]\n    __________________________________________________________________________________________________\n    dense_5 (Dense)                 (None, 4096)         16781312    p_re_lu_4[0][0]\n    __________________________________________________________________________________________________\n    batch_normalization_5 (BatchNor (None, 4096)         16384       dense_5[0][0]\n    __________________________________________________________________________________________________\n    p_re_lu_5 (PReLU)               (None, 4096)         4096        batch_normalization_5[0][0]\n    __________________________________________________________________________________________________\n    output (Dense)                  (None, 1)            4097        p_re_lu_5[0][0]\n    ==================================================================================================\n    Total params: 45,424,140\n    Trainable params: 45,404,684\n    Non-trainable params: 19,456\n    __________________________________________________________________________________________________",
    "333164": "By using pre-trained fasttext embedding and aggregated features now I can get around 0.224 LB score. ",
    "342205": "Neural Network 2 bagging with 5kfold is 0.2206.\nText feature is applied by CNN",
    "341695": "With a first shot we are at 0.2204 with simple single convolution for text. We will try more complicated (and time consuming ) structures next.\n\nUpdate: with RNN we are at 0.2196 ",
    "333230": "Single NN model at 0.2245 LB ",
    "341661": "Able to touch around 0.2215 with more epoch training and adjust dropout rate with GRU model plus dense feature",
    "334479": "my single nn model is now scoring 0.2219 on lb",
    "330926": "I am especially interested how to successfully process the description and title using neural networks"
  }
}