{
  "id": 56816,
  "title": "Best single model",
  "url": "/competitions/avito-demand-prediction/discussion/56816",
  "author_name": "Dieter",
  "post_date": "2018-05-15T11:17:09.201000",
  "votes": 47,
  "comment_count": 250,
  "views": 0,
  "content": "<p>I wonder what the best single model score is. I got my NN to 0.2240, but havn't included any image features nor the active and period files yet. How much is the improvement with image and co.?</p>",
  "messages": [
    {
      "id": 328917,
      "postDate": "2018-05-15T11:17:09.200Z",
      "content": "<p>I wonder what the best single model score is. I got my NN to 0.2240, but havn't included any image features nor the active and period files yet. How much is the improvement with image and co.?</p>",
      "rawMarkdown": "I wonder what the best single model score is. I got my NN to 0.2240, but havn't included any image features nor the active and period files yet. How much is the improvement with image and co.?",
      "votes": 47
    },
    {
      "id": 340360,
      "postDate": "2018-06-09T02:28:20.813Z",
      "content": "<p>I got 0.2181 with a NN network with 10 folds (UPDATED)</p>",
      "rawMarkdown": "I got 0.2181 with a NN network with 10 folds (UPDATED)\n",
      "votes": 17,
      "replies": [
        {
          "id": 340367,
          "postDate": "2018-06-09T02:45:00.490Z",
          "content": "<p>Wow, that's great!\nIs the feature amount effective? Or is the network special?</p>",
          "rawMarkdown": "Wow, that's great!\nIs the feature amount effective? Or is the network special?"
        },
        {
          "id": 340372,
          "postDate": "2018-06-09T03:01:40.660Z",
          "content": "<p>Wow,can I ask that using MLP or RNN or CNN?</p>",
          "rawMarkdown": "Wow,can I ask that using MLP or RNN or CNN?"
        },
        {
          "id": 340378,
          "postDate": "2018-06-09T03:23:17.113Z",
          "content": "<p>Nothing special. I use four kind features: continous (contains ImageNet score), category, text with embedding, raw image pixel. Four submodels for four kinds feature, concatenated to one tensor, serval Dense Layers. Heavily use BatchNormalization and Dropout. I believe there still have some room to improve because I rarely did feature engineering.</p>",
          "rawMarkdown": "Nothing special. I use four kind features: continous (contains ImageNet score), category, text with embedding, raw image pixel. Four submodels for four kinds feature, concatenated to one tensor, serval Dense Layers. Heavily use BatchNormalization and Dropout. I believe there still have some room to improve because I rarely did feature engineering.",
          "votes": 9
        },
        {
          "id": 340381,
          "postDate": "2018-06-09T03:32:07.747Z",
          "content": "<p>Thank you for sharing!\nMy nn is still 0.2222(TT)  I will do my best from here!</p>",
          "rawMarkdown": "Thank you for sharing!\nMy nn is still 0.2222(TT)  I will do my best from here!"
        },
        {
          "id": 340382,
          "postDate": "2018-06-09T03:42:58.173Z",
          "content": "<p>My LGB is still 0.2205 and I have no idea how to improve it. Poor feature engineer skills.</p>",
          "rawMarkdown": "My LGB is still 0.2205 and I have no idea how to improve it. Poor feature engineer skills.",
          "votes": 3
        },
        {
          "id": 340388,
          "postDate": "2018-06-09T04:12:33.540Z",
          "content": "<p>Wah!!! Great.. R u using self-embedding ? \"Raw image pixel\" - is it like using the raw pixel value of the image..How r u using it in the NN model?</p>",
          "rawMarkdown": "Wah!!! Great.. R u using self-embedding ? \"Raw image pixel\" - is it like using the raw pixel value of the image..How r u using it in the NN model?\n"
        },
        {
          "id": 340389,
          "postDate": "2018-06-09T04:13:28.803Z",
          "content": "<p>Congrats @Liu Jilong,</p>\n\n<blockquote>\n  <p>raw image pixel </p>\n</blockquote>\n\n<p>Do you mean the raw image without through any pretrained model? <br>\nIf yes, that is impressive. How long does it take to train?</p>",
          "rawMarkdown": "Congrats @Liu Jilong,\n\n&gt; raw image pixel \n\n\nDo you mean the raw image without through any pretrained model?  \nIf yes, that is impressive. How long does it take to train?",
          "votes": 1
        },
        {
          "id": 340391,
          "postDate": "2018-06-09T04:19:22.943Z",
          "content": "<p>Thanks for sharing, I have confused that you train four model for for  four kinds feature,then a final model based on the four predictions? </p>",
          "rawMarkdown": "Thanks for sharing, I have confused that you train four model for for  four kinds feature,then a final model based on the four predictions? "
        },
        {
          "id": 340400,
          "postDate": "2018-06-09T04:58:48.003Z",
          "content": "<p>Just like any other image network, serval Conv2D layers.</p>",
          "rawMarkdown": "Just like any other image network, serval Conv2D layers."
        },
        {
          "id": 340401,
          "postDate": "2018-06-09T04:59:20.403Z",
          "content": "<p>UPDATED: about 15min per epoch, 2.5h per fold.</p>",
          "rawMarkdown": "UPDATED: about 15min per epoch, 2.5h per fold.",
          "votes": 1
        },
        {
          "id": 340403,
          "postDate": "2018-06-09T05:08:52.107Z",
          "content": "<p>Thanks Liu... Any insight on text pre-processing / embedding?</p>",
          "rawMarkdown": "Thanks Liu... Any insight on text pre-processing / embedding?"
        },
        {
          "id": 340405,
          "postDate": "2018-06-09T05:19:17.890Z",
          "content": "<p>Just as many kagglers shared, I use self-trained word2vec embedding.  I have no idea about russian, so no text pre-processing is used.</p>",
          "rawMarkdown": "Just as many kagglers shared, I use self-trained word2vec embedding.  I have no idea about russian, so no text pre-processing is used."
        },
        {
          "id": 340407,
          "postDate": "2018-06-09T05:26:39.460Z",
          "content": "<p>Thank you for sharing.</p>\n\n<p>My situation is the opposite of you. I just started trying NN today,  and it didn't look good. </p>\n\n<p>Very valuable experience, thanks :).</p>",
          "rawMarkdown": "Thank you for sharing.\n\nMy situation is the opposite of you. I just started trying NN today,  and it didn't look good. \n\nVery valuable experience, thanks :)."
        },
        {
          "id": 340408,
          "postDate": "2018-06-09T05:31:06.123Z",
          "content": "<p>Many tricks to make NN work well. Good luck.</p>",
          "rawMarkdown": "Many tricks to make NN work well. Good luck.",
          "votes": 1
        },
        {
          "id": 340410,
          "postDate": "2018-06-09T05:32:25.727Z",
          "content": "<p>Thanks. Appreciate your help...</p>",
          "rawMarkdown": "Thanks. Appreciate your help..."
        },
        {
          "id": 340430,
          "postDate": "2018-06-09T06:21:50.623Z",
          "content": "<p>Thanks. wish to see experience sharing for NN model after this competition.I have no idea to train a good NN model.My NN model is poor score.</p>",
          "rawMarkdown": "Thanks. wish to see experience sharing for NN model after this competition.I have no idea to train a good NN model.My NN model is poor score."
        },
        {
          "id": 340688,
          "postDate": "2018-06-10T02:41:29.237Z",
          "content": "<p>Thanks for your sharing. I have a memory problem of training a CNN on images. Since I don't have enough memory to load all images, currently I use a generator to read images batch by batch from disk and feed them to the CNN, which makes the training process very slow (1 hour per epoch). Is there any better way to make it faster?</p>",
          "rawMarkdown": "Thanks for your sharing. I have a memory problem of training a CNN on images. Since I don't have enough memory to load all images, currently I use a generator to read images batch by batch from disk and feed them to the CNN, which makes the training process very slow (1 hour per epoch). Is there any better way to make it faster?"
        },
        {
          "id": 340691,
          "postDate": "2018-06-10T03:20:04.987Z",
          "content": "<p>I also tried fit_generator, but fit_generator's result is worse than fit, which is strange and I did not find the reason. As for time-consuming problem, you can try fit_generator's multi-processing parameter.</p>",
          "rawMarkdown": "I also tried fit_generator, but fit_generator's result is worse than fit, which is strange and I did not find the reason. As for time-consuming problem, you can try fit_generator's multi-processing parameter."
        },
        {
          "id": 340699,
          "postDate": "2018-06-10T03:47:12.887Z",
          "content": "<p>Wondering 15min/epoch for what kind of GPU?</p>",
          "rawMarkdown": "Wondering 15min/epoch for what kind of GPU?"
        },
        {
          "id": 340702,
          "postDate": "2018-06-10T04:02:07.473Z",
          "content": "<p>gtx1070</p>",
          "rawMarkdown": "gtx1070"
        },
        {
          "id": 341040,
          "postDate": "2018-06-10T21:38:04.990Z",
          "content": "<p>Hi Liu,\nThanks for your sharing, it's very encouraging. I wonder if you did not use  fit_generator, how can you load 1300k images to memory? Thanks!</p>",
          "rawMarkdown": "Hi Liu,\nThanks for your sharing, it's very encouraging. I wonder if you did not use  fit_generator, how can you load 1300k images to memory? Thanks!"
        },
        {
          "id": 341651,
          "postDate": "2018-06-12T03:03:38.347Z",
          "content": "<p>Hi Jilong, is your NN model end2end training? Did you use keras concatenate four kinds feature layers? \nThx for your sharing.</p>",
          "rawMarkdown": "Hi Jilong, is your NN model end2end training? Did you use keras concatenate four kinds feature layers? \nThx for your sharing."
        },
        {
          "id": 342079,
          "postDate": "2018-06-12T19:46:52.283Z",
          "content": "<blockquote>\n  <p>I also tried fit_generator, but fit_generator's result is worse than fit, which is strange and I did not find the reason. As for time-consuming problem, you can try fit_generator's multi-processing parameter.</p>\n</blockquote>\n\n<p>Fit automatically does a few things that are critical for NNs ( such as randomizing input per epoch) that you have to explicitly do with your own fit_generator. Fit_generator automatically should load info on a new thread- so unless you have something crazy expensive in your generator (like a garbage collecting call) it shouldn't be a bottleneck at all-loading is faster than training.</p>",
          "rawMarkdown": "&gt; I also tried fit_generator, but fit_generator's result is worse than fit, which is strange and I did not find the reason. As for time-consuming problem, you can try fit_generator's multi-processing parameter.\n\nFit automatically does a few things that are critical for NNs ( such as randomizing input per epoch) that you have to explicitly do with your own fit_generator. Fit_generator automatically should load info on a new thread- so unless you have something crazy expensive in your generator (like a garbage collecting call) it shouldn't be a bottleneck at all-loading is faster than training.\n\n",
          "votes": 3
        },
        {
          "id": 343689,
          "postDate": "2018-06-15T20:30:19.343Z",
          "content": "<p>I've had several issues with using the image data. I think one of the most major ones is that it turns out keras cannot do multiprocessing on windows so it isnt actually loading in images properly beforehand. Second hand is I am trying to do the image resize and padding in line. This ends up costing roughly .05 per image. I worked a little around this using concurrent. Futures, but a full epoch still takes much much more than 15 minutes. Closer to 2-3 hours.. One avenue I spent some time researching is rewriting the images as a hdf5 dataset, but I'm unsure how to load it in alongside my other features, but that has various issues as well regarding shuffling and getting everything aligned</p>",
          "rawMarkdown": "I've had several issues with using the image data. I think one of the most major ones is that it turns out keras cannot do multiprocessing on windows so it isnt actually loading in images properly beforehand. Second hand is I am trying to do the image resize and padding in line. This ends up costing roughly .05 per image. I worked a little around this using concurrent. Futures, but a full epoch still takes much much more than 15 minutes. Closer to 2-3 hours.. One avenue I spent some time researching is rewriting the images as a hdf5 dataset, but I'm unsure how to load it in alongside my other features, but that has various issues as well regarding shuffling and getting everything aligned"
        },
        {
          "id": 344203,
          "postDate": "2018-06-17T07:59:08.747Z",
          "content": "<p>Train with image pixels end to end is indeed tough, considering the million count. I did some trick to make it works, also i encountered some strange questions. I will share my workflow after competition ends and hoping for answers.</p>",
          "rawMarkdown": "Train with image pixels end to end is indeed tough, considering the million count. I did some trick to make it works, also i encountered some strange questions. I will share my workflow after competition ends and hoping for answers.",
          "votes": 1
        },
        {
          "id": 344290,
          "postDate": "2018-06-17T14:33:53.243Z",
          "content": "<p>thank you!  I am looking forward to your model</p>",
          "rawMarkdown": "thank you!  I am looking forward to your model"
        },
        {
          "id": 344749,
          "postDate": "2018-06-18T16:10:00.313Z",
          "content": "<blockquote>\n  <p>I think one of the most major ones is that it turns out Keras cannot do multiprocessing on Windows</p>\n</blockquote>\n\n<p>This is interesting. Can you do your own threading process? maybe something with threading and queue like:</p>\n\n<pre><code>        def data_loader(q,):\n            for start in range(0, len(train_imgs), batch_size):\n                #do your batch creation code here               \n                q.put([feat_batch,  y_batch])\n        q = queue.Queue(maxsize=q_size)\n        t1 = threading.Thread(target=data_loader, name='DataLoader', args=(q,))\n        t1.start()\n        t1.join()\n</code></pre>\n\n<p>And then just read the generator batches from the queue? I can't really check ( I'm running Linux atm).\nIf that's not possible, I recommend Linux...</p>",
          "rawMarkdown": "&gt; I think one of the most major ones is that it turns out Keras cannot do multiprocessing on Windows\n\nThis is interesting. Can you do your own threading process? maybe something with threading and queue like:\n\n            def data_loader(q,):\n                for start in range(0, len(train_imgs), batch_size):\n                    #do your batch creation code here               \n                    q.put([feat_batch,  y_batch])\n            q = queue.Queue(maxsize=q_size)\n            t1 = threading.Thread(target=data_loader, name='DataLoader', args=(q,))\n            t1.start()\n            t1.join()\n\n\nAnd then just read the generator batches from the queue? I can't really check ( I'm running Linux atm).\nIf that's not possible, I recommend Linux...",
          "votes": 2
        },
        {
          "id": 344921,
          "postDate": "2018-06-18T23:26:44.383Z",
          "content": "<p>I really should move over to linux at some point, but I am eventually planning on dedicating one system to work like this and the other for games and less demanding functions so dont want to dual boot right away. That being said, I have found some way to kind of work around this using the concurrent.futures builtin python to load in the images with threading and I also preprocessed the images with padding and downscaling so I dont have to do that function within the generator. </p>",
          "rawMarkdown": "I really should move over to linux at some point, but I am eventually planning on dedicating one system to work like this and the other for games and less demanding functions so dont want to dual boot right away. That being said, I have found some way to kind of work around this using the concurrent.futures builtin python to load in the images with threading and I also preprocessed the images with padding and downscaling so I dont have to do that function within the generator. ",
          "votes": 1
        },
        {
          "id": 345074,
          "postDate": "2018-06-19T06:41:18.443Z",
          "content": "<p>Hi Jilong, In my efforts to include text + image + tabular features, I find that it heavily overfits even before it finishes one epoch, does heavy use of batch normalization help to tackle overfitting? Thanks</p>",
          "rawMarkdown": "Hi Jilong, In my efforts to include text + image + tabular features, I find that it heavily overfits even before it finishes one epoch, does heavy use of batch normalization help to tackle overfitting? Thanks"
        },
        {
          "id": 345152,
          "postDate": "2018-06-19T09:28:40.913Z",
          "content": "<p>Yes, i heavily use BN and dropout, you may try add a BN before every Dense Layer.</p>",
          "rawMarkdown": "Yes, i heavily use BN and dropout, you may try add a BN before every Dense Layer.",
          "votes": 1
        },
        {
          "id": 345417,
          "postDate": "2018-06-19T20:57:11.850Z",
          "content": "<p>@Liu Jilong, there are few rows in test set where images are missing, how are you handling those when you say you are using imagent features?</p>",
          "rawMarkdown": "@Liu Jilong, there are few rows in test set where images are missing, how are you handling those when you say you are using imagent features?\n"
        },
        {
          "id": 345512,
          "postDate": "2018-06-20T01:14:15.283Z",
          "content": "<p>I just give them an all-black image for row image pixels and zero score for image net scores. It may have some approaches to improve, but i do not have enough time to handle this.</p>",
          "rawMarkdown": "I just give them an all-black image for row image pixels and zero score for image net scores. It may have some approaches to improve, but i do not have enough time to handle this.",
          "votes": 1
        },
        {
          "id": 346569,
          "postDate": "2018-06-22T01:17:55.993Z",
          "content": "<p>Have you tried training a model purely on images? Do you know roughly what training or validation loss that can get down to? I have been working on this for a while and cant seem to get much out of the images. </p>",
          "rawMarkdown": "Have you tried training a model purely on images? Do you know roughly what training or validation loss that can get down to? I have been working on this for a while and cant seem to get much out of the images. "
        },
        {
          "id": 346622,
          "postDate": "2018-06-22T04:09:39.683Z",
          "content": "<p>I tried several CNN models to learn only on images to predict deal probably. I used relu on last several dense layer, and I found out last dense layer (before Dense (1)) predict all zeros, which give me 0.133  rmse loss...</p>",
          "rawMarkdown": "I tried several CNN models to learn only on images to predict deal probably. I used relu on last several dense layer, and I found out last dense layer (before Dense (1)) predict all zeros, which give me 0.133  rmse loss...",
          "votes": 1
        },
        {
          "id": 346640,
          "postDate": "2018-06-22T04:56:09.560Z",
          "content": "<p>Mine are consistently getting to .066ish MSE but dont seem to contribute anything to my combined model. Havent inspected what the outputs or the weights areFor a while I thought there was something wrong with my generator but it seems to work fine without the images just fine so dont think that is the issues</p>",
          "rawMarkdown": "Mine are consistently getting to .066ish MSE but dont seem to contribute anything to my combined model. Havent inspected what the outputs or the weights areFor a while I thought there was something wrong with my generator but it seems to work fine without the images just fine so dont think that is the issues"
        }
      ]
    },
    {
      "id": 338921,
      "postDate": "2018-06-06T01:33:30.123Z",
      "content": "<p>I'm hitting .2190 with a 5-fold averaged LGBM</p>",
      "rawMarkdown": "I'm hitting .2190 with a 5-fold averaged LGBM",
      "votes": 9,
      "replies": [
        {
          "id": 338925,
          "postDate": "2018-06-06T01:49:48.870Z",
          "content": "<p>That's impressive, Congratulation!</p>",
          "rawMarkdown": "That's impressive, Congratulation!",
          "votes": 1
        },
        {
          "id": 338928,
          "postDate": "2018-06-06T01:55:08.110Z",
          "content": "<p>Wow,you must find some amazing features~</p>",
          "rawMarkdown": "Wow,you must find some amazing features~",
          "votes": 1
        },
        {
          "id": 338929,
          "postDate": "2018-06-06T01:57:05.233Z",
          "content": "<p>Seems like there is a magic feature for the big jump.</p>",
          "rawMarkdown": "Seems like there is a magic feature for the big jump.",
          "votes": 1
        },
        {
          "id": 338930,
          "postDate": "2018-06-06T02:09:09.757Z",
          "content": "<p>Thanks! :)</p>\n\n<p>I'm using 100+ engineered tabular features, but there are definitely some that have a particularly substantial contribution.</p>",
          "rawMarkdown": "Thanks! :)\n\nI'm using 100+ engineered tabular features, but there are definitely some that have a particularly substantial contribution.",
          "votes": 1
        },
        {
          "id": 338939,
          "postDate": "2018-06-06T02:34:12.020Z",
          "content": "<p>good job,did you use target encoding?</p>",
          "rawMarkdown": "good job,did you use target encoding?",
          "votes": 1
        },
        {
          "id": 339017,
          "postDate": "2018-06-06T06:43:56.893Z",
          "content": "<p>Wow, do you mean you didn't use text features at all?</p>",
          "rawMarkdown": "Wow, do you mean you didn't use text features at all?"
        },
        {
          "id": 339162,
          "postDate": "2018-06-06T12:37:55.033Z",
          "content": "<p>Sorry, I do use text features in addition to the tabular ones - could have been more clear</p>",
          "rawMarkdown": "Sorry, I do use text features in addition to the tabular ones - could have been more clear",
          "votes": 3
        },
        {
          "id": 339181,
          "postDate": "2018-06-06T13:13:56.817Z",
          "content": "<p>Thank you a lot for a clarification!</p>",
          "rawMarkdown": "Thank you a lot for a clarification!",
          "votes": 2
        },
        {
          "id": 340151,
          "postDate": "2018-06-08T13:24:03.263Z",
          "content": "<p>great work! did you use the train_active csv ?</p>",
          "rawMarkdown": "great work! did you use the train_active csv ?"
        }
      ]
    },
    {
      "id": 333988,
      "postDate": "2018-05-26T08:13:34.680Z",
      "content": "<p>LGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2200 (updated)</p>\n\n<p>XGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2216 (updated)</p>\n\n<p>NN (Categorical, Numerical, W2V, Images meta-data + ImageNet scoring): 0.2198 (updated)</p>\n\n<p>All with CV5.</p>",
      "rawMarkdown": "LGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2200 (updated)\n\nXGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2216 (updated)\n\nNN (Categorical, Numerical, W2V, Images meta-data + ImageNet scoring): 0.2198 (updated)\n\nAll with CV5.",
      "votes": 8,
      "replies": [
        {
          "id": 334098,
          "postDate": "2018-05-26T14:27:42.823Z",
          "content": "<p>may i ask what is image meta?</p>",
          "rawMarkdown": "may i ask what is image meta?"
        },
        {
          "id": 334132,
          "postDate": "2018-05-26T16:01:46.743Z",
          "content": "<p>What numerical features are working for you?</p>",
          "rawMarkdown": "What numerical features are working for you?"
        },
        {
          "id": 334182,
          "postDate": "2018-05-26T17:42:30.647Z",
          "content": "<p>Images meta-data is width, height. I've also included the average of VGG16, VGG19, InceptionV3, Xception, ResNet50 ImageNet scoring for each image but it does not really help. It only improves by 0.0002. And for numerical, only the basic ones, log1p(item_seq_number) and log1p(price). I'm trying different target-encoding but currently, as already reported below, it overfits too much.</p>",
          "rawMarkdown": "Images meta-data is width, height. I've also included the average of VGG16, VGG19, InceptionV3, Xception, ResNet50 ImageNet scoring for each image but it does not really help. It only improves by 0.0002. And for numerical, only the basic ones, log1p(item_seq_number) and log1p(price). I'm trying different target-encoding but currently, as already reported below, it overfits too much.",
          "votes": 8
        },
        {
          "id": 336600,
          "postDate": "2018-06-01T02:28:34.793Z",
          "content": "<p>does NIMA stand for Neural Image Assessment?</p>",
          "rawMarkdown": "does NIMA stand for Neural Image Assessment?"
        },
        {
          "id": 336711,
          "postDate": "2018-06-01T06:40:25.510Z",
          "content": "<p>Yes, I used this implementation:\n<a href=\"https://github.com/titu1994/neural-image-assessment\">https://github.com/titu1994/neural-image-assessment</a></p>\n\n<p>It improves LGB LB by 0.0002 and XGB by 0.0003 (only). I included mean/std as features from MobileNet and InceptionResNetv2. It does not improve NN (which should mean the NLP part is better).</p>",
          "rawMarkdown": "Yes, I used this implementation:\nhttps://github.com/titu1994/neural-image-assessment\n\nIt improves LGB LB by 0.0002 and XGB by 0.0003 (only). I included mean/std as features from MobileNet and InceptionResNetv2. It does not improve NN (which should mean the NLP part is better)."
        }
      ]
    },
    {
      "id": 344107,
      "postDate": "2018-06-17T02:10:41.900Z",
      "content": "<p>LGB 10-fold 0.2179</p>",
      "rawMarkdown": "LGB 10-fold 0.2179",
      "votes": 5,
      "replies": [
        {
          "id": 344108,
          "postDate": "2018-06-17T02:14:32.873Z",
          "content": "<p>Excellent achieve！Can you show some black magic you used？</p>",
          "rawMarkdown": "Excellent achieve！Can you show some black magic you used？"
        },
        {
          "id": 344128,
          "postDate": "2018-06-17T03:33:02.557Z",
          "content": "<p>no magic,features are almost from kernels like aggregations,images,tfidf,tsvd,impute...just try to use more features.</p>",
          "rawMarkdown": "no magic,features are almost from kernels like aggregations,images,tfidf,tsvd,impute...just try to use more features.",
          "votes": 1
        },
        {
          "id": 344129,
          "postDate": "2018-06-17T03:36:12.897Z",
          "content": "<p><a href=\"/senkin13\">@senkin13</a>: Is that a CV score or a LB score? Does your model contain a Ridge?</p>",
          "rawMarkdown": "@senkin13: Is that a CV score or a LB score? Does your model contain a Ridge?"
        },
        {
          "id": 344130,
          "postDate": "2018-06-17T03:39:13.767Z",
          "content": "<p>LB score without ridge,I prefer to use ridge for stacking instead of lgb features.</p>",
          "rawMarkdown": "LB score without ridge,I prefer to use ridge for stacking instead of lgb features.",
          "votes": 1
        },
        {
          "id": 344132,
          "postDate": "2018-06-17T03:52:27.573Z",
          "content": "<p>What do you mean by tsvd and impute? Any reference kernels or information?</p>",
          "rawMarkdown": "What do you mean by tsvd and impute? Any reference kernels or information?"
        },
        {
          "id": 344134,
          "postDate": "2018-06-17T04:07:04.590Z",
          "content": "<p>here: <a href=\"https://www.kaggle.com/krithi07/baseline-model-with-new-features\">https://www.kaggle.com/krithi07/baseline-model-with-new-features</a></p>",
          "rawMarkdown": "here: https://www.kaggle.com/krithi07/baseline-model-with-new-features",
          "votes": 4
        },
        {
          "id": 344146,
          "postDate": "2018-06-17T04:52:48.563Z",
          "content": "<p>Congrats @Senkin13, which image features did you use. The ones I've got makes my scores worse... both CV and LB.</p>",
          "rawMarkdown": "Congrats @Senkin13, which image features did you use. The ones I've got makes my scores worse... both CV and LB."
        },
        {
          "id": 344172,
          "postDate": "2018-06-17T06:21:42.553Z",
          "content": "<p>A few simple image meta features,not too much improvement compared to Peter Hurford(0.002 improvment)</p>",
          "rawMarkdown": "A few simple image meta features,not too much improvement compared to Peter Hurford(0.002 improvment)"
        },
        {
          "id": 344178,
          "postDate": "2018-06-17T06:47:15.090Z",
          "content": "<p>I know you called my improvement from images impressive, but I'm not coming anywhere close to 0.2179 single model LGB... Hey, I wonder what would happen if we teamed up and pooled our feature engineering?</p>",
          "rawMarkdown": "I know you called my improvement from images impressive, but I'm not coming anywhere close to 0.2179 single model LGB... Hey, I wonder what would happen if we teamed up and pooled our feature engineering?"
        },
        {
          "id": 344179,
          "postDate": "2018-06-17T06:52:11.893Z",
          "content": "<p>Hello, I have tried tsvd to tfidf of description and title,but my score became to worse,I choose 32 component tsvd.</p>",
          "rawMarkdown": "Hello, I have tried tsvd to tfidf of description and title,but my score became to worse,I choose 32 component tsvd.",
          "votes": 1
        },
        {
          "id": 344185,
          "postDate": "2018-06-17T07:13:08.823Z",
          "content": "<p>Maybe I dont choose the right n_ component </p>",
          "rawMarkdown": "Maybe I dont choose the right n_ component "
        },
        {
          "id": 344187,
          "postDate": "2018-06-17T07:14:05.520Z",
          "content": "<p>@Peter Hurford Your proposal is attractive ,there still have 4 days ,let's wait and see.</p>",
          "rawMarkdown": "@Peter Hurford Your proposal is attractive ,there still have 4 days ,let's wait and see."
        },
        {
          "id": 344188,
          "postDate": "2018-06-17T07:15:09.567Z",
          "content": "<p>@Johnny You can always check sum of explained_variance_ratio_. If it is very low, it means you not only lose noise but also information</p>",
          "rawMarkdown": "@Johnny You can always check sum of explained_variance_ratio_. If it is very low, it means you not only lose noise but also information",
          "votes": 1
        },
        {
          "id": 344189,
          "postDate": "2018-06-17T07:16:47.913Z",
          "content": "<p>@Johnny Liu I only get a little improvement from tsvd,yes should try different n_component and tfidf parameters.</p>",
          "rawMarkdown": "@Johnny Liu I only get a little improvement from tsvd,yes should try different n_component and tfidf parameters.",
          "votes": 1
        },
        {
          "id": 344191,
          "postDate": "2018-06-17T07:18:52.707Z",
          "content": "<p>very solid work! </p>",
          "rawMarkdown": "very solid work! "
        },
        {
          "id": 344201,
          "postDate": "2018-06-17T07:47:13.687Z",
          "content": "<p>thanks @AhmetErdem <a href=\"/senkin13\">@senkin13</a>, now I sum explained_variance_ratio_ to find the best component</p>",
          "rawMarkdown": "thanks @AhmetErdem @senkin13, now I sum explained_variance_ratio_ to find the best component\n\n"
        }
      ]
    },
    {
      "id": 338049,
      "postDate": "2018-06-04T10:21:58.890Z",
      "content": "<p>LGBM 10-fold avg. - 0.2203 (categorical, numerical, target encoding, text and features predicted using all csv data)</p>",
      "rawMarkdown": "LGBM 10-fold avg. - 0.2203 (categorical, numerical, target encoding, text and features predicted using all csv data)",
      "votes": 5,
      "replies": [
        {
          "id": 338078,
          "postDate": "2018-06-04T11:43:39.543Z",
          "content": "<p>Hello, does  target encoding contribute your  score a lot?</p>",
          "rawMarkdown": "Hello, does  target encoding contribute your  score a lot?"
        },
        {
          "id": 338083,
          "postDate": "2018-06-04T12:00:17.947Z",
          "content": "<p>When you say \"features predicted using all csv data\", does that mean that you found some interesting features leveraging train_active/test_active/periods_train/periods_test beyond the ones identified by Benjamin Minixhofer?</p>",
          "rawMarkdown": "When you say \"features predicted using all csv data\", does that mean that you found some interesting features leveraging train_active/test_active/periods_train/periods_test beyond the ones identified by Benjamin Minixhofer?",
          "votes": 1
        },
        {
          "id": 338100,
          "postDate": "2018-06-04T12:28:06.580Z",
          "content": "<p>Previous lgbm was at 0.2215 but TE was not the only features added so it is hard to say how much TE contributed to score. Still half of features in top-100 are TEs. </p>",
          "rawMarkdown": "Previous lgbm was at 0.2215 but TE was not the only features added so it is hard to say how much TE contributed to score. Still half of features in top-100 are TEs. ",
          "votes": 2
        },
        {
          "id": 338104,
          "postDate": "2018-06-04T12:32:30.757Z",
          "content": "<p>@Larry Currently I am using only 'times up' and 'days up' predictions, but I think it is worth trying to fill other missing values by predicting them from all csv data.</p>",
          "rawMarkdown": "@Larry Currently I am using only 'times up' and 'days up' predictions, but I think it is worth trying to fill other missing values by predicting them from all csv data."
        },
        {
          "id": 338111,
          "postDate": "2018-06-04T12:48:16.607Z",
          "content": "<p>@Andrii Do you mean you trained a model to estimate 'times up' and 'days up' from the other features? Or you just used the features described in Benjamin's kernel?</p>",
          "rawMarkdown": "@Andrii Do you mean you trained a model to estimate 'times up' and 'days up' from the other features? Or you just used the features described in Benjamin's kernel?"
        },
        {
          "id": 338116,
          "postDate": "2018-06-04T12:59:20.747Z",
          "content": "<p>@Dmitriy I am using both - averages from kernel and model predictions.</p>",
          "rawMarkdown": "@Dmitriy I am using both - averages from kernel and model predictions.",
          "votes": 2
        },
        {
          "id": 338427,
          "postDate": "2018-06-05T03:24:17.920Z",
          "content": "<p>wow, thanks for your replying.I have not try to using target encoding now</p>",
          "rawMarkdown": "wow, thanks for your replying.I have not try to using target encoding now"
        }
      ]
    },
    {
      "id": 333337,
      "postDate": "2018-05-24T23:45:27.317Z",
      "content": "<p>NN: <br>\n0.2218 Fasttext <br>\n0.2215 Self-training embedding  </p>\n\n<p>Thanks @Dieter</p>",
      "rawMarkdown": "NN:  \n0.2218 Fasttext  \n0.2215 Self-training embedding  \n\nThanks @Dieter",
      "votes": 5,
      "replies": [
        {
          "id": 333362,
          "postDate": "2018-05-25T02:50:23.843Z",
          "content": "<p>Wow your fasttext result is great. Are u using the crawl one or wiki one? Did you play a lot with different NN structures?</p>",
          "rawMarkdown": "Wow your fasttext result is great. Are u using the crawl one or wiki one? Did you play a lot with different NN structures?"
        },
        {
          "id": 333366,
          "postDate": "2018-05-25T03:09:20.937Z",
          "content": "<p>I use wiki 300D. <br>\nI tried some NLP parts such as: RNN, Bi-GRU, Bi-LSTM, CNN, ... (with Attention and without Attention).</p>",
          "rawMarkdown": "I use wiki 300D.  \nI tried some NLP parts such as: RNN, Bi-GRU, Bi-LSTM, CNN, ... (with Attention and without Attention)."
        },
        {
          "id": 333367,
          "postDate": "2018-05-25T03:16:23.640Z",
          "content": "<p>@Ghost can you tell us how much do you think contributes the NLP part to your score? Let's say you take a naive approach with a single gru layer, how much would that worsen your result?</p>",
          "rawMarkdown": "@Ghost can you tell us how much do you think contributes the NLP part to your score? Let's say you take a naive approach with a single gru layer, how much would that worsen your result?"
        },
        {
          "id": 333372,
          "postDate": "2018-05-25T03:35:58.927Z",
          "content": "<p>@Dieter <br>\nI am not sure that I understand your question correctly. Please feel free to ask. <br>\nWithout NLP, my model is really bad. It is arround 0.2240. Then I tried to add NLP.</p>\n\n<p>I have not ran all of the NLP parts above, so I don't have results of each part. <br>\nTo determine which NLP part is good or not. I have a training plan as follows:  </p>\n\n<ol>\n<li>Select fixed training set and validation set. Ex: First fold of 5-Fold split</li>\n<li>Replace NLP part and run for some epochs (Ex: 5 epochs)  </li>\n</ol>\n\n<p>You will have list of cross-validation of model in different NLP architectures (GRU, CNN, LSTM,...).\nSelect the best one and go fully training.  </p>\n\n<p>My top 3 best NLP parts are:\nCapsule --&gt; CNN --&gt; Bi-GRU --&gt; ... </p>",
          "rawMarkdown": "@Dieter  \nI am not sure that I understand your question correctly. Please feel free to ask.  \nWithout NLP, my model is really bad. It is arround 0.2240. Then I tried to add NLP.\n\nI have not ran all of the NLP parts above, so I don't have results of each part.  \nTo determine which NLP part is good or not. I have a training plan as follows:  \n\n 1. Select fixed training set and validation set. Ex: First fold of 5-Fold split\n 2. Replace NLP part and run for some epochs (Ex: 5 epochs)  \n\nYou will have list of cross-validation of model in different NLP architectures (GRU, CNN, LSTM,...).\nSelect the best one and go fully training.  \n\nMy top 3 best NLP parts are:\nCapsule --&gt; CNN --&gt; Bi-GRU --&gt; ... \n\n",
          "votes": 3
        },
        {
          "id": 333377,
          "postDate": "2018-05-25T03:52:36.307Z",
          "content": "<blockquote>\n  <p>Without NLP</p>\n</blockquote>\n\n<p>you mean skipping text completely?</p>",
          "rawMarkdown": "&gt; Without NLP\n\nyou mean skipping text completely?"
        },
        {
          "id": 333379,
          "postDate": "2018-05-25T03:56:39.793Z",
          "content": "<p>Yeap. <br>\nWithout NLP, I only have numeric and categorical features.</p>",
          "rawMarkdown": "Yeap.  \nWithout NLP, I only have numeric and categorical features."
        },
        {
          "id": 333380,
          "postDate": "2018-05-25T03:58:59.237Z",
          "content": "<p>Thank you, that helps :) </p>",
          "rawMarkdown": "Thank you, that helps :) "
        },
        {
          "id": 333499,
          "postDate": "2018-05-25T09:17:35.427Z",
          "content": "<p>0.2218 is this a single model or average of 5 fold?</p>",
          "rawMarkdown": "0.2218 is this a single model or average of 5 fold?"
        },
        {
          "id": 334030,
          "postDate": "2018-05-26T10:47:38.253Z",
          "content": "<p>@Ghost you must be very talented. You joined kaggle 2 days ago, and already have models that score top 50 in LB. I can only salute.</p>",
          "rawMarkdown": "@Ghost you must be very talented. You joined kaggle 2 days ago, and already have models that score top 50 in LB. I can only salute.",
          "votes": 1
        },
        {
          "id": 334099,
          "postDate": "2018-05-26T14:32:58.730Z",
          "content": "<p>0.2240 without Text is a strong model I think. \nDid you use image meta features in this?</p>",
          "rawMarkdown": "0.2240 without Text is a strong model I think. \nDid you use image meta features in this?"
        },
        {
          "id": 334592,
          "postDate": "2018-05-27T23:48:41.247Z",
          "content": "<p>Sorry for the late reply. I have not tried with image meta-features. The result is after CV.</p>",
          "rawMarkdown": "Sorry for the late reply. I have not tried with image meta-features. The result is after CV."
        }
      ]
    },
    {
      "id": 329320,
      "postDate": "2018-05-16T08:25:08.463Z",
      "content": "<p>I am working exclusively with NN for now..</p>\n\n<p>So far,  best single NN model : LB = 0.2246 ( 5 folds CV RMSE : 0.2215xxx ), no image features either</p>\n\n<p>The model was scoring 0.244xxx on LB when I started it ,   5 days ago...(Some ideas from kernels and discussion topics helped me in the improvement) </p>\n\n<p><strong>Edit :</strong>  The NN model is now scoring 0.2223 on LB </p>",
      "rawMarkdown": "I am working exclusively with NN for now..\n\nSo far,  best single NN model : LB = 0.2246 ( 5 folds CV RMSE : 0.2215xxx ), no image features either\n\nThe model was scoring 0.244xxx on LB when I started it ,   5 days ago...(Some ideas from kernels and discussion topics helped me in the improvement) \n\n**Edit :**  The NN model is now scoring 0.2223 on LB ",
      "votes": 6,
      "replies": [
        {
          "id": 330506,
          "postDate": "2018-05-19T01:47:12.347Z",
          "content": "<p>Nice. <br>\nMy model now is 0.2222 LB. Can I know which pretrained embedding are you using?</p>",
          "rawMarkdown": "Nice.  \nMy model now is 0.2222 LB. Can I know which pretrained embedding are you using?"
        },
        {
          "id": 330560,
          "postDate": "2018-05-19T05:30:45.840Z",
          "content": "<p>Thanks :)</p>\n\n<p>I am using self-trained embeddings (following the idea suggested by Dieter )</p>",
          "rawMarkdown": "Thanks :)\n\nI am using self-trained embeddings (following the idea suggested by Dieter )"
        },
        {
          "id": 330561,
          "postDate": "2018-05-19T05:32:38.377Z",
          "content": "<p>Amazing guy. <br>\nI could not get the high result with this embedding :D</p>",
          "rawMarkdown": "Amazing guy.  \nI could not get the high result with this embedding :D"
        },
        {
          "id": 330562,
          "postDate": "2018-05-19T05:34:16.830Z",
          "content": "<p>@ Totoro - Did you try both options ( trainable = True and Trainable = False )??</p>",
          "rawMarkdown": "@ Totoro - Did you try both options ( trainable = True and Trainable = False )??"
        },
        {
          "id": 330567,
          "postDate": "2018-05-19T05:51:58.893Z",
          "content": "<p>I use pretrained embedding, so trainble = False. I tried trainable = True, but it is worse for me</p>",
          "rawMarkdown": "I use pretrained embedding, so trainble = False. I tried trainable = True, but it is worse for me"
        },
        {
          "id": 330568,
          "postDate": "2018-05-19T05:54:44.660Z",
          "content": "<p>If it's any consolation...the biggest improvements in my model  are not from the embedding itself :)</p>",
          "rawMarkdown": "If it's any consolation...the biggest improvements in my model  are not from the embedding itself :)"
        },
        {
          "id": 330571,
          "postDate": "2018-05-19T05:59:16.977Z",
          "content": "<p>Interesting <br>\nHope to see your model when the competition ends :D</p>",
          "rawMarkdown": "Interesting  \nHope to see your model when the competition ends :D"
        },
        {
          "id": 334014,
          "postDate": "2018-05-26T09:58:11.267Z",
          "content": "<p>Thank you</p>",
          "rawMarkdown": "Thank you"
        }
      ]
    },
    {
      "id": 342176,
      "postDate": "2018-06-13T03:26:59.310Z",
      "content": "<p>Single model LightGBM 0.2199, with 5seed-avg 0.2195</p>",
      "rawMarkdown": "Single model LightGBM 0.2199, with 5seed-avg 0.2195",
      "votes": 3,
      "replies": [
        {
          "id": 343951,
          "postDate": "2018-06-16T15:11:08.827Z",
          "content": "<p>could you show some hints about how you achieve this?</p>",
          "rawMarkdown": "could you show some hints about how you achieve this?\n "
        },
        {
          "id": 345827,
          "postDate": "2018-06-20T14:44:08.317Z",
          "content": "<p>Sure, it will be in my write up if I get gold :)</p>",
          "rawMarkdown": "Sure, it will be in my write up if I get gold :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 339507,
      "postDate": "2018-06-07T04:21:41.537Z",
      "content": "<p>Finally! Able to achieve 2210-XGB and 2217-LGB. No image features though.\nI was stuck earlier and started from scratch, seems the right decision. Best of luck guys.</p>",
      "rawMarkdown": "Finally! Able to achieve 2210-XGB and 2217-LGB. No image features though.\nI was stuck earlier and started from scratch, seems the right decision. Best of luck guys.",
      "votes": 3,
      "replies": [
        {
          "id": 339528,
          "postDate": "2018-06-07T04:58:37.717Z",
          "content": "<p>Wow,Congratulation! Please sharing your solution after competition,I'm every interested how to achieve the score.</p>",
          "rawMarkdown": "Wow,Congratulation! Please sharing your solution after competition,I'm every interested how to achieve the score.",
          "votes": 3
        },
        {
          "id": 339547,
          "postDate": "2018-06-07T05:56:38.403Z",
          "content": "<p>Is your XGB model using the same features as LGBM?  I am finding that my LGBM models are both faster and perform better.  I am able to get to 2217-LGB but not even 2217-XGB.  </p>",
          "rawMarkdown": "Is your XGB model using the same features as LGBM?  I am finding that my LGBM models are both faster and perform better.  I am able to get to 2217-LGB but not even 2217-XGB.  ",
          "votes": 2
        },
        {
          "id": 339568,
          "postDate": "2018-06-07T06:27:00.560Z",
          "content": "<p>Congratulations @Nooh on getting those scores. Any hints you can share?</p>\n\n<p>@Larry Freeman, I have the same issue with you. My best LGBM is at 0.2215 but xgb is 0.2232. Both with single runs. The speed difference is not surprising as LGBM is always faster than xgb but in my experience their performance is comparable. Hence my surprise that I am not able to get xgb to  match lgb performance here.</p>",
          "rawMarkdown": "Congratulations @Nooh on getting those scores. Any hints you can share?\n\n@Larry Freeman, I have the same issue with you. My best LGBM is at 0.2215 but xgb is 0.2232. Both with single runs. The speed difference is not surprising as LGBM is always faster than xgb but in my experience their performance is comparable. Hence my surprise that I am not able to get xgb to  match lgb performance here.",
          "votes": 2
        },
        {
          "id": 339648,
          "postDate": "2018-06-07T09:18:25.063Z",
          "content": "<p>Thanks Liu, though I'm really bad in documentation and my code is a mess right now but will surely post a brief if I don't end up over-fitting. </p>",
          "rawMarkdown": "Thanks Liu, though I'm really bad in documentation and my code is a mess right now but will surely post a brief if I don't end up over-fitting. ",
          "votes": 1
        },
        {
          "id": 339649,
          "postDate": "2018-06-07T09:20:37.430Z",
          "content": "<p>Yes Larry, I can say XGB has 98% same features as of my LGB. \nMy XGB took around 15-hours to train. :/</p>",
          "rawMarkdown": "Yes Larry, I can say XGB has 98% same features as of my LGB. \nMy XGB took around 15-hours to train. :/",
          "votes": 2
        },
        {
          "id": 339651,
          "postDate": "2018-06-07T09:28:44.263Z",
          "content": "<p>Thanks a lot YaGana Sheriff-Hussaini!</p>\n\n<p>Nothing special in features though. I'm using nearly all the features used by <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">this kernel</a>. Played with some statistical features around deal_prob and Price. And another tip, try to thoroughly analyze image_top_1 and item_seq_no; I handcrafted something using them. After all that I tuned the hyperparams.</p>\n\n<p>Best of luck buddy!</p>",
          "rawMarkdown": "Thanks a lot YaGana Sheriff-Hussaini!\n\nNothing special in features though. I'm using nearly all the features used by [this kernel][1]. Played with some statistical features around deal_prob and Price. And another tip, try to thoroughly analyze image_top_1 and item_seq_no; I handcrafted something using them. After all that I tuned the hyperparams.\n\nBest of luck buddy!\n\n\n  [1]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm",
          "votes": 7
        },
        {
          "id": 339863,
          "postDate": "2018-06-07T19:44:29.180Z",
          "content": "<p>Added OOF ridge in LGB and now hitting 2211. :) </p>",
          "rawMarkdown": "Added OOF ridge in LGB and now hitting 2211. :) ",
          "votes": 3
        },
        {
          "id": 340353,
          "postDate": "2018-06-09T01:29:08.370Z",
          "content": "<p>Thanks @Nooh and best of luck to you too.</p>\n\n<p>I just managed to get my lgb to LB=0.2212 but still have some optimization to do. You are right the xgb takes a long time for me as well, hence I am trying to optimize my features with LGB before running the xgb model again. Like you, I use the same features for  both.</p>",
          "rawMarkdown": "Thanks @Nooh and best of luck to you too.\n\nI just managed to get my lgb to LB=0.2212 but still have some optimization to do. You are right the xgb takes a long time for me as well, hence I am trying to optimize my features with LGB before running the xgb model again. Like you, I use the same features for  both.",
          "votes": 1
        },
        {
          "id": 340584,
          "postDate": "2018-06-09T17:43:15.437Z",
          "content": "<p>Okay, so I'm convinced now that text features are bringing in some improvements in my models. Text normalization seems promising in my case. I just did some very basic normalization on title+desc and it is progressive. I should quote the original author here i.e; ololo (who stood 5th in the previous avito competition). I just stumbled upon his code and there is a lot to scavenge from his repo <a href=\"https://github.com/alexeygrigorev/avito-duplicates-kaggle\">here</a> </p>\n\n<p>I used <a href=\"https://raw.githubusercontent.com/alexeygrigorev/avito-duplicates-kaggle/master/prepare_text_features.py\">this</a> script to get some text features. Hope this helps.</p>",
          "rawMarkdown": "Okay, so I'm convinced now that text features are bringing in some improvements in my models. Text normalization seems promising in my case. I just did some very basic normalization on title+desc and it is progressive. I should quote the original author here i.e; ololo (who stood 5th in the previous avito competition). I just stumbled upon his code and there is a lot to scavenge from his repo [here][1] \n\nI used [this][2] script to get some text features. Hope this helps.\n\n\n  [1]: https://github.com/alexeygrigorev/avito-duplicates-kaggle\n  [2]: https://raw.githubusercontent.com/alexeygrigorev/avito-duplicates-kaggle/master/prepare_text_features.py",
          "votes": 7
        },
        {
          "id": 340759,
          "postDate": "2018-06-10T08:21:03.113Z",
          "content": "<p>Thanks @Nooh. I have a simple text normalization but after reading your comment, I tried ololo's i.e. \"def normalize(text):\" and I get the following error:</p>\n\n<blockquote>\n  <p>File \"../src/script.py\", line 98\n      text = re.sub(ur'(?&lt;=[^а-я])(' + shortenings + ')[.]', ur'\\1 ', text)\n                                   ^\n  SyntaxError: invalid syntax</p>\n</blockquote>",
          "rawMarkdown": "Thanks @Nooh. I have a simple text normalization but after reading your comment, I tried ololo's i.e. \"def normalize(text):\" and I get the following error:\n\n&gt;   File \"../src/script.py\", line 98\n    text = re.sub(ur'(?&lt;=[^а-я])(' + shortenings + ')[.]', ur'\\1 ', text)\n                                 ^\nSyntaxError: invalid syntax",
          "votes": 2
        },
        {
          "id": 340775,
          "postDate": "2018-06-10T09:24:03.940Z",
          "content": "<p>Hi YaGana Sheriff-Hussaini. His code is for Python 2.x\nAre you on Python3x? If so, make sure you port it accordingly.\nI didn't used it as is actually instead get ideas from his code and made my own normalizations.</p>",
          "rawMarkdown": "Hi YaGana Sheriff-Hussaini. His code is for Python 2.x\nAre you on Python3x? If so, make sure you port it accordingly.\nI didn't used it as is actually instead get ideas from his code and made my own normalizations.",
          "votes": 2
        },
        {
          "id": 341220,
          "postDate": "2018-06-11T08:22:41.477Z",
          "content": "<p>@Nooh you are right Ololo's code is for python2.x. I am using python3 and wrote a better text normalization than the basic one I had. This one improved both my CV and LB scores.</p>",
          "rawMarkdown": "@Nooh you are right Ololo's code is for python2.x. I am using python3 and wrote a better text normalization than the basic one I had. This one improved both my CV and LB scores.",
          "votes": 1
        },
        {
          "id": 341293,
          "postDate": "2018-06-11T11:25:21.387Z",
          "content": "<p>Great bro!, happy to hear that. I'm making up some image meta features' code. I hope that will bring in some improvement. If so, I'll share the code+data. \nFingers crossed. :)</p>",
          "rawMarkdown": "Great bro!, happy to hear that. I'm making up some image meta features' code. I hope that will bring in some improvement. If so, I'll share the code+data. \nFingers crossed. :)",
          "votes": 1
        },
        {
          "id": 341351,
          "postDate": "2018-06-11T13:25:45.630Z",
          "content": "<p>I hope you get a better results with image features than I do. I have a weak machine hence using a different LGBM model with reduced features that includes image ones but only scores 0.21777 CV and 0.2233 LB.</p>",
          "rawMarkdown": "I hope you get a better results with image features than I do. I have a weak machine hence using a different LGBM model with reduced features that includes image ones but only scores 0.21777 CV and 0.2233 LB."
        },
        {
          "id": 341522,
          "postDate": "2018-06-11T18:12:12.117Z",
          "content": "<p>Hopefully yeah. I'm done with test set images and training set is running for 1 day now. I think it will take a couple of more hours.\nHow many models have you ensembled to get the current lb score you have? I have 4 of them, 3 lgbms and 1 xgb.\nTry training different lgbms with different features and different hyperparams. You can get all your features in that way i guess. </p>",
          "rawMarkdown": "Hopefully yeah. I'm done with test set images and training set is running for 1 day now. I think it will take a couple of more hours.\nHow many models have you ensembled to get the current lb score you have? I have 4 of them, 3 lgbms and 1 xgb.\nTry training different lgbms with different features and different hyperparams. You can get all your features in that way i guess. "
        },
        {
          "id": 341598,
          "postDate": "2018-06-11T21:21:31.863Z",
          "content": "<p>Six.</p>",
          "rawMarkdown": "Six."
        },
        {
          "id": 344448,
          "postDate": "2018-06-18T03:05:43.033Z",
          "content": "<p>which image meta features will you care for？</p>",
          "rawMarkdown": "which image meta features will you care for？"
        }
      ]
    },
    {
      "id": 337277,
      "postDate": "2018-06-02T12:40:10.413Z",
      "content": "<p>We are at</p>\n\n<ul>\n<li>LGB 0.2200 </li>\n<li>XGB 0.2213 </li>\n<li>Catboost 0.2208</li>\n<li>NN 0.2194</li>\n</ul>",
      "rawMarkdown": "We are at\n\n - LGB 0.2200 \n - XGB 0.2213 \n - Catboost 0.2208\n - NN 0.2194",
      "votes": 3,
      "replies": [
        {
          "id": 337303,
          "postDate": "2018-06-02T14:09:27.843Z",
          "content": "<p>Impressive @Dieter. Are your results for single or N-Folds. If so is it 3 or 5 folds?</p>\n\n<p>My best single model is LGBM with LB = 0.2217 single run. I have not run it for N-Folds yet.</p>",
          "rawMarkdown": "Impressive @Dieter. Are your results for single or N-Folds. If so is it 3 or 5 folds?\n\nMy best single model is LGBM with LB = 0.2217 single run. I have not run it for N-Folds yet.",
          "votes": 1
        },
        {
          "id": 337916,
          "postDate": "2018-06-04T03:47:50.800Z",
          "content": "<p>5 fold</p>",
          "rawMarkdown": "5 fold",
          "votes": 2
        },
        {
          "id": 338044,
          "postDate": "2018-06-04T10:04:21.810Z",
          "content": "<p>What features are you using for LGB .2206?  That is very impressive.   </p>",
          "rawMarkdown": "What features are you using for LGB .2206?  That is very impressive.   "
        },
        {
          "id": 338101,
          "postDate": "2018-06-04T12:29:12.633Z",
          "content": "<p>@Dieter, if you can share , what is the difference in score between a single run and 5-Fold?</p>",
          "rawMarkdown": "@Dieter, if you can share , what is the difference in score between a single run and 5-Fold?"
        },
        {
          "id": 338426,
          "postDate": "2018-06-05T03:15:44.950Z",
          "content": "<p>@Larry, I guess some part of the good score are image features</p>\n\n<p>@YaGana Sheriff-Hussaini difference is about 0.001</p>",
          "rawMarkdown": "@Larry, I guess some part of the good score are image features\n\n@YaGana Sheriff-Hussaini difference is about 0.001",
          "votes": 2
        },
        {
          "id": 342435,
          "postDate": "2018-06-13T14:10:32.643Z",
          "content": "<p>@Dieter, the difference for me is 0.0003</p>",
          "rawMarkdown": "@Dieter, the difference for me is 0.0003"
        }
      ]
    },
    {
      "id": 335359,
      "postDate": "2018-05-29T16:31:37.513Z",
      "content": "<p>My best so far is an LGB model with .2216 (recently improved) after a 10-fold average.</p>",
      "rawMarkdown": "My best so far is an LGB model with .2216 (recently improved) after a 10-fold average.",
      "votes": 3,
      "replies": [
        {
          "id": 335485,
          "postDate": "2018-05-29T20:05:05.627Z",
          "content": "<p>Good job! Your public score is 0.2195, how did you earn those 0.002 on your score ? with some blending ?\nSame thing @AlexTru there is a significant gap between your LB and this LGB model. I'd understand if you don't want to share...</p>",
          "rawMarkdown": "Good job! Your public score is 0.2195, how did you earn those 0.002 on your score ? with some blending ?\nSame thing @AlexTru there is a significant gap between your LB and this LGB model. I'd understand if you don't want to share..."
        },
        {
          "id": 335487,
          "postDate": "2018-05-29T20:09:31.143Z",
          "content": "<p>I used lessons learned from here:\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557</a></p>",
          "rawMarkdown": "I used lessons learned from here:\nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557",
          "votes": 2
        },
        {
          "id": 335510,
          "postDate": "2018-05-29T20:50:30.400Z",
          "content": "<p>Thanks for sharing this link !</p>",
          "rawMarkdown": "Thanks for sharing this link !"
        },
        {
          "id": 335518,
          "postDate": "2018-05-29T21:21:05.710Z",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> Are you using train/test-time augments? I assume you're not doing psuedo-labeling?</p>",
          "rawMarkdown": "@larryfreeman Are you using train/test-time augments? I assume you're not doing psuedo-labeling?"
        },
        {
          "id": 335542,
          "postDate": "2018-05-29T22:17:33.987Z",
          "content": "<p>I am making attempts to use multiple methods (train/test augmentation, pseudo-labeling, and (less robust) cv + stacking framework).  I think its mostly stacking that's contributing to  score.    </p>\n\n<p>@PeterHurford, congrats to you on your excellent score.  Any reason why you assume that I'm not doing pseudo-labeling?  Is this something that you have tried?  I have not gotten too much mileage from it yet but I am working on it.  :-)</p>",
          "rawMarkdown": "I am making attempts to use multiple methods (train/test augmentation, pseudo-labeling, and (less robust) cv + stacking framework).  I think its mostly stacking that's contributing to  score.    \n\n@PeterHurford, congrats to you on your excellent score.  Any reason why you assume that I'm not doing pseudo-labeling?  Is this something that you have tried?  I have not gotten too much mileage from it yet but I am working on it.  :-)"
        },
        {
          "id": 335546,
          "postDate": "2018-05-29T22:23:42.073Z",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> Thanks for the congrats - I hope it holds! :)</p>\n\n<p>...My assumption is that pseudo-labeling is only worthwhile when you have really accurate models. We were lucky to be in that case in the toxic competition, though I stupidly rejected the idea and didn't try it there. With RMSE in the &gt;0.2 range, I don't think our models are accurate enough, so I'm rejecting the idea again. I have not tried it. I hope it's the right call this time. ;)</p>",
          "rawMarkdown": "@larryfreeman Thanks for the congrats - I hope it holds! :)\n\n...My assumption is that pseudo-labeling is only worthwhile when you have really accurate models. We were lucky to be in that case in the toxic competition, though I stupidly rejected the idea and didn't try it there. With RMSE in the &gt;0.2 range, I don't think our models are accurate enough, so I'm rejecting the idea again. I have not tried it. I hope it's the right call this time. ;)",
          "votes": 3
        },
        {
          "id": 335549,
          "postDate": "2018-05-29T22:27:27.060Z",
          "content": "<p>additionally train and test distributions are quite close, which reduces the benefit from pseudo labeling</p>",
          "rawMarkdown": "additionally train and test distributions are quite close, which reduces the benefit from pseudo labeling",
          "votes": 6
        },
        {
          "id": 335553,
          "postDate": "2018-05-29T22:30:46.693Z",
          "content": "<p>Thanks, @Peter and @Dieter.  I am still developing my intuition about pseudo labeling.  Cheers.</p>",
          "rawMarkdown": "Thanks, @Peter and @Dieter.  I am still developing my intuition about pseudo labeling.  Cheers."
        },
        {
          "id": 335555,
          "postDate": "2018-05-29T22:31:39.537Z",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> I'd also be really curious to hear if augmentation is helpful. My assumption is that with 10x the data here compared to toxic, plus the impact of categoricals/images in addition to text, the benefits of adding translated data would be much lower now. ...But I could also be wrong about that.</p>",
          "rawMarkdown": "@larryfreeman I'd also be really curious to hear if augmentation is helpful. My assumption is that with 10x the data here compared to toxic, plus the impact of categoricals/images in addition to text, the benefits of adding translated data would be much lower now. ...But I could also be wrong about that.",
          "votes": 1
        },
        {
          "id": 335556,
          "postDate": "2018-05-29T22:35:47.767Z",
          "content": "<p>I share your assumption on data augmentation.  At this point, I'm trying everything that I can think of just to see what happens and to see if it gives me any other ideas to try.</p>",
          "rawMarkdown": "I share your assumption on data augmentation.  At this point, I'm trying everything that I can think of just to see what happens and to see if it gives me any other ideas to try.",
          "votes": 1
        },
        {
          "id": 335559,
          "postDate": "2018-05-29T22:38:52.247Z",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> It's awesome to try stuff. This competition is overwhelming in the amount of things there are to try. I hope you get lucky and one of these ideas pans out for you. :)</p>",
          "rawMarkdown": "@larryfreeman It's awesome to try stuff. This competition is overwhelming in the amount of things there are to try. I hope you get lucky and one of these ideas pans out for you. :)",
          "votes": 1
        },
        {
          "id": 335584,
          "postDate": "2018-05-30T00:50:49.670Z",
          "content": "<p>Thanks, @Peter.  Best of luck to you too.  :-)</p>",
          "rawMarkdown": "Thanks, @Peter.  Best of luck to you too.  :-)"
        }
      ]
    },
    {
      "id": 328983,
      "postDate": "2018-05-15T13:29:53.330Z",
      "content": "<p>Got 0.2222 with NN just train and test</p>",
      "rawMarkdown": "Got 0.2222 with NN just train and test",
      "votes": 3,
      "replies": [
        {
          "id": 329135,
          "postDate": "2018-05-15T20:13:18.727Z",
          "content": "<p>So the NN is for the NLP of description or more like for the classifier for different features just like LGBM</p>",
          "rawMarkdown": "So the NN is for the NLP of description or more like for the classifier for different features just like LGBM"
        },
        {
          "id": 329244,
          "postDate": "2018-05-16T04:40:09.050Z",
          "content": "<p>I guess both. At least I use it for both. It makes it much easier for me to connect different types of information. </p>",
          "rawMarkdown": "I guess both. At least I use it for both. It makes it much easier for me to connect different types of information. "
        }
      ]
    },
    {
      "id": 341626,
      "postDate": "2018-06-12T00:36:43.723Z",
      "content": "<p>After a lot of hard work I got my single LGB (with 5 fold) down to 0.2200 on LB. Yay !</p>",
      "rawMarkdown": "After a lot of hard work I got my single LGB (with 5 fold) down to 0.2200 on LB. Yay !",
      "votes": 4,
      "replies": [
        {
          "id": 341667,
          "postDate": "2018-06-12T04:12:58.910Z",
          "content": "<p>Good going buddy!</p>",
          "rawMarkdown": "Good going buddy!"
        },
        {
          "id": 342197,
          "postDate": "2018-06-13T04:31:23.570Z",
          "content": "<p>Good work, buddy!</p>",
          "rawMarkdown": "Good work, buddy!"
        }
      ]
    },
    {
      "id": 340365,
      "postDate": "2018-06-09T02:38:39.960Z",
      "content": "<p>0.2190 10-fold LGB</p>",
      "rawMarkdown": "0.2190 10-fold LGB",
      "votes": 4,
      "replies": [
        {
          "id": 340375,
          "postDate": "2018-06-09T03:15:45.817Z",
          "content": "<p>Do you known what is lb difference between 5-fold and 10-fold?</p>",
          "rawMarkdown": "Do you known what is lb difference between 5-fold and 10-fold?",
          "votes": 1
        },
        {
          "id": 340379,
          "postDate": "2018-06-09T03:23:30.757Z",
          "content": "<p>Approximately 0.0001-0.0002</p>",
          "rawMarkdown": "Approximately 0.0001-0.0002",
          "votes": 1
        },
        {
          "id": 340406,
          "postDate": "2018-06-09T05:22:56.690Z",
          "content": "<p>@Liu Jilong, I've only tried 1fold and 10fold.</p>\n\n<p>But I think it should be around <code>0.0002</code></p>",
          "rawMarkdown": "@Liu Jilong, I've only tried 1fold and 10fold.\n\nBut I think it should be around `0.0002`"
        },
        {
          "id": 340427,
          "postDate": "2018-06-09T06:19:30.583Z",
          "content": "<p>Congratulation!! you guys are so cool. upvote for you guys.</p>",
          "rawMarkdown": "Congratulation!! you guys are so cool. upvote for you guys."
        },
        {
          "id": 340689,
          "postDate": "2018-06-10T02:44:00.383Z",
          "content": "<p>Wow, this small gap is quite impressive. Mine are alway around 0.004 (20 times of yours)</p>",
          "rawMarkdown": "Wow, this small gap is quite impressive. Mine are alway around 0.004 (20 times of yours)"
        }
      ]
    },
    {
      "id": 342452,
      "postDate": "2018-06-13T14:39:15.370Z",
      "content": "<p>Silly question.</p>\n\n<p>What do you guys exacly mean when you say, for example, 10-fold LGB?\nYes, you can use CV to better tune hyperparameters. Particularly, n_estimators. Is that what you mean? </p>\n\n<p>And when you find optimal number of estimators do you retrain on whole dataset? </p>",
      "rawMarkdown": "Silly question.\n\nWhat do you guys exacly mean when you say, for example, 10-fold LGB?\nYes, you can use CV to better tune hyperparameters. Particularly, n_estimators. Is that what you mean? \n\nAnd when you find optimal number of estimators do you retrain on whole dataset? ",
      "votes": 1,
      "replies": [
        {
          "id": 342463,
          "postDate": "2018-06-13T15:01:05.843Z",
          "content": "<p>I guess people mean average the prediction of 10 LGB models in CV instead of retrain and get one model.  The advantage of avg is obvious :) </p>",
          "rawMarkdown": "I guess people mean average the prediction of 10 LGB models in CV instead of retrain and get one model.  The advantage of avg is obvious :) ",
          "votes": 1
        },
        {
          "id": 342521,
          "postDate": "2018-06-13T16:28:38.163Z",
          "content": "<p>Oh I see.\nYou can also achieve that with different random seeds. \nBut subsets of training data probably make models more different than random seeds. And that is good. </p>",
          "rawMarkdown": "Oh I see.\nYou can also achieve that with different random seeds. \nBut subsets of training data probably make models more different than random seeds. And that is good. "
        },
        {
          "id": 342532,
          "postDate": "2018-06-13T16:49:00.773Z",
          "content": "<p>But that’s not cv tho. </p>",
          "rawMarkdown": "But that’s not cv tho. "
        },
        {
          "id": 342578,
          "postDate": "2018-06-13T18:29:51.537Z",
          "content": "<p>10-fold is not used for hyper parameters tuning but just as CV technic. When running 10 fold, it's similar to running 10 times your LGB with 10 different seed and then doing an average. The advantage of using 10 fold, is that you can be sure that all data were used for train and for validation. (which might not be the case if you take random seeds)</p>",
          "rawMarkdown": "10-fold is not used for hyper parameters tuning but just as CV technic. When running 10 fold, it's similar to running 10 times your LGB with 10 different seed and then doing an average. The advantage of using 10 fold, is that you can be sure that all data were used for train and for validation. (which might not be the case if you take random seeds)",
          "votes": 1
        },
        {
          "id": 342610,
          "postDate": "2018-06-13T19:50:04.500Z",
          "content": "<p>I hope to clear up some confusion here - it's not a silly question at all. It's common conflation of terminology on kaggle to refer to partitioned fold averaging as \"CV\", when it is a distinct method from CV that shares the partitioning structure in common. </p>\n\n<p>Here are the two methods:</p>\n\n<ol>\n<li><p>Cross-validation - partition the data into k folds for train/validation splitting to exhaustively use all of the data for validation, with the purpose of evaluating model and hyperparameter performance.  You use the evaluation results to guide your model/hyperparameter selection.</p></li>\n<li><p>K-fold averaging - partition the data into k folds (using the same partitioning as in CV), train the same model on each fold, and average the results to help reduce the variance of a single model. You could add out of fold validation to this process without using it directly for selection, or you can just run with the result of the averaging if you feel good about the model/hyperparameter selection.</p></li>\n</ol>\n\n<p>The reason the terminology gets confused is that many people take a shortcut and do both things at the same time. They do one k-fold split, use out of fold sets to select something like n_estimators for gradient boosting models with early stopping, then average the resulting models across the k train/validation runs. In theory, this is poor practice. Early stopping overfits to the particular choice of validation fold, so models averaged in this manner within CV inevitably overfit. In practice, the averaging helps mitigate some of the overfitting and the overall harm may be small, but it's something worth keeping in mind. It certainly causes your local score estimates to be at least slightly optimistic. </p>\n\n<p>This is a common pitfall in use of CV more generally. People don't realize that by using certain validation sets to guide their model selection, they've reduced the legitimacy of using those same validation sets for estimating generalization error. The best estimates of generalization error are given by sets that have had no influence on model selection. This is why doing CV <strong>and</strong> setting aside a true hold-out set that only gets used once for evaluation is a common and sound practice.    </p>",
          "rawMarkdown": "I hope to clear up some confusion here - it's not a silly question at all. It's common conflation of terminology on kaggle to refer to partitioned fold averaging as \"CV\", when it is a distinct method from CV that shares the partitioning structure in common. \n\nHere are the two methods:\n\n1. Cross-validation - partition the data into k folds for train/validation splitting to exhaustively use all of the data for validation, with the purpose of evaluating model and hyperparameter performance.  You use the evaluation results to guide your model/hyperparameter selection.\n\n2. K-fold averaging - partition the data into k folds (using the same partitioning as in CV), train the same model on each fold, and average the results to help reduce the variance of a single model. You could add out of fold validation to this process without using it directly for selection, or you can just run with the result of the averaging if you feel good about the model/hyperparameter selection.\n\nThe reason the terminology gets confused is that many people take a shortcut and do both things at the same time. They do one k-fold split, use out of fold sets to select something like n_estimators for gradient boosting models with early stopping, then average the resulting models across the k train/validation runs. In theory, this is poor practice. Early stopping overfits to the particular choice of validation fold, so models averaged in this manner within CV inevitably overfit. In practice, the averaging helps mitigate some of the overfitting and the overall harm may be small, but it's something worth keeping in mind. It certainly causes your local score estimates to be at least slightly optimistic. \n\nThis is a common pitfall in use of CV more generally. People don't realize that by using certain validation sets to guide their model selection, they've reduced the legitimacy of using those same validation sets for estimating generalization error. The best estimates of generalization error are given by sets that have had no influence on model selection. This is why doing CV **and** setting aside a true hold-out set that only gets used once for evaluation is a common and sound practice.    ",
          "votes": 26
        },
        {
          "id": 342652,
          "postDate": "2018-06-13T20:59:59.003Z",
          "content": "<p>One practical reason why k-fold cross validation is famous around kagglers, is that you can do this approach for several layers, without needing a new hold-out set for each level. This is especially helpful when working together in a team or stacking different models.</p>",
          "rawMarkdown": "One practical reason why k-fold cross validation is famous around kagglers, is that you can do this approach for several layers, without needing a new hold-out set for each level. This is especially helpful when working together in a team or stacking different models.",
          "votes": 4
        },
        {
          "id": 342692,
          "postDate": "2018-06-13T22:43:44.870Z",
          "content": "<p>@Joe Eddy: nice overview!</p>\n\n<blockquote>\n  <p>This is why doing CV and setting aside a true hold-out set that only\n  gets used once for evaluation is a common and sound practice.</p>\n</blockquote>\n\n<p>In this competition at least, I consider the public leaderboard to be my true hold-out set.  I prefer to use K-Fold CV to generate my predictions, with K models each built on K-1 folds having their predictions averaged together.  I see this as form of de facto model averaging.  </p>\n\n<p>I use early stopping.  I agree that it overfits.  But it also seems to me that each K-1 folds are best fit by a different # of epochs or trees, so guessing this # without early stopping is also misfitting (either underfitting or overfitting).  So the misfitting washes out.</p>\n\n<p>The big bonus for me is using the out-of-fold predictions for model selection and to optimize model stacking.</p>",
          "rawMarkdown": "@Joe Eddy: nice overview!\n\n&gt; This is why doing CV and setting aside a true hold-out set that only\n&gt; gets used once for evaluation is a common and sound practice.\n\nIn this competition at least, I consider the public leaderboard to be my true hold-out set.  I prefer to use K-Fold CV to generate my predictions, with K models each built on K-1 folds having their predictions averaged together.  I see this as form of de facto model averaging.  \n\nI use early stopping.  I agree that it overfits.  But it also seems to me that each K-1 folds are best fit by a different # of epochs or trees, so guessing this # without early stopping is also misfitting (either underfitting or overfitting).  So the misfitting washes out.\n\nThe big bonus for me is using the out-of-fold predictions for model selection and to optimize model stacking.\n",
          "votes": 6
        },
        {
          "id": 342713,
          "postDate": "2018-06-14T00:00:36.640Z",
          "content": "<p>I agree with @Harlan Seymour. The reason you stated is why I am using CV this way and I believe most people in this competition are as well.</p>",
          "rawMarkdown": "I agree with @Harlan Seymour. The reason you stated is why I am using CV this way and I believe most people in this competition are as well.\n",
          "votes": 2
        },
        {
          "id": 342722,
          "postDate": "2018-06-14T00:31:55.450Z",
          "content": "<p>@Harlan and YaGana, I definitely agree it's very convenient and that overfitting is mostly mitigated. It works well enough to produce good submissions and stacking results. But it's also <em>definitely not optimal</em>. When stacking, generating out of fold predictions on the same folds you early stop on introduces an optimistic bias that your meta learner may overfit to (these predictions are no longer out of fold in a true sense of the model not having seen the data, it now has glimpsed at that data). If that optimistic bias exists across all the models you stack it likely washes out, but if some models were trained differently (e.g. a non-boosting model that doesn't have the same early stopping concept), it might be problematic for the meta learner.</p>\n\n<p>There's no reason you can't get the best of both worlds, at the cost of extra computation time. You can run a 5-fold CV for model selection, then run another 5-fold process on a different seed to generate true OOF predictions to stack with. You can take n_estimators as the mean of the 5-fold CV early-stopped n_estimators, or use a similar heuristic. I can't find it right now, but I believe there's a post somewhere where Laurae (LGBM contributor) recommends the heuristic 110% of mean early-stopped n_estimators.     </p>",
          "rawMarkdown": "@Harlan and YaGana, I definitely agree it's very convenient and that overfitting is mostly mitigated. It works well enough to produce good submissions and stacking results. But it's also *definitely not optimal*. When stacking, generating out of fold predictions on the same folds you early stop on introduces an optimistic bias that your meta learner may overfit to (these predictions are no longer out of fold in a true sense of the model not having seen the data, it now has glimpsed at that data). If that optimistic bias exists across all the models you stack it likely washes out, but if some models were trained differently (e.g. a non-boosting model that doesn't have the same early stopping concept), it might be problematic for the meta learner.\n\nThere's no reason you can't get the best of both worlds, at the cost of extra computation time. You can run a 5-fold CV for model selection, then run another 5-fold process on a different seed to generate true OOF predictions to stack with. You can take n_estimators as the mean of the 5-fold CV early-stopped n_estimators, or use a similar heuristic. I can't find it right now, but I believe there's a post somewhere where Laurae (LGBM contributor) recommends the heuristic 110% of mean early-stopped n_estimators.     ",
          "votes": 2
        },
        {
          "id": 342758,
          "postDate": "2018-06-14T03:16:38.033Z",
          "content": "<p>@ Joe, you've definitely given me food for thought.  Take for example XGBoost.  What do you think of this compromise to save compute time? </p>\n\n<ol>\n<li>Save models and record ntree_limit used in early stopping for all K folds</li>\n<li>Calculate mean of ntree_limit used in early stopping over the K folds (call it M)</li>\n<li>Reload models with &lt; 110% x M trees and train them up to 110% x M</li>\n<li>Call predict() for each model to generate test and out-of-fold predictions, passing 110% x M for ntree_limit for each model.</li>\n</ol>\n\n<p>Note that I might prefer something like 95% x M to 110% x M.  Val score changes very little for the final trees before early stopping while train score goes down significantly i.e. lots of overfitting at the end before early stopping.</p>\n\n<p>Doesn't this almost entirely mitigate your overfitting worry without having to do a double run?  Note that XGBoost takes a long for me in this competition.  </p>\n\n<p>@Joe, I have another question it would be interesting to get your input on.  Do you create a model on the entire training set, using in XGBoost, say, 110% x M x (K / (K-1)) for ntree_limit?  This model would be used just for test predictions.  Maybe average its test predictions with average of K fold test predictions?</p>",
          "rawMarkdown": "@ Joe, you've definitely given me food for thought.  Take for example XGBoost.  What do you think of this compromise to save compute time? \n \n1. Save models and record ntree_limit used in early stopping for all K folds\n2. Calculate mean of ntree_limit used in early stopping over the K folds (call it M)\n3. Reload models with &lt; 110% x M trees and train them up to 110% x M\n4. Call predict() for each model to generate test and out-of-fold predictions, passing 110% x M for ntree_limit for each model.\n\nNote that I might prefer something like 95% x M to 110% x M.  Val score changes very little for the final trees before early stopping while train score goes down significantly i.e. lots of overfitting at the end before early stopping.\n\nDoesn't this almost entirely mitigate your overfitting worry without having to do a double run?  Note that XGBoost takes a long for me in this competition.  \n\n@Joe, I have another question it would be interesting to get your input on.  Do you create a model on the entire training set, using in XGBoost, say, 110% x M x (K / (K-1)) for ntree_limit?  This model would be used just for test predictions.  Maybe average its test predictions with average of K fold test predictions?",
          "votes": 1
        },
        {
          "id": 342763,
          "postDate": "2018-06-14T03:57:22.260Z",
          "content": "<p>My personal preference is to not use early stopping. Use trial and error or some initial early stopping to get a certain number of rounds, and then use that number of rounds consistently with all the folds. Also remember to pay attention to the relationship between choice of learning rate and number of rounds needed.</p>",
          "rawMarkdown": "My personal preference is to not use early stopping. Use trial and error or some initial early stopping to get a certain number of rounds, and then use that number of rounds consistently with all the folds. Also remember to pay attention to the relationship between choice of learning rate and number of rounds needed.",
          "votes": 2
        },
        {
          "id": 343080,
          "postDate": "2018-06-14T17:06:36.897Z",
          "content": "<p>@Joe Eddy</p>\n\n<p>I agree with a lot of this but if we are talking about <em>optimal</em> then using a heuristic such as 110% of the mean of the early stopped n_estimators certainly isn't optimal.</p>\n\n<p>P.S. Does anyone know why when I reply to, say, Joe Eddy, the forum thinks I am replying to the person he replied to?</p>\n\n<p>P.P.S Some of this is covered <a href=\"http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\">here</a> which is useful reading.</p>",
          "rawMarkdown": "@Joe Eddy\n\nI agree with a lot of this but if we are talking about *optimal* then using a heuristic such as 110% of the mean of the early stopped n_estimators certainly isn't optimal.\n\nP.S. Does anyone know why when I reply to, say, Joe Eddy, the forum thinks I am replying to the person he replied to?\n\nP.P.S Some of this is covered [here][1] which is useful reading.\n\n\n  [1]: http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/",
          "votes": 1
        },
        {
          "id": 343102,
          "postDate": "2018-06-14T17:37:55.190Z",
          "content": "<p><a href=\"/maw501\">@maw501</a> that's fair, there's no such this as optimality in statistical modeling of distributions that we can't have perfect knowledge of, everything is relative :) </p>\n\n<p>But on a practical level, I don't know what you would expect (both in the normal sense of expect, and <em>on expectation</em> in the statistical sense) to work better for generalization performance of n_estimators than something close to the mean of early stopped rounds. Here's a <a href=\"https://www.kaggle.com/ogrellier/can-early-stopping-generalize/notebook\">nice experiment</a> in favor of using the mean to estimate the optimal number of rounds from olivier. Do you have in mind something that works better / what would your choice be?</p>\n\n<p>@Harlan I like that idea a lot, it's not something I've thought of before. I can't think of a reason why it wouldn't mitigate the overfitting issue, sounds like a great time saver and worth a shot to me. </p>\n\n<p>Re: your second question, I think you could definitely do that, but the gains might be pretty small if you're already using all of the data for training by averaging test predictions across k-folds. I tend to skip that to save time, but using all the training data in one model might be worth it. It aligns with the more traditional way of stacking - generate oof predictions in k-fold, retrain model on entire dataset to generate test predictions (without averaging of models). With relatively smaller datasets like this one my instinct is that averaging on k-folds gives a notable improvement via variance reduction, so that's what I go with.</p>",
          "rawMarkdown": "@maw501 that's fair, there's no such this as optimality in statistical modeling of distributions that we can't have perfect knowledge of, everything is relative :) \n\nBut on a practical level, I don't know what you would expect (both in the normal sense of expect, and *on expectation* in the statistical sense) to work better for generalization performance of n_estimators than something close to the mean of early stopped rounds. Here's a [nice experiment][1] in favor of using the mean to estimate the optimal number of rounds from olivier. Do you have in mind something that works better / what would your choice be?\n\n @Harlan I like that idea a lot, it's not something I've thought of before. I can't think of a reason why it wouldn't mitigate the overfitting issue, sounds like a great time saver and worth a shot to me. \n\nRe: your second question, I think you could definitely do that, but the gains might be pretty small if you're already using all of the data for training by averaging test predictions across k-folds. I tend to skip that to save time, but using all the training data in one model might be worth it. It aligns with the more traditional way of stacking - generate oof predictions in k-fold, retrain model on entire dataset to generate test predictions (without averaging of models). With relatively smaller datasets like this one my instinct is that averaging on k-folds gives a notable improvement via variance reduction, so that's what I go with.\n\n\n\n  [1]: https://www.kaggle.com/ogrellier/can-early-stopping-generalize/notebook",
          "votes": 4
        },
        {
          "id": 343163,
          "postDate": "2018-06-14T20:01:47.250Z",
          "content": "<p>+1 for variance reduction. </p>",
          "rawMarkdown": "+1 for variance reduction. "
        },
        {
          "id": 343188,
          "postDate": "2018-06-14T21:07:04.853Z",
          "content": "<p>I have a question about stacking here in Kaggle and for that competition:\nMy current way of doing it is:\n-For 5 folds, train 5 models then predict on the OOF to construct train meta data for that model.\n-Train a full model and create a test meta data.\n(then later on train a meta model etc.)</p>\n\n<p>What I'm wondering is if it makes sense to instead use the average of the 5 models predictions on test data. The same way I (and others) currently test my single models on the LB.\nDoes it make sense to think that the test meta is more stable that way and closer to the train meta (on average) compared to have a different model that created test meta ?</p>",
          "rawMarkdown": "I have a question about stacking here in Kaggle and for that competition:\nMy current way of doing it is:\n-For 5 folds, train 5 models then predict on the OOF to construct train meta data for that model.\n-Train a full model and create a test meta data.\n(then later on train a meta model etc.)\n\nWhat I'm wondering is if it makes sense to instead use the average of the 5 models predictions on test data. The same way I (and others) currently test my single models on the LB.\nDoes it make sense to think that the test meta is more stable that way and closer to the train meta (on average) compared to have a different model that created test meta ?",
          "votes": 1
        },
        {
          "id": 343216,
          "postDate": "2018-06-14T22:30:54.153Z",
          "content": "<p>Hi Joe,</p>\n\n<p>Again, thanks for a thoughtful reply. I definitely agree with your first comment re. optimality and this alongside scepticism is a reasonable grounding for most problems. I also think you somewhat answer your own point when asking for what I would <em>expect</em> to perform better when you point out the elusiveness of optimality. I don't expect any method to be consistently better than any other (for a reasonable choice of methods).</p>\n\n<p>This is more complicated clearly and I think a simple (self-evident) answer is that it's entirely problem dependent. I've yet to see any ML technique/approach (including state of the art) that one feels entirely comfortable applying in all contexts. My last two competitions I worried far too much about whether things were theoretically justifiable/sound and paid the price. Worrying about reliably estimating generalization error (which we care about in the real world) vs. performing well on a Kaggle competition vs. having robust statistical backing for one's ideas are all different things.</p>\n\n<p>My choice is usually to try most of the approaches listed with a (more recent) favouring to average across the predictions from the folds for the base models. The biggest problem with using a simple heuristic like 110% of the mean of the early stopped n_estimators is that given the extra data it feels a little too much like guesswork for me. Yes under some assumptions I'm sure it's fine, but what are they? Can we even state them?</p>\n\n<p>Further if you change from 5-fold to 10-fold to any <strong>k</strong>-fold - what then? Add in that one should perhaps favour higher entropy validation methods then you should accept some folds will be harder than others for your model and early stopping at least deals with this (of course, at some risk of overfitting) in a consistent manner - additionally you have the benefit of averaging too. If you have a diverse set of base models I don't particularly worry about any single model overfitting too much.</p>\n\n<p>I do like the idea of using different random seeds for model iteration and oof predictions and will certainly try this.</p>\n\n<p>Edit extra: Just wondering if we can resolve via a thought experiment. For instance, take an extreme 2-fold example with one really easy fold needing 500 rounds and the other really difficult needing 1500 - using the heuristic we would train on the whole data now with 1, 100 rounds. I'm not sure I'm happy with that vs. averaging the predictions from both which at least predicts at the point when it thinks it has discovered all the information from the data. I'm not saying the difference will be big, or one is consistently better but I like getting variance reduction for free and some comfort in knowing the number of rounds. I guess to some extent it's a function of the dissimilarity in the folds/strength of signal in each but I'm not sure how one would formalise this - so back to my point about not expecting any method to do better on average.</p>",
          "rawMarkdown": "Hi Joe,\n\nAgain, thanks for a thoughtful reply. I definitely agree with your first comment re. optimality and this alongside scepticism is a reasonable grounding for most problems. I also think you somewhat answer your own point when asking for what I would *expect* to perform better when you point out the elusiveness of optimality. I don't expect any method to be consistently better than any other (for a reasonable choice of methods).\n\nThis is more complicated clearly and I think a simple (self-evident) answer is that it's entirely problem dependent. I've yet to see any ML technique/approach (including state of the art) that one feels entirely comfortable applying in all contexts. My last two competitions I worried far too much about whether things were theoretically justifiable/sound and paid the price. Worrying about reliably estimating generalization error (which we care about in the real world) vs. performing well on a Kaggle competition vs. having robust statistical backing for one's ideas are all different things.\n\nMy choice is usually to try most of the approaches listed with a (more recent) favouring to average across the predictions from the folds for the base models. The biggest problem with using a simple heuristic like 110% of the mean of the early stopped n_estimators is that given the extra data it feels a little too much like guesswork for me. Yes under some assumptions I'm sure it's fine, but what are they? Can we even state them?\n\nFurther if you change from 5-fold to 10-fold to any **k**-fold - what then? Add in that one should perhaps favour higher entropy validation methods then you should accept some folds will be harder than others for your model and early stopping at least deals with this (of course, at some risk of overfitting) in a consistent manner - additionally you have the benefit of averaging too. If you have a diverse set of base models I don't particularly worry about any single model overfitting too much.\n\nI do like the idea of using different random seeds for model iteration and oof predictions and will certainly try this.\n\nEdit extra: Just wondering if we can resolve via a thought experiment. For instance, take an extreme 2-fold example with one really easy fold needing 500 rounds and the other really difficult needing 1500 - using the heuristic we would train on the whole data now with 1, 100 rounds. I'm not sure I'm happy with that vs. averaging the predictions from both which at least predicts at the point when it thinks it has discovered all the information from the data. I'm not saying the difference will be big, or one is consistently better but I like getting variance reduction for free and some comfort in knowing the number of rounds. I guess to some extent it's a function of the dissimilarity in the folds/strength of signal in each but I'm not sure how one would formalise this - so back to my point about not expecting any method to do better on average.",
          "votes": 1
        },
        {
          "id": 343670,
          "postDate": "2018-06-15T19:47:05.077Z",
          "content": "<p>@Joe: As an experiment I applied the procedure I outlined to my latest XGB per-fold models, syncing them to all to be 100% of the mean of the early stopping ntree_limits.  My local CV score went up minutely by +0.00001.  The newly submitted test predictions had the same RMSE score on the public leaderboard, but ordering \"My Submissions\" by score showed it was not as good as the original.</p>",
          "rawMarkdown": "@Joe: As an experiment I applied the procedure I outlined to my latest XGB per-fold models, syncing them to all to be 100% of the mean of the early stopping ntree_limits.  My local CV score went up minutely by +0.00001.  The newly submitted test predictions had the same RMSE score on the public leaderboard, but ordering \"My Submissions\" by score showed it was not as good as the original.",
          "votes": 1
        },
        {
          "id": 343696,
          "postDate": "2018-06-15T20:54:37.490Z",
          "content": "<p>Thanks @Joe Eddy for sparking this useful discussion. And thanks @Harlan Seymour for the update on your experiment. </p>\n\n<p><a href=\"/maw501\">@maw501</a>, I would like to know the result of the thought experiment on testing on 2 extreme cases you mentioned. If you ever do that, please report back if you can.</p>\n\n<p>Averaging is what I am comfortable with at the moment to reduce variance. </p>",
          "rawMarkdown": "Thanks @Joe Eddy for sparking this useful discussion. And thanks @Harlan Seymour for the update on your experiment. \n\n@maw501, I would like to know the result of the thought experiment on testing on 2 extreme cases you mentioned. If you ever do that, please report back if you can.\n\nAveraging is what I am comfortable with at the moment to reduce variance. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 336590,
      "postDate": "2018-06-01T01:59:22.720Z",
      "content": "<p>LGB 0.2223 5-fold average</p>\n\n<p><strong>update</strong>:\nnew feature add, 0.2217 without 5 folder</p>\n\n<p>0.2210 with 5 folder</p>\n\n<p><strong>update</strong>  6-15:</p>\n\n<p>0.2208 with 5 folder</p>\n\n<p><strong>update</strong>  6-20:</p>\n\n<p>0.2205 with 5 folder</p>\n\n<p>final:</p>\n\n<p>0.2204 with 10 folder</p>",
      "rawMarkdown": "LGB 0.2223 5-fold average\n\n**update**:\nnew feature add, 0.2217 without 5 folder\n\n\n0.2210 with 5 folder\n\n\n**update**  6-15:\n\n0.2208 with 5 folder\n\n**update**  6-20:\n\n0.2205 with 5 folder\n\nfinal:\n\n0.2204 with 10 folder",
      "votes": 1,
      "replies": [
        {
          "id": 336595,
          "postDate": "2018-06-01T02:17:59.517Z",
          "content": "<p>劉sir,  您是用Stack跳到0.2210的麼?</p>",
          "rawMarkdown": "劉sir,  您是用Stack跳到0.2210的麼?",
          "votes": -2
        },
        {
          "id": 337169,
          "postDate": "2018-06-02T06:24:52.117Z",
          "content": "<p>还没stacking，简单的blend几个public kernel</p>",
          "rawMarkdown": "还没stacking，简单的blend几个public kernel",
          "votes": -2
        },
        {
          "id": 337275,
          "postDate": "2018-06-02T12:38:43.800Z",
          "content": "<p>Translation: \nLiu sir , are you jumping to 0.2210 with Stack?\n - No stacking, simple blend several public kernel</p>",
          "rawMarkdown": "Translation: \nLiu sir , are you jumping to 0.2210 with Stack?\n - No stacking, simple blend several public kernel",
          "votes": 9
        },
        {
          "id": 338644,
          "postDate": "2018-06-05T13:44:03.090Z",
          "content": "<p>有啥技巧么，我blend一个0.2212和0.2216结果也就0.2212</p>",
          "rawMarkdown": "有啥技巧么，我blend一个0.2212和0.2216结果也就0.2212",
          "votes": -5
        },
        {
          "id": 338654,
          "postDate": "2018-06-05T14:17:22.973Z",
          "content": "<p>I guess there is not any techniques. Just use weight to blend public with private kernel. I used blending, too</p>",
          "rawMarkdown": "I guess there is not any techniques. Just use weight to blend public with private kernel. I used blending, too"
        },
        {
          "id": 339058,
          "postDate": "2018-06-06T08:01:20.183Z",
          "content": "<p>you need try to improve your single model,then try to blend or stacking</p>",
          "rawMarkdown": "you need try to improve your single model,then try to blend or stacking"
        }
      ]
    },
    {
      "id": 330553,
      "postDate": "2018-05-19T05:14:52.653Z",
      "content": "<p>Edit :  0.2249</p>\n\n<p>Edit: 0.2252 with Image_top_1 feature included </p>\n\n<p>Edit: 0.2259 with RNN </p>\n\n<p>0.2276 with RNN, Fast Text Embeddings . - No Tuning done,  A lot of Scope for improvement and tuning. \nLink to Public kernel is here </p>\n\n<p><a href=\"https://www.kaggle.com/shanth84/avito-fast-text-keras-model/code\">0.2276 RNN - Good for Blend :)</a></p>",
      "rawMarkdown": "Edit :  0.2249\n\nEdit: 0.2252 with Image_top_1 feature included \n\nEdit: 0.2259 with RNN \n\n0.2276 with RNN, Fast Text Embeddings . - No Tuning done,  A lot of Scope for improvement and tuning. \nLink to Public kernel is here \n\n[0.2276 RNN - Good for Blend :)][1]\n\n\n  [1]: https://www.kaggle.com/shanth84/avito-fast-text-keras-model/code",
      "votes": 1
    },
    {
      "id": 329707,
      "postDate": "2018-05-17T02:38:18.170Z",
      "content": "<p>0.2223 with 5fold LGB model (all features are taken from public kernels). </p>",
      "rawMarkdown": "0.2223 with 5fold LGB model (all features are taken from public kernels). ",
      "votes": 1
    },
    {
      "id": 329210,
      "postDate": "2018-05-16T02:38:27.313Z",
      "content": "<p>Single lgbm 0.223. </p>",
      "rawMarkdown": "Single lgbm 0.223. ",
      "votes": 1
    },
    {
      "id": 329185,
      "postDate": "2018-05-16T00:01:24.690Z",
      "content": "<p>Single lgbm 0.2215.\nBut using image features from <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">https://www.kaggle.com/bguberfain/vgg16-train-features</a> , my score got worse.\nI would like to know how to use image feature.</p>",
      "rawMarkdown": "Single lgbm 0.2215.\nBut using image features from https://www.kaggle.com/bguberfain/vgg16-train-features , my score got worse.\nI would like to know how to use image feature.",
      "votes": 1,
      "replies": [
        {
          "id": 329196,
          "postDate": "2018-05-16T00:39:06.463Z",
          "content": "<p>Wow, that is so high for a single lgbm model! Would you mind if share some ideas how to improve from 0.223 level to 0.2215 ?</p>",
          "rawMarkdown": "Wow, that is so high for a single lgbm model! Would you mind if share some ideas how to improve from 0.223 level to 0.2215 ?"
        },
        {
          "id": 329209,
          "postDate": "2018-05-16T02:34:48.440Z",
          "content": "<p>Price and item-seq-number is important feature. I refer to <a href=\"https://www.kaggle.com/tapioca/item-seq-number-vs-deal-probability/code\">https://www.kaggle.com/tapioca/item-seq-number-vs-deal-probability/code</a> , and make some features. Good luck!</p>",
          "rawMarkdown": "Price and item-seq-number is important feature. I refer to https://www.kaggle.com/tapioca/item-seq-number-vs-deal-probability/code , and make some features. Good luck!",
          "votes": 2
        },
        {
          "id": 329213,
          "postDate": "2018-05-16T02:43:55.210Z",
          "content": "<p>Thanks! For the image feature, my thought is: the features from the pre-trained Imagenet model are designed for classification. For example, we want our VGG16 distinguish bicycle from car. However, we already have the features like item_id and category, so directly use imgenet features may not helps. Maybe we can try fine tune those features with the target for the deal probabilities.  This is just my naive thought and I haven't try any image features</p>",
          "rawMarkdown": "Thanks! For the image feature, my thought is: the features from the pre-trained Imagenet model are designed for classification. For example, we want our VGG16 distinguish bicycle from car. However, we already have the features like item_id and category, so directly use imgenet features may not helps. Maybe we can try fine tune those features with the target for the deal probabilities.  This is just my naive thought and I haven't try any image features",
          "votes": 1
        },
        {
          "id": 329214,
          "postDate": "2018-05-16T02:51:07.357Z",
          "content": "<p>Wow! Your thought of image feature is very cool! I try to use color base feature because it is easy to use. Thank you!</p>",
          "rawMarkdown": "Wow! Your thought of image feature is very cool! I try to use color base feature because it is easy to use. Thank you!"
        },
        {
          "id": 329451,
          "postDate": "2018-05-16T13:45:34.857Z",
          "content": "<p>And also may ask how many models you are using to achieve your LB position now?</p>",
          "rawMarkdown": "And also may ask how many models you are using to achieve your LB position now?"
        },
        {
          "id": 329466,
          "postDate": "2018-05-16T14:10:34.853Z",
          "content": "<p>LB 0.2209 by single lgbm. Add some features from active.csv and periods data.</p>",
          "rawMarkdown": "LB 0.2209 by single lgbm. Add some features from active.csv and periods data.",
          "votes": 7
        },
        {
          "id": 329646,
          "postDate": "2018-05-16T21:35:57.487Z",
          "content": "<p>@takuoko : item-seq-number seems to be exactly the same as doing a count after a group by user no ? I was able to get good features from the price, but not by combining item_seq_number with another feature... Thanks for sharing !</p>",
          "rawMarkdown": "@takuoko : item-seq-number seems to be exactly the same as doing a count after a group by user no ? I was able to get good features from the price, but not by combining item_seq_number with another feature... Thanks for sharing !"
        },
        {
          "id": 329770,
          "postDate": "2018-05-17T06:39:04.460Z",
          "content": "<p>Thanks for your comment, Antoine! By using item-seq-number feature, 0.00005 better val score in my 5 folds. So it may be ineffective. But I think item-seq-number is important. So I continue trying to make featrue about that. </p>",
          "rawMarkdown": "Thanks for your comment, Antoine! By using item-seq-number feature, 0.00005 better val score in my 5 folds. So it may be ineffective. But I think item-seq-number is important. So I continue trying to make featrue about that. ",
          "votes": 1
        },
        {
          "id": 329895,
          "postDate": "2018-05-17T13:36:33.833Z",
          "content": "<p>simple feature from VGG16 is not helpful, because  the prediction is not precise nor suitable. I have write a kernel <a href=\"https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label\">https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label</a></p>",
          "rawMarkdown": "simple feature from VGG16 is not helpful, because  the prediction is not precise nor suitable. I have write a kernel https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label"
        }
      ]
    },
    {
      "id": 329122,
      "postDate": "2018-05-15T19:17:42.043Z",
      "content": "<p>Single lgbm 0.2229</p>",
      "rawMarkdown": "Single lgbm 0.2229",
      "votes": 1
    },
    {
      "id": 329052,
      "postDate": "2018-05-15T16:02:51.120Z",
      "content": "<p>LB 0.2249 with only sample train and test dataset via lightgbm</p>",
      "rawMarkdown": "LB 0.2249 with only sample train and test dataset via lightgbm",
      "votes": 1
    },
    {
      "id": 328924,
      "postDate": "2018-05-15T11:41:07.980Z",
      "content": "<p>My best single NN is 0.2233 without image nor active and period files.</p>",
      "rawMarkdown": "My best single NN is 0.2233 without image nor active and period files.",
      "votes": 1,
      "replies": [
        {
          "id": 329641,
          "postDate": "2018-05-16T21:24:57.023Z",
          "content": "<p>Update: <br>\n0.2226 for NN</p>",
          "rawMarkdown": "Update:  \n0.2226 for NN",
          "votes": 1
        },
        {
          "id": 329741,
          "postDate": "2018-05-17T04:58:21.573Z",
          "content": "<p>Could you give a hint about architecture of your NN?</p>",
          "rawMarkdown": "Could you give a hint about architecture of your NN?"
        },
        {
          "id": 329745,
          "postDate": "2018-05-17T05:15:39.803Z",
          "content": "<p>As @Strideradu commented above,</p>\n\n<p>&gt; So the NN is for the NLP of description or more like for the classifier for different features just like LGBM</p>\n\n<p>I use both. </p>\n\n<ol>\n<li><p>Category + Numeric part <br>\nYou can refer from fastai or some solution in this competition or similar one: \n<a href=\"https://www.kaggle.com/c/donorschoose-application-screening\">https://www.kaggle.com/c/donorschoose-application-screening</a></p></li>\n<li><p>NLP <br>\nYou can refer from previous competitions like: <br>\n<a href=\"https://www.kaggle.com/c/donorschoose-application-screening\">https://www.kaggle.com/c/donorschoose-application-screening</a> <br>\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge</a>\n...\nto know how they dealed with NLP.</p></li>\n</ol>",
          "rawMarkdown": "As @Strideradu commented above,\n\n&gt; So the NN is for the NLP of description or more like for the classifier for different features just like LGBM\n\nI use both. \n\n 1. Category + Numeric part  \nYou can refer from fastai or some solution in this competition or similar one: \nhttps://www.kaggle.com/c/donorschoose-application-screening\n\n 2. NLP  \nYou can refer from previous competitions like:  \nhttps://www.kaggle.com/c/donorschoose-application-screening  \nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\n...\nto know how they dealed with NLP.\n\n",
          "votes": 5
        },
        {
          "id": 329748,
          "postDate": "2018-05-17T05:20:15.037Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 336777,
      "postDate": "2018-06-01T09:02:02.143Z",
      "content": "<p>NN:\n0.2213 (Categorical, Numerical, W2V)\n10-fold</p>",
      "rawMarkdown": "NN:\n0.2213 (Categorical, Numerical, W2V)\n10-fold",
      "votes": 2
    },
    {
      "id": 335073,
      "postDate": "2018-05-29T05:48:58.033Z",
      "content": "<p>one LightGBM  model 5kfold - 0,2211 (UPD)</p>",
      "rawMarkdown": "one LightGBM  model 5kfold - 0,2211 (UPD)",
      "votes": 2,
      "replies": [
        {
          "id": 335075,
          "postDate": "2018-05-29T05:55:18.377Z",
          "content": "<p>incredible</p>",
          "rawMarkdown": "incredible"
        },
        {
          "id": 335214,
          "postDate": "2018-05-29T11:41:39.043Z",
          "content": "<p>Almost all ideas i take from public kernels :)</p>",
          "rawMarkdown": "Almost all ideas i take from public kernels :)",
          "votes": 1
        },
        {
          "id": 338415,
          "postDate": "2018-06-05T02:30:16.977Z",
          "content": "<p>@AlexTru, did you use image features? If so which ones if you can share? They seem to worsen both my CV and LB.</p>",
          "rawMarkdown": "@AlexTru, did you use image features? If so which ones if you can share? They seem to worsen both my CV and LB."
        },
        {
          "id": 338447,
          "postDate": "2018-06-05T04:28:42.777Z",
          "content": "<p>Yes, i use features from this two kernels. <a href=\"https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf\">this</a>  and <a href=\"https://www.kaggle.com/peterhurford/image-feature-engineering\">this</a></p>",
          "rawMarkdown": "Yes, i use features from this two kernels. [this][1]  and [this][2]\n\n\n  [1]: https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf\n  [2]: https://www.kaggle.com/peterhurford/image-feature-engineering"
        },
        {
          "id": 339187,
          "postDate": "2018-06-06T13:30:03.850Z",
          "content": "<p>UPD 22.07</p>",
          "rawMarkdown": "UPD 22.07",
          "votes": 1
        }
      ]
    },
    {
      "id": 331503,
      "postDate": "2018-05-21T12:22:52.577Z",
      "content": "<p>Update: 0.2220 with variation of the 'ridge trick', 5fold lgbm.</p>",
      "rawMarkdown": "Update: 0.2220 with variation of the 'ridge trick', 5fold lgbm.",
      "votes": 2,
      "replies": [
        {
          "id": 331505,
          "postDate": "2018-05-21T12:28:21.953Z",
          "content": "<p>What's a \"ridge trick\"?</p>",
          "rawMarkdown": "What's a \"ridge trick\"?"
        },
        {
          "id": 331511,
          "postDate": "2018-05-21T12:37:26.997Z",
          "content": "<p>Creating an oof prediction from a ridge model and using it as a feature - along the lines of <a href=\"https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0\">https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0</a></p>",
          "rawMarkdown": "Creating an oof prediction from a ridge model and using it as a feature - along the lines of https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0",
          "votes": 1
        },
        {
          "id": 331531,
          "postDate": "2018-05-21T13:05:26.643Z",
          "content": "<p>Cool. But does that really count as a single model? Though I'd concede at this point I'm just being pedantic. \"Single model\" is a bit of a vanity metric because of things like this, and it's really just about finding models that are individually good but blend together.</p>\n\n<p>My LGB has Ridges in it so I wasn't posting about it here. But I'm currently at 0.21611 CV / 0.2200 LB with a single LGB with multiple Ridge models inside. The best individual Ridge is at 0.22674 CV.</p>",
          "rawMarkdown": "Cool. But does that really count as a single model? Though I'd concede at this point I'm just being pedantic. \"Single model\" is a bit of a vanity metric because of things like this, and it's really just about finding models that are individually good but blend together.\n\nMy LGB has Ridges in it so I wasn't posting about it here. But I'm currently at 0.21611 CV / 0.2200 LB with a single LGB with multiple Ridge models inside. The best individual Ridge is at 0.22674 CV.",
          "votes": 5
        },
        {
          "id": 331551,
          "postDate": "2018-05-21T13:30:21.487Z",
          "content": "<p>Nice one - your gap cv - lb is a bit smaller than mine (~ 0.006). If you don't mind sharing, have you tuned the lgb a lot? I've been rolling mostly with parameters close to the ones from the public kernels so far.</p>",
          "rawMarkdown": "Nice one - your gap cv - lb is a bit smaller than mine (~ 0.006). If you don't mind sharing, have you tuned the lgb a lot? I've been rolling mostly with parameters close to the ones from the public kernels so far."
        },
        {
          "id": 331556,
          "postDate": "2018-05-21T13:39:06.703Z",
          "content": "<p>...I sure hope my smaller CV gap does not mean that my LB position will come crashing down at the end. :/</p>\n\n<p>I have spent some amount of time tuning the LGB by hand to the point where I can no longer easily improve it. I still think I have some room to non-easily improve it, perhaps with Hyperopt or something.</p>",
          "rawMarkdown": "...I sure hope my smaller CV gap does not mean that my LB position will come crashing down at the end. :/\n\nI have spent some amount of time tuning the LGB by hand to the point where I can no longer easily improve it. I still think I have some room to non-easily improve it, perhaps with Hyperopt or something."
        },
        {
          "id": 331561,
          "postDate": "2018-05-21T13:43:35.087Z",
          "content": "<p>@Konrad: Are you seeing 0.216XX CV as well? Do you have a lot of other features beyond those in the kernels?</p>",
          "rawMarkdown": "@Konrad: Are you seeing 0.216XX CV as well? Do you have a lot of other features beyond those in the kernels?"
        },
        {
          "id": 331596,
          "postDate": "2018-05-21T14:48:27.193Z",
          "content": "<p>Yes and yes - with the latter, i went to town with target encoding. I guess tuning is my next step, as i have barely touched it.</p>",
          "rawMarkdown": "Yes and yes - with the latter, i went to town with target encoding. I guess tuning is my next step, as i have barely touched it.",
          "votes": 1
        },
        {
          "id": 331642,
          "postDate": "2018-05-21T16:14:54.077Z",
          "content": "<p>Hey I have tried target encoding but it was not helping much as passing the categorical values into LGBM is doing. I think keeping the number of features high is helping the tree to build better.\nwhat do you think?</p>",
          "rawMarkdown": "Hey I have tried target encoding but it was not helping much as passing the categorical values into LGBM is doing. I think keeping the number of features high is helping the tree to build better.\nwhat do you think?",
          "votes": 1
        },
        {
          "id": 331665,
          "postDate": "2018-05-21T17:04:40.073Z",
          "content": "<p>Interesting - I have exact opposite experience to yours: once I mean-encoded the categoricals (created new numerical cols) and then dropped the originals, my score improved significantly. </p>",
          "rawMarkdown": "Interesting - I have exact opposite experience to yours: once I mean-encoded the categoricals (created new numerical cols) and then dropped the originals, my score improved significantly. ",
          "votes": 3
        },
        {
          "id": 331669,
          "postDate": "2018-05-21T17:09:44.993Z",
          "content": "<p>maybe I am doing somewhere wrong then. I tried using this <br>\n<a href=\"https://www.kaggle.com/tnarik/likelihood-encoding-of-categorical-features\">https://www.kaggle.com/tnarik/likelihood-encoding-of-categorical-features</a>. I am encoding all the categorical values .</p>",
          "rawMarkdown": "maybe I am doing somewhere wrong then. I tried using this  \nhttps://www.kaggle.com/tnarik/likelihood-encoding-of-categorical-features. I am encoding all the categorical values .",
          "votes": 1
        },
        {
          "id": 331682,
          "postDate": "2018-05-21T17:22:01.073Z",
          "content": "<p><a href=\"/konrad\">@konrad</a> How do you mean encode? You may be at risk of overfitting your CV. When I did mean encoding I ended up in overfit hell for a few days and it took ten submits to get to the bottom of it. So I'm on team LGB native encoding. I might try doing a second LGB with mean encoding though and average the two approaches and see if that helps... I'll calculate it within the CV folds to monitor overfitting better.</p>",
          "rawMarkdown": "@konrad How do you mean encode? You may be at risk of overfitting your CV. When I did mean encoding I ended up in overfit hell for a few days and it took ten submits to get to the bottom of it. So I'm on team LGB native encoding. I might try doing a second LGB with mean encoding though and average the two approaches and see if that helps... I'll calculate it within the CV folds to monitor overfitting better.",
          "votes": 2
        },
        {
          "id": 331687,
          "postDate": "2018-05-21T17:35:05.907Z",
          "content": "<p>yeah, my CV score was good enough but LB was worst.</p>",
          "rawMarkdown": "yeah, my CV score was good enough but LB was worst.",
          "votes": 2
        },
        {
          "id": 331695,
          "postDate": "2018-05-21T17:45:46.247Z",
          "content": "<p>@Peter: I did the mean-encoding inside the cv and added heavy regularization, so I think I'm fine (especially that cv - lb is consistently around 0.006, and changes are directionally correct).</p>",
          "rawMarkdown": "@Peter: I did the mean-encoding inside the cv and added heavy regularization, so I think I'm fine (especially that cv - lb is consistently around 0.006, and changes are directionally correct).",
          "votes": 2
        },
        {
          "id": 333343,
          "postDate": "2018-05-25T00:31:58.560Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 340394,
      "postDate": "2018-06-09T04:46:18.230Z",
      "content": "<p>10-fold LGB, LB 0.2214.</p>\n\n<p>Training a model including some image features. Hopefully it would give me a boost. Finger crossed.  </p>\n\n<p>Update: Including image featuers, 10-fold LGB, LB 0.2210</p>\n\n<p>6-15 Update: manage to improve my LGB to LB 0.2202</p>\n\n<p>6-19 Update: finally got an LGB with LB 0.2199, XGB with LB 0.2209.  </p>",
      "rawMarkdown": "10-fold LGB, LB 0.2214.\n\nTraining a model including some image features. Hopefully it would give me a boost. Finger crossed.  \n\nUpdate: Including image featuers, 10-fold LGB, LB 0.2210\n\n6-15 Update: manage to improve my LGB to LB 0.2202\n\n6-19 Update: finally got an LGB with LB 0.2199, XGB with LB 0.2209.  ",
      "replies": [
        {
          "id": 343000,
          "postDate": "2018-06-14T14:21:53.030Z",
          "content": "<p>10-fold LGB avg or min ？</p>",
          "rawMarkdown": "10-fold LGB avg or min ？"
        },
        {
          "id": 343002,
          "postDate": "2018-06-14T14:24:21.427Z",
          "content": "<p>10-fold LGB avg.</p>",
          "rawMarkdown": "10-fold LGB avg."
        }
      ]
    },
    {
      "id": 334596,
      "postDate": "2018-05-28T00:37:22.033Z",
      "content": "<p>Right now I have 3 models in parallel. A catboost at 0.2237, a LGB at 0.2228 and a Neural network at 0.2240. I'm really impressed by people achieving 0.222X with neural network on that data set. Maybe I'm missing something fundamental or some kind of clever encoding of categoricals.</p>\n\n<p>Not yet used additional FE from satellite data sets.</p>",
      "rawMarkdown": "Right now I have 3 models in parallel. A catboost at 0.2237, a LGB at 0.2228 and a Neural network at 0.2240. I'm really impressed by people achieving 0.222X with neural network on that data set. Maybe I'm missing something fundamental or some kind of clever encoding of categoricals.\n\nNot yet used additional FE from satellite data sets.",
      "replies": [
        {
          "id": 334699,
          "postDate": "2018-05-28T08:32:27.917Z",
          "content": "<p>As for me , I am just using simple label encoding for categorical data . </p>\n\n<p>I think the NN architecture plays a key role and clever archtitecture can bring huge improvement. </p>\n\n<p>I tend to rely on deep models just to avoid lot of features engeneering. </p>",
          "rawMarkdown": "As for me , I am just using simple label encoding for categorical data . \n\nI think the NN architecture plays a key role and clever archtitecture can bring huge improvement. \n\nI tend to rely on deep models just to avoid lot of features engeneering. ",
          "votes": 4
        },
        {
          "id": 334725,
          "postDate": "2018-05-28T09:30:27.170Z",
          "content": "<p>Interesting. <br>\nIt is a surprise for me to know the simple label encoding for categorical data works well.  I will give it a try. <br>\nThank for your comment</p>",
          "rawMarkdown": "Interesting.  \nIt is a surprise for me to know the simple label encoding for categorical data works well.  I will give it a try.  \nThank for your comment"
        },
        {
          "id": 335045,
          "postDate": "2018-05-29T04:06:51.050Z",
          "content": "<p>Thanks. I'll play around with some extra ideas. My architecture is fairly straightforward so far, ConvNN for texts, embeddings for categoricals then concat everything and pass it through dense layers.</p>\n\n<p>Maybe I can try going deeper with the text part.</p>",
          "rawMarkdown": "Thanks. I'll play around with some extra ideas. My architecture is fairly straightforward so far, ConvNN for texts, embeddings for categoricals then concat everything and pass it through dense layers.\n\nMaybe I can try going deeper with the text part."
        },
        {
          "id": 335870,
          "postDate": "2018-05-30T14:50:52.780Z",
          "content": "<p>Hi @Serigne!\nYou really just encode cat features to numbers? or use one-hot encoding? </p>",
          "rawMarkdown": "Hi @Serigne!\nYou really just encode cat features to numbers? or use one-hot encoding? ",
          "votes": 1
        },
        {
          "id": 335874,
          "postDate": "2018-05-30T15:05:14.377Z",
          "content": "<p>Hi Alex... Yes just simple encoding to numbers..Then embed--&gt; concat--&gt; dense layers just like Arroqc..</p>\n\n<p>But that's just for categorical  features. </p>\n\n<p>For text features I try to follow these  4 principles : <a href=\"https://explosion.ai/blog/deep-learning-formula-nlp\">Embed, Encode, Attend and Predict</a></p>",
          "rawMarkdown": "Hi Alex... Yes just simple encoding to numbers..Then embed--&gt; concat--&gt; dense layers just like Arroqc..\n\nBut that's just for categorical  features. \n\nFor text features I try to follow these  4 principles : [Embed, Encode, Attend and Predict][1]\n\n\n  [1]: https://explosion.ai/blog/deep-learning-formula-nlp",
          "votes": 7
        },
        {
          "id": 335972,
          "postDate": "2018-05-30T19:51:50.463Z",
          "content": "<p>Interesting article Serigne. Thanks. Will try to see if attention helps.</p>",
          "rawMarkdown": "Interesting article Serigne. Thanks. Will try to see if attention helps.",
          "votes": 1
        }
      ]
    },
    {
      "id": 334448,
      "postDate": "2018-05-27T12:26:31.887Z",
      "content": "<p>I got 0.223 using just train.csv and test.csv,How to using train_active.csv which has no labels</p>",
      "rawMarkdown": "I got 0.223 using just train.csv and test.csv,How to using train_active.csv which has no labels",
      "replies": [
        {
          "id": 334827,
          "postDate": "2018-05-28T14:17:12.837Z",
          "content": "<p>@ unexceptednull </p>\n\n<p>You perform feature engineering with user_id, item_id and merge it to train using 'user_id'\nIt is available in Benjamin's awesome kernel here <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm/output\">BENJAMIN FEATURE ENGINEERING WITH TRAIN ACTIVE </a></p>\n\n<p>You can also refer to this kernel for building RNN on these additional features <a href=\"https://www.kaggle.com/shanth84/rnn-detailed-explanation-0-2246\">RNN - 0.2246</a></p>\n\n<p>Hope this helps \nCheers\nShanth</p>",
          "rawMarkdown": "@ unexceptednull \n\nYou perform feature engineering with user_id, item_id and merge it to train using 'user_id'\nIt is available in Benjamin's awesome kernel here [BENJAMIN FEATURE ENGINEERING WITH TRAIN ACTIVE ][1]\n\nYou can also refer to this kernel for building RNN on these additional features [RNN - 0.2246][2]\n\nHope this helps \nCheers\nShanth\n\n  [1]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm/output\n  [2]: https://www.kaggle.com/shanth84/rnn-detailed-explanation-0-2246"
        }
      ]
    },
    {
      "id": 330469,
      "postDate": "2018-05-18T21:18:22.597Z",
      "content": "<p>I've started with catboost and currently sits at 0.2246 with just train.csv. Edit: (Improved it to 0.2237)</p>\n\n<p>I'm currently hoping to get further with LGB because of the ability to use sparse matrix.</p>",
      "rawMarkdown": "I've started with catboost and currently sits at 0.2246 with just train.csv. Edit: (Improved it to 0.2237)\n\nI'm currently hoping to get further with LGB because of the ability to use sparse matrix.\n\n"
    },
    {
      "id": 329704,
      "postDate": "2018-05-17T02:28:57.377Z",
      "content": "<p>0.2222 with one lgb model</p>",
      "rawMarkdown": "0.2222 with one lgb model"
    },
    {
      "id": 329226,
      "postDate": "2018-05-16T03:34:41.703Z",
      "content": "<p>0.2228 with one lgb</p>",
      "rawMarkdown": "0.2228 with one lgb"
    },
    {
      "id": 329137,
      "postDate": "2018-05-15T20:13:42.170Z",
      "content": "<p>LGBM with 5-fold CV 0.2231</p>",
      "rawMarkdown": "LGBM with 5-fold CV 0.2231",
      "replies": [
        {
          "id": 329166,
          "postDate": "2018-05-15T22:25:15.630Z",
          "content": "<p>Can I know which score for one fold? Is 5CV better than one fold only?</p>",
          "rawMarkdown": "Can I know which score for one fold? Is 5CV better than one fold only?"
        },
        {
          "id": 329168,
          "postDate": "2018-05-15T22:30:51.207Z",
          "content": "<p>In my experiment, it boosts about 0.001 level I believed, but I did this when my LB score is around 0.2250</p>",
          "rawMarkdown": "In my experiment, it boosts about 0.001 level I believed, but I did this when my LB score is around 0.2250",
          "votes": 1
        },
        {
          "id": 332664,
          "postDate": "2018-05-23T14:41:55.867Z",
          "content": "<p>@Strideradu, I was not able to run 5fold CV with Lightgbm, I'm having an error :\nValueError: Supported target types are: ('binary', 'multiclass'). Got 'continuous' instead.</p>\n\n<p>I'm using the following code :</p>\n\n<p>lgb_clf = lgb.cv(\n    lgbm_params, \n    training_data, \n    num_boost_round=1000, \n    nfold=5,\n    verbose_eval=100,\n    early_stopping_rounds=50)</p>\n\n<p>Did it work for you ?\nThanks!</p>",
          "rawMarkdown": "@Strideradu, I was not able to run 5fold CV with Lightgbm, I'm having an error :\nValueError: Supported target types are: ('binary', 'multiclass'). Got 'continuous' instead.\n\nI'm using the following code :\n\nlgb_clf = lgb.cv(\n    lgbm_params, \n    training_data, \n    num_boost_round=1000, \n    nfold=5,\n    verbose_eval=100,\n    early_stopping_rounds=50)\n\nDid it work for you ?\nThanks!"
        },
        {
          "id": 332668,
          "postDate": "2018-05-23T14:49:33.353Z",
          "content": "<p>I didn't use CV from lightgbm I use the KFold from sklearn to generate train and validation for each fold then just run fit for each fold data</p>",
          "rawMarkdown": "I didn't use CV from lightgbm I use the KFold from sklearn to generate train and validation for each fold then just run fit for each fold data",
          "votes": 2
        },
        {
          "id": 332671,
          "postDate": "2018-05-23T14:53:20.553Z",
          "content": "<p>Ok thanks a lot !</p>",
          "rawMarkdown": "Ok thanks a lot !"
        },
        {
          "id": 332778,
          "postDate": "2018-05-23T18:25:35.923Z",
          "content": "<p>I got the same error when I used Stratified Kfold by mistake. A regular KFold  did the trick</p>",
          "rawMarkdown": "I got the same error when I used Stratified Kfold by mistake. A regular KFold  did the trick"
        },
        {
          "id": 332782,
          "postDate": "2018-05-23T18:30:05.807Z",
          "content": "<p>@ Strideradu </p>\n\n<p>I used a 3 Fold in my RNN and got only a 0.0002 boost.  Here's what I do </p>\n\n<ul>\n<li>For each Kfold initialize a new RNN model and train it. </li>\n<li>Make predictions for each fold using this model.  Delete the model and move to next Fold </li>\n<li>A simple average of all predictions </li>\n</ul>\n\n<p>Is there a better way to do it? Any improvements you can suggest?</p>",
          "rawMarkdown": "@ Strideradu \n\nI used a 3 Fold in my RNN and got only a 0.0002 boost.  Here's what I do \n\n -  For each Kfold initialize a new RNN model and train it. \n -  Make predictions for each fold using this model.  Delete the model and move to next Fold \n -  A simple average of all predictions \n\nIs there a better way to do it? Any improvements you can suggest?"
        },
        {
          "id": 332787,
          "postDate": "2018-05-23T18:39:50.697Z",
          "content": "<p>@Antoine - \nWhat are your lgb_params?</p>",
          "rawMarkdown": "@Antoine - \nWhat are your lgb_params?"
        },
        {
          "id": 332791,
          "postDate": "2018-05-23T18:52:05.387Z",
          "content": "<p>@AmirH, I use 0.2 learning rate,  250 num_leaves. Increasing the number of leaves clearly helped a lot...</p>\n\n<p>@Shanth, I did use a normal Kfold... with the code above and it failed:\nkf = KFold(X.shape[0], n_folds=5, random_state=None, shuffle=False)</p>\n\n<p>I'm trying to implement it as <a href=\"/strideradu\">@strideradu</a> advised, after doing .tocsr() on my sparse Matrix.</p>\n\n<p>Thanks for the help</p>",
          "rawMarkdown": "@AmirH, I use 0.2 learning rate,  250 num_leaves. Increasing the number of leaves clearly helped a lot...\n\n@Shanth, I did use a normal Kfold... with the code above and it failed:\nkf = KFold(X.shape[0], n_folds=5, random_state=None, shuffle=False)\n\nI'm trying to implement it as @strideradu advised, after doing .tocsr() on my sparse Matrix.\n\nThanks for the help"
        },
        {
          "id": 332807,
          "postDate": "2018-05-23T19:37:05.640Z",
          "content": "<p>If you want to use stratified Kfold then you need to convert the probabilities to the binned result, some people just round all probability to 0 and 1, and here you can find code for binning <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/56346#330580\">https://www.kaggle.com/c/avito-demand-prediction/discussion/56346#330580</a></p>",
          "rawMarkdown": "If you want to use stratified Kfold then you need to convert the probabilities to the binned result, some people just round all probability to 0 and 1, and here you can find code for binning https://www.kaggle.com/c/avito-demand-prediction/discussion/56346#330580",
          "votes": 1
        },
        {
          "id": 332820,
          "postDate": "2018-05-23T20:32:55.627Z",
          "content": "<p>@Antoine - make sure your lgb_params have : objective=\"regression\"  or something of that kind, and metric = \"rmse\".\nI successfully run lgb.cv  with these parameters</p>",
          "rawMarkdown": "@Antoine - make sure your lgb_params have : objective=\"regression\"  or something of that kind, and metric = \"rmse\".\nI successfully run lgb.cv  with these parameters"
        },
        {
          "id": 332827,
          "postDate": "2018-05-23T20:57:28.390Z",
          "content": "<p>Yes that's what I'm using too</p>\n\n<p>lgbm_params =  {\n    'boosting_type': 'gbdt',\n    'objective': 'regression',\n    'metric': 'rmse',</p>\n\n<p>Not sure what I did wrong...</p>",
          "rawMarkdown": "Yes that's what I'm using too\n\nlgbm_params =  {\n    'boosting_type': 'gbdt',\n    'objective': 'regression',\n    'metric': 'rmse',\n\nNot sure what I did wrong..."
        }
      ]
    },
    {
      "id": 334810,
      "postDate": "2018-05-28T13:27:01.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 330081,
      "postDate": "2018-05-18T02:21:52.027Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 333907,
      "postDate": "2018-05-26T02:34:02.520Z",
      "content": "<p>thank you</p>",
      "rawMarkdown": "thank you"
    }
  ],
  "comments": [
    {
      "id": 340360,
      "author_name": "Liu Jilong",
      "author_url": "",
      "post_date": "2018-06-09T02:28:20.813000",
      "content": "<p>I got 0.2181 with a NN network with 10 folds (UPDATED)</p>",
      "votes": 17,
      "replies": [
        {
          "id": 340367,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-06-09T02:45:00.490000",
          "content": "<p>Wow, that's great!\nIs the feature amount effective? Or is the network special?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340372,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-09T03:01:40.660000",
          "content": "<p>Wow,can I ask that using MLP or RNN or CNN?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340378,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-09T03:23:17.113000",
          "content": "<p>Nothing special. I use four kind features: continous (contains ImageNet score), category, text with embedding, raw image pixel. Four submodels for four kinds feature, concatenated to one tensor, serval Dense Layers. Heavily use BatchNormalization and Dropout. I believe there still have some room to improve because I rarely did feature engineering.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 340381,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-06-09T03:32:07.747000",
          "content": "<p>Thank you for sharing!\nMy nn is still 0.2222(TT)  I will do my best from here!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340382,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-09T03:42:58.173000",
          "content": "<p>My LGB is still 0.2205 and I have no idea how to improve it. Poor feature engineer skills.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 340388,
          "author_name": "SubikashPal",
          "author_url": "",
          "post_date": "2018-06-09T04:12:33.540000",
          "content": "<p>Wah!!! Great.. R u using self-embedding ? \"Raw image pixel\" - is it like using the raw pixel value of the image..How r u using it in the NN model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340389,
          "author_name": "Totoro (ごめん)",
          "author_url": "",
          "post_date": "2018-06-09T04:13:28.803000",
          "content": "<p>Congrats @Liu Jilong,</p>\n\n<blockquote>\n  <p>raw image pixel </p>\n</blockquote>\n\n<p>Do you mean the raw image without through any pretrained model? <br>\nIf yes, that is impressive. How long does it take to train?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340391,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-09T04:19:22.943000",
          "content": "<p>Thanks for sharing, I have confused that you train four model for for  four kinds feature,then a final model based on the four predictions? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340400,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-09T04:58:48.003000",
          "content": "<p>Just like any other image network, serval Conv2D layers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340401,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-09T04:59:20.403000",
          "content": "<p>UPDATED: about 15min per epoch, 2.5h per fold.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340403,
          "author_name": "SubikashPal",
          "author_url": "",
          "post_date": "2018-06-09T05:08:52.107000",
          "content": "<p>Thanks Liu... Any insight on text pre-processing / embedding?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340405,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-09T05:19:17.890000",
          "content": "<p>Just as many kagglers shared, I use self-trained word2vec embedding.  I have no idea about russian, so no text pre-processing is used.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340407,
          "author_name": "Hikkiiiiiiiii",
          "author_url": "",
          "post_date": "2018-06-09T05:26:39.460000",
          "content": "<p>Thank you for sharing.</p>\n\n<p>My situation is the opposite of you. I just started trying NN today,  and it didn't look good. </p>\n\n<p>Very valuable experience, thanks :).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340408,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-09T05:31:06.123000",
          "content": "<p>Many tricks to make NN work well. Good luck.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340410,
          "author_name": "SubikashPal",
          "author_url": "",
          "post_date": "2018-06-09T05:32:25.727000",
          "content": "<p>Thanks. Appreciate your help...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340430,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-09T06:21:50.623000",
          "content": "<p>Thanks. wish to see experience sharing for NN model after this competition.I have no idea to train a good NN model.My NN model is poor score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340688,
          "author_name": "Webber",
          "author_url": "",
          "post_date": "2018-06-10T02:41:29.237000",
          "content": "<p>Thanks for your sharing. I have a memory problem of training a CNN on images. Since I don't have enough memory to load all images, currently I use a generator to read images batch by batch from disk and feed them to the CNN, which makes the training process very slow (1 hour per epoch). Is there any better way to make it faster?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340691,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-10T03:20:04.987000",
          "content": "<p>I also tried fit_generator, but fit_generator's result is worse than fit, which is strange and I did not find the reason. As for time-consuming problem, you can try fit_generator's multi-processing parameter.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340699,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-06-10T03:47:12.887000",
          "content": "<p>Wondering 15min/epoch for what kind of GPU?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340702,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-10T04:02:07.473000",
          "content": "<p>gtx1070</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341040,
          "author_name": "YoungLamb",
          "author_url": "",
          "post_date": "2018-06-10T21:38:04.990000",
          "content": "<p>Hi Liu,\nThanks for your sharing, it's very encouraging. I wonder if you did not use  fit_generator, how can you load 1300k images to memory? Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341651,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2018-06-12T03:03:38.347000",
          "content": "<p>Hi Jilong, is your NN model end2end training? Did you use keras concatenate four kinds feature layers? \nThx for your sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 342079,
          "author_name": "To Train Them Is My Cause",
          "author_url": "",
          "post_date": "2018-06-12T19:46:52.283000",
          "content": "<blockquote>\n  <p>I also tried fit_generator, but fit_generator's result is worse than fit, which is strange and I did not find the reason. As for time-consuming problem, you can try fit_generator's multi-processing parameter.</p>\n</blockquote>\n\n<p>Fit automatically does a few things that are critical for NNs ( such as randomizing input per epoch) that you have to explicitly do with your own fit_generator. Fit_generator automatically should load info on a new thread- so unless you have something crazy expensive in your generator (like a garbage collecting call) it shouldn't be a bottleneck at all-loading is faster than training.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 343689,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2018-06-15T20:30:19.343000",
          "content": "<p>I've had several issues with using the image data. I think one of the most major ones is that it turns out keras cannot do multiprocessing on windows so it isnt actually loading in images properly beforehand. Second hand is I am trying to do the image resize and padding in line. This ends up costing roughly .05 per image. I worked a little around this using concurrent. Futures, but a full epoch still takes much much more than 15 minutes. Closer to 2-3 hours.. One avenue I spent some time researching is rewriting the images as a hdf5 dataset, but I'm unsure how to load it in alongside my other features, but that has various issues as well regarding shuffling and getting everything aligned</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344203,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-17T07:59:08.747000",
          "content": "<p>Train with image pixels end to end is indeed tough, considering the million count. I did some trick to make it works, also i encountered some strange questions. I will share my workflow after competition ends and hoping for answers.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344290,
          "author_name": "jetou Xu",
          "author_url": "",
          "post_date": "2018-06-17T14:33:53.243000",
          "content": "<p>thank you!  I am looking forward to your model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344749,
          "author_name": "To Train Them Is My Cause",
          "author_url": "",
          "post_date": "2018-06-18T16:10:00.313000",
          "content": "<blockquote>\n  <p>I think one of the most major ones is that it turns out Keras cannot do multiprocessing on Windows</p>\n</blockquote>\n\n<p>This is interesting. Can you do your own threading process? maybe something with threading and queue like:</p>\n\n<pre><code>        def data_loader(q,):\n            for start in range(0, len(train_imgs), batch_size):\n                #do your batch creation code here               \n                q.put([feat_batch,  y_batch])\n        q = queue.Queue(maxsize=q_size)\n        t1 = threading.Thread(target=data_loader, name='DataLoader', args=(q,))\n        t1.start()\n        t1.join()\n</code></pre>\n\n<p>And then just read the generator batches from the queue? I can't really check ( I'm running Linux atm).\nIf that's not possible, I recommend Linux...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 344921,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2018-06-18T23:26:44.383000",
          "content": "<p>I really should move over to linux at some point, but I am eventually planning on dedicating one system to work like this and the other for games and less demanding functions so dont want to dual boot right away. That being said, I have found some way to kind of work around this using the concurrent.futures builtin python to load in the images with threading and I also preprocessed the images with padding and downscaling so I dont have to do that function within the generator. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345074,
          "author_name": "Ee Kin Chin",
          "author_url": "",
          "post_date": "2018-06-19T06:41:18.443000",
          "content": "<p>Hi Jilong, In my efforts to include text + image + tabular features, I find that it heavily overfits even before it finishes one epoch, does heavy use of batch normalization help to tackle overfitting? Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345152,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-19T09:28:40.913000",
          "content": "<p>Yes, i heavily use BN and dropout, you may try add a BN before every Dense Layer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345417,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2018-06-19T20:57:11.850000",
          "content": "<p>@Liu Jilong, there are few rows in test set where images are missing, how are you handling those when you say you are using imagent features?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345512,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-20T01:14:15.283000",
          "content": "<p>I just give them an all-black image for row image pixels and zero score for image net scores. It may have some approaches to improve, but i do not have enough time to handle this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 346569,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2018-06-22T01:17:55.993000",
          "content": "<p>Have you tried training a model purely on images? Do you know roughly what training or validation loss that can get down to? I have been working on this for a while and cant seem to get much out of the images. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346622,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-06-22T04:09:39.683000",
          "content": "<p>I tried several CNN models to learn only on images to predict deal probably. I used relu on last several dense layer, and I found out last dense layer (before Dense (1)) predict all zeros, which give me 0.133  rmse loss...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 346640,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2018-06-22T04:56:09.560000",
          "content": "<p>Mine are consistently getting to .066ish MSE but dont seem to contribute anything to my combined model. Havent inspected what the outputs or the weights areFor a while I thought there was something wrong with my generator but it seems to work fine without the images just fine so dont think that is the issues</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 338921,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-06-06T01:33:30.123000",
      "content": "<p>I'm hitting .2190 with a 5-fold averaged LGBM</p>",
      "votes": 9,
      "replies": [
        {
          "id": 338925,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-06T01:49:48.870000",
          "content": "<p>That's impressive, Congratulation!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 338928,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-06T01:55:08.110000",
          "content": "<p>Wow,you must find some amazing features~</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 338929,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2018-06-06T01:57:05.233000",
          "content": "<p>Seems like there is a magic feature for the big jump.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 338930,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-06T02:09:09.757000",
          "content": "<p>Thanks! :)</p>\n\n<p>I'm using 100+ engineered tabular features, but there are definitely some that have a particularly substantial contribution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 338939,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-06-06T02:34:12.020000",
          "content": "<p>good job,did you use target encoding?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 339017,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2018-06-06T06:43:56.893000",
          "content": "<p>Wow, do you mean you didn't use text features at all?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 339162,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-06T12:37:55.033000",
          "content": "<p>Sorry, I do use text features in addition to the tabular ones - could have been more clear</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 339181,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2018-06-06T13:13:56.817000",
          "content": "<p>Thank you a lot for a clarification!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 340151,
          "author_name": "gufra",
          "author_url": "",
          "post_date": "2018-06-08T13:24:03.263000",
          "content": "<p>great work! did you use the train_active csv ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 333988,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2018-05-26T08:13:34.680000",
      "content": "<p>LGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2200 (updated)</p>\n\n<p>XGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2216 (updated)</p>\n\n<p>NN (Categorical, Numerical, W2V, Images meta-data + ImageNet scoring): 0.2198 (updated)</p>\n\n<p>All with CV5.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 334098,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-26T14:27:42.823000",
          "content": "<p>may i ask what is image meta?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334132,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-05-26T16:01:46.743000",
          "content": "<p>What numerical features are working for you?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334182,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2018-05-26T17:42:30.647000",
          "content": "<p>Images meta-data is width, height. I've also included the average of VGG16, VGG19, InceptionV3, Xception, ResNet50 ImageNet scoring for each image but it does not really help. It only improves by 0.0002. And for numerical, only the basic ones, log1p(item_seq_number) and log1p(price). I'm trying different target-encoding but currently, as already reported below, it overfits too much.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 336600,
          "author_name": "Daniel S.",
          "author_url": "",
          "post_date": "2018-06-01T02:28:34.793000",
          "content": "<p>does NIMA stand for Neural Image Assessment?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 336711,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2018-06-01T06:40:25.510000",
          "content": "<p>Yes, I used this implementation:\n<a href=\"https://github.com/titu1994/neural-image-assessment\">https://github.com/titu1994/neural-image-assessment</a></p>\n\n<p>It improves LGB LB by 0.0002 and XGB by 0.0003 (only). I included mean/std as features from MobileNet and InceptionResNetv2. It does not improve NN (which should mean the NLP part is better).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 344107,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "2018-06-17T02:10:41.900000",
      "content": "<p>LGB 10-fold 0.2179</p>",
      "votes": 5,
      "replies": [
        {
          "id": 344108,
          "author_name": "Siyuan Dang",
          "author_url": "",
          "post_date": "2018-06-17T02:14:32.873000",
          "content": "<p>Excellent achieve！Can you show some black magic you used？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344128,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-06-17T03:33:02.557000",
          "content": "<p>no magic,features are almost from kernels like aggregations,images,tfidf,tsvd,impute...just try to use more features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344129,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-06-17T03:36:12.897000",
          "content": "<p><a href=\"/senkin13\">@senkin13</a>: Is that a CV score or a LB score? Does your model contain a Ridge?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344130,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-06-17T03:39:13.767000",
          "content": "<p>LB score without ridge,I prefer to use ridge for stacking instead of lgb features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344132,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-17T03:52:27.573000",
          "content": "<p>What do you mean by tsvd and impute? Any reference kernels or information?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344134,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-06-17T04:07:04.590000",
          "content": "<p>here: <a href=\"https://www.kaggle.com/krithi07/baseline-model-with-new-features\">https://www.kaggle.com/krithi07/baseline-model-with-new-features</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 344146,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-17T04:52:48.563000",
          "content": "<p>Congrats @Senkin13, which image features did you use. The ones I've got makes my scores worse... both CV and LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344172,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-06-17T06:21:42.553000",
          "content": "<p>A few simple image meta features,not too much improvement compared to Peter Hurford(0.002 improvment)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344178,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-06-17T06:47:15.090000",
          "content": "<p>I know you called my improvement from images impressive, but I'm not coming anywhere close to 0.2179 single model LGB... Hey, I wonder what would happen if we teamed up and pooled our feature engineering?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344179,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-17T06:52:11.893000",
          "content": "<p>Hello, I have tried tsvd to tfidf of description and title,but my score became to worse,I choose 32 component tsvd.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344185,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-17T07:13:08.823000",
          "content": "<p>Maybe I dont choose the right n_ component </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344187,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-06-17T07:14:05.520000",
          "content": "<p>@Peter Hurford Your proposal is attractive ,there still have 4 days ,let's wait and see.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344188,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-06-17T07:15:09.567000",
          "content": "<p>@Johnny You can always check sum of explained_variance_ratio_. If it is very low, it means you not only lose noise but also information</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344189,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-06-17T07:16:47.913000",
          "content": "<p>@Johnny Liu I only get a little improvement from tsvd,yes should try different n_component and tfidf parameters.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344191,
          "author_name": "Cheng",
          "author_url": "",
          "post_date": "2018-06-17T07:18:52.707000",
          "content": "<p>very solid work! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344201,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-17T07:47:13.687000",
          "content": "<p>thanks @AhmetErdem <a href=\"/senkin13\">@senkin13</a>, now I sum explained_variance_ratio_ to find the best component</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 338049,
      "author_name": "Andrii Gogniat",
      "author_url": "",
      "post_date": "2018-06-04T10:21:58.890000",
      "content": "<p>LGBM 10-fold avg. - 0.2203 (categorical, numerical, target encoding, text and features predicted using all csv data)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 338078,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-04T11:43:39.543000",
          "content": "<p>Hello, does  target encoding contribute your  score a lot?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338083,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-06-04T12:00:17.947000",
          "content": "<p>When you say \"features predicted using all csv data\", does that mean that you found some interesting features leveraging train_active/test_active/periods_train/periods_test beyond the ones identified by Benjamin Minixhofer?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 338100,
          "author_name": "Andrii Gogniat",
          "author_url": "",
          "post_date": "2018-06-04T12:28:06.580000",
          "content": "<p>Previous lgbm was at 0.2215 but TE was not the only features added so it is hard to say how much TE contributed to score. Still half of features in top-100 are TEs. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 338104,
          "author_name": "Andrii Gogniat",
          "author_url": "",
          "post_date": "2018-06-04T12:32:30.757000",
          "content": "<p>@Larry Currently I am using only 'times up' and 'days up' predictions, but I think it is worth trying to fill other missing values by predicting them from all csv data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338111,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2018-06-04T12:48:16.607000",
          "content": "<p>@Andrii Do you mean you trained a model to estimate 'times up' and 'days up' from the other features? Or you just used the features described in Benjamin's kernel?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338116,
          "author_name": "Andrii Gogniat",
          "author_url": "",
          "post_date": "2018-06-04T12:59:20.747000",
          "content": "<p>@Dmitriy I am using both - averages from kernel and model predictions.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 338427,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-05T03:24:17.920000",
          "content": "<p>wow, thanks for your replying.I have not try to using target encoding now</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 333337,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2018-05-24T23:45:27.317000",
      "content": "<p>NN: <br>\n0.2218 Fasttext <br>\n0.2215 Self-training embedding  </p>\n\n<p>Thanks @Dieter</p>",
      "votes": 5,
      "replies": [
        {
          "id": 333362,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-25T02:50:23.843000",
          "content": "<p>Wow your fasttext result is great. Are u using the crawl one or wiki one? Did you play a lot with different NN structures?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333366,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2018-05-25T03:09:20.937000",
          "content": "<p>I use wiki 300D. <br>\nI tried some NLP parts such as: RNN, Bi-GRU, Bi-LSTM, CNN, ... (with Attention and without Attention).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333367,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-25T03:16:23.640000",
          "content": "<p>@Ghost can you tell us how much do you think contributes the NLP part to your score? Let's say you take a naive approach with a single gru layer, how much would that worsen your result?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333372,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2018-05-25T03:35:58.927000",
          "content": "<p>@Dieter <br>\nI am not sure that I understand your question correctly. Please feel free to ask. <br>\nWithout NLP, my model is really bad. It is arround 0.2240. Then I tried to add NLP.</p>\n\n<p>I have not ran all of the NLP parts above, so I don't have results of each part. <br>\nTo determine which NLP part is good or not. I have a training plan as follows:  </p>\n\n<ol>\n<li>Select fixed training set and validation set. Ex: First fold of 5-Fold split</li>\n<li>Replace NLP part and run for some epochs (Ex: 5 epochs)  </li>\n</ol>\n\n<p>You will have list of cross-validation of model in different NLP architectures (GRU, CNN, LSTM,...).\nSelect the best one and go fully training.  </p>\n\n<p>My top 3 best NLP parts are:\nCapsule --&gt; CNN --&gt; Bi-GRU --&gt; ... </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 333377,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-25T03:52:36.307000",
          "content": "<blockquote>\n  <p>Without NLP</p>\n</blockquote>\n\n<p>you mean skipping text completely?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333379,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2018-05-25T03:56:39.793000",
          "content": "<p>Yeap. <br>\nWithout NLP, I only have numeric and categorical features.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333380,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-25T03:58:59.237000",
          "content": "<p>Thank you, that helps :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333499,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-05-25T09:17:35.427000",
          "content": "<p>0.2218 is this a single model or average of 5 fold?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334030,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-26T10:47:38.253000",
          "content": "<p>@Ghost you must be very talented. You joined kaggle 2 days ago, and already have models that score top 50 in LB. I can only salute.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 334099,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-05-26T14:32:58.730000",
          "content": "<p>0.2240 without Text is a strong model I think. \nDid you use image meta features in this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334592,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2018-05-27T23:48:41.247000",
          "content": "<p>Sorry for the late reply. I have not tried with image meta-features. The result is after CV.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 329320,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-05-16T08:25:08.463000",
      "content": "<p>I am working exclusively with NN for now..</p>\n\n<p>So far,  best single NN model : LB = 0.2246 ( 5 folds CV RMSE : 0.2215xxx ), no image features either</p>\n\n<p>The model was scoring 0.244xxx on LB when I started it ,   5 days ago...(Some ideas from kernels and discussion topics helped me in the improvement) </p>\n\n<p><strong>Edit :</strong>  The NN model is now scoring 0.2223 on LB </p>",
      "votes": 6,
      "replies": [
        {
          "id": 330506,
          "author_name": "Totoro (ごめん)",
          "author_url": "",
          "post_date": "2018-05-19T01:47:12.347000",
          "content": "<p>Nice. <br>\nMy model now is 0.2222 LB. Can I know which pretrained embedding are you using?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330560,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-19T05:30:45.840000",
          "content": "<p>Thanks :)</p>\n\n<p>I am using self-trained embeddings (following the idea suggested by Dieter )</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330561,
          "author_name": "Totoro (ごめん)",
          "author_url": "",
          "post_date": "2018-05-19T05:32:38.377000",
          "content": "<p>Amazing guy. <br>\nI could not get the high result with this embedding :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330562,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-05-19T05:34:16.830000",
          "content": "<p>@ Totoro - Did you try both options ( trainable = True and Trainable = False )??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330567,
          "author_name": "Totoro (ごめん)",
          "author_url": "",
          "post_date": "2018-05-19T05:51:58.893000",
          "content": "<p>I use pretrained embedding, so trainble = False. I tried trainable = True, but it is worse for me</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330568,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-19T05:54:44.660000",
          "content": "<p>If it's any consolation...the biggest improvements in my model  are not from the embedding itself :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330571,
          "author_name": "Totoro (ごめん)",
          "author_url": "",
          "post_date": "2018-05-19T05:59:16.977000",
          "content": "<p>Interesting <br>\nHope to see your model when the competition ends :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334014,
          "author_name": "Mr C",
          "author_url": "",
          "post_date": "2018-05-26T09:58:11.267000",
          "content": "<p>Thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 342176,
      "author_name": "pocket",
      "author_url": "",
      "post_date": "2018-06-13T03:26:59.310000",
      "content": "<p>Single model LightGBM 0.2199, with 5seed-avg 0.2195</p>",
      "votes": 3,
      "replies": [
        {
          "id": 343951,
          "author_name": "Siyuan Dang",
          "author_url": "",
          "post_date": "2018-06-16T15:11:08.827000",
          "content": "<p>could you show some hints about how you achieve this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345827,
          "author_name": "pocket",
          "author_url": "",
          "post_date": "2018-06-20T14:44:08.317000",
          "content": "<p>Sure, it will be in my write up if I get gold :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 339507,
      "author_name": "Nooh",
      "author_url": "",
      "post_date": "2018-06-07T04:21:41.537000",
      "content": "<p>Finally! Able to achieve 2210-XGB and 2217-LGB. No image features though.\nI was stuck earlier and started from scratch, seems the right decision. Best of luck guys.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 339528,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-07T04:58:37.717000",
          "content": "<p>Wow,Congratulation! Please sharing your solution after competition,I'm every interested how to achieve the score.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 339547,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-06-07T05:56:38.403000",
          "content": "<p>Is your XGB model using the same features as LGBM?  I am finding that my LGBM models are both faster and perform better.  I am able to get to 2217-LGB but not even 2217-XGB.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 339568,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-07T06:27:00.560000",
          "content": "<p>Congratulations @Nooh on getting those scores. Any hints you can share?</p>\n\n<p>@Larry Freeman, I have the same issue with you. My best LGBM is at 0.2215 but xgb is 0.2232. Both with single runs. The speed difference is not surprising as LGBM is always faster than xgb but in my experience their performance is comparable. Hence my surprise that I am not able to get xgb to  match lgb performance here.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 339648,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-07T09:18:25.063000",
          "content": "<p>Thanks Liu, though I'm really bad in documentation and my code is a mess right now but will surely post a brief if I don't end up over-fitting. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 339649,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-07T09:20:37.430000",
          "content": "<p>Yes Larry, I can say XGB has 98% same features as of my LGB. \nMy XGB took around 15-hours to train. :/</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 339651,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-07T09:28:44.263000",
          "content": "<p>Thanks a lot YaGana Sheriff-Hussaini!</p>\n\n<p>Nothing special in features though. I'm using nearly all the features used by <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">this kernel</a>. Played with some statistical features around deal_prob and Price. And another tip, try to thoroughly analyze image_top_1 and item_seq_no; I handcrafted something using them. After all that I tuned the hyperparams.</p>\n\n<p>Best of luck buddy!</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 339863,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-07T19:44:29.180000",
          "content": "<p>Added OOF ridge in LGB and now hitting 2211. :) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 340353,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-09T01:29:08.370000",
          "content": "<p>Thanks @Nooh and best of luck to you too.</p>\n\n<p>I just managed to get my lgb to LB=0.2212 but still have some optimization to do. You are right the xgb takes a long time for me as well, hence I am trying to optimize my features with LGB before running the xgb model again. Like you, I use the same features for  both.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340584,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-09T17:43:15.437000",
          "content": "<p>Okay, so I'm convinced now that text features are bringing in some improvements in my models. Text normalization seems promising in my case. I just did some very basic normalization on title+desc and it is progressive. I should quote the original author here i.e; ololo (who stood 5th in the previous avito competition). I just stumbled upon his code and there is a lot to scavenge from his repo <a href=\"https://github.com/alexeygrigorev/avito-duplicates-kaggle\">here</a> </p>\n\n<p>I used <a href=\"https://raw.githubusercontent.com/alexeygrigorev/avito-duplicates-kaggle/master/prepare_text_features.py\">this</a> script to get some text features. Hope this helps.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 340759,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-10T08:21:03.113000",
          "content": "<p>Thanks @Nooh. I have a simple text normalization but after reading your comment, I tried ololo's i.e. \"def normalize(text):\" and I get the following error:</p>\n\n<blockquote>\n  <p>File \"../src/script.py\", line 98\n      text = re.sub(ur'(?&lt;=[^а-я])(' + shortenings + ')[.]', ur'\\1 ', text)\n                                   ^\n  SyntaxError: invalid syntax</p>\n</blockquote>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 340775,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-10T09:24:03.940000",
          "content": "<p>Hi YaGana Sheriff-Hussaini. His code is for Python 2.x\nAre you on Python3x? If so, make sure you port it accordingly.\nI didn't used it as is actually instead get ideas from his code and made my own normalizations.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 341220,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-11T08:22:41.477000",
          "content": "<p>@Nooh you are right Ololo's code is for python2.x. I am using python3 and wrote a better text normalization than the basic one I had. This one improved both my CV and LB scores.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 341293,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-11T11:25:21.387000",
          "content": "<p>Great bro!, happy to hear that. I'm making up some image meta features' code. I hope that will bring in some improvement. If so, I'll share the code+data. \nFingers crossed. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 341351,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-11T13:25:45.630000",
          "content": "<p>I hope you get a better results with image features than I do. I have a weak machine hence using a different LGBM model with reduced features that includes image ones but only scores 0.21777 CV and 0.2233 LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341522,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-11T18:12:12.117000",
          "content": "<p>Hopefully yeah. I'm done with test set images and training set is running for 1 day now. I think it will take a couple of more hours.\nHow many models have you ensembled to get the current lb score you have? I have 4 of them, 3 lgbms and 1 xgb.\nTry training different lgbms with different features and different hyperparams. You can get all your features in that way i guess. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341598,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-11T21:21:31.863000",
          "content": "<p>Six.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344448,
          "author_name": "high&mean",
          "author_url": "",
          "post_date": "2018-06-18T03:05:43.033000",
          "content": "<p>which image meta features will you care for？</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 337277,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-06-02T12:40:10.413000",
      "content": "<p>We are at</p>\n\n<ul>\n<li>LGB 0.2200 </li>\n<li>XGB 0.2213 </li>\n<li>Catboost 0.2208</li>\n<li>NN 0.2194</li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 337303,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-02T14:09:27.843000",
          "content": "<p>Impressive @Dieter. Are your results for single or N-Folds. If so is it 3 or 5 folds?</p>\n\n<p>My best single model is LGBM with LB = 0.2217 single run. I have not run it for N-Folds yet.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 337916,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-04T03:47:50.800000",
          "content": "<p>5 fold</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 338044,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-06-04T10:04:21.810000",
          "content": "<p>What features are you using for LGB .2206?  That is very impressive.   </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338101,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-04T12:29:12.633000",
          "content": "<p>@Dieter, if you can share , what is the difference in score between a single run and 5-Fold?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338426,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-05T03:15:44.950000",
          "content": "<p>@Larry, I guess some part of the good score are image features</p>\n\n<p>@YaGana Sheriff-Hussaini difference is about 0.001</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 342435,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-13T14:10:32.643000",
          "content": "<p>@Dieter, the difference for me is 0.0003</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 335359,
      "author_name": "Larry Freeman",
      "author_url": "",
      "post_date": "2018-05-29T16:31:37.513000",
      "content": "<p>My best so far is an LGB model with .2216 (recently improved) after a 10-fold average.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 335485,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-05-29T20:05:05.627000",
          "content": "<p>Good job! Your public score is 0.2195, how did you earn those 0.002 on your score ? with some blending ?\nSame thing @AlexTru there is a significant gap between your LB and this LGB model. I'd understand if you don't want to share...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335487,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-05-29T20:09:31.143000",
          "content": "<p>I used lessons learned from here:\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 335510,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-05-29T20:50:30.400000",
          "content": "<p>Thanks for sharing this link !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335518,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-29T21:21:05.710000",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> Are you using train/test-time augments? I assume you're not doing psuedo-labeling?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335542,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-05-29T22:17:33.987000",
          "content": "<p>I am making attempts to use multiple methods (train/test augmentation, pseudo-labeling, and (less robust) cv + stacking framework).  I think its mostly stacking that's contributing to  score.    </p>\n\n<p>@PeterHurford, congrats to you on your excellent score.  Any reason why you assume that I'm not doing pseudo-labeling?  Is this something that you have tried?  I have not gotten too much mileage from it yet but I am working on it.  :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335546,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-29T22:23:42.073000",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> Thanks for the congrats - I hope it holds! :)</p>\n\n<p>...My assumption is that pseudo-labeling is only worthwhile when you have really accurate models. We were lucky to be in that case in the toxic competition, though I stupidly rejected the idea and didn't try it there. With RMSE in the &gt;0.2 range, I don't think our models are accurate enough, so I'm rejecting the idea again. I have not tried it. I hope it's the right call this time. ;)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 335549,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-29T22:27:27.060000",
          "content": "<p>additionally train and test distributions are quite close, which reduces the benefit from pseudo labeling</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 335553,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-05-29T22:30:46.693000",
          "content": "<p>Thanks, @Peter and @Dieter.  I am still developing my intuition about pseudo labeling.  Cheers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335555,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-29T22:31:39.537000",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> I'd also be really curious to hear if augmentation is helpful. My assumption is that with 10x the data here compared to toxic, plus the impact of categoricals/images in addition to text, the benefits of adding translated data would be much lower now. ...But I could also be wrong about that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 335556,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-05-29T22:35:47.767000",
          "content": "<p>I share your assumption on data augmentation.  At this point, I'm trying everything that I can think of just to see what happens and to see if it gives me any other ideas to try.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 335559,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-29T22:38:52.247000",
          "content": "<p><a href=\"/larryfreeman\">@larryfreeman</a> It's awesome to try stuff. This competition is overwhelming in the amount of things there are to try. I hope you get lucky and one of these ideas pans out for you. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 335584,
          "author_name": "Larry Freeman",
          "author_url": "",
          "post_date": "2018-05-30T00:50:49.670000",
          "content": "<p>Thanks, @Peter.  Best of luck to you too.  :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 328983,
      "author_name": "ms",
      "author_url": "",
      "post_date": "2018-05-15T13:29:53.330000",
      "content": "<p>Got 0.2222 with NN just train and test</p>",
      "votes": 3,
      "replies": [
        {
          "id": 329135,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-15T20:13:18.727000",
          "content": "<p>So the NN is for the NLP of description or more like for the classifier for different features just like LGBM</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329244,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-05-16T04:40:09.050000",
          "content": "<p>I guess both. At least I use it for both. It makes it much easier for me to connect different types of information. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 341626,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2018-06-12T00:36:43.723000",
      "content": "<p>After a lot of hard work I got my single LGB (with 5 fold) down to 0.2200 on LB. Yay !</p>",
      "votes": 4,
      "replies": [
        {
          "id": 341667,
          "author_name": "Nooh",
          "author_url": "",
          "post_date": "2018-06-12T04:12:58.910000",
          "content": "<p>Good going buddy!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 342197,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-13T04:31:23.570000",
          "content": "<p>Good work, buddy!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 340365,
      "author_name": "Hikkiiiiiiiii",
      "author_url": "",
      "post_date": "2018-06-09T02:38:39.960000",
      "content": "<p>0.2190 10-fold LGB</p>",
      "votes": 4,
      "replies": [
        {
          "id": 340375,
          "author_name": "Liu Jilong",
          "author_url": "",
          "post_date": "2018-06-09T03:15:45.817000",
          "content": "<p>Do you known what is lb difference between 5-fold and 10-fold?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340379,
          "author_name": "KALE",
          "author_url": "",
          "post_date": "2018-06-09T03:23:30.757000",
          "content": "<p>Approximately 0.0001-0.0002</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340406,
          "author_name": "Hikkiiiiiiiii",
          "author_url": "",
          "post_date": "2018-06-09T05:22:56.690000",
          "content": "<p>@Liu Jilong, I've only tried 1fold and 10fold.</p>\n\n<p>But I think it should be around <code>0.0002</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340427,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-09T06:19:30.583000",
          "content": "<p>Congratulation!! you guys are so cool. upvote for you guys.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340689,
          "author_name": "Webber",
          "author_url": "",
          "post_date": "2018-06-10T02:44:00.383000",
          "content": "<p>Wow, this small gap is quite impressive. Mine are alway around 0.004 (20 times of yours)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 342452,
      "author_name": "Stanislav K.",
      "author_url": "",
      "post_date": "2018-06-13T14:39:15.370000",
      "content": "<p>Silly question.</p>\n\n<p>What do you guys exacly mean when you say, for example, 10-fold LGB?\nYes, you can use CV to better tune hyperparameters. Particularly, n_estimators. Is that what you mean? </p>\n\n<p>And when you find optimal number of estimators do you retrain on whole dataset? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 342463,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2018-06-13T15:01:05.843000",
          "content": "<p>I guess people mean average the prediction of 10 LGB models in CV instead of retrain and get one model.  The advantage of avg is obvious :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 342521,
          "author_name": "Stanislav K.",
          "author_url": "",
          "post_date": "2018-06-13T16:28:38.163000",
          "content": "<p>Oh I see.\nYou can also achieve that with different random seeds. \nBut subsets of training data probably make models more different than random seeds. And that is good. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 342532,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2018-06-13T16:49:00.773000",
          "content": "<p>But that’s not cv tho. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 342578,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-06-13T18:29:51.537000",
          "content": "<p>10-fold is not used for hyper parameters tuning but just as CV technic. When running 10 fold, it's similar to running 10 times your LGB with 10 different seed and then doing an average. The advantage of using 10 fold, is that you can be sure that all data were used for train and for validation. (which might not be the case if you take random seeds)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 342610,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-13T19:50:04.500000",
          "content": "<p>I hope to clear up some confusion here - it's not a silly question at all. It's common conflation of terminology on kaggle to refer to partitioned fold averaging as \"CV\", when it is a distinct method from CV that shares the partitioning structure in common. </p>\n\n<p>Here are the two methods:</p>\n\n<ol>\n<li><p>Cross-validation - partition the data into k folds for train/validation splitting to exhaustively use all of the data for validation, with the purpose of evaluating model and hyperparameter performance.  You use the evaluation results to guide your model/hyperparameter selection.</p></li>\n<li><p>K-fold averaging - partition the data into k folds (using the same partitioning as in CV), train the same model on each fold, and average the results to help reduce the variance of a single model. You could add out of fold validation to this process without using it directly for selection, or you can just run with the result of the averaging if you feel good about the model/hyperparameter selection.</p></li>\n</ol>\n\n<p>The reason the terminology gets confused is that many people take a shortcut and do both things at the same time. They do one k-fold split, use out of fold sets to select something like n_estimators for gradient boosting models with early stopping, then average the resulting models across the k train/validation runs. In theory, this is poor practice. Early stopping overfits to the particular choice of validation fold, so models averaged in this manner within CV inevitably overfit. In practice, the averaging helps mitigate some of the overfitting and the overall harm may be small, but it's something worth keeping in mind. It certainly causes your local score estimates to be at least slightly optimistic. </p>\n\n<p>This is a common pitfall in use of CV more generally. People don't realize that by using certain validation sets to guide their model selection, they've reduced the legitimacy of using those same validation sets for estimating generalization error. The best estimates of generalization error are given by sets that have had no influence on model selection. This is why doing CV <strong>and</strong> setting aside a true hold-out set that only gets used once for evaluation is a common and sound practice.    </p>",
          "votes": 26,
          "replies": []
        },
        {
          "id": 342652,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-13T20:59:59.003000",
          "content": "<p>One practical reason why k-fold cross validation is famous around kagglers, is that you can do this approach for several layers, without needing a new hold-out set for each level. This is especially helpful when working together in a team or stacking different models.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 342692,
          "author_name": "Harlan Seymour",
          "author_url": "",
          "post_date": "2018-06-13T22:43:44.870000",
          "content": "<p>@Joe Eddy: nice overview!</p>\n\n<blockquote>\n  <p>This is why doing CV and setting aside a true hold-out set that only\n  gets used once for evaluation is a common and sound practice.</p>\n</blockquote>\n\n<p>In this competition at least, I consider the public leaderboard to be my true hold-out set.  I prefer to use K-Fold CV to generate my predictions, with K models each built on K-1 folds having their predictions averaged together.  I see this as form of de facto model averaging.  </p>\n\n<p>I use early stopping.  I agree that it overfits.  But it also seems to me that each K-1 folds are best fit by a different # of epochs or trees, so guessing this # without early stopping is also misfitting (either underfitting or overfitting).  So the misfitting washes out.</p>\n\n<p>The big bonus for me is using the out-of-fold predictions for model selection and to optimize model stacking.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 342713,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-14T00:00:36.640000",
          "content": "<p>I agree with @Harlan Seymour. The reason you stated is why I am using CV this way and I believe most people in this competition are as well.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 342722,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-14T00:31:55.450000",
          "content": "<p>@Harlan and YaGana, I definitely agree it's very convenient and that overfitting is mostly mitigated. It works well enough to produce good submissions and stacking results. But it's also <em>definitely not optimal</em>. When stacking, generating out of fold predictions on the same folds you early stop on introduces an optimistic bias that your meta learner may overfit to (these predictions are no longer out of fold in a true sense of the model not having seen the data, it now has glimpsed at that data). If that optimistic bias exists across all the models you stack it likely washes out, but if some models were trained differently (e.g. a non-boosting model that doesn't have the same early stopping concept), it might be problematic for the meta learner.</p>\n\n<p>There's no reason you can't get the best of both worlds, at the cost of extra computation time. You can run a 5-fold CV for model selection, then run another 5-fold process on a different seed to generate true OOF predictions to stack with. You can take n_estimators as the mean of the 5-fold CV early-stopped n_estimators, or use a similar heuristic. I can't find it right now, but I believe there's a post somewhere where Laurae (LGBM contributor) recommends the heuristic 110% of mean early-stopped n_estimators.     </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 342758,
          "author_name": "Harlan Seymour",
          "author_url": "",
          "post_date": "2018-06-14T03:16:38.033000",
          "content": "<p>@ Joe, you've definitely given me food for thought.  Take for example XGBoost.  What do you think of this compromise to save compute time? </p>\n\n<ol>\n<li>Save models and record ntree_limit used in early stopping for all K folds</li>\n<li>Calculate mean of ntree_limit used in early stopping over the K folds (call it M)</li>\n<li>Reload models with &lt; 110% x M trees and train them up to 110% x M</li>\n<li>Call predict() for each model to generate test and out-of-fold predictions, passing 110% x M for ntree_limit for each model.</li>\n</ol>\n\n<p>Note that I might prefer something like 95% x M to 110% x M.  Val score changes very little for the final trees before early stopping while train score goes down significantly i.e. lots of overfitting at the end before early stopping.</p>\n\n<p>Doesn't this almost entirely mitigate your overfitting worry without having to do a double run?  Note that XGBoost takes a long for me in this competition.  </p>\n\n<p>@Joe, I have another question it would be interesting to get your input on.  Do you create a model on the entire training set, using in XGBoost, say, 110% x M x (K / (K-1)) for ntree_limit?  This model would be used just for test predictions.  Maybe average its test predictions with average of K fold test predictions?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 342763,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-06-14T03:57:22.260000",
          "content": "<p>My personal preference is to not use early stopping. Use trial and error or some initial early stopping to get a certain number of rounds, and then use that number of rounds consistently with all the folds. Also remember to pay attention to the relationship between choice of learning rate and number of rounds needed.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 343080,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-06-14T17:06:36.897000",
          "content": "<p>@Joe Eddy</p>\n\n<p>I agree with a lot of this but if we are talking about <em>optimal</em> then using a heuristic such as 110% of the mean of the early stopped n_estimators certainly isn't optimal.</p>\n\n<p>P.S. Does anyone know why when I reply to, say, Joe Eddy, the forum thinks I am replying to the person he replied to?</p>\n\n<p>P.P.S Some of this is covered <a href=\"http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\">here</a> which is useful reading.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343102,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-06-14T17:37:55.190000",
          "content": "<p><a href=\"/maw501\">@maw501</a> that's fair, there's no such this as optimality in statistical modeling of distributions that we can't have perfect knowledge of, everything is relative :) </p>\n\n<p>But on a practical level, I don't know what you would expect (both in the normal sense of expect, and <em>on expectation</em> in the statistical sense) to work better for generalization performance of n_estimators than something close to the mean of early stopped rounds. Here's a <a href=\"https://www.kaggle.com/ogrellier/can-early-stopping-generalize/notebook\">nice experiment</a> in favor of using the mean to estimate the optimal number of rounds from olivier. Do you have in mind something that works better / what would your choice be?</p>\n\n<p>@Harlan I like that idea a lot, it's not something I've thought of before. I can't think of a reason why it wouldn't mitigate the overfitting issue, sounds like a great time saver and worth a shot to me. </p>\n\n<p>Re: your second question, I think you could definitely do that, but the gains might be pretty small if you're already using all of the data for training by averaging test predictions across k-folds. I tend to skip that to save time, but using all the training data in one model might be worth it. It aligns with the more traditional way of stacking - generate oof predictions in k-fold, retrain model on entire dataset to generate test predictions (without averaging of models). With relatively smaller datasets like this one my instinct is that averaging on k-folds gives a notable improvement via variance reduction, so that's what I go with.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 343163,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2018-06-14T20:01:47.250000",
          "content": "<p>+1 for variance reduction. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343188,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2018-06-14T21:07:04.853000",
          "content": "<p>I have a question about stacking here in Kaggle and for that competition:\nMy current way of doing it is:\n-For 5 folds, train 5 models then predict on the OOF to construct train meta data for that model.\n-Train a full model and create a test meta data.\n(then later on train a meta model etc.)</p>\n\n<p>What I'm wondering is if it makes sense to instead use the average of the 5 models predictions on test data. The same way I (and others) currently test my single models on the LB.\nDoes it make sense to think that the test meta is more stable that way and closer to the train meta (on average) compared to have a different model that created test meta ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343216,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-06-14T22:30:54.153000",
          "content": "<p>Hi Joe,</p>\n\n<p>Again, thanks for a thoughtful reply. I definitely agree with your first comment re. optimality and this alongside scepticism is a reasonable grounding for most problems. I also think you somewhat answer your own point when asking for what I would <em>expect</em> to perform better when you point out the elusiveness of optimality. I don't expect any method to be consistently better than any other (for a reasonable choice of methods).</p>\n\n<p>This is more complicated clearly and I think a simple (self-evident) answer is that it's entirely problem dependent. I've yet to see any ML technique/approach (including state of the art) that one feels entirely comfortable applying in all contexts. My last two competitions I worried far too much about whether things were theoretically justifiable/sound and paid the price. Worrying about reliably estimating generalization error (which we care about in the real world) vs. performing well on a Kaggle competition vs. having robust statistical backing for one's ideas are all different things.</p>\n\n<p>My choice is usually to try most of the approaches listed with a (more recent) favouring to average across the predictions from the folds for the base models. The biggest problem with using a simple heuristic like 110% of the mean of the early stopped n_estimators is that given the extra data it feels a little too much like guesswork for me. Yes under some assumptions I'm sure it's fine, but what are they? Can we even state them?</p>\n\n<p>Further if you change from 5-fold to 10-fold to any <strong>k</strong>-fold - what then? Add in that one should perhaps favour higher entropy validation methods then you should accept some folds will be harder than others for your model and early stopping at least deals with this (of course, at some risk of overfitting) in a consistent manner - additionally you have the benefit of averaging too. If you have a diverse set of base models I don't particularly worry about any single model overfitting too much.</p>\n\n<p>I do like the idea of using different random seeds for model iteration and oof predictions and will certainly try this.</p>\n\n<p>Edit extra: Just wondering if we can resolve via a thought experiment. For instance, take an extreme 2-fold example with one really easy fold needing 500 rounds and the other really difficult needing 1500 - using the heuristic we would train on the whole data now with 1, 100 rounds. I'm not sure I'm happy with that vs. averaging the predictions from both which at least predicts at the point when it thinks it has discovered all the information from the data. I'm not saying the difference will be big, or one is consistently better but I like getting variance reduction for free and some comfort in knowing the number of rounds. I guess to some extent it's a function of the dissimilarity in the folds/strength of signal in each but I'm not sure how one would formalise this - so back to my point about not expecting any method to do better on average.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343670,
          "author_name": "Harlan Seymour",
          "author_url": "",
          "post_date": "2018-06-15T19:47:05.077000",
          "content": "<p>@Joe: As an experiment I applied the procedure I outlined to my latest XGB per-fold models, syncing them to all to be 100% of the mean of the early stopping ntree_limits.  My local CV score went up minutely by +0.00001.  The newly submitted test predictions had the same RMSE score on the public leaderboard, but ordering \"My Submissions\" by score showed it was not as good as the original.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343696,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-15T20:54:37.490000",
          "content": "<p>Thanks @Joe Eddy for sparking this useful discussion. And thanks @Harlan Seymour for the update on your experiment. </p>\n\n<p><a href=\"/maw501\">@maw501</a>, I would like to know the result of the thought experiment on testing on 2 extreme cases you mentioned. If you ever do that, please report back if you can.</p>\n\n<p>Averaging is what I am comfortable with at the moment to reduce variance. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 336590,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-06-01T01:59:22.720000",
      "content": "<p>LGB 0.2223 5-fold average</p>\n\n<p><strong>update</strong>:\nnew feature add, 0.2217 without 5 folder</p>\n\n<p>0.2210 with 5 folder</p>\n\n<p><strong>update</strong>  6-15:</p>\n\n<p>0.2208 with 5 folder</p>\n\n<p><strong>update</strong>  6-20:</p>\n\n<p>0.2205 with 5 folder</p>\n\n<p>final:</p>\n\n<p>0.2204 with 10 folder</p>",
      "votes": 1,
      "replies": [
        {
          "id": 336595,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2018-06-01T02:17:59.517000",
          "content": "<p>劉sir,  您是用Stack跳到0.2210的麼?</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 337169,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-02T06:24:52.117000",
          "content": "<p>还没stacking，简单的blend几个public kernel</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 337275,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-06-02T12:38:43.800000",
          "content": "<p>Translation: \nLiu sir , are you jumping to 0.2210 with Stack?\n - No stacking, simple blend several public kernel</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 338644,
          "author_name": "ea",
          "author_url": "",
          "post_date": "2018-06-05T13:44:03.090000",
          "content": "<p>有啥技巧么，我blend一个0.2212和0.2216结果也就0.2212</p>",
          "votes": -5,
          "replies": []
        },
        {
          "id": 338654,
          "author_name": "Diangu Xie",
          "author_url": "",
          "post_date": "2018-06-05T14:17:22.973000",
          "content": "<p>I guess there is not any techniques. Just use weight to blend public with private kernel. I used blending, too</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 339058,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-06-06T08:01:20.183000",
          "content": "<p>you need try to improve your single model,then try to blend or stacking</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 330553,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-05-19T05:14:52.653000",
      "content": "<p>Edit :  0.2249</p>\n\n<p>Edit: 0.2252 with Image_top_1 feature included </p>\n\n<p>Edit: 0.2259 with RNN </p>\n\n<p>0.2276 with RNN, Fast Text Embeddings . - No Tuning done,  A lot of Scope for improvement and tuning. \nLink to Public kernel is here </p>\n\n<p><a href=\"https://www.kaggle.com/shanth84/avito-fast-text-keras-model/code\">0.2276 RNN - Good for Blend :)</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 329707,
      "author_name": "Phat Hoang",
      "author_url": "",
      "post_date": "2018-05-17T02:38:18.170000",
      "content": "<p>0.2223 with 5fold LGB model (all features are taken from public kernels). </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 329210,
      "author_name": "Chau Huynh",
      "author_url": "",
      "post_date": "2018-05-16T02:38:27.313000",
      "content": "<p>Single lgbm 0.223. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 329185,
      "author_name": "takuoko",
      "author_url": "",
      "post_date": "2018-05-16T00:01:24.690000",
      "content": "<p>Single lgbm 0.2215.\nBut using image features from <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features\">https://www.kaggle.com/bguberfain/vgg16-train-features</a> , my score got worse.\nI would like to know how to use image feature.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 329196,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-16T00:39:06.463000",
          "content": "<p>Wow, that is so high for a single lgbm model! Would you mind if share some ideas how to improve from 0.223 level to 0.2215 ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329209,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-05-16T02:34:48.440000",
          "content": "<p>Price and item-seq-number is important feature. I refer to <a href=\"https://www.kaggle.com/tapioca/item-seq-number-vs-deal-probability/code\">https://www.kaggle.com/tapioca/item-seq-number-vs-deal-probability/code</a> , and make some features. Good luck!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 329213,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-16T02:43:55.210000",
          "content": "<p>Thanks! For the image feature, my thought is: the features from the pre-trained Imagenet model are designed for classification. For example, we want our VGG16 distinguish bicycle from car. However, we already have the features like item_id and category, so directly use imgenet features may not helps. Maybe we can try fine tune those features with the target for the deal probabilities.  This is just my naive thought and I haven't try any image features</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 329214,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-05-16T02:51:07.357000",
          "content": "<p>Wow! Your thought of image feature is very cool! I try to use color base feature because it is easy to use. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329451,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-05-16T13:45:34.857000",
          "content": "<p>And also may ask how many models you are using to achieve your LB position now?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329466,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-05-16T14:10:34.853000",
          "content": "<p>LB 0.2209 by single lgbm. Add some features from active.csv and periods data.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 329646,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-05-16T21:35:57.487000",
          "content": "<p>@takuoko : item-seq-number seems to be exactly the same as doing a count after a group by user no ? I was able to get good features from the price, but not by combining item_seq_number with another feature... Thanks for sharing !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329770,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-05-17T06:39:04.460000",
          "content": "<p>Thanks for your comment, Antoine! By using item-seq-number feature, 0.00005 better val score in my 5 folds. So it may be ineffective. But I think item-seq-number is important. So I continue trying to make featrue about that. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 329895,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-05-17T13:36:33.833000",
          "content": "<p>simple feature from VGG16 is not helpful, because  the prediction is not precise nor suitable. I have write a kernel <a href=\"https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label\">https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 329122,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2018-05-15T19:17:42.043000",
      "content": "<p>Single lgbm 0.2229</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 329052,
      "author_name": "Victor An",
      "author_url": "",
      "post_date": "2018-05-15T16:02:51.120000",
      "content": "<p>LB 0.2249 with only sample train and test dataset via lightgbm</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 328924,
      "author_name": "Totoro (ごめん)",
      "author_url": "",
      "post_date": "2018-05-15T11:41:07.980000",
      "content": "<p>My best single NN is 0.2233 without image nor active and period files.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 329641,
          "author_name": "Totoro (ごめん)",
          "author_url": "",
          "post_date": "2018-05-16T21:24:57.023000",
          "content": "<p>Update: <br>\n0.2226 for NN</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 329741,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2018-05-17T04:58:21.573000",
          "content": "<p>Could you give a hint about architecture of your NN?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329745,
          "author_name": "Totoro (ごめん)",
          "author_url": "",
          "post_date": "2018-05-17T05:15:39.803000",
          "content": "<p>As @Strideradu commented above,</p>\n\n<p>&gt; So the NN is for the NLP of description or more like for the classifier for different features just like LGBM</p>\n\n<p>I use both. </p>\n\n<ol>\n<li><p>Category + Numeric part <br>\nYou can refer from fastai or some solution in this competition or similar one: \n<a href=\"https://www.kaggle.com/c/donorschoose-application-screening\">https://www.kaggle.com/c/donorschoose-application-screening</a></p></li>\n<li><p>NLP <br>\nYou can refer from previous competitions like: <br>\n<a href=\"https://www.kaggle.com/c/donorschoose-application-screening\">https://www.kaggle.com/c/donorschoose-application-screening</a> <br>\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge</a>\n...\nto know how they dealed with NLP.</p></li>\n</ol>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 329748,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2018-05-17T05:20:15.037000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 336777,
      "author_name": "jetou Xu",
      "author_url": "",
      "post_date": "2018-06-01T09:02:02.143000",
      "content": "<p>NN:\n0.2213 (Categorical, Numerical, W2V)\n10-fold</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 335073,
      "author_name": "AlexTru",
      "author_url": "",
      "post_date": "2018-05-29T05:48:58.033000",
      "content": "<p>one LightGBM  model 5kfold - 0,2211 (UPD)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 335075,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2018-05-29T05:55:18.377000",
          "content": "<p>incredible</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335214,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-05-29T11:41:39.043000",
          "content": "<p>Almost all ideas i take from public kernels :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 338415,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-05T02:30:16.977000",
          "content": "<p>@AlexTru, did you use image features? If so which ones if you can share? They seem to worsen both my CV and LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338447,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-06-05T04:28:42.777000",
          "content": "<p>Yes, i use features from this two kernels. <a href=\"https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf\">this</a>  and <a href=\"https://www.kaggle.com/peterhurford/image-feature-engineering\">this</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 339187,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-06-06T13:30:03.850000",
          "content": "<p>UPD 22.07</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 331503,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2018-05-21T12:22:52.577000",
      "content": "<p>Update: 0.2220 with variation of the 'ridge trick', 5fold lgbm.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 331505,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-21T12:28:21.953000",
          "content": "<p>What's a \"ridge trick\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331511,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-21T12:37:26.997000",
          "content": "<p>Creating an oof prediction from a ridge model and using it as a feature - along the lines of <a href=\"https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0\">https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331531,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-21T13:05:26.643000",
          "content": "<p>Cool. But does that really count as a single model? Though I'd concede at this point I'm just being pedantic. \"Single model\" is a bit of a vanity metric because of things like this, and it's really just about finding models that are individually good but blend together.</p>\n\n<p>My LGB has Ridges in it so I wasn't posting about it here. But I'm currently at 0.21611 CV / 0.2200 LB with a single LGB with multiple Ridge models inside. The best individual Ridge is at 0.22674 CV.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 331551,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-21T13:30:21.487000",
          "content": "<p>Nice one - your gap cv - lb is a bit smaller than mine (~ 0.006). If you don't mind sharing, have you tuned the lgb a lot? I've been rolling mostly with parameters close to the ones from the public kernels so far.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331556,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-21T13:39:06.703000",
          "content": "<p>...I sure hope my smaller CV gap does not mean that my LB position will come crashing down at the end. :/</p>\n\n<p>I have spent some amount of time tuning the LGB by hand to the point where I can no longer easily improve it. I still think I have some room to non-easily improve it, perhaps with Hyperopt or something.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331561,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-21T13:43:35.087000",
          "content": "<p>@Konrad: Are you seeing 0.216XX CV as well? Do you have a lot of other features beyond those in the kernels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 331596,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-21T14:48:27.193000",
          "content": "<p>Yes and yes - with the latter, i went to town with target encoding. I guess tuning is my next step, as i have barely touched it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331642,
          "author_name": "Himanshu Chaudhary",
          "author_url": "",
          "post_date": "2018-05-21T16:14:54.077000",
          "content": "<p>Hey I have tried target encoding but it was not helping much as passing the categorical values into LGBM is doing. I think keeping the number of features high is helping the tree to build better.\nwhat do you think?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331665,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-21T17:04:40.073000",
          "content": "<p>Interesting - I have exact opposite experience to yours: once I mean-encoded the categoricals (created new numerical cols) and then dropped the originals, my score improved significantly. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 331669,
          "author_name": "Himanshu Chaudhary",
          "author_url": "",
          "post_date": "2018-05-21T17:09:44.993000",
          "content": "<p>maybe I am doing somewhere wrong then. I tried using this <br>\n<a href=\"https://www.kaggle.com/tnarik/likelihood-encoding-of-categorical-features\">https://www.kaggle.com/tnarik/likelihood-encoding-of-categorical-features</a>. I am encoding all the categorical values .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331682,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-21T17:22:01.073000",
          "content": "<p><a href=\"/konrad\">@konrad</a> How do you mean encode? You may be at risk of overfitting your CV. When I did mean encoding I ended up in overfit hell for a few days and it took ten submits to get to the bottom of it. So I'm on team LGB native encoding. I might try doing a second LGB with mean encoding though and average the two approaches and see if that helps... I'll calculate it within the CV folds to monitor overfitting better.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 331687,
          "author_name": "Himanshu Chaudhary",
          "author_url": "",
          "post_date": "2018-05-21T17:35:05.907000",
          "content": "<p>yeah, my CV score was good enough but LB was worst.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 331695,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-21T17:45:46.247000",
          "content": "<p>@Peter: I did the mean-encoding inside the cv and added heavy regularization, so I think I'm fine (especially that cv - lb is consistently around 0.006, and changes are directionally correct).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 333343,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-25T00:31:58.560000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 340394,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2018-06-09T04:46:18.230000",
      "content": "<p>10-fold LGB, LB 0.2214.</p>\n\n<p>Training a model including some image features. Hopefully it would give me a boost. Finger crossed.  </p>\n\n<p>Update: Including image featuers, 10-fold LGB, LB 0.2210</p>\n\n<p>6-15 Update: manage to improve my LGB to LB 0.2202</p>\n\n<p>6-19 Update: finally got an LGB with LB 0.2199, XGB with LB 0.2209.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 343000,
          "author_name": "high&mean",
          "author_url": "",
          "post_date": "2018-06-14T14:21:53.030000",
          "content": "<p>10-fold LGB avg or min ？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343002,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2018-06-14T14:24:21.427000",
          "content": "<p>10-fold LGB avg.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 334596,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2018-05-28T00:37:22.033000",
      "content": "<p>Right now I have 3 models in parallel. A catboost at 0.2237, a LGB at 0.2228 and a Neural network at 0.2240. I'm really impressed by people achieving 0.222X with neural network on that data set. Maybe I'm missing something fundamental or some kind of clever encoding of categoricals.</p>\n\n<p>Not yet used additional FE from satellite data sets.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 334699,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-28T08:32:27.917000",
          "content": "<p>As for me , I am just using simple label encoding for categorical data . </p>\n\n<p>I think the NN architecture plays a key role and clever archtitecture can bring huge improvement. </p>\n\n<p>I tend to rely on deep models just to avoid lot of features engeneering. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 334725,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2018-05-28T09:30:27.170000",
          "content": "<p>Interesting. <br>\nIt is a surprise for me to know the simple label encoding for categorical data works well.  I will give it a try. <br>\nThank for your comment</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335045,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2018-05-29T04:06:51.050000",
          "content": "<p>Thanks. I'll play around with some extra ideas. My architecture is fairly straightforward so far, ConvNN for texts, embeddings for categoricals then concat everything and pass it through dense layers.</p>\n\n<p>Maybe I can try going deeper with the text part.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335870,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-05-30T14:50:52.780000",
          "content": "<p>Hi @Serigne!\nYou really just encode cat features to numbers? or use one-hot encoding? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 335874,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-30T15:05:14.377000",
          "content": "<p>Hi Alex... Yes just simple encoding to numbers..Then embed--&gt; concat--&gt; dense layers just like Arroqc..</p>\n\n<p>But that's just for categorical  features. </p>\n\n<p>For text features I try to follow these  4 principles : <a href=\"https://explosion.ai/blog/deep-learning-formula-nlp\">Embed, Encode, Attend and Predict</a></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 335972,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2018-05-30T19:51:50.463000",
          "content": "<p>Interesting article Serigne. Thanks. Will try to see if attention helps.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 334448,
      "author_name": "ea",
      "author_url": "",
      "post_date": "2018-05-27T12:26:31.887000",
      "content": "<p>I got 0.223 using just train.csv and test.csv,How to using train_active.csv which has no labels</p>",
      "votes": 0,
      "replies": [
        {
          "id": 334827,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-05-28T14:17:12.837000",
          "content": "<p>@ unexceptednull </p>\n\n<p>You perform feature engineering with user_id, item_id and merge it to train using 'user_id'\nIt is available in Benjamin's awesome kernel here <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm/output\">BENJAMIN FEATURE ENGINEERING WITH TRAIN ACTIVE </a></p>\n\n<p>You can also refer to this kernel for building RNN on these additional features <a href=\"https://www.kaggle.com/shanth84/rnn-detailed-explanation-0-2246\">RNN - 0.2246</a></p>\n\n<p>Hope this helps \nCheers\nShanth</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 330469,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-18T21:18:22.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329704,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-17T02:28:57.377000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329226,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-16T03:34:41.703000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329137,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-15T20:13:42.170000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 329166,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-15T22:25:15.630000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329168,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-15T22:30:51.207000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 332664,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T14:41:55.867000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332668,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T14:49:33.353000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 332671,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T14:53:20.553000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332778,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T18:25:35.923000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332782,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T18:30:05.807000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332787,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T18:39:50.697000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332791,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T18:52:05.387000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332807,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T19:37:05.640000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 332820,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T20:32:55.627000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332827,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-23T20:57:28.390000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 334810,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-28T13:27:01.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 330081,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-18T02:21:52.027000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 333907,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-26T02:34:02.520000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "328917": "I wonder what the best single model score is. I got my NN to 0.2240, but havn't included any image features nor the active and period files yet. How much is the improvement with image and co.?",
    "340360": "I got 0.2181 with a NN network with 10 folds (UPDATED)\n",
    "338921": "I'm hitting .2190 with a 5-fold averaged LGBM",
    "333988": "LGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2200 (updated)\n\nXGB (Categorical, Numerical, TFIDF, Images meta-data + ImageNet + NIMA scoring): 0.2216 (updated)\n\nNN (Categorical, Numerical, W2V, Images meta-data + ImageNet scoring): 0.2198 (updated)\n\nAll with CV5.",
    "344107": "LGB 10-fold 0.2179",
    "338049": "LGBM 10-fold avg. - 0.2203 (categorical, numerical, target encoding, text and features predicted using all csv data)",
    "333337": "NN:  \n0.2218 Fasttext  \n0.2215 Self-training embedding  \n\nThanks @Dieter",
    "329320": "I am working exclusively with NN for now..\n\nSo far,  best single NN model : LB = 0.2246 ( 5 folds CV RMSE : 0.2215xxx ), no image features either\n\nThe model was scoring 0.244xxx on LB when I started it ,   5 days ago...(Some ideas from kernels and discussion topics helped me in the improvement) \n\n**Edit :**  The NN model is now scoring 0.2223 on LB ",
    "342176": "Single model LightGBM 0.2199, with 5seed-avg 0.2195",
    "339507": "Finally! Able to achieve 2210-XGB and 2217-LGB. No image features though.\nI was stuck earlier and started from scratch, seems the right decision. Best of luck guys.",
    "337277": "We are at\n\n - LGB 0.2200 \n - XGB 0.2213 \n - Catboost 0.2208\n - NN 0.2194",
    "335359": "My best so far is an LGB model with .2216 (recently improved) after a 10-fold average.",
    "328983": "Got 0.2222 with NN just train and test",
    "341626": "After a lot of hard work I got my single LGB (with 5 fold) down to 0.2200 on LB. Yay !",
    "340365": "0.2190 10-fold LGB",
    "342452": "Silly question.\n\nWhat do you guys exacly mean when you say, for example, 10-fold LGB?\nYes, you can use CV to better tune hyperparameters. Particularly, n_estimators. Is that what you mean? \n\nAnd when you find optimal number of estimators do you retrain on whole dataset? ",
    "336590": "LGB 0.2223 5-fold average\n\n**update**:\nnew feature add, 0.2217 without 5 folder\n\n\n0.2210 with 5 folder\n\n\n**update**  6-15:\n\n0.2208 with 5 folder\n\n**update**  6-20:\n\n0.2205 with 5 folder\n\nfinal:\n\n0.2204 with 10 folder",
    "330553": "Edit :  0.2249\n\nEdit: 0.2252 with Image_top_1 feature included \n\nEdit: 0.2259 with RNN \n\n0.2276 with RNN, Fast Text Embeddings . - No Tuning done,  A lot of Scope for improvement and tuning. \nLink to Public kernel is here \n\n[0.2276 RNN - Good for Blend :)][1]\n\n\n  [1]: https://www.kaggle.com/shanth84/avito-fast-text-keras-model/code",
    "329707": "0.2223 with 5fold LGB model (all features are taken from public kernels). ",
    "329210": "Single lgbm 0.223. ",
    "329185": "Single lgbm 0.2215.\nBut using image features from https://www.kaggle.com/bguberfain/vgg16-train-features , my score got worse.\nI would like to know how to use image feature.",
    "329122": "Single lgbm 0.2229",
    "329052": "LB 0.2249 with only sample train and test dataset via lightgbm",
    "328924": "My best single NN is 0.2233 without image nor active and period files.",
    "336777": "NN:\n0.2213 (Categorical, Numerical, W2V)\n10-fold",
    "335073": "one LightGBM  model 5kfold - 0,2211 (UPD)",
    "331503": "Update: 0.2220 with variation of the 'ridge trick', 5fold lgbm.",
    "340394": "10-fold LGB, LB 0.2214.\n\nTraining a model including some image features. Hopefully it would give me a boost. Finger crossed.  \n\nUpdate: Including image featuers, 10-fold LGB, LB 0.2210\n\n6-15 Update: manage to improve my LGB to LB 0.2202\n\n6-19 Update: finally got an LGB with LB 0.2199, XGB with LB 0.2209.  ",
    "334596": "Right now I have 3 models in parallel. A catboost at 0.2237, a LGB at 0.2228 and a Neural network at 0.2240. I'm really impressed by people achieving 0.222X with neural network on that data set. Maybe I'm missing something fundamental or some kind of clever encoding of categoricals.\n\nNot yet used additional FE from satellite data sets.",
    "334448": "I got 0.223 using just train.csv and test.csv,How to using train_active.csv which has no labels",
    "330469": "I've started with catboost and currently sits at 0.2246 with just train.csv. Edit: (Improved it to 0.2237)\n\nI'm currently hoping to get further with LGB because of the ability to use sparse matrix.\n\n",
    "329704": "0.2222 with one lgb model",
    "329226": "0.2228 with one lgb",
    "329137": "LGBM with 5-fold CV 0.2231",
    "334810": "",
    "330081": "",
    "333907": "thank you"
  }
}