{
  "id": 53634,
  "title": "[Updated] Single Best Model Scores : Motivation : Feature Engineering..!!",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53634",
  "author_name": "Rohit Mehra",
  "post_date": "2018-04-02T19:08:48.694000",
  "votes": 67,
  "comment_count": 395,
  "views": 0,
  "content": "<p>Everyone,</p>\n\n<p>A request to all the masters and people leading the board, can you share single best model scores? With all the blending going on, single scores would give us motivation.</p>\n\n<ol>\n<li>Mine ==&gt; LGB: All data (train + test, not test_supplement): 0.9697</li>\n<li><strong>Update 4-12-2018: All data (train + test, not test_supplement): 0.9774</strong> I have yet to see how the new features i.e. unique user instance features can be used with the full train and also I am not using any \"attribute-based freq\" features.</li>\n<li><strong>Update 4-27-2018: All data (train + test, not test_supplement): 0.9799</strong> Removing minute and seconds as features.</li>\n<li><strong>Update 4-28-2018: All data (train + test, not test_supplement): 0.9802</strong> Removed more features of type \"prevclick\". </li>\n<li><strong>Update 4-29-2018: All data (train + test, not test_supplement): 0.9803</strong> Hyperparameter tuning. More regularization.</li>\n</ol>\n\n<p>Below .98xx app had 10x more importance over the channel. Above .98xx channel has 3x more importance. I am unable to understand this. Does anyone have some views on this? What is your most important feature?</p>\n\n<p>Thanks.</p>\n\n<p>P.S. Upvote if you feel it will motivate you as well.. ;)</p>",
  "messages": [
    {
      "id": 307993,
      "postDate": "2018-04-02T19:08:48.693Z",
      "content": "<p>Everyone,</p>\n\n<p>A request to all the masters and people leading the board, can you share single best model scores? With all the blending going on, single scores would give us motivation.</p>\n\n<ol>\n<li>Mine ==&gt; LGB: All data (train + test, not test_supplement): 0.9697</li>\n<li><strong>Update 4-12-2018: All data (train + test, not test_supplement): 0.9774</strong> I have yet to see how the new features i.e. unique user instance features can be used with the full train and also I am not using any \"attribute-based freq\" features.</li>\n<li><strong>Update 4-27-2018: All data (train + test, not test_supplement): 0.9799</strong> Removing minute and seconds as features.</li>\n<li><strong>Update 4-28-2018: All data (train + test, not test_supplement): 0.9802</strong> Removed more features of type \"prevclick\". </li>\n<li><strong>Update 4-29-2018: All data (train + test, not test_supplement): 0.9803</strong> Hyperparameter tuning. More regularization.</li>\n</ol>\n\n<p>Below .98xx app had 10x more importance over the channel. Above .98xx channel has 3x more importance. I am unable to understand this. Does anyone have some views on this? What is your most important feature?</p>\n\n<p>Thanks.</p>\n\n<p>P.S. Upvote if you feel it will motivate you as well.. ;)</p>",
      "rawMarkdown": "Everyone,\n\nA request to all the masters and people leading the board, can you share single best model scores? With all the blending going on, single scores would give us motivation.\n\n1. Mine ==&gt; LGB: All data (train + test, not test_supplement): 0.9697\n2. **Update 4-12-2018: All data (train + test, not test_supplement): 0.9774** I have yet to see how the new features i.e. unique user instance features can be used with the full train and also I am not using any \"attribute-based freq\" features.\n3. **Update 4-27-2018: All data (train + test, not test_supplement): 0.9799** Removing minute and seconds as features.\n4.  **Update 4-28-2018: All data (train + test, not test_supplement): 0.9802** Removed more features of type \"prevclick\". \n5.  **Update 4-29-2018: All data (train + test, not test_supplement): 0.9803** Hyperparameter tuning. More regularization.\n\nBelow .98xx app had 10x more importance over the channel. Above .98xx channel has 3x more importance. I am unable to understand this. Does anyone have some views on this? What is your most important feature?\n\nThanks.\n\nP.S. Upvote if you feel it will motivate you as well.. ;)",
      "votes": 67
    },
    {
      "id": 317359,
      "postDate": "2018-04-21T09:28:12.150Z",
      "content": "<p>Single NN model, LB 0.9814</p>",
      "rawMarkdown": "Single NN model, LB 0.9814",
      "votes": 20,
      "replies": [
        {
          "id": 317368,
          "postDate": "2018-04-21T09:44:36.477Z",
          "content": "<p>Thanks, very encouraging!</p>",
          "rawMarkdown": "Thanks, very encouraging!"
        },
        {
          "id": 317388,
          "postDate": "2018-04-21T11:44:47.780Z",
          "content": "<p>Good to see a NN model....  Do you mind telling us the number of new features.... a single digit or a dozen of them.... :)</p>",
          "rawMarkdown": "Good to see a NN model....  Do you mind telling us the number of new features.... a single digit or a dozen of them.... :)"
        },
        {
          "id": 317394,
          "postDate": "2018-04-21T12:14:34.597Z",
          "content": "<p>Hi Samrat, its a single digit.</p>",
          "rawMarkdown": "Hi Samrat, its a single digit.",
          "votes": 2
        },
        {
          "id": 317411,
          "postDate": "2018-04-21T13:14:22.283Z",
          "content": "<p>Great!!!! Thanks for the reply....</p>",
          "rawMarkdown": "Great!!!! Thanks for the reply...."
        },
        {
          "id": 317621,
          "postDate": "2018-04-22T02:37:58.297Z",
          "content": "<p>GREAT!!!!!</p>",
          "rawMarkdown": "GREAT!!!!!"
        },
        {
          "id": 318015,
          "postDate": "2018-04-23T03:25:11.437Z",
          "content": "<p>owesome</p>",
          "rawMarkdown": "owesome"
        },
        {
          "id": 318256,
          "postDate": "2018-04-23T13:35:03.333Z",
          "content": "<p>Great to see that NN is performing so well. Are you using simple NN setup or combination of CNN/FC?\nCrossing 0.9750 using NN becomes a nightmare for me :)</p>",
          "rawMarkdown": "Great to see that NN is performing so well. Are you using simple NN setup or combination of CNN/FC?\nCrossing 0.9750 using NN becomes a nightmare for me :)"
        },
        {
          "id": 318373,
          "postDate": "2018-04-23T17:37:37.857Z",
          "content": "<p>Very nice. Looking forward to your kernel after the competition ends :)</p>",
          "rawMarkdown": "Very nice. Looking forward to your kernel after the competition ends :)"
        }
      ]
    },
    {
      "id": 314699,
      "postDate": "2018-04-16T06:06:16.850Z",
      "content": "<p>LB 0.9815 with single model and ~50 features and using all the train data</p>",
      "rawMarkdown": "LB 0.9815 with single model and ~50 features and using all the train data",
      "votes": 17,
      "replies": [
        {
          "id": 314700,
          "postDate": "2018-04-16T06:08:18.903Z",
          "content": "<p>It is really wonderfull</p>",
          "rawMarkdown": "It is really wonderfull"
        },
        {
          "id": 315009,
          "postDate": "2018-04-16T16:27:40.190Z",
          "content": "<p>@ Sameh....50 features is awesome. Are you hooked to a super computer? :)</p>",
          "rawMarkdown": "@ Sameh....50 features is awesome. Are you hooked to a super computer? :)"
        },
        {
          "id": 315082,
          "postDate": "2018-04-16T18:20:18.960Z",
          "content": "<p>i didnt measure (yet) the effect of including all the train data compared to only 1 day. but i assume the gain is not major. will report when i try it.</p>",
          "rawMarkdown": "i didnt measure (yet) the effect of including all the train data compared to only 1 day. but i assume the gain is not major. will report when i try it.",
          "votes": 1
        },
        {
          "id": 315096,
          "postDate": "2018-04-16T18:36:34.413Z",
          "content": "<p>Same assumption here. Now my score of 0.9805 is still based on day 9 with multiple groups of features and bagging. I'm still running something on day 9  and I think the limit of using a single day will be close to 0.981. </p>",
          "rawMarkdown": "Same assumption here. Now my score of 0.9805 is still based on day 9 with multiple groups of features and bagging. I'm still running something on day 9  and I think the limit of using a single day will be close to 0.981. ",
          "votes": 1
        },
        {
          "id": 315501,
          "postDate": "2018-04-17T07:29:37.077Z",
          "content": "<p>Hi Sameh,\nWhen you use ALL train data? do you not use CV?\nif so - when do you STOP training to make sure you are not overfitting?</p>\n\n<p>Thanks! =)</p>",
          "rawMarkdown": "Hi Sameh,\nWhen you use ALL train data? do you not use CV?\nif so - when do you STOP training to make sure you are not overfitting?\n\nThanks! =)"
        },
        {
          "id": 315506,
          "postDate": "2018-04-17T07:36:00.257Z",
          "content": "<p>@AmirH I guess it makes sense to first validate your model on day 9 and then train your model including day 9 with same features and parameters.</p>",
          "rawMarkdown": "@AmirH I guess it makes sense to first validate your model on day 9 and then train your model including day 9 with same features and parameters.",
          "votes": 1
        },
        {
          "id": 315580,
          "postDate": "2018-04-17T09:58:28.373Z",
          "content": "<p>@Sohaib</p>\n\n<p>Instead of retraining the model on all the data, why not just using the prediction for day 9 as OOF for stacking purpose ? </p>",
          "rawMarkdown": "@Sohaib\n\nInstead of retraining the model on all the data, why not just using the prediction for day 9 as OOF for stacking purpose ? ",
          "votes": 1
        },
        {
          "id": 315711,
          "postDate": "2018-04-17T14:06:04.220Z",
          "content": "<p>Yes, we can do that too. I will share stacking vs averaging(models trained including day 9) results with you once I am done with my single model. <br> EDIT: For me including day 9 in training increase AUC</p>",
          "rawMarkdown": "Yes, we can do that too. I will share stacking vs averaging(models trained including day 9) results with you once I am done with my single model. <br> EDIT: For me including day 9 in training increase AUC",
          "votes": 1
        },
        {
          "id": 315915,
          "postDate": "2018-04-17T19:41:02.293Z",
          "content": "<p>@congb if feasible, can you give an example of bagging feature?</p>",
          "rawMarkdown": "@congb if feasible, can you give an example of bagging feature?",
          "votes": 1
        },
        {
          "id": 316019,
          "postDate": "2018-04-18T02:54:05.993Z",
          "content": "<p>I’m just creating different sets of features on the same data set and bagging the results. Nothing fancy. Just a compromise because of the limited resource I have. </p>",
          "rawMarkdown": "I’m just creating different sets of features on the same data set and bagging the results. Nothing fancy. Just a compromise because of the limited resource I have. "
        },
        {
          "id": 316414,
          "postDate": "2018-04-19T01:37:43.657Z",
          "content": "<p>You mean retrain the model with the same parameters and features you used in validation, but when will you stop in retraining(like early_stopping stops when it doesn't improve) Thanks. </p>",
          "rawMarkdown": "You mean retrain the model with the same parameters and features you used in validation, but when will you stop in retraining(like early_stopping stops when it doesn't improve) Thanks. "
        },
        {
          "id": 316416,
          "postDate": "2018-04-19T01:39:29.540Z",
          "content": "<p>@bgm, yep I do use early stopping </p>",
          "rawMarkdown": "@bgm, yep I do use early stopping ",
          "votes": 1
        },
        {
          "id": 316419,
          "postDate": "2018-04-19T01:45:28.213Z",
          "content": "<p>OK thanks! So you should use validation in retraining, right?</p>",
          "rawMarkdown": "OK thanks! So you should use validation in retraining, right?"
        },
        {
          "id": 317140,
          "postDate": "2018-04-20T21:16:32.910Z",
          "content": "<p>@Sohaib Omar</p>\n\n<p>You mentioned that \"I guess it makes sense to first validate your model on day 9 and then train your model including day 9 with same features and parameters.\"</p>\n\n<p>I don't understand how to do it. Could you please explain the statement with some simple Python code.</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "@Sohaib Omar\n\nYou mentioned that \"I guess it makes sense to first validate your model on day 9 and then train your model including day 9 with same features and parameters.\"\n\nI don't understand how to do it. Could you please explain the statement with some simple Python code.\n\nThanks!\n"
        },
        {
          "id": 317207,
          "postDate": "2018-04-21T01:47:59.153Z",
          "content": "<pre><code>train = day7_8\nval = day9\nmodel.fit(train)\nval_score = model.predict(day9)\nif va_score &gt; my_best_val_score:\n    model.fit(train + val)\n</code></pre>\n\n<p>Hope it helps.</p>",
          "rawMarkdown": "    train = day7_8\n    val = day9\n    model.fit(train)\n    val_score = model.predict(day9)\n    if va_score &gt; my_best_val_score:\n        model.fit(train + val)\n\n\nHope it helps.",
          "votes": 3
        },
        {
          "id": 317636,
          "postDate": "2018-04-22T04:17:22.387Z",
          "content": "<p>thanks @Sohaib</p>",
          "rawMarkdown": "thanks @Sohaib",
          "votes": 1
        }
      ]
    },
    {
      "id": 308329,
      "postDate": "2018-04-03T11:03:39.880Z",
      "content": "<p>Single model LB 0.9756, 18 features including original columns and frequencies.</p>\n\n<p>Additional info: IP and ip frequencies were not used. Training was done on day 8. \nLocally I have similar results on day 7.</p>",
      "rawMarkdown": "Single model LB 0.9756, 18 features including original columns and frequencies.\n\nAdditional info: IP and ip frequencies were not used. Training was done on day 8. \nLocally I have similar results on day 7.",
      "votes": 16,
      "replies": [
        {
          "id": 308996,
          "postDate": "2018-04-04T13:26:51.660Z",
          "content": "<p>Thanks, Alexander. Have been doing feature engineering for past two days, now have upto 20 (freq and groupby) I was actually worried about IP and Ip derived features, thanks for clearing that out, and also the training strategy.</p>",
          "rawMarkdown": "Thanks, Alexander. Have been doing feature engineering for past two days, now have upto 20 (freq and groupby) I was actually worried about IP and Ip derived features, thanks for clearing that out, and also the training strategy."
        },
        {
          "id": 309107,
          "postDate": "2018-04-04T16:52:18.410Z",
          "content": "<p>Impressive results with just 18 features and 1 training day!</p>\n\n<p>Do you mind if I ask - when you say that ip frequencies weren't used, do you mean that you don't use any sort of groupby-count features derived from ip? For example, you don't use something like # of ip clicks in the hour (a feature commonly seen in the kernels)? </p>\n\n<p>Of course I understand if you don't want to share more :) </p>",
          "rawMarkdown": "Impressive results with just 18 features and 1 training day!\n\nDo you mind if I ask - when you say that ip frequencies weren't used, do you mean that you don't use any sort of groupby-count features derived from ip? For example, you don't use something like # of ip clicks in the hour (a feature commonly seen in the kernels)? \n\nOf course I understand if you don't want to share more :) ",
          "votes": 1
        },
        {
          "id": 309122,
          "postDate": "2018-04-04T17:23:10.147Z",
          "content": "<p>Same question with Joe, but it doesn't matter if you dont want to answer, we understand that.</p>",
          "rawMarkdown": "Same question with Joe, but it doesn't matter if you dont want to answer, we understand that."
        },
        {
          "id": 309126,
          "postDate": "2018-04-04T17:27:18.470Z",
          "content": "<p>I do not think this is impressive, I was lucky with some feature engineering early, but currently I am stuck. \nRegarding your question, I am using grouping of several columns including IP. I just do not use IP directly and IP frequencies (percent of is_attributed=1 for IP). There are several topics warning about using IP as a feature, so I decided not to use it (I may reconsider this later). \nThe problem here is that IP column has a lot of info, but test does not have all IPs that we know from train. I beleive it is crucial to extract info using IP without actually using IP address itself.</p>",
          "rawMarkdown": "I do not think this is impressive, I was lucky with some feature engineering early, but currently I am stuck. \nRegarding your question, I am using grouping of several columns including IP. I just do not use IP directly and IP frequencies (percent of is_attributed=1 for IP). There are several topics warning about using IP as a feature, so I decided not to use it (I may reconsider this later). \nThe problem here is that IP column has a lot of info, but test does not have all IPs that we know from train. I beleive it is crucial to extract info using IP without actually using IP address itself.",
          "votes": 12
        },
        {
          "id": 309134,
          "postDate": "2018-04-04T17:37:45.993Z",
          "content": "<p>Ok understood, thanks! That is also what I've been doing, and most of my current features are derived from groupings on ip. I definitely agree with you re: finding ways to extract info without using the explicit address. </p>\n\n<p>I'm also in the same boat, stuck on FE progress.</p>",
          "rawMarkdown": "Ok understood, thanks! That is also what I've been doing, and most of my current features are derived from groupings on ip. I definitely agree with you re: finding ways to extract info without using the explicit address. \n\nI'm also in the same boat, stuck on FE progress.",
          "votes": 2
        },
        {
          "id": 309139,
          "postDate": "2018-04-04T17:53:15.250Z",
          "content": "<p>Thank you very much Alex! </p>",
          "rawMarkdown": "Thank you very much Alex! "
        },
        {
          "id": 309168,
          "postDate": "2018-04-04T19:01:02.317Z",
          "content": "<p>@Alexander Firsov: Do you mind sharing why you choosed to train on only day 8 ? </p>",
          "rawMarkdown": "@Alexander Firsov: Do you mind sharing why you choosed to train on only day 8 ? "
        },
        {
          "id": 309206,
          "postDate": "2018-04-04T19:50:48Z",
          "content": "<p>I described my single best model, and it was from day 8, but I have similar result from day 7.\nI do not plan to train only on day 8. \nWhat I want to avoid is an effect when model works good for data from next hour, but not next day, that is why I started with first two days.</p>",
          "rawMarkdown": "I described my single best model, and it was from day 8, but I have similar result from day 7.\nI do not plan to train only on day 8. \nWhat I want to avoid is an effect when model works good for data from next hour, but not next day, that is why I started with first two days.\n"
        },
        {
          "id": 309326,
          "postDate": "2018-04-05T03:59:32.890Z",
          "content": "<p>So great</p>",
          "rawMarkdown": "So great"
        },
        {
          "id": 309829,
          "postDate": "2018-04-06T03:12:02.763Z",
          "content": "<p>@Alexander Firsov, thank you for sharing your method but what do you call frequency? Is the following variable avg(is_attributed) over (partition by feature1,feature2...) a frequency? Thank you by advance for your answer</p>",
          "rawMarkdown": "@Alexander Firsov, thank you for sharing your method but what do you call frequency? Is the following variable avg(is_attributed) over (partition by feature1,feature2...) a frequency? Thank you by advance for your answer"
        },
        {
          "id": 310736,
          "postDate": "2018-04-08T12:36:26.700Z",
          "content": "<p>Yes. It's commonly known as mean-encoding: <a href=\"https://www.coursera.org/learn/competitive-data-science/lecture/b5Gxv/concept-of-mean-encoding\">https://www.coursera.org/learn/competitive-data-science/lecture/b5Gxv/concept-of-mean-encoding</a></p>",
          "rawMarkdown": "Yes. It's commonly known as mean-encoding: https://www.coursera.org/learn/competitive-data-science/lecture/b5Gxv/concept-of-mean-encoding",
          "votes": 2
        },
        {
          "id": 310746,
          "postDate": "2018-04-08T13:25:16.723Z",
          "content": "<p>Authman, thanks for the reference. I used the term from <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">popular kernel</a></p>",
          "rawMarkdown": "Authman, thanks for the reference. I used the term from [popular kernel][1]\n\n\n  [1]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing",
          "votes": 1
        }
      ]
    },
    {
      "id": 318823,
      "postDate": "2018-04-24T14:45:53.010Z",
      "content": "<p>update: single model with 11 features trained on day 9 only (~60M rows) get to .9802</p>",
      "rawMarkdown": "update: single model with 11 features trained on day 9 only (~60M rows) get to .9802",
      "votes": 12
    },
    {
      "id": 313326,
      "postDate": "2018-04-13T05:44:14.783Z",
      "content": "<p>My current lb score is a single model. 0.9784 with 7 additional features - 2 groupby count features, 2 delta-time features, 3 confRate features.  Trained on day 9 with 0.1 validation. </p>\n\n<p>Edit: the new score of 0.9796 is also from a single model using 3 more features, still trained on day 9 only. </p>\n\n<p>Edit: 0.9798, single model with 18 features, trained on day 9 only - cant go any further with my 32GB ram... will try something else next.</p>",
      "rawMarkdown": "My current lb score is a single model. 0.9784 with 7 additional features - 2 groupby count features, 2 delta-time features, 3 confRate features.  Trained on day 9 with 0.1 validation. \n\nEdit: the new score of 0.9796 is also from a single model using 3 more features, still trained on day 9 only. \n\nEdit: 0.9798, single model with 18 features, trained on day 9 only - cant go any further with my 32GB ram... will try something else next.",
      "votes": 12,
      "replies": [
        {
          "id": 313332,
          "postDate": "2018-04-13T05:52:04.893Z",
          "content": "<p>well done. What does confRate mean here? do you mean target encoding related features?</p>",
          "rawMarkdown": "well done. What does confRate mean here? do you mean target encoding related features?"
        },
        {
          "id": 313344,
          "postDate": "2018-04-13T06:09:57.073Z",
          "content": "<p>confidence rate - see <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">the great FE kernel</a>.  But this is quite tricky because it is essentially just leaking. Some of the confRate features I tested actually worsen the result a lot (to sub 0.95).</p>",
          "rawMarkdown": "confidence rate - see [the great FE kernel][1].  But this is quite tricky because it is essentially just leaking. Some of the confRate features I tested actually worsen the result a lot (to sub 0.95).\n\n\n  [1]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing"
        },
        {
          "id": 313373,
          "postDate": "2018-04-13T07:11:16.230Z",
          "content": "<p>Great! And what does delta-time mean? May be you can suggest any kernel... Thanx!</p>",
          "rawMarkdown": "Great! And what does delta-time mean? May be you can suggest any kernel... Thanx!"
        },
        {
          "id": 313545,
          "postDate": "2018-04-13T12:56:23.820Z",
          "content": "<p>Idea is from <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361\">this thread</a>. </p>",
          "rawMarkdown": "Idea is from [this thread][1]. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361",
          "votes": 2
        },
        {
          "id": 313554,
          "postDate": "2018-04-13T13:24:30.440Z",
          "content": "<p>@congb, You're very kind to share that much.  </p>",
          "rawMarkdown": "@congb, You're very kind to share that much.  ",
          "votes": -2
        },
        {
          "id": 313587,
          "postDate": "2018-04-13T14:17:18.020Z",
          "content": "<p>@congb, thank you very much for your sharing! And I would like to discuss some questions with you guys here:  Is using day 9 as train and 0.1 as validation  will overfit to public lb?  If you guys remember that 0.9736 public kernel a few days ago, this kernel uses last N million rows and when other kagglers reproduce this kernel with more rows, the public lb score decreases. In terms of my experiments, I use day 7 and 8 as train and day 9 as validation, and I found only one of confRate features actually improves validation scores... </p>",
          "rawMarkdown": "@congb, thank you very much for your sharing! And I would like to discuss some questions with you guys here:  Is using day 9 as train and 0.1 as validation  will overfit to public lb?  If you guys remember that 0.9736 public kernel a few days ago, this kernel uses last N million rows and when other kagglers reproduce this kernel with more rows, the public lb score decreases. In terms of my experiments, I use day 7 and 8 as train and day 9 as validation, and I found only one of confRate features actually improves validation scores... "
        },
        {
          "id": 313591,
          "postDate": "2018-04-13T14:31:41.600Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 313593,
          "postDate": "2018-04-13T14:34:37.757Z",
          "content": "<p>@CPMP, No problem! I learned a lot from your input in multiple threads. </p>\n\n<p>@Snorlax, The reason I'm using only day 9 is for faster modeling (need to go back and forth tuning things). But I believe that the one you described is a better strategy and for final submission I will definitely do that. And I'm sure my .9784 submission is an overfitting to the public lb, to some extent. confRate features are tricky and I did spent quite some time to decide which ones to include - but still, I'm also very unconfident with the confidence rate...</p>",
          "rawMarkdown": "@CPMP, No problem! I learned a lot from your input in multiple threads. \n\n@Snorlax, The reason I'm using only day 9 is for faster modeling (need to go back and forth tuning things). But I believe that the one you described is a better strategy and for final submission I will definitely do that. And I'm sure my .9784 submission is an overfitting to the public lb, to some extent. confRate features are tricky and I did spent quite some time to decide which ones to include - but still, I'm also very unconfident with the confidence rate...",
          "votes": 3
        },
        {
          "id": 313594,
          "postDate": "2018-04-13T14:36:16.430Z",
          "content": "<p>Hi @Zijun Yao, thank you very much for your reply. Actually, I only found one confRate feature is useful(but not that useful) and I've looked through almost all kernels using confRate and I think they could achieve the same public lb score without using confRate based on what other features they already have.</p>",
          "rawMarkdown": "Hi @Zijun Yao, thank you very much for your reply. Actually, I only found one confRate feature is useful(but not that useful) and I've looked through almost all kernels using confRate and I think they could achieve the same public lb score without using confRate based on what other features they already have."
        },
        {
          "id": 313599,
          "postDate": "2018-04-13T14:43:52.890Z",
          "content": "<p>@Zijun Yao, may I ask your validation score on day 9 4:00 to 5:00 and validation score on day 9 after 5:00?</p>",
          "rawMarkdown": "@Zijun Yao, may I ask your validation score on day 9 4:00 to 5:00 and validation score on day 9 after 5:00?"
        },
        {
          "id": 313603,
          "postDate": "2018-04-13T14:50:10.957Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 313604,
          "postDate": "2018-04-13T14:57:31.863Z",
          "content": "<p>@Zijun Yao, I mean your best model's validation score without confRate. In terms of my model, it's validation score on data after day 9 5:00 is 0.003 higher than the validation score on data between day 9 4:00 to 5:00.  I ask this because data on day 9 4:00 to 5:00 mimics public lb data and data after day 9 5:00 mimics private lb data. And I really want to know what is the gap between validation score on 'simulated' public lb data and 'simulated' private lb data. </p>",
          "rawMarkdown": "@Zijun Yao, I mean your best model's validation score without confRate. In terms of my model, it's validation score on data after day 9 5:00 is 0.003 higher than the validation score on data between day 9 4:00 to 5:00.  I ask this because data on day 9 4:00 to 5:00 mimics public lb data and data after day 9 5:00 mimics private lb data. And I really want to know what is the gap between validation score on 'simulated' public lb data and 'simulated' private lb data. "
        },
        {
          "id": 313606,
          "postDate": "2018-04-13T15:01:20.437Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 313608,
          "postDate": "2018-04-13T15:04:34.600Z",
          "content": "<p>@Zijun Yao, Got it! Thank you so much! And thank you again for your sharing @congb! I'm very lucky to meet you guys who generously share your valuable experience. </p>",
          "rawMarkdown": "@Zijun Yao, Got it! Thank you so much! And thank you again for your sharing @congb! I'm very lucky to meet you guys who generously share your valuable experience. "
        },
        {
          "id": 314991,
          "postDate": "2018-04-16T15:56:33.430Z",
          "content": "<p>@Snorlax are you using the entire day 9 for validation?</p>",
          "rawMarkdown": "@Snorlax are you using the entire day 9 for validation?",
          "votes": 1
        },
        {
          "id": 314994,
          "postDate": "2018-04-16T15:58:25.983Z",
          "content": "<p>9/1 training/val split</p>",
          "rawMarkdown": "9/1 training/val split"
        },
        {
          "id": 315015,
          "postDate": "2018-04-16T16:32:19.397Z",
          "rawMarkdown": ""
        },
        {
          "id": 315021,
          "postDate": "2018-04-16T16:37:26.377Z",
          "content": "<p>I would agree with you on both points - Validation strategy and Conf rates. Day 9 training only is definitely a strong overfit. I like Pranav's latest kernel where he trains all the data in different chunks and averages the predictions. But of course, nothing like it If you can train it all in one go with enough resources. </p>\n\n<p>Did you train 7 and 8 separately and average predictions from two models or Day 7 first and use the same fit model on Day 8 ( If that is even a relevant approach )?</p>",
          "rawMarkdown": "I would agree with you on both points - Validation strategy and Conf rates. Day 9 training only is definitely a strong overfit. I like Pranav's latest kernel where he trains all the data in different chunks and averages the predictions. But of course, nothing like it If you can train it all in one go with enough resources. \n\nDid you train 7 and 8 separately and average predictions from two models or Day 7 first and use the same fit model on Day 8 ( If that is even a relevant approach )?"
        },
        {
          "id": 315080,
          "postDate": "2018-04-16T18:12:06.977Z",
          "content": "<p>@ Shanth， I'm still doing preprocesing and feature generation for day 7. But combining preds from day 8 and day 9  does help me a little （~0.0005)</p>",
          "rawMarkdown": "@ Shanth， I'm still doing preprocesing and feature generation for day 7. But combining preds from day 8 and day 9  does help me a little （~0.0005)"
        }
      ]
    },
    {
      "id": 320603,
      "postDate": "2018-04-29T08:31:21.047Z",
      "content": "<p>0.9822 with full data,60 features,lgbm</p>",
      "rawMarkdown": "0.9822 with full data,60 features,lgbm",
      "votes": 9,
      "replies": [
        {
          "id": 320618,
          "postDate": "2018-04-29T09:32:02.303Z",
          "content": "<p>are you using target encoding?</p>",
          "rawMarkdown": "are you using target encoding?"
        },
        {
          "id": 320620,
          "postDate": "2018-04-29T09:52:19.110Z",
          "content": "<p>not yet,will try </p>",
          "rawMarkdown": "not yet,will try "
        },
        {
          "id": 320674,
          "postDate": "2018-04-29T13:19:52.317Z",
          "content": "<p>Thanks, very encouraging.</p>\n\n<p>I'm at 0.9820 lgb with 33 features.  I only have to add 27 features and buy some RAM to meet you ;)</p>",
          "rawMarkdown": "Thanks, very encouraging.\n\nI'm at 0.9820 lgb with 33 features.  I only have to add 27 features and buy some RAM to meet you ;)",
          "votes": 5
        },
        {
          "id": 320689,
          "postDate": "2018-04-29T14:05:01.800Z",
          "content": "<p>0.9811 lgb with 14 features. Unfortunately, cannot improve anymore with other features..</p>",
          "rawMarkdown": "0.9811 lgb with 14 features. Unfortunately, cannot improve anymore with other features..",
          "votes": 1
        },
        {
          "id": 320738,
          "postDate": "2018-04-29T16:51:27.627Z",
          "content": "<p>May I ask how you guys are iterating with different features? For instance - are you first thinking of a feature, coding it, adding it to your main model, and then checking the validation auc? Or do you test out multiple features at a time, etc.?</p>\n\n<p>I've noticed that when I add in different features, other features suddenly don't seem to work as well. Based on this, I'm thinking that the interaction between all the features you have in your model is crucial, and, as a result, really the best way to figure out if your model performs well is constantly iterating on the main model. Am I going about this correctly?</p>",
          "rawMarkdown": "May I ask how you guys are iterating with different features? For instance - are you first thinking of a feature, coding it, adding it to your main model, and then checking the validation auc? Or do you test out multiple features at a time, etc.?\n\nI've noticed that when I add in different features, other features suddenly don't seem to work as well. Based on this, I'm thinking that the interaction between all the features you have in your model is crucial, and, as a result, really the best way to figure out if your model performs well is constantly iterating on the main model. Am I going about this correctly?",
          "votes": 3
        },
        {
          "id": 320759,
          "postDate": "2018-04-29T18:16:56.847Z",
          "content": "<p>I mostly do what you write:</p>\n\n<blockquote>\n  <p>are you first thinking of a feature, coding it, adding it to your main model, and then checking the validation auc? </p>\n</blockquote>\n\n<p>But sometimes adding more than one feature at a time works, while adding each of them degrades local cv.  This is a bit of black magic ;)</p>\n\n<p>An other way is to add a bunch of features, then try removing them one by one to see if they were useful.</p>",
          "rawMarkdown": "I mostly do what you write:\n\n&gt; are you first thinking of a feature, coding it, adding it to your main model, and then checking the validation auc? \n\nBut sometimes adding more than one feature at a time works, while adding each of them degrades local cv.  This is a bit of black magic ;)\n\nAn other way is to add a bunch of features, then try removing them one by one to see if they were useful.",
          "votes": 7
        },
        {
          "id": 320781,
          "postDate": "2018-04-29T20:18:17.303Z",
          "content": "<blockquote>\n  <p><strong>CPMP wrote</strong></p>\n  \n  <p>An other way is to add a bunch of features, then try removing them one by one to see if they were useful.</p>\n</blockquote>\n\n<p>That's the approach I use</p>",
          "rawMarkdown": "\n&gt; **CPMP wrote**\n\n&gt; An other way is to add a bunch of features, then try removing them one by one to see if they were useful.\n\nThat's the approach I use",
          "votes": 3
        },
        {
          "id": 320858,
          "postDate": "2018-04-30T03:26:03.430Z",
          "content": "<p>@CPMP @Serigne Thanks! Quick question - how would you interpret it if you train a model with a group of, say, 5 features and they seem to perform well according to validation auc. But then when you add them to your main model, the validation auc of the main model decreases slightly (like .0002) from where it was without those extra 5 features. Does that mean those 5 features are just bad or is it ok to still keep them in the model?</p>",
          "rawMarkdown": "@CPMP @Serigne Thanks! Quick question - how would you interpret it if you train a model with a group of, say, 5 features and they seem to perform well according to validation auc. But then when you add them to your main model, the validation auc of the main model decreases slightly (like .0002) from where it was without those extra 5 features. Does that mean those 5 features are just bad or is it ok to still keep them in the model?"
        },
        {
          "id": 320874,
          "postDate": "2018-04-30T04:31:46.057Z",
          "content": "<p>@CPMP I think you can get 0.9823 with 40 features.I don't want to check my features one by one,so just left them.</p>",
          "rawMarkdown": "@CPMP I think you can get 0.9823 with 40 features.I don't want to check my features one by one,so just left them.",
          "votes": 1
        },
        {
          "id": 320876,
          "postDate": "2018-04-30T04:38:32.330Z",
          "content": "<p>@Edward Chen  At an earlier stage,I can try features one by one using sample data(only 12 hours of 9th).Now every time I add 3 ~ 5 new features to main model to see the val score,if it got better I will submit .I found public learderboard score always match val score.</p>",
          "rawMarkdown": "@Edward Chen  At an earlier stage,I can try features one by one using sample data(only 12 hours of 9th).Now every time I add 3 ~ 5 new features to main model to see the val score,if it got better I will submit .I found public learderboard score always match val score.",
          "votes": 1
        },
        {
          "id": 320877,
          "postDate": "2018-04-30T04:43:31.973Z",
          "content": "<p>ecuse me, can you share how to setup your validation set</p>",
          "rawMarkdown": "ecuse me, can you share how to setup your validation set"
        },
        {
          "id": 320879,
          "postDate": "2018-04-30T04:49:27.517Z",
          "content": "<p><a href=\"/senkin13\">@senkin13</a> interesting. You mentioned that you add 3-5 new features at a time. What happens if it doesn't improve val score though? Do you remove them?</p>",
          "rawMarkdown": "@senkin13 interesting. You mentioned that you add 3-5 new features at a time. What happens if it doesn't improve val score though? Do you remove them?"
        },
        {
          "id": 320902,
          "postDate": "2018-04-30T06:35:45Z",
          "content": "<p>Well, this magic is driving me crazy...</p>",
          "rawMarkdown": "Well, this magic is driving me crazy...",
          "votes": 1
        },
        {
          "id": 320931,
          "postDate": "2018-04-30T08:34:06.597Z",
          "content": "<p>Hello senkin13. I wonder know what's your learning rate and model\"s runing time. Glad to see your reply. </p>",
          "rawMarkdown": "Hello senkin13. I wonder know what's your learning rate and model\"s runing time. Glad to see your reply. "
        },
        {
          "id": 320942,
          "postDate": "2018-04-30T08:53:05.113Z",
          "content": "<p>@EdwardChen, I test features by adding them to the model.  That's the only test that matters.  I only have one model.  Not sure why you speak of a 'main' model here.</p>",
          "rawMarkdown": "@EdwardChen, I test features by adding them to the model.  That's the only test that matters.  I only have one model.  Not sure why you speak of a 'main' model here."
        },
        {
          "id": 320944,
          "postDate": "2018-04-30T08:55:34.787Z",
          "content": "<p><a href=\"/senkin13\">@senkin13</a>, I agree.  </p>\n\n<p>I cannot afford 50 features I think, because of RAM.  That's why I perform feature elimination.  And yes, I hope to get to 0.9823 with less than 40 features.</p>",
          "rawMarkdown": "@senkin13, I agree.  \n\nI cannot afford 50 features I think, because of RAM.  That's why I perform feature elimination.  And yes, I hope to get to 0.9823 with less than 40 features."
        },
        {
          "id": 320949,
          "postDate": "2018-04-30T08:59:51.097Z",
          "content": "<p>random split(0.95) for train validation set.\nif new features didn't improve val score,will remove them expect they have meaningful buisness sense.\nlower learning rate got better score, 0.02=0.03&gt;0.05&gt;0.1.\nafter removing dtrain from lgbm parameter [valid_sets],can reduce running time from 6,7 hours to 2,3hours</p>",
          "rawMarkdown": "random split(0.95) for train validation set.\nif new features didn't improve val score,will remove them expect they have meaningful buisness sense.\nlower learning rate got better score, 0.02=0.03&gt;0.05&gt;0.1.\nafter removing dtrain from lgbm parameter [valid_sets],can reduce running time from 6,7 hours to 2,3hours\n",
          "votes": 2
        },
        {
          "id": 320950,
          "postDate": "2018-04-30T09:02:12.413Z",
          "content": "<p>Hello CPMP. I cannot afford with memory, too.  I wonder know what's your learning rate and model\"s runing time. Glad to see your reply.</p>",
          "rawMarkdown": "Hello CPMP. I cannot afford with memory, too.  I wonder know what's your learning rate and model\"s runing time. Glad to see your reply.\n\n"
        },
        {
          "id": 320952,
          "postDate": "2018-04-30T09:08:15.947Z",
          "content": "<p>Thanks senkin13. I was troubled by the time of training.Will you train two times when you add new features? First you determine the number of iterations and second use full data train.</p>",
          "rawMarkdown": "Thanks senkin13. I was troubled by the time of training.Will you train two times when you add new features? First you determine the number of iterations and second use full data train."
        },
        {
          "id": 320953,
          "postDate": "2018-04-30T09:08:35.680Z",
          "content": "<p>I didn't tune learning rate, I am using 0.1.  I keep these tuning for when I'll run out of fuel with feature engineering.  A larger learning rate means faster training, hence faster experiments.</p>\n\n<p>I don't do tuning often, because when you do you no longer can compare with previous runs.  Now I have a good setting, my validation score matches exactly my LB score since I reached 0.9818.</p>",
          "rawMarkdown": "I didn't tune learning rate, I am using 0.1.  I keep these tuning for when I'll run out of fuel with feature engineering.  A larger learning rate means faster training, hence faster experiments.\n\nI don't do tuning often, because when you do you no longer can compare with previous runs.  Now I have a good setting, my validation score matches exactly my LB score since I reached 0.9818.",
          "votes": 3
        },
        {
          "id": 320954,
          "postDate": "2018-04-30T09:11:40.873Z",
          "content": "<p>Thanks CPMP. I always learn a lot from you. :)</p>",
          "rawMarkdown": "Thanks CPMP. I always learn a lot from you. :)",
          "votes": 2
        },
        {
          "id": 320975,
          "postDate": "2018-04-30T09:59:28.957Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 321032,
          "postDate": "2018-04-30T12:45:04.383Z",
          "content": "<p>@CPMP Thanks for the reply and sorry for the confusion. Please ignore what I said about the \"main model\". Basically, I was just asking - if you're testing some features in your model and they don't happen to improve the val score, do you keep them in the model or hope that some interactions with future features will help to improve the val score with those features but I believe @senkin has answered that - so thanks as well @senkin</p>",
          "rawMarkdown": "@CPMP Thanks for the reply and sorry for the confusion. Please ignore what I said about the \"main model\". Basically, I was just asking - if you're testing some features in your model and they don't happen to improve the val score, do you keep them in the model or hope that some interactions with future features will help to improve the val score with those features but I believe @senkin has answered that - so thanks as well @senkin"
        },
        {
          "id": 321089,
          "postDate": "2018-04-30T15:03:48.700Z",
          "content": "<p>@Edward Chen, no need to apologize, no harm done at all!</p>\n\n<p>If the features don't improve my model I discard them.</p>",
          "rawMarkdown": "@Edward Chen, no need to apologize, no harm done at all!\n\nIf the features don't improve my model I discard them.",
          "votes": 1
        },
        {
          "id": 321494,
          "postDate": "2018-05-01T12:21:56.147Z",
          "content": "<p>@CPMP : with a learning rate  of 0.1 it probably takes you like more than 1 or 2 hours to fit your model on the whole training set. So are you doing this feature selection process on the full training set ? Or do you create a smaller dataset that is correlated to it ?</p>\n\n<p>Thanks for all the advices !</p>",
          "rawMarkdown": "@CPMP : with a learning rate  of 0.1 it probably takes you like more than 1 or 2 hours to fit your model on the whole training set. So are you doing this feature selection process on the full training set ? Or do you create a smaller dataset that is correlated to it ?\n\nThanks for all the advices !"
        },
        {
          "id": 321509,
          "postDate": "2018-05-01T13:09:03.573Z",
          "content": "<blockquote>\n  <p>are you doing this feature selection process on the full training set ? </p>\n</blockquote>\n\n<p>@Antoine, I don't.  First, you need some validation data, hence you cannot use all train data for training.  Second, using a subset is faster.  Key is to use a representative subset, so that improvement on it yields improvement when training on all the dataset.  That's what I focused on in the early days of the competition: be able to run reliable experiments as fast as possible.  </p>",
          "rawMarkdown": "&gt; are you doing this feature selection process on the full training set ? \n\n@Antoine, I don't.  First, you need some validation data, hence you cannot use all train data for training.  Second, using a subset is faster.  Key is to use a representative subset, so that improvement on it yields improvement when training on all the dataset.  That's what I focused on in the early days of the competition: be able to run reliable experiments as fast as possible.  ",
          "votes": 1
        },
        {
          "id": 322104,
          "postDate": "2018-05-02T12:25:35.960Z",
          "content": "<blockquote>\n  <p>@CPMP I think you can get 0.9823 with 40 features.</p>\n</blockquote>\n\n<p><a href=\"/senkin13\">@senkin13</a>, you are right, just got there ;)</p>",
          "rawMarkdown": "&gt; @CPMP I think you can get 0.9823 with 40 features.\n\n@senkin13, you are right, just got there ;)",
          "votes": -1
        },
        {
          "id": 322131,
          "postDate": "2018-05-02T13:00:24.940Z",
          "content": "<p>@CPMP : Merci ! I guess my main problem during this competition was to find a correct training / validation set for testing, and one for submitting my score. I guess with experience it becomes easier...</p>",
          "rawMarkdown": "@CPMP : Merci ! I guess my main problem during this competition was to find a correct training / validation set for testing, and one for submitting my score. I guess with experience it becomes easier...",
          "votes": 1
        },
        {
          "id": 322134,
          "postDate": "2018-05-02T13:07:11.407Z",
          "content": "<p>@Antoine, de rien!</p>\n\n<p>Having a reliable local validation setting is the single most important thing in a competition.  My best results were obtained when I found a good CV setting early in the competition.</p>",
          "rawMarkdown": "@Antoine, de rien!\n\nHaving a reliable local validation setting is the single most important thing in a competition.  My best results were obtained when I found a good CV setting early in the competition.",
          "votes": 1
        },
        {
          "id": 322194,
          "postDate": "2018-05-02T14:29:43.343Z",
          "content": "<p>@CPMP congrats!I think there are still many good features to mine,will you continue to improve single model or make others to do ensembling</p>",
          "rawMarkdown": "@CPMP congrats!I think there are still many good features to mine,will you continue to improve single model or make others to do ensembling"
        },
        {
          "id": 322201,
          "postDate": "2018-05-02T14:40:19.400Z",
          "content": "<p><a href=\"/senkin13\">@senkin13</a>, Thanks!  I think I'll move to ensembling as I am reaching the limits of my HW.  I wish I had bought that extra memory I thought about...  I can extend my swap, but the code is already quite slow, paging will kill it.</p>",
          "rawMarkdown": "@senkin13, Thanks!  I think I'll move to ensembling as I am reaching the limits of my HW.  I wish I had bought that extra memory I thought about...  I can extend my swap, but the code is already quite slow, paging will kill it."
        },
        {
          "id": 322377,
          "postDate": "2018-05-02T20:36:04.593Z",
          "content": "<blockquote>\n  <p>@CPMP: I cannot afford 50 features I think, because of RAM.</p>\n</blockquote>\n\n<p>How much ram do you have? I'm in the process of running my largest model (31 features) but:</p>\n\n<pre><code>➜  ~ free -g\n              total        used        free      shared  buff/cache   available\nMem:             62          16          45           0           0          45\nSwap:           127          21         106\n</code></pre>\n\n<p>Only takes 16GB of ram during training. Are you talking about RAM usage during lgb dataset conversion? If so, there's a thread out there which instructs to use <code>'two_round':True</code> lgb_config parameter, so your data is automatically written from pd-&gt;disk, then read from disk-&gt;lgb... rather than pandas-&gt;float64-&gt;lgb. On the other hand, if you're limited by FE, you can always write your features to disk intermediary while building them, then recompile. For example, I current build all my count features, write to disk then drop from the df, clickdelta features, write to disk then drop from the df, etc. Then reload all the feathers to reassemble.</p>",
          "rawMarkdown": "&gt; @CPMP: I cannot afford 50 features I think, because of RAM.\n\nHow much ram do you have? I'm in the process of running my largest model (31 features) but:\n\n    ➜  ~ free -g\n                  total        used        free      shared  buff/cache   available\n    Mem:             62          16          45           0           0          45\n    Swap:           127          21         106\n\nOnly takes 16GB of ram during training. Are you talking about RAM usage during lgb dataset conversion? If so, there's a thread out there which instructs to use `'two_round':True` lgb_config parameter, so your data is automatically written from pd-&gt;disk, then read from disk-&gt;lgb... rather than pandas-&gt;float64-&gt;lgb. On the other hand, if you're limited by FE, you can always write your features to disk intermediary while building them, then recompile. For example, I current build all my count features, write to disk then drop from the df, clickdelta features, write to disk then drop from the df, etc. Then reload all the feathers to reassemble.",
          "votes": 2
        },
        {
          "id": 322395,
          "postDate": "2018-05-02T21:11:21.190Z",
          "content": "<blockquote>\n  <p>Then reload all the feathers to reassemble.</p>\n</blockquote>\n\n<p>That's what I am doing ;)</p>\n\n<p>I'll try the two rounds, thanks for the tip.</p>",
          "rawMarkdown": "&gt; Then reload all the feathers to reassemble.\n\nThat's what I am doing ;)\n\nI'll try the two rounds, thanks for the tip."
        },
        {
          "id": 322459,
          "postDate": "2018-05-03T01:53:33.073Z",
          "content": "<p>@CPMP I'm using a similar approach but with csv... But I see that the merging of this data to the original df is talking lot of time... I mean hours... Is there an efficient way to do that?</p>",
          "rawMarkdown": "@CPMP I'm using a similar approach but with csv... But I see that the merging of this data to the original df is talking lot of time... I mean hours... Is there an efficient way to do that?"
        },
        {
          "id": 322494,
          "postDate": "2018-05-03T04:50:50.373Z",
          "content": "<p>If you save each feature as a single column cssv, then you don't need to use merge, you simply copy the values.  You need to merge if you change the order of the rows.  See the advice given by <a href=\"/spongebob\">@spongebob</a>: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55601\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55601</a></p>",
          "rawMarkdown": "If you save each feature as a single column cssv, then you don't need to use merge, you simply copy the values.  You need to merge if you change the order of the rows.  See the advice given by @spongebob: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55601",
          "votes": 1
        },
        {
          "id": 322497,
          "postDate": "2018-05-03T04:56:40.437Z",
          "content": "<p>You can merge multiple columns using that same technique:</p>\n\n<pre><code>load = pd.read_feather('./new_features.ftr')\ndf[load.columns] = load\n</code></pre>\n\n<p>Merging in 20 gb of column data should be doable in 5-20 secs. Def shouldn't take minutes, or--God forbid--hours.</p>",
          "rawMarkdown": "You can merge multiple columns using that same technique:\n\n    load = pd.read_feather('./new_features.ftr')\n    df[load.columns] = load\n\nMerging in 20 gb of column data should be doable in 5-20 secs. Def shouldn't take minutes, or--God forbid--hours.",
          "votes": 3
        },
        {
          "id": 322508,
          "postDate": "2018-05-03T05:27:57.157Z",
          "content": "<p>@authan, you are right, this is how it should be done.  </p>\n\n<p>I would not call it merge though as it may confuse readers given there is a merge() function for dataframes.  What you describe is a way to efficiently append a column, instead of using concat().</p>\n\n<p>I do it in a one liner, that way you don't keep a pointer on the loaded feather file:</p>\n\n<pre><code>df[load.columns] = pd.read_feather('./new_features.ftr')\n</code></pre>\n\n<p>This way the loaded file will be reclaimed by the gc.</p>",
          "rawMarkdown": "@authan, you are right, this is how it should be done.  \n\nI would not call it merge though as it may confuse readers given there is a merge() function for dataframes.  What you describe is a way to efficiently append a column, instead of using concat().\n\nI do it in a one liner, that way you don't keep a pointer on the loaded feather file:\n\n    df[load.columns] = pd.read_feather('./new_features.ftr')\n\nThis way the loaded file will be reclaimed by the gc."
        },
        {
          "id": 322534,
          "postDate": "2018-05-03T06:26:13.993Z",
          "content": "<p>I'm curious what you said about pd -&gt; disk -&gt; lgb, how would one do that and specify categorical features? I did a Google search and didn't find anything, can you please help? </p>",
          "rawMarkdown": "I'm curious what you said about pd -&gt; disk -&gt; lgb, how would one do that and specify categorical features? I did a Google search and didn't find anything, can you please help? "
        },
        {
          "id": 322540,
          "postDate": "2018-05-03T06:54:23.110Z",
          "content": "<p>@CPMP, <a href=\"/authman\">@authman</a> I think I'm generating the CSV file a bit different. </p>\n\n<p>Lets say I group on IP and OS and get the counts. So, the CSV file has 3 columns now, IP, OS, IP_OS_Count. So, now this CSV file has to be merged to the original DF..</p>\n\n<ul>\n<li>IP OS</li>\n<li>1 1 </li>\n<li>1 1</li>\n<li>2 2</li>\n</ul>\n\n<p>So, the new CSV File will be</p>\n\n<ul>\n<li>IP OS IP_OS_Count</li>\n<li>1 1 2 </li>\n<li>2 2 1</li>\n</ul>\n\n<p>So, now this count has to be added to each of the rows and the new df will be as below.</p>\n\n<ul>\n<li>IP OS IP_OS_Count</li>\n<li>1 1 2</li>\n<li>1 1 2</li>\n<li>2 2 1</li>\n</ul>\n\n<p>Am I doing it incorrectly?</p>",
          "rawMarkdown": "@CPMP, @authman I think I'm generating the CSV file a bit different. \n\nLets say I group on IP and OS and get the counts. So, the CSV file has 3 columns now, IP, OS, IP_OS_Count. So, now this CSV file has to be merged to the original DF..\n\n - IP OS\n - 1 1 \n - 1 1\n - 2 2\n\nSo, the new CSV File will be\n\n - IP OS IP_OS_Count\n - 1 1 2 \n - 2 2 1\n\nSo, now this count has to be added to each of the rows and the new df will be as below.\n\n - IP OS IP_OS_Count\n - 1 1 2\n - 1 1 2\n - 2 2 1\n\nAm I doing it incorrectly?"
        },
        {
          "id": 322555,
          "postDate": "2018-05-03T07:22:53.320Z",
          "content": "<p><a href=\"/samrat\">@samrat</a>, You are doing it correctly.  But once done, you should save the column you just added to your main dataframe as a separate file instead of saving the whole dataframe.  Some of the public kernels use that: they save each feature separately.  It is way easier to make experiment that way, you just assemble the features you want to test from the individual files.</p>",
          "rawMarkdown": "@samrat, You are doing it correctly.  But once done, you should save the column you just added to your main dataframe as a separate file instead of saving the whole dataframe.  Some of the public kernels use that: they save each feature separately.  It is way easier to make experiment that way, you just assemble the features you want to test from the individual files.",
          "votes": 3
        },
        {
          "id": 322567,
          "postDate": "2018-05-03T07:52:09.890Z",
          "content": "<p>@CPMP Thanks a lot.. Now I got it.. Wish I knew it much earlier :(</p>",
          "rawMarkdown": "@CPMP Thanks a lot.. Now I got it.. Wish I knew it much earlier :("
        },
        {
          "id": 322589,
          "postDate": "2018-05-03T08:42:30.187Z",
          "content": "<p>No pb, but you could have learned it from public kernels ;)</p>\n\n<p>I don't use public kernels as is, as they often overfit, but some contain good ideas, and reading them is really something I value.</p>",
          "rawMarkdown": "No pb, but you could have learned it from public kernels ;)\n\nI don't use public kernels as is, as they often overfit, but some contain good ideas, and reading them is really something I value."
        },
        {
          "id": 322676,
          "postDate": "2018-05-03T12:14:51.643Z",
          "content": "<p>@HuyenNguyen same as normal, it's handled for you internally in the lgb python wrapper. You don't have to roll your own file, just pass that parameter into your config. So you can either <code>df.col=df.col.astype('category')</code>, or alternatively, you can use the <code>categorical_features=[...]</code> argument of <code>lgb.train()</code>.</p>",
          "rawMarkdown": "@HuyenNguyen same as normal, it's handled for you internally in the lgb python wrapper. You don't have to roll your own file, just pass that parameter into your config. So you can either `df.col=df.col.astype('category')`, or alternatively, you can use the `categorical_features=[...]` argument of `lgb.train()`."
        },
        {
          "id": 322706,
          "postDate": "2018-05-03T13:24:00.587Z",
          "content": "<p>Hi@CPMP~ Do you find 'attributed_time' useful? Again, I'm glad if you refuse to answer this coz it's too personal, lol~</p>",
          "rawMarkdown": "Hi@CPMP~ Do you find 'attributed_time' useful? Again, I'm glad if you refuse to answer this coz it's too personal, lol~",
          "votes": 1
        },
        {
          "id": 322744,
          "postDate": "2018-05-03T15:11:02.727Z",
          "content": "<p>I am not using attributed_time</p>",
          "rawMarkdown": "I am not using attributed_time",
          "votes": 2
        },
        {
          "id": 322746,
          "postDate": "2018-05-03T15:17:11.700Z",
          "content": "<p>Upvote you and thank you!</p>",
          "rawMarkdown": "Upvote you and thank you!",
          "votes": 2
        }
      ]
    },
    {
      "id": 314150,
      "postDate": "2018-04-14T18:31:47.220Z",
      "content": "<p>Our best score 0.9797 is a single model with 9 features.  </p>",
      "rawMarkdown": "Our best score 0.9797 is a single model with 9 features.  ",
      "votes": 10,
      "replies": [
        {
          "id": 314156,
          "postDate": "2018-04-14T18:39:08.883Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 314157,
          "postDate": "2018-04-14T18:42:16.970Z",
          "content": "<p>Just 9 !!! Crazy..</p>",
          "rawMarkdown": "Just 9 !!! Crazy.."
        }
      ]
    },
    {
      "id": 310313,
      "postDate": "2018-04-07T04:28:35.440Z",
      "content": "<p>Got to 0.9711 with no parameter tuning of xgBoost using the features in <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">https://www.kaggle.com/nanomathias/feature-engineering-importance-testing</a>.</p>",
      "rawMarkdown": "Got to 0.9711 with no parameter tuning of xgBoost using the features in https://www.kaggle.com/nanomathias/feature-engineering-importance-testing.",
      "votes": 10,
      "replies": [
        {
          "id": 310888,
          "postDate": "2018-04-09T01:36:47.970Z",
          "content": "<p>Using some of these features and a few of my own, I’m at 9.770.  LGB and plenty of tuning though. </p>",
          "rawMarkdown": "Using some of these features and a few of my own, I’m at 9.770.  LGB and plenty of tuning though. ",
          "votes": 1
        },
        {
          "id": 310892,
          "postDate": "2018-04-09T01:40:29.347Z",
          "content": "<p>@Chad Gardner Great job Chad! Could I ask what kind of features your r using? I'm totally fine if you keep it secret~</p>",
          "rawMarkdown": "@Chad Gardner Great job Chad! Could I ask what kind of features your r using? I'm totally fine if you keep it secret~"
        }
      ]
    },
    {
      "id": 315303,
      "postDate": "2018-04-17T01:00:27.453Z",
      "content": "<p>LB 0.9804 with single LGB in R with 38 features. (including some features from <code>Baris' next_click kernel</code> ). </p>\n\n<p>I use only 5% data for validation. </p>",
      "rawMarkdown": "LB 0.9804 with single LGB in R with 38 features. (including some features from `Baris' next_click kernel` ). \n\nI use only 5% data for validation. ",
      "votes": 9,
      "replies": [
        {
          "id": 315459,
          "postDate": "2018-04-17T06:23:00.907Z",
          "content": "<p>Hi Pranav,\nDo you use the common day 9 hour 4 as CV? Or you generalize your cross validation and ignore the Public LB of hour 4? What is your CV AUC when training?</p>\n\n<p>Thanks!!</p>",
          "rawMarkdown": "Hi Pranav,\nDo you use the common day 9 hour 4 as CV? Or you generalize your cross validation and ignore the Public LB of hour 4? What is your CV AUC when training?\n\nThanks!!",
          "votes": 1
        },
        {
          "id": 315566,
          "postDate": "2018-04-17T09:24:58.077Z",
          "content": "<p>Nope! I use randomly shuffled split for validation while making sure that class unbalance stays exactly same in split. I deal with test hours separately. :)</p>",
          "rawMarkdown": "Nope! I use randomly shuffled split for validation while making sure that class unbalance stays exactly same in split. I deal with test hours separately. :)",
          "votes": 1
        },
        {
          "id": 315590,
          "postDate": "2018-04-17T10:22:31.223Z",
          "content": "<p>@Pranav Pandya\nThat's what I used to do but my CV AUC was always significantly higher than LB. \nDoes your CV AUC resemble in any way to the LB score?</p>\n\n<p>Thanks =) </p>",
          "rawMarkdown": "@Pranav Pandya\nThat's what I used to do but my CV AUC was always significantly higher than LB. \nDoes your CV AUC resemble in any way to the LB score?\n\nThanks =) "
        },
        {
          "id": 315603,
          "postDate": "2018-04-17T10:40:54.230Z",
          "content": "<p>Check out <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54673\">this post</a>  which will clear all the confusions about CV LB thingy thing. </p>",
          "rawMarkdown": "Check out [this post][1]  which will clear all the confusions about CV LB thingy thing. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54673",
          "votes": 2
        },
        {
          "id": 316417,
          "postDate": "2018-04-19T01:39:56.803Z",
          "content": "<p>is everyone using forward time deltas, because i tried backward time delta features, the difference between my day 9 CV and LB was huge, 0.9805 and 0.51</p>",
          "rawMarkdown": "is everyone using forward time deltas, because i tried backward time delta features, the difference between my day 9 CV and LB was huge, 0.9805 and 0.51"
        },
        {
          "id": 317870,
          "postDate": "2018-04-22T18:11:06.037Z",
          "content": "<p>Hi Pranav,</p>\n\n<p>I would like to avoid the last rows for cross validation, so I tried to randomly split my data. I'm not sure if I should still leave those data in the training set or no ? I'm afraid if I exclude them I'd remove some useful data from the training set ?</p>",
          "rawMarkdown": "Hi Pranav,\n\nI would like to avoid the last rows for cross validation, so I tried to randomly split my data. I'm not sure if I should still leave those data in the training set or no ? I'm afraid if I exclude them I'd remove some useful data from the training set ?"
        },
        {
          "id": 317916,
          "postDate": "2018-04-22T20:33:26.257Z",
          "content": "<p>Hello Antonie,</p>\n\n<p>My personal opinion is that chances of loosing highly predictive observations (close to test) is already reduced when validation split ist random shuffled. I wouldn't worry about re-running model with validation data in training set if I'm using full training data. </p>",
          "rawMarkdown": "Hello Antonie,\n\nMy personal opinion is that chances of loosing highly predictive observations (close to test) is already reduced when validation split ist random shuffled. I wouldn't worry about re-running model with validation data in training set if I'm using full training data. ",
          "votes": 1
        },
        {
          "id": 317952,
          "postDate": "2018-04-22T23:14:14.197Z",
          "content": "<p>Thanks for clarifying this. I'll do more testing by creating 2 distinct set to see if this helps.</p>",
          "rawMarkdown": "Thanks for clarifying this. I'll do more testing by creating 2 distinct set to see if this helps.",
          "votes": 1
        },
        {
          "id": 317991,
          "postDate": "2018-04-23T02:16:37.220Z",
          "content": "<p>is there any method or way to randomly split data by maintaining the class imbalance?</p>",
          "rawMarkdown": "is there any method or way to randomly split data by maintaining the class imbalance?"
        },
        {
          "id": 318000,
          "postDate": "2018-04-23T02:44:50.387Z",
          "content": "<p>I'm using a basic numpy random function to split the data into training / validation. Since we have a huge number of data, I expect the imbalance to be close in train and valid set.</p>",
          "rawMarkdown": "I'm using a basic numpy random function to split the data into training / validation. Since we have a huge number of data, I expect the imbalance to be close in train and valid set."
        },
        {
          "id": 318095,
          "postDate": "2018-04-23T07:16:29.497Z",
          "content": "<p>@nickhillator , you can use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html\">sklearn train_test_split</a> and passing target array as \"stratify\" argument you can use stratified train,test split. </p>",
          "rawMarkdown": "@nickhillator , you can use [sklearn train_test_split][1] and passing target array as \"stratify\" argument you can use stratified train,test split. \n\n\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html",
          "votes": 1
        },
        {
          "id": 318104,
          "postDate": "2018-04-23T07:33:18.397Z",
          "content": "<p>Here You go. \ntrain, test, y_train, y_test = train_test_split(train, y, test_size=0.2, random_state=42,stratify = y)</p>",
          "rawMarkdown": "Here You go. \ntrain, test, y_train, y_test = train_test_split(train, y, test_size=0.2, random_state=42,stratify = y)",
          "votes": 2
        }
      ]
    },
    {
      "id": 318986,
      "postDate": "2018-04-25T01:16:30.867Z",
      "content": "<p>update: 0.9813 with 20 features (single lgb run)</p>",
      "rawMarkdown": "update: 0.9813 with 20 features (single lgb run)",
      "votes": 7,
      "replies": [
        {
          "id": 318987,
          "postDate": "2018-04-25T01:23:08.897Z",
          "content": "<p>Cool, may I ask how much data do you use?</p>",
          "rawMarkdown": "Cool, may I ask how much data do you use?"
        },
        {
          "id": 318993,
          "postDate": "2018-04-25T01:31:22.817Z",
          "content": "<p>All.</p>",
          "rawMarkdown": "All.",
          "votes": 3
        },
        {
          "id": 319192,
          "postDate": "2018-04-25T13:35:32.143Z",
          "content": "<p>All means including all training data? What is the size of the memory you are using ?</p>",
          "rawMarkdown": "All means including all training data? What is the size of the memory you are using ?"
        },
        {
          "id": 319198,
          "postDate": "2018-04-25T13:39:38.143Z",
          "content": "<p>I'm traveling today hence must answer from memory.  </p>\n\n<p>All means all train data.</p>\n\n<p>I worked quite a bit to decrease memory consumption.  There is a spike for lgb dataset creation, but then it runs in less than 32GB, if not less 20GB.  I'll check tonight the peak and average memory use.  </p>\n\n<p>I also tried XGBoost but it requires more memory and I have not run it on full data.  </p>",
          "rawMarkdown": "I'm traveling today hence must answer from memory.  \n\nAll means all train data.\n\nI worked quite a bit to decrease memory consumption.  There is a spike for lgb dataset creation, but then it runs in less than 32GB, if not less 20GB.  I'll check tonight the peak and average memory use.  \n\nI also tried XGBoost but it requires more memory and I have not run it on full data.  ",
          "votes": 1
        },
        {
          "id": 319218,
          "postDate": "2018-04-25T14:29:57.927Z",
          "content": "<p>I understand. Currently I am confined to a 24GB and that's the reason I was asking about memory. I am also looking to ways to improve memory management so that I can fit both more rows from train data and more features. Thank you for the details.</p>",
          "rawMarkdown": "I understand. Currently I am confined to a 24GB and that's the reason I was asking about memory. I am also looking to ways to improve memory management so that I can fit both more rows from train data and more features. Thank you for the details."
        }
      ]
    },
    {
      "id": 311093,
      "postDate": "2018-04-09T11:14:18.207Z",
      "content": "<p>I can confirm, that 0.9786 is also reachable by single model single run. </p>\n\n<p>Would be interesting to hear from Top-10 if 0.98+ is also reachable by single model?</p>",
      "rawMarkdown": "I can confirm, that 0.9786 is also reachable by single model single run. \n\nWould be interesting to hear from Top-10 if 0.98+ is also reachable by single model?",
      "votes": 8,
      "replies": [
        {
          "id": 311094,
          "postDate": "2018-04-09T11:18:46.783Z",
          "content": "<p>My single model scores 0.9810. I think people above me also have single models.</p>",
          "rawMarkdown": "My single model scores 0.9810. I think people above me also have single models.",
          "votes": 12
        },
        {
          "id": 311111,
          "postDate": "2018-04-09T12:15:15.247Z",
          "content": "<p>How do you define single model?  Are you averaging several runs for instance?  Even if you do, then this is rather impressive.</p>",
          "rawMarkdown": "How do you define single model?  Are you averaging several runs for instance?  Even if you do, then this is rather impressive.",
          "votes": -3
        },
        {
          "id": 311117,
          "postDate": "2018-04-09T12:31:25.090Z",
          "content": "<p>No averaging = 0.9810.\nAverage of two runs: 0.9811. \nSince I use LGB and train using all the data, I don't get as much as benefit that I used to get ensembling different neural net runs in other competitions. My current model has very small variation.</p>",
          "rawMarkdown": "No averaging = 0.9810.\nAverage of two runs: 0.9811. \nSince I use LGB and train using all the data, I don't get as much as benefit that I used to get ensembling different neural net runs in other competitions. My current model has very small variation.",
          "votes": 8
        },
        {
          "id": 311190,
          "postDate": "2018-04-09T15:13:41.657Z",
          "content": "<p>@AhmetErdem Did you use test_supplement to calculate features for the model scoring 0.9810?</p>",
          "rawMarkdown": "@AhmetErdem Did you use test_supplement to calculate features for the model scoring 0.9810?"
        },
        {
          "id": 311453,
          "postDate": "2018-04-10T05:10:27.940Z",
          "content": "<p>I think this supplement is not necessary at all. I'm nearby 0.98 (0.9781) and I am just using IP frequencies. Besides, I noticed that some unrelevant features can destroy the prediction. I advise to use a good training set and smart features that are using the IP (it is the most important variable with app). I will add other features tomorrow and tell what's going on on thursday!</p>",
          "rawMarkdown": "I think this supplement is not necessary at all. I'm nearby 0.98 (0.9781) and I am just using IP frequencies. Besides, I noticed that some unrelevant features can destroy the prediction. I advise to use a good training set and smart features that are using the IP (it is the most important variable with app). I will add other features tomorrow and tell what's going on on thursday!",
          "votes": 4
        },
        {
          "id": 311797,
          "postDate": "2018-04-10T17:26:20.947Z",
          "content": "<p>@Badr - Did you use any target encoding at all?</p>",
          "rawMarkdown": "@Badr - Did you use any target encoding at all?"
        },
        {
          "id": 311830,
          "postDate": "2018-04-10T18:32:06.940Z",
          "content": "<p>@Badr IP frequencies are mostly overfitting !!! 0.9781 LB with overfit? :/</p>",
          "rawMarkdown": "@Badr IP frequencies are mostly overfitting !!! 0.9781 LB with overfit? :/"
        },
        {
          "id": 311838,
          "postDate": "2018-04-10T18:48:48.070Z",
          "content": "<p>@Rohit IP frequencies that I use can not be done with regular Python. I am using Vertica to compute them. By the way the variables that I use are far from overfitting and using small trees, I think it is hard to overfitt with the parameters proposed by the LGBM. You have to think rationally about the situation, what fraudulent machines are doing and that's all.</p>\n\n<p>@Shanth, no not at all. I tried and the score decreased considerably because of overfitting this time.</p>",
          "rawMarkdown": "@Rohit IP frequencies that I use can not be done with regular Python. I am using Vertica to compute them. By the way the variables that I use are far from overfitting and using small trees, I think it is hard to overfitt with the parameters proposed by the LGBM. You have to think rationally about the situation, what fraudulent machines are doing and that's all.\n\n@Shanth, no not at all. I tried and the score decreased considerably because of overfitting this time.",
          "votes": 3
        },
        {
          "id": 311841,
          "postDate": "2018-04-10T18:53:44.850Z",
          "content": "<p>@Rohit @Badr I think you just mean two different things by frequency - % attributed (in the past) vs. counts. </p>\n\n<p>@Badr what do you mean by frequencies that can't be done with python? Like a customized computation that base pandas doesn't support?</p>",
          "rawMarkdown": "@Rohit @Badr I think you just mean two different things by frequency - % attributed (in the past) vs. counts. \n\n@Badr what do you mean by frequencies that can't be done with python? Like a customized computation that base pandas doesn't support?",
          "votes": 1
        },
        {
          "id": 311859,
          "postDate": "2018-04-10T19:51:19.073Z",
          "content": "<p>@Badr Thanks for the reply. Also the target encoding was for App/ App + device combination ? Obviously target encoding for IP will be a bad idea. </p>",
          "rawMarkdown": "@Badr Thanks for the reply. Also the target encoding was for App/ App + device combination ? Obviously target encoding for IP will be a bad idea. "
        },
        {
          "id": 311867,
          "postDate": "2018-04-10T20:02:44.360Z",
          "content": "<p>@Joe you are right! For the question, I think it will be hard to do it with pandas. It is still possible but you need a really really powerful machine and a lot of patience I think!</p>\n\n<p>@Shanth, you welcome. I really like the Kaggle community, you are all very helpful. I think that the LGBM will manage to do it, if you increase the depth (risk of overfitting). I tried to use some combinations but my results were not really good. Sometimes it increases the AUC of 0.0001 which is not relevant at all. I think smart variables are enough to win the competition (+ the 4 categorical ones), I hope I will find some others during the week.</p>",
          "rawMarkdown": "@Joe you are right! For the question, I think it will be hard to do it with pandas. It is still possible but you need a really really powerful machine and a lot of patience I think!\n\n@Shanth, you welcome. I really like the Kaggle community, you are all very helpful. I think that the LGBM will manage to do it, if you increase the depth (risk of overfitting). I tried to use some combinations but my results were not really good. Sometimes it increases the AUC of 0.0001 which is not relevant at all. I think smart variables are enough to win the competition (+ the 4 categorical ones), I hope I will find some others during the week.",
          "votes": 1
        },
        {
          "id": 311884,
          "postDate": "2018-04-10T20:54:34.033Z",
          "content": "<p>Curious exactly what you mean, but I think I have a pretty good guess. If my guess is right, I agree that it is hard, but definitely doable in python ;) (even doable in pandas but with much slowness).</p>",
          "rawMarkdown": "Curious exactly what you mean, but I think I have a pretty good guess. If my guess is right, I agree that it is hard, but definitely doable in python ;) (even doable in pandas but with much slowness).",
          "votes": 1
        },
        {
          "id": 312499,
          "postDate": "2018-04-11T21:28:04.227Z",
          "content": "<p>Hi, I am curious that when you train on the all data, how do you define the validation set to determine the best iteration number? </p>",
          "rawMarkdown": "Hi, I am curious that when you train on the all data, how do you define the validation set to determine the best iteration number? ",
          "votes": 1
        },
        {
          "id": 312503,
          "postDate": "2018-04-11T21:40:23.983Z",
          "content": "<p>To determine the best iteration you need a validation set, so you don't literally train on all data. After you learn what the best iteration number is you can repeat the training on the full set stopping at the best iteration. \nAlternatively, you can do cross-validation with 5 or 10 folds.</p>",
          "rawMarkdown": "To determine the best iteration you need a validation set, so you don't literally train on all data. After you learn what the best iteration number is you can repeat the training on the full set stopping at the best iteration. \nAlternatively, you can do cross-validation with 5 or 10 folds."
        },
        {
          "id": 312520,
          "postDate": "2018-04-11T22:29:55.070Z",
          "content": "<p>How are you planning to do a kfold without affecting the time series nature of data set?</p>",
          "rawMarkdown": "How are you planning to do a kfold without affecting the time series nature of data set?"
        },
        {
          "id": 312528,
          "postDate": "2018-04-11T23:36:08.837Z",
          "content": "<p>@Alexey Thanks for your telling that. What I mean is that when all data are used to train, the ways of determining validation set may be kFold or randomly selecting some data. But the data set has time series nature.</p>",
          "rawMarkdown": "@Alexey Thanks for your telling that. What I mean is that when all data are used to train, the ways of determining validation set may be kFold or randomly selecting some data. But the data set has time series nature."
        },
        {
          "id": 312579,
          "postDate": "2018-04-12T02:02:43.743Z",
          "content": "<p>@Longchun Tao and shivraj <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325\">Here</a> there is a nice discussion about validation. You might want to read it.</p>",
          "rawMarkdown": "@Longchun Tao and shivraj [Here][1] there is a nice discussion about validation. You might want to read it.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325"
        },
        {
          "id": 317623,
          "postDate": "2018-04-22T02:43:39.037Z",
          "content": "<p>@Joe @Badr @Shanth May I ask what you guys mean by target encoding in this case? Also, is this what I did related to it: I used ip + app + device as a group and calculated the % attributed for each of those groups and merged them into test set. My training acc was ~.999 and my valid. score was much lower </p>",
          "rawMarkdown": "@Joe @Badr @Shanth May I ask what you guys mean by target encoding in this case? Also, is this what I did related to it: I used ip + app + device as a group and calculated the % attributed for each of those groups and merged them into test set. My training acc was ~.999 and my valid. score was much lower "
        },
        {
          "id": 318370,
          "postDate": "2018-04-23T17:27:08.590Z",
          "content": "<p>@ Edward Chen -  Yes that is target encoding. My current model does not include it since time delta/ unique counts etc proved to be more useful. It may be useful if one is using the entire data set ( which I am not doing for lack of resources). </p>\n\n<p>As for training accuracy - same here. I got an unusually high training accuracy but low validation score.  I also used a variation of Target encoding from NanoMathias's awesome script  <a>here</a>. This script scales the target encoding with a confidence score. I believe it helped his LB score but again he used the entire data set for training. </p>\n\n<p>Hope this helps \nRegards\nShanth</p>",
          "rawMarkdown": "@ Edward Chen -  Yes that is target encoding. My current model does not include it since time delta/ unique counts etc proved to be more useful. It may be useful if one is using the entire data set ( which I am not doing for lack of resources). \n\nAs for training accuracy - same here. I got an unusually high training accuracy but low validation score.  I also used a variation of Target encoding from NanoMathias's awesome script  [here][1]. This script scales the target encoding with a confidence score. I believe it helped his LB score but again he used the entire data set for training. \n\nHope this helps \nRegards\nShanth\n\n\n  [1]: http://%20https://www.kaggle.com/nanomathias/feature-engineering-importance-testing."
        }
      ]
    },
    {
      "id": 321414,
      "postDate": "2018-05-01T08:15:23.667Z",
      "content": "<p>This is an edit of my previous post.</p>\n\n<p>0.9808 with 13 features (lgbm), </p>\n\n<p>0.9811 with 25 features (lgbm, 110M raws), </p>\n\n<p>0.9813 with 28 features (lgbm, 110M raws)</p>\n\n<p>I iteratively make one greedy choice after another (group of a few features) and remove it if it ends up not improving my local auc score.</p>",
      "rawMarkdown": "This is an edit of my previous post.\n\n0.9808 with 13 features (lgbm), \n\n0.9811 with 25 features (lgbm, 110M raws), \n\n0.9813 with 28 features (lgbm, 110M raws)\n\nI iteratively make one greedy choice after another (group of a few features) and remove it if it ends up not improving my local auc score.",
      "votes": 5,
      "replies": [
        {
          "id": 321514,
          "postDate": "2018-05-01T13:19:14.493Z",
          "content": "<p>Wow, that's amazing! </p>",
          "rawMarkdown": "Wow, that's amazing! "
        },
        {
          "id": 322157,
          "postDate": "2018-05-02T13:36:14.067Z",
          "content": "<p>great work phillipe</p>",
          "rawMarkdown": "great work phillipe"
        }
      ]
    },
    {
      "id": 319439,
      "postDate": "2018-04-26T03:28:54.853Z",
      "content": "<p>0.9806 with lgbm, 18features</p>",
      "rawMarkdown": "0.9806 with lgbm, 18features",
      "votes": 5,
      "replies": [
        {
          "id": 319868,
          "postDate": "2018-04-27T02:08:11.850Z",
          "content": "<p>with all data?</p>",
          "rawMarkdown": "with all data?",
          "votes": 1
        },
        {
          "id": 320057,
          "postDate": "2018-04-27T11:21:12.113Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes"
        },
        {
          "id": 320098,
          "postDate": "2018-04-27T13:13:36.663Z",
          "content": "<p>How r u calculating time deltas with such a huge data set?</p>",
          "rawMarkdown": "How r u calculating time deltas with such a huge data set?"
        },
        {
          "id": 320243,
          "postDate": "2018-04-28T01:49:25.170Z",
          "content": "<p>df.groupby(cols).clicktime.shift - df.clicktime</p>",
          "rawMarkdown": "df.groupby(cols).clicktime.shift - df.clicktime",
          "votes": 4
        },
        {
          "id": 320817,
          "postDate": "2018-04-30T00:30:24.847Z",
          "content": "<p>cool</p>",
          "rawMarkdown": "cool"
        }
      ]
    },
    {
      "id": 314319,
      "postDate": "2018-04-15T08:26:29.660Z",
      "content": "<p>0.9791 with a single lightgbm on 25 features. Looking at other posts, I should be able to shrink that amount :-)</p>",
      "rawMarkdown": "0.9791 with a single lightgbm on 25 features. Looking at other posts, I should be able to shrink that amount :-)",
      "votes": 5,
      "replies": [
        {
          "id": 315002,
          "postDate": "2018-04-16T16:14:29.770Z",
          "content": "<p>I have also managed to get similar score(0.9792) using 20 features with single lightgmb, but  I am still unsuccessful in shrinking features and get the same or better score. I guess I still need to engineer some strong features.</p>",
          "rawMarkdown": "I have also managed to get similar score(0.9792) using 20 features with single lightgmb, but  I am still unsuccessful in shrinking features and get the same or better score. I guess I still need to engineer some strong features."
        },
        {
          "id": 315011,
          "postDate": "2018-04-16T16:28:56.107Z",
          "content": "<p>@ Shoaib...Is this Day 8 training only or all data?</p>",
          "rawMarkdown": "@ Shoaib...Is this Day 8 training only or all data?"
        },
        {
          "id": 315031,
          "postDate": "2018-04-16T16:53:10.847Z",
          "content": "<p>I train only on day 7,8 and validate on day 9.</p>",
          "rawMarkdown": "I train only on day 7,8 and validate on day 9."
        }
      ]
    },
    {
      "id": 310606,
      "postDate": "2018-04-08T04:51:15.410Z",
      "content": "<p>Almost 30 features and a score of 0.9762 The training day is really really important to reach this score. I'll try to improve in the next days.</p>",
      "rawMarkdown": "Almost 30 features and a score of 0.9762 The training day is really really important to reach this score. I'll try to improve in the next days.",
      "votes": 5,
      "replies": [
        {
          "id": 310976,
          "postDate": "2018-04-09T06:17:30.867Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 319372,
      "postDate": "2018-04-25T22:29:48.633Z",
      "content": "<p>It's my first competition ! For now I have 0.9806 with lightgbm, I'm using less than half of training data. If I have time / resources I'll try to run it on the full data set. I have around 20 features which makes my computation very slow.</p>",
      "rawMarkdown": "It's my first competition ! For now I have 0.9806 with lightgbm, I'm using less than half of training data. If I have time / resources I'll try to run it on the full data set. I have around 20 features which makes my computation very slow.",
      "votes": 6,
      "replies": [
        {
          "id": 319533,
          "postDate": "2018-04-26T08:39:58.673Z",
          "content": "<p>That's very promising, esp for a first competition.</p>",
          "rawMarkdown": "That's very promising, esp for a first competition.",
          "votes": 1
        },
        {
          "id": 319655,
          "postDate": "2018-04-26T14:26:44.593Z",
          "content": "<p>Thanks :)</p>",
          "rawMarkdown": "Thanks :)"
        }
      ]
    },
    {
      "id": 314701,
      "postDate": "2018-04-16T06:09:23.830Z",
      "content": "<p>LB 0.9769 with features from <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">here</a> and xgBoost parameters from <a href=\"https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769\">here</a></p>",
      "rawMarkdown": "LB 0.9769 with features from [here](https://www.kaggle.com/nanomathias/feature-engineering-importance-testing) and xgBoost parameters from [here](https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769)",
      "votes": 6,
      "replies": [
        {
          "id": 314791,
          "postDate": "2018-04-16T09:26:27.790Z",
          "content": "<p>Awesome!  What was your LB before Bayesian optimization?</p>",
          "rawMarkdown": "Awesome!  What was your LB before Bayesian optimization?"
        },
        {
          "id": 314795,
          "postDate": "2018-04-16T09:37:52.897Z",
          "content": "<p>I was at 0.9711 prior to the optimization :)</p>",
          "rawMarkdown": "I was at 0.9711 prior to the optimization :)",
          "votes": 1
        },
        {
          "id": 314799,
          "postDate": "2018-04-16T09:43:23.100Z",
          "content": "<p>I guess you just triggered LOTS of interest for Bayesian optimization ;)  Thanks for sharing.</p>",
          "rawMarkdown": "I guess you just triggered LOTS of interest for Bayesian optimization ;)  Thanks for sharing.",
          "votes": 1
        },
        {
          "id": 314839,
          "postDate": "2018-04-16T10:56:00.077Z",
          "content": "<p>Thanks, hopefully it'll help people get better scores :) .. especially now that it's so easy to do with the scikit-optimize package..</p>",
          "rawMarkdown": "Thanks, hopefully it'll help people get better scores :) .. especially now that it's so easy to do with the scikit-optimize package..",
          "votes": 1
        },
        {
          "id": 314850,
          "postDate": "2018-04-16T11:16:17.113Z",
          "content": "<p>May I ask why do we to find optimal parameters using kfold or stratified kfold when most of us are using different time based CV setting? </p>",
          "rawMarkdown": "May I ask why do we to find optimal parameters using kfold or stratified kfold when most of us are using different time based CV setting? "
        },
        {
          "id": 314855,
          "postDate": "2018-04-16T11:23:42.373Z",
          "content": "<p>Using StratifiedKfold() is just my generic first-approach, which turned out to work OK (in terms of LB score being pretty similar to my local CV score)..</p>\n\n<p>However, adding in another cross-validator should be easy, e.g. using <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html#sklearn.model_selection.TimeSeriesSplit\">TimeSeriesSplit</a>. I guess it'd be worth testing that out as an alternative to the StratifiedKFold() one. </p>\n\n<p>Right now I'm back to working on improving the feature engineering though, and I think afterwards computing power could potentially be better spent (for me at least) looking into feature selection</p>",
          "rawMarkdown": "Using StratifiedKfold() is just my generic first-approach, which turned out to work OK (in terms of LB score being pretty similar to my local CV score)..\n\nHowever, adding in another cross-validator should be easy, e.g. using [TimeSeriesSplit](http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html#sklearn.model_selection.TimeSeriesSplit). I guess it'd be worth testing that out as an alternative to the StratifiedKFold() one. \n\nRight now I'm back to working on improving the feature engineering though, and I think afterwards computing power could potentially be better spent (for me at least) looking into feature selection",
          "votes": 2
        },
        {
          "id": 314868,
          "postDate": "2018-04-16T11:56:41.117Z",
          "content": "<p>I will try this using TimeSeriesSplit and report back in few days.</p>",
          "rawMarkdown": "I will try this using TimeSeriesSplit and report back in few days."
        },
        {
          "id": 314870,
          "postDate": "2018-04-16T11:59:44.707Z",
          "content": "<p>Cool, looking forward to hearing if results are better with that approach</p>",
          "rawMarkdown": "Cool, looking forward to hearing if results are better with that approach"
        }
      ]
    },
    {
      "id": 308079,
      "postDate": "2018-04-02T23:13:21.763Z",
      "content": "<p>Currently, best single run at 0.9736</p>\n\n<p>Others have said they had single model at 0.9764</p>",
      "rawMarkdown": "Currently, best single run at 0.9736\n\nOthers have said they had single model at 0.9764",
      "votes": 5,
      "replies": [
        {
          "id": 308085,
          "postDate": "2018-04-02T23:23:26.863Z",
          "content": "<p>@CPMP, lead us to 0.98, go go go!</p>",
          "rawMarkdown": "@CPMP, lead us to 0.98, go go go!"
        },
        {
          "id": 308092,
          "postDate": "2018-04-02T23:41:57.240Z",
          "content": "<p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>",
          "rawMarkdown": "This whole .9764 with 4 features thing is going to haunt me for weeks",
          "votes": 4
        },
        {
          "id": 308096,
          "postDate": "2018-04-02T23:46:03.197Z",
          "content": "<p>A spectre is haunting TalkingData—the spectre of model with only 4 engineered features.</p>",
          "rawMarkdown": "A spectre is haunting TalkingData—the spectre of model with only 4 engineered features.",
          "votes": 1
        },
        {
          "id": 308099,
          "postDate": "2018-04-02T23:50:55.413Z",
          "content": "<blockquote>\n  <p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n</blockquote>\n\n<p>There are things that haunt me, but none of them have \"Kaggle challenge\" in their description. There will always be things that are obvious to others but not me, and <em>vice versa</em>.</p>",
          "rawMarkdown": "&gt; This whole .9764 with 4 features thing is going to haunt me for weeks\n\nThere are things that haunt me, but none of them have \"Kaggle challenge\" in their description. There will always be things that are obvious to others but not me, and *vice versa*.",
          "votes": 4
        },
        {
          "id": 308102,
          "postDate": "2018-04-02T23:54:09.423Z",
          "content": "<blockquote>\n  <p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n</blockquote>\n\n<p>Heck, my single best model is at 0.9617, and even that doesn't haunt me.</p>",
          "rawMarkdown": "&gt; This whole .9764 with 4 features thing is going to haunt me for weeks\n\nHeck, my single best model is at 0.9617, and even that doesn't haunt me.",
          "votes": 3
        },
        {
          "id": 308104,
          "postDate": "2018-04-02T23:56:30.447Z",
          "content": "<blockquote>\n  <p><strong>Tilii wrote</strong></p>\n  \n  <blockquote>\n    <p>&gt; This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n  </blockquote>\n  \n  <p>There are things that haunt me, but none of them have \"Kaggle challenge\" in their description. There will always be things that are obvious to others but not me, and <em>vice versa</em>.</p>\n</blockquote>\n\n<p>Yeah I mean it jokingly :) In truth it's a motivation to explore the data more and keep digging for features, not a negative.</p>",
          "rawMarkdown": "&gt; **Tilii wrote**\n&gt; \n&gt; &gt; &gt; This whole .9764 with 4 features thing is going to haunt me for weeks\n&gt; \n&gt; There are things that haunt me, but none of them have \"Kaggle challenge\" in their description. There will always be things that are obvious to others but not me, and *vice versa*.\n\nYeah I mean it jokingly :) In truth it's a motivation to explore the data more and keep digging for features, not a negative.",
          "votes": 1
        },
        {
          "id": 308176,
          "postDate": "2018-04-03T04:36:43.417Z",
          "content": "<blockquote>\n  <p><strong>Joe Eddy wrote</strong></p>\n  \n  <blockquote>\n    <p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n  </blockquote>\n</blockquote>\n\n<p>Becarefull!! XD the more you want to find it the more it will go away!</p>",
          "rawMarkdown": "\n&gt; **Joe Eddy wrote**\n&gt; \n&gt; &gt; This whole .9764 with 4 features thing is going to haunt me for weeks\n\nBecarefull!! XD the more you want to find it the more it will go away!",
          "votes": 2
        },
        {
          "id": 309007,
          "postDate": "2018-04-04T13:34:37.977Z",
          "content": "<blockquote>\n  <p>@CPMP, lead us to 0.98, go go go!</p>\n</blockquote>\n\n<p>Thanks for your trust, but it won't happen anytime soon... Found a flaw in my data preparation, I don't expect submission before few days now.</p>",
          "rawMarkdown": "&gt; @CPMP, lead us to 0.98, go go go!\n\nThanks for your trust, but it won't happen anytime soon... Found a flaw in my data preparation, I don't expect submission before few days now.",
          "votes": 1
        },
        {
          "id": 312349,
          "postDate": "2018-04-11T15:48:02.973Z",
          "content": "<p>@CPMP I'm really impressed by your model. Do you have any suggestion to improve my score using single models?\nI mean maybe not for this competition but how do you know which model is going to work, what do you do to tune your models or if you have some strategy for feature engineering? How do you learn all of this?\nMaybe it's a lot to ask haha but I will take my chances.</p>",
          "rawMarkdown": "@CPMP I'm really impressed by your model. Do you have any suggestion to improve my score using single models?\nI mean maybe not for this competition but how do you know which model is going to work, what do you do to tune your models or if you have some strategy for feature engineering? How do you learn all of this?\nMaybe it's a lot to ask haha but I will take my chances."
        },
        {
          "id": 312394,
          "postDate": "2018-04-11T17:22:37.497Z",
          "content": "<p>Hi, there are way better scores than mine, you should ask one of these guys above 0.98 ;)</p>\n\n<p>I learned in two ways: </p>\n\n<ol>\n<li><p>I followed my own advice to newbies, posted elsewhere on this forum.</p></li>\n<li><p>I read carefully what top performers share after each competition ends. </p></li>\n</ol>",
          "rawMarkdown": "Hi, there are way better scores than mine, you should ask one of these guys above 0.98 ;)\n\nI learned in two ways: \n\n 1. I followed my own advice to newbies, posted elsewhere on this forum.\n\n 2.  I read carefully what top performers share after each competition ends. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 321908,
      "postDate": "2018-05-02T06:10:04.410Z",
      "content": "<blockquote>\n  <p>Below .98xx app had 10x more importance over the channel. Above .98xx\n  channel has 3x more importance. I am unable to understand this</p>\n</blockquote>\n\n<p>I think this is because you are adding more features that has similar information as 'app'.</p>",
      "rawMarkdown": "&gt; Below .98xx app had 10x more importance over the channel. Above .98xx\n&gt; channel has 3x more importance. I am unable to understand this\n\nI think this is because you are adding more features that has similar information as 'app'.",
      "votes": 3
    },
    {
      "id": 320910,
      "postDate": "2018-04-30T07:00:06.867Z",
      "content": "<p>0.9809 with all data, 7 new features and no target encoding - lgbm.</p>",
      "rawMarkdown": "0.9809 with all data, 7 new features and no target encoding - lgbm.",
      "votes": 3,
      "replies": [
        {
          "id": 320947,
          "postDate": "2018-04-30T08:58:21.810Z",
          "content": "<p>Care to share the validation strategy you used?</p>",
          "rawMarkdown": "Care to share the validation strategy you used?"
        },
        {
          "id": 320974,
          "postDate": "2018-04-30T09:55:11.127Z",
          "content": "<p>I validated with day 9, a learn rate of 0.05 and early stopping rounds at 50. The features I am using are all time delta (time since x) and click counts (by different grouping variables). Hope that helps a bit and good luck 🙂</p>",
          "rawMarkdown": "I validated with day 9, a learn rate of 0.05 and early stopping rounds at 50. The features I am using are all time delta (time since x) and click counts (by different grouping variables). Hope that helps a bit and good luck 🙂",
          "votes": 8
        },
        {
          "id": 321020,
          "postDate": "2018-04-30T12:18:18.940Z",
          "content": "<p>Sorry to pry -- time since x, or time to x? It seems most have given up on the former entirely, so would be pretty interesting if thy were the basis to your secret sauce. Also, impressive score, very motivating!</p>",
          "rawMarkdown": "Sorry to pry -- time since x, or time to x? It seems most have given up on the former entirely, so would be pretty interesting if thy were the basis to your secret sauce. Also, impressive score, very motivating!"
        },
        {
          "id": 321028,
          "postDate": "2018-04-30T12:38:37.017Z",
          "content": "<p>Time since x  but with different grouping variables to many I have seen in the kernels on here. I have also found that time lagged features have a high feature importance in the lgbm model, e.g. clicks in the last hour/day although these take a long time to process in Pandas! </p>",
          "rawMarkdown": "Time since x  but with different grouping variables to many I have seen in the kernels on here. I have also found that time lagged features have a high feature importance in the lgbm model, e.g. clicks in the last hour/day although these take a long time to process in Pandas! ",
          "votes": 4
        },
        {
          "id": 321030,
          "postDate": "2018-04-30T12:41:18.443Z",
          "content": "<p>🙏</p>",
          "rawMarkdown": "🙏"
        },
        {
          "id": 321057,
          "postDate": "2018-04-30T13:49:53.400Z",
          "content": "<p>Hi @George Carmichael, good job and thank you very much for your hints! Do you find features such as clicks in the last hour/day  from kernel? I think I've noticed that some kernel has this kind of feature but forget which one.</p>",
          "rawMarkdown": "Hi @George Carmichael, good job and thank you very much for your hints! Do you find features such as clicks in the last hour/day  from kernel? I think I've noticed that some kernel has this kind of feature but forget which one.",
          "votes": 1
        },
        {
          "id": 321253,
          "postDate": "2018-04-30T22:05:50.337Z",
          "content": "<p>Hi, no problem! I haven't seen these features in kernels but pandas group by and rolling windows (with click time as index and window as '1h' ) do the job without too much complexity. </p>",
          "rawMarkdown": "Hi, no problem! I haven't seen these features in kernels but pandas group by and rolling windows (with click time as index and window as '1h' ) do the job without too much complexity. ",
          "votes": 3
        },
        {
          "id": 321314,
          "postDate": "2018-05-01T02:44:47.550Z",
          "content": "<p>Congrats George ! Only 7 features is amazing... </p>\n\n<p>I'm surprised with your feature :\"Time since x\". I tried so many different way to include the time since last click with different grouping, but was never able to improve my models with those feature.</p>",
          "rawMarkdown": "Congrats George ! Only 7 features is amazing... \n\nI'm surprised with your feature :\"Time since x\". I tried so many different way to include the time since last click with different grouping, but was never able to improve my models with those feature."
        },
        {
          "id": 321358,
          "postDate": "2018-05-01T04:59:02.157Z",
          "content": "<p><a href=\"/authman\">@authman</a> did the features end up working for you?</p>",
          "rawMarkdown": "@authman did the features end up working for you?"
        },
        {
          "id": 321407,
          "postDate": "2018-05-01T07:42:45.263Z",
          "content": "<p>@George,  this is great, but we all count the number of features we use, not just the engineered ones.  Even if I include all original features that would mean 13 features for 0.9809, which is less than what I have seen so far.  The closest is Philippe Lonjoux with 0.9808 for 13 features, see below.</p>\n\n<p>Congrats for being so selective, and for sharing what works for you.</p>",
          "rawMarkdown": "@George,  this is great, but we all count the number of features we use, not just the engineered ones.  Even if I include all original features that would mean 13 features for 0.9809, which is less than what I have seen so far.  The closest is Philippe Lonjoux with 0.9808 for 13 features, see below.\n\nCongrats for being so selective, and for sharing what works for you.",
          "votes": 1
        },
        {
          "id": 321491,
          "postDate": "2018-05-01T12:15:45.797Z",
          "content": "<p><a href=\"/areveillon\">@areveillon</a> not yet.</p>",
          "rawMarkdown": "@areveillon not yet."
        },
        {
          "id": 321516,
          "postDate": "2018-05-01T13:21:16.860Z",
          "content": "<p>@CPMP thanks, well done on your great score! Yes I have 13 features in total due to dropping features with little importance after validation. It is likely I will add a couple more aimed specifically at dealing with duplicates.</p>",
          "rawMarkdown": "@CPMP thanks, well done on your great score! Yes I have 13 features in total due to dropping features with little importance after validation. It is likely I will add a couple more aimed specifically at dealing with duplicates.",
          "votes": 2
        },
        {
          "id": 321811,
          "postDate": "2018-05-02T00:09:57.947Z",
          "content": "<p>@George Congrats on your amazing score! You mentioned that time lag features had high feature importance from your runs, but you didn't mention using them in your model. I'm just curious why</p>",
          "rawMarkdown": "@George Congrats on your amazing score! You mentioned that time lag features had high feature importance from your runs, but you didn't mention using them in your model. I'm just curious why"
        },
        {
          "id": 321915,
          "postDate": "2018-05-02T06:21:53.767Z",
          "content": "<p>Thank you! I removed the time lag features such as time since last attribution as they seemed to be causing overfitting  of the model. I may add them back in but currently trying to keep the model as simple as possible :)</p>",
          "rawMarkdown": "Thank you! I removed the time lag features such as time since last attribution as they seemed to be causing overfitting  of the model. I may add them back in but currently trying to keep the model as simple as possible :)"
        },
        {
          "id": 322125,
          "postDate": "2018-05-02T12:51:11.617Z",
          "content": "<p>Thanks for your reply! What about the time lag features such as num clicks in last hour/day that you mentioned? Were they overfitting as well?</p>",
          "rawMarkdown": "Thanks for your reply! What about the time lag features such as num clicks in last hour/day that you mentioned? Were they overfitting as well?"
        },
        {
          "id": 322132,
          "postDate": "2018-05-02T13:01:55.917Z",
          "content": "<p>Thank you George. Regarding your hints about lagged features, the test set only contains several hours in discontinuous blocks. Did you use the full test set supplement? </p>",
          "rawMarkdown": "Thank you George. Regarding your hints about lagged features, the test set only contains several hours in discontinuous blocks. Did you use the full test set supplement? "
        }
      ]
    },
    {
      "id": 320816,
      "postDate": "2018-04-30T00:29:36.953Z",
      "content": "<p>9820 lgbm</p>",
      "rawMarkdown": "9820 lgbm",
      "votes": 3
    },
    {
      "id": 319384,
      "postDate": "2018-04-25T23:45:50.007Z",
      "content": "<p>0.9811 with 35 features</p>",
      "rawMarkdown": "0.9811 with 35 features",
      "votes": 3,
      "replies": [
        {
          "id": 319386,
          "postDate": "2018-04-25T23:56:06.233Z",
          "content": "<p>did you use full training set as well ?</p>",
          "rawMarkdown": "did you use full training set as well ?"
        },
        {
          "id": 319388,
          "postDate": "2018-04-26T00:20:34.493Z",
          "content": "<p>Yes</p>",
          "rawMarkdown": "Yes"
        },
        {
          "id": 319989,
          "postDate": "2018-04-27T08:45:17.593Z",
          "content": "<p>Hello AmirH.How much improvement will be achieved using all data compared to using one day data?</p>",
          "rawMarkdown": "Hello AmirH.How much improvement will be achieved using all data compared to using one day data?\n"
        },
        {
          "id": 320052,
          "postDate": "2018-04-27T11:10:07.870Z",
          "content": "<p>Depends on your features of course, but i experienced quite a significant improvement</p>",
          "rawMarkdown": "Depends on your features of course, but i experienced quite a significant improvement"
        },
        {
          "id": 320069,
          "postDate": "2018-04-27T11:57:19.873Z",
          "content": "<p>Thanks.About how much?</p>",
          "rawMarkdown": "Thanks.About how much?"
        },
        {
          "id": 320275,
          "postDate": "2018-04-28T05:27:48.970Z",
          "content": "<p>Hi! what train\\valid strategy do you use? Thanks!</p>",
          "rawMarkdown": "Hi! what train\\valid strategy do you use? Thanks!"
        },
        {
          "id": 320326,
          "postDate": "2018-04-28T10:24:07.910Z",
          "content": "<p>Hi @AlexTru,\nI use days 7 + 8  as training and day 9 for Validation</p>",
          "rawMarkdown": "Hi @AlexTru,\nI use days 7 + 8  as training and day 9 for Validation"
        },
        {
          "id": 320422,
          "postDate": "2018-04-28T16:16:36.750Z",
          "content": "<p>Hey @AmirH, did you find that your validation score for day 9 and public lb score with all training data were similar? Or to what extent was the lb score higher/lower?</p>",
          "rawMarkdown": "Hey @AmirH, did you find that your validation score for day 9 and public lb score with all training data were similar? Or to what extent was the lb score higher/lower?"
        },
        {
          "id": 320438,
          "postDate": "2018-04-28T17:04:10.277Z",
          "content": "<p>@Edward Chen\nAbout 0.003 difference, the difference used to be bigger but using different features raised my score and also decreased the different between LB and CV AUC.</p>",
          "rawMarkdown": "@Edward Chen\nAbout 0.003 difference, the difference used to be bigger but using different features raised my score and also decreased the different between LB and CV AUC.",
          "votes": 1
        },
        {
          "id": 320442,
          "postDate": "2018-04-28T17:27:08.397Z",
          "content": "<p>@AmirH Thanks for the response! So you're saying the LB was .003 higher than CV AUC right? </p>",
          "rawMarkdown": "@AmirH Thanks for the response! So you're saying the LB was .003 higher than CV AUC right? "
        },
        {
          "id": 320471,
          "postDate": "2018-04-28T18:56:59.240Z",
          "content": "<p>@Edward Chen\nNo, the other way around, the CV AUC tends to be about 0.003 higher than the LB</p>",
          "rawMarkdown": "@Edward Chen\nNo, the other way around, the CV AUC tends to be about 0.003 higher than the LB",
          "votes": 2
        },
        {
          "id": 320521,
          "postDate": "2018-04-28T22:20:37.087Z",
          "content": "<p>@AmirH I see - that makes sense. Thanks!</p>",
          "rawMarkdown": "@AmirH I see - that makes sense. Thanks!"
        }
      ]
    },
    {
      "id": 319216,
      "postDate": "2018-04-25T14:23:55.953Z",
      "content": "<p><strong>Update - April 25th</strong></p>\n\n<ul>\n<li>Added 7 more features (20 new features in total) and the LB score improved just to 0.9801.</li>\n<li>Looking to work on the feature importance to remove some under performing features and some new features.</li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030\">More Details</a></p>",
      "rawMarkdown": "**Update - April 25th**\n\n- Added 7 more features (20 new features in total) and the LB score improved just to 0.9801.\n- Looking to work on the feature importance to remove some under performing features and some new features.\n\n[More Details][1]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030",
      "votes": 3
    },
    {
      "id": 317413,
      "postDate": "2018-04-21T13:17:11.397Z",
      "content": "<p>0.9798 with 9 new features and lgbm with complete data.</p>\n\n<p><strong>Update - April 23rd</strong></p>\n\n<p>Added 4 more features and the LB score improved to 0.9800.</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030\">More Details Here</a></p>",
      "rawMarkdown": "0.9798 with 9 new features and lgbm with complete data.\n\n**Update - April 23rd**\n\nAdded 4 more features and the LB score improved to 0.9800.\n\n[More Details Here][1]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030",
      "votes": 3,
      "replies": [
        {
          "id": 317428,
          "postDate": "2018-04-21T14:43:11.770Z",
          "content": "<p>may I ask what kind of features are working for you, time delta or frequency features?</p>",
          "rawMarkdown": "may I ask what kind of features are working for you, time delta or frequency features?"
        },
        {
          "id": 317430,
          "postDate": "2018-04-21T14:49:02.517Z",
          "content": "<p>I have a mix of both kinds.... 2 new frequency features helped me move up the lb..  Right now I'm adding 3 more similar features.. probably I might get a better understanding after the current run...</p>",
          "rawMarkdown": "I have a mix of both kinds.... 2 new frequency features helped me move up the lb..  Right now I'm adding 3 more similar features.. probably I might get a better understanding after the current run...",
          "votes": 2
        },
        {
          "id": 317657,
          "postDate": "2018-04-22T06:54:17.673Z",
          "content": "<p>Congrats! When you say \"frequency features\", are you referring to features that are groups of variables mapped to the attributed rate or like just general counts of the feature in the data without relation to is_attributed?</p>",
          "rawMarkdown": "Congrats! When you say \"frequency features\", are you referring to features that are groups of variables mapped to the attributed rate or like just general counts of the feature in the data without relation to is_attributed?",
          "votes": 1
        },
        {
          "id": 318131,
          "postDate": "2018-04-23T08:10:34.303Z",
          "content": "<p>I have both the general count features and also ratio's related to attribute....</p>",
          "rawMarkdown": "I have both the general count features and also ratio's related to attribute....",
          "votes": 1
        },
        {
          "id": 318708,
          "postDate": "2018-04-24T10:00:40.907Z",
          "content": "<p>Samrat -\nHow r u doing \"ratio's related to attribute\" for test set?</p>",
          "rawMarkdown": "Samrat -\nHow r u doing \"ratio's related to attribute\" for test set?"
        }
      ]
    },
    {
      "id": 317295,
      "postDate": "2018-04-21T06:13:22.483Z",
      "content": "<p>0.9807 lgbm with 17 features.</p>",
      "rawMarkdown": "0.9807 lgbm with 17 features.",
      "votes": 3,
      "replies": [
        {
          "id": 317429,
          "postDate": "2018-04-21T14:44:19.470Z",
          "content": "<p>That's very inspiring, <br> may I ask what kind of features are working best for you, time delta or frequency features? </p>",
          "rawMarkdown": "That's very inspiring, <br> may I ask what kind of features are working best for you, time delta or frequency features? "
        },
        {
          "id": 317432,
          "postDate": "2018-04-21T14:57:54.150Z",
          "content": "<p>You can ask but  I won't share before competition ends ;)  This competition is all about feature engineering, just try stuff.</p>",
          "rawMarkdown": "You can ask but  I won't share before competition ends ;)  This competition is all about feature engineering, just try stuff.",
          "votes": 7
        },
        {
          "id": 319174,
          "postDate": "2018-04-25T12:49:53.737Z",
          "content": "<p>are you using target encoding ?</p>",
          "rawMarkdown": "are you using target encoding ?"
        },
        {
          "id": 319191,
          "postDate": "2018-04-25T13:33:09.917Z",
          "content": "<p>You can ask but I won't share before competition ends ;) </p>",
          "rawMarkdown": "You can ask but I won't share before competition ends ;) "
        }
      ]
    },
    {
      "id": 316881,
      "postDate": "2018-04-20T06:18:05.833Z",
      "content": "<p>Gave up on actively competing, so instead have been trying to see how good I can get with a very small set of features - so far got LB 0.9727 with only <strong>4 features</strong> in total. I've still only selected features from <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">feature engineering notebook</a> and using the xgBoost parameters from <a href=\"https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769\">bayesian optimization notebook</a>. Right now I'm trying to add one or two more features, and then I think it can probably be improved a bit more from another round of bayesian tuning afterwards.. and perhaps finally I'll switch to lightGBM as well :)</p>",
      "rawMarkdown": "Gave up on actively competing, so instead have been trying to see how good I can get with a very small set of features - so far got LB 0.9727 with only **4 features** in total. I've still only selected features from [feature engineering notebook](https://www.kaggle.com/nanomathias/feature-engineering-importance-testing) and using the xgBoost parameters from [bayesian optimization notebook](https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769). Right now I'm trying to add one or two more features, and then I think it can probably be improved a bit more from another round of bayesian tuning afterwards.. and perhaps finally I'll switch to lightGBM as well :)",
      "votes": 3,
      "replies": [
        {
          "id": 316888,
          "postDate": "2018-04-20T06:35:05.833Z",
          "content": "<p>You should try 'tree_method'='hist', it helps. XGBoost will run way faster, and you may even get better results.</p>",
          "rawMarkdown": "You should try 'tree_method'='hist', it helps. XGBoost will run way faster, and you may even get better results.",
          "votes": 7
        },
        {
          "id": 316898,
          "postDate": "2018-04-20T06:50:23.763Z",
          "content": "<p>This is a very good comment. Indeed, there are many examples when carefully selected (and aggregated) engineered features could be more effective than multiple features as a way to improve a predictive model. Here is an example by <a href=\"https://www.kaggle.com/pliptor\">@Oskar Takeshita</a> with one single feature: <a href=\"https://www.kaggle.com/pliptor/divide-and-conquer-0-82296\">https://www.kaggle.com/pliptor/divide-and-conquer-0-82296</a></p>",
          "rawMarkdown": "This is a very good comment. Indeed, there are many examples when carefully selected (and aggregated) engineered features could be more effective than multiple features as a way to improve a predictive model. Here is an example by [@Oskar Takeshita][1] with one single feature: https://www.kaggle.com/pliptor/divide-and-conquer-0-82296\n\n\n  [1]: https://www.kaggle.com/pliptor"
        },
        {
          "id": 316940,
          "postDate": "2018-04-20T08:05:10.270Z",
          "content": "<p>Interesting, thanks, I'll be testing out the 'tree_method='hist'' in my next round of tuning</p>",
          "rawMarkdown": "Interesting, thanks, I'll be testing out the 'tree_method='hist'' in my next round of tuning",
          "votes": 1
        }
      ]
    },
    {
      "id": 313003,
      "postDate": "2018-04-12T16:47:30.590Z",
      "content": "<p>My best model is 0.9774 with 22 features and validate on day 9.</p>",
      "rawMarkdown": "My best model is 0.9774 with 22 features and validate on day 9.",
      "votes": 3,
      "replies": [
        {
          "id": 313128,
          "postDate": "2018-04-12T20:10:59.500Z",
          "content": "<p>Good Job! I think your score is one of the highest for a single model approach.</p>",
          "rawMarkdown": "Good Job! I think your score is one of the highest for a single model approach."
        },
        {
          "id": 313347,
          "postDate": "2018-04-13T06:13:10.377Z",
          "content": "<p>Nice! Does it mean that you use as train 7,8 days?</p>",
          "rawMarkdown": "Nice! Does it mean that you use as train 7,8 days?"
        },
        {
          "id": 313848,
          "postDate": "2018-04-14T00:18:34.883Z",
          "content": "<p>Just day 8 actually, due to the limited RAM i got. Not sure if training on different day then combining their prediction on day 9 to predict test set would be helpful, I will try it in the next few days.</p>",
          "rawMarkdown": "Just day 8 actually, due to the limited RAM i got. Not sure if training on different day then combining their prediction on day 9 to predict test set would be helpful, I will try it in the next few days."
        }
      ]
    },
    {
      "id": 311318,
      "postDate": "2018-04-09T20:24:18.700Z",
      "content": "<p>I just reached 0.9727 with my first try of single model with more RAM ... (thanks to the free credits of MS Azure)</p>\n\n<p>Yes high scores with single model are reachable but you'll definitely need more RAM  than on  kaggle kernel (or  on my poor local machine :p)</p>",
      "rawMarkdown": "I just reached 0.9727 with my first try of single model with more RAM ... (thanks to the free credits of MS Azure)\n\nYes high scores with single model are reachable but you'll definitely need more RAM  than on  kaggle kernel (or  on my poor local machine :p)\n",
      "votes": 3,
      "replies": [
        {
          "id": 311329,
          "postDate": "2018-04-09T20:53:30.623Z",
          "content": "<p>That's great.</p>",
          "rawMarkdown": "That's great.",
          "votes": 1
        },
        {
          "id": 311411,
          "postDate": "2018-04-10T02:18:53.173Z",
          "content": "<p>how much RAM to be precise, i am using GCP 40 gb ram, it is still messy at times while creating new features</p>",
          "rawMarkdown": "how much RAM to be precise, i am using GCP 40 gb ram, it is still messy at times while creating new features"
        },
        {
          "id": 312466,
          "postDate": "2018-04-11T19:48:57.257Z",
          "content": "<p>About 42 Gb at the pic and higher at lightgbm init training (you can <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53773\">fix this last issue</a> )</p>\n\n<p>I am not using all the data...just 30 millions rows for training and about the same size for validation ..</p>\n\n<p>After some modification on features used  , I got now 0.9750 on LB with the same model ...</p>\n\n<p>Some features have really huge impact on model performance...I think I need now to dive deeper into features engineering </p>",
          "rawMarkdown": "About 42 Gb at the pic and higher at lightgbm init training (you can [fix this last issue][1] )\n\nI am not using all the data...just 30 millions rows for training and about the same size for validation ..\n\nAfter some modification on features used  , I got now 0.9750 on LB with the same model ...\n\n\nSome features have really huge impact on model performance...I think I need now to dive deeper into features engineering \n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53773",
          "votes": 3
        },
        {
          "id": 312472,
          "postDate": "2018-04-11T20:08:01.030Z",
          "content": "<p>Are your CV and LB in sync ?\nMy concern with using last 30 million or N million rows is Overfitting . Not sure what is better , using last N rows or using Day 2 and Day3 as train and Day 4 as validation . </p>",
          "rawMarkdown": "Are your CV and LB in sync ?\nMy concern with using last 30 million or N million rows is Overfitting . Not sure what is better , using last N rows or using Day 2 and Day3 as train and Day 4 as validation . "
        },
        {
          "id": 312518,
          "postDate": "2018-04-11T22:12:45.387Z",
          "content": "<p>have you used time delta features?</p>",
          "rawMarkdown": "have you used time delta features?"
        },
        {
          "id": 313084,
          "postDate": "2018-04-12T19:03:15.883Z",
          "content": "<p>@MayankSoni\nI am using the second strategy  (using Day 2 and Day3 as train and Day 4 as validation)</p>\n\n<p>@nickhillator <br>\nNot yet but I am working on them now....Hopefully will jump a bit at next submission ^^</p>",
          "rawMarkdown": "@MayankSoni\nI am using the second strategy  (using Day 2 and Day3 as train and Day 4 as validation)\n\n@nickhillator  \nNot yet but I am working on them now....Hopefully will jump a bit at next submission ^^\n \n"
        },
        {
          "id": 313111,
          "postDate": "2018-04-12T19:40:17.620Z",
          "content": "<p>Cool thanks . I am going to try Day 2 Day 3 Train , Validate Day 4 and then Train on Day 3,4 and Predict Test.</p>",
          "rawMarkdown": "Cool thanks . I am going to try Day 2 Day 3 Train , Validate Day 4 and then Train on Day 3,4 and Predict Test."
        }
      ]
    },
    {
      "id": 308093,
      "postDate": "2018-04-02T23:42:22.210Z",
      "content": "<p>Single model 0.9734</p>",
      "rawMarkdown": "Single model 0.9734",
      "votes": 3,
      "replies": [
        {
          "id": 308100,
          "postDate": "2018-04-02T23:51:38.440Z",
          "content": "<p>More than 4 engineered features :(</p>",
          "rawMarkdown": "More than 4 engineered features :(",
          "votes": 3
        },
        {
          "id": 310648,
          "postDate": "2018-04-08T06:54:41.740Z",
          "content": "<p>Punch line: 4 features </p>",
          "rawMarkdown": "Punch line: 4 features "
        }
      ]
    },
    {
      "id": 317296,
      "postDate": "2018-04-21T06:24:46.897Z",
      "content": "<p>0.9808 with 13 features (lgbm)</p>\n\n<p>EDIT: 0.9811 with 25 features (lgbm, 110M raws)</p>",
      "rawMarkdown": "0.9808 with 13 features (lgbm)\n\nEDIT: 0.9811 with 25 features (lgbm, 110M raws)",
      "votes": 4,
      "replies": [
        {
          "id": 317301,
          "postDate": "2018-04-21T06:43:12.263Z",
          "content": "<p>are you using the entire training set?</p>",
          "rawMarkdown": "are you using the entire training set?"
        },
        {
          "id": 317415,
          "postDate": "2018-04-21T13:24:33.443Z",
          "content": "<p>110M rows only due to RAM limit. CV and LB scores improve as I use more data.</p>",
          "rawMarkdown": "110M rows only due to RAM limit. CV and LB scores improve as I use more data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 308002,
      "postDate": "2018-04-02T19:35:55.713Z",
      "content": "<p>Right now my best single model score is public. The <a href=\"https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python\">single kernel version</a> scores 0.9694.  The <a href=\"https://www.kaggle.com/aharless/ceci-n-est-pas-un-m-lange\">multiple-kernel version</a> (two kernels using the same model for different time periods and blended using guesstimated weights based on their LB scores) scores 0.9695.  I could run the model on the full dataset at home, but I haven't tried yet, and not sure I will with that particular model.  It has a history in public kernels, and I fear it is overfit to the public LB already.  I'm trying to focus more on validation and development, which is time-consuming given the large size of the dataset.  I'm not at the point where original candidates for a possible best score are ready to submit.</p>",
      "rawMarkdown": "Right now my best single model score is public. The [single kernel version][1] scores 0.9694.  The [multiple-kernel version][2] (two kernels using the same model for different time periods and blended using guesstimated weights based on their LB scores) scores 0.9695.  I could run the model on the full dataset at home, but I haven't tried yet, and not sure I will with that particular model.  It has a history in public kernels, and I fear it is overfit to the public LB already.  I'm trying to focus more on validation and development, which is time-consuming given the large size of the dataset.  I'm not at the point where original candidates for a possible best score are ready to submit.\n\n [1]: https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python\n [2]: https://www.kaggle.com/aharless/ceci-n-est-pas-un-m-lange",
      "votes": 3,
      "replies": [
        {
          "id": 308013,
          "postDate": "2018-04-02T20:00:04.393Z",
          "content": "<p>Thanks for your insight Andy. In my bucket list, next step could be either better generalization with subsets for training and validation or some exploration into Neural Network based approach or a hybrid with tensorflows estimators. I am not sure if anyone has tried that yet. </p>",
          "rawMarkdown": "Thanks for your insight Andy. In my bucket list, next step could be either better generalization with subsets for training and validation or some exploration into Neural Network based approach or a hybrid with tensorflows estimators. I am not sure if anyone has tried that yet. "
        },
        {
          "id": 308058,
          "postDate": "2018-04-02T21:51:44.053Z",
          "content": "<p>Thanks Andy. And no matter what score your have, you are the hero in this game. Learned a lot from you and thank you again.</p>",
          "rawMarkdown": "Thanks Andy. And no matter what score your have, you are the hero in this game. Learned a lot from you and thank you again.",
          "votes": 3
        }
      ]
    },
    {
      "id": 320834,
      "postDate": "2018-04-30T01:27:44.743Z",
      "content": "<p>9807, 22 features </p>",
      "rawMarkdown": "9807, 22 features ",
      "votes": 1,
      "replies": [
        {
          "id": 322566,
          "postDate": "2018-05-03T07:50:37.553Z",
          "content": "<p>All data ?</p>",
          "rawMarkdown": "All data ?"
        }
      ]
    },
    {
      "id": 320615,
      "postDate": "2018-04-29T09:18:42.200Z",
      "content": "<p>single lgb, 0.9793, 80M rows, 23 features</p>",
      "rawMarkdown": "single lgb, 0.9793, 80M rows, 23 features",
      "votes": 1,
      "replies": [
        {
          "id": 320829,
          "postDate": "2018-04-30T01:11:25.767Z",
          "content": "<p>update: 0.9804 now</p>",
          "rawMarkdown": "update: 0.9804 now"
        },
        {
          "id": 320863,
          "postDate": "2018-04-30T03:52:04.653Z",
          "content": "<p>Congrats! Are you using the last 80M rows? So mostly day 9 right?</p>",
          "rawMarkdown": "Congrats! Are you using the last 80M rows? So mostly day 9 right?"
        },
        {
          "id": 320883,
          "postDate": "2018-04-30T05:21:54.623Z",
          "content": "<p>Thanks. I'm currently using day 9 as validation, so the 80M rows mostly come from day 8.</p>",
          "rawMarkdown": "Thanks. I'm currently using day 9 as validation, so the 80M rows mostly come from day 8."
        }
      ]
    },
    {
      "id": 316604,
      "postDate": "2018-04-19T12:59:58.767Z",
      "content": "<p>Got LB 0.9791 with 23 features.Single Model.</p>",
      "rawMarkdown": "Got LB 0.9791 with 23 features.Single Model.",
      "votes": 1
    },
    {
      "id": 315224,
      "postDate": "2018-04-16T22:19:17.490Z",
      "content": "<p>Our team has 13 features to reach 0.9803 with a single model using days 7 &amp; 8 as training and 9 to evaluate.</p>",
      "rawMarkdown": "Our team has 13 features to reach 0.9803 with a single model using days 7 &amp; 8 as training and 9 to evaluate.",
      "votes": 1,
      "replies": [
        {
          "id": 315355,
          "postDate": "2018-04-17T02:18:38.580Z",
          "content": "<p>entire day 9? or 10% of day 9</p>",
          "rawMarkdown": "entire day 9? or 10% of day 9"
        },
        {
          "id": 315367,
          "postDate": "2018-04-17T02:43:27.900Z",
          "content": "<p>entire day 9 but the evaluation stays the evaluation, the training is using the entire days 7 and 8 and it is the most important. I am a DS who thinks that we don't need 100 features to have a good model. I prefer to use 6 and reach 0.976 rather than 100 to reach 0.986 where the difference in time consumption and complexity for the company is too big to realize. In BIG DATA we prefer low complexity, fast result and not really the perfect score (a good one is sufficient). I think that interpretation of the variables is important. </p>\n\n<p>Other clue: Do not think complicated, all the variables are very simple and USE the test_sup you'll definitively increase your score. I first was using something really complicated till I realize that it is more Ad Tracking than Fraud Detection...</p>",
          "rawMarkdown": "entire day 9 but the evaluation stays the evaluation, the training is using the entire days 7 and 8 and it is the most important. I am a DS who thinks that we don't need 100 features to have a good model. I prefer to use 6 and reach 0.976 rather than 100 to reach 0.986 where the difference in time consumption and complexity for the company is too big to realize. In BIG DATA we prefer low complexity, fast result and not really the perfect score (a good one is sufficient). I think that interpretation of the variables is important. \n\nOther clue: Do not think complicated, all the variables are very simple and USE the test_sup you'll definitively increase your score. I first was using something really complicated till I realize that it is more Ad Tracking than Fraud Detection...",
          "votes": 8
        },
        {
          "id": 315502,
          "postDate": "2018-04-17T07:31:17.623Z",
          "content": "<p>Hi Badr,\nWhen using day 9 as evaluation,  do you get close AUC between the CV and LB?\nor CV AUC is much higher and you just train to maximize it? \nDid you have occurances where you train for higher AUC in the CV and get lower LB Score?</p>\n\n<p>Thanks!! =)</p>",
          "rawMarkdown": "Hi Badr,\nWhen using day 9 as evaluation,  do you get close AUC between the CV and LB?\nor CV AUC is much higher and you just train to maximize it? \nDid you have occurances where you train for higher AUC in the CV and get lower LB Score?\n\nThanks!! =)",
          "votes": 2
        },
        {
          "id": 315771,
          "postDate": "2018-04-17T15:18:08.523Z",
          "content": "<p>@ Badr - Very well written. :) I like the way you want to approach this problem in terms of a stable logical solution and low complexity .</p>",
          "rawMarkdown": "@ Badr - Very well written. :) I like the way you want to approach this problem in terms of a stable logical solution and low complexity .",
          "votes": 1
        },
        {
          "id": 316012,
          "postDate": "2018-04-18T02:19:10.960Z",
          "content": "<p>Unfortunately, a difference of 0.01 differentiates between getting top 10% vs top 1%. kaggle is really about getting that incremental increase in accuracy with insane complexity.</p>",
          "rawMarkdown": "Unfortunately, a difference of 0.01 differentiates between getting top 10% vs top 1%. kaggle is really about getting that incremental increase in accuracy with insane complexity.",
          "votes": 2
        },
        {
          "id": 316038,
          "postDate": "2018-04-18T04:04:21.110Z",
          "content": "<p>@AmirH, the evaluation and test scores are really close but the training score is much more bigger.</p>\n\n<p>@Shanth, Thank you very much I really appreciate. I am much more thinking about the company use case and a real application rather than the perfect score for very high complexity.</p>\n\n<p>@leexa, yes totally insane that's why I think I will stop to work to reach the score of the tops of the leaderboard. I know that I can reach 0.9814 something like this using 20 more features involving different time series variables but as it will take too much time to be realized and as I think that the company is much more interested into Fraud Detection rather than Ad Tracking, I prefer to stop there. When you think about it who cares about predicting which click will be a downloading one? When people click you already have the result. In the other hand, it is a real challenging problem to detect fraudulent clicks. But fraudulent machines are also downloading just to be less suspected... If people are enjoying the competition and are happy to have a silver or gold rank, we can let them improve their resume (why not?)</p>",
          "rawMarkdown": "@AmirH, the evaluation and test scores are really close but the training score is much more bigger.\n\n@Shanth, Thank you very much I really appreciate. I am much more thinking about the company use case and a real application rather than the perfect score for very high complexity.\n\n@leexa, yes totally insane that's why I think I will stop to work to reach the score of the tops of the leaderboard. I know that I can reach 0.9814 something like this using 20 more features involving different time series variables but as it will take too much time to be realized and as I think that the company is much more interested into Fraud Detection rather than Ad Tracking, I prefer to stop there. When you think about it who cares about predicting which click will be a downloading one? When people click you already have the result. In the other hand, it is a real challenging problem to detect fraudulent clicks. But fraudulent machines are also downloading just to be less suspected... If people are enjoying the competition and are happy to have a silver or gold rank, we can let them improve their resume (why not?)",
          "votes": -4
        },
        {
          "id": 316983,
          "postDate": "2018-04-20T10:45:03.190Z",
          "content": "<blockquote>\n  <p>I know that I can reach 0.9814 </p>\n</blockquote>\n\n<p>How do you know that?  If you know without submitting then you found something quite interesting.</p>\n\n<p>Other than that, I agree one can move above 0.98 with less than 15 features, I know via my own submissions ;)</p>",
          "rawMarkdown": "&gt; I know that I can reach 0.9814 \n\nHow do you know that?  If you know without submitting then you found something quite interesting.\n\nOther than that, I agree one can move above 0.98 with less than 15 features, I know via my own submissions ;)",
          "votes": 8
        }
      ]
    },
    {
      "id": 310750,
      "postDate": "2018-04-08T13:33:35.873Z",
      "content": "<p>My best single model scores 0.9751 on LB.</p>",
      "rawMarkdown": "My best single model scores 0.9751 on LB.",
      "votes": 1,
      "replies": [
        {
          "id": 310890,
          "postDate": "2018-04-09T01:39:10.277Z",
          "content": "<p>nice job! Could I ask what is your estimated private lb score?</p>",
          "rawMarkdown": "nice job! Could I ask what is your estimated private lb score?"
        },
        {
          "id": 310893,
          "postDate": "2018-04-09T01:41:23.313Z",
          "content": "<p>Not sure how to estimate that...I’m new here...</p>",
          "rawMarkdown": "Not sure how to estimate that...I’m new here..."
        },
        {
          "id": 310894,
          "postDate": "2018-04-09T01:45:16.187Z",
          "content": "<p>@Chad, my training and validation procedure: train on day7 and day 8, validate on day 9. Then using all data as train data, build model, generate predictions for test. How's yours?</p>",
          "rawMarkdown": "@Chad, my training and validation procedure: train on day7 and day 8, validate on day 9. Then using all data as train data, build model, generate predictions for test. How's yours?"
        },
        {
          "id": 310899,
          "postDate": "2018-04-09T02:02:18.673Z",
          "content": "<p>I have pretty limited resources at home....this score was just train on 7, validate on 9, then use same model without retraining to make predictions on test. I hope to train on more of the data soon. I'll post any improvements!</p>",
          "rawMarkdown": "I have pretty limited resources at home....this score was just train on 7, validate on 9, then use same model without retraining to make predictions on test. I hope to train on more of the data soon. I'll post any improvements!",
          "votes": 2
        },
        {
          "id": 310986,
          "postDate": "2018-04-09T06:46:01.400Z",
          "content": "<p>any specific reason you train on day7 and not day 8?</p>",
          "rawMarkdown": "any specific reason you train on day7 and not day 8?"
        },
        {
          "id": 311114,
          "postDate": "2018-04-09T12:20:11.543Z",
          "content": "<p>Just hadn't done my feature processing on it :). I let it run last night (takes about 3 hours), so now I have both days to work with!</p>",
          "rawMarkdown": "Just hadn't done my feature processing on it :). I let it run last night (takes about 3 hours), so now I have both days to work with!"
        }
      ]
    },
    {
      "id": 308564,
      "postDate": "2018-04-03T17:31:08.523Z",
      "content": "<p>Currently my best model LB 0.9694\ni wish to go LB 0.97 ...</p>",
      "rawMarkdown": "Currently my best model LB 0.9694\ni wish to go LB 0.97 ...",
      "votes": 1
    },
    {
      "id": 308147,
      "postDate": "2018-04-03T02:59:26.777Z",
      "content": "<p>I'm wondering what is the best score for somebody using a subsample of train data.  I have no way of running a model on the full train set, so training everything on various subsets.  Wondering how far that can go...</p>",
      "rawMarkdown": "I'm wondering what is the best score for somebody using a subsample of train data.  I have no way of running a model on the full train set, so training everything on various subsets.  Wondering how far that can go...",
      "votes": 1,
      "replies": [
        {
          "id": 308162,
          "postDate": "2018-04-03T03:54:07.773Z",
          "content": "<p>I am running on a sub sample of data, very similar to Andy's Train and Validation Split. My best single model is 0.9704, based on what people are able to achieve, I really need to dig deep into features. Glad that this competition calls for great feature engineering.</p>",
          "rawMarkdown": "I am running on a sub sample of data, very similar to Andy's Train and Validation Split. My best single model is 0.9704, based on what people are able to achieve, I really need to dig deep into features. Glad that this competition calls for great feature engineering.",
          "votes": 3
        },
        {
          "id": 308164,
          "postDate": "2018-04-03T04:00:59.003Z",
          "content": "<p>Do you find a lot of disparity in results between different subsets?  Mine seem to fluctuate a lot based on whether I run them on day 2 or day 3 train (day 4 reserved for validation).  On local validation the same model gives 0.9642 trained on day 3 and 0.9695 trained on day 2.  Seems like a big difference...</p>",
          "rawMarkdown": "Do you find a lot of disparity in results between different subsets?  Mine seem to fluctuate a lot based on whether I run them on day 2 or day 3 train (day 4 reserved for validation).  On local validation the same model gives 0.9642 trained on day 3 and 0.9695 trained on day 2.  Seems like a big difference...",
          "votes": 1
        },
        {
          "id": 308165,
          "postDate": "2018-04-03T04:03:23.267Z",
          "content": "<p>Some kinds of models you can run on the full train set even with low memory, if they can process the data in mini-batches, as in several public kernels that use FTRL (<a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9681\">example</a>).  Or you can run a model separately on different parts of the data and take some kind of weighted average of the results, as in the \"multiple kernel version\" I referenced in my comment above.</p>",
          "rawMarkdown": "Some kinds of models you can run on the full train set even with low memory, if they can process the data in mini-batches, as in several public kernels that use FTRL ([example][1]).  Or you can run a model separately on different parts of the data and take some kind of weighted average of the results, as in the \"multiple kernel version\" I referenced in my comment above.\n\n  [1]: https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9681"
        },
        {
          "id": 308177,
          "postDate": "2018-04-03T04:38:02.323Z",
          "content": "<p>I followed this discussion and havent tried anything beyond this <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492#303536\">Validation</a>. Based on this, one can get a stable validation, however one might overfit LB based on a similar looking validation, so I am validating on entire day as well as done by Andy with his awesome kernels :D and proceeding only if both my validation gets a jump based on newly engineered feature or any other implementation.</p>",
          "rawMarkdown": "I followed this discussion and havent tried anything beyond this [Validation][1]. Based on this, one can get a stable validation, however one might overfit LB based on a similar looking validation, so I am validating on entire day as well as done by Andy with his awesome kernels :D and proceeding only if both my validation gets a jump based on newly engineered feature or any other implementation.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492#303536"
        },
        {
          "id": 308911,
          "postDate": "2018-04-04T10:15:09.153Z",
          "content": "<p>reply to @yulia: There is some difference between the days. Day 4 is somehow different  from the other days and more similar to the test day. I get better results when training only on day 4.  </p>",
          "rawMarkdown": "reply to @yulia: There is some difference between the days. Day 4 is somehow different  from the other days and more similar to the test day. I get better results when training only on day 4.  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 320082,
      "postDate": "2018-04-27T12:44:47.833Z",
      "content": "<p>Getting 0.9786 with 10 features, running on Kaggle Kernel with 40M rows</p>",
      "rawMarkdown": "Getting 0.9786 with 10 features, running on Kaggle Kernel with 40M rows",
      "votes": 2
    },
    {
      "id": 318576,
      "postDate": "2018-04-24T05:06:13.820Z",
      "content": "<p>My NN (based on <a href=\"https://www.kaggle.com/antmarakis/deep-learning-approach-validation-lb-0-9684\">https://www.kaggle.com/antmarakis/deep-learning-approach-validation-lb-0-9684</a>) trained on 100 m datas(day8 and day9) and got 0.977 lb score(local auc 0.9812). <br></p>\n\n<p><em><strong></strong></em><strong>__<em>_</em>_</strong><em>update</em><strong><em>_</em>__<em>_</em>__<em>_</em>__<em>_</em>__<em>_</em></strong> <br>\nNow my nn model can get 0.9777(local auc 0.982369).</p>",
      "rawMarkdown": "My NN (based on https://www.kaggle.com/antmarakis/deep-learning-approach-validation-lb-0-9684) trained on 100 m datas(day8 and day9) and got 0.977 lb score(local auc 0.9812). <br>\n\n_____________update____________________________ <br>\nNow my nn model can get 0.9777(local auc 0.982369).\n",
      "votes": 2,
      "replies": [
        {
          "id": 318583,
          "postDate": "2018-04-24T05:18:44.223Z",
          "content": "<p>Are many features are you using for NN? your current score(LB) seems like from other model.</p>",
          "rawMarkdown": "Are many features are you using for NN? your current score(LB) seems like from other model."
        },
        {
          "id": 318595,
          "postDate": "2018-04-24T05:32:21.260Z",
          "content": "<p>Yeah. I added 11 counting features.</p>",
          "rawMarkdown": "Yeah. I added 11 counting features."
        },
        {
          "id": 318762,
          "postDate": "2018-04-24T12:44:12.743Z",
          "content": "<p>Nice! If you don't mind sharing, your architecture is the same with all the count features passed to embeddings?</p>",
          "rawMarkdown": "Nice! If you don't mind sharing, your architecture is the same with all the count features passed to embeddings?"
        },
        {
          "id": 318775,
          "postDate": "2018-04-24T13:11:45.143Z",
          "content": "<p>I hope not, lol.</p>",
          "rawMarkdown": "I hope not, lol."
        },
        {
          "id": 318777,
          "postDate": "2018-04-24T13:14:09.537Z",
          "content": "<p><a href=\"/authman\">@authman</a> because it seems like it doesn't make sense to pass counts (numerical) to embeddings?</p>",
          "rawMarkdown": "@authman because it seems like it doesn't make sense to pass counts (numerical) to embeddings?"
        },
        {
          "id": 318780,
          "postDate": "2018-04-24T13:17:17.153Z",
          "content": "<p>Yeah. I used Embedding for all the count features and concatenate them  together just like Anthony's model. </p>",
          "rawMarkdown": "Yeah. I used Embedding for all the count features and concatenate them  together just like Anthony's model. ",
          "votes": 1
        },
        {
          "id": 318782,
          "postDate": "2018-04-24T13:22:17.170Z",
          "content": "<p>Thank you for sharing!</p>",
          "rawMarkdown": "Thank you for sharing!"
        },
        {
          "id": 318788,
          "postDate": "2018-04-24T13:32:08.533Z",
          "content": "<p>It seems like there'd be a lot of information loss with that setup. I guess score tumps all =) but the only way it'd make sense to me would be a ridiculously large dset where there's decent coverage for most represented count values, or if a DAE was used for pre-training which included tied weights between encoder/decoder.</p>",
          "rawMarkdown": "It seems like there'd be a lot of information loss with that setup. I guess score tumps all =) but the only way it'd make sense to me would be a ridiculously large dset where there's decent coverage for most represented count values, or if a DAE was used for pre-training which included tied weights between encoder/decoder.",
          "votes": 2
        },
        {
          "id": 318811,
          "postDate": "2018-04-24T14:13:27.237Z",
          "content": "<p>There's definitely information loss, but my theory is that there's a lot of relevant information in the smaller count values that are more likely to be shared across many different aggregates, and less relevant information in the larger count values that are less likely to be shared. Large count values may just be spammy and not really tell you much in terms of the range they fall into. Believe that <a href=\"https://www.kaggle.com/its7171/download-once-and-only-once\">this is related</a>. Losing information about large counts may actually be playing a regularization role here if that information is largely useless. </p>",
          "rawMarkdown": "There's definitely information loss, but my theory is that there's a lot of relevant information in the smaller count values that are more likely to be shared across many different aggregates, and less relevant information in the larger count values that are less likely to be shared. Large count values may just be spammy and not really tell you much in terms of the range they fall into. Believe that [this is related][1]. Losing information about large counts may actually be playing a regularization role here if that information is largely useless. \n\n\n  [1]: https://www.kaggle.com/its7171/download-once-and-only-once",
          "votes": 2
        }
      ]
    },
    {
      "id": 317298,
      "postDate": "2018-04-21T06:29:36.807Z",
      "content": "<p>added three features and LB score jumped from 0.9691 to 0.9799...with single lightgbm</p>\n\n<p><strong>update:</strong> 0.9810 using all data...</p>",
      "rawMarkdown": "added three features and LB score jumped from 0.9691 to 0.9799...with single lightgbm\n\n**update:** 0.9810 using all data...",
      "votes": 2,
      "replies": [
        {
          "id": 318017,
          "postDate": "2018-04-23T03:29:58.553Z",
          "content": "<p>Congrats! Do you have any tips for what features might work best?</p>",
          "rawMarkdown": "Congrats! Do you have any tips for what features might work best?"
        },
        {
          "id": 318165,
          "postDate": "2018-04-23T09:32:29.103Z",
          "content": "<p>Hi Ed, thx.\nApart from count, ratio features, i simply added \"time_delta\" features.</p>",
          "rawMarkdown": "Hi Ed, thx.\nApart from count, ratio features, i simply added \"time_delta\" features."
        },
        {
          "id": 318821,
          "postDate": "2018-04-24T14:33:05.533Z",
          "content": "<p>i notice how you said time delta features instead of feature, now i can work on it</p>",
          "rawMarkdown": "i notice how you said time delta features instead of feature, now i can work on it"
        },
        {
          "id": 320099,
          "postDate": "2018-04-27T13:17:12.823Z",
          "content": "<p>Hi Endi:\nAre you using single time delta features or multiple with different group by.</p>",
          "rawMarkdown": "Hi Endi:\nAre you using single time delta features or multiple with different group by."
        },
        {
          "id": 320667,
          "postDate": "2018-04-29T13:03:22.053Z",
          "content": "<p>Hi, i think the best part on kaggle competition is to try different ideas and get feedback from LB. Anyway, sth work well in my model doesn't guarantee it works in your model. ( ;</p>",
          "rawMarkdown": "Hi, i think the best part on kaggle competition is to try different ideas and get feedback from LB. Anyway, sth work well in my model doesn't guarantee it works in your model. ( ;",
          "votes": 1
        },
        {
          "id": 322544,
          "postDate": "2018-05-03T06:57:15.867Z",
          "content": "<p>do you use the test supplement or test set? </p>",
          "rawMarkdown": "do you use the test supplement or test set? "
        }
      ]
    },
    {
      "id": 308967,
      "postDate": "2018-04-04T12:48:22.673Z",
      "content": "<p>My current single model has a public LB score of 0.9710. It has 12 features (including the ones that come with the data set). Still working on it...</p>",
      "rawMarkdown": "My current single model has a public LB score of 0.9710. It has 12 features (including the ones that come with the data set). Still working on it...",
      "votes": 2,
      "replies": [
        {
          "id": 309842,
          "postDate": "2018-04-06T03:44:43.180Z",
          "content": "<p>Thanks you! Is your current public LB score from additional stacking then?</p>",
          "rawMarkdown": "Thanks you! Is your current public LB score from additional stacking then?"
        },
        {
          "id": 310021,
          "postDate": "2018-04-06T12:21:11.340Z",
          "content": "<p>No. The score improved because I recently added a couple of additional features.</p>",
          "rawMarkdown": "No. The score improved because I recently added a couple of additional features."
        },
        {
          "id": 310090,
          "postDate": "2018-04-06T15:32:40.743Z",
          "content": "<p>@ Alexey - This is a basic question so please do bear with me. How do you gauge the improvement/ decrease in model accuracy from the validation score.? ( Or in other words, how do you interpret the validation score )</p>\n\n<p>For ex: I added a feature --&gt; Got a 0.002X increase in training accuracy but it did not move me on the LB</p>",
          "rawMarkdown": "@ Alexey - This is a basic question so please do bear with me. How do you gauge the improvement/ decrease in model accuracy from the validation score.? ( Or in other words, how do you interpret the validation score )\n\nFor ex: I added a feature --&gt; Got a 0.002X increase in training accuracy but it did not move me on the LB"
        },
        {
          "id": 310094,
          "postDate": "2018-04-06T15:35:29.970Z",
          "content": "<p>training accuracy or validation accuracy?  The latter is the one to watch.</p>",
          "rawMarkdown": "training accuracy or validation accuracy?  The latter is the one to watch.",
          "votes": 2
        },
        {
          "id": 310095,
          "postDate": "2018-04-06T15:46:34.203Z",
          "content": "<p>In order to be able to draw any conclusions about your model's accuracy from the validation score, your validation set must resemble as close as possible the test data. In real life, the test set is not known to us ahead of time, but Kaggle's competitions are different in this respect -- here we are provided one. This opens up a possibility to apply some clever strategies for selecting your validation set. If you are really interested in this subject you might want to google \"adversarial validation\". It gives a interesting example of such a strategy. \nPersonally, I am still working on feature engineering and do not use any fancy validation strategy at the moment. I just randomly select 5% of the data from the training set and use them for validation. Of course, this does not guarantee that any changes in my validation score will always mimic the corresponding changing in the LB score but it gives me at least a rough idea.</p>",
          "rawMarkdown": "In order to be able to draw any conclusions about your model's accuracy from the validation score, your validation set must resemble as close as possible the test data. In real life, the test set is not known to us ahead of time, but Kaggle's competitions are different in this respect -- here we are provided one. This opens up a possibility to apply some clever strategies for selecting your validation set. If you are really interested in this subject you might want to google \"adversarial validation\". It gives a interesting example of such a strategy. \nPersonally, I am still working on feature engineering and do not use any fancy validation strategy at the moment. I just randomly select 5% of the data from the training set and use them for validation. Of course, this does not guarantee that any changes in my validation score will always mimic the corresponding changing in the LB score but it gives me at least a rough idea.",
          "votes": 7
        },
        {
          "id": 310107,
          "postDate": "2018-04-06T16:19:56.117Z",
          "content": "<p>Hi @CPMP, could I ask what validation strategy you are using? Is it validating on 9th 4:00 to 15:59 and training on data before validation set? And please forgive me if this question is too private and I'm totally fine if you want to keep your validation method secret. Thank you.</p>",
          "rawMarkdown": "Hi @CPMP, could I ask what validation strategy you are using? Is it validating on 9th 4:00 to 15:59 and training on data before validation set? And please forgive me if this question is too private and I'm totally fine if you want to keep your validation method secret. Thank you.",
          "votes": 1
        },
        {
          "id": 310128,
          "postDate": "2018-04-06T16:57:01.373Z",
          "content": "<p>I'm redoing everything as I goofed my data prep.  But so far I was using last day as validation.  </p>",
          "rawMarkdown": "I'm redoing everything as I goofed my data prep.  But so far I was using last day as validation.  ",
          "votes": 1
        },
        {
          "id": 310132,
          "postDate": "2018-04-06T16:58:04.393Z",
          "content": "<p>@CPMP, what I could contribute is an upvote and say thank you!</p>",
          "rawMarkdown": "@CPMP, what I could contribute is an upvote and say thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 308209,
      "postDate": "2018-04-03T06:06:42.970Z",
      "content": "<p>My best single model score is my current LB score:  0.9731 with 26 features.  I will dig deep into feature engineer after finished my graduation thesis.</p>",
      "rawMarkdown": "My best single model score is my current LB score:  0.9731 with 26 features.  I will dig deep into feature engineer after finished my graduation thesis.",
      "votes": 2
    },
    {
      "id": 308060,
      "postDate": "2018-04-02T22:01:14.790Z",
      "content": "<p>My best single model scores 0.9718 on public lb and I've created almost 40 new features.  I think I must be on a wrong way because Danijel said that he achieved 0.976+ with only four new features. And I've noticed that many top rankers jump from 0.97 to 0.98 very quickly thus I think there should be some magic features. Emmm...</p>",
      "rawMarkdown": "My best single model scores 0.9718 on public lb and I've created almost 40 new features.  I think I must be on a wrong way because Danijel said that he achieved 0.976+ with only four new features. And I've noticed that many top rankers jump from 0.97 to 0.98 very quickly thus I think there should be some magic features. Emmm...",
      "votes": 2,
      "replies": [
        {
          "id": 308063,
          "postDate": "2018-04-02T22:06:40.830Z",
          "content": "<p>There are always some magic features, or maybe the jump was because of better validation strategy or ensemble?  </p>\n\n<p>\"0.976+ with only four new features\" this is valuable.. </p>",
          "rawMarkdown": "There are always some magic features, or maybe the jump was because of better validation strategy or ensemble?  \n\n\"0.976+ with only four new features\" this is valuable.. "
        },
        {
          "id": 308074,
          "postDate": "2018-04-02T22:50:21.533Z",
          "content": "<p>well, what popular validation strategies do we have right now?</p>",
          "rawMarkdown": "well, what popular validation strategies do we have right now?"
        },
        {
          "id": 308105,
          "postDate": "2018-04-02T23:59:20.040Z",
          "content": "<p>So far Andy's seem the best. I'll be using that next. It's my guess that the leaders have something better for validation (I wish I had a clue about it..but then I am just starting..)</p>",
          "rawMarkdown": "So far Andy's seem the best. I'll be using that next. It's my guess that the leaders have something better for validation (I wish I had a clue about it..but then I am just starting..)"
        },
        {
          "id": 309112,
          "postDate": "2018-04-04T17:00:41.953Z",
          "content": "<p>@snorlax - How much data did you use?</p>",
          "rawMarkdown": "@snorlax - How much data did you use?"
        },
        {
          "id": 309123,
          "postDate": "2018-04-04T17:23:50.157Z",
          "content": "<p>all data. I tried to use small subset and the performance may go downward 0.01</p>",
          "rawMarkdown": "all data. I tried to use small subset and the performance may go downward 0.01",
          "votes": 1
        },
        {
          "id": 312595,
          "postDate": "2018-04-12T02:41:30.600Z",
          "content": "<p>do you mean, keep on changing the <code>random_state</code> of train_test_split  ?</p>",
          "rawMarkdown": "do you mean, keep on changing the `random_state` of train_test_split  ?"
        }
      ]
    },
    {
      "id": 307998,
      "postDate": "2018-04-02T19:23:17.713Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53250\">Have a look at this</a></p>",
      "rawMarkdown": "[Have a look at this][1]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53250",
      "votes": 2,
      "replies": [
        {
          "id": 308008,
          "postDate": "2018-04-02T19:48:58.260Z",
          "content": "<p>Looked,  Yes, this is a duplicate but the intention to ask is different.  Not asking for any help, tips or shortcuts.  Just simple info. I do not see any harm in it :)</p>",
          "rawMarkdown": "Looked,  Yes, this is a duplicate but the intention to ask is different.  Not asking for any help, tips or shortcuts.  Just simple info. I do not see any harm in it :)",
          "votes": 1
        },
        {
          "id": 308010,
          "postDate": "2018-04-02T19:56:08.360Z",
          "content": "<p>My best single model is my current LB score, and I believe it can improve, most of the top models in top 50 are single models I guess. </p>",
          "rawMarkdown": "My best single model is my current LB score, and I believe it can improve, most of the top models in top 50 are single models I guess. ",
          "votes": 1
        },
        {
          "id": 308016,
          "postDate": "2018-04-02T20:02:32.570Z",
          "content": "<p>That's great, I think validation is working for the best. Thanks.</p>",
          "rawMarkdown": "That's great, I think validation is working for the best. Thanks."
        }
      ]
    },
    {
      "id": 310887,
      "postDate": "2018-04-09T01:35:12.757Z",
      "content": "<p>14 features, LGB, hitting  9.770 (edit: 0.9770  Thanks for the catch @CPMP :)</p>",
      "rawMarkdown": "14 features, LGB, hitting  9.770 (edit: 0.9770  Thanks for the catch @CPMP :)",
      "replies": [
        {
          "id": 310900,
          "postDate": "2018-04-09T02:03:36.470Z",
          "content": "<blockquote>\n  <p>hitting 9.770</p>\n</blockquote>\n\n<p>Wow, 10x better than everyone else! ;)</p>",
          "rawMarkdown": "&gt; hitting 9.770\n\nWow, 10x better than everyone else! ;)",
          "votes": 1
        },
        {
          "id": 310909,
          "postDate": "2018-04-09T02:44:17.523Z",
          "content": "<p>Haha oops</p>",
          "rawMarkdown": "Haha oops",
          "votes": 1
        },
        {
          "id": 310911,
          "postDate": "2018-04-09T02:45:41.563Z",
          "content": "<p>WOW, amazing! You guys are soooooooooo smart!</p>",
          "rawMarkdown": "WOW, amazing! You guys are soooooooooo smart!",
          "votes": 1
        },
        {
          "id": 310914,
          "postDate": "2018-04-09T02:59:23.963Z",
          "content": "<p>Even with the edit you have a great start for your first competition!</p>",
          "rawMarkdown": "Even with the edit you have a great start for your first competition!"
        },
        {
          "id": 310916,
          "postDate": "2018-04-09T03:05:49.433Z",
          "content": "<p>Thank you! I'm learning a lot from your kernels and all the other great people here.</p>",
          "rawMarkdown": "Thank you! I'm learning a lot from your kernels and all the other great people here.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3015835,
      "postDate": "2024-10-13T02:50:37.783Z",
      "content": "<p>nicee       </p>",
      "rawMarkdown": "nicee       "
    },
    {
      "id": 578234,
      "postDate": "2019-07-17T13:52:53.417Z",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice"
    },
    {
      "id": 540213,
      "postDate": "2019-05-31T06:53:45.380Z",
      "content": "<p>Bagged trees haha</p>",
      "rawMarkdown": "Bagged trees haha"
    },
    {
      "id": 395367,
      "postDate": "2018-09-28T12:49:36.457Z",
      "content": "<p>wow ..nice!!</p>",
      "rawMarkdown": "wow ..nice!!"
    },
    {
      "id": 373053,
      "postDate": "2018-08-21T00:12:39.253Z",
      "content": "<p>gj!</p>",
      "rawMarkdown": "gj!"
    },
    {
      "id": 325144,
      "postDate": "2018-05-08T07:03:50.743Z",
      "content": "<p>lgbm with 18 plus more feature 120m data private score 0.9802</p>",
      "rawMarkdown": "lgbm with 18 plus more feature 120m data private score 0.9802"
    },
    {
      "id": 321825,
      "postDate": "2018-05-02T01:19:47.167Z",
      "content": "<p>0.9795 total 22 features </p>",
      "rawMarkdown": "0.9795 total 22 features "
    },
    {
      "id": 321438,
      "postDate": "2018-05-01T09:37:25.120Z",
      "content": "<p>0.9792, 7 new features with all training data. Still working on it!</p>",
      "rawMarkdown": "0.9792, 7 new features with all training data. Still working on it!"
    },
    {
      "id": 319717,
      "postDate": "2018-04-26T17:16:32.900Z",
      "content": "<p><strong>Update - April 26th</strong> - Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.</p>",
      "rawMarkdown": "**Update - April 26th** - Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.",
      "replies": [
        {
          "id": 319723,
          "postDate": "2018-04-26T17:48:13.907Z",
          "content": "<p>your local cross validation score looks very high ? are those data included in your training set, or is your validation set too small ? It seems that somehow you are overfitting it...</p>\n\n<p>To get to 0.9806 I have a local validation of 0.984668.</p>",
          "rawMarkdown": "your local cross validation score looks very high ? are those data included in your training set, or is your validation set too small ? It seems that somehow you are overfitting it...\n\nTo get to 0.9806 I have a local validation of 0.984668."
        },
        {
          "id": 319728,
          "postDate": "2018-04-26T17:57:51.650Z",
          "content": "<blockquote>\n  <p>your local cross validation score looks very high </p>\n</blockquote>\n\n<p>I concur.  To get 0.9817 I have a local validation of 0.9814.</p>",
          "rawMarkdown": "&gt; your local cross validation score looks very high \n\nI concur.  To get 0.9817 I have a local validation of 0.9814."
        },
        {
          "id": 319729,
          "postDate": "2018-04-26T18:10:52.207Z",
          "content": "<p>I think it depends on the validation size... </p>\n\n<p>May be he has very small validation size (compared to training data ) </p>",
          "rawMarkdown": "I think it depends on the validation size... \n\nMay be he has very small validation size (compared to training data ) "
        },
        {
          "id": 319732,
          "postDate": "2018-04-26T18:18:35.177Z",
          "content": "<p>@CPMP dang that val-lb is legiiiiiiit ;-) Does your val set remove hours not preset in test?</p>",
          "rawMarkdown": "@CPMP dang that val-lb is legiiiiiiit ;-) Does your val set remove hours not preset in test?",
          "votes": 1
        },
        {
          "id": 319743,
          "postDate": "2018-04-26T18:52:25.757Z",
          "content": "<p>That's day 9 hour 4 validation.  I also use validation on more data.</p>",
          "rawMarkdown": "That's day 9 hour 4 validation.  I also use validation on more data.",
          "votes": 1
        },
        {
          "id": 319748,
          "postDate": "2018-04-26T18:57:28.787Z",
          "content": "<p>@CPMP after validating, do you train on all data? And if yes how do you decide boosting rounds? <br>I use day 9 as a validation set, it scores 0.9818 and 0.9803 on LB. </p>",
          "rawMarkdown": "@CPMP after validating, do you train on all data? And if yes how do you decide boosting rounds? <br>I use day 9 as a validation set, it scores 0.9818 and 0.9803 on LB. "
        },
        {
          "id": 319754,
          "postDate": "2018-04-26T19:07:13.520Z",
          "content": "<p>Number of trees is one hyper parameter among others, its value should be found like the others: via local validation.  </p>\n\n<p>I train on all train data.</p>",
          "rawMarkdown": "Number of trees is one hyper parameter among others, its value should be found like the others: via local validation.  \n\nI train on all train data.",
          "votes": 1
        },
        {
          "id": 319764,
          "postDate": "2018-04-26T19:48:51.527Z",
          "content": "<p>Absolutely . Its this way : You have trained a model , now if i use it for predicting only 1 percent size of the train , the validation error maybe lower than train . Until there is a directional correlation between CV and LB , i think it is fine.</p>",
          "rawMarkdown": "Absolutely . Its this way : You have trained a model , now if i use it for predicting only 1 percent size of the train , the validation error maybe lower than train . Until there is a directional correlation between CV and LB , i think it is fine."
        },
        {
          "id": 319782,
          "postDate": "2018-04-26T20:24:45.453Z",
          "content": "<p>Hi@CPMP, may I ask if best number of trees when training on day 7 and day 8 is 100, then when using day7, 8, 9 as training set ,would you still use 100 as number of trees or use a larger number of trees?</p>",
          "rawMarkdown": "Hi@CPMP, may I ask if best number of trees when training on day 7 and day 8 is 100, then when using day7, 8, 9 as training set ,would you still use 100 as number of trees or use a larger number of trees?"
        },
        {
          "id": 319859,
          "postDate": "2018-04-27T01:44:27.433Z",
          "content": "<p>I'm using the last 2500000 rows for validation... As I see a correlation between local training score, validation score and the lb score I'm continuing with the same. The validation set is not used for training.</p>\n\n<p>Do you have any suggestion for the validation set? @CPMP, @Antoine, @Sohaib, @Snorlax</p>",
          "rawMarkdown": "I'm using the last 2500000 rows for validation... As I see a correlation between local training score, validation score and the lb score I'm continuing with the same. The validation set is not used for training.\n\nDo you have any suggestion for the validation set? @CPMP, @Antoine, @Sohaib, @Snorlax"
        },
        {
          "id": 319865,
          "postDate": "2018-04-27T01:54:13.087Z",
          "content": "<p>Hi@Samrat, my current strategy is using day 7 and 8 as train, and day 9 as validation. Currently, my training auc is about 0.987+, validation auc on day 9 hour 4 is 0.9803, validation auc on day 9 after hour 4 is 0.9836. Then train model on day 7,8,9, and corresponding public lb score is 0.9809</p>",
          "rawMarkdown": "Hi@Samrat, my current strategy is using day 7 and 8 as train, and day 9 as validation. Currently, my training auc is about 0.987+, validation auc on day 9 hour 4 is 0.9803, validation auc on day 9 after hour 4 is 0.9836. Then train model on day 7,8,9, and corresponding public lb score is 0.9809",
          "votes": 2
        },
        {
          "id": 319899,
          "postDate": "2018-04-27T04:18:04.993Z",
          "content": "<p><a href=\"/snorlax\">@snorlax</a>, yes, I use the same number of trees.  And your train/val auc numbers are similar to mine, except mine are a bit higher altogether. You're on the right track!</p>",
          "rawMarkdown": "@snorlax, yes, I use the same number of trees.  And your train/val auc numbers are similar to mine, except mine are a bit higher altogether. You're on the right track!",
          "votes": 1
        },
        {
          "id": 320363,
          "postDate": "2018-04-28T12:42:25.770Z",
          "content": "<p>Hi <a href=\"/snorlax\">@snorlax</a>! what do you mean by \"Then train model on day 7,8,9\"? You train without val set, with num_boost_round which you get when trained with val?)</p>",
          "rawMarkdown": "Hi @snorlax! what do you mean by \"Then train model on day 7,8,9\"? You train without val set, with num_boost_round which you get when trained with val?)"
        },
        {
          "id": 320972,
          "postDate": "2018-04-30T09:49:55.040Z",
          "content": "<p>Excuse me. CPMP. Do you use valid_set when use all data training? If not , how you use stop_round when predict?</p>",
          "rawMarkdown": "Excuse me. CPMP. Do you use valid_set when use all data training? If not , how you use stop_round when predict?"
        },
        {
          "id": 320993,
          "postDate": "2018-04-30T10:58:11.060Z",
          "content": "<p><a href=\"/zhixing\">@zhixing</a>, I answered that earlier in this thread.  </p>",
          "rawMarkdown": "@zhixing, I answered that earlier in this thread.  "
        }
      ]
    },
    {
      "id": 319112,
      "postDate": "2018-04-25T08:40:39.417Z",
      "content": "<p>single lgbm model 0.9794, 100m rows,  24 features </p>",
      "rawMarkdown": "single lgbm model 0.9794, 100m rows,  24 features "
    },
    {
      "id": 318632,
      "postDate": "2018-04-24T06:52:31.570Z",
      "content": "<h1>offtopic</h1>\n\n<p>can anyone explain how to utilize confRate features or use target encoding for test set. Thank you!</p>",
      "rawMarkdown": "#offtopic\ncan anyone explain how to utilize confRate features or use target encoding for test set. Thank you!",
      "replies": [
        {
          "id": 318645,
          "postDate": "2018-04-24T07:20:36.710Z",
          "content": "<p>@nickhillator </p>\n\n<p>Step 1 - Create target encoding for training data set. Say attribution ratio for each unique app \n             grp_app['Conf_feature'] = df.groupby(by = ['app']).is_attributed.mean()( Rename the column )</p>\n\n<p>Step 2 - Instead of deleting the group you can retain it to a new variable </p>\n\n<p>Step 3 - Create a different function for feature engineering for test and pass this group as a variable </p>\n\n<p>Step 4 - In the test pre processing function use this group to merge data into test. \n               Though test does not have attribute value, Step 4 will assign the Target encoding feature \n               calculated in train to corresponding records in test. \n               ie App 35 with an attribution ratio of say 0.71 was calculated for train \n                   Step 4 will simply assign this value to all records with app 35 in test. </p>\n\n<p>There might be a more memory efficient way to do this but this worked for me. \nHope this helps </p>\n\n<p>Regards\nShanth </p>",
          "rawMarkdown": "@nickhillator \n\nStep 1 - Create target encoding for training data set. Say attribution ratio for each unique app \n             grp_app['Conf_feature'] = df.groupby(by = ['app']).is_attributed.mean()( Rename the column )\n             \nStep 2 - Instead of deleting the group you can retain it to a new variable \n\nStep 3 - Create a different function for feature engineering for test and pass this group as a variable \n\nStep 4 - In the test pre processing function use this group to merge data into test. \n               Though test does not have attribute value, Step 4 will assign the Target encoding feature \n               calculated in train to corresponding records in test. \n               ie App 35 with an attribution ratio of say 0.71 was calculated for train \n                   Step 4 will simply assign this value to all records with app 35 in test. \n\nThere might be a more memory efficient way to do this but this worked for me. \nHope this helps \n\nRegards\nShanth \n             \n               \n"
        },
        {
          "id": 318648,
          "postDate": "2018-04-24T07:25:57.367Z",
          "content": "<p>so , it is better to apply it after splitting the training and test set, and cannot be applied in the merged set, okay thanks! is there any kernel apart from <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">https://www.kaggle.com/nanomathias/feature-engineering-importance-testing</a> for the same</p>",
          "rawMarkdown": "so , it is better to apply it after splitting the training and test set, and cannot be applied in the merged set, okay thanks! is there any kernel apart from https://www.kaggle.com/nanomathias/feature-engineering-importance-testing for the same"
        },
        {
          "id": 318649,
          "postDate": "2018-04-24T07:30:50.380Z",
          "content": "<p>Haven't seen other kernels.  I guess you could do it with merging as well. You can save memory by not carrying the groups </p>\n\n<ol>\n<li>Merge train and test to DF </li>\n<li>Identify train records alone, create groups for unique app and calculate target encoding </li>\n<li>Identify test records from DF and merge the groups to these test records </li>\n</ol>\n\n<p>Same method, but this way you can avoid carrying the variables around. </p>\n\n<p>Regards\nShanth</p>",
          "rawMarkdown": "Haven't seen other kernels.  I guess you could do it with merging as well. You can save memory by not carrying the groups \n\n1. Merge train and test to DF \n2. Identify train records alone, create groups for unique app and calculate target encoding \n3. Identify test records from DF and merge the groups to these test records \n\nSame method, but this way you can avoid carrying the variables around. \n\nRegards\nShanth",
          "votes": 2
        },
        {
          "id": 318715,
          "postDate": "2018-04-24T10:17:14.830Z",
          "content": "<p>have the attribution features helped you, are they important features in achieving results?</p>",
          "rawMarkdown": "have the attribution features helped you, are they important features in achieving results?"
        },
        {
          "id": 318751,
          "postDate": "2018-04-24T12:10:39.257Z",
          "content": "<p>CPMP gave a very good explanation of how to do target encoding <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55147#317893\">here</a>. I am currently playing with this idea and so far I have not had much luck with these \"frequency\" features. I might be missing something but so far they do not seem to be helpful at all.</p>",
          "rawMarkdown": "CPMP gave a very good explanation of how to do target encoding [here][1]. I am currently playing with this idea and so far I have not had much luck with these \"frequency\" features. I might be missing something but so far they do not seem to be helpful at all.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55147#317893"
        },
        {
          "id": 318754,
          "postDate": "2018-04-24T12:15:47.547Z",
          "content": "<p>some of them have been able to go beyond 0.98 with less than 15 features, i don't know how to use the frequency features anymore, the intution isn't helping, when the competition ends and they reveal their features, it might be a jaw dropping moment. \nI am trying my luck with time delta and target encoding now</p>",
          "rawMarkdown": "some of them have been able to go beyond 0.98 with less than 15 features, i don't know how to use the frequency features anymore, the intution isn't helping, when the competition ends and they reveal their features, it might be a jaw dropping moment. \nI am trying my luck with time delta and target encoding now"
        },
        {
          "id": 319971,
          "postDate": "2018-04-27T07:52:33.853Z",
          "content": "<p>@ Nickhillator. Congrats on your new rank. :) Did the target encoding help?  </p>",
          "rawMarkdown": "@ Nickhillator. Congrats on your new rank. :) Did the target encoding help?  "
        },
        {
          "id": 322470,
          "postDate": "2018-05-03T02:51:59.913Z",
          "content": "<p>@shanth sorry for the late reply, i haven't used yet, i am trying to use the best validation strategy as of now, and fiddle a bit with feature engineering</p>",
          "rawMarkdown": "@shanth sorry for the late reply, i haven't used yet, i am trying to use the best validation strategy as of now, and fiddle a bit with feature engineering"
        }
      ]
    },
    {
      "id": 318255,
      "postDate": "2018-04-23T13:31:25.510Z",
      "content": "<p>@CuteChibiko -\nGreat to see NN is performing well.  Are u using vanilla NN architecture or some sort of CNN/LSTM model.?</p>",
      "rawMarkdown": "@CuteChibiko -\nGreat to see NN is performing well.  Are u using vanilla NN architecture or some sort of CNN/LSTM model.?\n",
      "replies": [
        {
          "id": 318259,
          "postDate": "2018-04-23T13:43:06.913Z",
          "content": "<p>Hi SubikashPal, Thank you for asking, but I'd like to share you the detail after this competition ends. ;)</p>",
          "rawMarkdown": "Hi SubikashPal, Thank you for asking, but I'd like to share you the detail after this competition ends. ;)",
          "votes": 1
        },
        {
          "id": 318262,
          "postDate": "2018-04-23T13:52:48.170Z",
          "content": "<p>Nop. I understand. Even after the end , I would love to see it. I am really interested.</p>",
          "rawMarkdown": "Nop. I understand. Even after the end , I would love to see it. I am really interested."
        },
        {
          "id": 319967,
          "postDate": "2018-04-27T07:46:41.420Z",
          "content": "<p>@ CuteChibiko - Would love to see your NN after the competition. Good luck for the rest of the competition.</p>",
          "rawMarkdown": "@ CuteChibiko - Would love to see your NN after the competition. Good luck for the rest of the competition."
        }
      ]
    },
    {
      "id": 315687,
      "postDate": "2018-04-17T13:31:45.767Z",
      "content": "<h1>offtopic</h1>\n\n<p>how do you guys deal with the NaN values created while using time delta features, i.e the first iterations of each case</p>",
      "rawMarkdown": "#offtopic\nhow do you guys deal with the NaN values created while using time delta features, i.e the first iterations of each case",
      "replies": [
        {
          "id": 317047,
          "postDate": "2018-04-20T15:10:19.337Z",
          "content": "<p>@ nickhillator </p>\n\n<p>Listing out a few options</p>\n\n<ol>\n<li>Leaving NAN values as they are </li>\n<li>For every group, replace the NAN record with the avg time_delta value for that group. To me it sounds like the most logical thing to </li>\n<li>As an easier option, replace it with a global average which may skew the value a bit. </li>\n</ol>\n\n<p>Option 2 would be the best bet. I haven't explored it for lack of time. </p>\n\n<p>Cheers \nShanth</p>",
          "rawMarkdown": "@ nickhillator \n\nListing out a few options\n\n1. Leaving NAN values as they are \n2. For every group, replace the NAN record with the avg time_delta value for that group. To me it sounds like the most logical thing to \n3. As an easier option, replace it with a global average which may skew the value a bit. \n\nOption 2 would be the best bet. I haven't explored it for lack of time. \n\nCheers \nShanth",
          "votes": 1
        },
        {
          "id": 319382,
          "postDate": "2018-04-25T23:40:23.360Z",
          "content": "<p>Since LGBM is tree based, i simply replaced them with -9999 and the algorithm appears to ignore them. this might be totally wrong, but works for me =)</p>",
          "rawMarkdown": "Since LGBM is tree based, i simply replaced them with -9999 and the algorithm appears to ignore them. this might be totally wrong, but works for me =)"
        },
        {
          "id": 319484,
          "postDate": "2018-04-26T06:55:27.087Z",
          "content": "<p>I leave them as NaN.  lgb and xgb handle NaN.  You have to replace NaN if you use general other algorithms, for instance NNs, of factorization machines.  I have not yet looked into that (maybe I should start, time is flying...)</p>",
          "rawMarkdown": "I leave them as NaN.  lgb and xgb handle NaN.  You have to replace NaN if you use general other algorithms, for instance NNs, of factorization machines.  I have not yet looked into that (maybe I should start, time is flying...)",
          "votes": 1
        }
      ]
    },
    {
      "id": 312340,
      "postDate": "2018-04-11T15:42:37.150Z",
      "content": "<p>I'm having problems getting a better score than .9687 with a single model. Do you have any suggestions?\nI noticed that most of the people are using LGBM model so I'm doing the same. Probably I can improve my score using an other aproach in the feature engineering step.</p>",
      "rawMarkdown": "I'm having problems getting a better score than .9687 with a single model. Do you have any suggestions?\nI noticed that most of the people are using LGBM model so I'm doing the same. Probably I can improve my score using an other aproach in the feature engineering step.",
      "replies": [
        {
          "id": 312348,
          "postDate": "2018-04-11T15:46:10.677Z",
          "content": "<p>Something does not add up... Last time I checked, the top LB score was 0.9821, so I doubt that the score you provided is correct.</p>",
          "rawMarkdown": "Something does not add up... Last time I checked, the top LB score was 0.9821, so I doubt that the score you provided is correct.",
          "votes": 1
        },
        {
          "id": 312351,
          "postDate": "2018-04-11T15:49:38.930Z",
          "content": "<p>he might be referring to his local validation score</p>",
          "rawMarkdown": "he might be referring to his local validation score"
        },
        {
          "id": 312362,
          "postDate": "2018-04-11T16:20:49.623Z",
          "content": "<p>@nickhillator He might be. Even in this case, the score provided is not very informative without knowing his validation procedure.</p>",
          "rawMarkdown": "@nickhillator He might be. Even in this case, the score provided is not very informative without knowing his validation procedure."
        },
        {
          "id": 312455,
          "postDate": "2018-04-11T19:26:41.453Z",
          "content": "<p>Sorry it was .9687 </p>",
          "rawMarkdown": "Sorry it was .9687 "
        }
      ]
    },
    {
      "id": 312126,
      "postDate": "2018-04-11T08:34:57.817Z",
      "content": "<p>are all the top guys using time deltas in feature engineering, can anyone confirm if they achieved over 0.9710 LB without it?</p>",
      "rawMarkdown": "are all the top guys using time deltas in feature engineering, can anyone confirm if they achieved over 0.9710 LB without it?",
      "replies": [
        {
          "id": 312229,
          "postDate": "2018-04-11T12:31:02.917Z",
          "content": "<p>Yes, my current score (0.9734) is without time delta features, and <a href=\"https://www.kaggle.com/krishnakesavan/r-lgbm-single-model-40m-rows-lb-0-9736\">also see this kernel</a></p>",
          "rawMarkdown": "Yes, my current score (0.9734) is without time delta features, and [also see this kernel][1]\n\n\n  [1]: https://www.kaggle.com/krishnakesavan/r-lgbm-single-model-40m-rows-lb-0-9736"
        },
        {
          "id": 312232,
          "postDate": "2018-04-11T12:36:37.477Z",
          "content": "<p>I got 0.973x without time delta as well.</p>",
          "rawMarkdown": "I got 0.973x without time delta as well.",
          "votes": 1
        },
        {
          "id": 312235,
          "postDate": "2018-04-11T12:41:36.043Z",
          "content": "<p>Thanks for the link!</p>",
          "rawMarkdown": "Thanks for the link!"
        }
      ]
    },
    {
      "id": 311974,
      "postDate": "2018-04-11T02:21:01.770Z",
      "content": "<p>Best single LGB hits 0.9712, 12 features with half of the training set and a random split (0.1) for validation :P</p>",
      "rawMarkdown": "Best single LGB hits 0.9712, 12 features with half of the training set and a random split (0.1) for validation :P"
    },
    {
      "id": 311718,
      "postDate": "2018-04-10T15:07:22.380Z",
      "content": "<p>What do you mean by single model ?</p>",
      "rawMarkdown": "What do you mean by single model ?",
      "replies": [
        {
          "id": 311875,
          "postDate": "2018-04-10T20:21:46.640Z",
          "content": "<p>@ Nathan Single model means say a Single LGBM trained model is used to predict output. \nAs opposed to say using two LGBM models with totally different configurations to make predictions for the entire test set and then taking an average to be used for submission. ( Also popularly called as Blending/ Ensemble of models )</p>\n\n<p>Hope I answered your question. </p>",
          "rawMarkdown": "@ Nathan Single model means say a Single LGBM trained model is used to predict output. \nAs opposed to say using two LGBM models with totally different configurations to make predictions for the entire test set and then taking an average to be used for submission. ( Also popularly called as Blending/ Ensemble of models )\n\nHope I answered your question. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 310949,
      "postDate": "2018-04-09T04:47:41.373Z",
      "content": "<p>Some one try to use test_supplement?</p>",
      "rawMarkdown": "Some one try to use test_supplement?",
      "replies": [
        {
          "id": 310950,
          "postDate": "2018-04-09T04:54:11.633Z",
          "content": "<p>Suppose everybody :)</p>",
          "rawMarkdown": "Suppose everybody :)"
        },
        {
          "id": 311381,
          "postDate": "2018-04-10T00:21:03.613Z",
          "content": "<p>I'm still trying to understand - what can we use test_supplement for? It doesn't have 'attributed' value so we can't use it as extra training data right? Or am I missing something</p>",
          "rawMarkdown": "I'm still trying to understand - what can we use test_supplement for? It doesn't have 'attributed' value so we can't use it as extra training data right? Or am I missing something"
        },
        {
          "id": 311462,
          "postDate": "2018-04-10T05:38:19.110Z",
          "content": "<p>For example for crating new features</p>",
          "rawMarkdown": "For example for crating new features"
        },
        {
          "id": 311686,
          "postDate": "2018-04-10T14:00:20.197Z",
          "content": "<p>You can use test_supplement, for example, to aggregate features. For example, if you want to count number of clicks for a given IP address in a certain hour, the official test file will not give the right count for some hours, because it has only part of the hour, so you can use test_supplement.  Or if you want calculate the time since the last click by that IP address, test_supplement allows you to calculate a value for the first records of test.</p>",
          "rawMarkdown": "You can use test_supplement, for example, to aggregate features. For example, if you want to count number of clicks for a given IP address in a certain hour, the official test file will not give the right count for some hours, because it has only part of the hour, so you can use test_supplement.  Or if you want calculate the time since the last click by that IP address, test_supplement allows you to calculate a value for the first records of test.",
          "votes": 2
        },
        {
          "id": 311776,
          "postDate": "2018-04-10T16:50:19.037Z",
          "content": "<p>@Tetyana - I looked at Test supplement hoping to find IP's that are found in train ( basic test set as you may know has many new ip's compared to train ). This way the IP's can be used for better target encoding, but the IP's in test supplement is still different from that of train. </p>",
          "rawMarkdown": "@Tetyana - I looked at Test supplement hoping to find IP's that are found in train ( basic test set as you may know has many new ip's compared to train ). This way the IP's can be used for better target encoding, but the IP's in test supplement is still different from that of train. "
        },
        {
          "id": 315610,
          "postDate": "2018-04-17T10:44:11.380Z",
          "content": "<p>@ Tetyana </p>\n\n<p>Are you using test supplement in your current LB score?</p>",
          "rawMarkdown": "@ Tetyana \n\nAre you using test supplement in your current LB score?"
        },
        {
          "id": 317634,
          "postDate": "2018-04-22T04:03:28.967Z",
          "content": "<p>Not</p>",
          "rawMarkdown": "Not"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 317359,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2018-04-21T09:28:12.150000",
      "content": "<p>Single NN model, LB 0.9814</p>",
      "votes": 20,
      "replies": [
        {
          "id": 317368,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-21T09:44:36.477000",
          "content": "<p>Thanks, very encouraging!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317388,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-21T11:44:47.780000",
          "content": "<p>Good to see a NN model....  Do you mind telling us the number of new features.... a single digit or a dozen of them.... :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317394,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2018-04-21T12:14:34.597000",
          "content": "<p>Hi Samrat, its a single digit.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 317411,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-21T13:14:22.283000",
          "content": "<p>Great!!!! Thanks for the reply....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317621,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-04-22T02:37:58.297000",
          "content": "<p>GREAT!!!!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318015,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-23T03:25:11.437000",
          "content": "<p>owesome</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318256,
          "author_name": "SubikashPal",
          "author_url": "",
          "post_date": "2018-04-23T13:35:03.333000",
          "content": "<p>Great to see that NN is performing so well. Are you using simple NN setup or combination of CNN/FC?\nCrossing 0.9750 using NN becomes a nightmare for me :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318373,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-23T17:37:37.857000",
          "content": "<p>Very nice. Looking forward to your kernel after the competition ends :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 314699,
      "author_name": "Sameh Faidi",
      "author_url": "",
      "post_date": "2018-04-16T06:06:16.850000",
      "content": "<p>LB 0.9815 with single model and ~50 features and using all the train data</p>",
      "votes": 17,
      "replies": [
        {
          "id": 314700,
          "author_name": "Tetyana Yatsenko",
          "author_url": "",
          "post_date": "2018-04-16T06:08:18.903000",
          "content": "<p>It is really wonderfull</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315009,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-16T16:27:40.190000",
          "content": "<p>@ Sameh....50 features is awesome. Are you hooked to a super computer? :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315082,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-04-16T18:20:18.960000",
          "content": "<p>i didnt measure (yet) the effect of including all the train data compared to only 1 day. but i assume the gain is not major. will report when i try it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315096,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-16T18:36:34.413000",
          "content": "<p>Same assumption here. Now my score of 0.9805 is still based on day 9 with multiple groups of features and bagging. I'm still running something on day 9  and I think the limit of using a single day will be close to 0.981. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315501,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-17T07:29:37.077000",
          "content": "<p>Hi Sameh,\nWhen you use ALL train data? do you not use CV?\nif so - when do you STOP training to make sure you are not overfitting?</p>\n\n<p>Thanks! =)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315506,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-17T07:36:00.257000",
          "content": "<p>@AmirH I guess it makes sense to first validate your model on day 9 and then train your model including day 9 with same features and parameters.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315580,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-04-17T09:58:28.373000",
          "content": "<p>@Sohaib</p>\n\n<p>Instead of retraining the model on all the data, why not just using the prediction for day 9 as OOF for stacking purpose ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315711,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-17T14:06:04.220000",
          "content": "<p>Yes, we can do that too. I will share stacking vs averaging(models trained including day 9) results with you once I am done with my single model. <br> EDIT: For me including day 9 in training increase AUC</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315915,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-17T19:41:02.293000",
          "content": "<p>@congb if feasible, can you give an example of bagging feature?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 316019,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-18T02:54:05.993000",
          "content": "<p>I’m just creating different sets of features on the same data set and bagging the results. Nothing fancy. Just a compromise because of the limited resource I have. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316414,
          "author_name": "Perfect Is Shit",
          "author_url": "",
          "post_date": "2018-04-19T01:37:43.657000",
          "content": "<p>You mean retrain the model with the same parameters and features you used in validation, but when will you stop in retraining(like early_stopping stops when it doesn't improve) Thanks. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316416,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-19T01:39:29.540000",
          "content": "<p>@bgm, yep I do use early stopping </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 316419,
          "author_name": "Perfect Is Shit",
          "author_url": "",
          "post_date": "2018-04-19T01:45:28.213000",
          "content": "<p>OK thanks! So you should use validation in retraining, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317140,
          "author_name": "richinmind",
          "author_url": "",
          "post_date": "2018-04-20T21:16:32.910000",
          "content": "<p>@Sohaib Omar</p>\n\n<p>You mentioned that \"I guess it makes sense to first validate your model on day 9 and then train your model including day 9 with same features and parameters.\"</p>\n\n<p>I don't understand how to do it. Could you please explain the statement with some simple Python code.</p>\n\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317207,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-21T01:47:59.153000",
          "content": "<pre><code>train = day7_8\nval = day9\nmodel.fit(train)\nval_score = model.predict(day9)\nif va_score &gt; my_best_val_score:\n    model.fit(train + val)\n</code></pre>\n\n<p>Hope it helps.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 317636,
          "author_name": "Perfect Is Shit",
          "author_url": "",
          "post_date": "2018-04-22T04:17:22.387000",
          "content": "<p>thanks @Sohaib</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 308329,
      "author_name": "Alexander Firsov",
      "author_url": "",
      "post_date": "2018-04-03T11:03:39.880000",
      "content": "<p>Single model LB 0.9756, 18 features including original columns and frequencies.</p>\n\n<p>Additional info: IP and ip frequencies were not used. Training was done on day 8. \nLocally I have similar results on day 7.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 308996,
          "author_name": "Rohit Mehra",
          "author_url": "",
          "post_date": "2018-04-04T13:26:51.660000",
          "content": "<p>Thanks, Alexander. Have been doing feature engineering for past two days, now have upto 20 (freq and groupby) I was actually worried about IP and Ip derived features, thanks for clearing that out, and also the training strategy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309107,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-04T16:52:18.410000",
          "content": "<p>Impressive results with just 18 features and 1 training day!</p>\n\n<p>Do you mind if I ask - when you say that ip frequencies weren't used, do you mean that you don't use any sort of groupby-count features derived from ip? For example, you don't use something like # of ip clicks in the hour (a feature commonly seen in the kernels)? </p>\n\n<p>Of course I understand if you don't want to share more :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 309122,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-04T17:23:10.147000",
          "content": "<p>Same question with Joe, but it doesn't matter if you dont want to answer, we understand that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309126,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2018-04-04T17:27:18.470000",
          "content": "<p>I do not think this is impressive, I was lucky with some feature engineering early, but currently I am stuck. \nRegarding your question, I am using grouping of several columns including IP. I just do not use IP directly and IP frequencies (percent of is_attributed=1 for IP). There are several topics warning about using IP as a feature, so I decided not to use it (I may reconsider this later). \nThe problem here is that IP column has a lot of info, but test does not have all IPs that we know from train. I beleive it is crucial to extract info using IP without actually using IP address itself.</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 309134,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-04T17:37:45.993000",
          "content": "<p>Ok understood, thanks! That is also what I've been doing, and most of my current features are derived from groupings on ip. I definitely agree with you re: finding ways to extract info without using the explicit address. </p>\n\n<p>I'm also in the same boat, stuck on FE progress.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 309139,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-04T17:53:15.250000",
          "content": "<p>Thank you very much Alex! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309168,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-04-04T19:01:02.317000",
          "content": "<p>@Alexander Firsov: Do you mind sharing why you choosed to train on only day 8 ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309206,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2018-04-04T19:50:48",
          "content": "<p>I described my single best model, and it was from day 8, but I have similar result from day 7.\nI do not plan to train only on day 8. \nWhat I want to avoid is an effect when model works good for data from next hour, but not next day, that is why I started with first two days.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309326,
          "author_name": "Tetyana Yatsenko",
          "author_url": "",
          "post_date": "2018-04-05T03:59:32.890000",
          "content": "<p>So great</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309829,
          "author_name": "Badr Ouali",
          "author_url": "",
          "post_date": "2018-04-06T03:12:02.763000",
          "content": "<p>@Alexander Firsov, thank you for sharing your method but what do you call frequency? Is the following variable avg(is_attributed) over (partition by feature1,feature2...) a frequency? Thank you by advance for your answer</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310736,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-08T12:36:26.700000",
          "content": "<p>Yes. It's commonly known as mean-encoding: <a href=\"https://www.coursera.org/learn/competitive-data-science/lecture/b5Gxv/concept-of-mean-encoding\">https://www.coursera.org/learn/competitive-data-science/lecture/b5Gxv/concept-of-mean-encoding</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 310746,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2018-04-08T13:25:16.723000",
          "content": "<p>Authman, thanks for the reference. I used the term from <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">popular kernel</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 318823,
      "author_name": "Cong",
      "author_url": "",
      "post_date": "2018-04-24T14:45:53.010000",
      "content": "<p>update: single model with 11 features trained on day 9 only (~60M rows) get to .9802</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 313326,
      "author_name": "Cong",
      "author_url": "",
      "post_date": "2018-04-13T05:44:14.783000",
      "content": "<p>My current lb score is a single model. 0.9784 with 7 additional features - 2 groupby count features, 2 delta-time features, 3 confRate features.  Trained on day 9 with 0.1 validation. </p>\n\n<p>Edit: the new score of 0.9796 is also from a single model using 3 more features, still trained on day 9 only. </p>\n\n<p>Edit: 0.9798, single model with 18 features, trained on day 9 only - cant go any further with my 32GB ram... will try something else next.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 313332,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-04-13T05:52:04.893000",
          "content": "<p>well done. What does confRate mean here? do you mean target encoding related features?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313344,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-13T06:09:57.073000",
          "content": "<p>confidence rate - see <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">the great FE kernel</a>.  But this is quite tricky because it is essentially just leaking. Some of the confRate features I tested actually worsen the result a lot (to sub 0.95).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313373,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-04-13T07:11:16.230000",
          "content": "<p>Great! And what does delta-time mean? May be you can suggest any kernel... Thanx!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313545,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-13T12:56:23.820000",
          "content": "<p>Idea is from <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361\">this thread</a>. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 313554,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-13T13:24:30.440000",
          "content": "<p>@congb, You're very kind to share that much.  </p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 313587,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-13T14:17:18.020000",
          "content": "<p>@congb, thank you very much for your sharing! And I would like to discuss some questions with you guys here:  Is using day 9 as train and 0.1 as validation  will overfit to public lb?  If you guys remember that 0.9736 public kernel a few days ago, this kernel uses last N million rows and when other kagglers reproduce this kernel with more rows, the public lb score decreases. In terms of my experiments, I use day 7 and 8 as train and day 9 as validation, and I found only one of confRate features actually improves validation scores... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313591,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-13T14:31:41.600000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313593,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-13T14:34:37.757000",
          "content": "<p>@CPMP, No problem! I learned a lot from your input in multiple threads. </p>\n\n<p>@Snorlax, The reason I'm using only day 9 is for faster modeling (need to go back and forth tuning things). But I believe that the one you described is a better strategy and for final submission I will definitely do that. And I'm sure my .9784 submission is an overfitting to the public lb, to some extent. confRate features are tricky and I did spent quite some time to decide which ones to include - but still, I'm also very unconfident with the confidence rate...</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 313594,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-13T14:36:16.430000",
          "content": "<p>Hi @Zijun Yao, thank you very much for your reply. Actually, I only found one confRate feature is useful(but not that useful) and I've looked through almost all kernels using confRate and I think they could achieve the same public lb score without using confRate based on what other features they already have.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313599,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-13T14:43:52.890000",
          "content": "<p>@Zijun Yao, may I ask your validation score on day 9 4:00 to 5:00 and validation score on day 9 after 5:00?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313603,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-13T14:50:10.957000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 313604,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-13T14:57:31.863000",
          "content": "<p>@Zijun Yao, I mean your best model's validation score without confRate. In terms of my model, it's validation score on data after day 9 5:00 is 0.003 higher than the validation score on data between day 9 4:00 to 5:00.  I ask this because data on day 9 4:00 to 5:00 mimics public lb data and data after day 9 5:00 mimics private lb data. And I really want to know what is the gap between validation score on 'simulated' public lb data and 'simulated' private lb data. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313606,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-13T15:01:20.437000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 313608,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-13T15:04:34.600000",
          "content": "<p>@Zijun Yao, Got it! Thank you so much! And thank you again for your sharing @congb! I'm very lucky to meet you guys who generously share your valuable experience. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314991,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-16T15:56:33.430000",
          "content": "<p>@Snorlax are you using the entire day 9 for validation?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 314994,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-16T15:58:25.983000",
          "content": "<p>9/1 training/val split</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315015,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-16T16:32:19.397000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315021,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-16T16:37:26.377000",
          "content": "<p>I would agree with you on both points - Validation strategy and Conf rates. Day 9 training only is definitely a strong overfit. I like Pranav's latest kernel where he trains all the data in different chunks and averages the predictions. But of course, nothing like it If you can train it all in one go with enough resources. </p>\n\n<p>Did you train 7 and 8 separately and average predictions from two models or Day 7 first and use the same fit model on Day 8 ( If that is even a relevant approach )?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315080,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-16T18:12:06.977000",
          "content": "<p>@ Shanth， I'm still doing preprocesing and feature generation for day 7. But combining preds from day 8 and day 9  does help me a little （~0.0005)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 320603,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "2018-04-29T08:31:21.047000",
      "content": "<p>0.9822 with full data,60 features,lgbm</p>",
      "votes": 9,
      "replies": [
        {
          "id": 320618,
          "author_name": "SubikashPal",
          "author_url": "",
          "post_date": "2018-04-29T09:32:02.303000",
          "content": "<p>are you using target encoding?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320620,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-04-29T09:52:19.110000",
          "content": "<p>not yet,will try </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320674,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-29T13:19:52.317000",
          "content": "<p>Thanks, very encouraging.</p>\n\n<p>I'm at 0.9820 lgb with 33 features.  I only have to add 27 features and buy some RAM to meet you ;)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 320689,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2018-04-29T14:05:01.800000",
          "content": "<p>0.9811 lgb with 14 features. Unfortunately, cannot improve anymore with other features..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320738,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-29T16:51:27.627000",
          "content": "<p>May I ask how you guys are iterating with different features? For instance - are you first thinking of a feature, coding it, adding it to your main model, and then checking the validation auc? Or do you test out multiple features at a time, etc.?</p>\n\n<p>I've noticed that when I add in different features, other features suddenly don't seem to work as well. Based on this, I'm thinking that the interaction between all the features you have in your model is crucial, and, as a result, really the best way to figure out if your model performs well is constantly iterating on the main model. Am I going about this correctly?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 320759,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-29T18:16:56.847000",
          "content": "<p>I mostly do what you write:</p>\n\n<blockquote>\n  <p>are you first thinking of a feature, coding it, adding it to your main model, and then checking the validation auc? </p>\n</blockquote>\n\n<p>But sometimes adding more than one feature at a time works, while adding each of them degrades local cv.  This is a bit of black magic ;)</p>\n\n<p>An other way is to add a bunch of features, then try removing them one by one to see if they were useful.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 320781,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-04-29T20:18:17.303000",
          "content": "<blockquote>\n  <p><strong>CPMP wrote</strong></p>\n  \n  <p>An other way is to add a bunch of features, then try removing them one by one to see if they were useful.</p>\n</blockquote>\n\n<p>That's the approach I use</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 320858,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-30T03:26:03.430000",
          "content": "<p>@CPMP @Serigne Thanks! Quick question - how would you interpret it if you train a model with a group of, say, 5 features and they seem to perform well according to validation auc. But then when you add them to your main model, the validation auc of the main model decreases slightly (like .0002) from where it was without those extra 5 features. Does that mean those 5 features are just bad or is it ok to still keep them in the model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320874,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-04-30T04:31:46.057000",
          "content": "<p>@CPMP I think you can get 0.9823 with 40 features.I don't want to check my features one by one,so just left them.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320876,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-04-30T04:38:32.330000",
          "content": "<p>@Edward Chen  At an earlier stage,I can try features one by one using sample data(only 12 hours of 9th).Now every time I add 3 ~ 5 new features to main model to see the val score,if it got better I will submit .I found public learderboard score always match val score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320877,
          "author_name": "dingxianghu",
          "author_url": "",
          "post_date": "2018-04-30T04:43:31.973000",
          "content": "<p>ecuse me, can you share how to setup your validation set</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320879,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-30T04:49:27.517000",
          "content": "<p><a href=\"/senkin13\">@senkin13</a> interesting. You mentioned that you add 3-5 new features at a time. What happens if it doesn't improve val score though? Do you remove them?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320902,
          "author_name": "Laevatein",
          "author_url": "",
          "post_date": "2018-04-30T06:35:45",
          "content": "<p>Well, this magic is driving me crazy...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320931,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-30T08:34:06.597000",
          "content": "<p>Hello senkin13. I wonder know what's your learning rate and model\"s runing time. Glad to see your reply. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320942,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-30T08:53:05.113000",
          "content": "<p>@EdwardChen, I test features by adding them to the model.  That's the only test that matters.  I only have one model.  Not sure why you speak of a 'main' model here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320944,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-30T08:55:34.787000",
          "content": "<p><a href=\"/senkin13\">@senkin13</a>, I agree.  </p>\n\n<p>I cannot afford 50 features I think, because of RAM.  That's why I perform feature elimination.  And yes, I hope to get to 0.9823 with less than 40 features.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320949,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-04-30T08:59:51.097000",
          "content": "<p>random split(0.95) for train validation set.\nif new features didn't improve val score,will remove them expect they have meaningful buisness sense.\nlower learning rate got better score, 0.02=0.03&gt;0.05&gt;0.1.\nafter removing dtrain from lgbm parameter [valid_sets],can reduce running time from 6,7 hours to 2,3hours</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 320950,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-30T09:02:12.413000",
          "content": "<p>Hello CPMP. I cannot afford with memory, too.  I wonder know what's your learning rate and model\"s runing time. Glad to see your reply.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320952,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-30T09:08:15.947000",
          "content": "<p>Thanks senkin13. I was troubled by the time of training.Will you train two times when you add new features? First you determine the number of iterations and second use full data train.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320953,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-30T09:08:35.680000",
          "content": "<p>I didn't tune learning rate, I am using 0.1.  I keep these tuning for when I'll run out of fuel with feature engineering.  A larger learning rate means faster training, hence faster experiments.</p>\n\n<p>I don't do tuning often, because when you do you no longer can compare with previous runs.  Now I have a good setting, my validation score matches exactly my LB score since I reached 0.9818.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 320954,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-30T09:11:40.873000",
          "content": "<p>Thanks CPMP. I always learn a lot from you. :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 320975,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-30T09:59:28.957000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321032,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-30T12:45:04.383000",
          "content": "<p>@CPMP Thanks for the reply and sorry for the confusion. Please ignore what I said about the \"main model\". Basically, I was just asking - if you're testing some features in your model and they don't happen to improve the val score, do you keep them in the model or hope that some interactions with future features will help to improve the val score with those features but I believe @senkin has answered that - so thanks as well @senkin</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321089,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-30T15:03:48.700000",
          "content": "<p>@Edward Chen, no need to apologize, no harm done at all!</p>\n\n<p>If the features don't improve my model I discard them.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 321494,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-05-01T12:21:56.147000",
          "content": "<p>@CPMP : with a learning rate  of 0.1 it probably takes you like more than 1 or 2 hours to fit your model on the whole training set. So are you doing this feature selection process on the full training set ? Or do you create a smaller dataset that is correlated to it ?</p>\n\n<p>Thanks for all the advices !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321509,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-01T13:09:03.573000",
          "content": "<blockquote>\n  <p>are you doing this feature selection process on the full training set ? </p>\n</blockquote>\n\n<p>@Antoine, I don't.  First, you need some validation data, hence you cannot use all train data for training.  Second, using a subset is faster.  Key is to use a representative subset, so that improvement on it yields improvement when training on all the dataset.  That's what I focused on in the early days of the competition: be able to run reliable experiments as fast as possible.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 322104,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-02T12:25:35.960000",
          "content": "<blockquote>\n  <p>@CPMP I think you can get 0.9823 with 40 features.</p>\n</blockquote>\n\n<p><a href=\"/senkin13\">@senkin13</a>, you are right, just got there ;)</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 322131,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-05-02T13:00:24.940000",
          "content": "<p>@CPMP : Merci ! I guess my main problem during this competition was to find a correct training / validation set for testing, and one for submitting my score. I guess with experience it becomes easier...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 322134,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-02T13:07:11.407000",
          "content": "<p>@Antoine, de rien!</p>\n\n<p>Having a reliable local validation setting is the single most important thing in a competition.  My best results were obtained when I found a good CV setting early in the competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 322194,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2018-05-02T14:29:43.343000",
          "content": "<p>@CPMP congrats!I think there are still many good features to mine,will you continue to improve single model or make others to do ensembling</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322201,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-02T14:40:19.400000",
          "content": "<p><a href=\"/senkin13\">@senkin13</a>, Thanks!  I think I'll move to ensembling as I am reaching the limits of my HW.  I wish I had bought that extra memory I thought about...  I can extend my swap, but the code is already quite slow, paging will kill it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322377,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-02T20:36:04.593000",
          "content": "<blockquote>\n  <p>@CPMP: I cannot afford 50 features I think, because of RAM.</p>\n</blockquote>\n\n<p>How much ram do you have? I'm in the process of running my largest model (31 features) but:</p>\n\n<pre><code>➜  ~ free -g\n              total        used        free      shared  buff/cache   available\nMem:             62          16          45           0           0          45\nSwap:           127          21         106\n</code></pre>\n\n<p>Only takes 16GB of ram during training. Are you talking about RAM usage during lgb dataset conversion? If so, there's a thread out there which instructs to use <code>'two_round':True</code> lgb_config parameter, so your data is automatically written from pd-&gt;disk, then read from disk-&gt;lgb... rather than pandas-&gt;float64-&gt;lgb. On the other hand, if you're limited by FE, you can always write your features to disk intermediary while building them, then recompile. For example, I current build all my count features, write to disk then drop from the df, clickdelta features, write to disk then drop from the df, etc. Then reload all the feathers to reassemble.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 322395,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-02T21:11:21.190000",
          "content": "<blockquote>\n  <p>Then reload all the feathers to reassemble.</p>\n</blockquote>\n\n<p>That's what I am doing ;)</p>\n\n<p>I'll try the two rounds, thanks for the tip.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322459,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-05-03T01:53:33.073000",
          "content": "<p>@CPMP I'm using a similar approach but with csv... But I see that the merging of this data to the original df is talking lot of time... I mean hours... Is there an efficient way to do that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322494,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-03T04:50:50.373000",
          "content": "<p>If you save each feature as a single column cssv, then you don't need to use merge, you simply copy the values.  You need to merge if you change the order of the rows.  See the advice given by <a href=\"/spongebob\">@spongebob</a>: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55601\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55601</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 322497,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-03T04:56:40.437000",
          "content": "<p>You can merge multiple columns using that same technique:</p>\n\n<pre><code>load = pd.read_feather('./new_features.ftr')\ndf[load.columns] = load\n</code></pre>\n\n<p>Merging in 20 gb of column data should be doable in 5-20 secs. Def shouldn't take minutes, or--God forbid--hours.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 322508,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-03T05:27:57.157000",
          "content": "<p>@authan, you are right, this is how it should be done.  </p>\n\n<p>I would not call it merge though as it may confuse readers given there is a merge() function for dataframes.  What you describe is a way to efficiently append a column, instead of using concat().</p>\n\n<p>I do it in a one liner, that way you don't keep a pointer on the loaded feather file:</p>\n\n<pre><code>df[load.columns] = pd.read_feather('./new_features.ftr')\n</code></pre>\n\n<p>This way the loaded file will be reclaimed by the gc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322534,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-05-03T06:26:13.993000",
          "content": "<p>I'm curious what you said about pd -&gt; disk -&gt; lgb, how would one do that and specify categorical features? I did a Google search and didn't find anything, can you please help? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322540,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-05-03T06:54:23.110000",
          "content": "<p>@CPMP, <a href=\"/authman\">@authman</a> I think I'm generating the CSV file a bit different. </p>\n\n<p>Lets say I group on IP and OS and get the counts. So, the CSV file has 3 columns now, IP, OS, IP_OS_Count. So, now this CSV file has to be merged to the original DF..</p>\n\n<ul>\n<li>IP OS</li>\n<li>1 1 </li>\n<li>1 1</li>\n<li>2 2</li>\n</ul>\n\n<p>So, the new CSV File will be</p>\n\n<ul>\n<li>IP OS IP_OS_Count</li>\n<li>1 1 2 </li>\n<li>2 2 1</li>\n</ul>\n\n<p>So, now this count has to be added to each of the rows and the new df will be as below.</p>\n\n<ul>\n<li>IP OS IP_OS_Count</li>\n<li>1 1 2</li>\n<li>1 1 2</li>\n<li>2 2 1</li>\n</ul>\n\n<p>Am I doing it incorrectly?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322555,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-03T07:22:53.320000",
          "content": "<p><a href=\"/samrat\">@samrat</a>, You are doing it correctly.  But once done, you should save the column you just added to your main dataframe as a separate file instead of saving the whole dataframe.  Some of the public kernels use that: they save each feature separately.  It is way easier to make experiment that way, you just assemble the features you want to test from the individual files.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 322567,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-05-03T07:52:09.890000",
          "content": "<p>@CPMP Thanks a lot.. Now I got it.. Wish I knew it much earlier :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322589,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-03T08:42:30.187000",
          "content": "<p>No pb, but you could have learned it from public kernels ;)</p>\n\n<p>I don't use public kernels as is, as they often overfit, but some contain good ideas, and reading them is really something I value.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322676,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-03T12:14:51.643000",
          "content": "<p>@HuyenNguyen same as normal, it's handled for you internally in the lgb python wrapper. You don't have to roll your own file, just pass that parameter into your config. So you can either <code>df.col=df.col.astype('category')</code>, or alternatively, you can use the <code>categorical_features=[...]</code> argument of <code>lgb.train()</code>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322706,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-05-03T13:24:00.587000",
          "content": "<p>Hi@CPMP~ Do you find 'attributed_time' useful? Again, I'm glad if you refuse to answer this coz it's too personal, lol~</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 322744,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-03T15:11:02.727000",
          "content": "<p>I am not using attributed_time</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 322746,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-05-03T15:17:11.700000",
          "content": "<p>Upvote you and thank you!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 314150,
      "author_name": "Matt Motoki",
      "author_url": "",
      "post_date": "2018-04-14T18:31:47.220000",
      "content": "<p>Our best score 0.9797 is a single model with 9 features.  </p>",
      "votes": 10,
      "replies": [
        {
          "id": 314156,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-14T18:39:08.883000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314157,
          "author_name": "Rohit Mehra",
          "author_url": "",
          "post_date": "2018-04-14T18:42:16.970000",
          "content": "<p>Just 9 !!! Crazy..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 310313,
      "author_name": "NanoMathias",
      "author_url": "",
      "post_date": "2018-04-07T04:28:35.440000",
      "content": "<p>Got to 0.9711 with no parameter tuning of xgBoost using the features in <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">https://www.kaggle.com/nanomathias/feature-engineering-importance-testing</a>.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 310888,
          "author_name": "Chad Gardner",
          "author_url": "",
          "post_date": "2018-04-09T01:36:47.970000",
          "content": "<p>Using some of these features and a few of my own, I’m at 9.770.  LGB and plenty of tuning though. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 310892,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-09T01:40:29.347000",
          "content": "<p>@Chad Gardner Great job Chad! Could I ask what kind of features your r using? I'm totally fine if you keep it secret~</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 315303,
      "author_name": "Pranav Pandya",
      "author_url": "",
      "post_date": "2018-04-17T01:00:27.453000",
      "content": "<p>LB 0.9804 with single LGB in R with 38 features. (including some features from <code>Baris' next_click kernel</code> ). </p>\n\n<p>I use only 5% data for validation. </p>",
      "votes": 9,
      "replies": [
        {
          "id": 315459,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-17T06:23:00.907000",
          "content": "<p>Hi Pranav,\nDo you use the common day 9 hour 4 as CV? Or you generalize your cross validation and ignore the Public LB of hour 4? What is your CV AUC when training?</p>\n\n<p>Thanks!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315566,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-17T09:24:58.077000",
          "content": "<p>Nope! I use randomly shuffled split for validation while making sure that class unbalance stays exactly same in split. I deal with test hours separately. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315590,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-17T10:22:31.223000",
          "content": "<p>@Pranav Pandya\nThat's what I used to do but my CV AUC was always significantly higher than LB. \nDoes your CV AUC resemble in any way to the LB score?</p>\n\n<p>Thanks =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315603,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-17T10:40:54.230000",
          "content": "<p>Check out <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54673\">this post</a>  which will clear all the confusions about CV LB thingy thing. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 316417,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-19T01:39:56.803000",
          "content": "<p>is everyone using forward time deltas, because i tried backward time delta features, the difference between my day 9 CV and LB was huge, 0.9805 and 0.51</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317870,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-04-22T18:11:06.037000",
          "content": "<p>Hi Pranav,</p>\n\n<p>I would like to avoid the last rows for cross validation, so I tried to randomly split my data. I'm not sure if I should still leave those data in the training set or no ? I'm afraid if I exclude them I'd remove some useful data from the training set ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317916,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-22T20:33:26.257000",
          "content": "<p>Hello Antonie,</p>\n\n<p>My personal opinion is that chances of loosing highly predictive observations (close to test) is already reduced when validation split ist random shuffled. I wouldn't worry about re-running model with validation data in training set if I'm using full training data. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 317952,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-04-22T23:14:14.197000",
          "content": "<p>Thanks for clarifying this. I'll do more testing by creating 2 distinct set to see if this helps.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 317991,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-23T02:16:37.220000",
          "content": "<p>is there any method or way to randomly split data by maintaining the class imbalance?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318000,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-04-23T02:44:50.387000",
          "content": "<p>I'm using a basic numpy random function to split the data into training / validation. Since we have a huge number of data, I expect the imbalance to be close in train and valid set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318095,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-23T07:16:29.497000",
          "content": "<p>@nickhillator , you can use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html\">sklearn train_test_split</a> and passing target array as \"stratify\" argument you can use stratified train,test split. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318104,
          "author_name": "Mayank Soni",
          "author_url": "",
          "post_date": "2018-04-23T07:33:18.397000",
          "content": "<p>Here You go. \ntrain, test, y_train, y_test = train_test_split(train, y, test_size=0.2, random_state=42,stratify = y)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 318986,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-04-25T01:16:30.867000",
      "content": "<p>update: 0.9813 with 20 features (single lgb run)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 318987,
          "author_name": "Perfect Is Shit",
          "author_url": "",
          "post_date": "2018-04-25T01:23:08.897000",
          "content": "<p>Cool, may I ask how much data do you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318993,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-25T01:31:22.817000",
          "content": "<p>All.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 319192,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2018-04-25T13:35:32.143000",
          "content": "<p>All means including all training data? What is the size of the memory you are using ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319198,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-25T13:39:38.143000",
          "content": "<p>I'm traveling today hence must answer from memory.  </p>\n\n<p>All means all train data.</p>\n\n<p>I worked quite a bit to decrease memory consumption.  There is a spike for lgb dataset creation, but then it runs in less than 32GB, if not less 20GB.  I'll check tonight the peak and average memory use.  </p>\n\n<p>I also tried XGBoost but it requires more memory and I have not run it on full data.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319218,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2018-04-25T14:29:57.927000",
          "content": "<p>I understand. Currently I am confined to a 24GB and that's the reason I was asking about memory. I am also looking to ways to improve memory management so that I can fit both more rows from train data and more features. Thank you for the details.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 311093,
      "author_name": "alijs",
      "author_url": "",
      "post_date": "2018-04-09T11:14:18.207000",
      "content": "<p>I can confirm, that 0.9786 is also reachable by single model single run. </p>\n\n<p>Would be interesting to hear from Top-10 if 0.98+ is also reachable by single model?</p>",
      "votes": 8,
      "replies": [
        {
          "id": 311094,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-04-09T11:18:46.783000",
          "content": "<p>My single model scores 0.9810. I think people above me also have single models.</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 311111,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-09T12:15:15.247000",
          "content": "<p>How do you define single model?  Are you averaging several runs for instance?  Even if you do, then this is rather impressive.</p>",
          "votes": -3,
          "replies": []
        },
        {
          "id": 311117,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-04-09T12:31:25.090000",
          "content": "<p>No averaging = 0.9810.\nAverage of two runs: 0.9811. \nSince I use LGB and train using all the data, I don't get as much as benefit that I used to get ensembling different neural net runs in other competitions. My current model has very small variation.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 311190,
          "author_name": "Rohit Mehra",
          "author_url": "",
          "post_date": "2018-04-09T15:13:41.657000",
          "content": "<p>@AhmetErdem Did you use test_supplement to calculate features for the model scoring 0.9810?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311453,
          "author_name": "Badr Ouali",
          "author_url": "",
          "post_date": "2018-04-10T05:10:27.940000",
          "content": "<p>I think this supplement is not necessary at all. I'm nearby 0.98 (0.9781) and I am just using IP frequencies. Besides, I noticed that some unrelevant features can destroy the prediction. I advise to use a good training set and smart features that are using the IP (it is the most important variable with app). I will add other features tomorrow and tell what's going on on thursday!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 311797,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-10T17:26:20.947000",
          "content": "<p>@Badr - Did you use any target encoding at all?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311830,
          "author_name": "Rohit Mehra",
          "author_url": "",
          "post_date": "2018-04-10T18:32:06.940000",
          "content": "<p>@Badr IP frequencies are mostly overfitting !!! 0.9781 LB with overfit? :/</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311838,
          "author_name": "Badr Ouali",
          "author_url": "",
          "post_date": "2018-04-10T18:48:48.070000",
          "content": "<p>@Rohit IP frequencies that I use can not be done with regular Python. I am using Vertica to compute them. By the way the variables that I use are far from overfitting and using small trees, I think it is hard to overfitt with the parameters proposed by the LGBM. You have to think rationally about the situation, what fraudulent machines are doing and that's all.</p>\n\n<p>@Shanth, no not at all. I tried and the score decreased considerably because of overfitting this time.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 311841,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-10T18:53:44.850000",
          "content": "<p>@Rohit @Badr I think you just mean two different things by frequency - % attributed (in the past) vs. counts. </p>\n\n<p>@Badr what do you mean by frequencies that can't be done with python? Like a customized computation that base pandas doesn't support?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 311859,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-10T19:51:19.073000",
          "content": "<p>@Badr Thanks for the reply. Also the target encoding was for App/ App + device combination ? Obviously target encoding for IP will be a bad idea. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311867,
          "author_name": "Badr Ouali",
          "author_url": "",
          "post_date": "2018-04-10T20:02:44.360000",
          "content": "<p>@Joe you are right! For the question, I think it will be hard to do it with pandas. It is still possible but you need a really really powerful machine and a lot of patience I think!</p>\n\n<p>@Shanth, you welcome. I really like the Kaggle community, you are all very helpful. I think that the LGBM will manage to do it, if you increase the depth (risk of overfitting). I tried to use some combinations but my results were not really good. Sometimes it increases the AUC of 0.0001 which is not relevant at all. I think smart variables are enough to win the competition (+ the 4 categorical ones), I hope I will find some others during the week.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 311884,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-10T20:54:34.033000",
          "content": "<p>Curious exactly what you mean, but I think I have a pretty good guess. If my guess is right, I agree that it is hard, but definitely doable in python ;) (even doable in pandas but with much slowness).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 312499,
          "author_name": "Longchun Tao",
          "author_url": "",
          "post_date": "2018-04-11T21:28:04.227000",
          "content": "<p>Hi, I am curious that when you train on the all data, how do you define the validation set to determine the best iteration number? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 312503,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2018-04-11T21:40:23.983000",
          "content": "<p>To determine the best iteration you need a validation set, so you don't literally train on all data. After you learn what the best iteration number is you can repeat the training on the full set stopping at the best iteration. \nAlternatively, you can do cross-validation with 5 or 10 folds.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312520,
          "author_name": "shivraj",
          "author_url": "",
          "post_date": "2018-04-11T22:29:55.070000",
          "content": "<p>How are you planning to do a kfold without affecting the time series nature of data set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312528,
          "author_name": "Longchun Tao",
          "author_url": "",
          "post_date": "2018-04-11T23:36:08.837000",
          "content": "<p>@Alexey Thanks for your telling that. What I mean is that when all data are used to train, the ways of determining validation set may be kFold or randomly selecting some data. But the data set has time series nature.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312579,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2018-04-12T02:02:43.743000",
          "content": "<p>@Longchun Tao and shivraj <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325\">Here</a> there is a nice discussion about validation. You might want to read it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317623,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-22T02:43:39.037000",
          "content": "<p>@Joe @Badr @Shanth May I ask what you guys mean by target encoding in this case? Also, is this what I did related to it: I used ip + app + device as a group and calculated the % attributed for each of those groups and merged them into test set. My training acc was ~.999 and my valid. score was much lower </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318370,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-23T17:27:08.590000",
          "content": "<p>@ Edward Chen -  Yes that is target encoding. My current model does not include it since time delta/ unique counts etc proved to be more useful. It may be useful if one is using the entire data set ( which I am not doing for lack of resources). </p>\n\n<p>As for training accuracy - same here. I got an unusually high training accuracy but low validation score.  I also used a variation of Target encoding from NanoMathias's awesome script  <a>here</a>. This script scales the target encoding with a confidence score. I believe it helped his LB score but again he used the entire data set for training. </p>\n\n<p>Hope this helps \nRegards\nShanth</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 321414,
      "author_name": "Philippe Lonjoux",
      "author_url": "",
      "post_date": "2018-05-01T08:15:23.667000",
      "content": "<p>This is an edit of my previous post.</p>\n\n<p>0.9808 with 13 features (lgbm), </p>\n\n<p>0.9811 with 25 features (lgbm, 110M raws), </p>\n\n<p>0.9813 with 28 features (lgbm, 110M raws)</p>\n\n<p>I iteratively make one greedy choice after another (group of a few features) and remove it if it ends up not improving my local auc score.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 321514,
          "author_name": "Laevatein",
          "author_url": "",
          "post_date": "2018-05-01T13:19:14.493000",
          "content": "<p>Wow, that's amazing! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322157,
          "author_name": "Mahesh Kulkarni",
          "author_url": "",
          "post_date": "2018-05-02T13:36:14.067000",
          "content": "<p>great work phillipe</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319439,
      "author_name": "蓝鲤鱼",
      "author_url": "",
      "post_date": "2018-04-26T03:28:54.853000",
      "content": "<p>0.9806 with lgbm, 18features</p>",
      "votes": 5,
      "replies": [
        {
          "id": 319868,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-27T02:08:11.850000",
          "content": "<p>with all data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320057,
          "author_name": "蓝鲤鱼",
          "author_url": "",
          "post_date": "2018-04-27T11:21:12.113000",
          "content": "<p>yes</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320098,
          "author_name": "SubikashPal",
          "author_url": "",
          "post_date": "2018-04-27T13:13:36.663000",
          "content": "<p>How r u calculating time deltas with such a huge data set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320243,
          "author_name": "蓝鲤鱼",
          "author_url": "",
          "post_date": "2018-04-28T01:49:25.170000",
          "content": "<p>df.groupby(cols).clicktime.shift - df.clicktime</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 320817,
          "author_name": "spongebob",
          "author_url": "",
          "post_date": "2018-04-30T00:30:24.847000",
          "content": "<p>cool</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 314319,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2018-04-15T08:26:29.660000",
      "content": "<p>0.9791 with a single lightgbm on 25 features. Looking at other posts, I should be able to shrink that amount :-)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 315002,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-16T16:14:29.770000",
          "content": "<p>I have also managed to get similar score(0.9792) using 20 features with single lightgmb, but  I am still unsuccessful in shrinking features and get the same or better score. I guess I still need to engineer some strong features.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315011,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-16T16:28:56.107000",
          "content": "<p>@ Shoaib...Is this Day 8 training only or all data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315031,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-16T16:53:10.847000",
          "content": "<p>I train only on day 7,8 and validate on day 9.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 310606,
      "author_name": "Badr Ouali",
      "author_url": "",
      "post_date": "2018-04-08T04:51:15.410000",
      "content": "<p>Almost 30 features and a score of 0.9762 The training day is really really important to reach this score. I'll try to improve in the next days.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 310976,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T06:17:30.867000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319372,
      "author_name": "Antoine",
      "author_url": "",
      "post_date": "2018-04-25T22:29:48.633000",
      "content": "<p>It's my first competition ! For now I have 0.9806 with lightgbm, I'm using less than half of training data. If I have time / resources I'll try to run it on the full data set. I have around 20 features which makes my computation very slow.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 319533,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-26T08:39:58.673000",
          "content": "<p>That's very promising, esp for a first competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319655,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-04-26T14:26:44.593000",
          "content": "<p>Thanks :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 314701,
      "author_name": "NanoMathias",
      "author_url": "",
      "post_date": "2018-04-16T06:09:23.830000",
      "content": "<p>LB 0.9769 with features from <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">here</a> and xgBoost parameters from <a href=\"https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769\">here</a></p>",
      "votes": 6,
      "replies": [
        {
          "id": 314791,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-16T09:26:27.790000",
          "content": "<p>Awesome!  What was your LB before Bayesian optimization?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314795,
          "author_name": "NanoMathias",
          "author_url": "",
          "post_date": "2018-04-16T09:37:52.897000",
          "content": "<p>I was at 0.9711 prior to the optimization :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 314799,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-16T09:43:23.100000",
          "content": "<p>I guess you just triggered LOTS of interest for Bayesian optimization ;)  Thanks for sharing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 314839,
          "author_name": "NanoMathias",
          "author_url": "",
          "post_date": "2018-04-16T10:56:00.077000",
          "content": "<p>Thanks, hopefully it'll help people get better scores :) .. especially now that it's so easy to do with the scikit-optimize package..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 314850,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-16T11:16:17.113000",
          "content": "<p>May I ask why do we to find optimal parameters using kfold or stratified kfold when most of us are using different time based CV setting? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314855,
          "author_name": "NanoMathias",
          "author_url": "",
          "post_date": "2018-04-16T11:23:42.373000",
          "content": "<p>Using StratifiedKfold() is just my generic first-approach, which turned out to work OK (in terms of LB score being pretty similar to my local CV score)..</p>\n\n<p>However, adding in another cross-validator should be easy, e.g. using <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html#sklearn.model_selection.TimeSeriesSplit\">TimeSeriesSplit</a>. I guess it'd be worth testing that out as an alternative to the StratifiedKFold() one. </p>\n\n<p>Right now I'm back to working on improving the feature engineering though, and I think afterwards computing power could potentially be better spent (for me at least) looking into feature selection</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 314868,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-16T11:56:41.117000",
          "content": "<p>I will try this using TimeSeriesSplit and report back in few days.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314870,
          "author_name": "NanoMathias",
          "author_url": "",
          "post_date": "2018-04-16T11:59:44.707000",
          "content": "<p>Cool, looking forward to hearing if results are better with that approach</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 308079,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-04-02T23:13:21.763000",
      "content": "<p>Currently, best single run at 0.9736</p>\n\n<p>Others have said they had single model at 0.9764</p>",
      "votes": 5,
      "replies": [
        {
          "id": 308085,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-02T23:23:26.863000",
          "content": "<p>@CPMP, lead us to 0.98, go go go!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 308092,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-02T23:41:57.240000",
          "content": "<p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 308096,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-02T23:46:03.197000",
          "content": "<p>A spectre is haunting TalkingData—the spectre of model with only 4 engineered features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308099,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-04-02T23:50:55.413000",
          "content": "<blockquote>\n  <p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n</blockquote>\n\n<p>There are things that haunt me, but none of them have \"Kaggle challenge\" in their description. There will always be things that are obvious to others but not me, and <em>vice versa</em>.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 308102,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-04-02T23:54:09.423000",
          "content": "<blockquote>\n  <p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n</blockquote>\n\n<p>Heck, my single best model is at 0.9617, and even that doesn't haunt me.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 308104,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-02T23:56:30.447000",
          "content": "<blockquote>\n  <p><strong>Tilii wrote</strong></p>\n  \n  <blockquote>\n    <p>&gt; This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n  </blockquote>\n  \n  <p>There are things that haunt me, but none of them have \"Kaggle challenge\" in their description. There will always be things that are obvious to others but not me, and <em>vice versa</em>.</p>\n</blockquote>\n\n<p>Yeah I mean it jokingly :) In truth it's a motivation to explore the data more and keep digging for features, not a negative.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308176,
          "author_name": "Muhammad Alfiansyah",
          "author_url": "",
          "post_date": "2018-04-03T04:36:43.417000",
          "content": "<blockquote>\n  <p><strong>Joe Eddy wrote</strong></p>\n  \n  <blockquote>\n    <p>This whole .9764 with 4 features thing is going to haunt me for weeks</p>\n  </blockquote>\n</blockquote>\n\n<p>Becarefull!! XD the more you want to find it the more it will go away!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 309007,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-04T13:34:37.977000",
          "content": "<blockquote>\n  <p>@CPMP, lead us to 0.98, go go go!</p>\n</blockquote>\n\n<p>Thanks for your trust, but it won't happen anytime soon... Found a flaw in my data preparation, I don't expect submission before few days now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 312349,
          "author_name": "AlejandroCoronado",
          "author_url": "",
          "post_date": "2018-04-11T15:48:02.973000",
          "content": "<p>@CPMP I'm really impressed by your model. Do you have any suggestion to improve my score using single models?\nI mean maybe not for this competition but how do you know which model is going to work, what do you do to tune your models or if you have some strategy for feature engineering? How do you learn all of this?\nMaybe it's a lot to ask haha but I will take my chances.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312394,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-11T17:22:37.497000",
          "content": "<p>Hi, there are way better scores than mine, you should ask one of these guys above 0.98 ;)</p>\n\n<p>I learned in two ways: </p>\n\n<ol>\n<li><p>I followed my own advice to newbies, posted elsewhere on this forum.</p></li>\n<li><p>I read carefully what top performers share after each competition ends. </p></li>\n</ol>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 321908,
      "author_name": "Cheng",
      "author_url": "",
      "post_date": "2018-05-02T06:10:04.410000",
      "content": "<blockquote>\n  <p>Below .98xx app had 10x more importance over the channel. Above .98xx\n  channel has 3x more importance. I am unable to understand this</p>\n</blockquote>\n\n<p>I think this is because you are adding more features that has similar information as 'app'.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 320910,
      "author_name": "George Carmichael",
      "author_url": "",
      "post_date": "2018-04-30T07:00:06.867000",
      "content": "<p>0.9809 with all data, 7 new features and no target encoding - lgbm.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 320947,
          "author_name": "Sudeep Shukla",
          "author_url": "",
          "post_date": "2018-04-30T08:58:21.810000",
          "content": "<p>Care to share the validation strategy you used?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320974,
          "author_name": "George Carmichael",
          "author_url": "",
          "post_date": "2018-04-30T09:55:11.127000",
          "content": "<p>I validated with day 9, a learn rate of 0.05 and early stopping rounds at 50. The features I am using are all time delta (time since x) and click counts (by different grouping variables). Hope that helps a bit and good luck 🙂</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 321020,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-30T12:18:18.940000",
          "content": "<p>Sorry to pry -- time since x, or time to x? It seems most have given up on the former entirely, so would be pretty interesting if thy were the basis to your secret sauce. Also, impressive score, very motivating!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321028,
          "author_name": "George Carmichael",
          "author_url": "",
          "post_date": "2018-04-30T12:38:37.017000",
          "content": "<p>Time since x  but with different grouping variables to many I have seen in the kernels on here. I have also found that time lagged features have a high feature importance in the lgbm model, e.g. clicks in the last hour/day although these take a long time to process in Pandas! </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 321030,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-30T12:41:18.443000",
          "content": "<p>🙏</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321057,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-30T13:49:53.400000",
          "content": "<p>Hi @George Carmichael, good job and thank you very much for your hints! Do you find features such as clicks in the last hour/day  from kernel? I think I've noticed that some kernel has this kind of feature but forget which one.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 321253,
          "author_name": "George Carmichael",
          "author_url": "",
          "post_date": "2018-04-30T22:05:50.337000",
          "content": "<p>Hi, no problem! I haven't seen these features in kernels but pandas group by and rolling windows (with click time as index and window as '1h' ) do the job without too much complexity. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 321314,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-05-01T02:44:47.550000",
          "content": "<p>Congrats George ! Only 7 features is amazing... </p>\n\n<p>I'm surprised with your feature :\"Time since x\". I tried so many different way to include the time since last click with different grouping, but was never able to improve my models with those feature.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321358,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-05-01T04:59:02.157000",
          "content": "<p><a href=\"/authman\">@authman</a> did the features end up working for you?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321407,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-01T07:42:45.263000",
          "content": "<p>@George,  this is great, but we all count the number of features we use, not just the engineered ones.  Even if I include all original features that would mean 13 features for 0.9809, which is less than what I have seen so far.  The closest is Philippe Lonjoux with 0.9808 for 13 features, see below.</p>\n\n<p>Congrats for being so selective, and for sharing what works for you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 321491,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-01T12:15:45.797000",
          "content": "<p><a href=\"/areveillon\">@areveillon</a> not yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321516,
          "author_name": "George Carmichael",
          "author_url": "",
          "post_date": "2018-05-01T13:21:16.860000",
          "content": "<p>@CPMP thanks, well done on your great score! Yes I have 13 features in total due to dropping features with little importance after validation. It is likely I will add a couple more aimed specifically at dealing with duplicates.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 321811,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-05-02T00:09:57.947000",
          "content": "<p>@George Congrats on your amazing score! You mentioned that time lag features had high feature importance from your runs, but you didn't mention using them in your model. I'm just curious why</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321915,
          "author_name": "George Carmichael",
          "author_url": "",
          "post_date": "2018-05-02T06:21:53.767000",
          "content": "<p>Thank you! I removed the time lag features such as time since last attribution as they seemed to be causing overfitting  of the model. I may add them back in but currently trying to keep the model as simple as possible :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322125,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-05-02T12:51:11.617000",
          "content": "<p>Thanks for your reply! What about the time lag features such as num clicks in last hour/day that you mentioned? Were they overfitting as well?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322132,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-05-02T13:01:55.917000",
          "content": "<p>Thank you George. Regarding your hints about lagged features, the test set only contains several hours in discontinuous blocks. Did you use the full test set supplement? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 320816,
      "author_name": "spongebob",
      "author_url": "",
      "post_date": "2018-04-30T00:29:36.953000",
      "content": "<p>9820 lgbm</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 319384,
      "author_name": "AmirH",
      "author_url": "",
      "post_date": "2018-04-25T23:45:50.007000",
      "content": "<p>0.9811 with 35 features</p>",
      "votes": 3,
      "replies": [
        {
          "id": 319386,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-04-25T23:56:06.233000",
          "content": "<p>did you use full training set as well ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319388,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-26T00:20:34.493000",
          "content": "<p>Yes</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319989,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-27T08:45:17.593000",
          "content": "<p>Hello AmirH.How much improvement will be achieved using all data compared to using one day data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320052,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-27T11:10:07.870000",
          "content": "<p>Depends on your features of course, but i experienced quite a significant improvement</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320069,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-27T11:57:19.873000",
          "content": "<p>Thanks.About how much?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320275,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-04-28T05:27:48.970000",
          "content": "<p>Hi! what train\\valid strategy do you use? Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320326,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-28T10:24:07.910000",
          "content": "<p>Hi @AlexTru,\nI use days 7 + 8  as training and day 9 for Validation</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320422,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-28T16:16:36.750000",
          "content": "<p>Hey @AmirH, did you find that your validation score for day 9 and public lb score with all training data were similar? Or to what extent was the lb score higher/lower?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320438,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-28T17:04:10.277000",
          "content": "<p>@Edward Chen\nAbout 0.003 difference, the difference used to be bigger but using different features raised my score and also decreased the different between LB and CV AUC.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320442,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-28T17:27:08.397000",
          "content": "<p>@AmirH Thanks for the response! So you're saying the LB was .003 higher than CV AUC right? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320471,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-28T18:56:59.240000",
          "content": "<p>@Edward Chen\nNo, the other way around, the CV AUC tends to be about 0.003 higher than the LB</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 320521,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-28T22:20:37.087000",
          "content": "<p>@AmirH I see - that makes sense. Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319216,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-04-25T14:23:55.953000",
      "content": "<p><strong>Update - April 25th</strong></p>\n\n<ul>\n<li>Added 7 more features (20 new features in total) and the LB score improved just to 0.9801.</li>\n<li>Looking to work on the feature importance to remove some under performing features and some new features.</li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030\">More Details</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 317413,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-04-21T13:17:11.397000",
      "content": "<p>0.9798 with 9 new features and lgbm with complete data.</p>\n\n<p><strong>Update - April 23rd</strong></p>\n\n<p>Added 4 more features and the LB score improved to 0.9800.</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030\">More Details Here</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 317428,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-21T14:43:11.770000",
          "content": "<p>may I ask what kind of features are working for you, time delta or frequency features?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317430,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-21T14:49:02.517000",
          "content": "<p>I have a mix of both kinds.... 2 new frequency features helped me move up the lb..  Right now I'm adding 3 more similar features.. probably I might get a better understanding after the current run...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 317657,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-22T06:54:17.673000",
          "content": "<p>Congrats! When you say \"frequency features\", are you referring to features that are groups of variables mapped to the attributed rate or like just general counts of the feature in the data without relation to is_attributed?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318131,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-23T08:10:34.303000",
          "content": "<p>I have both the general count features and also ratio's related to attribute....</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318708,
          "author_name": "SubikashPal",
          "author_url": "",
          "post_date": "2018-04-24T10:00:40.907000",
          "content": "<p>Samrat -\nHow r u doing \"ratio's related to attribute\" for test set?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 317295,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-04-21T06:13:22.483000",
      "content": "<p>0.9807 lgbm with 17 features.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 317429,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-21T14:44:19.470000",
          "content": "<p>That's very inspiring, <br> may I ask what kind of features are working best for you, time delta or frequency features? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317432,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-21T14:57:54.150000",
          "content": "<p>You can ask but  I won't share before competition ends ;)  This competition is all about feature engineering, just try stuff.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 319174,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-25T12:49:53.737000",
          "content": "<p>are you using target encoding ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319191,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-25T13:33:09.917000",
          "content": "<p>You can ask but I won't share before competition ends ;) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 316881,
      "author_name": "NanoMathias",
      "author_url": "",
      "post_date": "2018-04-20T06:18:05.833000",
      "content": "<p>Gave up on actively competing, so instead have been trying to see how good I can get with a very small set of features - so far got LB 0.9727 with only <strong>4 features</strong> in total. I've still only selected features from <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">feature engineering notebook</a> and using the xgBoost parameters from <a href=\"https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769\">bayesian optimization notebook</a>. Right now I'm trying to add one or two more features, and then I think it can probably be improved a bit more from another round of bayesian tuning afterwards.. and perhaps finally I'll switch to lightGBM as well :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 316888,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-20T06:35:05.833000",
          "content": "<p>You should try 'tree_method'='hist', it helps. XGBoost will run way faster, and you may even get better results.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 316898,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2018-04-20T06:50:23.763000",
          "content": "<p>This is a very good comment. Indeed, there are many examples when carefully selected (and aggregated) engineered features could be more effective than multiple features as a way to improve a predictive model. Here is an example by <a href=\"https://www.kaggle.com/pliptor\">@Oskar Takeshita</a> with one single feature: <a href=\"https://www.kaggle.com/pliptor/divide-and-conquer-0-82296\">https://www.kaggle.com/pliptor/divide-and-conquer-0-82296</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316940,
          "author_name": "NanoMathias",
          "author_url": "",
          "post_date": "2018-04-20T08:05:10.270000",
          "content": "<p>Interesting, thanks, I'll be testing out the 'tree_method='hist'' in my next round of tuning</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 313003,
      "author_name": "KenMa",
      "author_url": "",
      "post_date": "2018-04-12T16:47:30.590000",
      "content": "<p>My best model is 0.9774 with 22 features and validate on day 9.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 313128,
          "author_name": "AlejandroCoronado",
          "author_url": "",
          "post_date": "2018-04-12T20:10:59.500000",
          "content": "<p>Good Job! I think your score is one of the highest for a single model approach.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313347,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-04-13T06:13:10.377000",
          "content": "<p>Nice! Does it mean that you use as train 7,8 days?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313848,
          "author_name": "KenMa",
          "author_url": "",
          "post_date": "2018-04-14T00:18:34.883000",
          "content": "<p>Just day 8 actually, due to the limited RAM i got. Not sure if training on different day then combining their prediction on day 9 to predict test set would be helpful, I will try it in the next few days.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 311318,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-04-09T20:24:18.700000",
      "content": "<p>I just reached 0.9727 with my first try of single model with more RAM ... (thanks to the free credits of MS Azure)</p>\n\n<p>Yes high scores with single model are reachable but you'll definitely need more RAM  than on  kaggle kernel (or  on my poor local machine :p)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 311329,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-09T20:53:30.623000",
          "content": "<p>That's great.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 311411,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-10T02:18:53.173000",
          "content": "<p>how much RAM to be precise, i am using GCP 40 gb ram, it is still messy at times while creating new features</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312466,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-04-11T19:48:57.257000",
          "content": "<p>About 42 Gb at the pic and higher at lightgbm init training (you can <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53773\">fix this last issue</a> )</p>\n\n<p>I am not using all the data...just 30 millions rows for training and about the same size for validation ..</p>\n\n<p>After some modification on features used  , I got now 0.9750 on LB with the same model ...</p>\n\n<p>Some features have really huge impact on model performance...I think I need now to dive deeper into features engineering </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 312472,
          "author_name": "Mayank Soni",
          "author_url": "",
          "post_date": "2018-04-11T20:08:01.030000",
          "content": "<p>Are your CV and LB in sync ?\nMy concern with using last 30 million or N million rows is Overfitting . Not sure what is better , using last N rows or using Day 2 and Day3 as train and Day 4 as validation . </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312518,
          "author_name": "nickhillator",
          "author_url": "",
          "post_date": "2018-04-11T22:12:45.387000",
          "content": "<p>have you used time delta features?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313084,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-04-12T19:03:15.883000",
          "content": "<p>@MayankSoni\nI am using the second strategy  (using Day 2 and Day3 as train and Day 4 as validation)</p>\n\n<p>@nickhillator <br>\nNot yet but I am working on them now....Hopefully will jump a bit at next submission ^^</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313111,
          "author_name": "Mayank Soni",
          "author_url": "",
          "post_date": "2018-04-12T19:40:17.620000",
          "content": "<p>Cool thanks . I am going to try Day 2 Day 3 Train , Validate Day 4 and then Train on Day 3,4 and Predict Test.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 308093,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-04-02T23:42:22.210000",
      "content": "<p>Single model 0.9734</p>",
      "votes": 3,
      "replies": [
        {
          "id": 308100,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-02T23:51:38.440000",
          "content": "<p>More than 4 engineered features :(</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 310648,
          "author_name": "Laevatein",
          "author_url": "",
          "post_date": "2018-04-08T06:54:41.740000",
          "content": "<p>Punch line: 4 features </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 317296,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-21T06:24:46.897000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 317301,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-21T06:43:12.263000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317415,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-21T13:24:33.443000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 308002,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-02T19:35:55.713000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 308013,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T20:00:04.393000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 308058,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T21:51:44.053000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 320834,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-30T01:27:44.743000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 322566,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-03T07:50:37.553000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 320615,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-29T09:18:42.200000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 320829,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-30T01:11:25.767000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320863,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-30T03:52:04.653000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320883,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-30T05:21:54.623000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 316604,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-19T12:59:58.767000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 315224,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-16T22:19:17.490000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 315355,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-17T02:18:38.580000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315367,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-17T02:43:27.900000",
          "content": "",
          "votes": 8,
          "replies": []
        },
        {
          "id": 315502,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-17T07:31:17.623000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 315771,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-17T15:18:08.523000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 316012,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-18T02:19:10.960000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 316038,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-18T04:04:21.110000",
          "content": "",
          "votes": -4,
          "replies": []
        },
        {
          "id": 316983,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-20T10:45:03.190000",
          "content": "",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 310750,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-08T13:33:35.873000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 310890,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T01:39:10.277000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310893,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T01:41:23.313000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310894,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T01:45:16.187000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310899,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T02:02:18.673000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 310986,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T06:46:01.400000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311114,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T12:20:11.543000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 308564,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-03T17:31:08.523000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 308147,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-03T02:59:26.777000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 308162,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-03T03:54:07.773000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 308164,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-03T04:00:59.003000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308165,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-03T04:03:23.267000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 308177,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-03T04:38:02.323000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 308911,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-04T10:15:09.153000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 320082,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-27T12:44:47.833000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 318576,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-24T05:06:13.820000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 318583,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T05:18:44.223000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318595,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T05:32:21.260000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318762,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T12:44:12.743000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318775,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T13:11:45.143000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318777,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T13:14:09.537000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318780,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T13:17:17.153000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318782,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T13:22:17.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318788,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T13:32:08.533000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 318811,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T14:13:27.237000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 317298,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-21T06:29:36.807000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 318017,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-23T03:29:58.553000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318165,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-23T09:32:29.103000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318821,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T14:33:05.533000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320099,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-27T13:17:12.823000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320667,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-29T13:03:22.053000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 322544,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-03T06:57:15.867000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 308967,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-04T12:48:22.673000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 309842,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T03:44:43.180000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310021,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T12:21:11.340000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310090,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T15:32:40.743000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310094,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T15:35:29.970000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 310095,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T15:46:34.203000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 310107,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T16:19:56.117000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 310128,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T16:57:01.373000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 310132,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-06T16:58:04.393000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 308209,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-03T06:06:42.970000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 308060,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-02T22:01:14.790000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 308063,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T22:06:40.830000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 308074,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T22:50:21.533000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 308105,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T23:59:20.040000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309112,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-04T17:00:41.953000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309123,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-04T17:23:50.157000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 312595,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-12T02:41:30.600000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 307998,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-02T19:23:17.713000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 308008,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T19:48:58.260000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308010,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T19:56:08.360000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308016,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-02T20:02:32.570000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 310887,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-09T01:35:12.757000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 310900,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T02:03:36.470000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 310909,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T02:44:17.523000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 310911,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T02:45:41.563000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 310914,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T02:59:23.963000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 310916,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T03:05:49.433000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3015835,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-13T02:50:37.783000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 578234,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T13:52:53.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 540213,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-31T06:53:45.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 395367,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-09-28T12:49:36.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 373053,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-21T00:12:39.253000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325144,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T07:03:50.743000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 321825,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-02T01:19:47.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 321438,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-01T09:37:25.120000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319717,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-26T17:16:32.900000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 319723,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T17:48:13.907000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319728,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T17:57:51.650000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319729,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T18:10:52.207000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319732,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T18:18:35.177000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319743,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T18:52:25.757000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319748,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T18:57:28.787000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319754,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T19:07:13.520000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319764,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T19:48:51.527000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319782,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T20:24:45.453000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319859,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-27T01:44:27.433000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319865,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-27T01:54:13.087000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 319899,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-27T04:18:04.993000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320363,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-28T12:42:25.770000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320972,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-30T09:49:55.040000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320993,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-30T10:58:11.060000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319112,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-25T08:40:39.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 318632,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-24T06:52:31.570000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 318645,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T07:20:36.710000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318648,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T07:25:57.367000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318649,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T07:30:50.380000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 318715,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T10:17:14.830000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318751,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T12:10:39.257000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318754,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-24T12:15:47.547000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319971,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-27T07:52:33.853000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322470,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-03T02:51:59.913000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 318255,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-23T13:31:25.510000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 318259,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-23T13:43:06.913000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318262,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-23T13:52:48.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319967,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-27T07:46:41.420000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 315687,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-17T13:31:45.767000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 317047,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-20T15:10:19.337000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319382,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-25T23:40:23.360000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319484,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T06:55:27.087000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 312340,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-11T15:42:37.150000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 312348,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T15:46:10.677000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 312351,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T15:49:38.930000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312362,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T16:20:49.623000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312455,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T19:26:41.453000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 312126,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-11T08:34:57.817000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 312229,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T12:31:02.917000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312232,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T12:36:37.477000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 312235,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T12:41:36.043000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 311974,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-11T02:21:01.770000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 311718,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-10T15:07:22.380000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 311875,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-10T20:21:46.640000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 310949,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-09T04:47:41.373000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 310950,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-09T04:54:11.633000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311381,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-10T00:21:03.613000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311462,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-10T05:38:19.110000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311686,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-10T14:00:20.197000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 311776,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-10T16:50:19.037000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315610,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-17T10:44:11.380000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317634,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-22T04:03:28.967000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "307993": "Everyone,\n\nA request to all the masters and people leading the board, can you share single best model scores? With all the blending going on, single scores would give us motivation.\n\n1. Mine ==&gt; LGB: All data (train + test, not test_supplement): 0.9697\n2. **Update 4-12-2018: All data (train + test, not test_supplement): 0.9774** I have yet to see how the new features i.e. unique user instance features can be used with the full train and also I am not using any \"attribute-based freq\" features.\n3. **Update 4-27-2018: All data (train + test, not test_supplement): 0.9799** Removing minute and seconds as features.\n4.  **Update 4-28-2018: All data (train + test, not test_supplement): 0.9802** Removed more features of type \"prevclick\". \n5.  **Update 4-29-2018: All data (train + test, not test_supplement): 0.9803** Hyperparameter tuning. More regularization.\n\nBelow .98xx app had 10x more importance over the channel. Above .98xx channel has 3x more importance. I am unable to understand this. Does anyone have some views on this? What is your most important feature?\n\nThanks.\n\nP.S. Upvote if you feel it will motivate you as well.. ;)",
    "317359": "Single NN model, LB 0.9814",
    "314699": "LB 0.9815 with single model and ~50 features and using all the train data",
    "308329": "Single model LB 0.9756, 18 features including original columns and frequencies.\n\nAdditional info: IP and ip frequencies were not used. Training was done on day 8. \nLocally I have similar results on day 7.",
    "318823": "update: single model with 11 features trained on day 9 only (~60M rows) get to .9802",
    "313326": "My current lb score is a single model. 0.9784 with 7 additional features - 2 groupby count features, 2 delta-time features, 3 confRate features.  Trained on day 9 with 0.1 validation. \n\nEdit: the new score of 0.9796 is also from a single model using 3 more features, still trained on day 9 only. \n\nEdit: 0.9798, single model with 18 features, trained on day 9 only - cant go any further with my 32GB ram... will try something else next.",
    "320603": "0.9822 with full data,60 features,lgbm",
    "314150": "Our best score 0.9797 is a single model with 9 features.  ",
    "310313": "Got to 0.9711 with no parameter tuning of xgBoost using the features in https://www.kaggle.com/nanomathias/feature-engineering-importance-testing.",
    "315303": "LB 0.9804 with single LGB in R with 38 features. (including some features from `Baris' next_click kernel` ). \n\nI use only 5% data for validation. ",
    "318986": "update: 0.9813 with 20 features (single lgb run)",
    "311093": "I can confirm, that 0.9786 is also reachable by single model single run. \n\nWould be interesting to hear from Top-10 if 0.98+ is also reachable by single model?",
    "321414": "This is an edit of my previous post.\n\n0.9808 with 13 features (lgbm), \n\n0.9811 with 25 features (lgbm, 110M raws), \n\n0.9813 with 28 features (lgbm, 110M raws)\n\nI iteratively make one greedy choice after another (group of a few features) and remove it if it ends up not improving my local auc score.",
    "319439": "0.9806 with lgbm, 18features",
    "314319": "0.9791 with a single lightgbm on 25 features. Looking at other posts, I should be able to shrink that amount :-)",
    "310606": "Almost 30 features and a score of 0.9762 The training day is really really important to reach this score. I'll try to improve in the next days.",
    "319372": "It's my first competition ! For now I have 0.9806 with lightgbm, I'm using less than half of training data. If I have time / resources I'll try to run it on the full data set. I have around 20 features which makes my computation very slow.",
    "314701": "LB 0.9769 with features from [here](https://www.kaggle.com/nanomathias/feature-engineering-importance-testing) and xgBoost parameters from [here](https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769)",
    "308079": "Currently, best single run at 0.9736\n\nOthers have said they had single model at 0.9764",
    "321908": "&gt; Below .98xx app had 10x more importance over the channel. Above .98xx\n&gt; channel has 3x more importance. I am unable to understand this\n\nI think this is because you are adding more features that has similar information as 'app'.",
    "320910": "0.9809 with all data, 7 new features and no target encoding - lgbm.",
    "320816": "9820 lgbm",
    "319384": "0.9811 with 35 features",
    "319216": "**Update - April 25th**\n\n- Added 7 more features (20 new features in total) and the LB score improved just to 0.9801.\n- Looking to work on the feature importance to remove some under performing features and some new features.\n\n[More Details][1]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030",
    "317413": "0.9798 with 9 new features and lgbm with complete data.\n\n**Update - April 23rd**\n\nAdded 4 more features and the LB score improved to 0.9800.\n\n[More Details Here][1]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55030",
    "317295": "0.9807 lgbm with 17 features.",
    "316881": "Gave up on actively competing, so instead have been trying to see how good I can get with a very small set of features - so far got LB 0.9727 with only **4 features** in total. I've still only selected features from [feature engineering notebook](https://www.kaggle.com/nanomathias/feature-engineering-importance-testing) and using the xgBoost parameters from [bayesian optimization notebook](https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769). Right now I'm trying to add one or two more features, and then I think it can probably be improved a bit more from another round of bayesian tuning afterwards.. and perhaps finally I'll switch to lightGBM as well :)",
    "313003": "My best model is 0.9774 with 22 features and validate on day 9.",
    "311318": "I just reached 0.9727 with my first try of single model with more RAM ... (thanks to the free credits of MS Azure)\n\nYes high scores with single model are reachable but you'll definitely need more RAM  than on  kaggle kernel (or  on my poor local machine :p)\n",
    "308093": "Single model 0.9734",
    "317296": "0.9808 with 13 features (lgbm)\n\nEDIT: 0.9811 with 25 features (lgbm, 110M raws)",
    "308002": "Right now my best single model score is public. The [single kernel version][1] scores 0.9694.  The [multiple-kernel version][2] (two kernels using the same model for different time periods and blended using guesstimated weights based on their LB scores) scores 0.9695.  I could run the model on the full dataset at home, but I haven't tried yet, and not sure I will with that particular model.  It has a history in public kernels, and I fear it is overfit to the public LB already.  I'm trying to focus more on validation and development, which is time-consuming given the large size of the dataset.  I'm not at the point where original candidates for a possible best score are ready to submit.\n\n [1]: https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python\n [2]: https://www.kaggle.com/aharless/ceci-n-est-pas-un-m-lange",
    "320834": "9807, 22 features ",
    "320615": "single lgb, 0.9793, 80M rows, 23 features",
    "316604": "Got LB 0.9791 with 23 features.Single Model.",
    "315224": "Our team has 13 features to reach 0.9803 with a single model using days 7 &amp; 8 as training and 9 to evaluate.",
    "310750": "My best single model scores 0.9751 on LB.",
    "308564": "Currently my best model LB 0.9694\ni wish to go LB 0.97 ...",
    "308147": "I'm wondering what is the best score for somebody using a subsample of train data.  I have no way of running a model on the full train set, so training everything on various subsets.  Wondering how far that can go...",
    "320082": "Getting 0.9786 with 10 features, running on Kaggle Kernel with 40M rows",
    "318576": "My NN (based on https://www.kaggle.com/antmarakis/deep-learning-approach-validation-lb-0-9684) trained on 100 m datas(day8 and day9) and got 0.977 lb score(local auc 0.9812). <br>\n\n_____________update____________________________ <br>\nNow my nn model can get 0.9777(local auc 0.982369).\n",
    "317298": "added three features and LB score jumped from 0.9691 to 0.9799...with single lightgbm\n\n**update:** 0.9810 using all data...",
    "308967": "My current single model has a public LB score of 0.9710. It has 12 features (including the ones that come with the data set). Still working on it...",
    "308209": "My best single model score is my current LB score:  0.9731 with 26 features.  I will dig deep into feature engineer after finished my graduation thesis.",
    "308060": "My best single model scores 0.9718 on public lb and I've created almost 40 new features.  I think I must be on a wrong way because Danijel said that he achieved 0.976+ with only four new features. And I've noticed that many top rankers jump from 0.97 to 0.98 very quickly thus I think there should be some magic features. Emmm...",
    "307998": "[Have a look at this][1]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53250",
    "310887": "14 features, LGB, hitting  9.770 (edit: 0.9770  Thanks for the catch @CPMP :)",
    "3015835": "nicee       ",
    "578234": "Nice",
    "540213": "Bagged trees haha",
    "395367": "wow ..nice!!",
    "373053": "gj!",
    "325144": "lgbm with 18 plus more feature 120m data private score 0.9802",
    "321825": "0.9795 total 22 features ",
    "321438": "0.9792, 7 new features with all training data. Still working on it!",
    "319717": "**Update - April 26th** - Added 5 more features (25 new features in total) and the local validation score jumped to 0.99145 and the LB score improved to 0.9802.",
    "319112": "single lgbm model 0.9794, 100m rows,  24 features ",
    "318632": "#offtopic\ncan anyone explain how to utilize confRate features or use target encoding for test set. Thank you!",
    "318255": "@CuteChibiko -\nGreat to see NN is performing well.  Are u using vanilla NN architecture or some sort of CNN/LSTM model.?\n",
    "315687": "#offtopic\nhow do you guys deal with the NaN values created while using time delta features, i.e the first iterations of each case",
    "312340": "I'm having problems getting a better score than .9687 with a single model. Do you have any suggestions?\nI noticed that most of the people are using LGBM model so I'm doing the same. Probably I can improve my score using an other aproach in the feature engineering step.",
    "312126": "are all the top guys using time deltas in feature engineering, can anyone confirm if they achieved over 0.9710 LB without it?",
    "311974": "Best single LGB hits 0.9712, 12 features with half of the training set and a random split (0.1) for validation :P",
    "311718": "What do you mean by single model ?",
    "310949": "Some one try to use test_supplement?"
  }
}