{
  "id": 54171,
  "title": "Best Single Model Kernel Score?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54171",
  "author_name": "Joe Eddy",
  "post_date": "2018-04-10T16:17:07.895000",
  "votes": 24,
  "comment_count": 91,
  "views": 0,
  "content": "<p>Already <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\">here</a> is a good discussion of single model scores.</p>\n\n<p>What I'm also interested in is how far people have been able to push the limits of kaggle kernel computation for single models, beyond visible public kernel scores. </p>\n\n<p>The best I've seen that's publicly visible is <a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711\">anttip's awesome FM_FTRL kernel</a> (.9711), and my personal best is a lgbm kernel on 50 million rows with 5 engineered features (.9707). </p>\n\n<p>Curious what others have found possible!</p>",
  "messages": [
    {
      "id": 311758,
      "postDate": "2018-04-10T16:17:07.897Z",
      "content": "<p>Already <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\">here</a> is a good discussion of single model scores.</p>\n\n<p>What I'm also interested in is how far people have been able to push the limits of kaggle kernel computation for single models, beyond visible public kernel scores. </p>\n\n<p>The best I've seen that's publicly visible is <a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711\">anttip's awesome FM_FTRL kernel</a> (.9711), and my personal best is a lgbm kernel on 50 million rows with 5 engineered features (.9707). </p>\n\n<p>Curious what others have found possible!</p>",
      "rawMarkdown": "Already [here][1] is a good discussion of single model scores.\n\nWhat I'm also interested in is how far people have been able to push the limits of kaggle kernel computation for single models, beyond visible public kernel scores. \n\nThe best I've seen that's publicly visible is [anttip's awesome FM_FTRL kernel][2] (.9711), and my personal best is a lgbm kernel on 50 million rows with 5 engineered features (.9707). \n\nCurious what others have found possible!\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\n  [2]: https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711",
      "votes": 24
    },
    {
      "id": 317344,
      "postDate": "2018-04-21T08:41:02.237Z",
      "content": "<p>9820，single lgb</p>",
      "rawMarkdown": "9820，single lgb",
      "votes": 10
    },
    {
      "id": 312004,
      "postDate": "2018-04-11T03:40:39.747Z",
      "content": "<p>40 Million training rows, R lightgbm, 4 new features added to Pranav's public kernel. LB 0.9736. Available in Public Kaggle kernel.</p>",
      "rawMarkdown": "40 Million training rows, R lightgbm, 4 new features added to Pranav's public kernel. LB 0.9736. Available in Public Kaggle kernel.",
      "votes": 9,
      "replies": [
        {
          "id": 312227,
          "postDate": "2018-04-11T12:26:39.047Z",
          "content": "<p>Great job! Thanks for sharing</p>",
          "rawMarkdown": "Great job! Thanks for sharing"
        }
      ]
    },
    {
      "id": 314360,
      "postDate": "2018-04-15T10:54:01.790Z",
      "content": "<p>0.979(5-9) with single lgb, (7-8) engineered features + different parameters</p>\n\n<p>the engineered features and parameters changed over runs but I think the range should be stable.</p>\n\n<p>Update: new model at .9803</p>",
      "rawMarkdown": "0.979(5-9) with single lgb, (7-8) engineered features + different parameters\n\nthe engineered features and parameters changed over runs but I think the range should be stable.\n\nUpdate: new model at .9803",
      "votes": 10,
      "replies": [
        {
          "id": 314586,
          "postDate": "2018-04-15T21:52:35.547Z",
          "content": "<p>Awesome!</p>",
          "rawMarkdown": "Awesome!"
        },
        {
          "id": 314620,
          "postDate": "2018-04-16T00:17:53.427Z",
          "content": "<p>Congrats! What are the typical number of rounds that you guys are using to train your lgbm models? I've been using ~1500 and am not sure if it's enough</p>",
          "rawMarkdown": "Congrats! What are the typical number of rounds that you guys are using to train your lgbm models? I've been using ~1500 and am not sure if it's enough"
        },
        {
          "id": 314624,
          "postDate": "2018-04-16T00:40:25.417Z",
          "content": "<p>it depends on your features + parameters. i use 600 rounds (because my kernel is derived from Andy Harless' kernel). my attempts to tune parameters gave me a range of 300 - 750 but the improvements were marginal (1 point at the 4th decimal place) so i just stuck with 600.</p>",
          "rawMarkdown": "it depends on your features + parameters. i use 600 rounds (because my kernel is derived from Andy Harless' kernel). my attempts to tune parameters gave me a range of 300 - 750 but the improvements were marginal (1 point at the 4th decimal place) so i just stuck with 600."
        },
        {
          "id": 315456,
          "postDate": "2018-04-17T06:21:10.307Z",
          "content": "<p>Do you use the common day 9 hour 4 as CV?\nOr you generalize your cross validation and ignore the Public LB of hour 4?\nWhat is your CV AUC when training?</p>\n\n<p>Thanks!!</p>",
          "rawMarkdown": "Do you use the common day 9 hour 4 as CV?\nOr you generalize your cross validation and ignore the Public LB of hour 4?\nWhat is your CV AUC when training?\n\nThanks!!",
          "votes": -1
        },
        {
          "id": 317473,
          "postDate": "2018-04-21T17:20:44.433Z",
          "content": "<p>It doesn't say much as it is depends on the learning rate, features and hyperparamers such as subsample. but ~240.</p>",
          "rawMarkdown": "It doesn't say much as it is depends on the learning rate, features and hyperparamers such as subsample. but ~240."
        }
      ]
    },
    {
      "id": 321192,
      "postDate": "2018-04-30T19:40:42.057Z",
      "content": "<p>0.9801 NN model run on Kaggle GPU in less than 1 hour..but hard to blend for now </p>",
      "rawMarkdown": "0.9801 NN model run on Kaggle GPU in less than 1 hour..but hard to blend for now ",
      "votes": 7,
      "replies": [
        {
          "id": 324520,
          "postDate": "2018-05-07T19:10:13.393Z",
          "content": "<p>Hi!</p>\n\n<p>I am really interested in your NN model. Can you share your knowledge when competition ends?\nMax what I can achieve is 0.9764 NN that is 0.0004 better than public kernel :(  I will surely try to improve my score on NN when competition ends and I wil have more time for it ;)</p>\n\n<p>Cheers and good luck! </p>",
          "rawMarkdown": "Hi!\n\nI am really interested in your NN model. Can you share your knowledge when competition ends?\nMax what I can achieve is 0.9764 NN that is 0.0004 better than public kernel :(  I will surely try to improve my score on NN when competition ends and I wil have more time for it ;)\n\nCheers and good luck! "
        }
      ]
    },
    {
      "id": 316468,
      "postDate": "2018-04-19T06:24:10.870Z",
      "content": "<p>9806</p>",
      "rawMarkdown": "9806",
      "votes": 5,
      "replies": [
        {
          "id": 316546,
          "postDate": "2018-04-19T10:06:21.563Z",
          "content": "<p>Great, wonder what guys above 0.98xx are doing.</p>",
          "rawMarkdown": "Great, wonder what guys above 0.98xx are doing."
        },
        {
          "id": 316798,
          "postDate": "2018-04-19T23:56:26.363Z",
          "content": "<p>Is this trained on the whole dataset?</p>",
          "rawMarkdown": "Is this trained on the whole dataset?"
        },
        {
          "id": 316802,
          "postDate": "2018-04-20T00:11:45.177Z",
          "content": "<p>no, just one day   </p>",
          "rawMarkdown": "no, just one day   ",
          "votes": 2
        },
        {
          "id": 316809,
          "postDate": "2018-04-20T00:47:18.333Z",
          "content": "<p>Nice work! It is impressive</p>",
          "rawMarkdown": "Nice work! It is impressive"
        },
        {
          "id": 316835,
          "postDate": "2018-04-20T03:35:28.210Z",
          "content": "<p>how about you</p>",
          "rawMarkdown": "how about you"
        },
        {
          "id": 316836,
          "postDate": "2018-04-20T03:36:48.690Z",
          "content": "<p>I always train with full data so don't know the performance for subset.</p>",
          "rawMarkdown": "I always train with full data so don't know the performance for subset."
        },
        {
          "id": 316915,
          "postDate": "2018-04-20T07:19:43.963Z",
          "content": "<p>so the full data?</p>",
          "rawMarkdown": "so the full data?"
        },
        {
          "id": 318633,
          "postDate": "2018-04-24T06:53:07.200Z",
          "content": "<p>Cheng how many extra features do you use?</p>",
          "rawMarkdown": "Cheng how many extra features do you use?"
        },
        {
          "id": 318640,
          "postDate": "2018-04-24T07:02:29.833Z",
          "content": "<p>less than 20</p>",
          "rawMarkdown": "less than 20"
        },
        {
          "id": 318847,
          "postDate": "2018-04-24T15:32:57.267Z",
          "content": "<p>who can more than。。。</p>",
          "rawMarkdown": "who can more than。。。"
        },
        {
          "id": 318959,
          "postDate": "2018-04-24T21:40:27.597Z",
          "content": "<p>Hi,  Peter. Do you mind tell me which day did you use?</p>",
          "rawMarkdown": "Hi,  Peter. Do you mind tell me which day did you use?"
        }
      ]
    },
    {
      "id": 323974,
      "postDate": "2018-05-06T20:18:58.553Z",
      "content": "<p>0.9819 with single LightGBM model - we entered the game too late, and probably don't have time to do stacking though :)</p>",
      "rawMarkdown": "0.9819 with single LightGBM model - we entered the game too late, and probably don't have time to do stacking though :)",
      "votes": 3,
      "replies": [
        {
          "id": 323976,
          "postDate": "2018-05-06T20:22:21.630Z",
          "content": "<p>@Yifan Xie how many features do you use?</p>",
          "rawMarkdown": "@Yifan Xie how many features do you use?"
        },
        {
          "id": 323978,
          "postDate": "2018-05-06T20:27:18.963Z",
          "content": "<p>36</p>",
          "rawMarkdown": "36",
          "votes": 1
        },
        {
          "id": 323994,
          "postDate": "2018-05-06T21:46:21.633Z",
          "content": "<p>@Yifan Xie That is impressive! May I know how you/your team handle ram while building the model in kernel? By my limited cs knowledge, 9 new features are the maximum I can generate in kernel.</p>",
          "rawMarkdown": "@Yifan Xie That is impressive! May I know how you/your team handle ram while building the model in kernel? By my limited cs knowledge, 9 new features are the maximum I can generate in kernel."
        },
        {
          "id": 323996,
          "postDate": "2018-05-06T21:51:57.030Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 323998,
          "postDate": "2018-05-06T21:56:32.607Z",
          "content": "<p>Thanks for the quick response. I do hope I can have a decent computer because my local machine only has 4gb RAM.  Kaggle kernel is the best resource I have for this competition. </p>",
          "rawMarkdown": "Thanks for the quick response. I do hope I can have a decent computer because my local machine only has 4gb RAM.  Kaggle kernel is the best resource I have for this competition. "
        },
        {
          "id": 324187,
          "postDate": "2018-05-07T10:32:08.700Z",
          "content": "<p>@Yifan,  you run this model on Kaggle kernels?  If so then this is really impressive.  If you run it on your own machine then you may want to compare with what people report in this other thread.</p>",
          "rawMarkdown": "@Yifan,  you run this model on Kaggle kernels?  If so then this is really impressive.  If you run it on your own machine then you may want to compare with what people report in this other thread."
        },
        {
          "id": 324191,
          "postDate": "2018-05-07T10:38:45.147Z",
          "content": "<p>lol, didn't pay attention this thread is for kernel only, yup I will have a go at that after the competition. </p>",
          "rawMarkdown": "lol, didn't pay attention this thread is for kernel only, yup I will have a go at that after the competition. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 312791,
      "postDate": "2018-04-12T10:45:24.233Z",
      "content": "<p>My update is that I managed to hit .9745 in a kernel with 7 engineered features. </p>\n\n<p>.9765 with only 4 engineered features..... maybe within reach!</p>",
      "rawMarkdown": "My update is that I managed to hit .9745 in a kernel with 7 engineered features. \n\n.9765 with only 4 engineered features..... maybe within reach!",
      "votes": 3,
      "replies": [
        {
          "id": 312797,
          "postDate": "2018-04-12T10:56:24.767Z",
          "content": "<p>I also think that this is possible. I got 0.9761 in a kernel  with 7  engineered features.</p>",
          "rawMarkdown": "I also think that this is possible. I got 0.9761 in a kernel  with 7  engineered features.",
          "votes": 3
        },
        {
          "id": 312936,
          "postDate": "2018-04-12T14:50:09.087Z",
          "content": "<p>Nice!</p>",
          "rawMarkdown": "Nice!"
        },
        {
          "id": 313103,
          "postDate": "2018-04-12T19:26:43.957Z",
          "content": "<p>.9787 is possible in a kernel with 7 engineered features :)</p>",
          "rawMarkdown": ".9787 is possible in a kernel with 7 engineered features :)",
          "votes": 5
        },
        {
          "id": 313107,
          "postDate": "2018-04-12T19:34:40.523Z",
          "content": "<p>Great job!</p>",
          "rawMarkdown": "Great job!"
        },
        {
          "id": 313116,
          "postDate": "2018-04-12T19:44:27.350Z",
          "content": "<p>That's Amazing @Joe, I have more than 10 features and I can not figure out how to cross 0.975x. <br> Looking forward to your feature engineering explanation.  </p>",
          "rawMarkdown": "That's Amazing @Joe, I have more than 10 features and I can not figure out how to cross 0.975x. <br> Looking forward to your feature engineering explanation.  ",
          "votes": 1
        },
        {
          "id": 313129,
          "postDate": "2018-04-12T20:11:52.110Z",
          "content": "<blockquote>\n  <p><strong>Joe Eddy wrote</strong></p>\n  \n  <blockquote>\n    <p>.9787 is possible in a kernel with 7 engineered features :)</p>\n  </blockquote>\n</blockquote>\n\n<p>Wow! that's impressive. May I ask if you're using R/Python?</p>",
          "rawMarkdown": "\n&gt; **Joe Eddy wrote**\n&gt; \n&gt; &gt; .9787 is possible in a kernel with 7 engineered features :)\n\nWow! that's impressive. May I ask if you're using R/Python?",
          "votes": 2
        },
        {
          "id": 313138,
          "postDate": "2018-04-12T20:29:16.420Z",
          "content": "<p>@Sohaib @Pranav Thanks both!</p>\n\n<p>I definitely plan to share my approach after the competition ends.</p>\n\n<p>I'm using python. And actually this kernel was built on top of the python port of your R kernel, so thank you for that awesome work.</p>",
          "rawMarkdown": "@Sohaib @Pranav Thanks both!\n\nI definitely plan to share my approach after the competition ends.\n\nI'm using python. And actually this kernel was built on top of the python port of your R kernel, so thank you for that awesome work.",
          "votes": 2
        },
        {
          "id": 313140,
          "postDate": "2018-04-12T20:33:06.250Z",
          "content": "<p>My pleasure! And thank you so much for the kind words. I really glad to hear that. </p>",
          "rawMarkdown": "My pleasure! And thank you so much for the kind words. I really glad to hear that. ",
          "votes": 1
        },
        {
          "id": 313333,
          "postDate": "2018-04-13T05:52:06.987Z",
          "content": "<p>may I ask you how you deal with ram explore while building new features? I have 5 features other than the base, ones, when I add the 6th new feature, kernel died.   </p>",
          "rawMarkdown": "may I ask you how you deal with ram explore while building new features? I have 5 features other than the base, ones, when I add the 6th new feature, kernel died.   "
        },
        {
          "id": 313489,
          "postDate": "2018-04-13T10:33:34.203Z",
          "content": "<p>@MengYe, If you're adding new features then best would be to reduce chunk size (observations) that can pass through without hitting memory limit. You may find <a href=\"https://www.kaggle.com/pranav84/double-xgb-hist-example-for-chunk-processing\">this kernel</a>  helpful in this regard. </p>",
          "rawMarkdown": "@MengYe, If you're adding new features then best would be to reduce chunk size (observations) that can pass through without hitting memory limit. You may find [this kernel][1]  helpful in this regard. \n\n\n  [1]: https://www.kaggle.com/pranav84/double-xgb-hist-example-for-chunk-processing"
        },
        {
          "id": 313823,
          "postDate": "2018-04-13T22:07:32.703Z",
          "content": "<p>@Joe Eddy congrats! May I ask on which subset of the training data you trained that model on?</p>",
          "rawMarkdown": "@Joe Eddy congrats! May I ask on which subset of the training data you trained that model on?"
        },
        {
          "id": 314587,
          "postDate": "2018-04-15T21:53:02.127Z",
          "content": "<p>@Edward thanks! That model was just trained on day 9.</p>",
          "rawMarkdown": "@Edward thanks! That model was just trained on day 9."
        }
      ]
    },
    {
      "id": 311955,
      "postDate": "2018-04-11T01:17:45.030Z",
      "content": "<p>0.9720 with 45 mln observations, 14 features, single model lgb in R. Tried to add a few more promising features and hit the kernel memory limit. Planning to try Python instead.</p>\n\n<p>Update: Python helped a little bit (plus some feature engineering, of course). The current result is 0.9789.</p>",
      "rawMarkdown": "0.9720 with 45 mln observations, 14 features, single model lgb in R. Tried to add a few more promising features and hit the kernel memory limit. Planning to try Python instead.\n\nUpdate: Python helped a little bit (plus some feature engineering, of course). The current result is 0.9789.",
      "votes": 3
    },
    {
      "id": 311800,
      "postDate": "2018-04-10T17:37:53.580Z",
      "content": "<p>0.9723 with 75M observations (single LGB in R). </p>",
      "rawMarkdown": "0.9723 with 75M observations (single LGB in R). ",
      "votes": 3
    },
    {
      "id": 311779,
      "postDate": "2018-04-10T16:53:03.120Z",
      "content": "<p>62M, simple lgb model, 15 features, 0.9743, run on the kernel. </p>",
      "rawMarkdown": "62M, simple lgb model, 15 features, 0.9743, run on the kernel. ",
      "votes": 3
    },
    {
      "id": 311761,
      "postDate": "2018-04-10T16:25:38.340Z",
      "content": "<p>My PB is 0.9724 on 75 millions training rows with 14 features including some of original ones. Still trying to hit higher LB score before diving into blending/stacking. </p>",
      "rawMarkdown": "My PB is 0.9724 on 75 millions training rows with 14 features including some of original ones. Still trying to hit higher LB score before diving into blending/stacking. ",
      "votes": 4
    },
    {
      "id": 325030,
      "postDate": "2018-05-08T03:35:03.663Z",
      "content": "<p>The competition has ended (with disappointment), may I know how you guys build models and achieve high scores with kernel all the way through? I really want to know because I too used kernel from the begining till the end. </p>",
      "rawMarkdown": "The competition has ended (with disappointment), may I know how you guys build models and achieve high scores with kernel all the way through? I really want to know because I too used kernel from the begining till the end. ",
      "votes": 1
    },
    {
      "id": 317471,
      "postDate": "2018-04-21T17:18:41.833Z",
      "content": "<p>0.9792 single lgb bested on baris Kanbar's kernel.</p>\n\n<p>I had to do better CV for feature selection, but I still use the last 5M rows for early stopping trigger.</p>",
      "rawMarkdown": "0.9792 single lgb bested on baris Kanbar's kernel.\n\nI had to do better CV for feature selection, but I still use the last 5M rows for early stopping trigger.",
      "votes": 1
    },
    {
      "id": 317175,
      "postDate": "2018-04-21T00:18:35.823Z",
      "content": "<p>0.9793, single lgb, 22 engineered features</p>",
      "rawMarkdown": "0.9793, single lgb, 22 engineered features",
      "votes": 1
    },
    {
      "id": 316453,
      "postDate": "2018-04-19T04:47:33.760Z",
      "content": "<p>0.9787, single lgb, few engineered features, 75M training, 5 M validation</p>",
      "rawMarkdown": "0.9787, single lgb, few engineered features, 75M training, 5 M validation",
      "votes": 1
    },
    {
      "id": 311992,
      "postDate": "2018-04-11T03:04:34.617Z",
      "content": "<p>0.9716 on PB with Full observations using lightgbm.  It's so frustrated found someone is over .98.</p>",
      "rawMarkdown": "0.9716 on PB with Full observations using lightgbm.  It's so frustrated found someone is over .98.",
      "votes": 1,
      "replies": [
        {
          "id": 311999,
          "postDate": "2018-04-11T03:26:21.307Z",
          "content": "<p>What method are you using to run the full data? Are you chunking?</p>",
          "rawMarkdown": "What method are you using to run the full data? Are you chunking?"
        },
        {
          "id": 312002,
          "postDate": "2018-04-11T03:33:39.353Z",
          "content": "<p>Sorry, I misunderstand the topic, I run it on my pc, not kaggle's kernel.</p>",
          "rawMarkdown": "Sorry, I misunderstand the topic, I run it on my pc, not kaggle's kernel."
        }
      ]
    },
    {
      "id": 311770,
      "postDate": "2018-04-10T16:38:21.277Z",
      "content": "<p>75 M, simple lgb model, few engineered features (just what had circulated around), 0.9692. Thank you for pointing to the other thread, it is very useful.   </p>\n\n<p>And good luck in the competition!</p>",
      "rawMarkdown": "75 M, simple lgb model, few engineered features (just what had circulated around), 0.9692. Thank you for pointing to the other thread, it is very useful.   \n  \nAnd good luck in the competition!",
      "votes": 1
    },
    {
      "id": 321270,
      "postDate": "2018-04-30T22:38:38.340Z",
      "content": "<p>0.9745 using XGBoost on the entire set. It seems most are using LightGBM. Does anyone have an insight on what makes LightGBM possibly better than XGBoost in this competition?</p>",
      "rawMarkdown": "0.9745 using XGBoost on the entire set. It seems most are using LightGBM. Does anyone have an insight on what makes LightGBM possibly better than XGBoost in this competition?",
      "votes": 2,
      "replies": [
        {
          "id": 321289,
          "postDate": "2018-04-30T23:50:21.897Z",
          "content": "<p>Check out this <a href=\"https://github.com/Microsoft/LightGBM/issues/211\">benchmarks</a>  which are basically from previous Kaggle competitions and compares performance of XGBoost and LightGBM. I actually, gave up working on XGBoost but I believe it's possible to go further with XGBoost histogram optimized version with proper tuning. </p>",
          "rawMarkdown": "Check out this [benchmarks][1]  which are basically from previous Kaggle competitions and compares performance of XGBoost and LightGBM. I actually, gave up working on XGBoost but I believe it's possible to go further with XGBoost histogram optimized version with proper tuning. \n\n  [1]: https://github.com/Microsoft/LightGBM/issues/211",
          "votes": 1
        },
        {
          "id": 321298,
          "postDate": "2018-05-01T00:40:19.453Z",
          "content": "<p>Thanks. I'll see then if the score improves by switching to LightGBM.</p>",
          "rawMarkdown": "Thanks. I'll see then if the score improves by switching to LightGBM.",
          "votes": 1
        },
        {
          "id": 321317,
          "postDate": "2018-05-01T02:58:52.653Z",
          "content": "<p>0.9738 with LightGBM with the exact same data, features and validation. Different hyper-parameters though.  </p>",
          "rawMarkdown": "0.9738 with LightGBM with the exact same data, features and validation. Different hyper-parameters though.  "
        }
      ]
    },
    {
      "id": 316394,
      "postDate": "2018-04-18T22:49:53.137Z",
      "content": "<p>@Joe Eddy: \nSomewhere you said, you are using 6,7 and 8  for training and day 9 for validation. When you set up your CV like that do you separately engineer feature for training and val set or do it together (6,7,8 and 9) and separate the data set later? I think doing separately makes sense as together will introduce future information into the training set? For example a groupby like this one, will you do it on 6+7+8 together and 9 separately or all together (6+7+8+9 day) for the above CV setup?</p>\n\n<pre><code>train_df[['ip','hour','channel']].groupby(by=['ip','hour'])[['channel']].count()\n</code></pre>",
      "rawMarkdown": "@Joe Eddy: \nSomewhere you said, you are using 6,7 and 8  for training and day 9 for validation. When you set up your CV like that do you separately engineer feature for training and val set or do it together (6,7,8 and 9) and separate the data set later? I think doing separately makes sense as together will introduce future information into the training set? For example a groupby like this one, will you do it on 6+7+8 together and 9 separately or all together (6+7+8+9 day) for the above CV setup?\n    \n    train_df[['ip','hour','channel']].groupby(by=['ip','hour'])[['channel']].count()",
      "votes": 2,
      "replies": [
        {
          "id": 316411,
          "postDate": "2018-04-19T01:20:09.043Z",
          "content": "<p>I break up the training data into 3 chunks that all cover the same time range as the test supplement, and run the same feature engineering loop on each chunk (and the test supplement). This ensures that count-based aggregations are fair comparisons and also makes feature engineering generally more manageable - when you start building out a lot of features, it's too RAM intensive to do it all at once (unless you have an incredible machine). </p>",
          "rawMarkdown": "I break up the training data into 3 chunks that all cover the same time range as the test supplement, and run the same feature engineering loop on each chunk (and the test supplement). This ensures that count-based aggregations are fair comparisons and also makes feature engineering generally more manageable - when you start building out a lot of features, it's too RAM intensive to do it all at once (unless you have an incredible machine). ",
          "votes": 4
        },
        {
          "id": 316413,
          "postDate": "2018-04-19T01:27:17.773Z",
          "content": "<p>Thank you so much for the answer. That make sense and i am trying that now to see if the problem of very different LB and local CV score go away. I do have a very powerful machine (256 RAM and 24 cores +GPU)</p>",
          "rawMarkdown": "Thank you so much for the answer. That make sense and i am trying that now to see if the problem of very different LB and local CV score go away. I do have a very powerful machine (256 RAM and 24 cores +GPU)",
          "votes": 2
        },
        {
          "id": 318639,
          "postDate": "2018-04-24T07:00:22.930Z",
          "content": "<p>Hi Joe, did it make a big difference when you selected the same time windows in the training set  as the test set? And why do you use the test supplement?</p>",
          "rawMarkdown": "Hi Joe, did it make a big difference when you selected the same time windows in the training set  as the test set? And why do you use the test supplement?"
        },
        {
          "id": 318769,
          "postDate": "2018-04-24T12:59:33.433Z",
          "content": "<p>The main reason I do that is for a proper validation setup - this way, when I train a model on day 8 and predict on the day 9 test hours, I feel pretty confident in my validation results. It also helps me make sure that I don't engineer features whose meaning get distorted on the test set. I've been using this setup from the beginning so I can't say whether it makes a big difference, but I think being able to trust your validation is very useful.</p>\n\n<p>I think using the test supplement is a must. Most of the features that seem to work well rely on looking over the entire range of the day. So if you only use the test set for feature engineering, you're missing a lot of information in those features and you should expect worse results. </p>",
          "rawMarkdown": "The main reason I do that is for a proper validation setup - this way, when I train a model on day 8 and predict on the day 9 test hours, I feel pretty confident in my validation results. It also helps me make sure that I don't engineer features whose meaning get distorted on the test set. I've been using this setup from the beginning so I can't say whether it makes a big difference, but I think being able to trust your validation is very useful.\n\nI think using the test supplement is a must. Most of the features that seem to work well rely on looking over the entire range of the day. So if you only use the test set for feature engineering, you're missing a lot of information in those features and you should expect worse results. ",
          "votes": 4
        },
        {
          "id": 318784,
          "postDate": "2018-04-24T13:24:31.883Z",
          "content": "<pre><code>Most of the features that seem to work well rely on looking over the entire range of the day. So if you only use the test set for feature engineering, you're missing a lot of information in those features and you should expect worse results.\n</code></pre>\n\n<p>I agree with you, but whenever I use test_supplement to make test features my LB drops by at-least 0.0004-0.0005, I ignore this calling random variation but I don't understand why the drop occurs. Have you noticed this difference when using test vs test_supplement?</p>",
          "rawMarkdown": "    Most of the features that seem to work well rely on looking over the entire range of the day. So if you only use the test set for feature engineering, you're missing a lot of information in those features and you should expect worse results.\n\nI agree with you, but whenever I use test_supplement to make test features my LB drops by at-least 0.0004-0.0005, I ignore this calling random variation but I don't understand why the drop occurs. Have you noticed this difference when using test vs test_supplement?"
        },
        {
          "id": 318826,
          "postDate": "2018-04-24T14:50:47.553Z",
          "content": "<p>That's surprising to me. I've mainly been using the supplement the entire time, but when I've added it to something done on just test I saw an improvement of at least .0007. I would double check your feature engineering/setup to see if there's something that assumes the more limited hours.</p>",
          "rawMarkdown": "That's surprising to me. I've mainly been using the supplement the entire time, but when I've added it to something done on just test I saw an improvement of at least .0007. I would double check your feature engineering/setup to see if there's something that assumes the more limited hours.",
          "votes": 1
        },
        {
          "id": 318857,
          "postDate": "2018-04-24T16:00:30.397Z",
          "content": "<p>I will review my feature engineering setup. <br> Thanks</p>",
          "rawMarkdown": "I will review my feature engineering setup. <br> Thanks"
        },
        {
          "id": 321087,
          "postDate": "2018-04-30T14:53:46.467Z",
          "content": "<p>Hi@Joe Eddy, could you elaborate a bit more on how to use the test supplement? Did you do feature engineering with train + test_supplement, and then merge engineered test_supplement with test, and then score on test? On what keys you use to merge? I think the ids in test and test_supplement are different and I don't wanna mess it up. </p>\n\n<p>Thanks in advance!!</p>",
          "rawMarkdown": "Hi@Joe Eddy, could you elaborate a bit more on how to use the test supplement? Did you do feature engineering with train + test_supplement, and then merge engineered test_supplement with test, and then score on test? On what keys you use to merge? I think the ids in test and test_supplement are different and I don't wanna mess it up. \n\nThanks in advance!!"
        },
        {
          "id": 321095,
          "postDate": "2018-04-30T15:09:28.587Z",
          "content": "<p>You can see this <a href=\"https://www.kaggle.com/alexfir/mapping-between-test-supplement-csv-and-test-csv\">notebook </a>.</p>",
          "rawMarkdown": "You can see this [notebook ][1].\n\n\n  [1]: https://www.kaggle.com/alexfir/mapping-between-test-supplement-csv-and-test-csv",
          "votes": 2
        },
        {
          "id": 321106,
          "postDate": "2018-04-30T16:04:26.310Z",
          "content": "<p>For the test features, I run my feature engineering loop on the test supplement then subset the supplement to the test clicks using Alexander Firsov's mapping - which jyyx has linked to above.</p>",
          "rawMarkdown": "For the test features, I run my feature engineering loop on the test supplement then subset the supplement to the test clicks using Alexander Firsov's mapping - which jyyx has linked to above.",
          "votes": 1
        }
      ]
    },
    {
      "id": 311848,
      "postDate": "2018-04-10T19:07:59.177Z",
      "content": "<p>My best single model is .9719 with 75M obs and 9 features using LGB in python.</p>",
      "rawMarkdown": "My best single model is .9719 with 75M obs and 9 features using LGB in python.",
      "votes": 2
    },
    {
      "id": 323952,
      "postDate": "2018-05-06T19:23:38.997Z",
      "content": "<p>0.9811 with single LGB . Trained on full data.</p>",
      "rawMarkdown": "0.9811 with single LGB . Trained on full data."
    },
    {
      "id": 323947,
      "postDate": "2018-05-06T19:10:15.300Z",
      "content": "<p>98.10 Single LGB</p>",
      "rawMarkdown": "98.10 Single LGB"
    },
    {
      "id": 321290,
      "postDate": "2018-05-01T00:03:41.470Z",
      "content": "<p>0.9797 with single LGB.</p>",
      "rawMarkdown": "0.9797 with single LGB."
    },
    {
      "id": 318734,
      "postDate": "2018-04-24T11:15:24.643Z",
      "content": "<p>0.9792 single LGB 12 features train on 55M rows</p>",
      "rawMarkdown": "0.9792 single LGB 12 features train on 55M rows"
    },
    {
      "id": 318371,
      "postDate": "2018-04-23T17:27:17.653Z",
      "content": "<p>9784 with 11 features using LGB. Trained on 40 M rows. </p>",
      "rawMarkdown": "9784 with 11 features using LGB. Trained on 40 M rows. "
    },
    {
      "id": 318234,
      "postDate": "2018-04-23T13:02:02.057Z",
      "content": "<p>9801,  lgbm</p>",
      "rawMarkdown": "9801,  lgbm"
    },
    {
      "id": 316410,
      "postDate": "2018-04-19T00:50:35.837Z",
      "content": "<p>Interesting question.</p>",
      "rawMarkdown": "Interesting question."
    },
    {
      "id": 312895,
      "postDate": "2018-04-12T14:07:09.783Z",
      "content": "<p>very nice</p>",
      "rawMarkdown": "very nice"
    },
    {
      "id": 312786,
      "postDate": "2018-04-12T10:32:26.313Z",
      "content": "<p>0.9754, single LGB, train on day 7,8 and validate on day 9.</p>",
      "rawMarkdown": "0.9754, single LGB, train on day 7,8 and validate on day 9.",
      "replies": [
        {
          "id": 312790,
          "postDate": "2018-04-12T10:44:07.453Z",
          "content": "<p>All in the kernel?! Averaging runs or managing to go through 120 million rows?</p>",
          "rawMarkdown": "All in the kernel?! Averaging runs or managing to go through 120 million rows?"
        },
        {
          "id": 312850,
          "postDate": "2018-04-12T13:03:07.433Z",
          "content": "<p>Sorry, I train on my local machine, No averaging.</p>",
          "rawMarkdown": "Sorry, I train on my local machine, No averaging."
        },
        {
          "id": 313339,
          "postDate": "2018-04-13T05:58:55.587Z",
          "content": "<p>Do you validate on whole 9 day?</p>",
          "rawMarkdown": "Do you validate on whole 9 day?"
        },
        {
          "id": 313402,
          "postDate": "2018-04-13T07:57:50.447Z",
          "content": "<p>@AlexTru\nI validate on whole day 9, then on hour 4 of day 9, then on private hours of day 9.</p>",
          "rawMarkdown": "@AlexTru\nI validate on whole day 9, then on hour 4 of day 9, then on private hours of day 9.",
          "votes": 1
        },
        {
          "id": 314802,
          "postDate": "2018-04-16T09:48:32.997Z",
          "content": "<p>hi,I'm new here, what's the meaning of 'day 7,8'?</p>",
          "rawMarkdown": "hi,I'm new here, what's the meaning of 'day 7,8'?"
        },
        {
          "id": 314954,
          "postDate": "2018-04-16T14:48:56.647Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 312447,
      "postDate": "2018-04-11T19:16:50.590Z",
      "content": "<p>Very impressive results! Colleagues, what validation scheme did you use? random X% from split? And did you just use last NN mln rows from train for training?</p>",
      "rawMarkdown": "Very impressive results! Colleagues, what validation scheme did you use? random X% from split? And did you just use last NN mln rows from train for training?",
      "replies": [
        {
          "id": 312494,
          "postDate": "2018-04-11T21:12:07.680Z",
          "content": "<p>I believe, most of us (kernel model users) are using:</p>\n\n<ul>\n<li>last observations that are <strong>close to test</strong>  as a training set</li>\n<li>validation set is <strong>5% or 10%</strong>  and either close to test observations or shuffled split</li>\n</ul>",
          "rawMarkdown": "I believe, most of us (kernel model users) are using:\n\n- last observations that are **close to test**  as a training set\n- validation set is **5% or 10%**  and either close to test observations or shuffled split",
          "votes": 4
        },
        {
          "id": 312567,
          "postDate": "2018-04-12T01:48:30.590Z",
          "content": "<p>Hi, you can check this link:\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325</a></p>",
          "rawMarkdown": "Hi, you can check this link:\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325",
          "votes": 1
        }
      ]
    },
    {
      "id": 311781,
      "postDate": "2018-04-10T16:55:36.800Z",
      "content": "<p>I wonder who the top 20 have score less than 0.001 from a single model</p>",
      "rawMarkdown": "I wonder who the top 20 have score less than 0.001 from a single model\n\n"
    }
  ],
  "comments": [
    {
      "id": 317344,
      "author_name": "spongebob",
      "author_url": "",
      "post_date": "2018-04-21T08:41:02.237000",
      "content": "<p>9820，single lgb</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 312004,
      "author_name": "Krishna",
      "author_url": "",
      "post_date": "2018-04-11T03:40:39.747000",
      "content": "<p>40 Million training rows, R lightgbm, 4 new features added to Pranav's public kernel. LB 0.9736. Available in Public Kaggle kernel.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 312227,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-11T12:26:39.047000",
          "content": "<p>Great job! Thanks for sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 314360,
      "author_name": "astraldawn",
      "author_url": "",
      "post_date": "2018-04-15T10:54:01.790000",
      "content": "<p>0.979(5-9) with single lgb, (7-8) engineered features + different parameters</p>\n\n<p>the engineered features and parameters changed over runs but I think the range should be stable.</p>\n\n<p>Update: new model at .9803</p>",
      "votes": 10,
      "replies": [
        {
          "id": 314586,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-15T21:52:35.547000",
          "content": "<p>Awesome!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314620,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-16T00:17:53.427000",
          "content": "<p>Congrats! What are the typical number of rounds that you guys are using to train your lgbm models? I've been using ~1500 and am not sure if it's enough</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314624,
          "author_name": "astraldawn",
          "author_url": "",
          "post_date": "2018-04-16T00:40:25.417000",
          "content": "<p>it depends on your features + parameters. i use 600 rounds (because my kernel is derived from Andy Harless' kernel). my attempts to tune parameters gave me a range of 300 - 750 but the improvements were marginal (1 point at the 4th decimal place) so i just stuck with 600.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315456,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-17T06:21:10.307000",
          "content": "<p>Do you use the common day 9 hour 4 as CV?\nOr you generalize your cross validation and ignore the Public LB of hour 4?\nWhat is your CV AUC when training?</p>\n\n<p>Thanks!!</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 317473,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-04-21T17:20:44.433000",
          "content": "<p>It doesn't say much as it is depends on the learning rate, features and hyperparamers such as subsample. but ~240.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 321192,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-04-30T19:40:42.057000",
      "content": "<p>0.9801 NN model run on Kaggle GPU in less than 1 hour..but hard to blend for now </p>",
      "votes": 7,
      "replies": [
        {
          "id": 324520,
          "author_name": "Meyk",
          "author_url": "",
          "post_date": "2018-05-07T19:10:13.393000",
          "content": "<p>Hi!</p>\n\n<p>I am really interested in your NN model. Can you share your knowledge when competition ends?\nMax what I can achieve is 0.9764 NN that is 0.0004 better than public kernel :(  I will surely try to improve my score on NN when competition ends and I wil have more time for it ;)</p>\n\n<p>Cheers and good luck! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 316468,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2018-04-19T06:24:10.870000",
      "content": "<p>9806</p>",
      "votes": 5,
      "replies": [
        {
          "id": 316546,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-19T10:06:21.563000",
          "content": "<p>Great, wonder what guys above 0.98xx are doing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316798,
          "author_name": "Cheng",
          "author_url": "",
          "post_date": "2018-04-19T23:56:26.363000",
          "content": "<p>Is this trained on the whole dataset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316802,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-04-20T00:11:45.177000",
          "content": "<p>no, just one day   </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 316809,
          "author_name": "Cheng",
          "author_url": "",
          "post_date": "2018-04-20T00:47:18.333000",
          "content": "<p>Nice work! It is impressive</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316835,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-04-20T03:35:28.210000",
          "content": "<p>how about you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316836,
          "author_name": "Cheng",
          "author_url": "",
          "post_date": "2018-04-20T03:36:48.690000",
          "content": "<p>I always train with full data so don't know the performance for subset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316915,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-04-20T07:19:43.963000",
          "content": "<p>so the full data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318633,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-04-24T06:53:07.200000",
          "content": "<p>Cheng how many extra features do you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318640,
          "author_name": "Cheng",
          "author_url": "",
          "post_date": "2018-04-24T07:02:29.833000",
          "content": "<p>less than 20</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318847,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-04-24T15:32:57.267000",
          "content": "<p>who can more than。。。</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318959,
          "author_name": "Siyuan Dang",
          "author_url": "",
          "post_date": "2018-04-24T21:40:27.597000",
          "content": "<p>Hi,  Peter. Do you mind tell me which day did you use?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 323974,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2018-05-06T20:18:58.553000",
      "content": "<p>0.9819 with single LightGBM model - we entered the game too late, and probably don't have time to do stacking though :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 323976,
          "author_name": "Nithinthakur",
          "author_url": "",
          "post_date": "2018-05-06T20:22:21.630000",
          "content": "<p>@Yifan Xie how many features do you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323978,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2018-05-06T20:27:18.963000",
          "content": "<p>36</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 323994,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-05-06T21:46:21.633000",
          "content": "<p>@Yifan Xie That is impressive! May I know how you/your team handle ram while building the model in kernel? By my limited cs knowledge, 9 new features are the maximum I can generate in kernel.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323996,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-06T21:51:57.030000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323998,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-05-06T21:56:32.607000",
          "content": "<p>Thanks for the quick response. I do hope I can have a decent computer because my local machine only has 4gb RAM.  Kaggle kernel is the best resource I have for this competition. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 324187,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-07T10:32:08.700000",
          "content": "<p>@Yifan,  you run this model on Kaggle kernels?  If so then this is really impressive.  If you run it on your own machine then you may want to compare with what people report in this other thread.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 324191,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2018-05-07T10:38:45.147000",
          "content": "<p>lol, didn't pay attention this thread is for kernel only, yup I will have a go at that after the competition. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 312791,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-04-12T10:45:24.233000",
      "content": "<p>My update is that I managed to hit .9745 in a kernel with 7 engineered features. </p>\n\n<p>.9765 with only 4 engineered features..... maybe within reach!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 312797,
          "author_name": "jyyx",
          "author_url": "",
          "post_date": "2018-04-12T10:56:24.767000",
          "content": "<p>I also think that this is possible. I got 0.9761 in a kernel  with 7  engineered features.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 312936,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-12T14:50:09.087000",
          "content": "<p>Nice!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313103,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-12T19:26:43.957000",
          "content": "<p>.9787 is possible in a kernel with 7 engineered features :)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 313107,
          "author_name": "jyyx",
          "author_url": "",
          "post_date": "2018-04-12T19:34:40.523000",
          "content": "<p>Great job!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313116,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-12T19:44:27.350000",
          "content": "<p>That's Amazing @Joe, I have more than 10 features and I can not figure out how to cross 0.975x. <br> Looking forward to your feature engineering explanation.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 313129,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-12T20:11:52.110000",
          "content": "<blockquote>\n  <p><strong>Joe Eddy wrote</strong></p>\n  \n  <blockquote>\n    <p>.9787 is possible in a kernel with 7 engineered features :)</p>\n  </blockquote>\n</blockquote>\n\n<p>Wow! that's impressive. May I ask if you're using R/Python?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 313138,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-12T20:29:16.420000",
          "content": "<p>@Sohaib @Pranav Thanks both!</p>\n\n<p>I definitely plan to share my approach after the competition ends.</p>\n\n<p>I'm using python. And actually this kernel was built on top of the python port of your R kernel, so thank you for that awesome work.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 313140,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-12T20:33:06.250000",
          "content": "<p>My pleasure! And thank you so much for the kind words. I really glad to hear that. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 313333,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-04-13T05:52:06.987000",
          "content": "<p>may I ask you how you deal with ram explore while building new features? I have 5 features other than the base, ones, when I add the 6th new feature, kernel died.   </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313489,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-13T10:33:34.203000",
          "content": "<p>@MengYe, If you're adding new features then best would be to reduce chunk size (observations) that can pass through without hitting memory limit. You may find <a href=\"https://www.kaggle.com/pranav84/double-xgb-hist-example-for-chunk-processing\">this kernel</a>  helpful in this regard. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313823,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-13T22:07:32.703000",
          "content": "<p>@Joe Eddy congrats! May I ask on which subset of the training data you trained that model on?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314587,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-15T21:53:02.127000",
          "content": "<p>@Edward thanks! That model was just trained on day 9.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 311955,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2018-04-11T01:17:45.030000",
      "content": "<p>0.9720 with 45 mln observations, 14 features, single model lgb in R. Tried to add a few more promising features and hit the kernel memory limit. Planning to try Python instead.</p>\n\n<p>Update: Python helped a little bit (plus some feature engineering, of course). The current result is 0.9789.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 311800,
      "author_name": "Pranav Pandya",
      "author_url": "",
      "post_date": "2018-04-10T17:37:53.580000",
      "content": "<p>0.9723 with 75M observations (single LGB in R). </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 311779,
      "author_name": "jyyx",
      "author_url": "",
      "post_date": "2018-04-10T16:53:03.120000",
      "content": "<p>62M, simple lgb model, 15 features, 0.9743, run on the kernel. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 311761,
      "author_name": "Zishi",
      "author_url": "",
      "post_date": "2018-04-10T16:25:38.340000",
      "content": "<p>My PB is 0.9724 on 75 millions training rows with 14 features including some of original ones. Still trying to hit higher LB score before diving into blending/stacking. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 325030,
      "author_name": "MengYe",
      "author_url": "",
      "post_date": "2018-05-08T03:35:03.663000",
      "content": "<p>The competition has ended (with disappointment), may I know how you guys build models and achieve high scores with kernel all the way through? I really want to know because I too used kernel from the begining till the end. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 317471,
      "author_name": "Yair Beer",
      "author_url": "",
      "post_date": "2018-04-21T17:18:41.833000",
      "content": "<p>0.9792 single lgb bested on baris Kanbar's kernel.</p>\n\n<p>I had to do better CV for feature selection, but I still use the last 5M rows for early stopping trigger.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 317175,
      "author_name": "Nithinthakur",
      "author_url": "",
      "post_date": "2018-04-21T00:18:35.823000",
      "content": "<p>0.9793, single lgb, 22 engineered features</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 316453,
      "author_name": "Gabriel Preda",
      "author_url": "",
      "post_date": "2018-04-19T04:47:33.760000",
      "content": "<p>0.9787, single lgb, few engineered features, 75M training, 5 M validation</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 311992,
      "author_name": "Marcus Lin",
      "author_url": "",
      "post_date": "2018-04-11T03:04:34.617000",
      "content": "<p>0.9716 on PB with Full observations using lightgbm.  It's so frustrated found someone is over .98.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 311999,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-11T03:26:21.307000",
          "content": "<p>What method are you using to run the full data? Are you chunking?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312002,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2018-04-11T03:33:39.353000",
          "content": "<p>Sorry, I misunderstand the topic, I run it on my pc, not kaggle's kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 311770,
      "author_name": "Gabriel Preda",
      "author_url": "",
      "post_date": "2018-04-10T16:38:21.277000",
      "content": "<p>75 M, simple lgb model, few engineered features (just what had circulated around), 0.9692. Thank you for pointing to the other thread, it is very useful.   </p>\n\n<p>And good luck in the competition!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 321270,
      "author_name": "Oscar Takeshita",
      "author_url": "",
      "post_date": "2018-04-30T22:38:38.340000",
      "content": "<p>0.9745 using XGBoost on the entire set. It seems most are using LightGBM. Does anyone have an insight on what makes LightGBM possibly better than XGBoost in this competition?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 321289,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-30T23:50:21.897000",
          "content": "<p>Check out this <a href=\"https://github.com/Microsoft/LightGBM/issues/211\">benchmarks</a>  which are basically from previous Kaggle competitions and compares performance of XGBoost and LightGBM. I actually, gave up working on XGBoost but I believe it's possible to go further with XGBoost histogram optimized version with proper tuning. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 321298,
          "author_name": "Oscar Takeshita",
          "author_url": "",
          "post_date": "2018-05-01T00:40:19.453000",
          "content": "<p>Thanks. I'll see then if the score improves by switching to LightGBM.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 321317,
          "author_name": "Oscar Takeshita",
          "author_url": "",
          "post_date": "2018-05-01T02:58:52.653000",
          "content": "<p>0.9738 with LightGBM with the exact same data, features and validation. Different hyper-parameters though.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 316394,
      "author_name": "Anon",
      "author_url": "",
      "post_date": "2018-04-18T22:49:53.137000",
      "content": "<p>@Joe Eddy: \nSomewhere you said, you are using 6,7 and 8  for training and day 9 for validation. When you set up your CV like that do you separately engineer feature for training and val set or do it together (6,7,8 and 9) and separate the data set later? I think doing separately makes sense as together will introduce future information into the training set? For example a groupby like this one, will you do it on 6+7+8 together and 9 separately or all together (6+7+8+9 day) for the above CV setup?</p>\n\n<pre><code>train_df[['ip','hour','channel']].groupby(by=['ip','hour'])[['channel']].count()\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 316411,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-19T01:20:09.043000",
          "content": "<p>I break up the training data into 3 chunks that all cover the same time range as the test supplement, and run the same feature engineering loop on each chunk (and the test supplement). This ensures that count-based aggregations are fair comparisons and also makes feature engineering generally more manageable - when you start building out a lot of features, it's too RAM intensive to do it all at once (unless you have an incredible machine). </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 316413,
          "author_name": "Anon",
          "author_url": "",
          "post_date": "2018-04-19T01:27:17.773000",
          "content": "<p>Thank you so much for the answer. That make sense and i am trying that now to see if the problem of very different LB and local CV score go away. I do have a very powerful machine (256 RAM and 24 cores +GPU)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 318639,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-04-24T07:00:22.930000",
          "content": "<p>Hi Joe, did it make a big difference when you selected the same time windows in the training set  as the test set? And why do you use the test supplement?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318769,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-24T12:59:33.433000",
          "content": "<p>The main reason I do that is for a proper validation setup - this way, when I train a model on day 8 and predict on the day 9 test hours, I feel pretty confident in my validation results. It also helps me make sure that I don't engineer features whose meaning get distorted on the test set. I've been using this setup from the beginning so I can't say whether it makes a big difference, but I think being able to trust your validation is very useful.</p>\n\n<p>I think using the test supplement is a must. Most of the features that seem to work well rely on looking over the entire range of the day. So if you only use the test set for feature engineering, you're missing a lot of information in those features and you should expect worse results. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 318784,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-24T13:24:31.883000",
          "content": "<pre><code>Most of the features that seem to work well rely on looking over the entire range of the day. So if you only use the test set for feature engineering, you're missing a lot of information in those features and you should expect worse results.\n</code></pre>\n\n<p>I agree with you, but whenever I use test_supplement to make test features my LB drops by at-least 0.0004-0.0005, I ignore this calling random variation but I don't understand why the drop occurs. Have you noticed this difference when using test vs test_supplement?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318826,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-24T14:50:47.553000",
          "content": "<p>That's surprising to me. I've mainly been using the supplement the entire time, but when I've added it to something done on just test I saw an improvement of at least .0007. I would double check your feature engineering/setup to see if there's something that assumes the more limited hours.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318857,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-24T16:00:30.397000",
          "content": "<p>I will review my feature engineering setup. <br> Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321087,
          "author_name": "Rui Li",
          "author_url": "",
          "post_date": "2018-04-30T14:53:46.467000",
          "content": "<p>Hi@Joe Eddy, could you elaborate a bit more on how to use the test supplement? Did you do feature engineering with train + test_supplement, and then merge engineered test_supplement with test, and then score on test? On what keys you use to merge? I think the ids in test and test_supplement are different and I don't wanna mess it up. </p>\n\n<p>Thanks in advance!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321095,
          "author_name": "jyyx",
          "author_url": "",
          "post_date": "2018-04-30T15:09:28.587000",
          "content": "<p>You can see this <a href=\"https://www.kaggle.com/alexfir/mapping-between-test-supplement-csv-and-test-csv\">notebook </a>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 321106,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-30T16:04:26.310000",
          "content": "<p>For the test features, I run my feature engineering loop on the test supplement then subset the supplement to the test clicks using Alexander Firsov's mapping - which jyyx has linked to above.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 311848,
      "author_name": "shivraj",
      "author_url": "",
      "post_date": "2018-04-10T19:07:59.177000",
      "content": "<p>My best single model is .9719 with 75M obs and 9 features using LGB in python.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 323952,
      "author_name": "Nithinthakur",
      "author_url": "",
      "post_date": "2018-05-06T19:23:38.997000",
      "content": "<p>0.9811 with single LGB . Trained on full data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 323947,
      "author_name": "Akarsh Raj",
      "author_url": "",
      "post_date": "2018-05-06T19:10:15.300000",
      "content": "<p>98.10 Single LGB</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 321290,
      "author_name": "Amit Kumar Jaiswal",
      "author_url": "",
      "post_date": "2018-05-01T00:03:41.470000",
      "content": "<p>0.9797 with single LGB.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 318734,
      "author_name": "HuyenNguyen",
      "author_url": "",
      "post_date": "2018-04-24T11:15:24.643000",
      "content": "<p>0.9792 single LGB 12 features train on 55M rows</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 318371,
      "author_name": "Sudeep Shukla",
      "author_url": "",
      "post_date": "2018-04-23T17:27:17.653000",
      "content": "<p>9784 with 11 features using LGB. Trained on 40 M rows. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 318234,
      "author_name": "Zishi",
      "author_url": "",
      "post_date": "2018-04-23T13:02:02.057000",
      "content": "<p>9801,  lgbm</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 316410,
      "author_name": "Artur Quirino",
      "author_url": "",
      "post_date": "2018-04-19T00:50:35.837000",
      "content": "<p>Interesting question.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 312895,
      "author_name": "Sai Bharath Dhanekula",
      "author_url": "",
      "post_date": "2018-04-12T14:07:09.783000",
      "content": "<p>very nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 312786,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-04-12T10:32:26.313000",
      "content": "<p>0.9754, single LGB, train on day 7,8 and validate on day 9.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 312790,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-12T10:44:07.453000",
          "content": "<p>All in the kernel?! Averaging runs or managing to go through 120 million rows?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312850,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-12T13:03:07.433000",
          "content": "<p>Sorry, I train on my local machine, No averaging.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313339,
          "author_name": "AlexTru",
          "author_url": "",
          "post_date": "2018-04-13T05:58:55.587000",
          "content": "<p>Do you validate on whole 9 day?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 313402,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-13T07:57:50.447000",
          "content": "<p>@AlexTru\nI validate on whole day 9, then on hour 4 of day 9, then on private hours of day 9.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 314802,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-04-16T09:48:32.997000",
          "content": "<p>hi,I'm new here, what's the meaning of 'day 7,8'?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314954,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-16T14:48:56.647000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 312447,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-11T19:16:50.590000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 312494,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-11T21:12:07.680000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 312567,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-12T01:48:30.590000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 311781,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-10T16:55:36.800000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "311758": "Already [here][1] is a good discussion of single model scores.\n\nWhat I'm also interested in is how far people have been able to push the limits of kaggle kernel computation for single models, beyond visible public kernel scores. \n\nThe best I've seen that's publicly visible is [anttip's awesome FM_FTRL kernel][2] (.9711), and my personal best is a lgbm kernel on 50 million rows with 5 engineered features (.9707). \n\nCurious what others have found possible!\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\n  [2]: https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711",
    "317344": "9820，single lgb",
    "312004": "40 Million training rows, R lightgbm, 4 new features added to Pranav's public kernel. LB 0.9736. Available in Public Kaggle kernel.",
    "314360": "0.979(5-9) with single lgb, (7-8) engineered features + different parameters\n\nthe engineered features and parameters changed over runs but I think the range should be stable.\n\nUpdate: new model at .9803",
    "321192": "0.9801 NN model run on Kaggle GPU in less than 1 hour..but hard to blend for now ",
    "316468": "9806",
    "323974": "0.9819 with single LightGBM model - we entered the game too late, and probably don't have time to do stacking though :)",
    "312791": "My update is that I managed to hit .9745 in a kernel with 7 engineered features. \n\n.9765 with only 4 engineered features..... maybe within reach!",
    "311955": "0.9720 with 45 mln observations, 14 features, single model lgb in R. Tried to add a few more promising features and hit the kernel memory limit. Planning to try Python instead.\n\nUpdate: Python helped a little bit (plus some feature engineering, of course). The current result is 0.9789.",
    "311800": "0.9723 with 75M observations (single LGB in R). ",
    "311779": "62M, simple lgb model, 15 features, 0.9743, run on the kernel. ",
    "311761": "My PB is 0.9724 on 75 millions training rows with 14 features including some of original ones. Still trying to hit higher LB score before diving into blending/stacking. ",
    "325030": "The competition has ended (with disappointment), may I know how you guys build models and achieve high scores with kernel all the way through? I really want to know because I too used kernel from the begining till the end. ",
    "317471": "0.9792 single lgb bested on baris Kanbar's kernel.\n\nI had to do better CV for feature selection, but I still use the last 5M rows for early stopping trigger.",
    "317175": "0.9793, single lgb, 22 engineered features",
    "316453": "0.9787, single lgb, few engineered features, 75M training, 5 M validation",
    "311992": "0.9716 on PB with Full observations using lightgbm.  It's so frustrated found someone is over .98.",
    "311770": "75 M, simple lgb model, few engineered features (just what had circulated around), 0.9692. Thank you for pointing to the other thread, it is very useful.   \n  \nAnd good luck in the competition!",
    "321270": "0.9745 using XGBoost on the entire set. It seems most are using LightGBM. Does anyone have an insight on what makes LightGBM possibly better than XGBoost in this competition?",
    "316394": "@Joe Eddy: \nSomewhere you said, you are using 6,7 and 8  for training and day 9 for validation. When you set up your CV like that do you separately engineer feature for training and val set or do it together (6,7,8 and 9) and separate the data set later? I think doing separately makes sense as together will introduce future information into the training set? For example a groupby like this one, will you do it on 6+7+8 together and 9 separately or all together (6+7+8+9 day) for the above CV setup?\n    \n    train_df[['ip','hour','channel']].groupby(by=['ip','hour'])[['channel']].count()",
    "311848": "My best single model is .9719 with 75M obs and 9 features using LGB in python.",
    "323952": "0.9811 with single LGB . Trained on full data.",
    "323947": "98.10 Single LGB",
    "321290": "0.9797 with single LGB.",
    "318734": "0.9792 single LGB 12 features train on 55M rows",
    "318371": "9784 with 11 features using LGB. Trained on 40 M rows. ",
    "318234": "9801,  lgbm",
    "316410": "Interesting question.",
    "312895": "very nice",
    "312786": "0.9754, single LGB, train on day 7,8 and validate on day 9.",
    "312447": "Very impressive results! Colleagues, what validation scheme did you use? random X% from split? And did you just use last NN mln rows from train for training?",
    "311781": "I wonder who the top 20 have score less than 0.001 from a single model\n\n"
  }
}