{
  "id": 56205,
  "title": "Does Anyone Have a 0.9835 Kernel?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56205",
  "author_name": "",
  "post_date": "2018-05-07T18:31:07.811051800Z",
  "votes": 35,
  "comment_count": 55,
  "views": 0,
  "content": "<p>We are preparing our final submission, and would hate to be excluded from potentially sharing the first place in case someone submits such a gem at the last moment. Thanks. </p>",
  "messages": [
    {
      "id": "324484",
      "postDate": "05/07/2018 18:31:07",
      "content": "<p>We are preparing our final submission, and would hate to be excluded from potentially sharing the first place in case someone submits such a gem at the last moment. Thanks. </p>",
      "rawMarkdown": "We are preparing our final submission, and would hate to be excluded from potentially sharing the first place in case someone submits such a gem at the last moment. Thanks.",
      "votes": null
    },
    {
      "id": "324486",
      "postDate": "05/07/2018 18:36:04",
      "content": "<p>What about 0.9836 kernels? Why be so inclusive to 0.9835 kernels?</p>",
      "rawMarkdown": "What about 0.9836 kernels? Why be so inclusive to 0.9835 kernels?",
      "votes": null
    },
    {
      "id": "324492",
      "postDate": "05/07/2018 18:43:15",
      "content": "<p>Wait for it. Almost there.</p>",
      "rawMarkdown": "Wait for it. Almost there.",
      "votes": null
    },
    {
      "id": "324507",
      "postDate": "05/07/2018 18:54:26",
      "content": "<p>My intuition tells me that 0.9835 is the absolute theoretical limit for this dataset.</p>",
      "rawMarkdown": "My intuition tells me that 0.9835 is the absolute theoretical limit for this dataset.",
      "votes": null
    },
    {
      "id": "324510",
      "postDate": "05/07/2018 18:58:51",
      "content": "<p><a href=\"/bestfitting\">@bestfitting</a> is going to pull something :-)</p>",
      "rawMarkdown": "bestfitting is going to pull something :-)",
      "votes": null
    },
    {
      "id": "324512",
      "postDate": "05/07/2018 18:59:25",
      "content": "<p>I'm training the full dataset with 1M features. once I've done, I put there !</p>",
      "rawMarkdown": "I'm training the full dataset with 1M features. once I've done, I put there !",
      "votes": null
    },
    {
      "id": "324516",
      "postDate": "05/07/2018 19:04:43",
      "content": "<p>The kernel would probably be useless ....maybe just sharing 0.9835 submission file ^^</p>",
      "rawMarkdown": "The kernel would probably be useless ....maybe just sharing 0.9835 submission file ^^",
      "votes": null
    },
    {
      "id": "324519",
      "postDate": "05/07/2018 19:09:35",
      "content": "<p>How often your intuition is right?</p>",
      "rawMarkdown": "How often your intuition is right?",
      "votes": null
    },
    {
      "id": "324525",
      "postDate": "05/07/2018 19:13:20",
      "content": "<p>Normally Bojan would have the probability of 200%  to be right.</p>",
      "rawMarkdown": "Normally Bojan would have the probability of 200%  to be right.",
      "votes": null
    },
    {
      "id": "324530",
      "postDate": "05/07/2018 19:19:55",
      "content": "<p>@Bojan In a more serious matter, what's your intuition for the private score? Will the wining score be below, above or at 0.9835? Strange nobody posted anything about shake ups yet.</p>",
      "rawMarkdown": "Bojan In a more serious matter, what's your intuition for the private score? Will the wining score be below, above or at 0.9835? Strange nobody posted anything about shake ups yet.",
      "votes": null
    },
    {
      "id": "324531",
      "postDate": "05/07/2018 19:20:23",
      "content": "<p>I have a notebook that has a local AUC of 0.9978799432371899 ... but it's useless :x</p>",
      "rawMarkdown": "I have a notebook that has a local AUC of 0.9978799432371899 ... but it's useless :x",
      "votes": null
    },
    {
      "id": "324533",
      "postDate": "05/07/2018 19:21:36",
      "content": "<p>@Oscar Takeshita, I guess the winning private lb score should be above 0.9860+</p>",
      "rawMarkdown": "Oscar Takeshita, I guess the winning private lb score should be above 0.9860+",
      "votes": null
    },
    {
      "id": "324536",
      "postDate": "05/07/2018 19:23:22",
      "content": "<p>@Snorlax My bet is below 0.9835. </p>",
      "rawMarkdown": "Snorlax My bet is below 0.9835.",
      "votes": null
    },
    {
      "id": "324543",
      "postDate": "05/07/2018 19:28:21",
      "content": "<p>I bet .9855-.986</p>",
      "rawMarkdown": "I bet .9855-.986",
      "votes": null
    },
    {
      "id": "324546",
      "postDate": "05/07/2018 19:29:42",
      "content": "<p>We did a bit of correlation analysis. Correlation among submissions drops as time progresses. I therefore expect private scores to be lower, understanding that the public scores are based on the beginning of the test set. </p>",
      "rawMarkdown": "We did a bit of correlation analysis. Correlation among submissions drops as time progresses. I therefore expect private scores to be lower, understanding that the public scores are based on the beginning of the test set.",
      "votes": null
    },
    {
      "id": "324550",
      "postDate": "05/07/2018 19:33:07",
      "content": "<p>@Snorlax @Joe Eddy That's interesting. Probably we shouldn't discuss why now but I'd like to hear the reasons after the competition.</p>",
      "rawMarkdown": "Snorlax @Joe Eddy That's interesting. Probably we shouldn't discuss why now but I'd like to hear the reasons after the competition.",
      "votes": null
    },
    {
      "id": "324554",
      "postDate": "05/07/2018 19:36:35",
      "content": "<p>almost finished ... don't waste your last submission lol</p>",
      "rawMarkdown": "almost finished ... don't waste your last submission lol",
      "votes": null
    },
    {
      "id": "324555",
      "postDate": "05/07/2018 19:36:36",
      "content": "<p>@Joe Eddy, happy to see you here, you are the one I want to thank to in this competition~ Would you mind sharing your estimation of your private lb score here? Here is mine, estimation of public lb = 0.9811, estimation of private lb = 0.9840, and the actual public lb score is 0.9812. I hope my validation schema works, then I would have a higher private lb compared with public lb score.</p>",
      "rawMarkdown": "Joe Eddy, happy to see you here, you are the one I want to thank to in this competition~ Would you mind sharing your estimation of your private lb score here? Here is mine, estimation of public lb = 0.9811, estimation of private lb = 0.9840, and the actual public lb score is 0.9812. I hope my validation schema works, then I would have a higher private lb compared with public lb score.",
      "votes": null
    },
    {
      "id": "324557",
      "postDate": "05/07/2018 19:38:30",
      "content": "<p>@ Bojan \nLOL..Good one</p>",
      "rawMarkdown": "Bojan \nLOL..Good one",
      "votes": null
    },
    {
      "id": "324558",
      "postDate": "05/07/2018 19:40:03",
      "content": "<p>Hahhahha!</p>",
      "rawMarkdown": "Hahhahha!",
      "votes": null
    },
    {
      "id": "324563",
      "postDate": "05/07/2018 19:48:09",
      "content": "<p>Working on it, just a sec...</p>",
      "rawMarkdown": "Working on it, just a sec...",
      "votes": null
    },
    {
      "id": "324599",
      "postDate": "05/07/2018 21:00:30",
      "content": "<p>@Oscar I'm happy to share my thoughts on that now, especially since the ideas have already come up in multiple other threads. @Snorlax it sounds like my best single model has very similar looking validation to yours (see below), then I'd estimate + ~.0006 or so from ensembling.</p>\n\n<p>The basis for my top score estimate comes from my validation setup. For this problem, a natural validation path is to train your model on days 7-8 and predict on the test hours in day 9. Doing that, I've seen hour 4 validation consistently fall within .0004 of public LB score, and I expect private LB hours validation to show similar consistency with actual private LB scores. I also would bet that top scorers use a similar validation style. </p>\n\n<p>So here's my (very unscientific) reasoning for my guess. My best single model validates at roughly .984 / .981 (public LB .9812). The gap between my best public LB and top of leaderboard is .0022, so if I extrapolate that to the private hours I get .9862. Toning that down to account for some diminishing returns on the larger sample, I estimate .9855-.986 for the top scorers. </p>",
      "rawMarkdown": "Oscar I'm happy to share my thoughts on that now, especially since the ideas have already come up in multiple other threads. @Snorlax it sounds like my best single model has very similar looking validation to yours (see below), then I'd estimate + ~.0006 or so from ensembling.\n\nThe basis for my top score estimate comes from my validation setup. For this problem, a natural validation path is to train your model on days 7-8 and predict on the test hours in day 9. Doing that, I've seen hour 4 validation consistently fall within .0004 of public LB score, and I expect private LB hours validation to show similar consistency with actual private LB scores. I also would bet that top scorers use a similar validation style. \n\nSo here's my (very unscientific) reasoning for my guess. My best single model validates at roughly .984 / .981 (public LB .9812). The gap between my best public LB and top of leaderboard is .0022, so if I extrapolate that to the private hours I get .9862. Toning that down to account for some diminishing returns on the larger sample, I estimate .9855-.986 for the top scorers.",
      "votes": null
    },
    {
      "id": "324603",
      "postDate": "05/07/2018 21:04:43",
      "content": "<p>Is there any reason or any result that shows 0.9835 is the limit for this dataset?</p>",
      "rawMarkdown": "Is there any reason or any result that shows 0.9835 is the limit for this dataset?",
      "votes": null
    },
    {
      "id": "324604",
      "postDate": "05/07/2018 21:07:18",
      "content": "<p>In all honesty, I don't have a good estimate or intuition for this one. Due to the huge imbalance in data, and out-of time test set, it has been hard to get any \"gut level\" feelings for this one. Blending has been tough, and only started working more or less consistently once our models started approaching 0.98. This to me suggests that minute deviations from the correct ordering could have really serious consequences. I fear that we might be in for a really big shakeup. </p>",
      "rawMarkdown": "In all honesty, I don't have a good estimate or intuition for this one. Due to the huge imbalance in data, and out-of time test set, it has been hard to get any \"gut level\" feelings for this one. Blending has been tough, and only started working more or less consistently once our models started approaching 0.98. This to me suggests that minute deviations from the correct ordering could have really serious consequences. I fear that we might be in for a really big shakeup.",
      "votes": null
    },
    {
      "id": "324618",
      "postDate": "05/07/2018 21:22:08",
      "content": "<p>Sadly we didn't manage to blend well Neural nets model even above 0.98xx</p>",
      "rawMarkdown": "Sadly we didn't manage to blend well Neural nets model even above 0.98xx",
      "votes": null
    },
    {
      "id": "324620",
      "postDate": "05/07/2018 21:24:49",
      "content": "<p>What is going on?</p>\n\n<p>I was in a meeting all morning, in between talks , I was glancing at my labtop checking the completion of my last model. All done, CV looked good, I did my ensemble and submitted with my fingers crossed as this is my last submission. Saw the LB and thought it was a mistake but had to keep quite and go on with the meeting. I just came out and realized there was a bomb shell kernel? This is really crazy. I am one of those with very weak HW and am relying on limited use of cloud service to do the best that I can and now this?</p>",
      "rawMarkdown": "What is going on?\n\nI was in a meeting all morning, in between talks , I was glancing at my labtop checking the completion of my last model. All done, CV looked good, I did my ensemble and submitted with my fingers crossed as this is my last submission. Saw the LB and thought it was a mistake but had to keep quite and go on with the meeting. I just came out and realized there was a bomb shell kernel? This is really crazy. I am one of those with very weak HW and am relying on limited use of cloud service to do the best that I can and now this?",
      "votes": null
    },
    {
      "id": "324629",
      "postDate": "05/07/2018 21:37:49",
      "content": "<p>My model's local 184,903,891-Fold CV has a mean AUROC of 1, should I publish?</p>",
      "rawMarkdown": "My model's local 184,903,891-Fold CV has a mean AUROC of 1, should I publish?",
      "votes": null
    },
    {
      "id": "324632",
      "postDate": "05/07/2018 21:38:46",
      "content": "<p>LOL, go for it!</p>",
      "rawMarkdown": "LOL, go for it!",
      "votes": null
    },
    {
      "id": "324634",
      "postDate": "05/07/2018 21:40:52",
      "content": "<p>@Joe Eddy. Thanks for sharing. Unfortunately I didn't get to play much with validation. Your numbers are compelling. I was trying to catch up with a 0.9800 LB looking at feature engineering (got a hair close but failed :( I'll try one more submission). </p>",
      "rawMarkdown": "Joe Eddy. Thanks for sharing. Unfortunately I didn't get to play much with validation. Your numbers are compelling. I was trying to catch up with a 0.9800 LB looking at feature engineering (got a hair close but failed :( I'll try one more submission).",
      "votes": null
    },
    {
      "id": "324636",
      "postDate": "05/07/2018 21:41:42",
      "content": "<p>Good luck.</p>",
      "rawMarkdown": "Good luck.",
      "votes": null
    },
    {
      "id": "324641",
      "postDate": "05/07/2018 21:47:26",
      "content": "<p>I am sure it is strictly Bayesian :)</p>",
      "rawMarkdown": "I am sure it is strictly Bayesian :)",
      "votes": null
    },
    {
      "id": "324642",
      "postDate": "05/07/2018 21:50:06",
      "content": "<p>Of course.  Frequentism would be inappropriate for such a respectable competition.</p>",
      "rawMarkdown": "Of course.  Frequentism would be inappropriate for such a respectable competition.",
      "votes": null
    },
    {
      "id": "324644",
      "postDate": "05/07/2018 21:52:16",
      "content": "<p>No, none whatsoever. :-)</p>",
      "rawMarkdown": "No, none whatsoever. :-)",
      "votes": null
    },
    {
      "id": "324649",
      "postDate": "05/07/2018 21:57:18",
      "content": "<p><a href=\"/bestfitting\">@bestfitting</a> did the last minute jump...but still 0.9834</p>\n\n<p>May be 0.9835 will be his next and he will share the submission file on the forum ^^</p>",
      "rawMarkdown": "bestfitting did the last minute jump...but still 0.9834\n\nMay be 0.9835 will be his next and he will share the submission file on the forum ^^",
      "votes": null
    },
    {
      "id": "324659",
      "postDate": "05/07/2018 22:01:21",
      "content": "<p>At least we know we're gonna get a good write-up post competition! </p>",
      "rawMarkdown": "At least we know we're gonna get a good write-up post competition!",
      "votes": null
    },
    {
      "id": "324665",
      "postDate": "05/07/2018 22:03:37",
      "content": "<p>Rumor has it that <a href=\"/bestfitting\">@bestfitting</a> is actually Mark Zuckerberg. </p>",
      "rawMarkdown": "Rumor has it that @bestfitting is actually Mark Zuckerberg.",
      "votes": null
    },
    {
      "id": "324669",
      "postDate": "05/07/2018 22:04:52",
      "content": "<p>Zuckerberg wouldn't do the post-competition write-up</p>",
      "rawMarkdown": "Zuckerberg wouldn't do the post-competition write-up",
      "votes": null
    },
    {
      "id": "324670",
      "postDate": "05/07/2018 22:05:43",
      "content": "<p>That's what he would want you to believe ....</p>",
      "rawMarkdown": "That's what he would want you to believe ....",
      "votes": null
    },
    {
      "id": "324673",
      "postDate": "05/07/2018 22:07:09",
      "content": "<p>Tell us how you build you intuition? </p>",
      "rawMarkdown": "Tell us how you build you intuition?",
      "votes": null
    },
    {
      "id": "324682",
      "postDate": "05/07/2018 22:17:23",
      "content": "<blockquote>\n  <p><strong>eric wrote</strong></p>\n  \n  <blockquote>\n    <p>Tell us how you build you intuition? </p>\n  </blockquote>\n</blockquote>\n\n<p>I make lots, and lots, and lots of mistakes. </p>",
      "rawMarkdown": "&gt; **eric wrote**\n&gt; \n&gt; &gt; Tell us how you build you intuition? \n\nI make lots, and lots, and lots of mistakes.",
      "votes": null
    },
    {
      "id": "324697",
      "postDate": "05/07/2018 22:29:26",
      "content": "<p>LOL</p>",
      "rawMarkdown": "LOL",
      "votes": null
    },
    {
      "id": "324699",
      "postDate": "05/07/2018 22:31:28",
      "content": "<p>please at me if you get one :)</p>",
      "rawMarkdown": "please at me if you get one :)",
      "votes": null
    },
    {
      "id": "324703",
      "postDate": "05/07/2018 22:36:55",
      "content": "<p>Bojan Tunguz, this is a good option! One of the best indeed. No one learns without trying! I fully agree</p>",
      "rawMarkdown": "Bojan Tunguz, this is a good option! One of the best indeed. No one learns without trying! I fully agree",
      "votes": null
    },
    {
      "id": "324716",
      "postDate": "05/07/2018 22:48:13",
      "content": "<p>I believe that the only way to achieve 0.9835 is through the use of <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689\">Generative Adversarial Denoising Autoencoder</a>. </p>",
      "rawMarkdown": "I believe that the only way to achieve 0.9835 is through the use of [Generative Adversarial Denoising Autoencoder][1]. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689",
      "votes": null
    },
    {
      "id": "324719",
      "postDate": "05/07/2018 22:49:23",
      "content": "<p>Reinforcement learning is the way to go. :-)</p>",
      "rawMarkdown": "Reinforcement learning is the way to go. :-)",
      "votes": null
    },
    {
      "id": "324727",
      "postDate": "05/07/2018 22:58:23",
      "content": "<p>I'd first try changing the RNG seed ;)</p>",
      "rawMarkdown": "I'd first try changing the RNG seed ;)",
      "votes": null
    },
    {
      "id": "324738",
      "postDate": "05/07/2018 23:07:12",
      "content": "<p>My lgb-xgb blend is finally done after 8 days and now I can't submit it because the file is too big to download on the platform I was hosting it on.</p>",
      "rawMarkdown": "My lgb-xgb blend is finally done after 8 days and now I can't submit it because the file is too big to download on the platform I was hosting it on.",
      "votes": null
    },
    {
      "id": "324741",
      "postDate": "05/07/2018 23:13:22",
      "content": "<p>Try compressing it first.</p>",
      "rawMarkdown": "Try compressing it first.",
      "votes": null
    },
    {
      "id": "324770",
      "postDate": "05/07/2018 23:41:21",
      "content": "<p>That worked, thanks! I'm working on submitting it right now... fingers crossed.</p>",
      "rawMarkdown": "That worked, thanks! I'm working on submitting it right now... fingers crossed.",
      "votes": null
    },
    {
      "id": "324787",
      "postDate": "05/08/2018 00:00:48",
      "content": "<p>AAAGH! IT ALMOST WORKED!  I was downloading it as a compressed file and right at the very end it would stop working... I'll see if I can get it working and where it stands on the private leaderboard afterwards.</p>\n\n<p>Definitely my fault for putting off ending training and uploading predictions until the last minute.</p>",
      "rawMarkdown": "AAAGH! IT ALMOST WORKED!  I was downloading it as a compressed file and right at the very end it would stop working... I'll see if I can get it working and where it stands on the private leaderboard afterwards.\n\nDefinitely my fault for putting off ending training and uploading predictions until the last minute.",
      "votes": null
    },
    {
      "id": "324822",
      "postDate": "05/08/2018 00:24:57",
      "content": "<p>ok, my top score estimate was real bad!</p>",
      "rawMarkdown": "ok, my top score estimate was real bad!",
      "votes": null
    },
    {
      "id": "324830",
      "postDate": "05/08/2018 00:30:10",
      "content": "<p>@Joe It's all good. I was wrong too and the PB went above the LB. Congrats for your position and I hope you could share how you chose your features! </p>",
      "rawMarkdown": "Joe It's all good. I was wrong too and the PB went above the LB. Congrats for your position and I hope you could share how you chose your features!",
      "votes": null
    },
    {
      "id": "324841",
      "postDate": "05/08/2018 00:36:41",
      "content": "<p>Thanks Oscar! I could share feature ideas, but our strong team result comes primarily from productive ensembling. My best single model ended up being below bronze.</p>",
      "rawMarkdown": "Thanks Oscar! I could share feature ideas, but our strong team result comes primarily from productive ensembling. My best single model ended up being below bronze.",
      "votes": null
    },
    {
      "id": "324863",
      "postDate": "05/08/2018 00:50:49",
      "content": "<p>Interesting. How many models did you ensemble? </p>",
      "rawMarkdown": "Interesting. How many models did you ensemble?",
      "votes": null
    },
    {
      "id": "324870",
      "postDate": "05/08/2018 00:58:31",
      "content": "<p>It's hard to give an exact count, but our best scoring submissions looked something like a blend of the following:</p>\n\n<ul>\n<li>A blend of individually contributed best solutions (single lgb models, averaged lgb models, a stacked model with maybe ~7 base models)</li>\n<li>A logistic regression stack of all base models</li>\n<li>Random forest stack of all base models</li>\n<li>Xgboost stack of all base models</li>\n</ul>\n\n<p>Our single best was just an equal weighted average of the last 3, but the private scores of variants of this were all within .00001 of each other.</p>",
      "rawMarkdown": "It's hard to give an exact count, but our best scoring submissions looked something like a blend of the following:\n\n - A blend of individually contributed best solutions (single lgb models, averaged lgb models, a stacked model with maybe ~7 base models)\n - A logistic regression stack of all base models\n - Random forest stack of all base models\n - Xgboost stack of all base models\n\nOur single best was just an equal weighted average of the last 3, but the private scores of variants of this were all within .00001 of each other.",
      "votes": null
    },
    {
      "id": "324911",
      "postDate": "05/08/2018 01:32:08",
      "content": "<p>Fairly impressive amount of work. Good job!</p>",
      "rawMarkdown": "Fairly impressive amount of work. Good job!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 324486,
      "author_name": "pliptor",
      "author_url": "",
      "post_date": "05/07/2018 18:36:04",
      "content": "<p>What about 0.9836 kernels? Why be so inclusive to 0.9835 kernels?</p>",
      "votes": null,
      "replies": [
        {
          "id": 324507,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 18:54:26",
          "content": "<p>My intuition tells me that 0.9835 is the absolute theoretical limit for this dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324519,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/07/2018 19:09:35",
          "content": "<p>How often your intuition is right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324525,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "05/07/2018 19:13:20",
          "content": "<p>Normally Bojan would have the probability of 200%  to be right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324530,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/07/2018 19:19:55",
          "content": "<p>@Bojan In a more serious matter, what's your intuition for the private score? Will the wining score be below, above or at 0.9835? Strange nobody posted anything about shake ups yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324533,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "05/07/2018 19:21:36",
          "content": "<p>@Oscar Takeshita, I guess the winning private lb score should be above 0.9860+</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324536,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/07/2018 19:23:22",
          "content": "<p>@Snorlax My bet is below 0.9835. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324543,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "05/07/2018 19:28:21",
          "content": "<p>I bet .9855-.986</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324546,
          "author_name": "andrenaef",
          "author_url": "",
          "post_date": "05/07/2018 19:29:42",
          "content": "<p>We did a bit of correlation analysis. Correlation among submissions drops as time progresses. I therefore expect private scores to be lower, understanding that the public scores are based on the beginning of the test set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324550,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/07/2018 19:33:07",
          "content": "<p>@Snorlax @Joe Eddy That's interesting. Probably we shouldn't discuss why now but I'd like to hear the reasons after the competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324555,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "05/07/2018 19:36:36",
          "content": "<p>@Joe Eddy, happy to see you here, you are the one I want to thank to in this competition~ Would you mind sharing your estimation of your private lb score here? Here is mine, estimation of public lb = 0.9811, estimation of private lb = 0.9840, and the actual public lb score is 0.9812. I hope my validation schema works, then I would have a higher private lb compared with public lb score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324599,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "05/07/2018 21:00:30",
          "content": "<p>@Oscar I'm happy to share my thoughts on that now, especially since the ideas have already come up in multiple other threads. @Snorlax it sounds like my best single model has very similar looking validation to yours (see below), then I'd estimate + ~.0006 or so from ensembling.</p>\n\n<p>The basis for my top score estimate comes from my validation setup. For this problem, a natural validation path is to train your model on days 7-8 and predict on the test hours in day 9. Doing that, I've seen hour 4 validation consistently fall within .0004 of public LB score, and I expect private LB hours validation to show similar consistency with actual private LB scores. I also would bet that top scorers use a similar validation style. </p>\n\n<p>So here's my (very unscientific) reasoning for my guess. My best single model validates at roughly .984 / .981 (public LB .9812). The gap between my best public LB and top of leaderboard is .0022, so if I extrapolate that to the private hours I get .9862. Toning that down to account for some diminishing returns on the larger sample, I estimate .9855-.986 for the top scorers. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324604,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 21:07:18",
          "content": "<p>In all honesty, I don't have a good estimate or intuition for this one. Due to the huge imbalance in data, and out-of time test set, it has been hard to get any \"gut level\" feelings for this one. Blending has been tough, and only started working more or less consistently once our models started approaching 0.98. This to me suggests that minute deviations from the correct ordering could have really serious consequences. I fear that we might be in for a really big shakeup. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324618,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/07/2018 21:22:08",
          "content": "<p>Sadly we didn't manage to blend well Neural nets model even above 0.98xx</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324634,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/07/2018 21:40:52",
          "content": "<p>@Joe Eddy. Thanks for sharing. Unfortunately I didn't get to play much with validation. Your numbers are compelling. I was trying to catch up with a 0.9800 LB looking at feature engineering (got a hair close but failed :( I'll try one more submission). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324673,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/07/2018 22:07:09",
          "content": "<p>Tell us how you build you intuition? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324682,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 22:17:23",
          "content": "<blockquote>\n  <p><strong>eric wrote</strong></p>\n  \n  <blockquote>\n    <p>Tell us how you build you intuition? </p>\n  </blockquote>\n</blockquote>\n\n<p>I make lots, and lots, and lots of mistakes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324703,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/07/2018 22:36:55",
          "content": "<p>Bojan Tunguz, this is a good option! One of the best indeed. No one learns without trying! I fully agree</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324719,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 22:49:23",
          "content": "<p>Reinforcement learning is the way to go. :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324822,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "05/08/2018 00:24:57",
          "content": "<p>ok, my top score estimate was real bad!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324830,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/08/2018 00:30:10",
          "content": "<p>@Joe It's all good. I was wrong too and the PB went above the LB. Congrats for your position and I hope you could share how you chose your features! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324841,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "05/08/2018 00:36:41",
          "content": "<p>Thanks Oscar! I could share feature ideas, but our strong team result comes primarily from productive ensembling. My best single model ended up being below bronze.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324863,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/08/2018 00:50:49",
          "content": "<p>Interesting. How many models did you ensemble? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324870,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "05/08/2018 00:58:31",
          "content": "<p>It's hard to give an exact count, but our best scoring submissions looked something like a blend of the following:</p>\n\n<ul>\n<li>A blend of individually contributed best solutions (single lgb models, averaged lgb models, a stacked model with maybe ~7 base models)</li>\n<li>A logistic regression stack of all base models</li>\n<li>Random forest stack of all base models</li>\n<li>Xgboost stack of all base models</li>\n</ul>\n\n<p>Our single best was just an equal weighted average of the last 3, but the private scores of variants of this were all within .00001 of each other.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324911,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/08/2018 01:32:08",
          "content": "<p>Fairly impressive amount of work. Good job!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324492,
      "author_name": "rrqqmm",
      "author_url": "",
      "post_date": "05/07/2018 18:43:15",
      "content": "<p>Wait for it. Almost there.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324510,
      "author_name": "authman",
      "author_url": "",
      "post_date": "05/07/2018 18:58:51",
      "content": "<p><a href=\"/bestfitting\">@bestfitting</a> is going to pull something :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 324649,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/07/2018 21:57:18",
          "content": "<p><a href=\"/bestfitting\">@bestfitting</a> did the last minute jump...but still 0.9834</p>\n\n<p>May be 0.9835 will be his next and he will share the submission file on the forum ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324659,
          "author_name": "authman",
          "author_url": "",
          "post_date": "05/07/2018 22:01:21",
          "content": "<p>At least we know we're gonna get a good write-up post competition! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324665,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 22:03:37",
          "content": "<p>Rumor has it that <a href=\"/bestfitting\">@bestfitting</a> is actually Mark Zuckerberg. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324669,
          "author_name": "authman",
          "author_url": "",
          "post_date": "05/07/2018 22:04:52",
          "content": "<p>Zuckerberg wouldn't do the post-competition write-up</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324670,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 22:05:43",
          "content": "<p>That's what he would want you to believe ....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324512,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "05/07/2018 18:59:25",
      "content": "<p>I'm training the full dataset with 1M features. once I've done, I put there !</p>",
      "votes": null,
      "replies": [
        {
          "id": 324554,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "05/07/2018 19:36:35",
          "content": "<p>almost finished ... don't waste your last submission lol</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324636,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/07/2018 21:41:42",
          "content": "<p>Good luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324516,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/07/2018 19:04:43",
      "content": "<p>The kernel would probably be useless ....maybe just sharing 0.9835 submission file ^^</p>",
      "votes": null,
      "replies": [
        {
          "id": 324558,
          "author_name": "konchar",
          "author_url": "",
          "post_date": "05/07/2018 19:40:03",
          "content": "<p>Hahhahha!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324531,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "05/07/2018 19:20:23",
      "content": "<p>I have a notebook that has a local AUC of 0.9978799432371899 ... but it's useless :x</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324557,
      "author_name": "shanth84",
      "author_url": "",
      "post_date": "05/07/2018 19:38:30",
      "content": "<p>@ Bojan \nLOL..Good one</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324563,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/07/2018 19:48:09",
      "content": "<p>Working on it, just a sec...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324603,
      "author_name": "dwchen",
      "author_url": "",
      "post_date": "05/07/2018 21:04:43",
      "content": "<p>Is there any reason or any result that shows 0.9835 is the limit for this dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 324644,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 21:52:16",
          "content": "<p>No, none whatsoever. :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324620,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/07/2018 21:24:49",
      "content": "<p>What is going on?</p>\n\n<p>I was in a meeting all morning, in between talks , I was glancing at my labtop checking the completion of my last model. All done, CV looked good, I did my ensemble and submitted with my fingers crossed as this is my last submission. Saw the LB and thought it was a mistake but had to keep quite and go on with the meeting. I just came out and realized there was a bomb shell kernel? This is really crazy. I am one of those with very weak HW and am relying on limited use of cloud service to do the best that I can and now this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324629,
      "author_name": "matthewa313",
      "author_url": "",
      "post_date": "05/07/2018 21:37:49",
      "content": "<p>My model's local 184,903,891-Fold CV has a mean AUROC of 1, should I publish?</p>",
      "votes": null,
      "replies": [
        {
          "id": 324632,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 21:38:46",
          "content": "<p>LOL, go for it!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324641,
          "author_name": "konchar",
          "author_url": "",
          "post_date": "05/07/2018 21:47:26",
          "content": "<p>I am sure it is strictly Bayesian :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324642,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/07/2018 21:50:06",
          "content": "<p>Of course.  Frequentism would be inappropriate for such a respectable competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324697,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "05/07/2018 22:29:26",
      "content": "<p>LOL</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324699,
      "author_name": "creamiracle",
      "author_url": "",
      "post_date": "05/07/2018 22:31:28",
      "content": "<p>please at me if you get one :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 324716,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "05/07/2018 22:48:13",
      "content": "<p>I believe that the only way to achieve 0.9835 is through the use of <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689\">Generative Adversarial Denoising Autoencoder</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 324727,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/07/2018 22:58:23",
          "content": "<p>I'd first try changing the RNG seed ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324738,
      "author_name": "matthewa313",
      "author_url": "",
      "post_date": "05/07/2018 23:07:12",
      "content": "<p>My lgb-xgb blend is finally done after 8 days and now I can't submit it because the file is too big to download on the platform I was hosting it on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 324741,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/07/2018 23:13:22",
          "content": "<p>Try compressing it first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324770,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/07/2018 23:41:21",
          "content": "<p>That worked, thanks! I'm working on submitting it right now... fingers crossed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324787,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/08/2018 00:00:48",
          "content": "<p>AAAGH! IT ALMOST WORKED!  I was downloading it as a compressed file and right at the very end it would stop working... I'll see if I can get it working and where it stands on the private leaderboard afterwards.</p>\n\n<p>Definitely my fault for putting off ending training and uploading predictions until the last minute.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "324484": "We are preparing our final submission, and would hate to be excluded from potentially sharing the first place in case someone submits such a gem at the last moment. Thanks.",
    "324486": "What about 0.9836 kernels? Why be so inclusive to 0.9835 kernels?",
    "324492": "Wait for it. Almost there.",
    "324507": "My intuition tells me that 0.9835 is the absolute theoretical limit for this dataset.",
    "324510": "bestfitting is going to pull something :-)",
    "324512": "I'm training the full dataset with 1M features. once I've done, I put there !",
    "324516": "The kernel would probably be useless ....maybe just sharing 0.9835 submission file ^^",
    "324519": "How often your intuition is right?",
    "324525": "Normally Bojan would have the probability of 200%  to be right.",
    "324530": "Bojan In a more serious matter, what's your intuition for the private score? Will the wining score be below, above or at 0.9835? Strange nobody posted anything about shake ups yet.",
    "324531": "I have a notebook that has a local AUC of 0.9978799432371899 ... but it's useless :x",
    "324533": "Oscar Takeshita, I guess the winning private lb score should be above 0.9860+",
    "324536": "Snorlax My bet is below 0.9835.",
    "324543": "I bet .9855-.986",
    "324546": "We did a bit of correlation analysis. Correlation among submissions drops as time progresses. I therefore expect private scores to be lower, understanding that the public scores are based on the beginning of the test set.",
    "324550": "Snorlax @Joe Eddy That's interesting. Probably we shouldn't discuss why now but I'd like to hear the reasons after the competition.",
    "324554": "almost finished ... don't waste your last submission lol",
    "324555": "Joe Eddy, happy to see you here, you are the one I want to thank to in this competition~ Would you mind sharing your estimation of your private lb score here? Here is mine, estimation of public lb = 0.9811, estimation of private lb = 0.9840, and the actual public lb score is 0.9812. I hope my validation schema works, then I would have a higher private lb compared with public lb score.",
    "324557": "Bojan \nLOL..Good one",
    "324558": "Hahhahha!",
    "324563": "Working on it, just a sec...",
    "324599": "Oscar I'm happy to share my thoughts on that now, especially since the ideas have already come up in multiple other threads. @Snorlax it sounds like my best single model has very similar looking validation to yours (see below), then I'd estimate + ~.0006 or so from ensembling.\n\nThe basis for my top score estimate comes from my validation setup. For this problem, a natural validation path is to train your model on days 7-8 and predict on the test hours in day 9. Doing that, I've seen hour 4 validation consistently fall within .0004 of public LB score, and I expect private LB hours validation to show similar consistency with actual private LB scores. I also would bet that top scorers use a similar validation style. \n\nSo here's my (very unscientific) reasoning for my guess. My best single model validates at roughly .984 / .981 (public LB .9812). The gap between my best public LB and top of leaderboard is .0022, so if I extrapolate that to the private hours I get .9862. Toning that down to account for some diminishing returns on the larger sample, I estimate .9855-.986 for the top scorers.",
    "324603": "Is there any reason or any result that shows 0.9835 is the limit for this dataset?",
    "324604": "In all honesty, I don't have a good estimate or intuition for this one. Due to the huge imbalance in data, and out-of time test set, it has been hard to get any \"gut level\" feelings for this one. Blending has been tough, and only started working more or less consistently once our models started approaching 0.98. This to me suggests that minute deviations from the correct ordering could have really serious consequences. I fear that we might be in for a really big shakeup.",
    "324618": "Sadly we didn't manage to blend well Neural nets model even above 0.98xx",
    "324620": "What is going on?\n\nI was in a meeting all morning, in between talks , I was glancing at my labtop checking the completion of my last model. All done, CV looked good, I did my ensemble and submitted with my fingers crossed as this is my last submission. Saw the LB and thought it was a mistake but had to keep quite and go on with the meeting. I just came out and realized there was a bomb shell kernel? This is really crazy. I am one of those with very weak HW and am relying on limited use of cloud service to do the best that I can and now this?",
    "324629": "My model's local 184,903,891-Fold CV has a mean AUROC of 1, should I publish?",
    "324632": "LOL, go for it!",
    "324634": "Joe Eddy. Thanks for sharing. Unfortunately I didn't get to play much with validation. Your numbers are compelling. I was trying to catch up with a 0.9800 LB looking at feature engineering (got a hair close but failed :( I'll try one more submission).",
    "324636": "Good luck.",
    "324641": "I am sure it is strictly Bayesian :)",
    "324642": "Of course.  Frequentism would be inappropriate for such a respectable competition.",
    "324644": "No, none whatsoever. :-)",
    "324649": "bestfitting did the last minute jump...but still 0.9834\n\nMay be 0.9835 will be his next and he will share the submission file on the forum ^^",
    "324659": "At least we know we're gonna get a good write-up post competition!",
    "324665": "Rumor has it that @bestfitting is actually Mark Zuckerberg.",
    "324669": "Zuckerberg wouldn't do the post-competition write-up",
    "324670": "That's what he would want you to believe ....",
    "324673": "Tell us how you build you intuition?",
    "324682": "&gt; **eric wrote**\n&gt; \n&gt; &gt; Tell us how you build you intuition? \n\nI make lots, and lots, and lots of mistakes.",
    "324697": "LOL",
    "324699": "please at me if you get one :)",
    "324703": "Bojan Tunguz, this is a good option! One of the best indeed. No one learns without trying! I fully agree",
    "324716": "I believe that the only way to achieve 0.9835 is through the use of [Generative Adversarial Denoising Autoencoder][1]. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689",
    "324719": "Reinforcement learning is the way to go. :-)",
    "324727": "I'd first try changing the RNG seed ;)",
    "324738": "My lgb-xgb blend is finally done after 8 days and now I can't submit it because the file is too big to download on the platform I was hosting it on.",
    "324741": "Try compressing it first.",
    "324770": "That worked, thanks! I'm working on submitting it right now... fingers crossed.",
    "324787": "AAAGH! IT ALMOST WORKED!  I was downloading it as a compressed file and right at the very end it would stop working... I'll see if I can get it working and where it stands on the private leaderboard afterwards.\n\nDefinitely my fault for putting off ending training and uploading predictions until the last minute.",
    "324822": "ok, my top score estimate was real bad!",
    "324830": "Joe It's all good. I was wrong too and the PB went above the LB. Congrats for your position and I hope you could share how you chose your features!",
    "324841": "Thanks Oscar! I could share feature ideas, but our strong team result comes primarily from productive ensembling. My best single model ended up being below bronze.",
    "324863": "Interesting. How many models did you ensemble?",
    "324870": "It's hard to give an exact count, but our best scoring submissions looked something like a blend of the following:\n\n - A blend of individually contributed best solutions (single lgb models, averaged lgb models, a stacked model with maybe ~7 base models)\n - A logistic regression stack of all base models\n - Random forest stack of all base models\n - Xgboost stack of all base models\n\nOur single best was just an equal weighted average of the last 3, but the private scores of variants of this were all within .00001 of each other.",
    "324911": "Fairly impressive amount of work. Good job!"
  },
  "source": "meta"
}