{
  "id": 347660,
  "title": "Share your shake up/down reasons",
  "url": "/competitions/amex-default-prediction/discussion/347660",
  "author_name": "",
  "post_date": "2022-08-25T02:55:12.471071600Z",
  "votes": 35,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Congratulations to all winners survived in such a shake competition.I worked hard for this competition but unfortunately I dropped from 15th to 90th,I want to share some possible reasons why I drop, hope you guys also can share yours.</p>\n<ol>\n<li>I focused on feature engineering/feature selection to improve my lightgbm cv ,best single model cv is 0.8005, private LB is 0.80763.I think the distribution difference between train and private-test and metric noise are huge that trust cv didn't work in this competition.</li>\n<li>I use ridge regressor for stacking that cause overfitting to train cv, maybe blend according to public lb is better choise.</li>\n<li>I created 8 neural network  models ,their ensemble's cv is 0.797 public LB 0.797,private LB 0.8049, after ensemble to gbdt models, the public score become better but private score worse,maybe nn models can not help ensemble.</li>\n<li>I only use raddar's clean data for feature engineering,I don't know if raw data is better or worse,please share your exeperiments results.(I found top rankers and high score kernel notebook use floor preprocess to achieve better private score)</li>\n<li>I merge train and test data then create very many features,not only groupby customer but also groupby category|month then calculate difference ,maybe merge train and test cause overfitting.</li>\n<li>After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.</li>\n</ol>",
  "messages": [
    {
      "id": "1912893",
      "postDate": "08/25/2022 02:55:12",
      "content": "<p>Congratulations to all winners survived in such a shake competition.I worked hard for this competition but unfortunately I dropped from 15th to 90th,I want to share some possible reasons why I drop, hope you guys also can share yours.</p>\n<ol>\n<li>I focused on feature engineering/feature selection to improve my lightgbm cv ,best single model cv is 0.8005, private LB is 0.80763.I think the distribution difference between train and private-test and metric noise are huge that trust cv didn't work in this competition.</li>\n<li>I use ridge regressor for stacking that cause overfitting to train cv, maybe blend according to public lb is better choise.</li>\n<li>I created 8 neural network  models ,their ensemble's cv is 0.797 public LB 0.797,private LB 0.8049, after ensemble to gbdt models, the public score become better but private score worse,maybe nn models can not help ensemble.</li>\n<li>I only use raddar's clean data for feature engineering,I don't know if raw data is better or worse,please share your exeperiments results.(I found top rankers and high score kernel notebook use floor preprocess to achieve better private score)</li>\n<li>I merge train and test data then create very many features,not only groupby customer but also groupby category|month then calculate difference ,maybe merge train and test cause overfitting.</li>\n<li>After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.</li>\n</ol>",
      "rawMarkdown": "Congratulations to all winners survived in such a shake competition.I worked hard for this competition but unfortunately I dropped from 15th to 90th,I want to share some possible reasons why I drop, hope you guys also can share yours.\n1. I focused on feature engineering/feature selection to improve my lightgbm cv ,best single model cv is 0.8005, private LB is 0.80763.I think the distribution difference between train and private-test and metric noise are huge that trust cv didn't work in this competition.\n2. I use ridge regressor for stacking that cause overfitting to train cv, maybe blend according to public lb is better choise.\n3. I created 8 neural network  models ,their ensemble's cv is 0.797 public LB 0.797,private LB 0.8049, after ensemble to gbdt models, the public score become better but private score worse,maybe nn models can not help ensemble.\n4. I only use raddar's clean data for feature engineering,I don't know if raw data is better or worse,please share your exeperiments results.(I found top rankers and high score kernel notebook use floor preprocess to achieve better private score)\n5. I merge train and test data then create very many features,not only groupby customer but also groupby category|month then calculate difference ,maybe merge train and test cause overfitting.\n6. After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.",
      "votes": null
    },
    {
      "id": "1912905",
      "postDate": "08/25/2022 03:11:34",
      "content": "<p>I feel skill/luck ratio for this competition is less than 0.5 ☹️ </p>",
      "rawMarkdown": "I feel skill/luck ratio for this competition is less than 0.5 ☹️",
      "votes": null
    },
    {
      "id": "1912909",
      "postDate": "08/25/2022 03:26:26",
      "content": "<p>\"best single model cv is 0.8005, private LB is 0.8007\" I think you means \"private LB is 0.8070\"?<br>\n I was too busy to spend much time in this competition. But from what I observed, one of the problem using stacking could be: when doing cv, prediction scales are quite different, which will have a huge effect on the performance. <br>\nI using lightgbm to  do the stacking. For one fold, the model is early stopped in round 24, and the max prediction is lower than 0.6(because it's not totally trained). For other folds, the problem exists still. So if you do nothing about the scaling problem, even the average amex_metric is great, the acutal result will be very bad. So after the prediction for each fold, I rescale the prediction to the range 0.0~1.0, which improve the result, but still it's not the perfect way, I guess.<br>\nAnyway, that's probably not the reason in your case to shake down. Maybe it's just bad luck.</p>",
      "rawMarkdown": "\"best single model cv is 0.8005, private LB is 0.8007\" I think you means \"private LB is 0.8070\"?\n I was too busy to spend much time in this competition. But from what I observed, one of the problem using stacking could be: when doing cv, prediction scales are quite different, which will have a huge effect on the performance. \nI using lightgbm to  do the stacking. For one fold, the model is early stopped in round 24, and the max prediction is lower than 0.6(because it's not totally trained). For other folds, the problem exists still. So if you do nothing about the scaling problem, even the average amex_metric is great, the acutal result will be very bad. So after the prediction for each fold, I rescale the prediction to the range 0.0~1.0, which improve the result, but still it's not the perfect way, I guess.\nAnyway, that's probably not the reason in your case to shake down. Maybe it's just bad luck.",
      "votes": null
    },
    {
      "id": "1912912",
      "postDate": "08/25/2022 03:28:53",
      "content": "<p>I hope so,but many GMs are at top rank,they should did some solid work</p>",
      "rawMarkdown": "I hope so,but many GMs are at top rank,they should did some solid work",
      "votes": null
    },
    {
      "id": "1912914",
      "postDate": "08/25/2022 03:30:15",
      "content": "<p>my best single model lgbm cv0.79932 public lb0.79980 private lb0.80731, found with tabnet ensemble seems to improve a lot. Unfortunately we ended up missing the best private LB</p>",
      "rawMarkdown": "my best single model lgbm cv0.79932 public lb0.79980 private lb0.80731, found with tabnet ensemble seems to improve a lot. Unfortunately we ended up missing the best private LB",
      "votes": null
    },
    {
      "id": "1912920",
      "postDate": "08/25/2022 03:36:34",
      "content": "<p>I also tried scale predictions values almost no changes.I prefer to think single blending is better.</p>",
      "rawMarkdown": "I also tried scale predictions values almost no changes.I prefer to think single blending is better.",
      "votes": null
    },
    {
      "id": "1912932",
      "postDate": "08/25/2022 03:54:13",
      "content": "<p>Thanks for sharing the solution. I worked on this competition for less than a week or so, I was busy with Feedback. I solely focused on bringing out diverse models without much focus on FE part as I was seeing just minor improvements. My currently ensemble includes variety of models 1dcnn, tabnet,NN, lgb,catboost and xgboost. Thankfully I got shakeup, I regret for not joining earlier now :)</p>",
      "rawMarkdown": "Thanks for sharing the solution. I worked on this competition for less than a week or so, I was busy with Feedback. I solely focused on bringing out diverse models without much focus on FE part as I was seeing just minor improvements. My currently ensemble includes variety of models 1dcnn, tabnet,NN, lgb,catboost and xgboost. Thankfully I got shakeup, I regret for not joining earlier now :)",
      "votes": null
    },
    {
      "id": "1913088",
      "postDate": "08/25/2022 06:14:36",
      "content": "<p>I notice that for my score, +/-0.0001, even the fourth digit, is +/- 15 slots or so.</p>\n<p>Besides pure luck, things that might've helped me jump higher:</p>\n<ul>\n<li>\"miss next payment\" prediction on private test data. This may have helped avoid some of the issues with using a model trained on a different time period?</li>\n<li>I dropped B_29.</li>\n</ul>\n<p>Re: #4 I used raddar's clean data as well.</p>",
      "rawMarkdown": "I notice that for my score, +/-0.0001, even the fourth digit, is +/- 15 slots or so.\n\nBesides pure luck, things that might've helped me jump higher:\n* \"miss next payment\" prediction on private test data. This may have helped avoid some of the issues with using a model trained on a different time period?\n* I dropped B_29.\n\nRe: #4 I used raddar's clean data as well.",
      "votes": null
    },
    {
      "id": "1913097",
      "postDate": "08/25/2022 06:21:41",
      "content": "<p>P.s. My best single model was, well, 49th overall. </p>\n<p>0.80798 private<br>\n0.79889 public</p>\n<p>Oh, and I trained on full dataset (4 times over) with no folds or early stopping. I can't think of any reason that would really help on private leaderboard more than public, though.</p>",
      "rawMarkdown": "P.s. My best single model was, well, 49th overall. \n\n0.80798 private\n0.79889 public\n\nOh, and I trained on full dataset (4 times over) with no folds or early stopping. I can't think of any reason that would really help on private leaderboard more than public, though.",
      "votes": null
    },
    {
      "id": "1913114",
      "postDate": "08/25/2022 06:38:32",
      "content": "<p>Once again, bad decision on submissions chosen for me 😓</p>\n<p>Went for best CV and best LB, which are both just under bronze on private.. ensemble of both was Silver.. Unselected it in the last few hours lol</p>",
      "rawMarkdown": "Once again, bad decision on submissions chosen for me 😓\n\nWent for best CV and best LB, which are both just under bronze on private.. ensemble of both was Silver.. Unselected it in the last few hours lol",
      "votes": null
    },
    {
      "id": "1913870",
      "postDate": "08/25/2022 14:53:08",
      "content": "<p>Our best submission CV score was 0.80002, Public score was 0.79956, Private score was 0.80739, we shook up over 1400th, not sure why.<br>\n(Our final submission was Stacking with 18 models, using ridge as Mr.@senkin13 did).</p>\n<p>Our Stacking used 11 models GBDT (out of 18 overall)<br>\nI guess that model was Fit for Private.<br>\n(Of those models, 5 are NN-based and 2 are Logistic regression.)</p>",
      "rawMarkdown": "Our best submission CV score was 0.80002, Public score was 0.79956, Private score was 0.80739, we shook up over 1400th, not sure why.\n(Our final submission was Stacking with 18 models, using ridge as Mr.@senkin13 did).\n\nOur Stacking used 11 models GBDT (out of 18 overall)\nI guess that model was Fit for Private.\n(Of those models, 5 are NN-based and 2 are Logistic regression.)",
      "votes": null
    },
    {
      "id": "1913881",
      "postDate": "08/25/2022 14:57:12",
      "content": "<p>I ran an experiment scaling/fix bias of features in train, public and private separatetly. It didn't changed train scores much, but it leaded to worse results in public and private.</p>",
      "rawMarkdown": "I ran an experiment scaling/fix bias of features in train, public and private separatetly. It didn't changed train scores much, but it leaded to worse results in public and private.",
      "votes": null
    },
    {
      "id": "1913884",
      "postDate": "08/25/2022 15:00:47",
      "content": "<p>Guess CV is not a good measurement for evaluating the performance in PriLB this time. Our CV improvement always align with publicLB but fall hard in private. We've tried several methods to estimate our performance in PrivateLB but no luck with it, you can read our solution post (maybe tomorrow or several hours later) to have more information about our crazy approaches 😃</p>",
      "rawMarkdown": "Guess CV is not a good measurement for evaluating the performance in PriLB this time. Our CV improvement always align with publicLB but fall hard in private. We've tried several methods to estimate our performance in PrivateLB but no luck with it, you can read our solution post (maybe tomorrow or several hours later) to have more information about our crazy approaches 😃",
      "votes": null
    },
    {
      "id": "1914012",
      "postDate": "08/25/2022 16:43:24",
      "content": "<p>Echoing <a href=\"https://www.kaggle.com/andrew60909\" target=\"_blank\">@andrew60909</a> -- my honest feeling is that a robust approach is necessary but not sufficient for reaching gold in this competition. Exploiting the test data well seems like a winning theme we could have used more, but I still think there is ultimately a significant amount of noise. Public LB and CV actually aligned really well for us but seemed meaningfully decoupled from Private LB. Maybe Public is too similar to training data and less similar to Private, but I'm skeptical -- I think that data drift in this problem domain is less of an issue than the raw volatility of the metric. We've been joking about how it's like high on the LB we're just fighting over flipping one customer on the right or wrong side of the threshold.</p>\n<p>Some of our solution stats--<br>\nCV .8006, Public LB 92nd, Private Gold<br>\nCV .8024, Public LB 6th, Higher Private Gold <br>\nCV .8031, Public LB 5th, Private .00009 off Gold</p>\n<p>5th decimal struggles!</p>",
      "rawMarkdown": "Echoing @andrew60909 -- my honest feeling is that a robust approach is necessary but not sufficient for reaching gold in this competition. Exploiting the test data well seems like a winning theme we could have used more, but I still think there is ultimately a significant amount of noise. Public LB and CV actually aligned really well for us but seemed meaningfully decoupled from Private LB. Maybe Public is too similar to training data and less similar to Private, but I'm skeptical -- I think that data drift in this problem domain is less of an issue than the raw volatility of the metric. We've been joking about how it's like high on the LB we're just fighting over flipping one customer on the right or wrong side of the threshold.\n\nSome of our solution stats--\nCV .8006, Public LB 92nd, Private Gold\nCV .8024, Public LB 6th, Higher Private Gold \nCV .8031, Public LB 5th, Private .00009 off Gold\n\n5th decimal struggles!",
      "votes": null
    },
    {
      "id": "1914139",
      "postDate": "08/25/2022 19:02:50",
      "content": "<p>For us features diversity was the key factor, then ensembling weights( calculated luck-based) and heavily correlated CV,LB. so we chose our public best and cv best.  </p>",
      "rawMarkdown": "For us features diversity was the key factor, then ensembling weights( calculated luck-based) and heavily correlated CV,LB. so we chose our public best and cv best.",
      "votes": null
    },
    {
      "id": "1914165",
      "postDate": "08/25/2022 19:41:59",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> Yes of course. It was a twitter-referenced half-joke. From the published solutions it seems leveraging test data in an innovative way provides a certain edge, so there's certainly a skill factor here. <br>\nBut in all seriousness, in terms of final ranking, luck can easily have an impact of maybe 50 place difference or more. Some shakeup teams believe their edge is diversity but I don't think that's really the case.  </p>",
      "rawMarkdown": "senkin13 Yes of course. It was a twitter-referenced half-joke. From the published solutions it seems leveraging test data in an innovative way provides a certain edge, so there's certainly a skill factor here. \nBut in all seriousness, in terms of final ranking, luck can easily have an impact of maybe 50 place difference or more. Some shakeup teams believe their edge is diversity but I don't think that's really the case.",
      "votes": null
    },
    {
      "id": "1914177",
      "postDate": "08/25/2022 19:54:33",
      "content": "<p>And it's probably a provable hypothesis if enough teams bothered to rerun their exact models with different seeds. See the variance first hand. </p>\n<p>That said, a couple thoughts:</p>\n<ul>\n<li>Seasonality? Public and train shared seasonality, but private was way different. Maximal difference, even (6mo offset). This may have been the number 1 issue, and even just unsupervised pretraining on private data (@cdeotte) or any other way of leveraging the raw data may be a common thread of those who avoided shakedown?</li>\n<li>Besides \"it's just luck\" there's the issue of, well, not truly over fitting CV or leaderboard, but simply having higher likelihood of inflated reported results. Which increases chance (relative to others) of shakedown. </li>\n</ul>\n<p>Let me give an example:<br>\nOne team tries 5 things that are all different, but unknowingly they are unimportant deltas. So it is like trying 5 different seeds. They haven't necessarily made any methodological mistake, and it doesn't hurt their final chances any, they haven't overfit in the sense of worse predictive power. But their equal predictive power on private gets inflated score on CV and public. </p>\n<p>A second team does the same, but 200 times, not 5 times. Ground truth they are equal, but CV and LB scores are inflated in favor the 200 experiment team. The 5 experiment team has higher chance of increased rank, even though - with fewer experiments - they might have the worse model. </p>\n<p>Point is there's statistical reason to expect a bit more inflated public scores (and CV scores) on the people doing the most optimization work EVEN if that optimization is the right approach. So they have odds working (slightly) against them when it comes shakedown time.</p>",
      "rawMarkdown": "And it's probably a provable hypothesis if enough teams bothered to rerun their exact models with different seeds. See the variance first hand. \n\nThat said, a couple thoughts:\n* Seasonality? Public and train shared seasonality, but private was way different. Maximal difference, even (6mo offset). This may have been the number 1 issue, and even just unsupervised pretraining on private data (@cdeotte) or any other way of leveraging the raw data may be a common thread of those who avoided shakedown?\n* Besides \"it's just luck\" there's the issue of, well, not truly over fitting CV or leaderboard, but simply having higher likelihood of inflated reported results. Which increases chance (relative to others) of shakedown. \n\nLet me give an example:\nOne team tries 5 things that are all different, but unknowingly they are unimportant deltas. So it is like trying 5 different seeds. They haven't necessarily made any methodological mistake, and it doesn't hurt their final chances any, they haven't overfit in the sense of worse predictive power. But their equal predictive power on private gets inflated score on CV and public. \n\nA second team does the same, but 200 times, not 5 times. Ground truth they are equal, but CV and LB scores are inflated in favor the 200 experiment team. The 5 experiment team has higher chance of increased rank, even though - with fewer experiments - they might have the worse model. \n\nPoint is there's statistical reason to expect a bit more inflated public scores (and CV scores) on the people doing the most optimization work EVEN if that optimization is the right approach. So they have odds working (slightly) against them when it comes shakedown time.",
      "votes": null
    },
    {
      "id": "1914180",
      "postDate": "08/25/2022 19:58:02",
      "content": "<p>I think the biggest numerical rank shakeups might be if you crossed a threshold compared with the best published public models? Just my guess. Ahead of the public on one LB, but behind it in the other. </p>",
      "rawMarkdown": "I think the biggest numerical rank shakeups might be if you crossed a threshold compared with the best published public models? Just my guess. Ahead of the public on one LB, but behind it in the other.",
      "votes": null
    },
    {
      "id": "1914217",
      "postDate": "08/25/2022 20:59:21",
      "content": "<p>The seasonality point is interesting, but I'm also skeptical that seasonality is significant in this specific problem (or that it is more significant than the level of metric noise).</p>\n<p>On your second point, I partially but don't completely agree. There is a big difference between shot in the dark and rigorous experiments, and you should bring strong priors to what experiments you try and how you evaluate them. It's one thing if trying different seeds improves your scores, but another if your new experiment is on good feature ideas or the addition of diverse models to an ensemble. There's statistical reason to expect inflated Public LB scores based on sheer experiment count, but also statistical reason to expect improvement across the board (CV, public, and private) with a large quantity of experiments if they are done well. </p>\n<p>At least for our subs, the experience here is that there is a meaningful correlation between better CV/LB and private LB, but it is weak enough that there can be a lot of fluctuation. So, our final changes may have been suboptimal experiments, but this could also mean that even though our solution improved on expectation that luck didn't favor us in the end. A clear example -- you would prefer a model that does somewhat better on 3/5 CV folds but a tiny bit worse on 2/5, but can plausibly expect that this model may underperform an older one by chance on a specific dataset.</p>",
      "rawMarkdown": "The seasonality point is interesting, but I'm also skeptical that seasonality is significant in this specific problem (or that it is more significant than the level of metric noise).\n\nOn your second point, I partially but don't completely agree. There is a big difference between shot in the dark and rigorous experiments, and you should bring strong priors to what experiments you try and how you evaluate them. It's one thing if trying different seeds improves your scores, but another if your new experiment is on good feature ideas or the addition of diverse models to an ensemble. There's statistical reason to expect inflated Public LB scores based on sheer experiment count, but also statistical reason to expect improvement across the board (CV, public, and private) with a large quantity of experiments if they are done well. \n\nAt least for our subs, the experience here is that there is a meaningful correlation between better CV/LB and private LB, but it is weak enough that there can be a lot of fluctuation. So, our final changes may have been suboptimal experiments, but this could also mean that even though our solution improved on expectation that luck didn't favor us in the end. A clear example -- you would prefer a model that does somewhat better on 3/5 CV folds but a tiny bit worse on 2/5, but can plausibly expect that this model may underperform an older one by chance on a specific dataset.",
      "votes": null
    },
    {
      "id": "1914226",
      "postDate": "08/25/2022 21:13:42",
      "content": "<p>Thank you for adding more nuance to the discussion. I completely agree. And in fact the last thing I want to do is sound like I'm lecturing the experts. Only trying to highlight a general principle for lurkers and in fact you and many others would know better than me exactly how to mitigate any such issues. </p>\n<p>As the noisy data became obvious, and as you got more tuned models, did you and do you ever do any multi seed variance studies to look at score mean and std dev in determining which experiments were the best? There's a big cost/benefit tradeoff, so curious if that's one way you help reduce risk of inflated scores?</p>",
      "rawMarkdown": "Thank you for adding more nuance to the discussion. I completely agree. And in fact the last thing I want to do is sound like I'm lecturing the experts. Only trying to highlight a general principle for lurkers and in fact you and many others would know better than me exactly how to mitigate any such issues. \n\nAs the noisy data became obvious, and as you got more tuned models, did you and do you ever do any multi seed variance studies to look at score mean and std dev in determining which experiments were the best? There's a big cost/benefit tradeoff, so curious if that's one way you help reduce risk of inflated scores?",
      "votes": null
    },
    {
      "id": "1914419",
      "postDate": "08/26/2022 05:03:25",
      "content": "<p>After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.</p>",
      "rawMarkdown": "After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.",
      "votes": null
    },
    {
      "id": "1914420",
      "postDate": "08/26/2022 05:05:09",
      "content": "<p>I think your shakeup come from better features that fit private data</p>",
      "rawMarkdown": "I think your shakeup come from better features that fit private data",
      "votes": null
    },
    {
      "id": "1914425",
      "postDate": "08/26/2022 05:10:16",
      "content": "<p>yes, diverse team work reduce the luck factor,congraluations for your huge shakeup.</p>",
      "rawMarkdown": "yes, diverse team work reduce the luck factor,congraluations for your huge shakeup.",
      "votes": null
    },
    {
      "id": "1914427",
      "postDate": "08/26/2022 05:11:56",
      "content": "<p>thanks, your approch is solid to fight against uncertainty.</p>",
      "rawMarkdown": "thanks, your approch is solid to fight against uncertainty.",
      "votes": null
    },
    {
      "id": "1914429",
      "postDate": "08/26/2022 05:17:07",
      "content": "<p>Hi Senkin13. I'm sorry to see you dropped on private LB. Like you say, your drop was probably caused by your features and/or types of models which overfitted to train and public LB. </p>\n<p>For me, I have 26 submissions over public LB 0.8010. They are all my same approach but have different seeds, different features, different architectures, etc. The CV scores all range from 0.7990 thru 0.8000. And all 26 of these submissions range from private LB 0.8082 thru LB 0.8085. So seeds and features doesn't change my final ranking too much. I think my approach which trained with test data worked very well on private LB and helped me maintain my public LB rank on private LB.</p>\n<p>This is kind of lucky. We didn't really know what approaches that worked on CV and public LB would also work on private LB.</p>",
      "rawMarkdown": "Hi Senkin13. I'm sorry to see you dropped on private LB. Like you say, your drop was probably caused by your features and/or types of models which overfitted to train and public LB. \n\nFor me, I have 26 submissions over public LB 0.8010. They are all my same approach but have different seeds, different features, different architectures, etc. The CV scores all range from 0.7990 thru 0.8000. And all 26 of these submissions range from private LB 0.8082 thru LB 0.8085. So seeds and features doesn't change my final ranking too much. I think my approach which trained with test data worked very well on private LB and helped me maintain my public LB rank on private LB.\n\nThis is kind of lucky. We didn't really know what approaches that worked on CV and public LB would also work on private LB.",
      "votes": null
    },
    {
      "id": "1914468",
      "postDate": "08/26/2022 05:43:40",
      "content": "<p>Hi Chris, thanks for your encourage.I found some teams including you win gold with very strong NN model.That is definitely no t luck, learn from you a lot.</p>",
      "rawMarkdown": "Hi Chris, thanks for your encourage.I found some teams including you win gold with very strong NN model.That is definitely no t luck, learn from you a lot.",
      "votes": null
    },
    {
      "id": "1914841",
      "postDate": "08/26/2022 13:20:57",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> <br>\nThank you for reply.I agree. I am sure that my model did not reach the public many Public fit Notebooks.<br>\n(I was sad because I mixed 18 models and couldn't beat the public Notebook…)</p>\n<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> <br>\nThank you for reply<br>\nI've checked but the single model made with any of the features was not a special Private fit. I really don't know…</p>",
      "rawMarkdown": "roberthatch \nThank you for reply.I agree. I am sure that my model did not reach the public many Public fit Notebooks.\n(I was sad because I mixed 18 models and couldn't beat the public Notebook...)\n\n@senkin13 \nThank you for reply\nI've checked but the single model made with any of the features was not a special Private fit. I really don't know...",
      "votes": null
    },
    {
      "id": "1915648",
      "postDate": "08/27/2022 06:41:05",
      "content": "<p>…a little more empirical evidence in support of the suggestion by Raddar of <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337609#1858410\" target=\"_blank\">\"<em>…second sub - best CV model/2 + best LB model/2…</em>\"</a></p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "...a little more empirical evidence in support of the suggestion by Raddar of [\"*...second sub - best CV model/2 + best LB model/2...*\"](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337609#1858410)\n\nAll the best,\ncarl",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1912905,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "08/25/2022 03:11:34",
      "content": "<p>I feel skill/luck ratio for this competition is less than 0.5 ☹️ </p>",
      "votes": null,
      "replies": [
        {
          "id": 1912912,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "08/25/2022 03:28:53",
          "content": "<p>I hope so,but many GMs are at top rank,they should did some solid work</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914165,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "08/25/2022 19:41:59",
          "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> Yes of course. It was a twitter-referenced half-joke. From the published solutions it seems leveraging test data in an innovative way provides a certain edge, so there's certainly a skill factor here. <br>\nBut in all seriousness, in terms of final ranking, luck can easily have an impact of maybe 50 place difference or more. Some shakeup teams believe their edge is diversity but I don't think that's really the case.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912909,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "08/25/2022 03:26:26",
      "content": "<p>\"best single model cv is 0.8005, private LB is 0.8007\" I think you means \"private LB is 0.8070\"?<br>\n I was too busy to spend much time in this competition. But from what I observed, one of the problem using stacking could be: when doing cv, prediction scales are quite different, which will have a huge effect on the performance. <br>\nI using lightgbm to  do the stacking. For one fold, the model is early stopped in round 24, and the max prediction is lower than 0.6(because it's not totally trained). For other folds, the problem exists still. So if you do nothing about the scaling problem, even the average amex_metric is great, the acutal result will be very bad. So after the prediction for each fold, I rescale the prediction to the range 0.0~1.0, which improve the result, but still it's not the perfect way, I guess.<br>\nAnyway, that's probably not the reason in your case to shake down. Maybe it's just bad luck.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912920,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "08/25/2022 03:36:34",
          "content": "<p>I also tried scale predictions values almost no changes.I prefer to think single blending is better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912914,
      "author_name": "tonymarkchris",
      "author_url": "",
      "post_date": "08/25/2022 03:30:15",
      "content": "<p>my best single model lgbm cv0.79932 public lb0.79980 private lb0.80731, found with tabnet ensemble seems to improve a lot. Unfortunately we ended up missing the best private LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912932,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "08/25/2022 03:54:13",
      "content": "<p>Thanks for sharing the solution. I worked on this competition for less than a week or so, I was busy with Feedback. I solely focused on bringing out diverse models without much focus on FE part as I was seeing just minor improvements. My currently ensemble includes variety of models 1dcnn, tabnet,NN, lgb,catboost and xgboost. Thankfully I got shakeup, I regret for not joining earlier now :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913088,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/25/2022 06:14:36",
      "content": "<p>I notice that for my score, +/-0.0001, even the fourth digit, is +/- 15 slots or so.</p>\n<p>Besides pure luck, things that might've helped me jump higher:</p>\n<ul>\n<li>\"miss next payment\" prediction on private test data. This may have helped avoid some of the issues with using a model trained on a different time period?</li>\n<li>I dropped B_29.</li>\n</ul>\n<p>Re: #4 I used raddar's clean data as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913097,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/25/2022 06:21:41",
          "content": "<p>P.s. My best single model was, well, 49th overall. </p>\n<p>0.80798 private<br>\n0.79889 public</p>\n<p>Oh, and I trained on full dataset (4 times over) with no folds or early stopping. I can't think of any reason that would really help on private leaderboard more than public, though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914427,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "08/26/2022 05:11:56",
          "content": "<p>thanks, your approch is solid to fight against uncertainty.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913114,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "08/25/2022 06:38:32",
      "content": "<p>Once again, bad decision on submissions chosen for me 😓</p>\n<p>Went for best CV and best LB, which are both just under bronze on private.. ensemble of both was Silver.. Unselected it in the last few hours lol</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915648,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "08/27/2022 06:41:05",
          "content": "<p>…a little more empirical evidence in support of the suggestion by Raddar of <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337609#1858410\" target=\"_blank\">\"<em>…second sub - best CV model/2 + best LB model/2…</em>\"</a></p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913870,
      "author_name": "takeshikobayashi",
      "author_url": "",
      "post_date": "08/25/2022 14:53:08",
      "content": "<p>Our best submission CV score was 0.80002, Public score was 0.79956, Private score was 0.80739, we shook up over 1400th, not sure why.<br>\n(Our final submission was Stacking with 18 models, using ridge as Mr.@senkin13 did).</p>\n<p>Our Stacking used 11 models GBDT (out of 18 overall)<br>\nI guess that model was Fit for Private.<br>\n(Of those models, 5 are NN-based and 2 are Logistic regression.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1914180,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/25/2022 19:58:02",
          "content": "<p>I think the biggest numerical rank shakeups might be if you crossed a threshold compared with the best published public models? Just my guess. Ahead of the public on one LB, but behind it in the other. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914420,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "08/26/2022 05:05:09",
          "content": "<p>I think your shakeup come from better features that fit private data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914841,
          "author_name": "takeshikobayashi",
          "author_url": "",
          "post_date": "08/26/2022 13:20:57",
          "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> <br>\nThank you for reply.I agree. I am sure that my model did not reach the public many Public fit Notebooks.<br>\n(I was sad because I mixed 18 models and couldn't beat the public Notebook…)</p>\n<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> <br>\nThank you for reply<br>\nI've checked but the single model made with any of the features was not a special Private fit. I really don't know…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913881,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "08/25/2022 14:57:12",
      "content": "<p>I ran an experiment scaling/fix bias of features in train, public and private separatetly. It didn't changed train scores much, but it leaded to worse results in public and private.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913884,
      "author_name": "andrew60909",
      "author_url": "",
      "post_date": "08/25/2022 15:00:47",
      "content": "<p>Guess CV is not a good measurement for evaluating the performance in PriLB this time. Our CV improvement always align with publicLB but fall hard in private. We've tried several methods to estimate our performance in PrivateLB but no luck with it, you can read our solution post (maybe tomorrow or several hours later) to have more information about our crazy approaches 😃</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914012,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "08/25/2022 16:43:24",
      "content": "<p>Echoing <a href=\"https://www.kaggle.com/andrew60909\" target=\"_blank\">@andrew60909</a> -- my honest feeling is that a robust approach is necessary but not sufficient for reaching gold in this competition. Exploiting the test data well seems like a winning theme we could have used more, but I still think there is ultimately a significant amount of noise. Public LB and CV actually aligned really well for us but seemed meaningfully decoupled from Private LB. Maybe Public is too similar to training data and less similar to Private, but I'm skeptical -- I think that data drift in this problem domain is less of an issue than the raw volatility of the metric. We've been joking about how it's like high on the LB we're just fighting over flipping one customer on the right or wrong side of the threshold.</p>\n<p>Some of our solution stats--<br>\nCV .8006, Public LB 92nd, Private Gold<br>\nCV .8024, Public LB 6th, Higher Private Gold <br>\nCV .8031, Public LB 5th, Private .00009 off Gold</p>\n<p>5th decimal struggles!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1914177,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/25/2022 19:54:33",
          "content": "<p>And it's probably a provable hypothesis if enough teams bothered to rerun their exact models with different seeds. See the variance first hand. </p>\n<p>That said, a couple thoughts:</p>\n<ul>\n<li>Seasonality? Public and train shared seasonality, but private was way different. Maximal difference, even (6mo offset). This may have been the number 1 issue, and even just unsupervised pretraining on private data (@cdeotte) or any other way of leveraging the raw data may be a common thread of those who avoided shakedown?</li>\n<li>Besides \"it's just luck\" there's the issue of, well, not truly over fitting CV or leaderboard, but simply having higher likelihood of inflated reported results. Which increases chance (relative to others) of shakedown. </li>\n</ul>\n<p>Let me give an example:<br>\nOne team tries 5 things that are all different, but unknowingly they are unimportant deltas. So it is like trying 5 different seeds. They haven't necessarily made any methodological mistake, and it doesn't hurt their final chances any, they haven't overfit in the sense of worse predictive power. But their equal predictive power on private gets inflated score on CV and public. </p>\n<p>A second team does the same, but 200 times, not 5 times. Ground truth they are equal, but CV and LB scores are inflated in favor the 200 experiment team. The 5 experiment team has higher chance of increased rank, even though - with fewer experiments - they might have the worse model. </p>\n<p>Point is there's statistical reason to expect a bit more inflated public scores (and CV scores) on the people doing the most optimization work EVEN if that optimization is the right approach. So they have odds working (slightly) against them when it comes shakedown time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914217,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "08/25/2022 20:59:21",
          "content": "<p>The seasonality point is interesting, but I'm also skeptical that seasonality is significant in this specific problem (or that it is more significant than the level of metric noise).</p>\n<p>On your second point, I partially but don't completely agree. There is a big difference between shot in the dark and rigorous experiments, and you should bring strong priors to what experiments you try and how you evaluate them. It's one thing if trying different seeds improves your scores, but another if your new experiment is on good feature ideas or the addition of diverse models to an ensemble. There's statistical reason to expect inflated Public LB scores based on sheer experiment count, but also statistical reason to expect improvement across the board (CV, public, and private) with a large quantity of experiments if they are done well. </p>\n<p>At least for our subs, the experience here is that there is a meaningful correlation between better CV/LB and private LB, but it is weak enough that there can be a lot of fluctuation. So, our final changes may have been suboptimal experiments, but this could also mean that even though our solution improved on expectation that luck didn't favor us in the end. A clear example -- you would prefer a model that does somewhat better on 3/5 CV folds but a tiny bit worse on 2/5, but can plausibly expect that this model may underperform an older one by chance on a specific dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914226,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/25/2022 21:13:42",
          "content": "<p>Thank you for adding more nuance to the discussion. I completely agree. And in fact the last thing I want to do is sound like I'm lecturing the experts. Only trying to highlight a general principle for lurkers and in fact you and many others would know better than me exactly how to mitigate any such issues. </p>\n<p>As the noisy data became obvious, and as you got more tuned models, did you and do you ever do any multi seed variance studies to look at score mean and std dev in determining which experiments were the best? There's a big cost/benefit tradeoff, so curious if that's one way you help reduce risk of inflated scores?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914419,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "08/26/2022 05:03:25",
          "content": "<p>After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914429,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/26/2022 05:17:07",
          "content": "<p>Hi Senkin13. I'm sorry to see you dropped on private LB. Like you say, your drop was probably caused by your features and/or types of models which overfitted to train and public LB. </p>\n<p>For me, I have 26 submissions over public LB 0.8010. They are all my same approach but have different seeds, different features, different architectures, etc. The CV scores all range from 0.7990 thru 0.8000. And all 26 of these submissions range from private LB 0.8082 thru LB 0.8085. So seeds and features doesn't change my final ranking too much. I think my approach which trained with test data worked very well on private LB and helped me maintain my public LB rank on private LB.</p>\n<p>This is kind of lucky. We didn't really know what approaches that worked on CV and public LB would also work on private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914468,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "08/26/2022 05:43:40",
          "content": "<p>Hi Chris, thanks for your encourage.I found some teams including you win gold with very strong NN model.That is definitely no t luck, learn from you a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1914139,
      "author_name": "chaudharypriyanshu",
      "author_url": "",
      "post_date": "08/25/2022 19:02:50",
      "content": "<p>For us features diversity was the key factor, then ensembling weights( calculated luck-based) and heavily correlated CV,LB. so we chose our public best and cv best.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1914425,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "08/26/2022 05:10:16",
          "content": "<p>yes, diverse team work reduce the luck factor,congraluations for your huge shakeup.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1912893": "Congratulations to all winners survived in such a shake competition.I worked hard for this competition but unfortunately I dropped from 15th to 90th,I want to share some possible reasons why I drop, hope you guys also can share yours.\n1. I focused on feature engineering/feature selection to improve my lightgbm cv ,best single model cv is 0.8005, private LB is 0.80763.I think the distribution difference between train and private-test and metric noise are huge that trust cv didn't work in this competition.\n2. I use ridge regressor for stacking that cause overfitting to train cv, maybe blend according to public lb is better choise.\n3. I created 8 neural network  models ,their ensemble's cv is 0.797 public LB 0.797,private LB 0.8049, after ensemble to gbdt models, the public score become better but private score worse,maybe nn models can not help ensemble.\n4. I only use raddar's clean data for feature engineering,I don't know if raw data is better or worse,please share your exeperiments results.(I found top rankers and high score kernel notebook use floor preprocess to achieve better private score)\n5. I merge train and test data then create very many features,not only groupby customer but also groupby category|month then calculate difference ,maybe merge train and test cause overfitting.\n6. After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.",
    "1912905": "I feel skill/luck ratio for this competition is less than 0.5 ☹️",
    "1912909": "\"best single model cv is 0.8005, private LB is 0.8007\" I think you means \"private LB is 0.8070\"?\n I was too busy to spend much time in this competition. But from what I observed, one of the problem using stacking could be: when doing cv, prediction scales are quite different, which will have a huge effect on the performance. \nI using lightgbm to  do the stacking. For one fold, the model is early stopped in round 24, and the max prediction is lower than 0.6(because it's not totally trained). For other folds, the problem exists still. So if you do nothing about the scaling problem, even the average amex_metric is great, the acutal result will be very bad. So after the prediction for each fold, I rescale the prediction to the range 0.0~1.0, which improve the result, but still it's not the perfect way, I guess.\nAnyway, that's probably not the reason in your case to shake down. Maybe it's just bad luck.",
    "1912912": "I hope so,but many GMs are at top rank,they should did some solid work",
    "1912914": "my best single model lgbm cv0.79932 public lb0.79980 private lb0.80731, found with tabnet ensemble seems to improve a lot. Unfortunately we ended up missing the best private LB",
    "1912920": "I also tried scale predictions values almost no changes.I prefer to think single blending is better.",
    "1912932": "Thanks for sharing the solution. I worked on this competition for less than a week or so, I was busy with Feedback. I solely focused on bringing out diverse models without much focus on FE part as I was seeing just minor improvements. My currently ensemble includes variety of models 1dcnn, tabnet,NN, lgb,catboost and xgboost. Thankfully I got shakeup, I regret for not joining earlier now :)",
    "1913088": "I notice that for my score, +/-0.0001, even the fourth digit, is +/- 15 slots or so.\n\nBesides pure luck, things that might've helped me jump higher:\n* \"miss next payment\" prediction on private test data. This may have helped avoid some of the issues with using a model trained on a different time period?\n* I dropped B_29.\n\nRe: #4 I used raddar's clean data as well.",
    "1913097": "P.s. My best single model was, well, 49th overall. \n\n0.80798 private\n0.79889 public\n\nOh, and I trained on full dataset (4 times over) with no folds or early stopping. I can't think of any reason that would really help on private leaderboard more than public, though.",
    "1913114": "Once again, bad decision on submissions chosen for me 😓\n\nWent for best CV and best LB, which are both just under bronze on private.. ensemble of both was Silver.. Unselected it in the last few hours lol",
    "1913870": "Our best submission CV score was 0.80002, Public score was 0.79956, Private score was 0.80739, we shook up over 1400th, not sure why.\n(Our final submission was Stacking with 18 models, using ridge as Mr.@senkin13 did).\n\nOur Stacking used 11 models GBDT (out of 18 overall)\nI guess that model was Fit for Private.\n(Of those models, 5 are NN-based and 2 are Logistic regression.)",
    "1913881": "I ran an experiment scaling/fix bias of features in train, public and private separatetly. It didn't changed train scores much, but it leaded to worse results in public and private.",
    "1913884": "Guess CV is not a good measurement for evaluating the performance in PriLB this time. Our CV improvement always align with publicLB but fall hard in private. We've tried several methods to estimate our performance in PrivateLB but no luck with it, you can read our solution post (maybe tomorrow or several hours later) to have more information about our crazy approaches 😃",
    "1914012": "Echoing @andrew60909 -- my honest feeling is that a robust approach is necessary but not sufficient for reaching gold in this competition. Exploiting the test data well seems like a winning theme we could have used more, but I still think there is ultimately a significant amount of noise. Public LB and CV actually aligned really well for us but seemed meaningfully decoupled from Private LB. Maybe Public is too similar to training data and less similar to Private, but I'm skeptical -- I think that data drift in this problem domain is less of an issue than the raw volatility of the metric. We've been joking about how it's like high on the LB we're just fighting over flipping one customer on the right or wrong side of the threshold.\n\nSome of our solution stats--\nCV .8006, Public LB 92nd, Private Gold\nCV .8024, Public LB 6th, Higher Private Gold \nCV .8031, Public LB 5th, Private .00009 off Gold\n\n5th decimal struggles!",
    "1914139": "For us features diversity was the key factor, then ensembling weights( calculated luck-based) and heavily correlated CV,LB. so we chose our public best and cv best.",
    "1914165": "senkin13 Yes of course. It was a twitter-referenced half-joke. From the published solutions it seems leveraging test data in an innovative way provides a certain edge, so there's certainly a skill factor here. \nBut in all seriousness, in terms of final ranking, luck can easily have an impact of maybe 50 place difference or more. Some shakeup teams believe their edge is diversity but I don't think that's really the case.",
    "1914177": "And it's probably a provable hypothesis if enough teams bothered to rerun their exact models with different seeds. See the variance first hand. \n\nThat said, a couple thoughts:\n* Seasonality? Public and train shared seasonality, but private was way different. Maximal difference, even (6mo offset). This may have been the number 1 issue, and even just unsupervised pretraining on private data (@cdeotte) or any other way of leveraging the raw data may be a common thread of those who avoided shakedown?\n* Besides \"it's just luck\" there's the issue of, well, not truly over fitting CV or leaderboard, but simply having higher likelihood of inflated reported results. Which increases chance (relative to others) of shakedown. \n\nLet me give an example:\nOne team tries 5 things that are all different, but unknowingly they are unimportant deltas. So it is like trying 5 different seeds. They haven't necessarily made any methodological mistake, and it doesn't hurt their final chances any, they haven't overfit in the sense of worse predictive power. But their equal predictive power on private gets inflated score on CV and public. \n\nA second team does the same, but 200 times, not 5 times. Ground truth they are equal, but CV and LB scores are inflated in favor the 200 experiment team. The 5 experiment team has higher chance of increased rank, even though - with fewer experiments - they might have the worse model. \n\nPoint is there's statistical reason to expect a bit more inflated public scores (and CV scores) on the people doing the most optimization work EVEN if that optimization is the right approach. So they have odds working (slightly) against them when it comes shakedown time.",
    "1914180": "I think the biggest numerical rank shakeups might be if you crossed a threshold compared with the best published public models? Just my guess. Ahead of the public on one LB, but behind it in the other.",
    "1914217": "The seasonality point is interesting, but I'm also skeptical that seasonality is significant in this specific problem (or that it is more significant than the level of metric noise).\n\nOn your second point, I partially but don't completely agree. There is a big difference between shot in the dark and rigorous experiments, and you should bring strong priors to what experiments you try and how you evaluate them. It's one thing if trying different seeds improves your scores, but another if your new experiment is on good feature ideas or the addition of diverse models to an ensemble. There's statistical reason to expect inflated Public LB scores based on sheer experiment count, but also statistical reason to expect improvement across the board (CV, public, and private) with a large quantity of experiments if they are done well. \n\nAt least for our subs, the experience here is that there is a meaningful correlation between better CV/LB and private LB, but it is weak enough that there can be a lot of fluctuation. So, our final changes may have been suboptimal experiments, but this could also mean that even though our solution improved on expectation that luck didn't favor us in the end. A clear example -- you would prefer a model that does somewhat better on 3/5 CV folds but a tiny bit worse on 2/5, but can plausibly expect that this model may underperform an older one by chance on a specific dataset.",
    "1914226": "Thank you for adding more nuance to the discussion. I completely agree. And in fact the last thing I want to do is sound like I'm lecturing the experts. Only trying to highlight a general principle for lurkers and in fact you and many others would know better than me exactly how to mitigate any such issues. \n\nAs the noisy data became obvious, and as you got more tuned models, did you and do you ever do any multi seed variance studies to look at score mean and std dev in determining which experiments were the best? There's a big cost/benefit tradeoff, so curious if that's one way you help reduce risk of inflated scores?",
    "1914419": "After I read many winners' solution,I feel my biggest issue is feature engineering,I created some train|public fitting features (more complex)first,then add some private fitting features(more simple),but cv become bad,then I removed them,that made my single model become worse and worse.It's too hard to solve unseen data drift and noise metric problem,maybe need some luck,or some different approches from team members.",
    "1914420": "I think your shakeup come from better features that fit private data",
    "1914425": "yes, diverse team work reduce the luck factor,congraluations for your huge shakeup.",
    "1914427": "thanks, your approch is solid to fight against uncertainty.",
    "1914429": "Hi Senkin13. I'm sorry to see you dropped on private LB. Like you say, your drop was probably caused by your features and/or types of models which overfitted to train and public LB. \n\nFor me, I have 26 submissions over public LB 0.8010. They are all my same approach but have different seeds, different features, different architectures, etc. The CV scores all range from 0.7990 thru 0.8000. And all 26 of these submissions range from private LB 0.8082 thru LB 0.8085. So seeds and features doesn't change my final ranking too much. I think my approach which trained with test data worked very well on private LB and helped me maintain my public LB rank on private LB.\n\nThis is kind of lucky. We didn't really know what approaches that worked on CV and public LB would also work on private LB.",
    "1914468": "Hi Chris, thanks for your encourage.I found some teams including you win gold with very strong NN model.That is definitely no t luck, learn from you a lot.",
    "1914841": "roberthatch \nThank you for reply.I agree. I am sure that my model did not reach the public many Public fit Notebooks.\n(I was sad because I mixed 18 models and couldn't beat the public Notebook...)\n\n@senkin13 \nThank you for reply\nI've checked but the single model made with any of the features was not a special Private fit. I really don't know...",
    "1915648": "...a little more empirical evidence in support of the suggestion by Raddar of [\"*...second sub - best CV model/2 + best LB model/2...*\"](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337609#1858410)\n\nAll the best,\ncarl"
  },
  "source": "meta"
}