{
  "id": 347540,
  "title": "Using high score notebook is risky?",
  "url": "/competitions/amex-default-prediction/discussion/347540",
  "author_name": "",
  "post_date": "2022-08-24T14:59:07.668684700Z",
  "votes": 6,
  "comment_count": 28,
  "views": 0,
  "content": "<p>I am wondering if I should trust CV or LB.</p>\n<p>The fact of the matter is that in this competition, the high score notebooks are in disarray, and combining them can get you into the top 10%.</p>\n<p>In my case, I use a combination of high score notebooks and my own model, but it doesn't seem like a good approach.</p>\n<p>Should I just use my own model as long as the cv is not measured?</p>\n<p>What do you guys think?</p>",
  "messages": [
    {
      "id": "1912174",
      "postDate": "08/24/2022 14:59:07",
      "content": "<p>I am wondering if I should trust CV or LB.</p>\n<p>The fact of the matter is that in this competition, the high score notebooks are in disarray, and combining them can get you into the top 10%.</p>\n<p>In my case, I use a combination of high score notebooks and my own model, but it doesn't seem like a good approach.</p>\n<p>Should I just use my own model as long as the cv is not measured?</p>\n<p>What do you guys think?</p>",
      "rawMarkdown": "I am wondering if I should trust CV or LB.\n\nThe fact of the matter is that in this competition, the high score notebooks are in disarray, and combining them can get you into the top 10%.\n\nIn my case, I use a combination of high score notebooks and my own model, but it doesn't seem like a good approach.\n\nShould I just use my own model as long as the cv is not measured?\n\nWhat do you guys think?",
      "votes": null
    },
    {
      "id": "1912193",
      "postDate": "08/24/2022 15:17:21",
      "content": "<p>If it is cv based ensemble, should be fine; otherwise, then it depends on luck. But no risk no fun, I think we should take the risk somehow, because the test dataset is significantly larger than the trainning dataset.</p>",
      "rawMarkdown": "If it is cv based ensemble, should be fine; otherwise, then it depends on luck. But no risk no fun, I think we should take the risk somehow, because the test dataset is significantly larger than the trainning dataset.",
      "votes": null
    },
    {
      "id": "1912196",
      "postDate": "08/24/2022 15:21:22",
      "content": "<p>Since we only get to pick two submissions, the usual approach is:</p>\n<blockquote>\n  <p>1 submission which is highest LB<br>\n  1 submission which is best CV</p>\n</blockquote>\n<p>I think picking submissions based off LB will work for this competition, because number of customers with 13 statements (easier to predict ones) is higher in test data than train..</p>",
      "rawMarkdown": "Since we only get to pick two submissions, the usual approach is:\n\n> 1 submission which is highest LB\n> 1 submission which is best CV\n\nI think picking submissions based off LB will work for this competition, because number of customers with 13 statements (easier to predict ones) is higher in test data than train..",
      "votes": null
    },
    {
      "id": "1912199",
      "postDate": "08/24/2022 15:24:51",
      "content": "<p>I feel that I have the lack of experience now and can't make a smart decision at the moment. On the one hand, ensembling is good for quality and public lb consists of almost half of data, which is a quite large amount, so it feels like that high scoring ensembles should be good on a private part as well. On the other hand, when you choose weights for an ensemble and look only on a public lb, there is a possibility that you might choose lucky weights for a good public score and a bad private one. So I can't advise anything, I am also very confused now, I've seen competitions with huge shake ups and without shake ups at all.</p>",
      "rawMarkdown": "I feel that I have the lack of experience now and can't make a smart decision at the moment. On the one hand, ensembling is good for quality and public lb consists of almost half of data, which is a quite large amount, so it feels like that high scoring ensembles should be good on a private part as well. On the other hand, when you choose weights for an ensemble and look only on a public lb, there is a possibility that you might choose lucky weights for a good public score and a bad private one. So I can't advise anything, I am also very confused now, I've seen competitions with huge shake ups and without shake ups at all.",
      "votes": null
    },
    {
      "id": "1912200",
      "postDate": "08/24/2022 15:25:23",
      "content": "<p>Even if an evil high score notebook was used?</p>",
      "rawMarkdown": "Even if an evil high score notebook was used?",
      "votes": null
    },
    {
      "id": "1912207",
      "postDate": "08/24/2022 15:31:38",
      "content": "<p>No comment on that one 👀😄</p>",
      "rawMarkdown": "No comment on that one 👀😄",
      "votes": null
    },
    {
      "id": "1912215",
      "postDate": "08/24/2022 15:35:03",
      "content": "<p>Your answer is one answer in a way.<br>\nhaha :)</p>",
      "rawMarkdown": "Your answer is one answer in a way.\nhaha :)",
      "votes": null
    },
    {
      "id": "1912237",
      "postDate": "08/24/2022 15:45:04",
      "content": "<p>I tend to go with a) best CV model b) best CV model/2 + best LB model/2. This way you hedge \"what if public LB is better choice\" scenario, as CV models usually underperform on public LB</p>",
      "rawMarkdown": "I tend to go with a) best CV model b) best CV model/2 + best LB model/2. This way you hedge \"what if public LB is better choice\" scenario, as CV models usually underperform on public LB",
      "votes": null
    },
    {
      "id": "1912241",
      "postDate": "08/24/2022 15:47:23",
      "content": "<p>I suggest one could choose the best model from the public leaderboard and another best CV model to perhaps complete the candidates for the final scoring. </p>",
      "rawMarkdown": "I suggest one could choose the best model from the public leaderboard and another best CV model to perhaps complete the candidates for the final scoring.",
      "votes": null
    },
    {
      "id": "1912253",
      "postDate": "08/24/2022 15:55:57",
      "content": "<p>Shakeup could happen</p>",
      "rawMarkdown": "Shakeup could happen",
      "votes": null
    },
    {
      "id": "1912256",
      "postDate": "08/24/2022 15:56:39",
      "content": "<p>Good luck solo gold!</p>",
      "rawMarkdown": "Good luck solo gold!",
      "votes": null
    },
    {
      "id": "1912262",
      "postDate": "08/24/2022 15:58:33",
      "content": "<p>Why do you think that?</p>",
      "rawMarkdown": "Why do you think that?",
      "votes": null
    },
    {
      "id": "1912266",
      "postDate": "08/24/2022 16:01:36",
      "content": "<p>You're 371st, great! My score is the same as you, but rank of the public leaderboard is different. What does it suggest?<br>\nYou would not need bronze medal anymore, therefore you should trust yourself for gold medal. If you behave like a lot, you would have no chance. Ah, silver and gold!</p>",
      "rawMarkdown": "You're 371st, great! My score is the same as you, but rank of the public leaderboard is different. What does it suggest?\nYou would not need bronze medal anymore, therefore you should trust yourself for gold medal. If you behave like a lot, you would have no chance. Ah, silver and gold!",
      "votes": null
    },
    {
      "id": "1912267",
      "postDate": "08/24/2022 16:03:09",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F3e1c78d6784fe95dd59401373ac52391%2Fusa_10years.png?generation=1661356599256534&amp;alt=media\" alt=\"\"></p>\n<p>This is just USA-10year bond chart.</p>\n<p>Unlike 2018(maybe public LB), in 2019(maybe private LB) the FOMC read signs of a economic downturn and cut rates sharply.<br>\nWhat does this have to do with personal debt?<br>\nThe host of this contest may want to see how the model responds to sudden volatility in the private sector.<br>\nIt may be better to strive for a robust model than a visible score.</p>\n<p>It's just my personal opinion. :)</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F3e1c78d6784fe95dd59401373ac52391%2Fusa_10years.png?generation=1661356599256534&alt=media)\n\nThis is just USA-10year bond chart.\n\nUnlike 2018(maybe public LB), in 2019(maybe private LB) the FOMC read signs of a economic downturn and cut rates sharply.\nWhat does this have to do with personal debt?\nThe host of this contest may want to see how the model responds to sudden volatility in the private sector.\nIt may be better to strive for a robust model than a visible score.\n\nIt's just my personal opinion. :)",
      "votes": null
    },
    {
      "id": "1912291",
      "postDate": "08/24/2022 16:17:30",
      "content": "<p>You're anytime great! I would like to know your general method and I like the former cat picture.</p>",
      "rawMarkdown": "You're anytime great! I would like to know your general method and I like the former cat picture.",
      "votes": null
    },
    {
      "id": "1912342",
      "postDate": "08/24/2022 16:49:01",
      "content": "<p>Haha, similar. I go with best LB + ( best LB + best cv ) / 2.</p>",
      "rawMarkdown": "Haha, similar. I go with best LB + ( best LB + best cv ) / 2.",
      "votes": null
    },
    {
      "id": "1912360",
      "postDate": "08/24/2022 17:04:39",
      "content": "<p>I also think that it will be. There are like 1800 people with scores 0.799***, the difference between them is very small, so evaluating on another fold of data can give another distribution of these small scores after first 3 digits.</p>",
      "rawMarkdown": "I also think that it will be. There are like 1800 people with scores 0.799***, the difference between them is very small, so evaluating on another fold of data can give another distribution of these small scores after first 3 digits.",
      "votes": null
    },
    {
      "id": "1912363",
      "postDate": "08/24/2022 17:07:05",
      "content": "<p>I strongly suspect that the popular public solutions are overfit to public LB. One way to think about this is to consider the pool of people working on and willing to share public solutions -- from this pool, the solutions that happen to score the best are disproportionately likely to be published and noticed within the pool of similar quality and rigor. There is extremely strong feedback from public LB, but much less feedback from CV, which is a recipe for instability with a noisy metric.</p>\n<p>What you aim for is the best of both worlds. If public approaches are truly robust, incorporating their ideas into your own methodology should also look good on CV (if there isn't enough time left for this, that's a better justification for a simpler blend - you could still see a diversity benefit from the approach being different). You should be able to get at least a rough directional correspondence between CV and LB. CV helps measure if your model is robust on contemporaneous data from the problem domain, while LB helps measure if your model generalizes to future data from the problem domain - both are important! In this competition I have consistently considered it to be a warning sign if either of CV or LB goes down, and our current best public LB and CV are the same submission. Maybe we are lucky (or unlucky, we'll see soon enough 😅), but I think it's been possible in this comp to have a reasonably stable framework for improvements.</p>\n<p>All that said, others have already mentioned good hedging strategies for combining your best personal work with public work. Those are good suggestions!</p>\n<p>Edit: one thing to add -- a good way to regularize your own public LB overfitting is to only submit when you see CV improvements. If LB also improves, great! But if it goes down, you can also evaluate depending on the size of the CV improvement whether there's noisiness (either in CV or LB). Even better might be to only submit when you see a substantial CV improvement, and then fully expect that to translate to an LB improvement - believe that's been true 100% of the time for me in this comp, where substantial is say ~.0004+ improvement.</p>",
      "rawMarkdown": "I strongly suspect that the popular public solutions are overfit to public LB. One way to think about this is to consider the pool of people working on and willing to share public solutions -- from this pool, the solutions that happen to score the best are disproportionately likely to be published and noticed within the pool of similar quality and rigor. There is extremely strong feedback from public LB, but much less feedback from CV, which is a recipe for instability with a noisy metric.\n\nWhat you aim for is the best of both worlds. If public approaches are truly robust, incorporating their ideas into your own methodology should also look good on CV (if there isn't enough time left for this, that's a better justification for a simpler blend - you could still see a diversity benefit from the approach being different). You should be able to get at least a rough directional correspondence between CV and LB. CV helps measure if your model is robust on contemporaneous data from the problem domain, while LB helps measure if your model generalizes to future data from the problem domain - both are important! In this competition I have consistently considered it to be a warning sign if either of CV or LB goes down, and our current best public LB and CV are the same submission. Maybe we are lucky (or unlucky, we'll see soon enough 😅), but I think it's been possible in this comp to have a reasonably stable framework for improvements.\n\nAll that said, others have already mentioned good hedging strategies for combining your best personal work with public work. Those are good suggestions!\n\nEdit: one thing to add -- a good way to regularize your own public LB overfitting is to only submit when you see CV improvements. If LB also improves, great! But if it goes down, you can also evaluate depending on the size of the CV improvement whether there's noisiness (either in CV or LB). Even better might be to only submit when you see a substantial CV improvement, and then fully expect that to translate to an LB improvement - believe that's been true 100% of the time for me in this comp, where substantial is say ~.0004+ improvement.",
      "votes": null
    },
    {
      "id": "1912378",
      "postDate": "08/24/2022 17:25:26",
      "content": "<p>thanks for the tip, I have something to do tonight.</p>",
      "rawMarkdown": "thanks for the tip, I have something to do tonight.",
      "votes": null
    },
    {
      "id": "1912433",
      "postDate": "08/24/2022 18:07:42",
      "content": "<p>You are indeed lucky that CV &amp; LB submission is the same! My last submission is also best CV and best LB - this really helps a lot when choosing final 2 :)</p>",
      "rawMarkdown": "You are indeed lucky that CV & LB submission is the same! My last submission is also best CV and best LB - this really helps a lot when choosing final 2 :)",
      "votes": null
    },
    {
      "id": "1912635",
      "postDate": "08/24/2022 21:55:44",
      "content": "<p>Forgot to try this tip think you recommended this long back 😐</p>",
      "rawMarkdown": "Forgot to try this tip think you recommended this long back 😐",
      "votes": null
    },
    {
      "id": "1912682",
      "postDate": "08/24/2022 23:31:02",
      "content": "<p>Are the high scoring models sensitive to seed?  I've been working out of R and haven't rerun a high scoring model with a different seed in python.  My efforts to use similar features and hyper parameters to the highest scoring models don't produce the same results, they're maybe 0.002 lower.</p>",
      "rawMarkdown": "Are the high scoring models sensitive to seed?  I've been working out of R and haven't rerun a high scoring model with a different seed in python.  My efforts to use similar features and hyper parameters to the highest scoring models don't produce the same results, they're maybe 0.002 lower.",
      "votes": null
    },
    {
      "id": "1912921",
      "postDate": "08/25/2022 03:38:14",
      "content": "<p>You chose correct. I made a wrong choice. If I had chosen an ensemble model with high score notebooks, I would have gotten a medal!</p>",
      "rawMarkdown": "You chose correct. I made a wrong choice. If I had chosen an ensemble model with high score notebooks, I would have gotten a medal!",
      "votes": null
    },
    {
      "id": "1912950",
      "postDate": "08/25/2022 04:11:10",
      "content": "<p>It is foolish to talk about what if we had done so after the answer is given!<br>\nI have thought deeply about the notebooks I use, and it is not a matter of consequence whether I did or did not do it.</p>",
      "rawMarkdown": "It is foolish to talk about what if we had done so after the answer is given!\nI have thought deeply about the notebooks I use, and it is not a matter of consequence whether I did or did not do it.",
      "votes": null
    },
    {
      "id": "1912954",
      "postDate": "08/25/2022 04:15:41",
      "content": "<p>Sorry, I just to say you're great!</p>",
      "rawMarkdown": "Sorry, I just to say you're great!",
      "votes": null
    },
    {
      "id": "1912994",
      "postDate": "08/25/2022 04:55:52",
      "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> You make a very important point about public notebooks being a large pool and the potential for overfitting. I want to highlight that similar logic can be used in reverse to decide which notebooks can be trusted a little more. The <em>first</em> innovative notebook -&gt; high trust, lower risk of overfit. The derivative notebooks - even with clever tweaks or twice the features or whatever -&gt; lower trust, higher risk of overfit.</p>\n<p>And read what the author says (or doesn't say) with regards to CV, seed, methodology, etc.</p>\n<p>My personal choice was to only ensemble with <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> 's model. Nothing 'higher' than it on the leaderboard seemed much better, and even the good ones seemed like they might have just happened to score higher, but not be any better quality. I stayed far away from the mega-ensembles, especially with what appeared to be minimal score increase even on public.</p>\n<p>(And in the end, my solo model beat even the <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> ensemble, as my choice of missed payment mega-feature luckily was way better on the private dataset than on the public for some reason. )</p>",
      "rawMarkdown": "aquatic You make a very important point about public notebooks being a large pool and the potential for overfitting. I want to highlight that similar logic can be used in reverse to decide which notebooks can be trusted a little more. The *first* innovative notebook -> high trust, lower risk of overfit. The derivative notebooks - even with clever tweaks or twice the features or whatever -> lower trust, higher risk of overfit.\n\nAnd read what the author says (or doesn't say) with regards to CV, seed, methodology, etc.\n\nMy personal choice was to only ensemble with @ragnar123 's model. Nothing 'higher' than it on the leaderboard seemed much better, and even the good ones seemed like they might have just happened to score higher, but not be any better quality. I stayed far away from the mega-ensembles, especially with what appeared to be minimal score increase even on public.\n\n(And in the end, my solo model beat even the @ragnar123 ensemble, as my choice of missed payment mega-feature luckily was way better on the private dataset than on the public for some reason. )",
      "votes": null
    },
    {
      "id": "1913729",
      "postDate": "08/25/2022 13:23:54",
      "content": "<p>I think the point I made here about PLB overfitting is generally true, but I was wrong in this case. Strong public solutions seem to still work well for private. Maybe the better takeaway here was that these solutions were robust, relatively simple models that generalized well into the future. Or maybe it played out that way by chance. I think it's hard to take any clear lesson from it, but it's interesting and I'm curious if anyone else has a clear take on this.</p>",
      "rawMarkdown": "I think the point I made here about PLB overfitting is generally true, but I was wrong in this case. Strong public solutions seem to still work well for private. Maybe the better takeaway here was that these solutions were robust, relatively simple models that generalized well into the future. Or maybe it played out that way by chance. I think it's hard to take any clear lesson from it, but it's interesting and I'm curious if anyone else has a clear take on this.",
      "votes": null
    },
    {
      "id": "1922190",
      "postDate": "09/01/2022 10:30:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yukio0201\" target=\"_blank\">@yukio0201</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @yukio0201, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    },
    {
      "id": "1923068",
      "postDate": "09/02/2022 01:34:21",
      "content": "<p>Hi Yang.<br>\nI I've filled out your survey.</p>",
      "rawMarkdown": "Hi Yang.\nI I've filled out your survey.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1912193,
      "author_name": "meli19",
      "author_url": "",
      "post_date": "08/24/2022 15:17:21",
      "content": "<p>If it is cv based ensemble, should be fine; otherwise, then it depends on luck. But no risk no fun, I think we should take the risk somehow, because the test dataset is significantly larger than the trainning dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912256,
          "author_name": "zhehaoliang",
          "author_url": "",
          "post_date": "08/24/2022 15:56:39",
          "content": "<p>Good luck solo gold!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912291,
          "author_name": "hdynamics",
          "author_url": "",
          "post_date": "08/24/2022 16:17:30",
          "content": "<p>You're anytime great! I would like to know your general method and I like the former cat picture.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912196,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "08/24/2022 15:21:22",
      "content": "<p>Since we only get to pick two submissions, the usual approach is:</p>\n<blockquote>\n  <p>1 submission which is highest LB<br>\n  1 submission which is best CV</p>\n</blockquote>\n<p>I think picking submissions based off LB will work for this competition, because number of customers with 13 statements (easier to predict ones) is higher in test data than train..</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912200,
          "author_name": "yukio0201",
          "author_url": "",
          "post_date": "08/24/2022 15:25:23",
          "content": "<p>Even if an evil high score notebook was used?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912207,
          "author_name": "julianmukaj",
          "author_url": "",
          "post_date": "08/24/2022 15:31:38",
          "content": "<p>No comment on that one 👀😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912215,
          "author_name": "yukio0201",
          "author_url": "",
          "post_date": "08/24/2022 15:35:03",
          "content": "<p>Your answer is one answer in a way.<br>\nhaha :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912266,
          "author_name": "hdynamics",
          "author_url": "",
          "post_date": "08/24/2022 16:01:36",
          "content": "<p>You're 371st, great! My score is the same as you, but rank of the public leaderboard is different. What does it suggest?<br>\nYou would not need bronze medal anymore, therefore you should trust yourself for gold medal. If you behave like a lot, you would have no chance. Ah, silver and gold!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912199,
      "author_name": "manwithaflower",
      "author_url": "",
      "post_date": "08/24/2022 15:24:51",
      "content": "<p>I feel that I have the lack of experience now and can't make a smart decision at the moment. On the one hand, ensembling is good for quality and public lb consists of almost half of data, which is a quite large amount, so it feels like that high scoring ensembles should be good on a private part as well. On the other hand, when you choose weights for an ensemble and look only on a public lb, there is a possibility that you might choose lucky weights for a good public score and a bad private one. So I can't advise anything, I am also very confused now, I've seen competitions with huge shake ups and without shake ups at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912237,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "08/24/2022 15:45:04",
      "content": "<p>I tend to go with a) best CV model b) best CV model/2 + best LB model/2. This way you hedge \"what if public LB is better choice\" scenario, as CV models usually underperform on public LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912342,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "08/24/2022 16:49:01",
          "content": "<p>Haha, similar. I go with best LB + ( best LB + best cv ) / 2.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912635,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "08/24/2022 21:55:44",
          "content": "<p>Forgot to try this tip think you recommended this long back 😐</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912241,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/24/2022 15:47:23",
      "content": "<p>I suggest one could choose the best model from the public leaderboard and another best CV model to perhaps complete the candidates for the final scoring. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912253,
      "author_name": "zhehaoliang",
      "author_url": "",
      "post_date": "08/24/2022 15:55:57",
      "content": "<p>Shakeup could happen</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912262,
          "author_name": "yukio0201",
          "author_url": "",
          "post_date": "08/24/2022 15:58:33",
          "content": "<p>Why do you think that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912360,
          "author_name": "manwithaflower",
          "author_url": "",
          "post_date": "08/24/2022 17:04:39",
          "content": "<p>I also think that it will be. There are like 1800 people with scores 0.799***, the difference between them is very small, so evaluating on another fold of data can give another distribution of these small scores after first 3 digits.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912267,
      "author_name": "yuuniekiri",
      "author_url": "",
      "post_date": "08/24/2022 16:03:09",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F3e1c78d6784fe95dd59401373ac52391%2Fusa_10years.png?generation=1661356599256534&amp;alt=media\" alt=\"\"></p>\n<p>This is just USA-10year bond chart.</p>\n<p>Unlike 2018(maybe public LB), in 2019(maybe private LB) the FOMC read signs of a economic downturn and cut rates sharply.<br>\nWhat does this have to do with personal debt?<br>\nThe host of this contest may want to see how the model responds to sudden volatility in the private sector.<br>\nIt may be better to strive for a robust model than a visible score.</p>\n<p>It's just my personal opinion. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912363,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "08/24/2022 17:07:05",
      "content": "<p>I strongly suspect that the popular public solutions are overfit to public LB. One way to think about this is to consider the pool of people working on and willing to share public solutions -- from this pool, the solutions that happen to score the best are disproportionately likely to be published and noticed within the pool of similar quality and rigor. There is extremely strong feedback from public LB, but much less feedback from CV, which is a recipe for instability with a noisy metric.</p>\n<p>What you aim for is the best of both worlds. If public approaches are truly robust, incorporating their ideas into your own methodology should also look good on CV (if there isn't enough time left for this, that's a better justification for a simpler blend - you could still see a diversity benefit from the approach being different). You should be able to get at least a rough directional correspondence between CV and LB. CV helps measure if your model is robust on contemporaneous data from the problem domain, while LB helps measure if your model generalizes to future data from the problem domain - both are important! In this competition I have consistently considered it to be a warning sign if either of CV or LB goes down, and our current best public LB and CV are the same submission. Maybe we are lucky (or unlucky, we'll see soon enough 😅), but I think it's been possible in this comp to have a reasonably stable framework for improvements.</p>\n<p>All that said, others have already mentioned good hedging strategies for combining your best personal work with public work. Those are good suggestions!</p>\n<p>Edit: one thing to add -- a good way to regularize your own public LB overfitting is to only submit when you see CV improvements. If LB also improves, great! But if it goes down, you can also evaluate depending on the size of the CV improvement whether there's noisiness (either in CV or LB). Even better might be to only submit when you see a substantial CV improvement, and then fully expect that to translate to an LB improvement - believe that's been true 100% of the time for me in this comp, where substantial is say ~.0004+ improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912378,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "08/24/2022 17:25:26",
          "content": "<p>thanks for the tip, I have something to do tonight.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912433,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "08/24/2022 18:07:42",
          "content": "<p>You are indeed lucky that CV &amp; LB submission is the same! My last submission is also best CV and best LB - this really helps a lot when choosing final 2 :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912994,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/25/2022 04:55:52",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> You make a very important point about public notebooks being a large pool and the potential for overfitting. I want to highlight that similar logic can be used in reverse to decide which notebooks can be trusted a little more. The <em>first</em> innovative notebook -&gt; high trust, lower risk of overfit. The derivative notebooks - even with clever tweaks or twice the features or whatever -&gt; lower trust, higher risk of overfit.</p>\n<p>And read what the author says (or doesn't say) with regards to CV, seed, methodology, etc.</p>\n<p>My personal choice was to only ensemble with <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> 's model. Nothing 'higher' than it on the leaderboard seemed much better, and even the good ones seemed like they might have just happened to score higher, but not be any better quality. I stayed far away from the mega-ensembles, especially with what appeared to be minimal score increase even on public.</p>\n<p>(And in the end, my solo model beat even the <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> ensemble, as my choice of missed payment mega-feature luckily was way better on the private dataset than on the public for some reason. )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1913729,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "08/25/2022 13:23:54",
          "content": "<p>I think the point I made here about PLB overfitting is generally true, but I was wrong in this case. Strong public solutions seem to still work well for private. Maybe the better takeaway here was that these solutions were robust, relatively simple models that generalized well into the future. Or maybe it played out that way by chance. I think it's hard to take any clear lesson from it, but it's interesting and I'm curious if anyone else has a clear take on this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912682,
      "author_name": "andyatkinson",
      "author_url": "",
      "post_date": "08/24/2022 23:31:02",
      "content": "<p>Are the high scoring models sensitive to seed?  I've been working out of R and haven't rerun a high scoring model with a different seed in python.  My efforts to use similar features and hyper parameters to the highest scoring models don't produce the same results, they're maybe 0.002 lower.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912921,
      "author_name": "hdynamics",
      "author_url": "",
      "post_date": "08/25/2022 03:38:14",
      "content": "<p>You chose correct. I made a wrong choice. If I had chosen an ensemble model with high score notebooks, I would have gotten a medal!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912950,
          "author_name": "yukio0201",
          "author_url": "",
          "post_date": "08/25/2022 04:11:10",
          "content": "<p>It is foolish to talk about what if we had done so after the answer is given!<br>\nI have thought deeply about the notebooks I use, and it is not a matter of consequence whether I did or did not do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1912954,
          "author_name": "hdynamics",
          "author_url": "",
          "post_date": "08/25/2022 04:15:41",
          "content": "<p>Sorry, I just to say you're great!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1922190,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 10:30:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yukio0201\" target=\"_blank\">@yukio0201</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1923068,
          "author_name": "yukio0201",
          "author_url": "",
          "post_date": "09/02/2022 01:34:21",
          "content": "<p>Hi Yang.<br>\nI I've filled out your survey.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1912174": "I am wondering if I should trust CV or LB.\n\nThe fact of the matter is that in this competition, the high score notebooks are in disarray, and combining them can get you into the top 10%.\n\nIn my case, I use a combination of high score notebooks and my own model, but it doesn't seem like a good approach.\n\nShould I just use my own model as long as the cv is not measured?\n\nWhat do you guys think?",
    "1912193": "If it is cv based ensemble, should be fine; otherwise, then it depends on luck. But no risk no fun, I think we should take the risk somehow, because the test dataset is significantly larger than the trainning dataset.",
    "1912196": "Since we only get to pick two submissions, the usual approach is:\n\n> 1 submission which is highest LB\n> 1 submission which is best CV\n\nI think picking submissions based off LB will work for this competition, because number of customers with 13 statements (easier to predict ones) is higher in test data than train..",
    "1912199": "I feel that I have the lack of experience now and can't make a smart decision at the moment. On the one hand, ensembling is good for quality and public lb consists of almost half of data, which is a quite large amount, so it feels like that high scoring ensembles should be good on a private part as well. On the other hand, when you choose weights for an ensemble and look only on a public lb, there is a possibility that you might choose lucky weights for a good public score and a bad private one. So I can't advise anything, I am also very confused now, I've seen competitions with huge shake ups and without shake ups at all.",
    "1912200": "Even if an evil high score notebook was used?",
    "1912207": "No comment on that one 👀😄",
    "1912215": "Your answer is one answer in a way.\nhaha :)",
    "1912237": "I tend to go with a) best CV model b) best CV model/2 + best LB model/2. This way you hedge \"what if public LB is better choice\" scenario, as CV models usually underperform on public LB",
    "1912241": "I suggest one could choose the best model from the public leaderboard and another best CV model to perhaps complete the candidates for the final scoring.",
    "1912253": "Shakeup could happen",
    "1912256": "Good luck solo gold!",
    "1912262": "Why do you think that?",
    "1912266": "You're 371st, great! My score is the same as you, but rank of the public leaderboard is different. What does it suggest?\nYou would not need bronze medal anymore, therefore you should trust yourself for gold medal. If you behave like a lot, you would have no chance. Ah, silver and gold!",
    "1912267": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F3e1c78d6784fe95dd59401373ac52391%2Fusa_10years.png?generation=1661356599256534&alt=media)\n\nThis is just USA-10year bond chart.\n\nUnlike 2018(maybe public LB), in 2019(maybe private LB) the FOMC read signs of a economic downturn and cut rates sharply.\nWhat does this have to do with personal debt?\nThe host of this contest may want to see how the model responds to sudden volatility in the private sector.\nIt may be better to strive for a robust model than a visible score.\n\nIt's just my personal opinion. :)",
    "1912291": "You're anytime great! I would like to know your general method and I like the former cat picture.",
    "1912342": "Haha, similar. I go with best LB + ( best LB + best cv ) / 2.",
    "1912360": "I also think that it will be. There are like 1800 people with scores 0.799***, the difference between them is very small, so evaluating on another fold of data can give another distribution of these small scores after first 3 digits.",
    "1912363": "I strongly suspect that the popular public solutions are overfit to public LB. One way to think about this is to consider the pool of people working on and willing to share public solutions -- from this pool, the solutions that happen to score the best are disproportionately likely to be published and noticed within the pool of similar quality and rigor. There is extremely strong feedback from public LB, but much less feedback from CV, which is a recipe for instability with a noisy metric.\n\nWhat you aim for is the best of both worlds. If public approaches are truly robust, incorporating their ideas into your own methodology should also look good on CV (if there isn't enough time left for this, that's a better justification for a simpler blend - you could still see a diversity benefit from the approach being different). You should be able to get at least a rough directional correspondence between CV and LB. CV helps measure if your model is robust on contemporaneous data from the problem domain, while LB helps measure if your model generalizes to future data from the problem domain - both are important! In this competition I have consistently considered it to be a warning sign if either of CV or LB goes down, and our current best public LB and CV are the same submission. Maybe we are lucky (or unlucky, we'll see soon enough 😅), but I think it's been possible in this comp to have a reasonably stable framework for improvements.\n\nAll that said, others have already mentioned good hedging strategies for combining your best personal work with public work. Those are good suggestions!\n\nEdit: one thing to add -- a good way to regularize your own public LB overfitting is to only submit when you see CV improvements. If LB also improves, great! But if it goes down, you can also evaluate depending on the size of the CV improvement whether there's noisiness (either in CV or LB). Even better might be to only submit when you see a substantial CV improvement, and then fully expect that to translate to an LB improvement - believe that's been true 100% of the time for me in this comp, where substantial is say ~.0004+ improvement.",
    "1912378": "thanks for the tip, I have something to do tonight.",
    "1912433": "You are indeed lucky that CV & LB submission is the same! My last submission is also best CV and best LB - this really helps a lot when choosing final 2 :)",
    "1912635": "Forgot to try this tip think you recommended this long back 😐",
    "1912682": "Are the high scoring models sensitive to seed?  I've been working out of R and haven't rerun a high scoring model with a different seed in python.  My efforts to use similar features and hyper parameters to the highest scoring models don't produce the same results, they're maybe 0.002 lower.",
    "1912921": "You chose correct. I made a wrong choice. If I had chosen an ensemble model with high score notebooks, I would have gotten a medal!",
    "1912950": "It is foolish to talk about what if we had done so after the answer is given!\nI have thought deeply about the notebooks I use, and it is not a matter of consequence whether I did or did not do it.",
    "1912954": "Sorry, I just to say you're great!",
    "1912994": "aquatic You make a very important point about public notebooks being a large pool and the potential for overfitting. I want to highlight that similar logic can be used in reverse to decide which notebooks can be trusted a little more. The *first* innovative notebook -> high trust, lower risk of overfit. The derivative notebooks - even with clever tweaks or twice the features or whatever -> lower trust, higher risk of overfit.\n\nAnd read what the author says (or doesn't say) with regards to CV, seed, methodology, etc.\n\nMy personal choice was to only ensemble with @ragnar123 's model. Nothing 'higher' than it on the leaderboard seemed much better, and even the good ones seemed like they might have just happened to score higher, but not be any better quality. I stayed far away from the mega-ensembles, especially with what appeared to be minimal score increase even on public.\n\n(And in the end, my solo model beat even the @ragnar123 ensemble, as my choice of missed payment mega-feature luckily was way better on the private dataset than on the public for some reason. )",
    "1913729": "I think the point I made here about PLB overfitting is generally true, but I was wrong in this case. Strong public solutions seem to still work well for private. Maybe the better takeaway here was that these solutions were robust, relatively simple models that generalized well into the future. Or maybe it played out that way by chance. I think it's hard to take any clear lesson from it, but it's interesting and I'm curious if anyone else has a clear take on this.",
    "1922190": "Hi @yukio0201, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1923068": "Hi Yang.\nI I've filled out your survey."
  },
  "source": "meta"
}