{
  "id": 347652,
  "title": "Did you make the right model choice?",
  "url": "/competitions/amex-default-prediction/discussion/347652",
  "author_name": "",
  "post_date": "2022-08-25T01:18:26.034807Z",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I see that my best model was actually submitted 2 1/2 weeks ago and would have placed #13. Do you have similar experiences / do you have a previous model which you did not select that would have worked better? </p>",
  "messages": [
    {
      "id": "1912803",
      "postDate": "08/25/2022 01:18:26",
      "content": "<p>I see that my best model was actually submitted 2 1/2 weeks ago and would have placed #13. Do you have similar experiences / do you have a previous model which you did not select that would have worked better? </p>",
      "rawMarkdown": "I see that my best model was actually submitted 2 1/2 weeks ago and would have placed #13. Do you have similar experiences / do you have a previous model which you did not select that would have worked better?",
      "votes": null
    },
    {
      "id": "1912810",
      "postDate": "08/25/2022 01:24:08",
      "content": "<p>Yes and no.</p>\n<p>I almost submitted an ensemble of my two best models, and an ensemble of those plus ragnar's LGBM dart. Switched to my single best (and more unusual) model which by itself got my best score, jumping to 49th.</p>\n<p>HOWEVER, it turns out that my blend of two models, which all the feedback I got was that it was a 'better' 0.798, was a 0.79890, vs the single model was a 0.79889. It was only 0.0001 better even on the public LB. If I had known that, I would've ensembled <em>only</em> that model with ragnar's. I haven't even submitted that one yet since my best model was only ready on the last day, heh. I'm curious if that would've gotten me any higher! Guess it's time to check… :)</p>",
      "rawMarkdown": "Yes and no.\n\nI almost submitted an ensemble of my two best models, and an ensemble of those plus ragnar's LGBM dart. Switched to my single best (and more unusual) model which by itself got my best score, jumping to 49th.\n\nHOWEVER, it turns out that my blend of two models, which all the feedback I got was that it was a 'better' 0.798, was a 0.79890, vs the single model was a 0.79889. It was only 0.0001 better even on the public LB. If I had known that, I would've ensembled *only* that model with ragnar's. I haven't even submitted that one yet since my best model was only ready on the last day, heh. I'm curious if that would've gotten me any higher! Guess it's time to check... :)",
      "votes": null
    },
    {
      "id": "1912811",
      "postDate": "08/25/2022 01:24:43",
      "content": "<p>Same here. We have quite a few 808 submissions and the highest one will get us either gold or the first place in silver (0.80838). However, even in hindsight I can’t convince myself to pick any of the 808 subs. None of them really stands out in any way. </p>",
      "rawMarkdown": "Same here. We have quite a few 808 submissions and the highest one will get us either gold or the first place in silver (0.80838). However, even in hindsight I can’t convince myself to pick any of the 808 subs. None of them really stands out in any way.",
      "votes": null
    },
    {
      "id": "1912815",
      "postDate": "08/25/2022 01:29:25",
      "content": "<p>That's right. Also, the competition metric seems a bit more volatile than log-loss so it doesn't really mean much, this is probably just within random fluctuation.</p>",
      "rawMarkdown": "That's right. Also, the competition metric seems a bit more volatile than log-loss so it doesn't really mean much, this is probably just within random fluctuation.",
      "votes": null
    },
    {
      "id": "1912816",
      "postDate": "08/25/2022 01:29:51",
      "content": "<p>Our best model was submitted about 22 days ago which would have landed us at around 200 (would have went from bronze to silver) on the private leaderboard. Kind of a bummer, haha but nevertheless I'm super stoked to read through all the different approaches this weekend.</p>",
      "rawMarkdown": "Our best model was submitted about 22 days ago which would have landed us at around 200 (would have went from bronze to silver) on the private leaderboard. Kind of a bummer, haha but nevertheless I'm super stoked to read through all the different approaches this weekend.",
      "votes": null
    },
    {
      "id": "1912817",
      "postDate": "08/25/2022 01:31:02",
      "content": "<p>Yes. I did not select one submission that would place me around 60th position. But it is what it is. </p>",
      "rawMarkdown": "Yes. I did not select one submission that would place me around 60th position. But it is what it is.",
      "votes": null
    },
    {
      "id": "1912831",
      "postDate": "08/25/2022 01:45:22",
      "content": "<p>Similar experience here. We had 6 subs that would have scored gold, with no evidence that any of them would have been better than the ones we chose -- all lower on both CV and LB.</p>",
      "rawMarkdown": "Similar experience here. We had 6 subs that would have scored gold, with no evidence that any of them would have been better than the ones we chose -- all lower on both CV and LB.",
      "votes": null
    },
    {
      "id": "1912836",
      "postDate": "08/25/2022 01:52:25",
      "content": "<p>Of my top 10 scoring models on private LB, 6 were made in the last 48 hours. I selected a model that was third best on private LB, but didn't miss by much. Its score was 0.80729, while the best model had 0.80732.</p>\n<p>I had two mega-blends, one with 14 models, and the other with 18. The first group included models that had LB-CV difference larger than 0.001. Something like <code>[0.794, 0.792889]</code> or <code>[0.795, 0.793405]</code>. That ranked blend/stack of models ended up being the best, yet I didn't select it (had overall CV of 0.7982). Then I had 18 models that had LB-CV difference smaller than 0.001 - something like <code>[0.798, 0,797541]</code> and <code>[0.797, 0.796619]</code>. This is what I ultimately selected (had CV of 0.7989), and it ended up being my third best model. In the latter category there wasn't a single model above 0.798 in private LB, yet it still ensembled to  0.79983.</p>",
      "rawMarkdown": "Of my top 10 scoring models on private LB, 6 were made in the last 48 hours. I selected a model that was third best on private LB, but didn't miss by much. Its score was 0.80729, while the best model had 0.80732.\n\nI had two mega-blends, one with 14 models, and the other with 18. The first group included models that had LB-CV difference larger than 0.001. Something like `[0.794, 0.792889]` or `[0.795, 0.793405]`. That ranked blend/stack of models ended up being the best, yet I didn't select it (had overall CV of 0.7982). Then I had 18 models that had LB-CV difference smaller than 0.001 - something like `[0.798, 0,797541]` and `[0.797, 0.796619]`. This is what I ultimately selected (had CV of 0.7989), and it ended up being my third best model. In the latter category there wasn't a single model above 0.798 in private LB, yet it still ensembled to  0.79983.",
      "votes": null
    },
    {
      "id": "1912841",
      "postDate": "08/25/2022 01:55:06",
      "content": "<p>The models I choose were .80704 and .080698, which were my 4th and 1st best models on the public LB, unfortunately I had ten, count em ten models that would've scored higher on the private LB, but scored lower than the ones I choose on the public one. The difference would've been being top 8% vs. top 15%. I had a hunch to pick the best one, but, it was an ensemble with some weaker models, had a poorer LB score, nothing about the better models suggested to me that they'd do better on the private LB…. except, some of them had higher scoring models CV wise. </p>\n<p>Lesson learned though: I have to study the test data more and think more strategically about picking my final models, not picking the best isn't surprising, but picking 11th and 12th stings, a lot. </p>",
      "rawMarkdown": "The models I choose were .80704 and .080698, which were my 4th and 1st best models on the public LB, unfortunately I had ten, count em ten models that would've scored higher on the private LB, but scored lower than the ones I choose on the public one. The difference would've been being top 8% vs. top 15%. I had a hunch to pick the best one, but, it was an ensemble with some weaker models, had a poorer LB score, nothing about the better models suggested to me that they'd do better on the private LB.... except, some of them had higher scoring models CV wise. \n\nLesson learned though: I have to study the test data more and think more strategically about picking my final models, not picking the best isn't surprising, but picking 11th and 12th stings, a lot.",
      "votes": null
    },
    {
      "id": "1913048",
      "postDate": "08/25/2022 05:31:43",
      "content": "<p>I made the right choice :) my final submission, best on CV, best on public LB, best on private LB. sadly - not good enough</p>",
      "rawMarkdown": "I made the right choice :) my final submission, best on CV, best on public LB, best on private LB. sadly - not good enough",
      "votes": null
    },
    {
      "id": "1913140",
      "postDate": "08/25/2022 07:05:53",
      "content": "<p>Yes! The blended model that I sent 8 hours before the deadline was the best that would have made me to 78th place but it was in the 14th place of my best public score. I had no means for an educated guess that it would be the best on the private LB :)</p>",
      "rawMarkdown": "Yes! The blended model that I sent 8 hours before the deadline was the best that would have made me to 78th place but it was in the 14th place of my best public score. I had no means for an educated guess that it would be the best on the private LB :)",
      "votes": null
    },
    {
      "id": "1913366",
      "postDate": "08/25/2022 09:49:51",
      "content": "<p>I was convinced my main edge would come from post processing to adapt for (covid) drift. I looked for external data and tried to de-anonymise categorical features. My final subs include some manual shift in prediction. They both perform worse than the unshifted preds.</p>",
      "rawMarkdown": "I was convinced my main edge would come from post processing to adapt for (covid) drift. I looked for external data and tried to de-anonymise categorical features. My final subs include some manual shift in prediction. They both perform worse than the unshifted preds.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1912810,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/25/2022 01:24:08",
      "content": "<p>Yes and no.</p>\n<p>I almost submitted an ensemble of my two best models, and an ensemble of those plus ragnar's LGBM dart. Switched to my single best (and more unusual) model which by itself got my best score, jumping to 49th.</p>\n<p>HOWEVER, it turns out that my blend of two models, which all the feedback I got was that it was a 'better' 0.798, was a 0.79890, vs the single model was a 0.79889. It was only 0.0001 better even on the public LB. If I had known that, I would've ensembled <em>only</em> that model with ragnar's. I haven't even submitted that one yet since my best model was only ready on the last day, heh. I'm curious if that would've gotten me any higher! Guess it's time to check… :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912811,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "08/25/2022 01:24:43",
      "content": "<p>Same here. We have quite a few 808 submissions and the highest one will get us either gold or the first place in silver (0.80838). However, even in hindsight I can’t convince myself to pick any of the 808 subs. None of them really stands out in any way. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1912815,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/25/2022 01:29:25",
          "content": "<p>That's right. Also, the competition metric seems a bit more volatile than log-loss so it doesn't really mean much, this is probably just within random fluctuation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912816,
      "author_name": "ol0fmeister",
      "author_url": "",
      "post_date": "08/25/2022 01:29:51",
      "content": "<p>Our best model was submitted about 22 days ago which would have landed us at around 200 (would have went from bronze to silver) on the private leaderboard. Kind of a bummer, haha but nevertheless I'm super stoked to read through all the different approaches this weekend.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912817,
      "author_name": "gogogopp",
      "author_url": "",
      "post_date": "08/25/2022 01:31:02",
      "content": "<p>Yes. I did not select one submission that would place me around 60th position. But it is what it is. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912831,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "08/25/2022 01:45:22",
      "content": "<p>Similar experience here. We had 6 subs that would have scored gold, with no evidence that any of them would have been better than the ones we chose -- all lower on both CV and LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912836,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "08/25/2022 01:52:25",
      "content": "<p>Of my top 10 scoring models on private LB, 6 were made in the last 48 hours. I selected a model that was third best on private LB, but didn't miss by much. Its score was 0.80729, while the best model had 0.80732.</p>\n<p>I had two mega-blends, one with 14 models, and the other with 18. The first group included models that had LB-CV difference larger than 0.001. Something like <code>[0.794, 0.792889]</code> or <code>[0.795, 0.793405]</code>. That ranked blend/stack of models ended up being the best, yet I didn't select it (had overall CV of 0.7982). Then I had 18 models that had LB-CV difference smaller than 0.001 - something like <code>[0.798, 0,797541]</code> and <code>[0.797, 0.796619]</code>. This is what I ultimately selected (had CV of 0.7989), and it ended up being my third best model. In the latter category there wasn't a single model above 0.798 in private LB, yet it still ensembled to  0.79983.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1912841,
      "author_name": "markhamlee",
      "author_url": "",
      "post_date": "08/25/2022 01:55:06",
      "content": "<p>The models I choose were .80704 and .080698, which were my 4th and 1st best models on the public LB, unfortunately I had ten, count em ten models that would've scored higher on the private LB, but scored lower than the ones I choose on the public one. The difference would've been being top 8% vs. top 15%. I had a hunch to pick the best one, but, it was an ensemble with some weaker models, had a poorer LB score, nothing about the better models suggested to me that they'd do better on the private LB…. except, some of them had higher scoring models CV wise. </p>\n<p>Lesson learned though: I have to study the test data more and think more strategically about picking my final models, not picking the best isn't surprising, but picking 11th and 12th stings, a lot. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913048,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "08/25/2022 05:31:43",
      "content": "<p>I made the right choice :) my final submission, best on CV, best on public LB, best on private LB. sadly - not good enough</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913140,
      "author_name": "rasoulmojtahedzadeh",
      "author_url": "",
      "post_date": "08/25/2022 07:05:53",
      "content": "<p>Yes! The blended model that I sent 8 hours before the deadline was the best that would have made me to 78th place but it was in the 14th place of my best public score. I had no means for an educated guess that it would be the best on the private LB :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913366,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "08/25/2022 09:49:51",
      "content": "<p>I was convinced my main edge would come from post processing to adapt for (covid) drift. I looked for external data and tried to de-anonymise categorical features. My final subs include some manual shift in prediction. They both perform worse than the unshifted preds.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912803": "I see that my best model was actually submitted 2 1/2 weeks ago and would have placed #13. Do you have similar experiences / do you have a previous model which you did not select that would have worked better?",
    "1912810": "Yes and no.\n\nI almost submitted an ensemble of my two best models, and an ensemble of those plus ragnar's LGBM dart. Switched to my single best (and more unusual) model which by itself got my best score, jumping to 49th.\n\nHOWEVER, it turns out that my blend of two models, which all the feedback I got was that it was a 'better' 0.798, was a 0.79890, vs the single model was a 0.79889. It was only 0.0001 better even on the public LB. If I had known that, I would've ensembled *only* that model with ragnar's. I haven't even submitted that one yet since my best model was only ready on the last day, heh. I'm curious if that would've gotten me any higher! Guess it's time to check... :)",
    "1912811": "Same here. We have quite a few 808 submissions and the highest one will get us either gold or the first place in silver (0.80838). However, even in hindsight I can’t convince myself to pick any of the 808 subs. None of them really stands out in any way.",
    "1912815": "That's right. Also, the competition metric seems a bit more volatile than log-loss so it doesn't really mean much, this is probably just within random fluctuation.",
    "1912816": "Our best model was submitted about 22 days ago which would have landed us at around 200 (would have went from bronze to silver) on the private leaderboard. Kind of a bummer, haha but nevertheless I'm super stoked to read through all the different approaches this weekend.",
    "1912817": "Yes. I did not select one submission that would place me around 60th position. But it is what it is.",
    "1912831": "Similar experience here. We had 6 subs that would have scored gold, with no evidence that any of them would have been better than the ones we chose -- all lower on both CV and LB.",
    "1912836": "Of my top 10 scoring models on private LB, 6 were made in the last 48 hours. I selected a model that was third best on private LB, but didn't miss by much. Its score was 0.80729, while the best model had 0.80732.\n\nI had two mega-blends, one with 14 models, and the other with 18. The first group included models that had LB-CV difference larger than 0.001. Something like `[0.794, 0.792889]` or `[0.795, 0.793405]`. That ranked blend/stack of models ended up being the best, yet I didn't select it (had overall CV of 0.7982). Then I had 18 models that had LB-CV difference smaller than 0.001 - something like `[0.798, 0,797541]` and `[0.797, 0.796619]`. This is what I ultimately selected (had CV of 0.7989), and it ended up being my third best model. In the latter category there wasn't a single model above 0.798 in private LB, yet it still ensembled to  0.79983.",
    "1912841": "The models I choose were .80704 and .080698, which were my 4th and 1st best models on the public LB, unfortunately I had ten, count em ten models that would've scored higher on the private LB, but scored lower than the ones I choose on the public one. The difference would've been being top 8% vs. top 15%. I had a hunch to pick the best one, but, it was an ensemble with some weaker models, had a poorer LB score, nothing about the better models suggested to me that they'd do better on the private LB.... except, some of them had higher scoring models CV wise. \n\nLesson learned though: I have to study the test data more and think more strategically about picking my final models, not picking the best isn't surprising, but picking 11th and 12th stings, a lot.",
    "1913048": "I made the right choice :) my final submission, best on CV, best on public LB, best on private LB. sadly - not good enough",
    "1913140": "Yes! The blended model that I sent 8 hours before the deadline was the best that would have made me to 78th place but it was in the 14th place of my best public score. I had no means for an educated guess that it would be the best on the private LB :)",
    "1913366": "I was convinced my main edge would come from post processing to adapt for (covid) drift. I looked for external data and tried to de-anonymise categorical features. My final subs include some manual shift in prediction. They both perform worse than the unshifted preds."
  },
  "source": "meta"
}