{
  "id": 477228,
  "title": "Conclusions from my experiments so far",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/477228",
  "author_name": "",
  "post_date": "2024-02-15T08:25:56.258346100Z",
  "votes": 38,
  "comment_count": 21,
  "views": 0,
  "content": "<h3>This competition promotes extreme levels of blending/ensembling</h3>\n<p>As we can see, the best public notebooks use blends of blends of blends (and even those are further blends - because they are e.g. Auto ML models consisting of NNs, GBDTs, and Random Forest already).<br>\nBlending is so far the only strategy that guarantees an increase in Public LB score.</p>\n<p><strong>This could be reflecting the fact that blending indeed increases the stability of predictions, and the current metric promotes it a lot.</strong></p>\n<h3>There seems to be a lot of difficulty in establishing a stable CV - LB framework</h3>\n<p>I have not found one so far. The results are all over the place. Because of pre-covid train data and post-covid LB data I think it will be difficult.\\<br>\nUPDATE: The reason also seems to be that hosts downsampled some data sources for the test period (see discussion below).</p>\n<h3>Metric hacking is x10 more important than Feature engineering</h3>\n<p>This was discussed extensively already and is a bit sad, and I think not something hosts are looking for. </p>\n<p>For perspective:</p>\n<ul>\n<li>Best <a href=\"https://www.kaggle.com/code/stechparme/home-credit-baseline-max-min-features\" target=\"_blank\">public single model</a>: 0.559 LB</li>\n<li>Best <a href=\"https://www.kaggle.com/code/andreynesterov/home-credit-baseline-inference\" target=\"_blank\">public blend</a>: 0.559 LB -&gt; 0.575 LB (+0.017)</li>\n<li><a href=\"https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference\" target=\"_blank\">Metric hacking</a> 0.575 -&gt; 0.620 LB (+0.045)</li>\n<li>My work on features engineering 0.559 LB -&gt; 0.564 LB (+0.005), but most of them failed to increase Public LB, while all of them increased CV scores a lot</li>\n</ul>\n<p>This is a bit disheartening that currently metic hacking brings SO MUCH more improvement than other strategies. It incentivizes folks to focus their effort on metric hacking. That said, hosts made an announcement they would fix it, which is great. Let's see what is the impact of the coming changes.</p>\n<p>Update 1:<br>\nAs discussed <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360\" target=\"_blank\">here</a>, it seems Credit Bureau A Data is not available in the future.</p>",
  "messages": [
    {
      "id": "2653128",
      "postDate": "02/15/2024 08:25:56",
      "content": "<h3>This competition promotes extreme levels of blending/ensembling</h3>\n<p>As we can see, the best public notebooks use blends of blends of blends (and even those are further blends - because they are e.g. Auto ML models consisting of NNs, GBDTs, and Random Forest already).<br>\nBlending is so far the only strategy that guarantees an increase in Public LB score.</p>\n<p><strong>This could be reflecting the fact that blending indeed increases the stability of predictions, and the current metric promotes it a lot.</strong></p>\n<h3>There seems to be a lot of difficulty in establishing a stable CV - LB framework</h3>\n<p>I have not found one so far. The results are all over the place. Because of pre-covid train data and post-covid LB data I think it will be difficult.\\<br>\nUPDATE: The reason also seems to be that hosts downsampled some data sources for the test period (see discussion below).</p>\n<h3>Metric hacking is x10 more important than Feature engineering</h3>\n<p>This was discussed extensively already and is a bit sad, and I think not something hosts are looking for. </p>\n<p>For perspective:</p>\n<ul>\n<li>Best <a href=\"https://www.kaggle.com/code/stechparme/home-credit-baseline-max-min-features\" target=\"_blank\">public single model</a>: 0.559 LB</li>\n<li>Best <a href=\"https://www.kaggle.com/code/andreynesterov/home-credit-baseline-inference\" target=\"_blank\">public blend</a>: 0.559 LB -&gt; 0.575 LB (+0.017)</li>\n<li><a href=\"https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference\" target=\"_blank\">Metric hacking</a> 0.575 -&gt; 0.620 LB (+0.045)</li>\n<li>My work on features engineering 0.559 LB -&gt; 0.564 LB (+0.005), but most of them failed to increase Public LB, while all of them increased CV scores a lot</li>\n</ul>\n<p>This is a bit disheartening that currently metic hacking brings SO MUCH more improvement than other strategies. It incentivizes folks to focus their effort on metric hacking. That said, hosts made an announcement they would fix it, which is great. Let's see what is the impact of the coming changes.</p>\n<p>Update 1:<br>\nAs discussed <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360\" target=\"_blank\">here</a>, it seems Credit Bureau A Data is not available in the future.</p>",
      "rawMarkdown": "### This competition promotes extreme levels of blending/ensembling\n\nAs we can see, the best public notebooks use blends of blends of blends (and even those are further blends - because they are e.g. Auto ML models consisting of NNs, GBDTs, and Random Forest already).\nBlending is so far the only strategy that guarantees an increase in Public LB score.\n\n**This could be reflecting the fact that blending indeed increases the stability of predictions, and the current metric promotes it a lot.**\n\n### There seems to be a lot of difficulty in establishing a stable CV - LB framework\n\nI have not found one so far. The results are all over the place. Because of pre-covid train data and post-covid LB data I think it will be difficult.\\\nUPDATE: The reason also seems to be that hosts downsampled some data sources for the test period (see discussion below).\n\n### Metric hacking is x10 more important than Feature engineering\n\nThis was discussed extensively already and is a bit sad, and I think not something hosts are looking for. \n\nFor perspective:\n- Best [public single model](https://www.kaggle.com/code/stechparme/home-credit-baseline-max-min-features): 0.559 LB\n- Best [public blend](https://www.kaggle.com/code/andreynesterov/home-credit-baseline-inference): 0.559 LB -> 0.575 LB (+0.017)\n- [Metric hacking](https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference) 0.575 -> 0.620 LB (+0.045)\n- My work on features engineering 0.559 LB -> 0.564 LB (+0.005), but most of them failed to increase Public LB, while all of them increased CV scores a lot\n\nThis is a bit disheartening that currently metic hacking brings SO MUCH more improvement than other strategies. It incentivizes folks to focus their effort on metric hacking. That said, hosts made an announcement they would fix it, which is great. Let's see what is the impact of the coming changes.\n\nUpdate 1:\nAs discussed [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360), it seems Credit Bureau A Data is not available in the future.",
      "votes": null
    },
    {
      "id": "2653144",
      "postDate": "02/15/2024 08:35:02",
      "content": "<blockquote>\n  <p>Metric hacking 0.575 -&gt; 0.620 LB (+0.045)</p>\n</blockquote>\n<p>You can get even more, it just makes no sense to spend additional time on that, because it has nothing to do with modeling. I guess it is worth to focus on ROC AUC improvements through CV and not even submit to LB until the metric is fixed.</p>",
      "rawMarkdown": ">Metric hacking 0.575 -> 0.620 LB (+0.045)\n\nYou can get even more, it just makes no sense to spend additional time on that, because it has nothing to do with modeling. I guess it is worth to focus on ROC AUC improvements through CV and not even submit to LB until the metric is fixed.",
      "votes": null
    },
    {
      "id": "2653145",
      "postDate": "02/15/2024 08:36:40",
      "content": "<p>True.</p>\n<p>At the same time we all love to see our name high on LB :D </p>",
      "rawMarkdown": "True.\n\nAt the same time we all love to see our name high on LB :D",
      "votes": null
    },
    {
      "id": "2653261",
      "postDate": "02/15/2024 10:25:30",
      "content": "<p>i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.</p>\n<p>this doesnt take long and there is nothing more to gain after that, i, for example just have five actual submissions, one of my model and 4 more to roughly get the problematic term to 0, the rest are just \"notebook threw exception\" due to memory errors.</p>",
      "rawMarkdown": "i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.\n\nthis doesnt take long and there is nothing more to gain after that, i, for example just have five actual submissions, one of my model and 4 more to roughly get the problematic term to 0, the rest are just \"notebook threw exception\" due to memory errors.",
      "votes": null
    },
    {
      "id": "2653275",
      "postDate": "02/15/2024 10:32:02",
      "content": "<blockquote>\n  <p>i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.</p>\n</blockquote>\n<p>Yes, but how to make it 0 and not hurt your avg gini at the same time is the key question i guess in metric hacking.</p>\n<blockquote>\n  <p>this doesnt take long and there is nothing more to gain after that,</p>\n</blockquote>\n<p>Hopefully yes. Let's see what the changes will bring.</p>",
      "rawMarkdown": "> i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.\n\nYes, but how to make it 0 and not hurt your avg gini at the same time is the key question i guess in metric hacking.\n\n> this doesnt take long and there is nothing more to gain after that,\n\nHopefully yes. Let's see what the changes will bring.",
      "votes": null
    },
    {
      "id": "2653339",
      "postDate": "02/15/2024 11:10:51",
      "content": "<p>well i think no matter how you do it it will always hurt the other term roughly the same… at least in my (albeit) very limited testing i have not found any meaningful difference no matter how i do it</p>\n<p>so i would wager it would make at most a difference of maybe a very negligible 0.001 between a very bad way and an optimal way, which i guess could be a very meaningful amount at the end of the competition judging from past competitions like this</p>",
      "rawMarkdown": "well i think no matter how you do it it will always hurt the other term roughly the same... at least in my (albeit) very limited testing i have not found any meaningful difference no matter how i do it\n\nso i would wager it would make at most a difference of maybe a very negligible 0.001 between a very bad way and an optimal way, which i guess could be a very meaningful amount at the end of the competition judging from past competitions like this",
      "votes": null
    },
    {
      "id": "2653366",
      "postDate": "02/15/2024 11:48:27",
      "content": "<p>probably there is even something mathematical here, where you can proof that it makes exactly 0 difference how you get this term to 0, at least that would explain the results i see in my experiments</p>",
      "rawMarkdown": "probably there is even something mathematical here, where you can proof that it makes exactly 0 difference how you get this term to 0, at least that would explain the results i see in my experiments",
      "votes": null
    },
    {
      "id": "2653905",
      "postDate": "02/15/2024 17:50:27",
      "content": "<p>Interesting. Myself, I have resisted the temptation so far to play with the metric, so I can't really testify to that. I only see ppl high on LB with names like \"Metric Hacking is all you Need\"</p>",
      "rawMarkdown": "Interesting. Myself, I have resisted the temptation so far to play with the metric, so I can't really testify to that. I only see ppl high on LB with names like \"Metric Hacking is all you Need\"",
      "votes": null
    },
    {
      "id": "2653917",
      "postDate": "02/15/2024 18:07:08",
      "content": "<p>yeah, they must mostly be using some form of the public notebooks and they seem to get the term to 0 at a shift of around -0.025 with the method i am using (where i am making the model worse for the first half of all the weeks)</p>\n<p>i get the term to 0 at a shift of around -0.035</p>\n<p>i am not using anything of the public notebooks (except for how to format the submission_csv) and my shift being higher just means that my model falls off harder but has an overall higher mean(gini)</p>\n<p>which also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks</p>",
      "rawMarkdown": "yeah, they must mostly be using some form of the public notebooks and they seem to get the term to 0 at a shift of around -0.025 with the method i am using (where i am making the model worse for the first half of all the weeks)\n\ni get the term to 0 at a shift of around -0.035\n\ni am not using anything of the public notebooks (except for how to format the submission_csv) and my shift being higher just means that my model falls off harder but has an overall higher mean(gini)\n\nwhich also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks",
      "votes": null
    },
    {
      "id": "2653990",
      "postDate": "02/15/2024 18:55:25",
      "content": "<p>but yeah, you can do it anyway you like… for example, just make the predictions really bad for the first 5 weeks or use any other kind of  transformation to make your predictions worse at the start</p>\n<p>from what i can tell, in the end, the score seems to be the same for the same model when you get the term to 0 no matter how you do it</p>",
      "rawMarkdown": "but yeah, you can do it anyway you like... for example, just make the predictions really bad for the first 5 weeks or use any other kind of  transformation to make your predictions worse at the start\n\nfrom what i can tell, in the end, the score seems to be the same for the same model when you get the term to 0 no matter how you do it",
      "votes": null
    },
    {
      "id": "2655281",
      "postDate": "02/16/2024 19:13:49",
      "content": "<p>Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.</p>\n<p>What does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? If that is the case, asking for score stability is not a very reasonable request given that input data is unstable.</p>\n<p>Or it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?</p>\n<p>Either way looks like current problem formulation is asking for something impossible or at least contradictory.</p>",
      "rawMarkdown": "Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.\n\nWhat does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? If that is the case, asking for score stability is not a very reasonable request given that input data is unstable.\n\nOr it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?\n\nEither way looks like current problem formulation is asking for something impossible or at least contradictory.",
      "votes": null
    },
    {
      "id": "2655298",
      "postDate": "02/16/2024 19:25:37",
      "content": "<p>It may be a combination of both issues. Normally, the behaviour of our clients changes over time. </p>",
      "rawMarkdown": "It may be a combination of both issues. Normally, the behaviour of our clients changes over time.",
      "votes": null
    },
    {
      "id": "2655329",
      "postDate": "02/16/2024 19:45:24",
      "content": "<blockquote>\n  <p>Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.</p>\n</blockquote>\n<p>I have the same observation.</p>",
      "rawMarkdown": "> Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.\n\nI have the same observation.",
      "votes": null
    },
    {
      "id": "2655334",
      "postDate": "02/16/2024 19:49:10",
      "content": "<p>As for </p>\n<blockquote>\n  <p>What does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?</p>\n<p>Regarding:</p>\n<blockquote>\n  <p>Or it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?</p>\n</blockquote>\n<p>-&gt; This I understand is the key challenge the hosts want us to solve.</p>",
      "rawMarkdown": "As for \n> What does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? \n\n@jetakow Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?\n\nRegarding:\n> Or it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?\n\n-> This I understand is the key challenge the hosts want us to solve.",
      "votes": null
    },
    {
      "id": "2655373",
      "postDate": "02/16/2024 20:08:37",
      "content": "<blockquote>\n  <p>Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?</p>\n</blockquote>\n<p>I am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day. </p>\n<blockquote>\n  <p>This I understand is the key challenge the hosts want us to solve.</p>\n</blockquote>\n<p>Exactly, that is why the metric is done in such a fashion and that is the reason why we don't want to disclose anything further about the test set. </p>",
      "rawMarkdown": ">Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?\n\nI am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day. \n\n>This I understand is the key challenge the hosts want us to solve.\n\nExactly, that is why the metric is done in such a fashion and that is the reason why we don't want to disclose anything further about the test set.",
      "votes": null
    },
    {
      "id": "2655481",
      "postDate": "02/16/2024 22:59:02",
      "content": "<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a><br>\nMaybe I'm missing something, but I don't understand how subtracting an offset from all predictions for a week worsens the Gini for that week. It must be due to the truncation to zero, and what worsens the metric is that cases with low risk all collapse to zero. Is that correct?\"</p>",
      "rawMarkdown": "at7459\nMaybe I'm missing something, but I don't understand how subtracting an offset from all predictions for a week worsens the Gini for that week. It must be due to the truncation to zero, and what worsens the metric is that cases with low risk all collapse to zero. Is that correct?\"",
      "votes": null
    },
    {
      "id": "2655488",
      "postDate": "02/16/2024 23:14:43",
      "content": "<p><a href=\"https://www.kaggle.com/blindape\" target=\"_blank\">@blindape</a> you are correct, check out this comment: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2648946\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2648946</a></p>",
      "rawMarkdown": "blindape you are correct, check out this comment: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2648946",
      "votes": null
    },
    {
      "id": "2655739",
      "postDate": "02/17/2024 07:20:07",
      "content": "<blockquote>\n  <p>I am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day.</p>\n</blockquote>\n<p>Well, I am not asking you about the split details, etc. I am asking why you downsampled heavily the <code>applprev</code> data for example during the test period. It is equivalent to asking: \"Why have you filled this very strong feature with Null values during the test period?\". I think it is fair to be curious about it and still cannot fathom a reason why you did so.</p>",
      "rawMarkdown": "> I am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day.\n\nWell, I am not asking you about the split details, etc. I am asking why you downsampled heavily the `applprev` data for example during the test period. It is equivalent to asking: \"Why have you filled this very strong feature with Null values during the test period?\". I think it is fair to be curious about it and still cannot fathom a reason why you did so.",
      "votes": null
    },
    {
      "id": "2655909",
      "postDate": "02/17/2024 09:22:16",
      "content": "<blockquote>\n  <p>which also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks</p>\n</blockquote>\n<p>what is the CV framework you are using? for me, I also have CV scores round 0.02 higher than public notebooks but it leads to lower public LB scores</p>",
      "rawMarkdown": ">which also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks\n\nwhat is the CV framework you are using? for me, I also have CV scores round 0.02 higher than public notebooks but it leads to lower public LB scores",
      "votes": null
    },
    {
      "id": "2655925",
      "postDate": "02/17/2024 09:53:30",
      "content": "<p>sorry, what i wrote there is misleading and makes no sense, i dont know what i was thinking there, actually, i dont know what the score of the public notebooks is for my split</p>\n<p>i was using WEEK_NUM&lt;60 for training and &gt;=60 for validation<br>\nit's around 0.863 for the roc_auc and around 0.7 for stability</p>\n<p>if youre seeing some inconsistencies you might have some unintended drift somewhere, but i dont know whether the cv/lb has  a good correlation or not because i only have 1 model submitted so far, so take what i say with a grain of salt</p>",
      "rawMarkdown": "sorry, what i wrote there is misleading and makes no sense, i dont know what i was thinking there, actually, i dont know what the score of the public notebooks is for my split\n\ni was using WEEK_NUM<60 for training and >=60 for validation\nit's around 0.863 for the roc_auc and around 0.7 for stability\n\nif youre seeing some inconsistencies you might have some unintended drift somewhere, but i dont know whether the cv/lb has  a good correlation or not because i only have 1 model submitted so far, so take what i say with a grain of salt",
      "votes": null
    },
    {
      "id": "2655941",
      "postDate": "02/17/2024 10:14:27",
      "content": "<p>Thanks for sharing. 0.7 without any postprocessing (hacking), right?</p>",
      "rawMarkdown": "Thanks for sharing. 0.7 without any postprocessing (hacking), right?",
      "votes": null
    },
    {
      "id": "2655946",
      "postDate": "02/17/2024 10:17:12",
      "content": "<p>yes, from what i could tell the term with the 88 weight is always 0 during training anyway.</p>",
      "rawMarkdown": "yes, from what i could tell the term with the 88 weight is always 0 during training anyway.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2653144,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "02/15/2024 08:35:02",
      "content": "<blockquote>\n  <p>Metric hacking 0.575 -&gt; 0.620 LB (+0.045)</p>\n</blockquote>\n<p>You can get even more, it just makes no sense to spend additional time on that, because it has nothing to do with modeling. I guess it is worth to focus on ROC AUC improvements through CV and not even submit to LB until the metric is fixed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2653145,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/15/2024 08:36:40",
          "content": "<p>True.</p>\n<p>At the same time we all love to see our name high on LB :D </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2653261,
      "author_name": "at7459",
      "author_url": "",
      "post_date": "02/15/2024 10:25:30",
      "content": "<p>i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.</p>\n<p>this doesnt take long and there is nothing more to gain after that, i, for example just have five actual submissions, one of my model and 4 more to roughly get the problematic term to 0, the rest are just \"notebook threw exception\" due to memory errors.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2653275,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/15/2024 10:32:02",
          "content": "<blockquote>\n  <p>i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.</p>\n</blockquote>\n<p>Yes, but how to make it 0 and not hurt your avg gini at the same time is the key question i guess in metric hacking.</p>\n<blockquote>\n  <p>this doesnt take long and there is nothing more to gain after that,</p>\n</blockquote>\n<p>Hopefully yes. Let's see what the changes will bring.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2653339,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "02/15/2024 11:10:51",
              "content": "<p>well i think no matter how you do it it will always hurt the other term roughly the same… at least in my (albeit) very limited testing i have not found any meaningful difference no matter how i do it</p>\n<p>so i would wager it would make at most a difference of maybe a very negligible 0.001 between a very bad way and an optimal way, which i guess could be a very meaningful amount at the end of the competition judging from past competitions like this</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2653366,
                  "author_name": "at7459",
                  "author_url": "",
                  "post_date": "02/15/2024 11:48:27",
                  "content": "<p>probably there is even something mathematical here, where you can proof that it makes exactly 0 difference how you get this term to 0, at least that would explain the results i see in my experiments</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2653905,
                      "author_name": "narsil",
                      "author_url": "",
                      "post_date": "02/15/2024 17:50:27",
                      "content": "<p>Interesting. Myself, I have resisted the temptation so far to play with the metric, so I can't really testify to that. I only see ppl high on LB with names like \"Metric Hacking is all you Need\"</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2653917,
                          "author_name": "at7459",
                          "author_url": "",
                          "post_date": "02/15/2024 18:07:08",
                          "content": "<p>yeah, they must mostly be using some form of the public notebooks and they seem to get the term to 0 at a shift of around -0.025 with the method i am using (where i am making the model worse for the first half of all the weeks)</p>\n<p>i get the term to 0 at a shift of around -0.035</p>\n<p>i am not using anything of the public notebooks (except for how to format the submission_csv) and my shift being higher just means that my model falls off harder but has an overall higher mean(gini)</p>\n<p>which also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2655909,
                              "author_name": "narsil",
                              "author_url": "",
                              "post_date": "02/17/2024 09:22:16",
                              "content": "<blockquote>\n  <p>which also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks</p>\n</blockquote>\n<p>what is the CV framework you are using? for me, I also have CV scores round 0.02 higher than public notebooks but it leads to lower public LB scores</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2655925,
                                  "author_name": "at7459",
                                  "author_url": "",
                                  "post_date": "02/17/2024 09:53:30",
                                  "content": "<p>sorry, what i wrote there is misleading and makes no sense, i dont know what i was thinking there, actually, i dont know what the score of the public notebooks is for my split</p>\n<p>i was using WEEK_NUM&lt;60 for training and &gt;=60 for validation<br>\nit's around 0.863 for the roc_auc and around 0.7 for stability</p>\n<p>if youre seeing some inconsistencies you might have some unintended drift somewhere, but i dont know whether the cv/lb has  a good correlation or not because i only have 1 model submitted so far, so take what i say with a grain of salt</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2655941,
                                      "author_name": "blindape",
                                      "author_url": "",
                                      "post_date": "02/17/2024 10:14:27",
                                      "content": "<p>Thanks for sharing. 0.7 without any postprocessing (hacking), right?</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2655946,
                                          "author_name": "at7459",
                                          "author_url": "",
                                          "post_date": "02/17/2024 10:17:12",
                                          "content": "<p>yes, from what i could tell the term with the 88 weight is always 0 during training anyway.</p>",
                                          "votes": null,
                                          "replies": []
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        },
                        {
                          "id": 2653990,
                          "author_name": "at7459",
                          "author_url": "",
                          "post_date": "02/15/2024 18:55:25",
                          "content": "<p>but yeah, you can do it anyway you like… for example, just make the predictions really bad for the first 5 weeks or use any other kind of  transformation to make your predictions worse at the start</p>\n<p>from what i can tell, in the end, the score seems to be the same for the same model when you get the term to 0 no matter how you do it</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2655481,
                              "author_name": "blindape",
                              "author_url": "",
                              "post_date": "02/16/2024 22:59:02",
                              "content": "<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a><br>\nMaybe I'm missing something, but I don't understand how subtracting an offset from all predictions for a week worsens the Gini for that week. It must be due to the truncation to zero, and what worsens the metric is that cases with low risk all collapse to zero. Is that correct?\"</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2655488,
                                  "author_name": "kononenko",
                                  "author_url": "",
                                  "post_date": "02/16/2024 23:14:43",
                                  "content": "<p><a href=\"https://www.kaggle.com/blindape\" target=\"_blank\">@blindape</a> you are correct, check out this comment: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2648946\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2648946</a></p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2655281,
      "author_name": "ymatioun",
      "author_url": "",
      "post_date": "02/16/2024 19:13:49",
      "content": "<p>Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.</p>\n<p>What does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? If that is the case, asking for score stability is not a very reasonable request given that input data is unstable.</p>\n<p>Or it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?</p>\n<p>Either way looks like current problem formulation is asking for something impossible or at least contradictory.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2655298,
          "author_name": "jetakow",
          "author_url": "",
          "post_date": "02/16/2024 19:25:37",
          "content": "<p>It may be a combination of both issues. Normally, the behaviour of our clients changes over time. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2655334,
              "author_name": "narsil",
              "author_url": "",
              "post_date": "02/16/2024 19:49:10",
              "content": "<p>As for </p>\n<blockquote>\n  <p>What does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?</p>\n<p>Regarding:</p>\n<blockquote>\n  <p>Or it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?</p>\n</blockquote>\n<p>-&gt; This I understand is the key challenge the hosts want us to solve.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2655373,
                  "author_name": "jetakow",
                  "author_url": "",
                  "post_date": "02/16/2024 20:08:37",
                  "content": "<blockquote>\n  <p>Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?</p>\n</blockquote>\n<p>I am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day. </p>\n<blockquote>\n  <p>This I understand is the key challenge the hosts want us to solve.</p>\n</blockquote>\n<p>Exactly, that is why the metric is done in such a fashion and that is the reason why we don't want to disclose anything further about the test set. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2655739,
                      "author_name": "narsil",
                      "author_url": "",
                      "post_date": "02/17/2024 07:20:07",
                      "content": "<blockquote>\n  <p>I am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day.</p>\n</blockquote>\n<p>Well, I am not asking you about the split details, etc. I am asking why you downsampled heavily the <code>applprev</code> data for example during the test period. It is equivalent to asking: \"Why have you filled this very strong feature with Null values during the test period?\". I think it is fair to be curious about it and still cannot fathom a reason why you did so.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2655329,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/16/2024 19:45:24",
          "content": "<blockquote>\n  <p>Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.</p>\n</blockquote>\n<p>I have the same observation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2653128": "### This competition promotes extreme levels of blending/ensembling\n\nAs we can see, the best public notebooks use blends of blends of blends (and even those are further blends - because they are e.g. Auto ML models consisting of NNs, GBDTs, and Random Forest already).\nBlending is so far the only strategy that guarantees an increase in Public LB score.\n\n**This could be reflecting the fact that blending indeed increases the stability of predictions, and the current metric promotes it a lot.**\n\n### There seems to be a lot of difficulty in establishing a stable CV - LB framework\n\nI have not found one so far. The results are all over the place. Because of pre-covid train data and post-covid LB data I think it will be difficult.\\\nUPDATE: The reason also seems to be that hosts downsampled some data sources for the test period (see discussion below).\n\n### Metric hacking is x10 more important than Feature engineering\n\nThis was discussed extensively already and is a bit sad, and I think not something hosts are looking for. \n\nFor perspective:\n- Best [public single model](https://www.kaggle.com/code/stechparme/home-credit-baseline-max-min-features): 0.559 LB\n- Best [public blend](https://www.kaggle.com/code/andreynesterov/home-credit-baseline-inference): 0.559 LB -> 0.575 LB (+0.017)\n- [Metric hacking](https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference) 0.575 -> 0.620 LB (+0.045)\n- My work on features engineering 0.559 LB -> 0.564 LB (+0.005), but most of them failed to increase Public LB, while all of them increased CV scores a lot\n\nThis is a bit disheartening that currently metic hacking brings SO MUCH more improvement than other strategies. It incentivizes folks to focus their effort on metric hacking. That said, hosts made an announcement they would fix it, which is great. Let's see what is the impact of the coming changes.\n\nUpdate 1:\nAs discussed [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360), it seems Credit Bureau A Data is not available in the future.",
    "2653144": ">Metric hacking 0.575 -> 0.620 LB (+0.045)\n\nYou can get even more, it just makes no sense to spend additional time on that, because it has nothing to do with modeling. I guess it is worth to focus on ROC AUC improvements through CV and not even submit to LB until the metric is fixed.",
    "2653145": "True.\n\nAt the same time we all love to see our name high on LB :D",
    "2653261": "i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.\n\nthis doesnt take long and there is nothing more to gain after that, i, for example just have five actual submissions, one of my model and 4 more to roughly get the problematic term to 0, the rest are just \"notebook threw exception\" due to memory errors.",
    "2653275": "> i dont think you need to spend that much time on metric hacking, due to the high weight of the term you just need to make it so that this term becomes approximately 0.\n\nYes, but how to make it 0 and not hurt your avg gini at the same time is the key question i guess in metric hacking.\n\n> this doesnt take long and there is nothing more to gain after that,\n\nHopefully yes. Let's see what the changes will bring.",
    "2653339": "well i think no matter how you do it it will always hurt the other term roughly the same... at least in my (albeit) very limited testing i have not found any meaningful difference no matter how i do it\n\nso i would wager it would make at most a difference of maybe a very negligible 0.001 between a very bad way and an optimal way, which i guess could be a very meaningful amount at the end of the competition judging from past competitions like this",
    "2653366": "probably there is even something mathematical here, where you can proof that it makes exactly 0 difference how you get this term to 0, at least that would explain the results i see in my experiments",
    "2653905": "Interesting. Myself, I have resisted the temptation so far to play with the metric, so I can't really testify to that. I only see ppl high on LB with names like \"Metric Hacking is all you Need\"",
    "2653917": "yeah, they must mostly be using some form of the public notebooks and they seem to get the term to 0 at a shift of around -0.025 with the method i am using (where i am making the model worse for the first half of all the weeks)\n\ni get the term to 0 at a shift of around -0.035\n\ni am not using anything of the public notebooks (except for how to format the submission_csv) and my shift being higher just means that my model falls off harder but has an overall higher mean(gini)\n\nwhich also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks",
    "2653990": "but yeah, you can do it anyway you like... for example, just make the predictions really bad for the first 5 weeks or use any other kind of  transformation to make your predictions worse at the start\n\nfrom what i can tell, in the end, the score seems to be the same for the same model when you get the term to 0 no matter how you do it",
    "2655281": "Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.\n\nWhat does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? If that is the case, asking for score stability is not a very reasonable request given that input data is unstable.\n\nOr it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?\n\nEither way looks like current problem formulation is asking for something impossible or at least contradictory.",
    "2655298": "It may be a combination of both issues. Normally, the behaviour of our clients changes over time.",
    "2655329": "> Exactly. There is large CV/LB difference using only depth=0 data, and it gets even larger when also using depth=1 data.\n\nI have the same observation.",
    "2655334": "As for \n> What does that mean? Depth=1 data is not always available in the prediction period(as stated by the holds)? \n\n@jetakow Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?\n\nRegarding:\n> Or it means that pattern of test data is materially different from train data - meaning that train data is unsuitable for building predictive models?\n\n-> This I understand is the key challenge the hosts want us to solve.",
    "2655373": ">Did you decrease the amount of depth 1 and 2 data on purpose for the test period? What was the purpose?\n\nI am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day. \n\n>This I understand is the key challenge the hosts want us to solve.\n\nExactly, that is why the metric is done in such a fashion and that is the reason why we don't want to disclose anything further about the test set.",
    "2655481": "at7459\nMaybe I'm missing something, but I don't understand how subtracting an offset from all predictions for a week worsens the Gini for that week. It must be due to the truncation to zero, and what worsens the metric is that cases with low risk all collapse to zero. Is that correct?\"",
    "2655488": "blindape you are correct, check out this comment: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2648946",
    "2655739": "> I am sure you understand we won't disclose any further information about test set. I get questions about test set almost every day.\n\nWell, I am not asking you about the split details, etc. I am asking why you downsampled heavily the `applprev` data for example during the test period. It is equivalent to asking: \"Why have you filled this very strong feature with Null values during the test period?\". I think it is fair to be curious about it and still cannot fathom a reason why you did so.",
    "2655909": ">which also makes sense because my roc_auc cv is around 0.02 higher than the public notebooks\n\nwhat is the CV framework you are using? for me, I also have CV scores round 0.02 higher than public notebooks but it leads to lower public LB scores",
    "2655925": "sorry, what i wrote there is misleading and makes no sense, i dont know what i was thinking there, actually, i dont know what the score of the public notebooks is for my split\n\ni was using WEEK_NUM<60 for training and >=60 for validation\nit's around 0.863 for the roc_auc and around 0.7 for stability\n\nif youre seeing some inconsistencies you might have some unintended drift somewhere, but i dont know whether the cv/lb has  a good correlation or not because i only have 1 model submitted so far, so take what i say with a grain of salt",
    "2655941": "Thanks for sharing. 0.7 without any postprocessing (hacking), right?",
    "2655946": "yes, from what i could tell the term with the 88 weight is always 0 during training anyway."
  },
  "source": "meta"
}