{
  "id": 501172,
  "title": "Another way to hack the metric",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501172",
  "author_name": "Ern711",
  "post_date": "2024-05-08T10:44:25.570000",
  "votes": 39,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Since this has turned into a hacking competition, here's another strategy:</p>\n<ul>\n<li>Train a good model (like AUC of 0.85)</li>\n<li>Train a slightly worse (or perhaps much worse is better) model (like a regressor trained on the training probabilities of the first model)</li>\n<li>Add an extra column target_new in training data and give label 0 to all obs</li>\n<li>Load in the test set and add target_new column to this as well but with label 1</li>\n<li>Concatenate the datasets and train a model to learn target_new, and let it predict on the data corresponding to the test set</li>\n</ul>\n<p>If the probability is high enough, use the better model, if the probability is lower (more similar to training data) use the worse model - assuming observations further into the future are more dissimilar.</p>\n<p>I tested with a very simple (and far from optimized) version of this, and it still increased the leaderboard with like 0.013. Doing something like this but for a bunch of models and probability thresholds, one could probably achieve much higher gains…</p>\n<p>There you have one way! </p>",
  "messages": [
    {
      "id": 2800753,
      "postDate": "2024-05-08T10:44:25.570Z",
      "content": "<p>Since this has turned into a hacking competition, here's another strategy:</p>\n<ul>\n<li>Train a good model (like AUC of 0.85)</li>\n<li>Train a slightly worse (or perhaps much worse is better) model (like a regressor trained on the training probabilities of the first model)</li>\n<li>Add an extra column target_new in training data and give label 0 to all obs</li>\n<li>Load in the test set and add target_new column to this as well but with label 1</li>\n<li>Concatenate the datasets and train a model to learn target_new, and let it predict on the data corresponding to the test set</li>\n</ul>\n<p>If the probability is high enough, use the better model, if the probability is lower (more similar to training data) use the worse model - assuming observations further into the future are more dissimilar.</p>\n<p>I tested with a very simple (and far from optimized) version of this, and it still increased the leaderboard with like 0.013. Doing something like this but for a bunch of models and probability thresholds, one could probably achieve much higher gains…</p>\n<p>There you have one way! </p>",
      "rawMarkdown": "Since this has turned into a hacking competition, here's another strategy:\n\n* Train a good model (like AUC of 0.85)\n* Train a slightly worse (or perhaps much worse is better) model (like a regressor trained on the training probabilities of the first model)\n* Add an extra column target_new in training data and give label 0 to all obs\n* Load in the test set and add target_new column to this as well but with label 1\n* Concatenate the datasets and train a model to learn target_new, and let it predict on the data corresponding to the test set\n\nIf the probability is high enough, use the better model, if the probability is lower (more similar to training data) use the worse model - assuming observations further into the future are more dissimilar.\n\nI tested with a very simple (and far from optimized) version of this, and it still increased the leaderboard with like 0.013. Doing something like this but for a bunch of models and probability thresholds, one could probably achieve much higher gains...\n\nThere you have one way! \n\n",
      "votes": 39
    },
    {
      "id": 2801707,
      "postDate": "2024-05-08T17:47:26.837Z",
      "content": "<p>Thanks for sharing, the ways to hack this metric are unlimited :)</p>\n<p>I cant wait to learn more of these so I can use the learnings to my real job :)</p>",
      "rawMarkdown": "Thanks for sharing, the ways to hack this metric are unlimited :)\n\nI cant wait to learn more of these so I can use the learnings to my real job :)",
      "votes": 3,
      "replies": [
        {
          "id": 2801824,
          "postDate": "2024-05-08T18:50:19.353Z",
          "content": "<p>Lol. </p>\n<p>As I mentioned before, I think a simpler and better metric would be to evaluate AUC on a test set where all data points are at least 1 year (or a similar duration) after the latest training data point. That would measure which model generalizes best X months / years into the future, which kind of feels like what they are looking for anyway. But I did not construct this metric, so I have no idea! </p>\n<p>Even though I understand it would be weird to change the metric now, and that it would also be super hard to go through all solutions and try to determine if hacking was used, it's a bit disappointing that it turned out like this. Since the original idea behind the metric was quite interesting.</p>",
          "rawMarkdown": "Lol. \n\nAs I mentioned before, I think a simpler and better metric would be to evaluate AUC on a test set where all data points are at least 1 year (or a similar duration) after the latest training data point. That would measure which model generalizes best X months / years into the future, which kind of feels like what they are looking for anyway. But I did not construct this metric, so I have no idea! \n\nEven though I understand it would be weird to change the metric now, and that it would also be super hard to go through all solutions and try to determine if hacking was used, it's a bit disappointing that it turned out like this. Since the original idea behind the metric was quite interesting.",
          "votes": 2,
          "replies": [
            {
              "id": 2801924,
              "postDate": "2024-05-08T19:44:41.900Z",
              "content": "<p>I think they could have simply kept AUC score as the metric and then assessed the top 100 solutions separately for stability and awarded a separate prize for the model stability akin to the efficiency leaderboard prize in other competitions <a href=\"https://www.kaggle.com/ern711\" target=\"_blank\">@ern711</a> </p>\n<p>Sometimes overthinking and over-engineering is also dangerous!</p>",
              "rawMarkdown": "I think they could have simply kept AUC score as the metric and then assessed the top 100 solutions separately for stability and awarded a separate prize for the model stability akin to the efficiency leaderboard prize in other competitions @ern711 \n\nSometimes overthinking and over-engineering is also dangerous!",
              "votes": 4
            },
            {
              "id": 2802230,
              "postDate": "2024-05-08T23:34:35.243Z",
              "content": "<p>Given how it turned out I would prefer that over the current metric for sure :) </p>",
              "rawMarkdown": "Given how it turned out I would prefer that over the current metric for sure :) ",
              "votes": 1
            }
          ]
        },
        {
          "id": 2802279,
          "postDate": "2024-05-09T00:52:43.997Z",
          "content": "<p>Ha-ha  ^_^ </p>",
          "rawMarkdown": "Ha-ha  ^_^ ",
          "votes": 1
        },
        {
          "id": 2803444,
          "postDate": "2024-05-09T14:07:40.333Z",
          "content": "<p><a href=\"https://www.kaggle.com/carloshuertas\" target=\"_blank\">@carloshuertas</a> I wonder how one could use *metric hacking in a job environment. I also work in the same domain and this information will be useful for me too!</p>",
          "rawMarkdown": "@carloshuertas I wonder how one could use *metric hacking in a job environment. I also work in the same domain and this information will be useful for me too!",
          "replies": [
            {
              "id": 2803541,
              "postDate": "2024-05-09T14:49:48.357Z",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> On the next salary review: </p>\n<p>\"But look how reliable the model is, performance has even increased after 6 months\" 😅😂</p>",
              "rawMarkdown": "@ravi20076 On the next salary review: \n\n\"But look how reliable the model is, performance has even increased after 6 months\" 😅😂",
              "votes": 3
            },
            {
              "id": 2803677,
              "postDate": "2024-05-09T15:36:15.097Z",
              "content": "<p>Apologies, I was being 101% sarcastic, this is completely nonsense, I think the organizers just came up with a random metric they think is good. It's not. Their problem is very interesting and real, but they rushed their metric, that's why science is usually peer reviewed, they got reviewed and failed here, yet here we are.</p>",
              "rawMarkdown": "Apologies, I was being 101% sarcastic, this is completely nonsense, I think the organizers just came up with a random metric they think is good. It's not. Their problem is very interesting and real, but they rushed their metric, that's why science is usually peer reviewed, they got reviewed and failed here, yet here we are.",
              "votes": 7
            },
            {
              "id": 2803834,
              "postDate": "2024-05-09T16:36:27.933Z",
              "content": "<p>The metric is not good even if hacking was impossible since it will never tell you which model performs best over a certain time period (which should really be what is relevant). Instead it more or less tells you which model deteriorates the least, which is something completely different. </p>\n<p>The assumption that the model which deteriorates the least in the first x months will continuing being more robust in the upcoming y months is probably incorrect (at least in many cases) and hence not very useful.</p>\n<p>If having a model where the deterioration over time is predictable is the goal, maybe one could try something like:</p>\n<ol>\n<li><p>Train a model for AUC the normal way (or to be good on some specific time period).</p></li>\n<li><p>Label training data 0 and newer production data 1, and train a model on this (as suggested for the hacking above) to detect data drift.</p></li>\n<li><p>Use the the drift defined by the probability in 2. as a feature to train a third model to predict the mae/rmse of the first models predictions over time. </p></li>\n</ol>\n<p>Here one could maybe divide probabilities into bins/intervals when computing the mae/rmse. That should perhaps give some indication of uncertainty over time. But likely such an estimate would be quite unreliable in itself, so it might not even work and just introduces lots of extra work/validation processes of the additional models and might hence turn out completely useless in the end (but it seems like a fun experiment) 😅</p>",
              "rawMarkdown": "The metric is not good even if hacking was impossible since it will never tell you which model performs best over a certain time period (which should really be what is relevant). Instead it more or less tells you which model deteriorates the least, which is something completely different. \n\nThe assumption that the model which deteriorates the least in the first x months will continuing being more robust in the upcoming y months is probably incorrect (at least in many cases) and hence not very useful.\n\nIf having a model where the deterioration over time is predictable is the goal, maybe one could try something like:\n\n1. Train a model for AUC the normal way (or to be good on some specific time period).\n\n2. Label training data 0 and newer production data 1, and train a model on this (as suggested for the hacking above) to detect data drift.\n\n3. Use the the drift defined by the probability in 2. as a feature to train a third model to predict the mae/rmse of the first models predictions over time. \n\nHere one could maybe divide probabilities into bins/intervals when computing the mae/rmse. That should perhaps give some indication of uncertainty over time. But likely such an estimate would be quite unreliable in itself, so it might not even work and just introduces lots of extra work/validation processes of the additional models and might hence turn out completely useless in the end (but it seems like a fun experiment) 😅\n",
              "votes": 2
            },
            {
              "id": 2805359,
              "postDate": "2024-05-10T14:11:04.747Z",
              "content": "<p><a href=\"https://www.kaggle.com/carloshuertas\" target=\"_blank\">@carloshuertas</a> <br>\nI simply played along your joke and sarcasm! </p>",
              "rawMarkdown": "@carloshuertas \nI simply played along your joke and sarcasm! "
            },
            {
              "id": 2806068,
              "postDate": "2024-05-10T21:03:52.167Z",
              "content": "<p>I have been wondering for a while if there was a way to incorporate data drift as a feature into the model and train it so that it can stay stable for a longer period of time. But, given the fact that the test data is post-covid while the train data is re-covid, I wonder if it would help at all?</p>",
              "rawMarkdown": "I have been wondering for a while if there was a way to incorporate data drift as a feature into the model and train it so that it can stay stable for a longer period of time. But, given the fact that the test data is post-covid while the train data is re-covid, I wonder if it would help at all?",
              "votes": 1
            },
            {
              "id": 2806073,
              "postDate": "2024-05-10T21:11:29.947Z",
              "content": "<p>I did not think of including it in the original model, but rather have another model predicting how reliable the predictions of the first model are - with drift as the main feature. Though I also wonder if it will work at all (even without something like covid)</p>",
              "rawMarkdown": "I did not think of including it in the original model, but rather have another model predicting how reliable the predictions of the first model are - with drift as the main feature. Though I also wonder if it will work at all (even without something like covid)"
            }
          ]
        }
      ]
    },
    {
      "id": 2804612,
      "postDate": "2024-05-10T06:00:31.553Z",
      "content": "<p>It seems that it is possible to get a improvement of 0.04 from restoring WEEK_NUM. That's an incredible advantage lol. Though I have no idea how to restore WEEK_NUM.</p>",
      "rawMarkdown": "It seems that it is possible to get a improvement of 0.04 from restoring WEEK_NUM. That's an incredible advantage lol. Though I have no idea how to restore WEEK_NUM.",
      "votes": 1,
      "replies": [
        {
          "id": 2805437,
          "postDate": "2024-05-10T14:49:53.180Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2814228,
          "postDate": "2024-05-15T08:15:38.383Z",
          "content": "<p>seems like now you know :)</p>",
          "rawMarkdown": "seems like now you know :)"
        }
      ]
    },
    {
      "id": 2815260,
      "postDate": "2024-05-15T19:06:28.453Z",
      "content": "<p>Has 'date_decision' only been re-formatted in new transformation? If so can't we just do</p>\n<pre><code> dateutil  parser\n\n ():\n    parsed_date = parser.parse(date_str)\n     parsed_date.strftime()\n</code></pre>\n<p>Or if they shifted it a lot there is other columns like lastapplicationdate_877D…</p>\n<p>Ultimately it just sounds like you need to approximate the time series ordering so other columns may be good enough if high enough correlation to original WEEK_NUM.</p>",
      "rawMarkdown": "Has 'date_decision' only been re-formatted in new transformation? If so can't we just do\n\n```python\nfrom dateutil import parser\n\ndef reformat_date(date_str):\n    parsed_date = parser.parse(date_str)\n    return parsed_date.strftime('%Y-%m-%d')\n```\n\nOr if they shifted it a lot there is other columns like lastapplicationdate_877D...\n\nUltimately it just sounds like you need to approximate the time series ordering so other columns may be good enough if high enough correlation to original WEEK_NUM."
    },
    {
      "id": 2814052,
      "postDate": "2024-05-15T06:19:21.423Z",
      "content": "<blockquote>\n  <ul>\n  <li>Concatenate the datasets </li>\n  </ul>\n</blockquote>\n<p>when load both train and test, there are out of memory issues and how to solve it? maybe sample or use less features?</p>",
      "rawMarkdown": "> - Concatenate the datasets \n\nwhen load both train and test, there are out of memory issues and how to solve it? maybe sample or use less features?",
      "replies": [
        {
          "id": 2814478,
          "postDate": "2024-05-15T10:40:59.430Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501170\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501170</a></p>",
          "rawMarkdown": "https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501170\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2808252,
      "postDate": "2024-05-12T05:32:20.763Z",
      "content": "<p>Hello, I'd like to ask if you've reset datedecision to its initial value now? Or are you using other methods to improve the LB score?</p>",
      "rawMarkdown": "\nHello, I'd like to ask if you've reset datedecision to its initial value now? Or are you using other methods to improve the LB score?",
      "replies": [
        {
          "id": 2808973,
          "postDate": "2024-05-12T13:28:28.643Z",
          "content": "<p>Without saying exactly what I'm doing, there are multiple ways of exploiting vulnerabilities in this metric. Some things might work better than others of course.</p>\n<p>If will carry over to the private lb, or if the host will get what they want in the end can be discussed, to put it mildly 😅</p>",
          "rawMarkdown": "Without saying exactly what I'm doing, there are multiple ways of exploiting vulnerabilities in this metric. Some things might work better than others of course.\n\nIf will carry over to the private lb, or if the host will get what they want in the end can be discussed, to put it mildly 😅"
        }
      ]
    },
    {
      "id": 2805416,
      "postDate": "2024-05-10T14:36:58.353Z",
      "content": "<p>This strategy proposes a method to manipulate metrics by incorporating a secondary model to predict when to use a better or worse model based on data dissimilarity. While it may temporarily boost leaderboard rankings, it risks compromising the integrity of the evaluation process and lacks ethical considerations regarding fair competition and genuine performance improvement.</p>",
      "rawMarkdown": "This strategy proposes a method to manipulate metrics by incorporating a secondary model to predict when to use a better or worse model based on data dissimilarity. While it may temporarily boost leaderboard rankings, it risks compromising the integrity of the evaluation process and lacks ethical considerations regarding fair competition and genuine performance improvement.",
      "replies": [
        {
          "id": 2806019,
          "postDate": "2024-05-10T20:20:00.233Z",
          "content": "<p><a href=\"https://www.kaggle.com/harold107\" target=\"_blank\">@harold107</a> So many people did warn the host multiple times regarding the flaws of this metric, and still, they decided to keep it.</p>\n<p>At first, they did, however, say that they would not allow \"hacking\" in terms of detecting early cases and reducing their performance on purpose. Unfortunately, that changed a couple of days ago when they clearly stated that it's not against the rules and therefore is allowed.</p>\n<p>All the top notebooks on the public LB are already doing something similar to this anyway, so I don't see the problem. Sure, the host does not get what they wanted in the end, but it's in this case self-inflicted. I understand that they had good intentions when deciding to use this metric (and I also understand the issues with changing the metric in the middle of the competition etc), but a bad metric is a bad metric, and that's reality :) </p>",
          "rawMarkdown": "@harold107 So many people did warn the host multiple times regarding the flaws of this metric, and still, they decided to keep it.\n\nAt first, they did, however, say that they would not allow \"hacking\" in terms of detecting early cases and reducing their performance on purpose. Unfortunately, that changed a couple of days ago when they clearly stated that it's not against the rules and therefore is allowed.\n\nAll the top notebooks on the public LB are already doing something similar to this anyway, so I don't see the problem. Sure, the host does not get what they wanted in the end, but it's in this case self-inflicted. I understand that they had good intentions when deciding to use this metric (and I also understand the issues with changing the metric in the middle of the competition etc), but a bad metric is a bad metric, and that's reality :) \n"
        }
      ]
    },
    {
      "id": 2803348,
      "postDate": "2024-05-09T12:55:08.370Z",
      "content": "<p>How do you define 'high enough probability' and 'lower probability'? What threshold do you set?</p>",
      "rawMarkdown": "How do you define 'high enough probability' and 'lower probability'? What threshold do you set?",
      "replies": [
        {
          "id": 2803375,
          "postDate": "2024-05-09T13:23:07.963Z",
          "content": "<p>I guess that (as well as what models to use) is something to experiment with. I only tried one version with a probability threshold of 0.5 where the \"bad\" model was just slightly worse. </p>\n<p>In addition, the thresholds should depend on the parameters of the classifier used (since you are predicting training probabilities so advanced models will overfit more and vice verse)</p>\n<p>Currently I have no GPU left, but when I get it back I plan to experiment a bit with this to get an idea on how much of an increase you could get by using this method. Though I imagine there might be lots of other (and likely better) ways to hack this metric as well :) </p>",
          "rawMarkdown": "I guess that (as well as what models to use) is something to experiment with. I only tried one version with a probability threshold of 0.5 where the \"bad\" model was just slightly worse. \n\nIn addition, the thresholds should depend on the parameters of the classifier used (since you are predicting training probabilities so advanced models will overfit more and vice verse)\n\nCurrently I have no GPU left, but when I get it back I plan to experiment a bit with this to get an idea on how much of an increase you could get by using this method. Though I imagine there might be lots of other (and likely better) ways to hack this metric as well :) ",
          "votes": 1,
          "replies": [
            {
              "id": 2803443,
              "postDate": "2024-05-09T14:07:31.143Z",
              "content": "<p>thanks！！First, I trained a good LGB model. <br>\nThen, I used the probability predictions of the good model as labels to train the bad model. <br>\nAfter that, I trained a model to learn target_new. I also set 0.5 as the threshold. However, the result decreased by around 0.02. Could you please advise where I might have gone wrong?</p>",
              "rawMarkdown": "thanks！！First, I trained a good LGB model. \nThen, I used the probability predictions of the good model as labels to train the bad model. \nAfter that, I trained a model to learn target_new. I also set 0.5 as the threshold. However, the result decreased by around 0.02. Could you please advise where I might have gone wrong?"
            },
            {
              "id": 2803531,
              "postDate": "2024-05-09T14:44:45.170Z",
              "content": "<p>I think the easiest is just to experiment with different thresholds and models. And using the probabilities of the first model as target for the second model is not really necessary, it was just something and tried out  for fun. The important part should be that the second model performs worse (and how much worse is not really clear)</p>\n<p>Test a bunch of thresholds, and a bunch of models? Also I'm not sure how well this strategy will work in general, but since I got an increase by 0.013 quite easily I guess it will be relatively easy to get a much higher increase by picking suitable thresholds, and degrade performance enough to more or less remove the negative slope. But who knows 🤷‍♂️</p>\n<p>Or perhaps look at some strategy directly aimed at figuring out the original week number! </p>",
              "rawMarkdown": "I think the easiest is just to experiment with different thresholds and models. And using the probabilities of the first model as target for the second model is not really necessary, it was just something and tried out  for fun. The important part should be that the second model performs worse (and how much worse is not really clear)\n\nTest a bunch of thresholds, and a bunch of models? Also I'm not sure how well this strategy will work in general, but since I got an increase by 0.013 quite easily I guess it will be relatively easy to get a much higher increase by picking suitable thresholds, and degrade performance enough to more or less remove the negative slope. But who knows 🤷‍♂️\n\nOr perhaps look at some strategy directly aimed at figuring out the original week number! ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2810678,
      "postDate": "2024-05-13T11:52:50.313Z",
      "content": "<p>Thanks for sharing! Interesting.</p>",
      "rawMarkdown": "Thanks for sharing! Interesting.",
      "votes": 1
    },
    {
      "id": 2807268,
      "postDate": "2024-05-11T15:34:21.373Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 2805777,
      "postDate": "2024-05-10T18:25:49.823Z",
      "content": "<p>Good work, thanks</p>",
      "rawMarkdown": "Good work, thanks",
      "votes": 1
    },
    {
      "id": 2804751,
      "postDate": "2024-05-10T07:27:51.113Z",
      "content": "<p>nice idea! thanks</p>",
      "rawMarkdown": "nice idea! thanks",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2801707,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2024-05-08T17:47:26.837000",
      "content": "<p>Thanks for sharing, the ways to hack this metric are unlimited :)</p>\n<p>I cant wait to learn more of these so I can use the learnings to my real job :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2801824,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-05-08T18:50:19.353000",
          "content": "<p>Lol. </p>\n<p>As I mentioned before, I think a simpler and better metric would be to evaluate AUC on a test set where all data points are at least 1 year (or a similar duration) after the latest training data point. That would measure which model generalizes best X months / years into the future, which kind of feels like what they are looking for anyway. But I did not construct this metric, so I have no idea! </p>\n<p>Even though I understand it would be weird to change the metric now, and that it would also be super hard to go through all solutions and try to determine if hacking was used, it's a bit disappointing that it turned out like this. Since the original idea behind the metric was quite interesting.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2801924,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-05-08T19:44:41.900000",
              "content": "<p>I think they could have simply kept AUC score as the metric and then assessed the top 100 solutions separately for stability and awarded a separate prize for the model stability akin to the efficiency leaderboard prize in other competitions <a href=\"https://www.kaggle.com/ern711\" target=\"_blank\">@ern711</a> </p>\n<p>Sometimes overthinking and over-engineering is also dangerous!</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2802230,
              "author_name": "Ern711",
              "author_url": "",
              "post_date": "2024-05-08T23:34:35.243000",
              "content": "<p>Given how it turned out I would prefer that over the current metric for sure :) </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2802279,
          "author_name": "Bruce",
          "author_url": "",
          "post_date": "2024-05-09T00:52:43.997000",
          "content": "<p>Ha-ha  ^_^ </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2803444,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-05-09T14:07:40.333000",
          "content": "<p><a href=\"https://www.kaggle.com/carloshuertas\" target=\"_blank\">@carloshuertas</a> I wonder how one could use *metric hacking in a job environment. I also work in the same domain and this information will be useful for me too!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2803541,
              "author_name": "Ern711",
              "author_url": "",
              "post_date": "2024-05-09T14:49:48.357000",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> On the next salary review: </p>\n<p>\"But look how reliable the model is, performance has even increased after 6 months\" 😅😂</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2803677,
              "author_name": "NxGTR",
              "author_url": "",
              "post_date": "2024-05-09T15:36:15.097000",
              "content": "<p>Apologies, I was being 101% sarcastic, this is completely nonsense, I think the organizers just came up with a random metric they think is good. It's not. Their problem is very interesting and real, but they rushed their metric, that's why science is usually peer reviewed, they got reviewed and failed here, yet here we are.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2803834,
              "author_name": "Ern711",
              "author_url": "",
              "post_date": "2024-05-09T16:36:27.933000",
              "content": "<p>The metric is not good even if hacking was impossible since it will never tell you which model performs best over a certain time period (which should really be what is relevant). Instead it more or less tells you which model deteriorates the least, which is something completely different. </p>\n<p>The assumption that the model which deteriorates the least in the first x months will continuing being more robust in the upcoming y months is probably incorrect (at least in many cases) and hence not very useful.</p>\n<p>If having a model where the deterioration over time is predictable is the goal, maybe one could try something like:</p>\n<ol>\n<li><p>Train a model for AUC the normal way (or to be good on some specific time period).</p></li>\n<li><p>Label training data 0 and newer production data 1, and train a model on this (as suggested for the hacking above) to detect data drift.</p></li>\n<li><p>Use the the drift defined by the probability in 2. as a feature to train a third model to predict the mae/rmse of the first models predictions over time. </p></li>\n</ol>\n<p>Here one could maybe divide probabilities into bins/intervals when computing the mae/rmse. That should perhaps give some indication of uncertainty over time. But likely such an estimate would be quite unreliable in itself, so it might not even work and just introduces lots of extra work/validation processes of the additional models and might hence turn out completely useless in the end (but it seems like a fun experiment) 😅</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2805359,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-05-10T14:11:04.747000",
              "content": "<p><a href=\"https://www.kaggle.com/carloshuertas\" target=\"_blank\">@carloshuertas</a> <br>\nI simply played along your joke and sarcasm! </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2806068,
              "author_name": "Varuni Rao",
              "author_url": "",
              "post_date": "2024-05-10T21:03:52.167000",
              "content": "<p>I have been wondering for a while if there was a way to incorporate data drift as a feature into the model and train it so that it can stay stable for a longer period of time. But, given the fact that the test data is post-covid while the train data is re-covid, I wonder if it would help at all?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2806073,
              "author_name": "Ern711",
              "author_url": "",
              "post_date": "2024-05-10T21:11:29.947000",
              "content": "<p>I did not think of including it in the original model, but rather have another model predicting how reliable the predictions of the first model are - with drift as the main feature. Though I also wonder if it will work at all (even without something like covid)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2804612,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-05-10T06:00:31.553000",
      "content": "<p>It seems that it is possible to get a improvement of 0.04 from restoring WEEK_NUM. That's an incredible advantage lol. Though I have no idea how to restore WEEK_NUM.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2805437,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-10T14:49:53.180000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2814228,
          "author_name": "Andrey Chankin",
          "author_url": "",
          "post_date": "2024-05-15T08:15:38.383000",
          "content": "<p>seems like now you know :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2815260,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2024-05-15T19:06:28.453000",
      "content": "<p>Has 'date_decision' only been re-formatted in new transformation? If so can't we just do</p>\n<pre><code> dateutil  parser\n\n ():\n    parsed_date = parser.parse(date_str)\n     parsed_date.strftime()\n</code></pre>\n<p>Or if they shifted it a lot there is other columns like lastapplicationdate_877D…</p>\n<p>Ultimately it just sounds like you need to approximate the time series ordering so other columns may be good enough if high enough correlation to original WEEK_NUM.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2814052,
      "author_name": "Yayue Rao",
      "author_url": "",
      "post_date": "2024-05-15T06:19:21.423000",
      "content": "<blockquote>\n  <ul>\n  <li>Concatenate the datasets </li>\n  </ul>\n</blockquote>\n<p>when load both train and test, there are out of memory issues and how to solve it? maybe sample or use less features?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2814478,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-05-15T10:40:59.430000",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501170\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501170</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2808252,
      "author_name": "Jiaxiu Zou",
      "author_url": "",
      "post_date": "2024-05-12T05:32:20.763000",
      "content": "<p>Hello, I'd like to ask if you've reset datedecision to its initial value now? Or are you using other methods to improve the LB score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2808973,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-05-12T13:28:28.643000",
          "content": "<p>Without saying exactly what I'm doing, there are multiple ways of exploiting vulnerabilities in this metric. Some things might work better than others of course.</p>\n<p>If will carry over to the private lb, or if the host will get what they want in the end can be discussed, to put it mildly 😅</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2805416,
      "author_name": "harold107",
      "author_url": "",
      "post_date": "2024-05-10T14:36:58.353000",
      "content": "<p>This strategy proposes a method to manipulate metrics by incorporating a secondary model to predict when to use a better or worse model based on data dissimilarity. While it may temporarily boost leaderboard rankings, it risks compromising the integrity of the evaluation process and lacks ethical considerations regarding fair competition and genuine performance improvement.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2806019,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-05-10T20:20:00.233000",
          "content": "<p><a href=\"https://www.kaggle.com/harold107\" target=\"_blank\">@harold107</a> So many people did warn the host multiple times regarding the flaws of this metric, and still, they decided to keep it.</p>\n<p>At first, they did, however, say that they would not allow \"hacking\" in terms of detecting early cases and reducing their performance on purpose. Unfortunately, that changed a couple of days ago when they clearly stated that it's not against the rules and therefore is allowed.</p>\n<p>All the top notebooks on the public LB are already doing something similar to this anyway, so I don't see the problem. Sure, the host does not get what they wanted in the end, but it's in this case self-inflicted. I understand that they had good intentions when deciding to use this metric (and I also understand the issues with changing the metric in the middle of the competition etc), but a bad metric is a bad metric, and that's reality :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2803348,
      "author_name": "Jiaxiu Zou",
      "author_url": "",
      "post_date": "2024-05-09T12:55:08.370000",
      "content": "<p>How do you define 'high enough probability' and 'lower probability'? What threshold do you set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2803375,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-05-09T13:23:07.963000",
          "content": "<p>I guess that (as well as what models to use) is something to experiment with. I only tried one version with a probability threshold of 0.5 where the \"bad\" model was just slightly worse. </p>\n<p>In addition, the thresholds should depend on the parameters of the classifier used (since you are predicting training probabilities so advanced models will overfit more and vice verse)</p>\n<p>Currently I have no GPU left, but when I get it back I plan to experiment a bit with this to get an idea on how much of an increase you could get by using this method. Though I imagine there might be lots of other (and likely better) ways to hack this metric as well :) </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2803443,
              "author_name": "Jiaxiu Zou",
              "author_url": "",
              "post_date": "2024-05-09T14:07:31.143000",
              "content": "<p>thanks！！First, I trained a good LGB model. <br>\nThen, I used the probability predictions of the good model as labels to train the bad model. <br>\nAfter that, I trained a model to learn target_new. I also set 0.5 as the threshold. However, the result decreased by around 0.02. Could you please advise where I might have gone wrong?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2803531,
              "author_name": "Ern711",
              "author_url": "",
              "post_date": "2024-05-09T14:44:45.170000",
              "content": "<p>I think the easiest is just to experiment with different thresholds and models. And using the probabilities of the first model as target for the second model is not really necessary, it was just something and tried out  for fun. The important part should be that the second model performs worse (and how much worse is not really clear)</p>\n<p>Test a bunch of thresholds, and a bunch of models? Also I'm not sure how well this strategy will work in general, but since I got an increase by 0.013 quite easily I guess it will be relatively easy to get a much higher increase by picking suitable thresholds, and degrade performance enough to more or less remove the negative slope. But who knows 🤷‍♂️</p>\n<p>Or perhaps look at some strategy directly aimed at figuring out the original week number! </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2810678,
      "author_name": "chesman",
      "author_url": "",
      "post_date": "2024-05-13T11:52:50.313000",
      "content": "<p>Thanks for sharing! Interesting.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2807268,
      "author_name": "Sheema Zain",
      "author_url": "",
      "post_date": "2024-05-11T15:34:21.373000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2805777,
      "author_name": "Uzair Kath",
      "author_url": "",
      "post_date": "2024-05-10T18:25:49.823000",
      "content": "<p>Good work, thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2804751,
      "author_name": "Dibin Ke",
      "author_url": "",
      "post_date": "2024-05-10T07:27:51.113000",
      "content": "<p>nice idea! thanks</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2800753": "Since this has turned into a hacking competition, here's another strategy:\n\n* Train a good model (like AUC of 0.85)\n* Train a slightly worse (or perhaps much worse is better) model (like a regressor trained on the training probabilities of the first model)\n* Add an extra column target_new in training data and give label 0 to all obs\n* Load in the test set and add target_new column to this as well but with label 1\n* Concatenate the datasets and train a model to learn target_new, and let it predict on the data corresponding to the test set\n\nIf the probability is high enough, use the better model, if the probability is lower (more similar to training data) use the worse model - assuming observations further into the future are more dissimilar.\n\nI tested with a very simple (and far from optimized) version of this, and it still increased the leaderboard with like 0.013. Doing something like this but for a bunch of models and probability thresholds, one could probably achieve much higher gains...\n\nThere you have one way! \n\n",
    "2801707": "Thanks for sharing, the ways to hack this metric are unlimited :)\n\nI cant wait to learn more of these so I can use the learnings to my real job :)",
    "2804612": "It seems that it is possible to get a improvement of 0.04 from restoring WEEK_NUM. That's an incredible advantage lol. Though I have no idea how to restore WEEK_NUM.",
    "2815260": "Has 'date_decision' only been re-formatted in new transformation? If so can't we just do\n\n```python\nfrom dateutil import parser\n\ndef reformat_date(date_str):\n    parsed_date = parser.parse(date_str)\n    return parsed_date.strftime('%Y-%m-%d')\n```\n\nOr if they shifted it a lot there is other columns like lastapplicationdate_877D...\n\nUltimately it just sounds like you need to approximate the time series ordering so other columns may be good enough if high enough correlation to original WEEK_NUM.",
    "2814052": "> - Concatenate the datasets \n\nwhen load both train and test, there are out of memory issues and how to solve it? maybe sample or use less features?",
    "2808252": "\nHello, I'd like to ask if you've reset datedecision to its initial value now? Or are you using other methods to improve the LB score?",
    "2805416": "This strategy proposes a method to manipulate metrics by incorporating a secondary model to predict when to use a better or worse model based on data dissimilarity. While it may temporarily boost leaderboard rankings, it risks compromising the integrity of the evaluation process and lacks ethical considerations regarding fair competition and genuine performance improvement.",
    "2803348": "How do you define 'high enough probability' and 'lower probability'? What threshold do you set?",
    "2810678": "Thanks for sharing! Interesting.",
    "2807268": "Thanks for sharing!",
    "2805777": "Good work, thanks",
    "2804751": "nice idea! thanks"
  }
}