{
  "id": 501612,
  "title": "What is metric hacking？",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501612",
  "author_name": "",
  "post_date": "2024-05-10T03:32:07.717210500Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I read some discussions and found the term metric hacking, which everyone seems to hate very much. I wonder what this means and what harm it will be?🤔</p>",
  "messages": [
    {
      "id": "2804383",
      "postDate": "05/10/2024 03:32:07",
      "content": "<p>I read some discussions and found the term metric hacking, which everyone seems to hate very much. I wonder what this means and what harm it will be?🤔</p>",
      "rawMarkdown": "I read some discussions and found the term metric hacking, which everyone seems to hate very much. I wonder what this means and what harm it will be?🤔",
      "votes": null
    },
    {
      "id": "2804477",
      "postDate": "05/10/2024 04:56:14",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/zexiuw\" target=\"_blank\">@zexiuw</a> </p>\n<p>That is a good question, I would suggest that  \"metric hacking\" could be defined as applying post-processing to the results of an ML model so as to lead to a more favorable score.</p>\n<p>For example, imagine one has a (probabilistic) binary classification task that is assessed via the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html\" target=\"_blank\">log-loss</a> metric. Now say for a given row of data the model predicts a probability of 0.95. This should be the end of the story, however it is known that at the end of the day the ground truth values are either 0 or 1. Thus to obtain a more favorable score the data scientist decides to change (hack) the output of 0.95 to be 1.00, and thus obtain a (slightly) better result.</p>\n<p>However, this \"trick\" can seriously back-fire. Imagine this particular row of data were misclassified, and the ground truth label was actually 0 rather than 1. In this case the loss incurred will be infinite (unless the metric is clipped), and the overall \"hacked\" score will be a total disaster. Indeed the log-loss is one of the few examples of a small set of <a href=\"https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf\" target=\"_blank\">strictly proper scoring rules</a>. On the whole these metrics are un-hackable and the optimal game is to be honest and stick with the original model output.</p>\n<p>However, in this particular competition the so-called \"stability metric\" is <strong>not</strong> one of the strictly proper scoring rules, and thus post-processing the model output can indeed lead to a genuine and lasting improvement in ones score, and if on Kaggle it can be done - it <em>will</em> be done!</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @zexiuw \n\nThat is a good question, I would suggest that  \"metric hacking\" could be defined as applying post-processing to the results of an ML model so as to lead to a more favorable score.\n\nFor example, imagine one has a (probabilistic) binary classification task that is assessed via the [log-loss](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html) metric. Now say for a given row of data the model predicts a probability of 0.95. This should be the end of the story, however it is known that at the end of the day the ground truth values are either 0 or 1. Thus to obtain a more favorable score the data scientist decides to change (hack) the output of 0.95 to be 1.00, and thus obtain a (slightly) better result.\n\nHowever, this \"trick\" can seriously back-fire. Imagine this particular row of data were misclassified, and the ground truth label was actually 0 rather than 1. In this case the loss incurred will be infinite (unless the metric is clipped), and the overall \"hacked\" score will be a total disaster. Indeed the log-loss is one of the few examples of a small set of [strictly proper scoring rules](https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf). On the whole these metrics are un-hackable and the optimal game is to be honest and stick with the original model output.\n\nHowever, in this particular competition the so-called \"stability metric\" is **not** one of the strictly proper scoring rules, and thus post-processing the model output can indeed lead to a genuine and lasting improvement in ones score, and if on Kaggle it can be done - it *will* be done!\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "2804631",
      "postDate": "05/10/2024 06:13:06",
      "content": "<p>I am also new to Kaggle, the way I see it is:</p>\n<ul>\n<li>To achieve good score by apply various \"tricks\" without meeting the competition goals.</li>\n</ul>\n<p>These \"tricks\" could be as simple as to produce a mean value for every input, or it could a complex post-processing.</p>\n<p>For the “trick” to be classified as a hack it must fulfil the criterion, i.e, it must not meet the competition goals. For example, ensembling is also post-processing, but it is not hacking since it meets the competition goals. </p>\n<p>Why everyone hates it?<br>\nKaggle is a platform which is run by a symbiotic relationship between organizations and kagglers. </p>\n<p>Now,<br>\nIf hosts stops making effort to make competition more transparent for everyone than soon most of the competitions will be won by \"tricks\" this will make other organization stop posting problems (they don't want tricks in exchange for 100s and 1000s of dollars) and it will also evaporate the trust of kagglers to put so much effort in solving a problem just to see it failing against a hack. So, in short, it's a slow poison and will slowly but surely kill kaggle.</p>",
      "rawMarkdown": "I am also new to Kaggle, the way I see it is:\n- To achieve good score by apply various \"tricks\" without meeting the competition goals.\n\nThese \"tricks\" could be as simple as to produce a mean value for every input, or it could a complex post-processing.\n\nFor the “trick” to be classified as a hack it must fulfil the criterion, i.e, it must not meet the competition goals. For example, ensembling is also post-processing, but it is not hacking since it meets the competition goals. \n\nWhy everyone hates it?\nKaggle is a platform which is run by a symbiotic relationship between organizations and kagglers. \n\nNow,\nIf hosts stops making effort to make competition more transparent for everyone than soon most of the competitions will be won by \"tricks\" this will make other organization stop posting problems (they don't want tricks in exchange for 100s and 1000s of dollars) and it will also evaporate the trust of kagglers to put so much effort in solving a problem just to see it failing against a hack. So, in short, it's a slow poison and will slowly but surely kill kaggle.",
      "votes": null
    },
    {
      "id": "2804828",
      "postDate": "05/10/2024 08:22:06",
      "content": "<p>Ensembling is not post-processing, it predicts a result like any single model. But taking these results and manipulate with them - ex. multiply by some factors in dependence of the row number - is a post processing. In any case the hosts gave the green light to do whatever we want and use any tricks <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412</a><br>\nThat's why the scores exploded in LB the last days and I don't bother anymore to improve my model because of no intereset in hacking (it's not ML already), so it became a hacking competition now. When everybody were afraid because of the manual verification - the scores were under 0.600</p>",
      "rawMarkdown": "Ensembling is not post-processing, it predicts a result like any single model. But taking these results and manipulate with them - ex. multiply by some factors in dependence of the row number - is a post processing. In any case the hosts gave the green light to do whatever we want and use any tricks https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412\nThat's why the scores exploded in LB the last days and I don't bother anymore to improve my model because of no intereset in hacking (it's not ML already), so it became a hacking competition now. When everybody were afraid because of the manual verification - the scores were under 0.600",
      "votes": null
    },
    {
      "id": "2804978",
      "postDate": "05/10/2024 09:40:53",
      "content": "<p>You are right the score exploded after the announcement, it is remorseful. Maybe if a proper hack is made public, it will relatively improve the situation.</p>",
      "rawMarkdown": "You are right the score exploded after the announcement, it is remorseful. Maybe if a proper hack is made public, it will relatively improve the situation.",
      "votes": null
    },
    {
      "id": "2805000",
      "postDate": "05/10/2024 10:01:13",
      "content": "<p>The goal of this competition was the model stability, it just drifted to hacking. To even try to improve the situation the hosts could keep their first word - to manually verify TOP 100 and disqualify the user and not just the submission, because each cheater was ready to use one cheating solution and one normal one as a backup, but we have what we have.</p>",
      "rawMarkdown": "The goal of this competition was the model stability, it just drifted to hacking. To even try to improve the situation the hosts could keep their first word - to manually verify TOP 100 and disqualify the user and not just the submission, because each cheater was ready to use one cheating solution and one normal one as a backup, but we have what we have.",
      "votes": null
    },
    {
      "id": "2807642",
      "postDate": "05/11/2024 18:59:02",
      "content": "<p>Oh, metric hacking is like sneaky score-boosting in real life! Picture someone tinkering with numbers to make things look better than they really are, whether in business or online platforms. It's the kind of trickery nobody likes because it messes with the truth and can lead to all sorts of misunderstandings. Think of it as trying to drive with a broken GPS through thick fog—it's easy to veer off course without even realizing it! So yeah, it's definitely something folks frown upon because it messes with trust, damages reputations, and can throw everything out of whack.</p>",
      "rawMarkdown": "Oh, metric hacking is like sneaky score-boosting in real life! Picture someone tinkering with numbers to make things look better than they really are, whether in business or online platforms. It's the kind of trickery nobody likes because it messes with the truth and can lead to all sorts of misunderstandings. Think of it as trying to drive with a broken GPS through thick fog—it's easy to veer off course without even realizing it! So yeah, it's definitely something folks frown upon because it messes with trust, damages reputations, and can throw everything out of whack.",
      "votes": null
    },
    {
      "id": "2824765",
      "postDate": "05/20/2024 02:48:10",
      "content": "<p>Metric hacking is the practice of manipulating metrics to create the appearance of improved performance without genuine improvement. It can lead to misleading performance indicators, erosion of trust, and misallocation of resources. This often prioritizes short-term gains over long-term goals and stifles innovation. Combating metric hacking involves using balanced scorecards, focusing on outcomes, regular reviews, and incorporating stakeholder feedback.</p>",
      "rawMarkdown": "Metric hacking is the practice of manipulating metrics to create the appearance of improved performance without genuine improvement. It can lead to misleading performance indicators, erosion of trust, and misallocation of resources. This often prioritizes short-term gains over long-term goals and stifles innovation. Combating metric hacking involves using balanced scorecards, focusing on outcomes, regular reviews, and incorporating stakeholder feedback.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2804477,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "05/10/2024 04:56:14",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/zexiuw\" target=\"_blank\">@zexiuw</a> </p>\n<p>That is a good question, I would suggest that  \"metric hacking\" could be defined as applying post-processing to the results of an ML model so as to lead to a more favorable score.</p>\n<p>For example, imagine one has a (probabilistic) binary classification task that is assessed via the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html\" target=\"_blank\">log-loss</a> metric. Now say for a given row of data the model predicts a probability of 0.95. This should be the end of the story, however it is known that at the end of the day the ground truth values are either 0 or 1. Thus to obtain a more favorable score the data scientist decides to change (hack) the output of 0.95 to be 1.00, and thus obtain a (slightly) better result.</p>\n<p>However, this \"trick\" can seriously back-fire. Imagine this particular row of data were misclassified, and the ground truth label was actually 0 rather than 1. In this case the loss incurred will be infinite (unless the metric is clipped), and the overall \"hacked\" score will be a total disaster. Indeed the log-loss is one of the few examples of a small set of <a href=\"https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf\" target=\"_blank\">strictly proper scoring rules</a>. On the whole these metrics are un-hackable and the optimal game is to be honest and stick with the original model output.</p>\n<p>However, in this particular competition the so-called \"stability metric\" is <strong>not</strong> one of the strictly proper scoring rules, and thus post-processing the model output can indeed lead to a genuine and lasting improvement in ones score, and if on Kaggle it can be done - it <em>will</em> be done!</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2804631,
      "author_name": "cemuzzamil",
      "author_url": "",
      "post_date": "05/10/2024 06:13:06",
      "content": "<p>I am also new to Kaggle, the way I see it is:</p>\n<ul>\n<li>To achieve good score by apply various \"tricks\" without meeting the competition goals.</li>\n</ul>\n<p>These \"tricks\" could be as simple as to produce a mean value for every input, or it could a complex post-processing.</p>\n<p>For the “trick” to be classified as a hack it must fulfil the criterion, i.e, it must not meet the competition goals. For example, ensembling is also post-processing, but it is not hacking since it meets the competition goals. </p>\n<p>Why everyone hates it?<br>\nKaggle is a platform which is run by a symbiotic relationship between organizations and kagglers. </p>\n<p>Now,<br>\nIf hosts stops making effort to make competition more transparent for everyone than soon most of the competitions will be won by \"tricks\" this will make other organization stop posting problems (they don't want tricks in exchange for 100s and 1000s of dollars) and it will also evaporate the trust of kagglers to put so much effort in solving a problem just to see it failing against a hack. So, in short, it's a slow poison and will slowly but surely kill kaggle.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2804828,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "05/10/2024 08:22:06",
          "content": "<p>Ensembling is not post-processing, it predicts a result like any single model. But taking these results and manipulate with them - ex. multiply by some factors in dependence of the row number - is a post processing. In any case the hosts gave the green light to do whatever we want and use any tricks <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412</a><br>\nThat's why the scores exploded in LB the last days and I don't bother anymore to improve my model because of no intereset in hacking (it's not ML already), so it became a hacking competition now. When everybody were afraid because of the manual verification - the scores were under 0.600</p>",
          "votes": null,
          "replies": [
            {
              "id": 2804978,
              "author_name": "cemuzzamil",
              "author_url": "",
              "post_date": "05/10/2024 09:40:53",
              "content": "<p>You are right the score exploded after the announcement, it is remorseful. Maybe if a proper hack is made public, it will relatively improve the situation.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2805000,
                  "author_name": "eu1234",
                  "author_url": "",
                  "post_date": "05/10/2024 10:01:13",
                  "content": "<p>The goal of this competition was the model stability, it just drifted to hacking. To even try to improve the situation the hosts could keep their first word - to manually verify TOP 100 and disqualify the user and not just the submission, because each cheater was ready to use one cheating solution and one normal one as a backup, but we have what we have.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2807642,
      "author_name": "hragsousani",
      "author_url": "",
      "post_date": "05/11/2024 18:59:02",
      "content": "<p>Oh, metric hacking is like sneaky score-boosting in real life! Picture someone tinkering with numbers to make things look better than they really are, whether in business or online platforms. It's the kind of trickery nobody likes because it messes with the truth and can lead to all sorts of misunderstandings. Think of it as trying to drive with a broken GPS through thick fog—it's easy to veer off course without even realizing it! So yeah, it's definitely something folks frown upon because it messes with trust, damages reputations, and can throw everything out of whack.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2824765,
      "author_name": "mianwu2024",
      "author_url": "",
      "post_date": "05/20/2024 02:48:10",
      "content": "<p>Metric hacking is the practice of manipulating metrics to create the appearance of improved performance without genuine improvement. It can lead to misleading performance indicators, erosion of trust, and misallocation of resources. This often prioritizes short-term gains over long-term goals and stifles innovation. Combating metric hacking involves using balanced scorecards, focusing on outcomes, regular reviews, and incorporating stakeholder feedback.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2804383": "I read some discussions and found the term metric hacking, which everyone seems to hate very much. I wonder what this means and what harm it will be?🤔",
    "2804477": "Dear @zexiuw \n\nThat is a good question, I would suggest that  \"metric hacking\" could be defined as applying post-processing to the results of an ML model so as to lead to a more favorable score.\n\nFor example, imagine one has a (probabilistic) binary classification task that is assessed via the [log-loss](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html) metric. Now say for a given row of data the model predicts a probability of 0.95. This should be the end of the story, however it is known that at the end of the day the ground truth values are either 0 or 1. Thus to obtain a more favorable score the data scientist decides to change (hack) the output of 0.95 to be 1.00, and thus obtain a (slightly) better result.\n\nHowever, this \"trick\" can seriously back-fire. Imagine this particular row of data were misclassified, and the ground truth label was actually 0 rather than 1. In this case the loss incurred will be infinite (unless the metric is clipped), and the overall \"hacked\" score will be a total disaster. Indeed the log-loss is one of the few examples of a small set of [strictly proper scoring rules](https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf). On the whole these metrics are un-hackable and the optimal game is to be honest and stick with the original model output.\n\nHowever, in this particular competition the so-called \"stability metric\" is **not** one of the strictly proper scoring rules, and thus post-processing the model output can indeed lead to a genuine and lasting improvement in ones score, and if on Kaggle it can be done - it *will* be done!\n\nAll the best,\ncarl",
    "2804631": "I am also new to Kaggle, the way I see it is:\n- To achieve good score by apply various \"tricks\" without meeting the competition goals.\n\nThese \"tricks\" could be as simple as to produce a mean value for every input, or it could a complex post-processing.\n\nFor the “trick” to be classified as a hack it must fulfil the criterion, i.e, it must not meet the competition goals. For example, ensembling is also post-processing, but it is not hacking since it meets the competition goals. \n\nWhy everyone hates it?\nKaggle is a platform which is run by a symbiotic relationship between organizations and kagglers. \n\nNow,\nIf hosts stops making effort to make competition more transparent for everyone than soon most of the competitions will be won by \"tricks\" this will make other organization stop posting problems (they don't want tricks in exchange for 100s and 1000s of dollars) and it will also evaporate the trust of kagglers to put so much effort in solving a problem just to see it failing against a hack. So, in short, it's a slow poison and will slowly but surely kill kaggle.",
    "2804828": "Ensembling is not post-processing, it predicts a result like any single model. But taking these results and manipulate with them - ex. multiply by some factors in dependence of the row number - is a post processing. In any case the hosts gave the green light to do whatever we want and use any tricks https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412\nThat's why the scores exploded in LB the last days and I don't bother anymore to improve my model because of no intereset in hacking (it's not ML already), so it became a hacking competition now. When everybody were afraid because of the manual verification - the scores were under 0.600",
    "2804978": "You are right the score exploded after the announcement, it is remorseful. Maybe if a proper hack is made public, it will relatively improve the situation.",
    "2805000": "The goal of this competition was the model stability, it just drifted to hacking. To even try to improve the situation the hosts could keep their first word - to manually verify TOP 100 and disqualify the user and not just the submission, because each cheater was ready to use one cheating solution and one normal one as a backup, but we have what we have.",
    "2807642": "Oh, metric hacking is like sneaky score-boosting in real life! Picture someone tinkering with numbers to make things look better than they really are, whether in business or online platforms. It's the kind of trickery nobody likes because it messes with the truth and can lead to all sorts of misunderstandings. Think of it as trying to drive with a broken GPS through thick fog—it's easy to veer off course without even realizing it! So yeah, it's definitely something folks frown upon because it messes with trust, damages reputations, and can throw everything out of whack.",
    "2824765": "Metric hacking is the practice of manipulating metrics to create the appearance of improved performance without genuine improvement. It can lead to misleading performance indicators, erosion of trust, and misallocation of resources. This often prioritizes short-term gains over long-term goals and stifles innovation. Combating metric hacking involves using balanced scorecards, focusing on outcomes, regular reviews, and incorporating stakeholder feedback."
  },
  "source": "meta"
}