{
  "id": 329447,
  "title": "Ensembling models focusing on parts of the competition metric",
  "url": "/competitions/amex-default-prediction/discussion/329447",
  "author_name": "",
  "post_date": "2022-06-06T18:02:30.539246100Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Ensembling models focusing on parts of the competition metric</h1>\n<p>After again reading the amazing <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">ambrosm</a> I started thinking about methods for taking the metric into consideration.</p>\n<p>I came up with the following idea:</p>\n<p><strong>What if we can train models that are specifically trained to be better on each part of the metric function and then ensemble them together?</strong></p>\n<p>This comes from understanding what the metric is trying to do. It's a pretty simple idea really: it wants us to \"catch\" as many positive predictions as possible while also making as few false positives predictions as possible. The first part is captured by the default rate at 4% and the second part by the Gini coefficient. So our goal should be to build an ensemble of models that strike a good balance between these two.</p>\n<p>Another way of putting it might be that we need models which are both good at predicting defaults and \"usually right\" in the rest of their predictions.</p>\n<ul>\n<li><p>One way of doing this is to use an ensemble of models, each of which is better at one task or the other. We can control this by calibrating a threshold (or a \"soft threshold\" function) and making some models more sensitive than others. If we have a robust CV strategy we might even be able to optimize this threshold based on it. </p></li>\n<li><p>Another way we can make our models behave this way is by <strong>training them deliberately under/over the sampled training set</strong>, making them learn \"a skewed\" distribution of the data. <br>\n<code>This will also work when training NN using balanced batches.</code></p></li>\n<li><p>We can also simply use early stopping based on each one of the metric components thus creating models that are highly optimized for each part of the metric. </p></li>\n</ul>\n<p>I'm not saying that one option is better than the other, but the point to take away is that it might be possible to get better performance by optimizing for multiple objectives simultaneously since the metric overall is a combination of multiple metrics. The key, however, will be to find some way to combine these models, so that their strengths are amplified and their weaknesses are muted.</p>\n<p>Just as an example, we might want to first try and take the avg/median and then we should probably just try to use some linear combination of the model's predictions. It might give us an ensemble that performs well on both tasks if we can assign weights to them correctly.</p>\n<p>Which approach is better will depend on the data and the models involved, so it's something we'll have to experiment with.</p>\n<p>What do you think? <br>\nDid anyone try something like this?</p>",
  "messages": [
    {
      "id": "1813336",
      "postDate": "06/06/2022 18:02:30",
      "content": "<h1>Ensembling models focusing on parts of the competition metric</h1>\n<p>After again reading the amazing <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">ambrosm</a> I started thinking about methods for taking the metric into consideration.</p>\n<p>I came up with the following idea:</p>\n<p><strong>What if we can train models that are specifically trained to be better on each part of the metric function and then ensemble them together?</strong></p>\n<p>This comes from understanding what the metric is trying to do. It's a pretty simple idea really: it wants us to \"catch\" as many positive predictions as possible while also making as few false positives predictions as possible. The first part is captured by the default rate at 4% and the second part by the Gini coefficient. So our goal should be to build an ensemble of models that strike a good balance between these two.</p>\n<p>Another way of putting it might be that we need models which are both good at predicting defaults and \"usually right\" in the rest of their predictions.</p>\n<ul>\n<li><p>One way of doing this is to use an ensemble of models, each of which is better at one task or the other. We can control this by calibrating a threshold (or a \"soft threshold\" function) and making some models more sensitive than others. If we have a robust CV strategy we might even be able to optimize this threshold based on it. </p></li>\n<li><p>Another way we can make our models behave this way is by <strong>training them deliberately under/over the sampled training set</strong>, making them learn \"a skewed\" distribution of the data. <br>\n<code>This will also work when training NN using balanced batches.</code></p></li>\n<li><p>We can also simply use early stopping based on each one of the metric components thus creating models that are highly optimized for each part of the metric. </p></li>\n</ul>\n<p>I'm not saying that one option is better than the other, but the point to take away is that it might be possible to get better performance by optimizing for multiple objectives simultaneously since the metric overall is a combination of multiple metrics. The key, however, will be to find some way to combine these models, so that their strengths are amplified and their weaknesses are muted.</p>\n<p>Just as an example, we might want to first try and take the avg/median and then we should probably just try to use some linear combination of the model's predictions. It might give us an ensemble that performs well on both tasks if we can assign weights to them correctly.</p>\n<p>Which approach is better will depend on the data and the models involved, so it's something we'll have to experiment with.</p>\n<p>What do you think? <br>\nDid anyone try something like this?</p>",
      "rawMarkdown": "# Ensembling models focusing on parts of the competition metric\n\nAfter again reading the amazing [post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464) by [ambrosm](https://www.kaggle.com/ambrosm) I started thinking about methods for taking the metric into consideration.\n\nI came up with the following idea:\n\n**What if we can train models that are specifically trained to be better on each part of the metric function and then ensemble them together?**\n\nThis comes from understanding what the metric is trying to do. It's a pretty simple idea really: it wants us to \"catch\" as many positive predictions as possible while also making as few false positives predictions as possible. The first part is captured by the default rate at 4% and the second part by the Gini coefficient. So our goal should be to build an ensemble of models that strike a good balance between these two.\n\nAnother way of putting it might be that we need models which are both good at predicting defaults and \"usually right\" in the rest of their predictions.\n\n- One way of doing this is to use an ensemble of models, each of which is better at one task or the other. We can control this by calibrating a threshold (or a \"soft threshold\" function) and making some models more sensitive than others. If we have a robust CV strategy we might even be able to optimize this threshold based on it. \n\n- Another way we can make our models behave this way is by **training them deliberately under/over the sampled training set**, making them learn \"a skewed\" distribution of the data. \n```This will also work when training NN using balanced batches.```\n\n- We can also simply use early stopping based on each one of the metric components thus creating models that are highly optimized for each part of the metric. \n\nI'm not saying that one option is better than the other, but the point to take away is that it might be possible to get better performance by optimizing for multiple objectives simultaneously since the metric overall is a combination of multiple metrics. The key, however, will be to find some way to combine these models, so that their strengths are amplified and their weaknesses are muted.\n\nJust as an example, we might want to first try and take the avg/median and then we should probably just try to use some linear combination of the model's predictions. It might give us an ensemble that performs well on both tasks if we can assign weights to them correctly.\n\nWhich approach is better will depend on the data and the models involved, so it's something we'll have to experiment with.\n\nWhat do you think? \nDid anyone try something like this?",
      "votes": null
    },
    {
      "id": "1813459",
      "postDate": "06/06/2022 20:52:14",
      "content": "<p>Besides over-sampling positives, are there any off-the-shelf evaluation criteria that would be more optimized towards correctly picking the positives? The other half - optimizing for AUC - is already there.</p>\n<p>I noticed that XGBoost allows multiple evaluation criteria, so if there were a good choice for the \"default rate at 4%\" metric, you could also try a model that uses both 'auc' and that second criteria.</p>",
      "rawMarkdown": "Besides over-sampling positives, are there any off-the-shelf evaluation criteria that would be more optimized towards correctly picking the positives? The other half - optimizing for AUC - is already there.\n\nI noticed that XGBoost allows multiple evaluation criteria, so if there were a good choice for the \"default rate at 4%\" metric, you could also try a model that uses both 'auc' and that second criteria.",
      "votes": null
    },
    {
      "id": "1814221",
      "postDate": "06/07/2022 16:14:40",
      "content": "<p>A couple more thoughts:</p>\n<ul>\n<li><p>We should stop (only) printing \"The Kaggle Metric\", and print both sub-components. This night give insight if and when an experiment improves one while hurting the other, or improves one by more then the other.  I suspect they might not vary too much from each other though. </p></li>\n<li><p>From the perspective of the (skewed) data given, it's not exactly as extreme as \"4%\". It's between 3% (0.0 score) to 19% (0.7 score) to 27% (1.0 score) in terms of the visible data. Since each negative represents 20 misses, the worse you do, the fewer predictions you get to check and count for your score. Probably pretty close if you think of it as going until 2% incorrect predictions are reached. (eg out of 400,000, pick your best guess for defaults until ~8000 incorrect guesses are reached.)</p></li>\n</ul>",
      "rawMarkdown": "A couple more thoughts:\n* We should stop (only) printing \"The Kaggle Metric\", and print both sub-components. This night give insight if and when an experiment improves one while hurting the other, or improves one by more then the other.  I suspect they might not vary too much from each other though. \n\n* From the perspective of the (skewed) data given, it's not exactly as extreme as \"4%\". It's between 3% (0.0 score) to 19% (0.7 score) to 27% (1.0 score) in terms of the visible data. Since each negative represents 20 misses, the worse you do, the fewer predictions you get to check and count for your score. Probably pretty close if you think of it as going until 2% incorrect predictions are reached. (eg out of 400,000, pick your best guess for defaults until ~8000 incorrect guesses are reached.)",
      "votes": null
    },
    {
      "id": "1814413",
      "postDate": "06/07/2022 21:20:18",
      "content": "<p>For XGB, this gets similar results to binary:logistic + logloss, and might better optimize for the 4% metric, so might be worth experimenting with:<br>\n    'eval_metric':'auc',<br>\n    'objective':'rank:pairwise',</p>",
      "rawMarkdown": "For XGB, this gets similar results to binary:logistic + logloss, and might better optimize for the 4% metric, so might be worth experimenting with:\n    'eval_metric':'auc',\n    'objective':'rank:pairwise',",
      "votes": null
    },
    {
      "id": "1869999",
      "postDate": "07/25/2022 07:28:36",
      "content": "<p>I think early stopping separately on the components won't add any value. They were 99.7+ correlated with the one early stopped on amex metric when I'd tried it.</p>",
      "rawMarkdown": "I think early stopping separately on the components won't add any value. They were 99.7+ correlated with the one early stopped on amex metric when I'd tried it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1813459,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "06/06/2022 20:52:14",
      "content": "<p>Besides over-sampling positives, are there any off-the-shelf evaluation criteria that would be more optimized towards correctly picking the positives? The other half - optimizing for AUC - is already there.</p>\n<p>I noticed that XGBoost allows multiple evaluation criteria, so if there were a good choice for the \"default rate at 4%\" metric, you could also try a model that uses both 'auc' and that second criteria.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1814413,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "06/07/2022 21:20:18",
          "content": "<p>For XGB, this gets similar results to binary:logistic + logloss, and might better optimize for the 4% metric, so might be worth experimenting with:<br>\n    'eval_metric':'auc',<br>\n    'objective':'rank:pairwise',</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1814221,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "06/07/2022 16:14:40",
      "content": "<p>A couple more thoughts:</p>\n<ul>\n<li><p>We should stop (only) printing \"The Kaggle Metric\", and print both sub-components. This night give insight if and when an experiment improves one while hurting the other, or improves one by more then the other.  I suspect they might not vary too much from each other though. </p></li>\n<li><p>From the perspective of the (skewed) data given, it's not exactly as extreme as \"4%\". It's between 3% (0.0 score) to 19% (0.7 score) to 27% (1.0 score) in terms of the visible data. Since each negative represents 20 misses, the worse you do, the fewer predictions you get to check and count for your score. Probably pretty close if you think of it as going until 2% incorrect predictions are reached. (eg out of 400,000, pick your best guess for defaults until ~8000 incorrect guesses are reached.)</p></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1869999,
      "author_name": "celiker",
      "author_url": "",
      "post_date": "07/25/2022 07:28:36",
      "content": "<p>I think early stopping separately on the components won't add any value. They were 99.7+ correlated with the one early stopped on amex metric when I'd tried it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1813336": "# Ensembling models focusing on parts of the competition metric\n\nAfter again reading the amazing [post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464) by [ambrosm](https://www.kaggle.com/ambrosm) I started thinking about methods for taking the metric into consideration.\n\nI came up with the following idea:\n\n**What if we can train models that are specifically trained to be better on each part of the metric function and then ensemble them together?**\n\nThis comes from understanding what the metric is trying to do. It's a pretty simple idea really: it wants us to \"catch\" as many positive predictions as possible while also making as few false positives predictions as possible. The first part is captured by the default rate at 4% and the second part by the Gini coefficient. So our goal should be to build an ensemble of models that strike a good balance between these two.\n\nAnother way of putting it might be that we need models which are both good at predicting defaults and \"usually right\" in the rest of their predictions.\n\n- One way of doing this is to use an ensemble of models, each of which is better at one task or the other. We can control this by calibrating a threshold (or a \"soft threshold\" function) and making some models more sensitive than others. If we have a robust CV strategy we might even be able to optimize this threshold based on it. \n\n- Another way we can make our models behave this way is by **training them deliberately under/over the sampled training set**, making them learn \"a skewed\" distribution of the data. \n```This will also work when training NN using balanced batches.```\n\n- We can also simply use early stopping based on each one of the metric components thus creating models that are highly optimized for each part of the metric. \n\nI'm not saying that one option is better than the other, but the point to take away is that it might be possible to get better performance by optimizing for multiple objectives simultaneously since the metric overall is a combination of multiple metrics. The key, however, will be to find some way to combine these models, so that their strengths are amplified and their weaknesses are muted.\n\nJust as an example, we might want to first try and take the avg/median and then we should probably just try to use some linear combination of the model's predictions. It might give us an ensemble that performs well on both tasks if we can assign weights to them correctly.\n\nWhich approach is better will depend on the data and the models involved, so it's something we'll have to experiment with.\n\nWhat do you think? \nDid anyone try something like this?",
    "1813459": "Besides over-sampling positives, are there any off-the-shelf evaluation criteria that would be more optimized towards correctly picking the positives? The other half - optimizing for AUC - is already there.\n\nI noticed that XGBoost allows multiple evaluation criteria, so if there were a good choice for the \"default rate at 4%\" metric, you could also try a model that uses both 'auc' and that second criteria.",
    "1814221": "A couple more thoughts:\n* We should stop (only) printing \"The Kaggle Metric\", and print both sub-components. This night give insight if and when an experiment improves one while hurting the other, or improves one by more then the other.  I suspect they might not vary too much from each other though. \n\n* From the perspective of the (skewed) data given, it's not exactly as extreme as \"4%\". It's between 3% (0.0 score) to 19% (0.7 score) to 27% (1.0 score) in terms of the visible data. Since each negative represents 20 misses, the worse you do, the fewer predictions you get to check and count for your score. Probably pretty close if you think of it as going until 2% incorrect predictions are reached. (eg out of 400,000, pick your best guess for defaults until ~8000 incorrect guesses are reached.)",
    "1814413": "For XGB, this gets similar results to binary:logistic + logloss, and might better optimize for the 4% metric, so might be worth experimenting with:\n    'eval_metric':'auc',\n    'objective':'rank:pairwise',",
    "1869999": "I think early stopping separately on the components won't add any value. They were 99.7+ correlated with the one early stopped on amex metric when I'd tried it."
  },
  "source": "meta"
}