{
  "id": 330949,
  "title": "Question regarding Default rate at 4% ",
  "url": "/competitions/amex-default-prediction/discussion/330949",
  "author_name": "",
  "post_date": "2022-06-15T05:02:05.030790600Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi ,<br>\nfrom the notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/inversion/amex-competition-metric-python</a> <br>\nmetric used for the default rate4% - (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum(). </p>\n<p>for this metric the max value is 0.04/0.25 = 0.16, assuming the test population has 25% of the defaulters.</p>\n<p>is it not better to have (df_cutoff['target'] == 1).sum() / len(df_cutoff) =&gt; i.e percent of defaults in the top4% instead of percent of defaults in the top4%  compared to all the population. <br>\nthe  max value for the above would be 0.04/0.04 = 1.0 (if all the 4% mentioned are defaulter.)</p>",
  "messages": [
    {
      "id": "1820905",
      "postDate": "06/15/2022 05:02:05",
      "content": "<p>Hi ,<br>\nfrom the notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/inversion/amex-competition-metric-python</a> <br>\nmetric used for the default rate4% - (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum(). </p>\n<p>for this metric the max value is 0.04/0.25 = 0.16, assuming the test population has 25% of the defaulters.</p>\n<p>is it not better to have (df_cutoff['target'] == 1).sum() / len(df_cutoff) =&gt; i.e percent of defaults in the top4% instead of percent of defaults in the top4%  compared to all the population. <br>\nthe  max value for the above would be 0.04/0.04 = 1.0 (if all the 4% mentioned are defaulter.)</p>",
      "rawMarkdown": "Hi ,\nfrom the notebook [https://www.kaggle.com/code/inversion/amex-competition-metric-python](url) \nmetric used for the default rate4% - (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum(). \n\nfor this metric the max value is 0.04/0.25 = 0.16, assuming the test population has 25% of the defaulters.\n\n\nis it not better to have (df_cutoff['target'] == 1).sum() / len(df_cutoff) => i.e percent of defaults in the top4% instead of percent of defaults in the top4%  compared to all the population. \nthe  max value for the above would be 0.04/0.04 = 1.0 (if all the 4% mentioned are defaulter.)",
      "votes": null
    },
    {
      "id": "1820960",
      "postDate": "06/15/2022 06:28:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">@narendra</a> The negative samples are weighted with a factor of 20 so that the population has only 1.7 % of defaulters, and with the weighting the maximum possible value for the metric is 1.0. See <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">here</a> for a calculation example.</p>",
      "rawMarkdown": "Hi @narendra The negative samples are weighted with a factor of 20 so that the population has only 1.7 % of defaulters, and with the weighting the maximum possible value for the metric is 1.0. See [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464) for a calculation example.",
      "votes": null
    },
    {
      "id": "1821004",
      "postDate": "06/15/2022 07:12:58",
      "content": "<p>Best to think of weighting as reversing the undersampling of negative cases. In other words, taking the published dataset we must oversample the negative cases 20x to get back to the reality</p>",
      "rawMarkdown": "Best to think of weighting as reversing the undersampling of negative cases. In other words, taking the published dataset we must oversample the negative cases 20x to get back to the reality",
      "votes": null
    },
    {
      "id": "1828399",
      "postDate": "06/21/2022 16:48:59",
      "content": "<p>thanks, missed to look on weights earlier. </p>",
      "rawMarkdown": "thanks, missed to look on weights earlier.",
      "votes": null
    },
    {
      "id": "1828401",
      "postDate": "06/21/2022 16:54:59",
      "content": "<p>very informative thnks.</p>",
      "rawMarkdown": "very informative thnks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1820960,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "06/15/2022 06:28:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">@narendra</a> The negative samples are weighted with a factor of 20 so that the population has only 1.7 % of defaulters, and with the weighting the maximum possible value for the metric is 1.0. See <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">here</a> for a calculation example.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1828401,
          "author_name": "narendra",
          "author_url": "",
          "post_date": "06/21/2022 16:54:59",
          "content": "<p>very informative thnks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1821004,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "06/15/2022 07:12:58",
      "content": "<p>Best to think of weighting as reversing the undersampling of negative cases. In other words, taking the published dataset we must oversample the negative cases 20x to get back to the reality</p>",
      "votes": null,
      "replies": [
        {
          "id": 1828399,
          "author_name": "narendra",
          "author_url": "",
          "post_date": "06/21/2022 16:48:59",
          "content": "<p>thanks, missed to look on weights earlier. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1820905": "Hi ,\nfrom the notebook [https://www.kaggle.com/code/inversion/amex-competition-metric-python](url) \nmetric used for the default rate4% - (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum(). \n\nfor this metric the max value is 0.04/0.25 = 0.16, assuming the test population has 25% of the defaulters.\n\n\nis it not better to have (df_cutoff['target'] == 1).sum() / len(df_cutoff) => i.e percent of defaults in the top4% instead of percent of defaults in the top4%  compared to all the population. \nthe  max value for the above would be 0.04/0.04 = 1.0 (if all the 4% mentioned are defaulter.)",
    "1820960": "Hi @narendra The negative samples are weighted with a factor of 20 so that the population has only 1.7 % of defaulters, and with the weighting the maximum possible value for the metric is 1.0. See [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464) for a calculation example.",
    "1821004": "Best to think of weighting as reversing the undersampling of negative cases. In other words, taking the published dataset we must oversample the negative cases 20x to get back to the reality",
    "1828399": "thanks, missed to look on weights earlier.",
    "1828401": "very informative thnks."
  },
  "source": "meta"
}