{
  "id": 337736,
  "title": "What does the weighting in the scoring metric mean ? ",
  "url": "/competitions/amex-default-prediction/discussion/337736",
  "author_name": "",
  "post_date": "2022-07-17T11:58:23.824368400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I did not understand the following statement in the Data description:</p>\n<p><strong>Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.</strong></p>\n<p>What does it mean ?</p>\n<p>Thank you in advance 👍.</p>",
  "messages": [
    {
      "id": "1859090",
      "postDate": "07/17/2022 11:58:23",
      "content": "<p>Hi,</p>\n<p>I did not understand the following statement in the Data description:</p>\n<p><strong>Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.</strong></p>\n<p>What does it mean ?</p>\n<p>Thank you in advance 👍.</p>",
      "rawMarkdown": "Hi,\n\nI did not understand the following statement in the Data description:\n\n**Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.**\n\nWhat does it mean ?\n\nThank you in advance 👍.",
      "votes": null
    },
    {
      "id": "1859328",
      "postDate": "07/17/2022 14:36:29",
      "content": "<p>it means that the data provider reduced the size of class 0 in order to balance the dataset.</p>",
      "rawMarkdown": "it means that the data provider reduced the size of class 0 in order to balance the dataset.",
      "votes": null
    },
    {
      "id": "1859677",
      "postDate": "07/17/2022 20:52:20",
      "content": "<p>The negative class means no-default. In Amex original data, they are the vast majority. It is difficult to train a model to recognise defaulters in an ocean of well-behaved customers.<br>\nTo increase the density of defaulters in the data, Amex removed 95 out of 100 non-defaulters (5% have been kept, hence the 5% subsample). This allows to have more default signal for the same file size.<br>\nAfter this subsampling, the density of defaulters increased to approximately 25%. These are the files we are training our models on. (train and label).<br>\nTest has been subsampled as well.  But when applying the metrics, the original conditions need to be somehow recreated. To do this, each non defaulter receives a factor x20 when calculating the metrics. </p>",
      "rawMarkdown": "The negative class means no-default. In Amex original data, they are the vast majority. It is difficult to train a model to recognise defaulters in an ocean of well-behaved customers.\nTo increase the density of defaulters in the data, Amex removed 95 out of 100 non-defaulters (5% have been kept, hence the 5% subsample). This allows to have more default signal for the same file size.\nAfter this subsampling, the density of defaulters increased to approximately 25%. These are the files we are training our models on. (train and label).\nTest has been subsampled as well.  But when applying the metrics, the original conditions need to be somehow recreated. To do this, each non defaulter receives a factor x20 when calculating the metrics.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1859328,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/17/2022 14:36:29",
      "content": "<p>it means that the data provider reduced the size of class 0 in order to balance the dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1859677,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "07/17/2022 20:52:20",
      "content": "<p>The negative class means no-default. In Amex original data, they are the vast majority. It is difficult to train a model to recognise defaulters in an ocean of well-behaved customers.<br>\nTo increase the density of defaulters in the data, Amex removed 95 out of 100 non-defaulters (5% have been kept, hence the 5% subsample). This allows to have more default signal for the same file size.<br>\nAfter this subsampling, the density of defaulters increased to approximately 25%. These are the files we are training our models on. (train and label).<br>\nTest has been subsampled as well.  But when applying the metrics, the original conditions need to be somehow recreated. To do this, each non defaulter receives a factor x20 when calculating the metrics. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1859090": "Hi,\n\nI did not understand the following statement in the Data description:\n\n**Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.**\n\nWhat does it mean ?\n\nThank you in advance 👍.",
    "1859328": "it means that the data provider reduced the size of class 0 in order to balance the dataset.",
    "1859677": "The negative class means no-default. In Amex original data, they are the vast majority. It is difficult to train a model to recognise defaulters in an ocean of well-behaved customers.\nTo increase the density of defaulters in the data, Amex removed 95 out of 100 non-defaulters (5% have been kept, hence the 5% subsample). This allows to have more default signal for the same file size.\nAfter this subsampling, the density of defaulters increased to approximately 25%. These are the files we are training our models on. (train and label).\nTest has been subsampled as well.  But when applying the metrics, the original conditions need to be somehow recreated. To do this, each non defaulter receives a factor x20 when calculating the metrics."
  },
  "source": "meta"
}