{
  "id": 333118,
  "title": "what does \"Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\" mean?",
  "url": "/competitions/amex-default-prediction/discussion/333118",
  "author_name": "",
  "post_date": "2022-06-24T18:57:30.540124600Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>what does the hosts mean by this and how should it help us?</p>\n<blockquote>\n  <p>Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.</p>\n</blockquote>",
  "messages": [
    {
      "id": "1832168",
      "postDate": "06/24/2022 18:57:30",
      "content": "<p>what does the hosts mean by this and how should it help us?</p>\n<blockquote>\n  <p>Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.</p>\n</blockquote>",
      "rawMarkdown": "what does the hosts mean by this and how should it help us?\n\n> Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.",
      "votes": null
    },
    {
      "id": "1832224",
      "postDate": "06/24/2022 20:04:07",
      "content": "<p>This has been answered several times already:</p>\n<blockquote>\n  <p>Of the 458913 customer_IDs in the training data, 340000 (74 %) have a label of 0 (good customer, no default) and 119000 (26 %) have a label of 1 (bad customer, default).</p>\n  <p>In reality, however, American Express has twenty times more good customers (they have 6.8 million non-defaulting customers). 98 % (6.8 million of 6.9 million) of the customers are good; 2 % (0.1 million of 6.9 million) are bad.</p>\n  <p>Why did they subsample the non-defaulting customers? They wanted to be nice and give us a dataset which fits into memory. If the dataset contained 6.9 million customers, every notebook would crash when you try to load the dataset, and almost nobody would participate in the competition.</p>\n</blockquote>\n<p>Answered by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> in following discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234</a></p>",
      "rawMarkdown": "This has been answered several times already:\n\n> Of the 458913 customer_IDs in the training data, 340000 (74 %) have a label of 0 (good customer, no default) and 119000 (26 %) have a label of 1 (bad customer, default).\n\n> In reality, however, American Express has twenty times more good customers (they have 6.8 million non-defaulting customers). 98 % (6.8 million of 6.9 million) of the customers are good; 2 % (0.1 million of 6.9 million) are bad.\n\n> Why did they subsample the non-defaulting customers? They wanted to be nice and give us a dataset which fits into memory. If the dataset contained 6.9 million customers, every notebook would crash when you try to load the dataset, and almost nobody would participate in the competition.\n\nAnswered by @ambrosm in following discussion: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234",
      "votes": null
    },
    {
      "id": "1832330",
      "postDate": "06/25/2022 00:29:41",
      "content": "<p>Hello there ! I would recommend the following discussion topics about metrics:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/330949\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/330949</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328890\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328890</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331886\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/331886</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234</a></p>\n<p>They cover a broad range of subjects, from the basic definition to deeper understandings of Gini coefficient to implementation.</p>",
      "rawMarkdown": "Hello there ! I would recommend the following discussion topics about metrics:\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/330949\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328890\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/331886\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\n\nThey cover a broad range of subjects, from the basic definition to deeper understandings of Gini coefficient to implementation.",
      "votes": null
    },
    {
      "id": "1838935",
      "postDate": "07/01/2022 02:28:00",
      "content": "<p>Great explanation! In addition to memory limitations, they may also performed the sub-sampling to balance the dataset.</p>",
      "rawMarkdown": "Great explanation! In addition to memory limitations, they may also performed the sub-sampling to balance the dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1832224,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "06/24/2022 20:04:07",
      "content": "<p>This has been answered several times already:</p>\n<blockquote>\n  <p>Of the 458913 customer_IDs in the training data, 340000 (74 %) have a label of 0 (good customer, no default) and 119000 (26 %) have a label of 1 (bad customer, default).</p>\n  <p>In reality, however, American Express has twenty times more good customers (they have 6.8 million non-defaulting customers). 98 % (6.8 million of 6.9 million) of the customers are good; 2 % (0.1 million of 6.9 million) are bad.</p>\n  <p>Why did they subsample the non-defaulting customers? They wanted to be nice and give us a dataset which fits into memory. If the dataset contained 6.9 million customers, every notebook would crash when you try to load the dataset, and almost nobody would participate in the competition.</p>\n</blockquote>\n<p>Answered by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> in following discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1838935,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/01/2022 02:28:00",
          "content": "<p>Great explanation! In addition to memory limitations, they may also performed the sub-sampling to balance the dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1832330,
      "author_name": "carlosasdesouza",
      "author_url": "",
      "post_date": "06/25/2022 00:29:41",
      "content": "<p>Hello there ! I would recommend the following discussion topics about metrics:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/330949\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/330949</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328890\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328890</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/329607</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331886\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/331886</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234</a></p>\n<p>They cover a broad range of subjects, from the basic definition to deeper understandings of Gini coefficient to implementation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1832168": "what does the hosts mean by this and how should it help us?\n\n> Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.",
    "1832224": "This has been answered several times already:\n\n> Of the 458913 customer_IDs in the training data, 340000 (74 %) have a label of 0 (good customer, no default) and 119000 (26 %) have a label of 1 (bad customer, default).\n\n> In reality, however, American Express has twenty times more good customers (they have 6.8 million non-defaulting customers). 98 % (6.8 million of 6.9 million) of the customers are good; 2 % (0.1 million of 6.9 million) are bad.\n\n> Why did they subsample the non-defaulting customers? They wanted to be nice and give us a dataset which fits into memory. If the dataset contained 6.9 million customers, every notebook would crash when you try to load the dataset, and almost nobody would participate in the competition.\n\nAnswered by @ambrosm in following discussion: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234",
    "1832330": "Hello there ! I would recommend the following discussion topics about metrics:\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/330949\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328890\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/329607\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/331886\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\n\nThey cover a broad range of subjects, from the basic definition to deeper understandings of Gini coefficient to implementation.",
    "1838935": "Great explanation! In addition to memory limitations, they may also performed the sub-sampling to balance the dataset."
  },
  "source": "meta"
}