{
  "id": 337836,
  "title": "“problem description” I have a problem",
  "url": "/competitions/amex-default-prediction/discussion/337836",
  "author_name": "",
  "post_date": "2022-07-18T01:09:46.188235900Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Note that the negative class has been subsampled for this dataset at 5%，这句话什么意思，subsampled 是从全部数据集（包含我们看不到的测试集）次采样？</p>",
  "messages": [
    {
      "id": "1859791",
      "postDate": "07/18/2022 01:09:46",
      "content": "<p>Note that the negative class has been subsampled for this dataset at 5%，这句话什么意思，subsampled 是从全部数据集（包含我们看不到的测试集）次采样？</p>",
      "rawMarkdown": "Note that the negative class has been subsampled for this dataset at 5%，这句话什么意思，subsampled 是从全部数据集（包含我们看不到的测试集）次采样？",
      "votes": null
    },
    {
      "id": "1859799",
      "postDate": "07/18/2022 01:37:13",
      "content": "<p>No. Both train and test data have both been subsampled at 5%. </p>\n<p>They have been subsampled from the real world. In the real world, customers default with a rate of <code>0.016%</code>. AMEX customers default 1 out of 61 customers (i.e. 1 customer defaults and 60 customers do not default). In this competition, they only provide 5% of customers who do not default. Thus they changed the <code>odds ratio</code> from 1:60 to 1:3. That is why the dataset has default mean equal 25%.</p>\n<p>If they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB! in order to provide the same number of positive samples for our models to learn the pattern of positive samples.</p>",
      "rawMarkdown": "No. Both train and test data have both been subsampled at 5%. \n\nThey have been subsampled from the real world. In the real world, customers default with a rate of `0.016%`. AMEX customers default 1 out of 61 customers (i.e. 1 customer defaults and 60 customers do not default). In this competition, they only provide 5% of customers who do not default. Thus they changed the `odds ratio` from 1:60 to 1:3. That is why the dataset has default mean equal 25%.\n\nIf they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB! in order to provide the same number of positive samples for our models to learn the pattern of positive samples.",
      "votes": null
    },
    {
      "id": "1859806",
      "postDate": "07/18/2022 01:48:27",
      "content": "<p>I understand, Thanks👍👍👍</p>",
      "rawMarkdown": "I understand, Thanks👍👍👍",
      "votes": null
    },
    {
      "id": "1859928",
      "postDate": "07/18/2022 05:01:19",
      "content": "<p>Here is a little numerical exercise since that statement still seems to be confusing.</p>\n<p>If we start with 7 million customers, at default rate of 0.0164 we will have 114800 defaults (class 1, or positive) and 6885200 non-defaults (class 0, or negative). After subsampling the negative class to 5% of its starting value, we get 344260 customers in the negative group. That would be 344260+114800 total customers, which is similar to our actual train data (458913). The same calculation applies to test data, except that all numbers other than default rate need to be multiplied by 2.</p>\n<blockquote>\n  <p>If they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB!</p>\n</blockquote>\n<p>I don't think this was the primary reason for downsampling. If they didn't do it, we would have had a dreadfully imbalanced dataset. Since we'd have a 60:1 ration of 0s to 1s, it would be very difficult to coerce a model to predict a 1 without massive sample weighting. Downsampling made this a more manageable dataset as 3:1 imbalance is not dramatic.</p>\n<p>If their main concern was dataset size, they could have given us 5% of all data without any downsampling, and that still would have been only ~50 Gb.</p>",
      "rawMarkdown": "Here is a little numerical exercise since that statement still seems to be confusing.\n\nIf we start with 7 million customers, at default rate of 0.0164 we will have 114800 defaults (class 1, or positive) and 6885200 non-defaults (class 0, or negative). After subsampling the negative class to 5% of its starting value, we get 344260 customers in the negative group. That would be 344260+114800 total customers, which is similar to our actual train data (458913). The same calculation applies to test data, except that all numbers other than default rate need to be multiplied by 2.\n\n> If they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB!\n\nI don't think this was the primary reason for downsampling. If they didn't do it, we would have had a dreadfully imbalanced dataset. Since we'd have a 60:1 ration of 0s to 1s, it would be very difficult to coerce a model to predict a 1 without massive sample weighting. Downsampling made this a more manageable dataset as 3:1 imbalance is not dramatic.\n\nIf their main concern was dataset size, they could have given us 5% of all data without any downsampling, and that still would have been only ~50 Gb.",
      "votes": null
    },
    {
      "id": "1861327",
      "postDate": "07/19/2022 01:11:40",
      "content": "<p>Thanks for your explanation, I was inspired.😄🎉</p>",
      "rawMarkdown": "Thanks for your explanation, I was inspired.😄🎉",
      "votes": null
    },
    {
      "id": "1883606",
      "postDate": "08/04/2022 00:33:40",
      "content": "<p>The whole point of building a better model for credit default detection is for early risk evaluation of potential default customers. Those customers have similar credit scores, shopping behavior, or number of credit applications on file. Those are the features that are anonymized in the data set. Potential default customers are what we are really interested in, and not those that will not default. So, having more default people data in the set is better for our models to evaluate all the features. Those people who don't default, are not so significant for the research and we can only observe reduced numbers of their records. Another thing to consider when you look at the multiple unique customer_ID records is that it might mean those applicants were rejected multiple times, but they keep applying for credit. This might affect the way those records are aggregated.</p>",
      "rawMarkdown": "The whole point of building a better model for credit default detection is for early risk evaluation of potential default customers. Those customers have similar credit scores, shopping behavior, or number of credit applications on file. Those are the features that are anonymized in the data set. Potential default customers are what we are really interested in, and not those that will not default. So, having more default people data in the set is better for our models to evaluate all the features. Those people who don't default, are not so significant for the research and we can only observe reduced numbers of their records. Another thing to consider when you look at the multiple unique customer_ID records is that it might mean those applicants were rejected multiple times, but they keep applying for credit. This might affect the way those records are aggregated.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1859799,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/18/2022 01:37:13",
      "content": "<p>No. Both train and test data have both been subsampled at 5%. </p>\n<p>They have been subsampled from the real world. In the real world, customers default with a rate of <code>0.016%</code>. AMEX customers default 1 out of 61 customers (i.e. 1 customer defaults and 60 customers do not default). In this competition, they only provide 5% of customers who do not default. Thus they changed the <code>odds ratio</code> from 1:60 to 1:3. That is why the dataset has default mean equal 25%.</p>\n<p>If they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB! in order to provide the same number of positive samples for our models to learn the pattern of positive samples.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1859806,
          "author_name": "fangxin123",
          "author_url": "",
          "post_date": "07/18/2022 01:48:27",
          "content": "<p>I understand, Thanks👍👍👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1859928,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "07/18/2022 05:01:19",
      "content": "<p>Here is a little numerical exercise since that statement still seems to be confusing.</p>\n<p>If we start with 7 million customers, at default rate of 0.0164 we will have 114800 defaults (class 1, or positive) and 6885200 non-defaults (class 0, or negative). After subsampling the negative class to 5% of its starting value, we get 344260 customers in the negative group. That would be 344260+114800 total customers, which is similar to our actual train data (458913). The same calculation applies to test data, except that all numbers other than default rate need to be multiplied by 2.</p>\n<blockquote>\n  <p>If they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB!</p>\n</blockquote>\n<p>I don't think this was the primary reason for downsampling. If they didn't do it, we would have had a dreadfully imbalanced dataset. Since we'd have a 60:1 ration of 0s to 1s, it would be very difficult to coerce a model to predict a 1 without massive sample weighting. Downsampling made this a more manageable dataset as 3:1 imbalance is not dramatic.</p>\n<p>If their main concern was dataset size, they could have given us 5% of all data without any downsampling, and that still would have been only ~50 Gb.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1861327,
          "author_name": "fangxin123",
          "author_url": "",
          "post_date": "07/19/2022 01:11:40",
          "content": "<p>Thanks for your explanation, I was inspired.😄🎉</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1883606,
      "author_name": "jdeneva",
      "author_url": "",
      "post_date": "08/04/2022 00:33:40",
      "content": "<p>The whole point of building a better model for credit default detection is for early risk evaluation of potential default customers. Those customers have similar credit scores, shopping behavior, or number of credit applications on file. Those are the features that are anonymized in the data set. Potential default customers are what we are really interested in, and not those that will not default. So, having more default people data in the set is better for our models to evaluate all the features. Those people who don't default, are not so significant for the research and we can only observe reduced numbers of their records. Another thing to consider when you look at the multiple unique customer_ID records is that it might mean those applicants were rejected multiple times, but they keep applying for credit. This might affect the way those records are aggregated.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1859791": "Note that the negative class has been subsampled for this dataset at 5%，这句话什么意思，subsampled 是从全部数据集（包含我们看不到的测试集）次采样？",
    "1859799": "No. Both train and test data have both been subsampled at 5%. \n\nThey have been subsampled from the real world. In the real world, customers default with a rate of `0.016%`. AMEX customers default 1 out of 61 customers (i.e. 1 customer defaults and 60 customers do not default). In this competition, they only provide 5% of customers who do not default. Thus they changed the `odds ratio` from 1:60 to 1:3. That is why the dataset has default mean equal 25%.\n\nIf they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB! in order to provide the same number of positive samples for our models to learn the pattern of positive samples.",
    "1859806": "I understand, Thanks👍👍👍",
    "1859928": "Here is a little numerical exercise since that statement still seems to be confusing.\n\nIf we start with 7 million customers, at default rate of 0.0164 we will have 114800 defaults (class 1, or positive) and 6885200 non-defaults (class 0, or negative). After subsampling the negative class to 5% of its starting value, we get 344260 customers in the negative group. That would be 344260+114800 total customers, which is similar to our actual train data (458913). The same calculation applies to test data, except that all numbers other than default rate need to be multiplied by 2.\n\n> If they didn't do this, then the dataset would have 15x more rows and the dataset would be 750GB instead of 50GB!\n\nI don't think this was the primary reason for downsampling. If they didn't do it, we would have had a dreadfully imbalanced dataset. Since we'd have a 60:1 ration of 0s to 1s, it would be very difficult to coerce a model to predict a 1 without massive sample weighting. Downsampling made this a more manageable dataset as 3:1 imbalance is not dramatic.\n\nIf their main concern was dataset size, they could have given us 5% of all data without any downsampling, and that still would have been only ~50 Gb.",
    "1861327": "Thanks for your explanation, I was inspired.😄🎉",
    "1883606": "The whole point of building a better model for credit default detection is for early risk evaluation of potential default customers. Those customers have similar credit scores, shopping behavior, or number of credit applications on file. Those are the features that are anonymized in the data set. Potential default customers are what we are really interested in, and not those that will not default. So, having more default people data in the set is better for our models to evaluate all the features. Those people who don't default, are not so significant for the research and we can only observe reduced numbers of their records. Another thing to consider when you look at the multiple unique customer_ID records is that it might mean those applicants were rejected multiple times, but they keep applying for credit. This might affect the way those records are aggregated."
  },
  "source": "meta"
}