{
  "id": 336193,
  "title": "Customers with fewer statements have a higher probability of default",
  "url": "/competitions/amex-default-prediction/discussion/336193",
  "author_name": "",
  "post_date": "2022-07-09T20:35:56.048618800Z",
  "votes": 22,
  "comment_count": 6,
  "views": 0,
  "content": "<p>As you may have noticed,<br>\nthere's a relation between the number of each customer's statements and its probability of default.</p>\n<p>In general, customers who continuously use credit cards are considered to be more creditworthy.<br>\nTherefore, it can be assumed that the default probability is relatively higher for customers who use less frequently.</p>\n<p>In this competition, the number of statements for each customer ranges from <strong>1</strong> to <strong>13</strong>. (<em>Train Data</em>)<br>\nAs shown in the figure below, you can find the fact.</p>\n<p><em>Left: Customers with 13-full statements account for a significant proportion of total customers.</em><br>\n<em>Right: There's a difference in ratio between <strong>13-full</strong> and <strong>the rest ([12-1])</strong>.</em></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5146692%2F29bd2d67e5332064884dea2123bca278%2Ffig_1.png?generation=1657396043953058&amp;alt=media\" alt=\"\"></p>\n<p>So, customers with <strong>fewer statements</strong> tend to have <strong>a higher default probability</strong>.</p>\n<p>See <a href=\"https://www.kaggle.com/code/shusky/amex-relation-b-w-number-of-statements-default\" target=\"_blank\">this notebook</a> for details!<br>\nThank you :D</p>",
  "messages": [
    {
      "id": "1849823",
      "postDate": "07/09/2022 20:35:56",
      "content": "<p>As you may have noticed,<br>\nthere's a relation between the number of each customer's statements and its probability of default.</p>\n<p>In general, customers who continuously use credit cards are considered to be more creditworthy.<br>\nTherefore, it can be assumed that the default probability is relatively higher for customers who use less frequently.</p>\n<p>In this competition, the number of statements for each customer ranges from <strong>1</strong> to <strong>13</strong>. (<em>Train Data</em>)<br>\nAs shown in the figure below, you can find the fact.</p>\n<p><em>Left: Customers with 13-full statements account for a significant proportion of total customers.</em><br>\n<em>Right: There's a difference in ratio between <strong>13-full</strong> and <strong>the rest ([12-1])</strong>.</em></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5146692%2F29bd2d67e5332064884dea2123bca278%2Ffig_1.png?generation=1657396043953058&amp;alt=media\" alt=\"\"></p>\n<p>So, customers with <strong>fewer statements</strong> tend to have <strong>a higher default probability</strong>.</p>\n<p>See <a href=\"https://www.kaggle.com/code/shusky/amex-relation-b-w-number-of-statements-default\" target=\"_blank\">this notebook</a> for details!<br>\nThank you :D</p>",
      "rawMarkdown": "As you may have noticed,\nthere's a relation between the number of each customer's statements and its probability of default.\n\nIn general, customers who continuously use credit cards are considered to be more creditworthy.\nTherefore, it can be assumed that the default probability is relatively higher for customers who use less frequently.\n\nIn this competition, the number of statements for each customer ranges from **1** to **13**. (*Train Data*)\nAs shown in the figure below, you can find the fact.\n\n*Left: Customers with 13-full statements account for a significant proportion of total customers.*\n*Right: There's a difference in ratio between **13-full** and **the rest ([12-1])**.*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5146692%2F29bd2d67e5332064884dea2123bca278%2Ffig_1.png?generation=1657396043953058&alt=media)\n\nSo, customers with **fewer statements** tend to have **a higher default probability**.\n\nSee [this notebook](https://www.kaggle.com/code/shusky/amex-relation-b-w-number-of-statements-default) for details!\nThank you :D",
      "votes": null
    },
    {
      "id": "1850202",
      "postDate": "07/10/2022 07:38:11",
      "content": "<p>Thanks for sharing your work. This is great! what do you think about the number of events in each category? for example, those who have only one record are minority and might not have enough data to conclude that they have higher probability of default. In addition, the right bar chart does not prove your claim. Can you elaborate more on how you support your claim?</p>",
      "rawMarkdown": "Thanks for sharing your work. This is great! what do you think about the number of events in each category? for example, those who have only one record are minority and might not have enough data to conclude that they have higher probability of default. In addition, the right bar chart does not prove your claim. Can you elaborate more on how you support your claim?",
      "votes": null
    },
    {
      "id": "1850519",
      "postDate": "07/10/2022 13:18:01",
      "content": "<p>Fewer statements could mean younger accounts…</p>",
      "rawMarkdown": "Fewer statements could mean younger accounts...",
      "votes": null
    },
    {
      "id": "1850600",
      "postDate": "07/10/2022 14:21:04",
      "content": "<p>Hi,<br>\nI noticed the same thing, I don't have a graph/plot of it but, here is the summary of what I found.<br>\nCUSTOMER ROWS 1 DEFAULT 0.34 %<br>\nCUSTOMER ROWS 2 DEFAULT 0.32 %<br>\nCUSTOMER ROWS 3 DEFAULT 0.36 %<br>\nCUSTOMER ROWS 4 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 5 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 6 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 7 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 8 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 9 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 10 DEFAULT 0.46 %<br>\nCUSTOMER ROWS 11 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 12 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 13 DEFAULT 0.23 %</p>\n<p>customers with the full 13 rows have around 19% of the entire defaults in the dataset, and customers with less than 13 rows have around 6% in total. I tried a little experiment with hyper parameter tuning and noticed you could get around +0.790 with customers with 13 rows, and around +0.600 with customers with less than 13 rows. Was it harder to predict the customers with less than 13 rows as well?</p>",
      "rawMarkdown": "Hi,\nI noticed the same thing, I don't have a graph/plot of it but, here is the summary of what I found.\nCUSTOMER ROWS 1 DEFAULT 0.34 %\nCUSTOMER ROWS 2 DEFAULT 0.32 %\nCUSTOMER ROWS 3 DEFAULT 0.36 %\nCUSTOMER ROWS 4 DEFAULT 0.42 %\nCUSTOMER ROWS 5 DEFAULT 0.39 %\nCUSTOMER ROWS 6 DEFAULT 0.39 %\nCUSTOMER ROWS 7 DEFAULT 0.42 %\nCUSTOMER ROWS 8 DEFAULT 0.45 %\nCUSTOMER ROWS 9 DEFAULT 0.45 %\nCUSTOMER ROWS 10 DEFAULT 0.46 %\nCUSTOMER ROWS 11 DEFAULT 0.45 %\nCUSTOMER ROWS 12 DEFAULT 0.39 %\nCUSTOMER ROWS 13 DEFAULT 0.23 %\n\ncustomers with the full 13 rows have around 19% of the entire defaults in the dataset, and customers with less than 13 rows have around 6% in total. I tried a little experiment with hyper parameter tuning and noticed you could get around +0.790 with customers with 13 rows, and around +0.600 with customers with less than 13 rows. Was it harder to predict the customers with less than 13 rows as well?",
      "votes": null
    },
    {
      "id": "1852675",
      "postDate": "07/12/2022 08:36:27",
      "content": "<p>Agree. Indeed you can confirm they are younger by looking at the delinquency variables. I don’t remember which one, but one of them goes up by a static increment each month, occasionally resetting to zero (presumably on default). Thus it’s a proxy for age of customer and allows you to separate new customers from customers who happen to be missing some statements.</p>",
      "rawMarkdown": "Agree. Indeed you can confirm they are younger by looking at the delinquency variables. I don’t remember which one, but one of them goes up by a static increment each month, occasionally resetting to zero (presumably on default). Thus it’s a proxy for age of customer and allows you to separate new customers from customers who happen to be missing some statements.",
      "votes": null
    },
    {
      "id": "1853632",
      "postDate": "07/13/2022 03:03:49",
      "content": "<p>Yeah, it's interesting! I posted about it awhile back, too lazy to find it on my phone though ;)</p>\n<p>Note that if they have the same min month, but fewer total statements (\"gap\" customers), they have barely higher risk than full 13 statement customers, and much lower than customers with a newer \"oldest\" statement. </p>\n<p>I also don't understand why (or if insignificant, or - most likely? - if related to US economic periods) but customers with fewer than 13 statements were more risk if they were closer to 13 rather than if closer to zero. </p>",
      "rawMarkdown": "Yeah, it's interesting! I posted about it awhile back, too lazy to find it on my phone though ;)\n\nNote that if they have the same min month, but fewer total statements (\"gap\" customers), they have barely higher risk than full 13 statement customers, and much lower than customers with a newer \"oldest\" statement. \n\nI also don't understand why (or if insignificant, or - most likely? - if related to US economic periods) but customers with fewer than 13 statements were more risk if they were closer to 13 rather than if closer to zero.",
      "votes": null
    },
    {
      "id": "1853709",
      "postDate": "07/13/2022 04:54:58",
      "content": "<p><a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a> Interesting! I'll have to see if I can find that, because of the oddity of 10-12 month long customers having higher default rate than 1-9 month, I was wondering if it increases with age up to a point in the 12-18 month range, then decreases if even older than that point (even bad customers need time before maxing out?)</p>",
      "rawMarkdown": "burritodan Interesting! I'll have to see if I can find that, because of the oddity of 10-12 month long customers having higher default rate than 1-9 month, I was wondering if it increases with age up to a point in the 12-18 month range, then decreases if even older than that point (even bad customers need time before maxing out?)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1850202,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/10/2022 07:38:11",
      "content": "<p>Thanks for sharing your work. This is great! what do you think about the number of events in each category? for example, those who have only one record are minority and might not have enough data to conclude that they have higher probability of default. In addition, the right bar chart does not prove your claim. Can you elaborate more on how you support your claim?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1850519,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "07/10/2022 13:18:01",
      "content": "<p>Fewer statements could mean younger accounts…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1852675,
          "author_name": "burritodan",
          "author_url": "",
          "post_date": "07/12/2022 08:36:27",
          "content": "<p>Agree. Indeed you can confirm they are younger by looking at the delinquency variables. I don’t remember which one, but one of them goes up by a static increment each month, occasionally resetting to zero (presumably on default). Thus it’s a proxy for age of customer and allows you to separate new customers from customers who happen to be missing some statements.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1853709,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/13/2022 04:54:58",
          "content": "<p><a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a> Interesting! I'll have to see if I can find that, because of the oddity of 10-12 month long customers having higher default rate than 1-9 month, I was wondering if it increases with age up to a point in the 12-18 month range, then decreases if even older than that point (even bad customers need time before maxing out?)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1850600,
      "author_name": "tarrasque9",
      "author_url": "",
      "post_date": "07/10/2022 14:21:04",
      "content": "<p>Hi,<br>\nI noticed the same thing, I don't have a graph/plot of it but, here is the summary of what I found.<br>\nCUSTOMER ROWS 1 DEFAULT 0.34 %<br>\nCUSTOMER ROWS 2 DEFAULT 0.32 %<br>\nCUSTOMER ROWS 3 DEFAULT 0.36 %<br>\nCUSTOMER ROWS 4 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 5 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 6 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 7 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 8 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 9 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 10 DEFAULT 0.46 %<br>\nCUSTOMER ROWS 11 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 12 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 13 DEFAULT 0.23 %</p>\n<p>customers with the full 13 rows have around 19% of the entire defaults in the dataset, and customers with less than 13 rows have around 6% in total. I tried a little experiment with hyper parameter tuning and noticed you could get around +0.790 with customers with 13 rows, and around +0.600 with customers with less than 13 rows. Was it harder to predict the customers with less than 13 rows as well?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1853632,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/13/2022 03:03:49",
      "content": "<p>Yeah, it's interesting! I posted about it awhile back, too lazy to find it on my phone though ;)</p>\n<p>Note that if they have the same min month, but fewer total statements (\"gap\" customers), they have barely higher risk than full 13 statement customers, and much lower than customers with a newer \"oldest\" statement. </p>\n<p>I also don't understand why (or if insignificant, or - most likely? - if related to US economic periods) but customers with fewer than 13 statements were more risk if they were closer to 13 rather than if closer to zero. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1849823": "As you may have noticed,\nthere's a relation between the number of each customer's statements and its probability of default.\n\nIn general, customers who continuously use credit cards are considered to be more creditworthy.\nTherefore, it can be assumed that the default probability is relatively higher for customers who use less frequently.\n\nIn this competition, the number of statements for each customer ranges from **1** to **13**. (*Train Data*)\nAs shown in the figure below, you can find the fact.\n\n*Left: Customers with 13-full statements account for a significant proportion of total customers.*\n*Right: There's a difference in ratio between **13-full** and **the rest ([12-1])**.*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5146692%2F29bd2d67e5332064884dea2123bca278%2Ffig_1.png?generation=1657396043953058&alt=media)\n\nSo, customers with **fewer statements** tend to have **a higher default probability**.\n\nSee [this notebook](https://www.kaggle.com/code/shusky/amex-relation-b-w-number-of-statements-default) for details!\nThank you :D",
    "1850202": "Thanks for sharing your work. This is great! what do you think about the number of events in each category? for example, those who have only one record are minority and might not have enough data to conclude that they have higher probability of default. In addition, the right bar chart does not prove your claim. Can you elaborate more on how you support your claim?",
    "1850519": "Fewer statements could mean younger accounts...",
    "1850600": "Hi,\nI noticed the same thing, I don't have a graph/plot of it but, here is the summary of what I found.\nCUSTOMER ROWS 1 DEFAULT 0.34 %\nCUSTOMER ROWS 2 DEFAULT 0.32 %\nCUSTOMER ROWS 3 DEFAULT 0.36 %\nCUSTOMER ROWS 4 DEFAULT 0.42 %\nCUSTOMER ROWS 5 DEFAULT 0.39 %\nCUSTOMER ROWS 6 DEFAULT 0.39 %\nCUSTOMER ROWS 7 DEFAULT 0.42 %\nCUSTOMER ROWS 8 DEFAULT 0.45 %\nCUSTOMER ROWS 9 DEFAULT 0.45 %\nCUSTOMER ROWS 10 DEFAULT 0.46 %\nCUSTOMER ROWS 11 DEFAULT 0.45 %\nCUSTOMER ROWS 12 DEFAULT 0.39 %\nCUSTOMER ROWS 13 DEFAULT 0.23 %\n\ncustomers with the full 13 rows have around 19% of the entire defaults in the dataset, and customers with less than 13 rows have around 6% in total. I tried a little experiment with hyper parameter tuning and noticed you could get around +0.790 with customers with 13 rows, and around +0.600 with customers with less than 13 rows. Was it harder to predict the customers with less than 13 rows as well?",
    "1852675": "Agree. Indeed you can confirm they are younger by looking at the delinquency variables. I don’t remember which one, but one of them goes up by a static increment each month, occasionally resetting to zero (presumably on default). Thus it’s a proxy for age of customer and allows you to separate new customers from customers who happen to be missing some statements.",
    "1853632": "Yeah, it's interesting! I posted about it awhile back, too lazy to find it on my phone though ;)\n\nNote that if they have the same min month, but fewer total statements (\"gap\" customers), they have barely higher risk than full 13 statement customers, and much lower than customers with a newer \"oldest\" statement. \n\nI also don't understand why (or if insignificant, or - most likely? - if related to US economic periods) but customers with fewer than 13 statements were more risk if they were closer to 13 rather than if closer to zero.",
    "1853709": "burritodan Interesting! I'll have to see if I can find that, because of the oddity of 10-12 month long customers having higher default rate than 1-9 month, I was wondering if it increases with age up to a point in the 12-18 month range, then decreases if even older than that point (even bad customers need time before maxing out?)"
  },
  "source": "meta"
}