{
  "id": 337329,
  "title": "Feature Engineering - Inverse relationship between number of samples at one timestamp versus the overall default rate",
  "url": "/competitions/amex-default-prediction/discussion/337329",
  "author_name": "",
  "post_date": "2022-07-15T14:10:09.670191600Z",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Fig1 is a chart of the number of data points at each timestamp and the mean of the target at that same time stamp. Although the target mean fluctuates around 0.25 (as the detailed in the metric description about upsampling) it does bounce around, also seems like the less samples gathered the higher the percentage of defaults for that observation. The correlation between these two variables are -36%!</p>\n<p>Kind of makes me wonder what is going on here - I think this might be a result of the upsampling where they need to artificially gather more data on a certain timestamp to meet that threshold. This can be used for feature engineering as we see that in the testing dataset the number of observations at each time stamp also exhibits a periodical pattern (see fig2).</p>\n<p>Just thought this was kind of interesting, and please let me know if you have any thoughts on this!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2F1e4602c4cad2e6c2e5ffa597be723556%2Ftrain_observations_vs_target_avg.png?generation=1657893670829571&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2Ff536ff745f6645bb97c85939f5a61e8b%2Ftesting_observation_count.png?generation=1657894193824962&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1856691",
      "postDate": "07/15/2022 14:10:09",
      "content": "<p>Fig1 is a chart of the number of data points at each timestamp and the mean of the target at that same time stamp. Although the target mean fluctuates around 0.25 (as the detailed in the metric description about upsampling) it does bounce around, also seems like the less samples gathered the higher the percentage of defaults for that observation. The correlation between these two variables are -36%!</p>\n<p>Kind of makes me wonder what is going on here - I think this might be a result of the upsampling where they need to artificially gather more data on a certain timestamp to meet that threshold. This can be used for feature engineering as we see that in the testing dataset the number of observations at each time stamp also exhibits a periodical pattern (see fig2).</p>\n<p>Just thought this was kind of interesting, and please let me know if you have any thoughts on this!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2F1e4602c4cad2e6c2e5ffa597be723556%2Ftrain_observations_vs_target_avg.png?generation=1657893670829571&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2Ff536ff745f6645bb97c85939f5a61e8b%2Ftesting_observation_count.png?generation=1657894193824962&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Fig1 is a chart of the number of data points at each timestamp and the mean of the target at that same time stamp. Although the target mean fluctuates around 0.25 (as the detailed in the metric description about upsampling) it does bounce around, also seems like the less samples gathered the higher the percentage of defaults for that observation. The correlation between these two variables are -36%!\n\nKind of makes me wonder what is going on here - I think this might be a result of the upsampling where they need to artificially gather more data on a certain timestamp to meet that threshold. This can be used for feature engineering as we see that in the testing dataset the number of observations at each time stamp also exhibits a periodical pattern (see fig2).\n\nJust thought this was kind of interesting, and please let me know if you have any thoughts on this!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2F1e4602c4cad2e6c2e5ffa597be723556%2Ftrain_observations_vs_target_avg.png?generation=1657893670829571&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2Ff536ff745f6645bb97c85939f5a61e8b%2Ftesting_observation_count.png?generation=1657894193824962&alt=media)",
      "votes": null
    },
    {
      "id": "1856946",
      "postDate": "07/15/2022 18:18:21",
      "content": "<p>Great observations! By the way, I think weekday is very important in this research because Sundays have the least observations and the highest average target. Also, if we do the same research but look at one day a week, we have the same result (negative correlation), except Sunday (+25%)</p>",
      "rawMarkdown": "Great observations! By the way, I think weekday is very important in this research because Sundays have the least observations and the highest average target. Also, if we do the same research but look at one day a week, we have the same result (negative correlation), except Sunday (+25%)",
      "votes": null
    },
    {
      "id": "1857551",
      "postDate": "07/16/2022 07:37:36",
      "content": "<p>It's because theres a deliberate mix of existing and new customers (whose first statement is after the start of the observation period). You can tell they're new customers by looking at some of the deliquency variables <code>D_XX</code>. I don't remember which one of the top of my head, but there is on which goes up by a fixed amount each month, which gives you a good idea how old the customer is.</p>",
      "rawMarkdown": "It's because theres a deliberate mix of existing and new customers (whose first statement is after the start of the observation period). You can tell they're new customers by looking at some of the deliquency variables `D_XX`. I don't remember which one of the top of my head, but there is on which goes up by a fixed amount each month, which gives you a good idea how old the customer is.",
      "votes": null
    },
    {
      "id": "1857733",
      "postDate": "07/16/2022 10:38:03",
      "content": "<p>Thanks for the note. Let me look into the data/posts regarding this. Appreciate it.</p>",
      "rawMarkdown": "Thanks for the note. Let me look into the data/posts regarding this. Appreciate it.",
      "votes": null
    },
    {
      "id": "1857734",
      "postDate": "07/16/2022 10:38:41",
      "content": "<p>Ah thanks for the observation! I'll make sure to look into that as well!</p>",
      "rawMarkdown": "Ah thanks for the observation! I'll make sure to look into that as well!",
      "votes": null
    },
    {
      "id": "1885816",
      "postDate": "08/05/2022 11:46:41",
      "content": "<p>This is because customers with 13 statements have higher default rate. </p>",
      "rawMarkdown": "This is because customers with 13 statements have higher default rate.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1856946,
      "author_name": "olehkondratenko",
      "author_url": "",
      "post_date": "07/15/2022 18:18:21",
      "content": "<p>Great observations! By the way, I think weekday is very important in this research because Sundays have the least observations and the highest average target. Also, if we do the same research but look at one day a week, we have the same result (negative correlation), except Sunday (+25%)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1857734,
          "author_name": "translucen7",
          "author_url": "",
          "post_date": "07/16/2022 10:38:41",
          "content": "<p>Ah thanks for the observation! I'll make sure to look into that as well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1857551,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "07/16/2022 07:37:36",
      "content": "<p>It's because theres a deliberate mix of existing and new customers (whose first statement is after the start of the observation period). You can tell they're new customers by looking at some of the deliquency variables <code>D_XX</code>. I don't remember which one of the top of my head, but there is on which goes up by a fixed amount each month, which gives you a good idea how old the customer is.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1857733,
          "author_name": "translucen7",
          "author_url": "",
          "post_date": "07/16/2022 10:38:03",
          "content": "<p>Thanks for the note. Let me look into the data/posts regarding this. Appreciate it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1885816,
      "author_name": "evan918",
      "author_url": "",
      "post_date": "08/05/2022 11:46:41",
      "content": "<p>This is because customers with 13 statements have higher default rate. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1856691": "Fig1 is a chart of the number of data points at each timestamp and the mean of the target at that same time stamp. Although the target mean fluctuates around 0.25 (as the detailed in the metric description about upsampling) it does bounce around, also seems like the less samples gathered the higher the percentage of defaults for that observation. The correlation between these two variables are -36%!\n\nKind of makes me wonder what is going on here - I think this might be a result of the upsampling where they need to artificially gather more data on a certain timestamp to meet that threshold. This can be used for feature engineering as we see that in the testing dataset the number of observations at each time stamp also exhibits a periodical pattern (see fig2).\n\nJust thought this was kind of interesting, and please let me know if you have any thoughts on this!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2F1e4602c4cad2e6c2e5ffa597be723556%2Ftrain_observations_vs_target_avg.png?generation=1657893670829571&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1068837%2Ff536ff745f6645bb97c85939f5a61e8b%2Ftesting_observation_count.png?generation=1657894193824962&alt=media)",
    "1856946": "Great observations! By the way, I think weekday is very important in this research because Sundays have the least observations and the highest average target. Also, if we do the same research but look at one day a week, we have the same result (negative correlation), except Sunday (+25%)",
    "1857551": "It's because theres a deliberate mix of existing and new customers (whose first statement is after the start of the observation period). You can tell they're new customers by looking at some of the deliquency variables `D_XX`. I don't remember which one of the top of my head, but there is on which goes up by a fixed amount each month, which gives you a good idea how old the customer is.",
    "1857733": "Thanks for the note. Let me look into the data/posts regarding this. Appreciate it.",
    "1857734": "Ah thanks for the observation! I'll make sure to look into that as well!",
    "1885816": "This is because customers with 13 statements have higher default rate."
  },
  "source": "meta"
}