{
  "id": 338121,
  "title": "📆 Last statement dates (S_2): default rate seasonality in train set March 2018",
  "url": "/competitions/amex-default-prediction/discussion/338121",
  "author_name": "",
  "post_date": "2022-07-19T09:20:32.318738800Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all!<br>\nI just finished check of 4 Hypothesis for feature eng. from statement dates and found out a bit strange pattern for # of statements and default rate for March 2018. <br>\nIe - strong intra-week seasonality for # of statements and variation in default rate by last statement date for the train period.<br>\nIn addition to that, I check default rate stationarity (assuming <strong>y=default rate</strong> is a time series) with Augmented Dickey-Fuller test and it didn't find any trend component with significance level &lt;1%. So we might have only seasonal components. \"Might\" - as it's still one month only data with 31 data points(</p>\n<p>All of this raises a questions:<br>\n<strong>1)</strong> How exactly was the train set compiled for March 2018? Is it some kind of segment based resampling (by lifetime, product etc) or it's a natural variation based on AMEX sales pattern (most of statement dates derived from sales/card activation date)?<br>\n<strong>2)</strong> How stable is this resampling across March 2018, April 2019 (Public LB) and October 2018 (Private LB)?</p>\n<p>Public notebook with a code: <a href=\"https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use/notebook\" target=\"_blank\">https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use/notebook</a></p>",
  "messages": [
    {
      "id": "1861833",
      "postDate": "07/19/2022 09:20:32",
      "content": "<p>Hi all!<br>\nI just finished check of 4 Hypothesis for feature eng. from statement dates and found out a bit strange pattern for # of statements and default rate for March 2018. <br>\nIe - strong intra-week seasonality for # of statements and variation in default rate by last statement date for the train period.<br>\nIn addition to that, I check default rate stationarity (assuming <strong>y=default rate</strong> is a time series) with Augmented Dickey-Fuller test and it didn't find any trend component with significance level &lt;1%. So we might have only seasonal components. \"Might\" - as it's still one month only data with 31 data points(</p>\n<p>All of this raises a questions:<br>\n<strong>1)</strong> How exactly was the train set compiled for March 2018? Is it some kind of segment based resampling (by lifetime, product etc) or it's a natural variation based on AMEX sales pattern (most of statement dates derived from sales/card activation date)?<br>\n<strong>2)</strong> How stable is this resampling across March 2018, April 2019 (Public LB) and October 2018 (Private LB)?</p>\n<p>Public notebook with a code: <a href=\"https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use/notebook\" target=\"_blank\">https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use/notebook</a></p>",
      "rawMarkdown": "Hi all!\nI just finished check of 4 Hypothesis for feature eng. from statement dates and found out a bit strange pattern for # of statements and default rate for March 2018. \nIe - strong intra-week seasonality for # of statements and variation in default rate by last statement date for the train period.\nIn addition to that, I check default rate stationarity (assuming **y=default rate** is a time series) with Augmented Dickey-Fuller test and it didn't find any trend component with significance level <1%. So we might have only seasonal components. \"Might\" - as it's still one month only data with 31 data points(\n\nAll of this raises a questions:\n**1)** How exactly was the train set compiled for March 2018? Is it some kind of segment based resampling (by lifetime, product etc) or it's a natural variation based on AMEX sales pattern (most of statement dates derived from sales/card activation date)?\n**2)** How stable is this resampling across March 2018, April 2019 (Public LB) and October 2018 (Private LB)?\n\nPublic notebook with a code: https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use/notebook",
      "votes": null
    },
    {
      "id": "1861994",
      "postDate": "07/19/2022 11:32:09",
      "content": "<p>I believe it's a deliberate mix of existing customers and new customers. See <code>D_45</code> and <code>D_47</code> for hints about the customer's age.</p>",
      "rawMarkdown": "I believe it's a deliberate mix of existing customers and new customers. See `D_45` and `D_47` for hints about the customer's age.",
      "votes": null
    },
    {
      "id": "1864617",
      "postDate": "07/21/2022 07:35:56",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a> </p>\n<p>I am intrigued by your observation that <code>D_45</code> and <code>D_47</code> are related to the customer's age. Could I be so bold as to enquire as to your reasoning?</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @burritodan \n\nI am intrigued by your observation that `D_45` and `D_47` are related to the customer's age. Could I be so bold as to enquire as to your reasoning?\n\nAll the best,\ncarl",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1861994,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "07/19/2022 11:32:09",
      "content": "<p>I believe it's a deliberate mix of existing customers and new customers. See <code>D_45</code> and <code>D_47</code> for hints about the customer's age.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1864617,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "07/21/2022 07:35:56",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a> </p>\n<p>I am intrigued by your observation that <code>D_45</code> and <code>D_47</code> are related to the customer's age. Could I be so bold as to enquire as to your reasoning?</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1861833": "Hi all!\nI just finished check of 4 Hypothesis for feature eng. from statement dates and found out a bit strange pattern for # of statements and default rate for March 2018. \nIe - strong intra-week seasonality for # of statements and variation in default rate by last statement date for the train period.\nIn addition to that, I check default rate stationarity (assuming **y=default rate** is a time series) with Augmented Dickey-Fuller test and it didn't find any trend component with significance level <1%. So we might have only seasonal components. \"Might\" - as it's still one month only data with 31 data points(\n\nAll of this raises a questions:\n**1)** How exactly was the train set compiled for March 2018? Is it some kind of segment based resampling (by lifetime, product etc) or it's a natural variation based on AMEX sales pattern (most of statement dates derived from sales/card activation date)?\n**2)** How stable is this resampling across March 2018, April 2019 (Public LB) and October 2018 (Private LB)?\n\nPublic notebook with a code: https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use/notebook",
    "1861994": "I believe it's a deliberate mix of existing customers and new customers. See `D_45` and `D_47` for hints about the customer's age.",
    "1864617": "Dear @burritodan \n\nI am intrigued by your observation that `D_45` and `D_47` are related to the customer's age. Could I be so bold as to enquire as to your reasoning?\n\nAll the best,\ncarl"
  },
  "source": "meta"
}