{
  "id": 336714,
  "title": "Variable Explanation",
  "url": "/competitions/amex-default-prediction/discussion/336714",
  "author_name": "",
  "post_date": "2022-07-12T16:43:48.717538500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I was wondering if anyone knew what are the definitions of these variables mean? For instance, P_2 is the biggest driving factor for the predictive model but does anyone know what P_2 is suppose to mean? Any sorts of data dictionary would be great as well.</p>\n<p>Thank you</p>\n<p>-Mudassir Ali</p>",
  "messages": [
    {
      "id": "1853154",
      "postDate": "07/12/2022 16:43:48",
      "content": "<p>Hello,</p>\n<p>I was wondering if anyone knew what are the definitions of these variables mean? For instance, P_2 is the biggest driving factor for the predictive model but does anyone know what P_2 is suppose to mean? Any sorts of data dictionary would be great as well.</p>\n<p>Thank you</p>\n<p>-Mudassir Ali</p>",
      "rawMarkdown": "Hello,\n\nI was wondering if anyone knew what are the definitions of these variables mean? For instance, P_2 is the biggest driving factor for the predictive model but does anyone know what P_2 is suppose to mean? Any sorts of data dictionary would be great as well.\n\nThank you\n\n-Mudassir Ali",
      "votes": null
    },
    {
      "id": "1853318",
      "postDate": "07/12/2022 19:56:32",
      "content": "<p>In the Data tab, the host states:</p>\n<blockquote>\n  <p>The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:</p>\n  <p>D_* = Delinquency variables<br>\n  S_* = Spend variables<br>\n  P_* = Payment variables<br>\n  B_* = Balance variables<br>\n  R_* = Risk variables</p>\n</blockquote>\n<p>This means we are not supposed to know what the features are exactly.</p>",
      "rawMarkdown": "In the Data tab, the host states:\n\n> The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n> \n> D_* = Delinquency variables\n> S_* = Spend variables\n> P_* = Payment variables\n> B_* = Balance variables\n> R_* = Risk variables\n\nThis means we are not supposed to know what the features are exactly.",
      "votes": null
    },
    {
      "id": "1853626",
      "postDate": "07/13/2022 02:48:30",
      "content": "<p>See this thread <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332574\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332574</a><br>\nraddar talks about P_2 &amp; D_39</p>",
      "rawMarkdown": "See this thread https://www.kaggle.com/competitions/amex-default-prediction/discussion/332574\nraddar talks about P_2 & D_39",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1853318,
      "author_name": "fritzcremer",
      "author_url": "",
      "post_date": "07/12/2022 19:56:32",
      "content": "<p>In the Data tab, the host states:</p>\n<blockquote>\n  <p>The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:</p>\n  <p>D_* = Delinquency variables<br>\n  S_* = Spend variables<br>\n  P_* = Payment variables<br>\n  B_* = Balance variables<br>\n  R_* = Risk variables</p>\n</blockquote>\n<p>This means we are not supposed to know what the features are exactly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1853626,
      "author_name": "hamonk",
      "author_url": "",
      "post_date": "07/13/2022 02:48:30",
      "content": "<p>See this thread <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332574\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332574</a><br>\nraddar talks about P_2 &amp; D_39</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1853154": "Hello,\n\nI was wondering if anyone knew what are the definitions of these variables mean? For instance, P_2 is the biggest driving factor for the predictive model but does anyone know what P_2 is suppose to mean? Any sorts of data dictionary would be great as well.\n\nThank you\n\n-Mudassir Ali",
    "1853318": "In the Data tab, the host states:\n\n> The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n> \n> D_* = Delinquency variables\n> S_* = Spend variables\n> P_* = Payment variables\n> B_* = Balance variables\n> R_* = Risk variables\n\nThis means we are not supposed to know what the features are exactly.",
    "1853626": "See this thread https://www.kaggle.com/competitions/amex-default-prediction/discussion/332574\nraddar talks about P_2 & D_39"
  },
  "source": "meta"
}