{
  "id": 327367,
  "title": "About the nature of the target",
  "url": "/competitions/amex-default-prediction/discussion/327367",
  "author_name": "mavillan",
  "post_date": "2022-05-26T22:24:04.422000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I'm a bit confused about how the target was created. If you run this, </p>\n<pre><code>pd.merge(train.query(\"target == 1\")[[\"customer_ID\"]], train.query(\"target == 0\")[[\"customer_ID\"]])\n</code></pre>\n<p>you get 0 intersection, it means that <code>customer_ID</code>s are labeled as defaulter (o not) for the entire period of the dataset. </p>\n<p>So, even if for some months a customer didn't commited default, he will be labeled as defaulter anyways. Is that right? Or I'm missing something.  </p>",
  "messages": [
    {
      "id": 1802560,
      "postDate": "2022-05-26T22:24:04.423Z",
      "content": "<p>Hi all,</p>\n<p>I'm a bit confused about how the target was created. If you run this, </p>\n<pre><code>pd.merge(train.query(\"target == 1\")[[\"customer_ID\"]], train.query(\"target == 0\")[[\"customer_ID\"]])\n</code></pre>\n<p>you get 0 intersection, it means that <code>customer_ID</code>s are labeled as defaulter (o not) for the entire period of the dataset. </p>\n<p>So, even if for some months a customer didn't commited default, he will be labeled as defaulter anyways. Is that right? Or I'm missing something.  </p>",
      "rawMarkdown": "Hi all,\n\nI'm a bit confused about how the target was created. If you run this, \n\n```python\r\npd.merge(train.query(\"target == 1\")[[\"customer_ID\"]], train.query(\"target == 0\")[[\"customer_ID\"]])\n```\n\nyou get 0 intersection, it means that `customer_ID`s are labeled as defaulter (o not) for the entire period of the dataset. \n\nSo, even if for some months a customer didn't commited default, he will be labeled as defaulter anyways. Is that right? Or I'm missing something.  ",
      "votes": 6
    },
    {
      "id": 1802661,
      "postDate": "2022-05-27T02:37:40Z",
      "content": "<p>In one of the discussion topics someone with domain knowledge indicated that default will typically occur after 120 days or a similar period since the customer last made a payment.  So somewhere in the customer's history they stopped paying - would still gets several months of statements afterwards until declared as defaulter.  Need to look at data to see if statements continue after a default is declared  - I seem to recall in someones EDA that statements probably continue in the data set even if the default was related to the earliest time frame statement.</p>\n<p>I assume if we look closely at the data we might be able figure out the last statement where they paid the bill.</p>",
      "rawMarkdown": "In one of the discussion topics someone with domain knowledge indicated that default will typically occur after 120 days or a similar period since the customer last made a payment.  So somewhere in the customer's history they stopped paying - would still gets several months of statements afterwards until declared as defaulter.  Need to look at data to see if statements continue after a default is declared  - I seem to recall in someones EDA that statements probably continue in the data set even if the default was related to the earliest time frame statement.\n\nI assume if we look closely at the data we might be able figure out the last statement where they paid the bill.",
      "votes": 1
    },
    {
      "id": 1802900,
      "postDate": "2022-05-27T09:39:44.150Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1802661,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2022-05-27T02:37:40",
      "content": "<p>In one of the discussion topics someone with domain knowledge indicated that default will typically occur after 120 days or a similar period since the customer last made a payment.  So somewhere in the customer's history they stopped paying - would still gets several months of statements afterwards until declared as defaulter.  Need to look at data to see if statements continue after a default is declared  - I seem to recall in someones EDA that statements probably continue in the data set even if the default was related to the earliest time frame statement.</p>\n<p>I assume if we look closely at the data we might be able figure out the last statement where they paid the bill.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1802900,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-27T09:39:44.150000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1802560": "Hi all,\n\nI'm a bit confused about how the target was created. If you run this, \n\n```python\r\npd.merge(train.query(\"target == 1\")[[\"customer_ID\"]], train.query(\"target == 0\")[[\"customer_ID\"]])\n```\n\nyou get 0 intersection, it means that `customer_ID`s are labeled as defaulter (o not) for the entire period of the dataset. \n\nSo, even if for some months a customer didn't commited default, he will be labeled as defaulter anyways. Is that right? Or I'm missing something.  ",
    "1802661": "In one of the discussion topics someone with domain knowledge indicated that default will typically occur after 120 days or a similar period since the customer last made a payment.  So somewhere in the customer's history they stopped paying - would still gets several months of statements afterwards until declared as defaulter.  Need to look at data to see if statements continue after a default is declared  - I seem to recall in someones EDA that statements probably continue in the data set even if the default was related to the earliest time frame statement.\n\nI assume if we look closely at the data we might be able figure out the last statement where they paid the bill.",
    "1802900": ""
  }
}