{
  "id": 335998,
  "title": "No customer falls into default within the 13-month period",
  "url": "/competitions/amex-default-prediction/discussion/335998",
  "author_name": "",
  "post_date": "2022-07-08T20:32:11.155136200Z",
  "votes": -5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Either the customer is in default or it isn´t. No 0's and 1's for the same person.</p>\n<p>Here is a little code to prove it.</p>\n<p>`D = pd.DataFrame(train, columns =[\"customer_ID\", \"target\"])<br>\nD = D.loc[D[\"target\"] == 1]</p>\n<p>ysum = D.groupby(\"customer_ID\", as_index=False).sum()<br>\nysize = D.groupby(\"customer_ID\",as_index=False).size()</p>\n<p>ysum['target'].equals(ysize['size'])`</p>",
  "messages": [
    {
      "id": "1848692",
      "postDate": "07/08/2022 20:32:11",
      "content": "<p>Either the customer is in default or it isn´t. No 0's and 1's for the same person.</p>\n<p>Here is a little code to prove it.</p>\n<p>`D = pd.DataFrame(train, columns =[\"customer_ID\", \"target\"])<br>\nD = D.loc[D[\"target\"] == 1]</p>\n<p>ysum = D.groupby(\"customer_ID\", as_index=False).sum()<br>\nysize = D.groupby(\"customer_ID\",as_index=False).size()</p>\n<p>ysum['target'].equals(ysize['size'])`</p>",
      "rawMarkdown": "Either the customer is in default or it isn´t. No 0's and 1's for the same person.\n\nHere is a little code to prove it.\n\n`D = pd.DataFrame(train, columns =[\"customer_ID\", \"target\"])\nD = D.loc[D[\"target\"] == 1]\n\nysum = D.groupby(\"customer_ID\", as_index=False).sum()\nysize = D.groupby(\"customer_ID\",as_index=False).size()\n\nysum['target'].equals(ysize['size'])`",
      "votes": null
    },
    {
      "id": "1848748",
      "postDate": "07/08/2022 21:27:33",
      "content": "<p>There is only a single target label for each customer which observes the 18 months following their latest credit card statement in the data to see if they are ever past due by 120 days</p>\n<p>train_labels.csv - target label for each customer_ID</p>\n<p>However, if you did want to use more rows than just the last one, I think this notebook by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> explored the possibility of using D_39 variable to assign labels on the intermediate rows for the customers</p>\n<p><a href=\"https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex\" target=\"_blank\">https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex</a></p>",
      "rawMarkdown": "There is only a single target label for each customer which observes the 18 months following their latest credit card statement in the data to see if they are ever past due by 120 days\n\ntrain_labels.csv - target label for each customer_ID\n\nHowever, if you did want to use more rows than just the last one, I think this notebook by @raddar explored the possibility of using D_39 variable to assign labels on the intermediate rows for the customers\n\nhttps://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex",
      "votes": null
    },
    {
      "id": "1848879",
      "postDate": "07/09/2022 02:53:16",
      "content": "<p>Hi. Thanks for your comment and for the suggestion about the <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> notebook. </p>\n<p>It is important to learn about each customer's behavior by looking at all of his/her statements.</p>\n<p>Regarding my comment, I just wanted to check the pattern of zeroes and ones per customer. Now I know there are only two patterns: all zeroes or all ones. </p>\n<p>Gracias!</p>",
      "rawMarkdown": "Hi. Thanks for your comment and for the suggestion about the @raddar notebook. \n\nIt is important to learn about each customer's behavior by looking at all of his/her statements.\n\nRegarding my comment, I just wanted to check the pattern of zeroes and ones per customer. Now I know there are only two patterns: all zeroes or all ones. \n\nGracias!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1848748,
      "author_name": "illidan7",
      "author_url": "",
      "post_date": "07/08/2022 21:27:33",
      "content": "<p>There is only a single target label for each customer which observes the 18 months following their latest credit card statement in the data to see if they are ever past due by 120 days</p>\n<p>train_labels.csv - target label for each customer_ID</p>\n<p>However, if you did want to use more rows than just the last one, I think this notebook by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> explored the possibility of using D_39 variable to assign labels on the intermediate rows for the customers</p>\n<p><a href=\"https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex\" target=\"_blank\">https://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1848879,
          "author_name": "jsmithperera",
          "author_url": "",
          "post_date": "07/09/2022 02:53:16",
          "content": "<p>Hi. Thanks for your comment and for the suggestion about the <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> notebook. </p>\n<p>It is important to learn about each customer's behavior by looking at all of his/her statements.</p>\n<p>Regarding my comment, I just wanted to check the pattern of zeroes and ones per customer. Now I know there are only two patterns: all zeroes or all ones. </p>\n<p>Gracias!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1848692": "Either the customer is in default or it isn´t. No 0's and 1's for the same person.\n\nHere is a little code to prove it.\n\n`D = pd.DataFrame(train, columns =[\"customer_ID\", \"target\"])\nD = D.loc[D[\"target\"] == 1]\n\nysum = D.groupby(\"customer_ID\", as_index=False).sum()\nysize = D.groupby(\"customer_ID\",as_index=False).size()\n\nysum['target'].equals(ysize['size'])`",
    "1848748": "There is only a single target label for each customer which observes the 18 months following their latest credit card statement in the data to see if they are ever past due by 120 days\n\ntrain_labels.csv - target label for each customer_ID\n\nHowever, if you did want to use more rows than just the last one, I think this notebook by @raddar explored the possibility of using D_39 variable to assign labels on the intermediate rows for the customers\n\nhttps://www.kaggle.com/code/raddar/deanonymized-days-overdue-feat-amex",
    "1848879": "Hi. Thanks for your comment and for the suggestion about the @raddar notebook. \n\nIt is important to learn about each customer's behavior by looking at all of his/her statements.\n\nRegarding my comment, I just wanted to check the pattern of zeroes and ones per customer. Now I know there are only two patterns: all zeroes or all ones. \n\nGracias!"
  },
  "source": "meta"
}