{
  "id": 327094,
  "title": "Last month per customer",
  "url": "/competitions/amex-default-prediction/discussion/327094",
  "author_name": "inversion",
  "post_date": "2022-05-25T15:09:50.650000",
  "votes": 64,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Here's a way using pandas to just keep the last statement month per customer.</p>\n<pre><code>X_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n</code></pre>",
  "messages": [
    {
      "id": 1801266,
      "postDate": "2022-05-25T15:09:50.650Z",
      "content": "<p>Here's a way using pandas to just keep the last statement month per customer.</p>\n<pre><code>X_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n</code></pre>",
      "rawMarkdown": "Here's a way using pandas to just keep the last statement month per customer.\n\n```\nX_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n```",
      "votes": 64
    },
    {
      "id": 1803730,
      "postDate": "2022-05-28T05:19:56.027Z",
      "content": "<p>If or after the dataframe has been sorted you could also write:</p>\n<p><code>X_train = X_train.drop_duplicates(subset=[\"customer_ID\"],  keep=\"last\")</code></p>",
      "rawMarkdown": "If or after the dataframe has been sorted you could also write:\n\n`X_train = X_train.drop_duplicates(subset=[\"customer_ID\"],  keep=\"last\")`",
      "votes": 12
    },
    {
      "id": 1802614,
      "postDate": "2022-05-27T00:38:40.670Z",
      "content": "<p>And you want to keep the last one just because it helps process the data due to hardware limitations? Or is there some logic behind that?</p>",
      "rawMarkdown": "And you want to keep the last one just because it helps process the data due to hardware limitations? Or is there some logic behind that?",
      "votes": 12,
      "replies": [
        {
          "id": 1846216,
          "postDate": "2022-07-07T00:01:47.803Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1851500,
          "postDate": "2022-07-11T09:56:23.850Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense?scriptVersionId=97329957&amp;cellId=17\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense?scriptVersionId=97329957&amp;cellId=17</a></p>",
          "rawMarkdown": "https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense?scriptVersionId=97329957&cellId=17"
        },
        {
          "id": 1851517,
          "postDate": "2022-07-11T10:13:06.260Z",
          "content": "<p>It makes sense to train our models based only on the latest statement available for each customer. 😎</p>\n<p>Because, the provided target variables describe the latest default status of each customer. So it will be wise to remove old statements from training dataset. ☀️</p>",
          "rawMarkdown": "It makes sense to train our models based only on the latest statement available for each customer. 😎\n\nBecause, the provided target variables describe the latest default status of each customer. So it will be wise to remove old statements from training dataset. ☀️",
          "votes": 3
        }
      ]
    },
    {
      "id": 1872849,
      "postDate": "2022-07-27T09:17:22.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1803730,
      "author_name": "Thomas Meißner",
      "author_url": "",
      "post_date": "2022-05-28T05:19:56.027000",
      "content": "<p>If or after the dataframe has been sorted you could also write:</p>\n<p><code>X_train = X_train.drop_duplicates(subset=[\"customer_ID\"],  keep=\"last\")</code></p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 1802614,
      "author_name": "Gad Benram",
      "author_url": "",
      "post_date": "2022-05-27T00:38:40.670000",
      "content": "<p>And you want to keep the last one just because it helps process the data due to hardware limitations? Or is there some logic behind that?</p>",
      "votes": 12,
      "replies": [
        {
          "id": 1846216,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-07T00:01:47.803000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1851500,
          "author_name": "kazuchimo",
          "author_url": "",
          "post_date": "2022-07-11T09:56:23.850000",
          "content": "<p><a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense?scriptVersionId=97329957&amp;cellId=17\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense?scriptVersionId=97329957&amp;cellId=17</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1851517,
          "author_name": "Muhammad Irfan Azam",
          "author_url": "",
          "post_date": "2022-07-11T10:13:06.260000",
          "content": "<p>It makes sense to train our models based only on the latest statement available for each customer. 😎</p>\n<p>Because, the provided target variables describe the latest default status of each customer. So it will be wise to remove old statements from training dataset. ☀️</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1872849,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-27T09:17:22.550000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801266": "Here's a way using pandas to just keep the last statement month per customer.\n\n```\nX_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n```",
    "1803730": "If or after the dataframe has been sorted you could also write:\n\n`X_train = X_train.drop_duplicates(subset=[\"customer_ID\"],  keep=\"last\")`",
    "1802614": "And you want to keep the last one just because it helps process the data due to hardware limitations? Or is there some logic behind that?",
    "1872849": ""
  }
}