{
  "id": 327361,
  "title": "Another way to keep only the last month without groupby",
  "url": "/competitions/amex-default-prediction/discussion/327361",
  "author_name": "João Pedro Peinado",
  "post_date": "2022-05-26T20:26:17.319000",
  "votes": 34,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi guys, I was working with the dataset and the way provided by <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> to keep only the last statement was causing a big memory usage on because of groupby in the code. So I came up with another way to keep the last statement without groupby.</p>\n<p>Inversion way:<br>\n<code>train.groupby('customer_ID').tail(1).drop(['S_2'], axis='columns')</code><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094\" target=\"_blank\">Inversion post</a></p>\n<p>My way:<br>\n<code>train.drop_duplicates(subset=['customer_ID'], keep='last').drop(['S_2'], axis='columns')</code></p>\n<p>With this way I'm able to create more features with less memory available.</p>\n<p>Hope it helps!</p>",
  "messages": [
    {
      "id": 1802526,
      "postDate": "2022-05-26T20:26:17.320Z",
      "content": "<p>Hi guys, I was working with the dataset and the way provided by <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> to keep only the last statement was causing a big memory usage on because of groupby in the code. So I came up with another way to keep the last statement without groupby.</p>\n<p>Inversion way:<br>\n<code>train.groupby('customer_ID').tail(1).drop(['S_2'], axis='columns')</code><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094\" target=\"_blank\">Inversion post</a></p>\n<p>My way:<br>\n<code>train.drop_duplicates(subset=['customer_ID'], keep='last').drop(['S_2'], axis='columns')</code></p>\n<p>With this way I'm able to create more features with less memory available.</p>\n<p>Hope it helps!</p>",
      "rawMarkdown": "Hi guys, I was working with the dataset and the way provided by @inversion to keep only the last statement was causing a big memory usage on because of groupby in the code. So I came up with another way to keep the last statement without groupby.\n\nInversion way:\n`train.groupby('customer_ID').tail(1).drop(['S_2'], axis='columns')`\n[Inversion post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094)\n\nMy way:\n`train.drop_duplicates(subset=['customer_ID'], keep='last').drop(['S_2'], axis='columns')`\n\nWith this way I'm able to create more features with less memory available.\n\nHope it helps!",
      "votes": 33
    },
    {
      "id": 1807908,
      "postDate": "2022-06-01T12:36:54.353Z",
      "content": "<p><a href=\"https://www.kaggle.com/joaopmpeinado\" target=\"_blank\">@joaopmpeinado</a> we are assuming that dates are sorted (luckily this is the case here) with ascending date. </p>",
      "rawMarkdown": "@joaopmpeinado we are assuming that dates are sorted (luckily this is the case here) with ascending date. ",
      "votes": 1,
      "replies": [
        {
          "id": 1808154,
          "postDate": "2022-06-01T15:11:50.543Z",
          "content": "<p><a href=\"https://www.kaggle.com/kmmohsin\" target=\"_blank\">@kmmohsin</a> yes, but this is also true if you are using the groupby version that were provided by the organizers</p>",
          "rawMarkdown": "@kmmohsin yes, but this is also true if you are using the groupby version that were provided by the organizers",
          "votes": 2
        }
      ]
    },
    {
      "id": 1838955,
      "postDate": "2022-07-01T02:54:39.367Z",
      "content": "<p>nice approach! </p>",
      "rawMarkdown": "nice approach! "
    },
    {
      "id": 1833762,
      "postDate": "2022-06-26T10:12:21.320Z",
      "content": "<p>You might want to play with things like average balance over the period. (New to this one so I'm exploring different options.)</p>",
      "rawMarkdown": "You might want to play with things like average balance over the period. (New to this one so I'm exploring different options.)"
    },
    {
      "id": 1804322,
      "postDate": "2022-05-28T19:03:27.943Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1807908,
      "author_name": "Mohsin",
      "author_url": "",
      "post_date": "2022-06-01T12:36:54.353000",
      "content": "<p><a href=\"https://www.kaggle.com/joaopmpeinado\" target=\"_blank\">@joaopmpeinado</a> we are assuming that dates are sorted (luckily this is the case here) with ascending date. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1808154,
          "author_name": "João Pedro Peinado",
          "author_url": "",
          "post_date": "2022-06-01T15:11:50.543000",
          "content": "<p><a href=\"https://www.kaggle.com/kmmohsin\" target=\"_blank\">@kmmohsin</a> yes, but this is also true if you are using the groupby version that were provided by the organizers</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1838955,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-07-01T02:54:39.367000",
      "content": "<p>nice approach! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1833762,
      "author_name": "NotGaussian",
      "author_url": "",
      "post_date": "2022-06-26T10:12:21.320000",
      "content": "<p>You might want to play with things like average balance over the period. (New to this one so I'm exploring different options.)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1804322,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-28T19:03:27.943000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1802526": "Hi guys, I was working with the dataset and the way provided by @inversion to keep only the last statement was causing a big memory usage on because of groupby in the code. So I came up with another way to keep the last statement without groupby.\n\nInversion way:\n`train.groupby('customer_ID').tail(1).drop(['S_2'], axis='columns')`\n[Inversion post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094)\n\nMy way:\n`train.drop_duplicates(subset=['customer_ID'], keep='last').drop(['S_2'], axis='columns')`\n\nWith this way I'm able to create more features with less memory available.\n\nHope it helps!",
    "1807908": "@joaopmpeinado we are assuming that dates are sorted (luckily this is the case here) with ascending date. ",
    "1838955": "nice approach! ",
    "1833762": "You might want to play with things like average balance over the period. (New to this one so I'm exploring different options.)",
    "1804322": ""
  }
}