{
  "id": 337005,
  "title": "Too time-consuming to make agg-like features",
  "url": "/competitions/amex-default-prediction/discussion/337005",
  "author_name": "Joseph Zhou",
  "post_date": "2022-07-14T03:57:34.268000",
  "votes": 3,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'm trying to make some features by user define functions, one of them likes below:</p>\n<pre><code>def drawup_duration(series):\n    series = np.asarray(series.ffill().bfill().fillna(0))\n    if len(series)&lt;2:\n        return 0\n    series=-series\n    k = np.argmax(np.maximum.accumulate(series) - series)\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i]) == 0:\n        j=k\n    else:\n        j = np.argmax(series[:i])\n    return k-j\n</code></pre>\n<p>Then I groupby customer_ID and use agg:</p>\n<pre><code>tmp_agg = train.groupby(\"customer_ID\")[some_features].agg([drawup_duration, func_x, func_y, func_z, ......])\n</code></pre>\n<p>But it is extremely time-consuming, 4 functions on train data cost nearly 3 hours(test data maybe more and more longer). Because the grouper numbers(customer_ID) is too large? And is there any method to accelerate the process? Look forward to your answers sincerely, thanks.</p>",
  "messages": [
    {
      "id": 1854816,
      "postDate": "2022-07-14T03:57:34.270Z",
      "content": "<p>I'm trying to make some features by user define functions, one of them likes below:</p>\n<pre><code>def drawup_duration(series):\n    series = np.asarray(series.ffill().bfill().fillna(0))\n    if len(series)&lt;2:\n        return 0\n    series=-series\n    k = np.argmax(np.maximum.accumulate(series) - series)\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i]) == 0:\n        j=k\n    else:\n        j = np.argmax(series[:i])\n    return k-j\n</code></pre>\n<p>Then I groupby customer_ID and use agg:</p>\n<pre><code>tmp_agg = train.groupby(\"customer_ID\")[some_features].agg([drawup_duration, func_x, func_y, func_z, ......])\n</code></pre>\n<p>But it is extremely time-consuming, 4 functions on train data cost nearly 3 hours(test data maybe more and more longer). Because the grouper numbers(customer_ID) is too large? And is there any method to accelerate the process? Look forward to your answers sincerely, thanks.</p>",
      "rawMarkdown": "I'm trying to make some features by user define functions, one of them likes below:\n```\ndef drawup_duration(series):\n    series = np.asarray(series.ffill().bfill().fillna(0))\n    if len(series)<2:\n        return 0\n    series=-series\n    k = np.argmax(np.maximum.accumulate(series) - series)\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i]) == 0:\n        j=k\n    else:\n        j = np.argmax(series[:i])\n    return k-j\n```\nThen I groupby customer_ID and use agg:\n```\ntmp_agg = train.groupby(\"customer_ID\")[some_features].agg([drawup_duration, func_x, func_y, func_z, ......])\n```\nBut it is extremely time-consuming, 4 functions on train data cost nearly 3 hours(test data maybe more and more longer). Because the grouper numbers(customer_ID) is too large? And is there any method to accelerate the process? Look forward to your answers sincerely, thanks.",
      "votes": 3
    },
    {
      "id": 1855009,
      "postDate": "2022-07-14T08:12:46.067Z",
      "content": "<p>import  cudf <br>\nimport  cupy<br>\n👀👀👀👀👀</p>",
      "rawMarkdown": "import  cudf \nimport  cupy\n👀👀👀👀👀",
      "votes": 2,
      "replies": [
        {
          "id": 1855034,
          "postDate": "2022-07-14T08:37:15.987Z",
          "content": "<p>I'v thought about it, but I don't use the kaggle kernel by the reason of RAM limitation. Let me find out how to install cudf, thanks a lot😄😄😄</p>",
          "rawMarkdown": "I'v thought about it, but I don't use the kaggle kernel by the reason of RAM limitation. Let me find out how to install cudf, thanks a lot😄😄😄"
        },
        {
          "id": 1855051,
          "postDate": "2022-07-14T08:49:43.893Z",
          "content": "<p>Perhaps I can do this part on kaggle kernel and download to my platform😂</p>",
          "rawMarkdown": "Perhaps I can do this part on kaggle kernel and download to my platform😂"
        },
        {
          "id": 1855166,
          "postDate": "2022-07-14T10:34:50.327Z",
          "content": "<p>I also suggest using RAPIDS cudf and cupy because using GPU can speed up computation by a factor of 100X or more!</p>\n<p>There is an example notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> which uses cudf. Furthermore, you should consider processing train and test in chunks if you have GPU VRAM memory errors. The linked notebook also shows how to process test in chunks. (And if we follow test's template, we could also process train in chunks if we're adding lots of new columns)</p>",
          "rawMarkdown": "I also suggest using RAPIDS cudf and cupy because using GPU can speed up computation by a factor of 100X or more!\n\nThere is an example notebook [here][1] which uses cudf. Furthermore, you should consider processing train and test in chunks if you have GPU VRAM memory errors. The linked notebook also shows how to process test in chunks. (And if we follow test's template, we could also process train in chunks if we're adding lots of new columns)\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
          "votes": 2
        },
        {
          "id": 1855381,
          "postDate": "2022-07-14T14:30:23.993Z",
          "content": "<p>I think only req with cudf is it you have a small GPU I think then it's more difficult to fit too many features. But I recommend in case going in pandas try dask or once fe data is generated store it in parquet or pickle file for training. You could create an uber fe file and then take features you want later during training. Which will save time as you iterate</p>",
          "rawMarkdown": "I think only req with cudf is it you have a small GPU I think then it's more difficult to fit too many features. But I recommend in case going in pandas try dask or once fe data is generated store it in parquet or pickle file for training. You could create an uber fe file and then take features you want later during training. Which will save time as you iterate"
        },
        {
          "id": 1855511,
          "postDate": "2022-07-14T16:36:53.320Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1855514,
          "postDate": "2022-07-14T16:38:10.077Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Truly an impressive work, Thanks! Now I am trying \"Pandarallel\" to make speeding up.</p>",
          "rawMarkdown": "@cdeotte Truly an impressive work, Thanks! Now I am trying \"Pandarallel\" to make speeding up.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1855101,
      "postDate": "2022-07-14T09:32:55.623Z",
      "content": "<p>have you tried groupby and lambda functions?</p>",
      "rawMarkdown": "have you tried groupby and lambda functions?",
      "replies": [
        {
          "id": 1855382,
          "postDate": "2022-07-14T14:31:44.083Z",
          "content": "<p>for example tp calculate last - first in each group of customers you can use:<br>\ndf.groupby('customer_ID').apply(lambda x: x[-1] - x[0]</p>",
          "rawMarkdown": "for example tp calculate last - first in each group of customers you can use:\ndf.groupby('customer_ID').apply(lambda x: x[-1] - x[0]"
        },
        {
          "id": 1855393,
          "postDate": "2022-07-14T14:43:53.277Z",
          "content": "<p>Thanks a lot, but some functions are more complicated, and could only be used like:</p>\n<pre><code>df.groupby('customer_ID')[some_features].apply(func)\n</code></pre>\n<p>😿😿😿</p>",
          "rawMarkdown": "Thanks a lot, but some functions are more complicated, and could only be used like:\n```\ndf.groupby('customer_ID')[some_features].apply(func)\n```\n😿😿😿",
          "votes": 1
        },
        {
          "id": 1855413,
          "postDate": "2022-07-14T15:01:12.953Z",
          "content": "<p>I tried it in my FEing and it was okay (speed wise)</p>",
          "rawMarkdown": "I tried it in my FEing and it was okay (speed wise)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1855009,
      "author_name": "kgxiao",
      "author_url": "",
      "post_date": "2022-07-14T08:12:46.067000",
      "content": "<p>import  cudf <br>\nimport  cupy<br>\n👀👀👀👀👀</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1855034,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-07-14T08:37:15.987000",
          "content": "<p>I'v thought about it, but I don't use the kaggle kernel by the reason of RAM limitation. Let me find out how to install cudf, thanks a lot😄😄😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1855051,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-07-14T08:49:43.893000",
          "content": "<p>Perhaps I can do this part on kaggle kernel and download to my platform😂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1855166,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-14T10:34:50.327000",
          "content": "<p>I also suggest using RAPIDS cudf and cupy because using GPU can speed up computation by a factor of 100X or more!</p>\n<p>There is an example notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> which uses cudf. Furthermore, you should consider processing train and test in chunks if you have GPU VRAM memory errors. The linked notebook also shows how to process test in chunks. (And if we follow test's template, we could also process train in chunks if we're adding lots of new columns)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1855381,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-07-14T14:30:23.993000",
          "content": "<p>I think only req with cudf is it you have a small GPU I think then it's more difficult to fit too many features. But I recommend in case going in pandas try dask or once fe data is generated store it in parquet or pickle file for training. You could create an uber fe file and then take features you want later during training. Which will save time as you iterate</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1855511,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-14T16:36:53.320000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1855514,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-07-14T16:38:10.077000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Truly an impressive work, Thanks! Now I am trying \"Pandarallel\" to make speeding up.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1855101,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-07-14T09:32:55.623000",
      "content": "<p>have you tried groupby and lambda functions?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1855382,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-07-14T14:31:44.083000",
          "content": "<p>for example tp calculate last - first in each group of customers you can use:<br>\ndf.groupby('customer_ID').apply(lambda x: x[-1] - x[0]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1855393,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-07-14T14:43:53.277000",
          "content": "<p>Thanks a lot, but some functions are more complicated, and could only be used like:</p>\n<pre><code>df.groupby('customer_ID')[some_features].apply(func)\n</code></pre>\n<p>😿😿😿</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1855413,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-07-14T15:01:12.953000",
          "content": "<p>I tried it in my FEing and it was okay (speed wise)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1854816": "I'm trying to make some features by user define functions, one of them likes below:\n```\ndef drawup_duration(series):\n    series = np.asarray(series.ffill().bfill().fillna(0))\n    if len(series)<2:\n        return 0\n    series=-series\n    k = np.argmax(np.maximum.accumulate(series) - series)\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i]) == 0:\n        j=k\n    else:\n        j = np.argmax(series[:i])\n    return k-j\n```\nThen I groupby customer_ID and use agg:\n```\ntmp_agg = train.groupby(\"customer_ID\")[some_features].agg([drawup_duration, func_x, func_y, func_z, ......])\n```\nBut it is extremely time-consuming, 4 functions on train data cost nearly 3 hours(test data maybe more and more longer). Because the grouper numbers(customer_ID) is too large? And is there any method to accelerate the process? Look forward to your answers sincerely, thanks.",
    "1855009": "import  cudf \nimport  cupy\n👀👀👀👀👀",
    "1855101": "have you tried groupby and lambda functions?"
  }
}