{
  "id": 333940,
  "title": "RAPIDS Feature Engineering",
  "url": "/competitions/amex-default-prediction/discussion/333940",
  "author_name": "",
  "post_date": "2022-06-29T04:33:28.231628200Z",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm really enjoying using the RAPIDS library, but I'm struggling with implementing some things and unable to find good resources online to learn how to do them. So I'm hoping that the RAPIDS natives in the Kaggle community can help me out. Couple of things I am unable to do:</p>\n<p>1) Compute date difference on the S_2 column</p>\n<p><code>df['S_2_diff'] = df['S_2'].diff().dt.days</code></p>\n<blockquote>\n  <p>NotImplementedError: Diff currently only supports numeric dtypes</p>\n</blockquote>\n<p>2) Apply lambda functions on group by statements</p>\n<p><code>df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())</code></p>\n<blockquote>\n  <p>/opt/conda/lib/python3.7/site-packages/cudf/core/groupby/groupby.py:495: UserWarning: GroupBy.apply() performance scales poorly with number of groups. Got 458913 groups.<br>\n    f\"GroupBy.apply() performance scales poorly with \"</p>\n</blockquote>\n<p>Are these cases where you have to bite the bullet and use pandas instead?</p>\n<p>Look forward to any help on this!</p>",
  "messages": [
    {
      "id": "1836782",
      "postDate": "06/29/2022 04:33:28",
      "content": "<p>I'm really enjoying using the RAPIDS library, but I'm struggling with implementing some things and unable to find good resources online to learn how to do them. So I'm hoping that the RAPIDS natives in the Kaggle community can help me out. Couple of things I am unable to do:</p>\n<p>1) Compute date difference on the S_2 column</p>\n<p><code>df['S_2_diff'] = df['S_2'].diff().dt.days</code></p>\n<blockquote>\n  <p>NotImplementedError: Diff currently only supports numeric dtypes</p>\n</blockquote>\n<p>2) Apply lambda functions on group by statements</p>\n<p><code>df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())</code></p>\n<blockquote>\n  <p>/opt/conda/lib/python3.7/site-packages/cudf/core/groupby/groupby.py:495: UserWarning: GroupBy.apply() performance scales poorly with number of groups. Got 458913 groups.<br>\n    f\"GroupBy.apply() performance scales poorly with \"</p>\n</blockquote>\n<p>Are these cases where you have to bite the bullet and use pandas instead?</p>\n<p>Look forward to any help on this!</p>",
      "rawMarkdown": "I'm really enjoying using the RAPIDS library, but I'm struggling with implementing some things and unable to find good resources online to learn how to do them. So I'm hoping that the RAPIDS natives in the Kaggle community can help me out. Couple of things I am unable to do:\n\n1) Compute date difference on the S_2 column\n\n`df['S_2_diff'] = df['S_2'].diff().dt.days`\n\n> NotImplementedError: Diff currently only supports numeric dtypes\n\n2) Apply lambda functions on group by statements\n\n`df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())`\n\n> /opt/conda/lib/python3.7/site-packages/cudf/core/groupby/groupby.py:495: UserWarning: GroupBy.apply() performance scales poorly with number of groups. Got 458913 groups.\n  f\"GroupBy.apply() performance scales poorly with \"\n\n\nAre these cases where you have to bite the bullet and use pandas instead?\n\nLook forward to any help on this!",
      "votes": null
    },
    {
      "id": "1837518",
      "postDate": "06/29/2022 16:35:22",
      "content": "<p>Are you using the latest version <code>22.06.00</code> of RAPIDS cuDF offline? Or are you using version <code>21.10.01</code> online at Kaggle?</p>\n<p>With the latest RAPIDS cuDF, you can groupby and diff with datetime. (But note you cannot diff datetime without groupby first). So for example, you can compute time difference with </p>\n<pre><code># DATAFRAME MUST BE SORTED BY TIME\n# ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY TIME\n# SO WE DON'T NEED TO SORT AGAIN\n# df = df.sort_values(['customer_ID','S_2'])\n\ndf['S_2_diff'] = df[['S_2','customer_ID']].groupby('customer_ID').S_2.diff().dt.days\n</code></pre>\n<p>I verified that the above line of code works on RAPIDS cuDF <code>22.04.00</code> and perhaps earlier versions but not Kaggle's <code>21.10.01</code> version</p>\n<p>========<br>\nNote if you must make it work in Kaggle's notebooks, there are usually tricks to accomplish anything. For example the code below will run on RAPIDS cuDF version <code>21.10.01</code> and accomplish the same as above:</p>\n<pre><code># DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n# ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n# SO WE DON'T NEED TO SORT AGAIN\n# df = df.sort_values(['customer_ID','S_2'])\n\ndf['days_since_1970'] = df.S_2.astype('int64')/1e9/(60*60*24)\ndf['S_2_diff'] = df.days_since_1970.diff()\ndf['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\ndf.loc[df.x==0,'S_2_diff'] = cudf.NA\ndf = df.drop(['days_since_1970','x'], axis=1)\n</code></pre>",
      "rawMarkdown": "Are you using the latest version `22.06.00` of RAPIDS cuDF offline? Or are you using version `21.10.01` online at Kaggle?\n\nWith the latest RAPIDS cuDF, you can groupby and diff with datetime. (But note you cannot diff datetime without groupby first). So for example, you can compute time difference with \n\n    # DATAFRAME MUST BE SORTED BY TIME\n    # ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY TIME\n    # SO WE DON'T NEED TO SORT AGAIN\n    # df = df.sort_values(['customer_ID','S_2'])\n\n    df['S_2_diff'] = df[['S_2','customer_ID']].groupby('customer_ID').S_2.diff().dt.days\n\nI verified that the above line of code works on RAPIDS cuDF `22.04.00` and perhaps earlier versions but not Kaggle's `21.10.01` version\n\n========\nNote if you must make it work in Kaggle's notebooks, there are usually tricks to accomplish anything. For example the code below will run on RAPIDS cuDF version `21.10.01` and accomplish the same as above:\n\n    # DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n    # ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n    # SO WE DON'T NEED TO SORT AGAIN\n    # df = df.sort_values(['customer_ID','S_2'])\n\n    df['days_since_1970'] = df.S_2.astype('int64')/1e9/(60*60*24)\n    df['S_2_diff'] = df.days_since_1970.diff()\n    df['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\n    df.loc[df.x==0,'S_2_diff'] = cudf.NA\n    df = df.drop(['days_since_1970','x'], axis=1)",
      "votes": null
    },
    {
      "id": "1837724",
      "postDate": "06/29/2022 20:43:10",
      "content": "<p>Yup I was using Kaggle notebooks. That's an awesome workaround, thank you for providing that example! 😃</p>\n<p>Do you have any inputs regarding the apply functions? </p>\n<p>I found the following link from the Rapids documentation. Still playing around with it, have not gotten it to work yet. </p>\n<p><a href=\"https://docs.rapids.ai/api/cudf/stable/user_guide/guide-to-udfs.html#groupby-dataframe-udfs\" target=\"_blank\">https://docs.rapids.ai/api/cudf/stable/user_guide/guide-to-udfs.html#groupby-dataframe-udfs</a></p>",
      "rawMarkdown": "Yup I was using Kaggle notebooks. That's an awesome workaround, thank you for providing that example! 😃\n\nDo you have any inputs regarding the apply functions? \n\nI found the following link from the Rapids documentation. Still playing around with it, have not gotten it to work yet. \n\nhttps://docs.rapids.ai/api/cudf/stable/user_guide/guide-to-udfs.html#groupby-dataframe-udfs",
      "votes": null
    },
    {
      "id": "1837754",
      "postDate": "06/29/2022 21:39:33",
      "content": "<p>I haven't written UDF in a while, so i would need to review before commenting on UDF. However most of your <code>apply</code> needs can be solved with tricks. For example, your specific question</p>\n<pre><code>df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())\n</code></pre>\n<p>Is just the groupby <code>diff()</code> function. In recent versions of RAPIDS cuDF, you can do it with</p>\n<pre><code>df.groupby('customer_ID')[col].diff()\n</code></pre>\n<p>In a Kaggle notebook, the RAPIDS version doesn't support groupby <code>diff</code> yet, so you would need to use <code>diff</code> without groupby after you sort by customer and time to simulate groupby:</p>\n<pre><code># DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n# ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n# SO WE DON'T NEED TO SORT AGAIN\n# df = df.sort_values(['customer_ID','S_2'])\n\ndf[f'{col}_diff'] = df[col].diff()\ndf['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\ndf.loc[df.x==0,f'{col}_diff'] = cudf.NA\ndf = df.drop(['x'], axis=1)\n</code></pre>",
      "rawMarkdown": "I haven't written UDF in a while, so i would need to review before commenting on UDF. However most of your `apply` needs can be solved with tricks. For example, your specific question\n\n    df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())\n\nIs just the groupby `diff()` function. In recent versions of RAPIDS cuDF, you can do it with\n\n    df.groupby('customer_ID')[col].diff()\n\nIn a Kaggle notebook, the RAPIDS version doesn't support groupby `diff` yet, so you would need to use `diff` without groupby after you sort by customer and time to simulate groupby:\n\n    # DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n    # ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n    # SO WE DON'T NEED TO SORT AGAIN\n    # df = df.sort_values(['customer_ID','S_2'])\n\n    df[f'{col}_diff'] = df[col].diff()\n    df['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\n    df.loc[df.x==0,f'{col}_diff'] = cudf.NA\n    df = df.drop(['x'], axis=1)",
      "votes": null
    },
    {
      "id": "1837852",
      "postDate": "06/30/2022 02:16:22",
      "content": "<p>Thank you for the responses <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Very helpful insights</p>\n<p>Yeah I messed with these UDFs and apply_grouped for a little while but I've given up on it for now. Will look for work arounds to implement some of the more complex ones or go crawling back to pandas</p>",
      "rawMarkdown": "Thank you for the responses @cdeotte. Very helpful insights\n\nYeah I messed with these UDFs and apply_grouped for a little while but I've given up on it for now. Will look for work arounds to implement some of the more complex ones or go crawling back to pandas",
      "votes": null
    },
    {
      "id": "1837903",
      "postDate": "06/30/2022 03:20:24",
      "content": "<p>Note that when using an older version of RAPIDS in Kaggle notebooks, you can use both RAPIDS cuDF and Pandas. For example do everything you can in cuDF and then if something doesn't work then just add <code>to_pandas()</code> to the cuDF command:</p>\n<pre><code>df = cudf.read_parquet('train.parquet')\ndf['new_feature'] = df.to_pandas().groupby('customer_ID')[col].diff()\n</code></pre>\n<p>This way you will accelerate as much as you can with GPU RAPIDS cuDF and only use Pandas when needed.</p>",
      "rawMarkdown": "Note that when using an older version of RAPIDS in Kaggle notebooks, you can use both RAPIDS cuDF and Pandas. For example do everything you can in cuDF and then if something doesn't work then just add `to_pandas()` to the cuDF command:\n\n    df = cudf.read_parquet('train.parquet')\n    df['new_feature'] = df.to_pandas().groupby('customer_ID')[col].diff()\n\nThis way you will accelerate as much as you can with GPU RAPIDS cuDF and only use Pandas when needed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1837518,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/29/2022 16:35:22",
      "content": "<p>Are you using the latest version <code>22.06.00</code> of RAPIDS cuDF offline? Or are you using version <code>21.10.01</code> online at Kaggle?</p>\n<p>With the latest RAPIDS cuDF, you can groupby and diff with datetime. (But note you cannot diff datetime without groupby first). So for example, you can compute time difference with </p>\n<pre><code># DATAFRAME MUST BE SORTED BY TIME\n# ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY TIME\n# SO WE DON'T NEED TO SORT AGAIN\n# df = df.sort_values(['customer_ID','S_2'])\n\ndf['S_2_diff'] = df[['S_2','customer_ID']].groupby('customer_ID').S_2.diff().dt.days\n</code></pre>\n<p>I verified that the above line of code works on RAPIDS cuDF <code>22.04.00</code> and perhaps earlier versions but not Kaggle's <code>21.10.01</code> version</p>\n<p>========<br>\nNote if you must make it work in Kaggle's notebooks, there are usually tricks to accomplish anything. For example the code below will run on RAPIDS cuDF version <code>21.10.01</code> and accomplish the same as above:</p>\n<pre><code># DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n# ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n# SO WE DON'T NEED TO SORT AGAIN\n# df = df.sort_values(['customer_ID','S_2'])\n\ndf['days_since_1970'] = df.S_2.astype('int64')/1e9/(60*60*24)\ndf['S_2_diff'] = df.days_since_1970.diff()\ndf['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\ndf.loc[df.x==0,'S_2_diff'] = cudf.NA\ndf = df.drop(['days_since_1970','x'], axis=1)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1837724,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "06/29/2022 20:43:10",
          "content": "<p>Yup I was using Kaggle notebooks. That's an awesome workaround, thank you for providing that example! 😃</p>\n<p>Do you have any inputs regarding the apply functions? </p>\n<p>I found the following link from the Rapids documentation. Still playing around with it, have not gotten it to work yet. </p>\n<p><a href=\"https://docs.rapids.ai/api/cudf/stable/user_guide/guide-to-udfs.html#groupby-dataframe-udfs\" target=\"_blank\">https://docs.rapids.ai/api/cudf/stable/user_guide/guide-to-udfs.html#groupby-dataframe-udfs</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1837754,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/29/2022 21:39:33",
          "content": "<p>I haven't written UDF in a while, so i would need to review before commenting on UDF. However most of your <code>apply</code> needs can be solved with tricks. For example, your specific question</p>\n<pre><code>df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())\n</code></pre>\n<p>Is just the groupby <code>diff()</code> function. In recent versions of RAPIDS cuDF, you can do it with</p>\n<pre><code>df.groupby('customer_ID')[col].diff()\n</code></pre>\n<p>In a Kaggle notebook, the RAPIDS version doesn't support groupby <code>diff</code> yet, so you would need to use <code>diff</code> without groupby after you sort by customer and time to simulate groupby:</p>\n<pre><code># DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n# ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n# SO WE DON'T NEED TO SORT AGAIN\n# df = df.sort_values(['customer_ID','S_2'])\n\ndf[f'{col}_diff'] = df[col].diff()\ndf['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\ndf.loc[df.x==0,f'{col}_diff'] = cudf.NA\ndf = df.drop(['x'], axis=1)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1837852,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "06/30/2022 02:16:22",
          "content": "<p>Thank you for the responses <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Very helpful insights</p>\n<p>Yeah I messed with these UDFs and apply_grouped for a little while but I've given up on it for now. Will look for work arounds to implement some of the more complex ones or go crawling back to pandas</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1837903,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/30/2022 03:20:24",
          "content": "<p>Note that when using an older version of RAPIDS in Kaggle notebooks, you can use both RAPIDS cuDF and Pandas. For example do everything you can in cuDF and then if something doesn't work then just add <code>to_pandas()</code> to the cuDF command:</p>\n<pre><code>df = cudf.read_parquet('train.parquet')\ndf['new_feature'] = df.to_pandas().groupby('customer_ID')[col].diff()\n</code></pre>\n<p>This way you will accelerate as much as you can with GPU RAPIDS cuDF and only use Pandas when needed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1836782": "I'm really enjoying using the RAPIDS library, but I'm struggling with implementing some things and unable to find good resources online to learn how to do them. So I'm hoping that the RAPIDS natives in the Kaggle community can help me out. Couple of things I am unable to do:\n\n1) Compute date difference on the S_2 column\n\n`df['S_2_diff'] = df['S_2'].diff().dt.days`\n\n> NotImplementedError: Diff currently only supports numeric dtypes\n\n2) Apply lambda functions on group by statements\n\n`df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())`\n\n> /opt/conda/lib/python3.7/site-packages/cudf/core/groupby/groupby.py:495: UserWarning: GroupBy.apply() performance scales poorly with number of groups. Got 458913 groups.\n  f\"GroupBy.apply() performance scales poorly with \"\n\n\nAre these cases where you have to bite the bullet and use pandas instead?\n\nLook forward to any help on this!",
    "1837518": "Are you using the latest version `22.06.00` of RAPIDS cuDF offline? Or are you using version `21.10.01` online at Kaggle?\n\nWith the latest RAPIDS cuDF, you can groupby and diff with datetime. (But note you cannot diff datetime without groupby first). So for example, you can compute time difference with \n\n    # DATAFRAME MUST BE SORTED BY TIME\n    # ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY TIME\n    # SO WE DON'T NEED TO SORT AGAIN\n    # df = df.sort_values(['customer_ID','S_2'])\n\n    df['S_2_diff'] = df[['S_2','customer_ID']].groupby('customer_ID').S_2.diff().dt.days\n\nI verified that the above line of code works on RAPIDS cuDF `22.04.00` and perhaps earlier versions but not Kaggle's `21.10.01` version\n\n========\nNote if you must make it work in Kaggle's notebooks, there are usually tricks to accomplish anything. For example the code below will run on RAPIDS cuDF version `21.10.01` and accomplish the same as above:\n\n    # DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n    # ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n    # SO WE DON'T NEED TO SORT AGAIN\n    # df = df.sort_values(['customer_ID','S_2'])\n\n    df['days_since_1970'] = df.S_2.astype('int64')/1e9/(60*60*24)\n    df['S_2_diff'] = df.days_since_1970.diff()\n    df['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\n    df.loc[df.x==0,'S_2_diff'] = cudf.NA\n    df = df.drop(['days_since_1970','x'], axis=1)",
    "1837724": "Yup I was using Kaggle notebooks. That's an awesome workaround, thank you for providing that example! 😃\n\nDo you have any inputs regarding the apply functions? \n\nI found the following link from the Rapids documentation. Still playing around with it, have not gotten it to work yet. \n\nhttps://docs.rapids.ai/api/cudf/stable/user_guide/guide-to-udfs.html#groupby-dataframe-udfs",
    "1837754": "I haven't written UDF in a while, so i would need to review before commenting on UDF. However most of your `apply` needs can be solved with tricks. For example, your specific question\n\n    df.groupby('customer_ID')[col].apply(lambda x: x-x.shift())\n\nIs just the groupby `diff()` function. In recent versions of RAPIDS cuDF, you can do it with\n\n    df.groupby('customer_ID')[col].diff()\n\nIn a Kaggle notebook, the RAPIDS version doesn't support groupby `diff` yet, so you would need to use `diff` without groupby after you sort by customer and time to simulate groupby:\n\n    # DATAFRAME MUST BE SORTED BY CUSTOMER AND TIME\n    # ORIGINAL CSV AND RADDAR'S PARQUET ARE SORTED BY CUSTOMER AND TIME\n    # SO WE DON'T NEED TO SORT AGAIN\n    # df = df.sort_values(['customer_ID','S_2'])\n\n    df[f'{col}_diff'] = df[col].diff()\n    df['x'] = df.groupby('customer_ID').S_2.agg('cumcount')\n    df.loc[df.x==0,f'{col}_diff'] = cudf.NA\n    df = df.drop(['x'], axis=1)",
    "1837852": "Thank you for the responses @cdeotte. Very helpful insights\n\nYeah I messed with these UDFs and apply_grouped for a little while but I've given up on it for now. Will look for work arounds to implement some of the more complex ones or go crawling back to pandas",
    "1837903": "Note that when using an older version of RAPIDS in Kaggle notebooks, you can use both RAPIDS cuDF and Pandas. For example do everything you can in cuDF and then if something doesn't work then just add `to_pandas()` to the cuDF command:\n\n    df = cudf.read_parquet('train.parquet')\n    df['new_feature'] = df.to_pandas().groupby('customer_ID')[col].diff()\n\nThis way you will accelerate as much as you can with GPU RAPIDS cuDF and only use Pandas when needed."
  },
  "source": "meta"
}