{
  "id": 344178,
  "title": "A faster implementation to create diff columns",
  "url": "/competitions/amex-default-prediction/discussion/344178",
  "author_name": "",
  "post_date": "2022-08-14T10:27:47.928953800Z",
  "votes": 19,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I saw that some of the notebooks are using <strong>for loops</strong> to create diff features, python <strong>for loops</strong> are really slow and they are taking <strong>10-12 minutes</strong> to create diff features.<br>\nHere is the faster code to implement them</p>\n<pre><code>def difference(groups,num_features,shift):\n    data=(groups[num_features].nth(-1)-groups[num_features].nth(-1*shift)).rename(columns={f: f\"{f}_diff{shift-1}\" for f in num_features})\n    return data\ndef get_difference(data,num_features):\n    groups=data.groupby('customer_ID')\n    df1=difference(groups,num_features,2)\n    df1.reset_index(inplace=True)\n    return df1\n</code></pre>\n<p>The above code takes   <strong>12-20 seconds</strong>  to create diff features</p>",
  "messages": [
    {
      "id": "1898126",
      "postDate": "08/14/2022 10:27:47",
      "content": "<p>I saw that some of the notebooks are using <strong>for loops</strong> to create diff features, python <strong>for loops</strong> are really slow and they are taking <strong>10-12 minutes</strong> to create diff features.<br>\nHere is the faster code to implement them</p>\n<pre><code>def difference(groups,num_features,shift):\n    data=(groups[num_features].nth(-1)-groups[num_features].nth(-1*shift)).rename(columns={f: f\"{f}_diff{shift-1}\" for f in num_features})\n    return data\ndef get_difference(data,num_features):\n    groups=data.groupby('customer_ID')\n    df1=difference(groups,num_features,2)\n    df1.reset_index(inplace=True)\n    return df1\n</code></pre>\n<p>The above code takes   <strong>12-20 seconds</strong>  to create diff features</p>",
      "rawMarkdown": "I saw that some of the notebooks are using **for loops** to create diff features, python **for loops** are really slow and they are taking **10-12 minutes** to create diff features.\nHere is the faster code to implement them\n```\ndef difference(groups,num_features,shift):\n    data=(groups[num_features].nth(-1)-groups[num_features].nth(-1*shift)).rename(columns={f: f\"{f}_diff{shift-1}\" for f in num_features})\n    return data\ndef get_difference(data,num_features):\n    groups=data.groupby('customer_ID')\n    df1=difference(groups,num_features,2)\n    df1.reset_index(inplace=True)\n    return df1\n```\nThe above code takes   **12-20 seconds**  to create diff features",
      "votes": null
    },
    {
      "id": "1898192",
      "postDate": "08/14/2022 11:45:01",
      "content": "<p>Great hack!! This is highly effective! Thanks for sharing!!</p>",
      "rawMarkdown": "Great hack!! This is highly effective! Thanks for sharing!!",
      "votes": null
    },
    {
      "id": "1900150",
      "postDate": "08/15/2022 18:34:38",
      "content": "<p>this is great .. I was doing the trick for diff one only but this helps to do for any level </p>",
      "rawMarkdown": "this is great .. I was doing the trick for diff one only but this helps to do for any level",
      "votes": null
    },
    {
      "id": "1900915",
      "postDate": "08/16/2022 11:03:54",
      "content": "<p>Great job on this.<br>\nDid you validate that you get the same results (by comparing the outputs - to be sure)? </p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Great job on this.\nDid you validate that you get the same results (by comparing the outputs - to be sure)? \n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "1900960",
      "postDate": "08/16/2022 11:33:56",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, I checked it for random <strong>customer_ID</strong> it seems good to me although. Only one issue that can be possible and also will be there for public notebooks is that <strong>Pandas Groupby</strong>  functions drop nan columns by default. If you have found anything else that's suspicious, please let me know.<br>\nRegards</p>",
      "rawMarkdown": "Hi @thedevastator, I checked it for random **customer_ID** it seems good to me although. Only one issue that can be possible and also will be there for public notebooks is that **Pandas Groupby**  functions drop nan columns by default. If you have found anything else that's suspicious, please let me know.\nRegards",
      "votes": null
    },
    {
      "id": "1901735",
      "postDate": "08/16/2022 21:12:22",
      "content": "<p>May I ask why you sort values by customer_ID? I don't think it is necessary.</p>",
      "rawMarkdown": "May I ask why you sort values by customer_ID? I don't think it is necessary.",
      "votes": null
    },
    {
      "id": "1903082",
      "postDate": "08/17/2022 05:01:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a>  thanks for the feedback I put it there just to be on the safer side if i try concatenating it somewhere, however you are right it is not required.</p>",
      "rawMarkdown": "Hi @rasoulmojtahedzadeh  thanks for the feedback I put it there just to be on the safer side if i try concatenating it somewhere, however you are right it is not required.",
      "votes": null
    },
    {
      "id": "1904913",
      "postDate": "08/18/2022 15:28:33",
      "content": "<p>You can also just diff to the previous row without the group, but null it out when previous row has a different customer ID. </p>",
      "rawMarkdown": "You can also just diff to the previous row without the group, but null it out when previous row has a different customer ID.",
      "votes": null
    },
    {
      "id": "1905556",
      "postDate": "08/19/2022 06:23:20",
      "content": "<p>Thanks, this is great.</p>",
      "rawMarkdown": "Thanks, this is great.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1898192,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/14/2022 11:45:01",
      "content": "<p>Great hack!! This is highly effective! Thanks for sharing!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1900150,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "08/15/2022 18:34:38",
      "content": "<p>this is great .. I was doing the trick for diff one only but this helps to do for any level </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1900915,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "08/16/2022 11:03:54",
      "content": "<p>Great job on this.<br>\nDid you validate that you get the same results (by comparing the outputs - to be sure)? </p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1900960,
          "author_name": "chaudharypriyanshu",
          "author_url": "",
          "post_date": "08/16/2022 11:33:56",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, I checked it for random <strong>customer_ID</strong> it seems good to me although. Only one issue that can be possible and also will be there for public notebooks is that <strong>Pandas Groupby</strong>  functions drop nan columns by default. If you have found anything else that's suspicious, please let me know.<br>\nRegards</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1901735,
      "author_name": "rasoulmojtahedzadeh",
      "author_url": "",
      "post_date": "08/16/2022 21:12:22",
      "content": "<p>May I ask why you sort values by customer_ID? I don't think it is necessary.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1903082,
          "author_name": "chaudharypriyanshu",
          "author_url": "",
          "post_date": "08/17/2022 05:01:26",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a>  thanks for the feedback I put it there just to be on the safer side if i try concatenating it somewhere, however you are right it is not required.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1904913,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "08/18/2022 15:28:33",
      "content": "<p>You can also just diff to the previous row without the group, but null it out when previous row has a different customer ID. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1905556,
      "author_name": "krishnapriya18",
      "author_url": "",
      "post_date": "08/19/2022 06:23:20",
      "content": "<p>Thanks, this is great.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1898126": "I saw that some of the notebooks are using **for loops** to create diff features, python **for loops** are really slow and they are taking **10-12 minutes** to create diff features.\nHere is the faster code to implement them\n```\ndef difference(groups,num_features,shift):\n    data=(groups[num_features].nth(-1)-groups[num_features].nth(-1*shift)).rename(columns={f: f\"{f}_diff{shift-1}\" for f in num_features})\n    return data\ndef get_difference(data,num_features):\n    groups=data.groupby('customer_ID')\n    df1=difference(groups,num_features,2)\n    df1.reset_index(inplace=True)\n    return df1\n```\nThe above code takes   **12-20 seconds**  to create diff features",
    "1898192": "Great hack!! This is highly effective! Thanks for sharing!!",
    "1900150": "this is great .. I was doing the trick for diff one only but this helps to do for any level",
    "1900915": "Great job on this.\nDid you validate that you get the same results (by comparing the outputs - to be sure)? \n\nThe Devastator.",
    "1900960": "Hi @thedevastator, I checked it for random **customer_ID** it seems good to me although. Only one issue that can be possible and also will be there for public notebooks is that **Pandas Groupby**  functions drop nan columns by default. If you have found anything else that's suspicious, please let me know.\nRegards",
    "1901735": "May I ask why you sort values by customer_ID? I don't think it is necessary.",
    "1903082": "Hi @rasoulmojtahedzadeh  thanks for the feedback I put it there just to be on the safer side if i try concatenating it somewhere, however you are right it is not required.",
    "1904913": "You can also just diff to the previous row without the group, but null it out when previous row has a different customer ID.",
    "1905556": "Thanks, this is great."
  },
  "source": "meta"
}