{
  "id": 208912,
  "title": "Quick Pandas tips for feature engineering",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208912",
  "author_name": "",
  "post_date": "2021-01-05T15:06:59.525350Z",
  "votes": 12,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I thought I'd share two Pandas tricks that might be useful for you.</p>\n<ul>\n<li><p>To avoid data leakage, I saw this suggestion in the comments of a discussion here <code>train.groupby([\"user_id\",\"task_container_id\"])[\"feature\"].transform(\"first\")</code> which is quite elegant, but this one is less buggy: </p>\n<pre><code>train[\"task_con\"] = (train.timestamp != train.timestamp.shift()).cumsum()\n train[\"feature\"] = train.groupby([\"user_id\",\"task_con\"])[\"feature\"].transform(\"first\")\n</code></pre></li>\n</ul>\n<p>This is better for two reasons: timestamp is a more reliable way to track concurrent<br>\n tasks than task_container_id, the latter is not always in order and continuous; second, task_con would be unique for every task container all over the dataset, which would simplify feature engineering for features that uses task_container.</p>\n<ul>\n<li><p>A faster way to compute rolling average. Pandas standard way:  <code>df.groupby(\"user_id\")[\"feature\"].rolling(window=5, min_periods=0).mean().reset_index(0,drop=True)</code> which takes forever on big dataframes. So I came up with this:</p>\n<pre><code>  window = 5\n  denom = (df.groupby(\"user_id\")[\"answered_correctly\"].cumcount() + 1).clip(upper=window)\n  df[\"rolled_feature\"] = (df.groupby(\"user_id\")[\"answered_correctly\"].cumsum() - df.groupby(\"user_id\")[\"answered_correctly\"].cumsum().groupby(df[\"user_id\"]).shift(window).fillna(0))/denom\n</code></pre></li>\n</ul>\n<p>It seems complicated but it would make sense if you look at it long enough. In terms of speed: 2m 16s vs 1.14s for 5 Million rows, tried to test it on 100 row but it took +∞ for the first method and 23.9 for mine lol.</p>\n<p>Two more days, best of luck!</p>",
  "messages": [
    {
      "id": "1139692",
      "postDate": "01/05/2021 15:06:59",
      "content": "<p>I thought I'd share two Pandas tricks that might be useful for you.</p>\n<ul>\n<li><p>To avoid data leakage, I saw this suggestion in the comments of a discussion here <code>train.groupby([\"user_id\",\"task_container_id\"])[\"feature\"].transform(\"first\")</code> which is quite elegant, but this one is less buggy: </p>\n<pre><code>train[\"task_con\"] = (train.timestamp != train.timestamp.shift()).cumsum()\n train[\"feature\"] = train.groupby([\"user_id\",\"task_con\"])[\"feature\"].transform(\"first\")\n</code></pre></li>\n</ul>\n<p>This is better for two reasons: timestamp is a more reliable way to track concurrent<br>\n tasks than task_container_id, the latter is not always in order and continuous; second, task_con would be unique for every task container all over the dataset, which would simplify feature engineering for features that uses task_container.</p>\n<ul>\n<li><p>A faster way to compute rolling average. Pandas standard way:  <code>df.groupby(\"user_id\")[\"feature\"].rolling(window=5, min_periods=0).mean().reset_index(0,drop=True)</code> which takes forever on big dataframes. So I came up with this:</p>\n<pre><code>  window = 5\n  denom = (df.groupby(\"user_id\")[\"answered_correctly\"].cumcount() + 1).clip(upper=window)\n  df[\"rolled_feature\"] = (df.groupby(\"user_id\")[\"answered_correctly\"].cumsum() - df.groupby(\"user_id\")[\"answered_correctly\"].cumsum().groupby(df[\"user_id\"]).shift(window).fillna(0))/denom\n</code></pre></li>\n</ul>\n<p>It seems complicated but it would make sense if you look at it long enough. In terms of speed: 2m 16s vs 1.14s for 5 Million rows, tried to test it on 100 row but it took +∞ for the first method and 23.9 for mine lol.</p>\n<p>Two more days, best of luck!</p>",
      "rawMarkdown": "I thought I'd share two Pandas tricks that might be useful for you.\n\n- To avoid data leakage, I saw this suggestion in the comments of a discussion here `train.groupby([\"user_id\",\"task_container_id\"])[\"feature\"].transform(\"first\")` which is quite elegant, but this one is less buggy: \n     \n        train[\"task_con\"] = (train.timestamp != train.timestamp.shift()).cumsum()\n         train[\"feature\"] = train.groupby([\"user_id\",\"task_con\"])[\"feature\"].transform(\"first\")\n\nThis is better for two reasons: timestamp is a more reliable way to track concurrent\n tasks than task_container_id, the latter is not always in order and continuous; second, task_con would be unique for every task container all over the dataset, which would simplify feature engineering for features that uses task_container.\n\n\n- A faster way to compute rolling average. Pandas standard way:  `df.groupby(\"user_id\")[\"feature\"].rolling(window=5, min_periods=0).mean().reset_index(0,drop=True)` which takes forever on big dataframes. So I came up with this:\n\n          window = 5\n          denom = (df.groupby(\"user_id\")[\"answered_correctly\"].cumcount() + 1).clip(upper=window)\n          df[\"rolled_feature\"] = (df.groupby(\"user_id\")[\"answered_correctly\"].cumsum() - df.groupby(\"user_id\")[\"answered_correctly\"].cumsum().groupby(df[\"user_id\"]).shift(window).fillna(0))/denom\n\nIt seems complicated but it would make sense if you look at it long enough. In terms of speed: 2m 16s vs 1.14s for 5 Million rows, tried to test it on 100 row but it took +∞ for the first method and 23.9 for mine lol.\n\nTwo more days, best of luck!",
      "votes": null
    },
    {
      "id": "1139961",
      "postDate": "01/05/2021 18:08:32",
      "content": "<p>What does \"con\" mean (in \"task_con\")?</p>",
      "rawMarkdown": "What does \"con\" mean (in \"task_con\")?",
      "votes": null
    },
    {
      "id": "1139988",
      "postDate": "01/05/2021 18:23:00",
      "content": "<p>Just curious, did rolling avg's help you?</p>",
      "rawMarkdown": "Just curious, did rolling avg's help you?",
      "votes": null
    },
    {
      "id": "1140004",
      "postDate": "01/05/2021 18:29:32",
      "content": "<p>I used few rolling average window values as features, as well as exponentially weighted average, and yes, for my case they helped.</p>",
      "rawMarkdown": "I used few rolling average window values as features, as well as exponentially weighted average, and yes, for my case they helped.",
      "votes": null
    },
    {
      "id": "1140008",
      "postDate": "01/05/2021 18:30:37",
      "content": "<p>Its just a random name but I meant it as a shorthand version of task container</p>",
      "rawMarkdown": "Its just a random name but I meant it as a shorthand version of task container",
      "votes": null
    },
    {
      "id": "1141025",
      "postDate": "01/06/2021 12:44:20",
      "content": "<p><code>a is df</code>, right?</p>",
      "rawMarkdown": "`a is df`, right?",
      "votes": null
    },
    {
      "id": "1141036",
      "postDate": "01/06/2021 12:53:38",
      "content": "<p>thanks for sharing these ideas</p>",
      "rawMarkdown": "thanks for sharing these ideas",
      "votes": null
    },
    {
      "id": "1141042",
      "postDate": "01/06/2021 12:55:56",
      "content": "<p>Yes exactly, I have updated my post, thanks for pointing this out!</p>",
      "rawMarkdown": "Yes exactly, I have updated my post, thanks for pointing this out!",
      "votes": null
    },
    {
      "id": "1141045",
      "postDate": "01/06/2021 12:56:19",
      "content": "<p>You're very much welcome, good luck!</p>",
      "rawMarkdown": "You're very much welcome, good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1139961,
      "author_name": "nyakaggle",
      "author_url": "",
      "post_date": "01/05/2021 18:08:32",
      "content": "<p>What does \"con\" mean (in \"task_con\")?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1140008,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/05/2021 18:30:37",
          "content": "<p>Its just a random name but I meant it as a shorthand version of task container</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139988,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/05/2021 18:23:00",
      "content": "<p>Just curious, did rolling avg's help you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1140004,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/05/2021 18:29:32",
          "content": "<p>I used few rolling average window values as features, as well as exponentially weighted average, and yes, for my case they helped.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1141025,
      "author_name": "neustroev",
      "author_url": "",
      "post_date": "01/06/2021 12:44:20",
      "content": "<p><code>a is df</code>, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1141042,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/06/2021 12:55:56",
          "content": "<p>Yes exactly, I have updated my post, thanks for pointing this out!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1141036,
      "author_name": "mohamadnawfal",
      "author_url": "",
      "post_date": "01/06/2021 12:53:38",
      "content": "<p>thanks for sharing these ideas</p>",
      "votes": null,
      "replies": [
        {
          "id": 1141045,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/06/2021 12:56:19",
          "content": "<p>You're very much welcome, good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1139692": "I thought I'd share two Pandas tricks that might be useful for you.\n\n- To avoid data leakage, I saw this suggestion in the comments of a discussion here `train.groupby([\"user_id\",\"task_container_id\"])[\"feature\"].transform(\"first\")` which is quite elegant, but this one is less buggy: \n     \n        train[\"task_con\"] = (train.timestamp != train.timestamp.shift()).cumsum()\n         train[\"feature\"] = train.groupby([\"user_id\",\"task_con\"])[\"feature\"].transform(\"first\")\n\nThis is better for two reasons: timestamp is a more reliable way to track concurrent\n tasks than task_container_id, the latter is not always in order and continuous; second, task_con would be unique for every task container all over the dataset, which would simplify feature engineering for features that uses task_container.\n\n\n- A faster way to compute rolling average. Pandas standard way:  `df.groupby(\"user_id\")[\"feature\"].rolling(window=5, min_periods=0).mean().reset_index(0,drop=True)` which takes forever on big dataframes. So I came up with this:\n\n          window = 5\n          denom = (df.groupby(\"user_id\")[\"answered_correctly\"].cumcount() + 1).clip(upper=window)\n          df[\"rolled_feature\"] = (df.groupby(\"user_id\")[\"answered_correctly\"].cumsum() - df.groupby(\"user_id\")[\"answered_correctly\"].cumsum().groupby(df[\"user_id\"]).shift(window).fillna(0))/denom\n\nIt seems complicated but it would make sense if you look at it long enough. In terms of speed: 2m 16s vs 1.14s for 5 Million rows, tried to test it on 100 row but it took +∞ for the first method and 23.9 for mine lol.\n\nTwo more days, best of luck!",
    "1139961": "What does \"con\" mean (in \"task_con\")?",
    "1139988": "Just curious, did rolling avg's help you?",
    "1140004": "I used few rolling average window values as features, as well as exponentially weighted average, and yes, for my case they helped.",
    "1140008": "Its just a random name but I meant it as a shorthand version of task container",
    "1141025": "`a is df`, right?",
    "1141036": "thanks for sharing these ideas",
    "1141042": "Yes exactly, I have updated my post, thanks for pointing this out!",
    "1141045": "You're very much welcome, good luck!"
  },
  "source": "meta"
}