{
  "id": 205016,
  "title": "Time's close, let's speed up your `groupby()` operation!",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205016",
  "author_name": "",
  "post_date": "2020-12-18T02:53:49.871637300Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<pre><code>a = pd.DataFrame(np.random.randn(100000, 3), columns=['a', 'b', 'c'])\na['a'] = np.random.randint(0, 800, 100000)\na['b'] = np.random.randint(0, 800, 100000)\n\na['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n\n%timeit a.groupby(['a', 'b'])['c'].cumsum()\n%timeit a.groupby(['a_b'])['c'].cumsum()\n%timeit a.groupby(['a_b_category'])['c'].cumsum()\n\n&gt;&gt;&gt;\n17.6 ms ± 157 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n112 ms ± 4.87 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n6.28 ms ± 107 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n</code></pre>\n<ul>\n<li><code>a</code> could be <code>user_id</code></li>\n<li><code>b</code> could be <code>content_id</code></li>\n<li><code>c</code> could be <code>answered_correctly</code></li>\n</ul>\n<p><strong>Conclusion: Concatenating the features and make it <code>category</code> type will save your time!</strong></p>",
  "messages": [
    {
      "id": "1117414",
      "postDate": "12/18/2020 02:53:49",
      "content": "<pre><code>a = pd.DataFrame(np.random.randn(100000, 3), columns=['a', 'b', 'c'])\na['a'] = np.random.randint(0, 800, 100000)\na['b'] = np.random.randint(0, 800, 100000)\n\na['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n\n%timeit a.groupby(['a', 'b'])['c'].cumsum()\n%timeit a.groupby(['a_b'])['c'].cumsum()\n%timeit a.groupby(['a_b_category'])['c'].cumsum()\n\n&gt;&gt;&gt;\n17.6 ms ± 157 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n112 ms ± 4.87 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n6.28 ms ± 107 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n</code></pre>\n<ul>\n<li><code>a</code> could be <code>user_id</code></li>\n<li><code>b</code> could be <code>content_id</code></li>\n<li><code>c</code> could be <code>answered_correctly</code></li>\n</ul>\n<p><strong>Conclusion: Concatenating the features and make it <code>category</code> type will save your time!</strong></p>",
      "rawMarkdown": "```\na = pd.DataFrame(np.random.randn(100000, 3), columns=['a', 'b', 'c'])\na['a'] = np.random.randint(0, 800, 100000)\na['b'] = np.random.randint(0, 800, 100000)\n\na['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n\n%timeit a.groupby(['a', 'b'])['c'].cumsum()\n%timeit a.groupby(['a_b'])['c'].cumsum()\n%timeit a.groupby(['a_b_category'])['c'].cumsum()\n\n>>>\n17.6 ms ± 157 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n112 ms ± 4.87 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n6.28 ms ± 107 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n\n```\n\n- `a` could be `user_id`\n- `b` could be `content_id`\n- `c` could be `answered_correctly`\n\n\n**Conclusion: Concatenating the features and make it `category` type will save your time!**",
      "votes": null
    },
    {
      "id": "1117615",
      "postDate": "12/18/2020 09:03:40",
      "content": "<p>For a one-off groupby like this, I think it's better add the preprocessing time to the benchmarking as well.  That is, don't just measure the time of</p>\n<pre><code>a.groupby(['a_b_category'])['c'].cumsum()\n</code></pre>\n<p>but also measure the time it takes to do the</p>\n<pre><code>a['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n</code></pre>\n<p>And maybe also measure the memory consumption.</p>",
      "rawMarkdown": "For a one-off groupby like this, I think it's better add the preprocessing time to the benchmarking as well.  That is, don't just measure the time of\n```python\na.groupby(['a_b_category'])['c'].cumsum()\n```\nbut also measure the time it takes to do the\n```\na['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n```\n\nAnd maybe also measure the memory consumption.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1117615,
      "author_name": "christoffer",
      "author_url": "",
      "post_date": "12/18/2020 09:03:40",
      "content": "<p>For a one-off groupby like this, I think it's better add the preprocessing time to the benchmarking as well.  That is, don't just measure the time of</p>\n<pre><code>a.groupby(['a_b_category'])['c'].cumsum()\n</code></pre>\n<p>but also measure the time it takes to do the</p>\n<pre><code>a['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n</code></pre>\n<p>And maybe also measure the memory consumption.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1117414": "```\na = pd.DataFrame(np.random.randn(100000, 3), columns=['a', 'b', 'c'])\na['a'] = np.random.randint(0, 800, 100000)\na['b'] = np.random.randint(0, 800, 100000)\n\na['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n\n%timeit a.groupby(['a', 'b'])['c'].cumsum()\n%timeit a.groupby(['a_b'])['c'].cumsum()\n%timeit a.groupby(['a_b_category'])['c'].cumsum()\n\n>>>\n17.6 ms ± 157 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n112 ms ± 4.87 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n6.28 ms ± 107 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)\n\n```\n\n- `a` could be `user_id`\n- `b` could be `content_id`\n- `c` could be `answered_correctly`\n\n\n**Conclusion: Concatenating the features and make it `category` type will save your time!**",
    "1117615": "For a one-off groupby like this, I think it's better add the preprocessing time to the benchmarking as well.  That is, don't just measure the time of\n```python\na.groupby(['a_b_category'])['c'].cumsum()\n```\nbut also measure the time it takes to do the\n```\na['a_b'] = a['a'].astype(\"str\") + \"_\" + a['b'].astype(\"str\")\na['a_b_category'] = a['a_b'].astype(\"category\")\n```\n\nAnd maybe also measure the memory consumption."
  },
  "source": "meta"
}