{
  "id": 71398,
  "title": "Write efficient code!",
  "url": "/competitions/PLAsTiCC-2018/discussion/71398",
  "author_name": "CPMP",
  "post_date": "2018-11-13T11:27:44.931000",
  "votes": 60,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Feature engineering is the crux of how to solve the problem if you're not using deep learning.  Tuning feature engineering code for speed is key:\n- Feature engineering can dramatically increase your running time if you aren't tuning your code. <br>\n- Having a faster code lets you do more experiments within the time you can devote to the competition.  Therefore you can try more ideas.</p>\n\n<p>For instance, my current best model takes about 12 seconds on my machine (i7 4.2 GHz) to generate about 110 features for the training data, from scratch.</p>\n\n<p>Here is an example of how to tune code.  An example that many did not see probably as it is <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538\">buried in another topic</a>.  I reproduce it here.  </p>\n\n<p>The feature we want to compute was disclosed by Grzegorz Sionkovski.  It is the difference between the latest and earliest time at which a flux is detected.  Here is how he stated it:</p>\n\n<pre><code>dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n</code></pre>\n\n<p>Let's see how one can implement it with pandas.</p>\n\n<p>First thing is to create a feature to represent the times when the flux is detected:</p>\n\n<pre><code>df['mjd_detected'] = np.NaN\ndf.loc[df.detected == 1, 'mjd_detected'] = df.loc[df.detected == 1, 'mjd']\n</code></pre>\n\n<p>Then we must compute the difference between its min and max values, for each object_id.</p>\n\n<p>A seemingly nice way to do it is to use <code>groupby()</code> and <code>apply()</code>:</p>\n\n<pre><code>train['mjd_diff'] = train.groupby('object_id').mjd_detected.apply(lambda x: x.ptp())\n</code></pre>\n\n<p>where <code>ptp</code> stands for \"peak to peak\".</p>\n\n<p>Issue is that <code>apply()</code> is slow and should be avoided as plague. If you want a faster code, here is one:</p>\n\n<pre><code>gr_mjd = train.groupby('object_id').mjd_detected\ntrain['mjd_diff']  = gr_mjd.transform('max') - gr_mjd.transform('min')\n</code></pre>\n\n<p>There are two optimizations here: use <code>transform()</code> instead of <code>apply()</code>, and compute <code>groupby('object_id').mjd_detected</code> only once.  </p>\n\n<p>On my machine, the code with apply takes 941 ms while without it it takes 30 ms. A 30x speedup ;)</p>\n\n<p>Unfortunately we cannot use <code>transform('ptp')</code>. It would be even faster. Not sure why it isn't supported.</p>",
  "messages": [
    {
      "id": 420268,
      "postDate": "2018-11-13T11:27:44.930Z",
      "content": "<p>Feature engineering is the crux of how to solve the problem if you're not using deep learning.  Tuning feature engineering code for speed is key:\n- Feature engineering can dramatically increase your running time if you aren't tuning your code. <br>\n- Having a faster code lets you do more experiments within the time you can devote to the competition.  Therefore you can try more ideas.</p>\n\n<p>For instance, my current best model takes about 12 seconds on my machine (i7 4.2 GHz) to generate about 110 features for the training data, from scratch.</p>\n\n<p>Here is an example of how to tune code.  An example that many did not see probably as it is <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538\">buried in another topic</a>.  I reproduce it here.  </p>\n\n<p>The feature we want to compute was disclosed by Grzegorz Sionkovski.  It is the difference between the latest and earliest time at which a flux is detected.  Here is how he stated it:</p>\n\n<pre><code>dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n</code></pre>\n\n<p>Let's see how one can implement it with pandas.</p>\n\n<p>First thing is to create a feature to represent the times when the flux is detected:</p>\n\n<pre><code>df['mjd_detected'] = np.NaN\ndf.loc[df.detected == 1, 'mjd_detected'] = df.loc[df.detected == 1, 'mjd']\n</code></pre>\n\n<p>Then we must compute the difference between its min and max values, for each object_id.</p>\n\n<p>A seemingly nice way to do it is to use <code>groupby()</code> and <code>apply()</code>:</p>\n\n<pre><code>train['mjd_diff'] = train.groupby('object_id').mjd_detected.apply(lambda x: x.ptp())\n</code></pre>\n\n<p>where <code>ptp</code> stands for \"peak to peak\".</p>\n\n<p>Issue is that <code>apply()</code> is slow and should be avoided as plague. If you want a faster code, here is one:</p>\n\n<pre><code>gr_mjd = train.groupby('object_id').mjd_detected\ntrain['mjd_diff']  = gr_mjd.transform('max') - gr_mjd.transform('min')\n</code></pre>\n\n<p>There are two optimizations here: use <code>transform()</code> instead of <code>apply()</code>, and compute <code>groupby('object_id').mjd_detected</code> only once.  </p>\n\n<p>On my machine, the code with apply takes 941 ms while without it it takes 30 ms. A 30x speedup ;)</p>\n\n<p>Unfortunately we cannot use <code>transform('ptp')</code>. It would be even faster. Not sure why it isn't supported.</p>",
      "rawMarkdown": "Feature engineering is the crux of how to solve the problem if you're not using deep learning.  Tuning feature engineering code for speed is key:\n- Feature engineering can dramatically increase your running time if you aren't tuning your code.  \n- Having a faster code lets you do more experiments within the time you can devote to the competition.  Therefore you can try more ideas.\n\nFor instance, my current best model takes about 12 seconds on my machine (i7 4.2 GHz) to generate about 110 features for the training data, from scratch.\n\nHere is an example of how to tune code.  An example that many did not see probably as it is [buried in another topic][1].  I reproduce it here.  \n\nThe feature we want to compute was disclosed by Grzegorz Sionkovski.  It is the difference between the latest and earliest time at which a flux is detected.  Here is how he stated it:\n\n    dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n\nLet's see how one can implement it with pandas.\n\nFirst thing is to create a feature to represent the times when the flux is detected:\n\n    df['mjd_detected'] = np.NaN\n    df.loc[df.detected == 1, 'mjd_detected'] = df.loc[df.detected == 1, 'mjd']\n\nThen we must compute the difference between its min and max values, for each object_id.\n\nA seemingly nice way to do it is to use `groupby()` and `apply()`:\n\n    train['mjd_diff'] = train.groupby('object_id').mjd_detected.apply(lambda x: x.ptp())\n\nwhere `ptp` stands for \"peak to peak\".\n\nIssue is that `apply()` is slow and should be avoided as plague. If you want a faster code, here is one:\n\n    gr_mjd = train.groupby('object_id').mjd_detected\n    train['mjd_diff']  = gr_mjd.transform('max') - gr_mjd.transform('min')\n\nThere are two optimizations here: use `transform()` instead of `apply()`, and compute `groupby('object_id').mjd_detected` only once.  \n\nOn my machine, the code with apply takes 941 ms while without it it takes 30 ms. A 30x speedup ;)\n\nUnfortunately we cannot use `transform('ptp')`. It would be even faster. Not sure why it isn't supported.\n\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538",
      "votes": 60
    },
    {
      "id": 420413,
      "postDate": "2018-11-13T15:38:20.917Z",
      "content": "<p>Great sage advice <a href=\"/cpmpml\">@cpmpml</a> .</p>\n\n<p>Some other tricks I use are:</p>\n\n<ul>\n<li>The <code>sort=False</code> flag on <code>.groupby()</code>; no need to burn cycles sorting by the group unless desired.</li>\n<li>Also <code>as_index=False</code> flag if the aggregate I'm creating needs to be further worked on in its own table (as opposed to <code>.reset_index()</code> or <code>.to_frame()</code>.</li>\n<li>Finally, <code>in_place=True</code> wherever supported to reduce memcopying; but be careful and make sure you really wanna do it in place, especially if you have further analysis that depends on the original data.</li>\n</ul>",
      "rawMarkdown": "Great sage advice @cpmpml .\n\nSome other tricks I use are:\n\n- The `sort=False` flag on `.groupby()`; no need to burn cycles sorting by the group unless desired.\n- Also `as_index=False` flag if the aggregate I'm creating needs to be further worked on in its own table (as opposed to `.reset_index()` or `.to_frame()`.\n- Finally, `in_place=True` wherever supported to reduce memcopying; but be careful and make sure you really wanna do it in place, especially if you have further analysis that depends on the original data.",
      "votes": 13
    },
    {
      "id": 421544,
      "postDate": "2018-11-15T05:36:49.823Z",
      "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> for the trick.</p>\n\n<p>Just in case you want to go a little faster, no need to create the <code>mjd_detected</code> feature</p>\n\n<pre><code>grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].groupby('object_id')['mjd']\ndf['test']  = grp.transform('max') - grp.transform('min')\n</code></pre>\n\n<p>92ms on kaggle kernel instead of 120ms.</p>\n\n<p>For people using linux <a href=\"https://github.com/modin-project/modin\">modin</a> may help you a bit, I believe it provides parallel apply.</p>",
      "rawMarkdown": "Thanks @cpmpml for the trick.\n\nJust in case you want to go a little faster, no need to create the `mjd_detected` feature\n\n    grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].groupby('object_id')['mjd']\n    df['test']  = grp.transform('max') - grp.transform('min')\n\n\n92ms on kaggle kernel instead of 120ms.\n\nFor people using linux [modin](https://github.com/modin-project/modin) may help you a bit, I believe it provides parallel apply.",
      "votes": 8,
      "replies": [
        {
          "id": 421556,
          "postDate": "2018-11-15T05:52:22.127Z",
          "content": "<p>Thanks, This runs in 49 ms vs 66 ms on my machine.  I've learned something else today!</p>",
          "rawMarkdown": "Thanks, This runs in 49 ms vs 66 ms on my machine.  I've learned something else today!",
          "votes": 2
        }
      ]
    },
    {
      "id": 422137,
      "postDate": "2018-11-15T20:33:17.173Z",
      "content": "<p>Code optimization reminds me my youth. \nRegarding calculating mjd_diff, may I suggest you an approach using the properties of the data? As far as I know, simulated measurements are sorted by objectid and mjd, so you do not have to find max and min values of mjd, but just take last and first one:</p>\n\n<pre><code>dt[detected==1,mjd_diff:=mjd[length(mjd)]-mjd[1],by=object_id]\n</code></pre>",
      "rawMarkdown": "Code optimization reminds me my youth. \nRegarding calculating mjd_diff, may I suggest you an approach using the properties of the data? As far as I know, simulated measurements are sorted by objectid and mjd, so you do not have to find max and min values of mjd, but just take last and first one:\n\n    dt[detected==1,mjd_diff:=mjd[length(mjd)]-mjd[1],by=object_id]",
      "votes": 5,
      "replies": [
        {
          "id": 422373,
          "postDate": "2018-11-16T06:14:03.843Z",
          "content": "<p>Sure, but this works only for this specific feature.  The discussion is about writing efficient pandas code in general.</p>",
          "rawMarkdown": "Sure, but this works only for this specific feature.  The discussion is about writing efficient pandas code in general."
        }
      ]
    },
    {
      "id": 421430,
      "postDate": "2018-11-15T01:59:18.547Z",
      "content": "<p>Thanks for the reminder CPMP.</p>\n\n<p>I'm addicted to pivot_table rather than group_by and my code is about as fast (faster?) as your optimized code. And more intuitive to me.  It's probably not optimally written.</p>\n\n<p>This is what I've been using:</p>\n\n<pre><code> result1 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=[min,max])\n\n result1.columns = ['min_mjd','max_mjd']\n\n result1 ['detected_duration'] = result1['max_mjd'] - result1['min_mjd']\n</code></pre>\n\n<p>Interestingly, both of the following are slow:</p>\n\n<pre><code>result2 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: x.ptp())\n\ndf['mjd_diff'] = df.groupby('object_id').mjd_detected.apply(np.max)\n</code></pre>\n\n<p>And then this code:</p>\n\n<pre><code> train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=np.max)\n</code></pre>\n\n<p>is a lot faster than this one:</p>\n\n<p>train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: np.max(x))</p>\n\n<p>Not sure why! I guess the conclusion is to find your bottlenecks and to test alternatives when needed.</p>\n\n<p>Also, is there a need to compute mjd_diff for each train row in your code? I think we would only need the value for each object_id.</p>",
      "rawMarkdown": "Thanks for the reminder CPMP.\n\nI'm addicted to pivot_table rather than group_by and my code is about as fast (faster?) as your optimized code. And more intuitive to me.  It's probably not optimally written.\n\nThis is what I've been using:\n\n     result1 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=[min,max])\n    \n     result1.columns = ['min_mjd','max_mjd']\n    \n     result1 ['detected_duration'] = result1['max_mjd'] - result1['min_mjd']\n\nInterestingly, both of the following are slow:\n\n    result2 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: x.ptp())\n\n    df['mjd_diff'] = df.groupby('object_id').mjd_detected.apply(np.max)\n\nAnd then this code:\n\n     train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=np.max)\n\nis a lot faster than this one:\n\n train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: np.max(x))\n\n\nNot sure why! I guess the conclusion is to find your bottlenecks and to test alternatives when needed.\n\nAlso, is there a need to compute mjd_diff for each train row in your code? I think we would only need the value for each object_id.\n\n\n\n",
      "votes": 5,
      "replies": [
        {
          "id": 421547,
          "postDate": "2018-11-15T05:42:10.790Z",
          "content": "<p>Thanks. Pivot_table is significantly faster than groupby.  I learned something today!  </p>\n\n<p>&gt; Also, is there a need to compute mjddiff for each train row in your code?</p>\n\n<p>Not in my code, but the code I saw proposed was doing that with apply().  To your point, I only compute it once per object_id, with:</p>\n\n<pre><code>df['mjd_detected'] = np.NaN\ndf.loc[df.detected == 1, 'mjd_detected'] = df.loc[df.detected == 1, 'mjd']\nagg = df.groupby('object_id').agg({'mjd_detected':['min','max']})\nagg.columns = ['mjd_detected_min', 'mjd_detected_max']\nagg['mjd_diff'] = agg['mjd_detected_max'] - agg['mjd_detected_min']\ndel agg['mjd_detected_max'],  agg['mjd_detected_min']\n</code></pre>\n\n<p>The equivalent way with pivot_table is</p>\n\n<pre><code>result1 = df[df['detected']==1].pivot_table('mjd','object_id',aggfunc=[min,max])\nresult1.columns = ['min_mjd','max_mjd']\nresult1 ['detected_duration'] = result1['max_mjd'] - result1['min_mjd']\ndel result1['max_mjd'], result1['min_mjd']\n</code></pre>\n\n<p>The latter runs in 19.6 ms while the former runs in 33 ms.  </p>\n\n<p>numpy/pandas also has specific code for the most common aggregations.  It is why using <code>'max'</code>, or <code>np.max</code> is faster than <code>lambda(x: np.max(x))</code></p>\n\n<p>PS.  I don't know how you cut and pasted your code, but it seems underscore got transformed into something weird...</p>",
          "rawMarkdown": "Thanks. Pivot_table is significantly faster than groupby.  I learned something today!  \n\n&gt; Also, is there a need to compute mjddiff for each train row in your code?\n\nNot in my code, but the code I saw proposed was doing that with apply().  To your point, I only compute it once per object_id, with:\n\n    df['mjd_detected'] = np.NaN\n    df.loc[df.detected == 1, 'mjd_detected'] = df.loc[df.detected == 1, 'mjd']\n    agg = df.groupby('object_id').agg({'mjd_detected':['min','max']})\n    agg.columns = ['mjd_detected_min', 'mjd_detected_max']\n    agg['mjd_diff'] = agg['mjd_detected_max'] - agg['mjd_detected_min']\n    del agg['mjd_detected_max'],  agg['mjd_detected_min']\n\nThe equivalent way with pivot_table is\n\n    result1 = df[df['detected']==1].pivot_table('mjd','object_id',aggfunc=[min,max])\n    result1.columns = ['min_mjd','max_mjd']\n    result1 ['detected_duration'] = result1['max_mjd'] - result1['min_mjd']\n    del result1['max_mjd'], result1['min_mjd']\n\nThe latter runs in 19.6 ms while the former runs in 33 ms.  \n\nnumpy/pandas also has specific code for the most common aggregations.  It is why using `'max'`, or `np.max` is faster than `lambda(x: np.max(x))`\n\nPS.  I don't know how you cut and pasted your code, but it seems underscore got transformed into something weird...",
          "votes": 2
        },
        {
          "id": 421865,
          "postDate": "2018-11-15T14:11:32.610Z",
          "content": "<p>(look like I'm having trouble with Markdown styling; my blockquotes aren't showing anymore,and underscores create Italic)</p>\n\n<p>Thanks for the feedback.</p>\n\n<p>Looks like you found yet another way to run this, without transform. </p>\n\n<p>I think pivot_table uses groupby under the hood. So you should be able to get the same performance.</p>\n\n<p>Try not creating that extra column (the 2nd line of your code is slow):</p>\n\n<pre><code> agg = df[df.detected == 1].groupby('object_id').agg({'mjd':['min','max']})\n\n agg.columns = ['mjd_detected_min', 'mjd_detected_max']\n\n agg['mjd_diff'] = agg['mjd_detected_max'] - agg['mjd_detected_min']\n\n del agg['mjd_detected_max'],  agg['mjd_detected_min']\n</code></pre>\n\n<p>Edit: Even though we save half the time here by removing that step, I wonder if that will translate to the same savings when we run the on the bigger test data.</p>",
          "rawMarkdown": "(look like I'm having trouble with Markdown styling; my blockquotes aren't showing anymore,and underscores create Italic)\n\nThanks for the feedback.\n\nLooks like you found yet another way to run this, without transform. \n\nI think pivot_table uses groupby under the hood. So you should be able to get the same performance.\n\nTry not creating that extra column (the 2nd line of your code is slow):\n\n     agg = df[df.detected == 1].groupby('object_id').agg({'mjd':['min','max']})\n    \n     agg.columns = ['mjd_detected_min', 'mjd_detected_max']\n    \n     agg['mjd_diff'] = agg['mjd_detected_max'] - agg['mjd_detected_min']\n    \n     del agg['mjd_detected_max'],  agg['mjd_detected_min']\n\n\n\nEdit: Even though we save half the time here by removing that step, I wonder if that will translate to the same savings when we run the on the bigger test data.",
          "votes": 1
        },
        {
          "id": 421926,
          "postDate": "2018-11-15T15:19:21.653Z",
          "content": "<p>Thanks, you found why my code was slower.</p>\n\n<p>All the great feedback I got here from you and others was worth the time spent in writing the post :)</p>\n\n<p>Re markdown, I need to type backquote twice to see it happen.  And to get code style I need to prepend 4 white spaces.</p>",
          "rawMarkdown": "Thanks, you found why my code was slower.\n\nAll the great feedback I got here from you and others was worth the time spent in writing the post :)\n\nRe markdown, I need to type backquote twice to see it happen.  And to get code style I need to prepend 4 white spaces."
        },
        {
          "id": 425036,
          "postDate": "2018-11-21T03:34:43.983Z",
          "content": "<p>SD, I love your pivot table suggestion so I tried it out on the code in my <a href=\"https://www.kaggle.com/jasonduncanwilson/grouping-contiguous-time-series-data\">Grouping contiguous time series data kernel</a> in step 6. Surprisingly I saw that the groupby was actually faster than the pivot_table.</p>\n\n<p>The groupby averaged around 7ms for execution time.\n<code>\nelec_data_grouped = elec_data.groupby(['group_id','dayofweek','month'], as_index=False).agg({'IT_solar_generation':sum, 'utc_timestamp':\"count\"})\n</code></p>\n\n<p>The pivot_table averaged around 10 ms for execution time.\n<code>\nelec_data_grouped = elec_data.pivot_table(index=['group_id','dayofweek','month'], values=['IT_solar_generation', 'utc_timestamp'], aggfunc={'IT_solar_generation':np.sum, 'utc_timestamp':\"count\"})\n</code></p>\n\n<p>I know this is a different variant than the thread here but figured it was worth mentioning that I'm seeing a case where groupby still performs better.</p>",
          "rawMarkdown": "SD, I love your pivot table suggestion so I tried it out on the code in my [Grouping contiguous time series data kernel](https://www.kaggle.com/jasonduncanwilson/grouping-contiguous-time-series-data) in step 6. Surprisingly I saw that the groupby was actually faster than the pivot_table.\n\nThe groupby averaged around 7ms for execution time.\n```\nelec_data_grouped = elec_data.groupby(['group_id','dayofweek','month'], as_index=False).agg({'IT_solar_generation':sum, 'utc_timestamp':\"count\"})\n```\n\nThe pivot_table averaged around 10 ms for execution time.\n```\nelec_data_grouped = elec_data.pivot_table(index=['group_id','dayofweek','month'], values=['IT_solar_generation', 'utc_timestamp'], aggfunc={'IT_solar_generation':np.sum, 'utc_timestamp':\"count\"})\n```\n\nI know this is a different variant than the thread here but figured it was worth mentioning that I'm seeing a case where groupby still performs better.",
          "votes": 1
        },
        {
          "id": 425080,
          "postDate": "2018-11-21T05:29:51.670Z",
          "content": "<p>Thanks for shairng.  Can you try 'sum' instead of a function?  In your code you use funciton sum while with pivot you use np.sum.  That may explain the difference.</p>",
          "rawMarkdown": "Thanks for shairng.  Can you try 'sum' instead of a function?  In your code you use funciton sum while with pivot you use np.sum.  That may explain the difference.",
          "votes": 1
        },
        {
          "id": 425334,
          "postDate": "2018-11-21T13:07:34.300Z",
          "content": "<p>Good suggestion CPMP. I ran it again with and replaced the np.sum with a straight sum in the pivot table example. It did speed up the execution time by roughly 1 ms but that was all the gain I saw. So the groupby is still outperforming the pivot table by about 2 ms.</p>",
          "rawMarkdown": "Good suggestion CPMP. I ran it again with and replaced the np.sum with a straight sum in the pivot table example. It did speed up the execution time by roughly 1 ms but that was all the gain I saw. So the groupby is still outperforming the pivot table by about 2 ms.",
          "votes": 1
        }
      ]
    },
    {
      "id": 420427,
      "postDate": "2018-11-13T15:59:19.297Z",
      "content": "<p>I guess another example of writing efficient code is the loss function. By setting the weights properly, we do not need to write new custom objective and evaluation functions during training. Please correct me if I'm wrong. Thanks!</p>",
      "rawMarkdown": "I guess another example of writing efficient code is the loss function. By setting the weights properly, we do not need to write new custom objective and evaluation functions during training. Please correct me if I'm wrong. Thanks!",
      "votes": 3,
      "replies": [
        {
          "id": 420430,
          "postDate": "2018-11-13T16:11:14.353Z",
          "content": "<p>You are right!</p>",
          "rawMarkdown": "You are right!"
        },
        {
          "id": 420450,
          "postDate": "2018-11-13T16:38:04.317Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 420452,
          "postDate": "2018-11-13T16:39:23.640Z",
          "content": "<p>Thanks for your confirmation and great suggestions!</p>",
          "rawMarkdown": "Thanks for your confirmation and great suggestions!"
        },
        {
          "id": 421910,
          "postDate": "2018-11-15T14:58:00.970Z",
          "content": "<p>Do you mean using the sample_weight parameter for sklearn.metrics.log_loss ?</p>",
          "rawMarkdown": "Do you mean using the sample_weight parameter for sklearn.metrics.log_loss ?"
        },
        {
          "id": 422136,
          "postDate": "2018-11-15T20:32:03.857Z",
          "content": "<p>I also just found out about the weight parameter when creating lgbm datasets:</p>\n\n<p>train_data = lgb.Dataset(data, label=label, weight=w)</p>\n\n<p>So I guess we could use 'metric': 'multi_logloss', add the weights in the dataset, and voila?</p>",
          "rawMarkdown": "I also just found out about the weight parameter when creating lgbm datasets:\n\ntrain_data = lgb.Dataset(data, label=label, weight=w)\n\nSo I guess we could use 'metric': 'multi_logloss', add the weights in the dataset, and voila?"
        }
      ]
    },
    {
      "id": 430428,
      "postDate": "2018-11-30T11:17:02.820Z",
      "content": "<p>Thanks a lot for these tipps. I have been a notorious <code>groupby()</code>-<code>apply()</code>-offender.</p>",
      "rawMarkdown": "Thanks a lot for these tipps. I have been a notorious `groupby()`-`apply()`-offender.",
      "votes": 1
    },
    {
      "id": 424574,
      "postDate": "2018-11-20T10:45:53.957Z",
      "content": "<p>Thanks for the tips. I'm also curious about the total running time for a single model. What is your total running time for a single model in your own computer?</p>",
      "rawMarkdown": "Thanks for the tips. I'm also curious about the total running time for a single model. What is your total running time for a single model in your own computer?",
      "votes": 1,
      "replies": [
        {
          "id": 424591,
          "postDate": "2018-11-20T11:36:30.150Z",
          "content": "<p>Few hours from scratch, much less is I reuse previously computed features.</p>",
          "rawMarkdown": "Few hours from scratch, much less is I reuse previously computed features.",
          "votes": 1
        }
      ]
    },
    {
      "id": 422013,
      "postDate": "2018-11-15T17:09:05.030Z",
      "content": "<p>This is a great thread; thanks for kicking it off!</p>",
      "rawMarkdown": "This is a great thread; thanks for kicking it off!",
      "votes": 1,
      "replies": [
        {
          "id": 422066,
          "postDate": "2018-11-15T18:29:05.403Z",
          "content": "<p>Thanks, you're welcome!</p>",
          "rawMarkdown": "Thanks, you're welcome!"
        }
      ]
    },
    {
      "id": 421950,
      "postDate": "2018-11-15T15:43:58.163Z",
      "content": "<p>One other option is too use numpy along with numba. That opens door to usage of more numpy functions , e.g. np.ptp</p>\n\n<p>For example, this code </p>\n\n<p>```</p>\n\n<pre><code>@numba.jit\n\ndef get_splits(a):\n\n    m = np.concatenate([[True], a[1:] != a[:-1], [True]])\n\n    m = np.flatnonzero(m)\n\n    return m\n\n @numba.jit(parallel=True)\n\n def grp_ptp(a, b):\n\n     m = get_splits(a)\n\n     n = len(m)-1\n\n     c = np.empty((n, ), dtype=np.float32)\n\n    for i in prange(n):\n\n        x = b[m[i]:m[i+1]]\n\n        c[i] = np.ptp(x)\n\n    return c \n</code></pre>\n\n<p>```</p>\n\n<p>runs in 95.1 s ± 2.1 ms on kaggle kernel compared to 1.53 s ± 31.8 ms for df.groupby(\"object_id\")[\"mjd\"].apply(np.ptp)</p>\n\n<p>similar implementation of max-min using numpy + numba runs in same time as df.groupby(cola)[colb].max() - df.groupby(cola)[colb].min()</p>\n\n<p>On the flip side, one has to implement more code (one for each feature) and compile using numba :P</p>",
      "rawMarkdown": "One other option is too use numpy along with numba. That opens door to usage of more numpy functions , e.g. np.ptp\n\nFor example, this code \n\n```\n\n    @numba.jit\n\n    def get_splits(a):\n\n        m = np.concatenate([[True], a[1:] != a[:-1], [True]])\n\n        m = np.flatnonzero(m)\n\n        return m\n\n     @numba.jit(parallel=True)\n\n     def grp_ptp(a, b):\n\n         m = get_splits(a)\n\n         n = len(m)-1\n\n         c = np.empty((n, ), dtype=np.float32)\n\n        for i in prange(n):\n\n            x = b[m[i]:m[i+1]]\n\n            c[i] = np.ptp(x)\n\n        return c \n\n```\n\nruns in 95.1 s ± 2.1 ms on kaggle kernel compared to 1.53 s ± 31.8 ms for df.groupby(\"object_id\")[\"mjd\"].apply(np.ptp)\n\n\nsimilar implementation of max-min using numpy + numba runs in same time as df.groupby(cola)[colb].max() - df.groupby(cola)[colb].min()\n\nOn the flip side, one has to implement more code (one for each feature) and compile using numba :P",
      "votes": 2,
      "replies": [
        {
          "id": 421952,
          "postDate": "2018-11-15T15:49:20.247Z",
          "content": "<p>From what you say compared to apply, then using groupby without apply as discussed elsewhere in this topic is probably way faster than your numba code.  But using numba when groupby isn't applicable is certainly the way to go.</p>",
          "rawMarkdown": "From what you say compared to apply, then using groupby without apply as discussed elsewhere in this topic is probably way faster than your numba code.  But using numba when groupby isn't applicable is certainly the way to go."
        },
        {
          "id": 421961,
          "postDate": "2018-11-15T15:58:35.603Z",
          "content": "<p>```</p>\n\n<pre><code>@numba.jit(parallel=True)\ndef grp_max_min(a, b):\n    m = get_splits(a)\n    n = len(m)-1\n    c = np.empty((n, ), dtype=np.float32)\n    for i in prange(n):\n        x = b[m[i]:m[i+1]]\n        c[i] = np.max(x) - np.min(x)\n    return c\n\ngrp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].values\nmjddiff = grp_max_min(grp[:, 0], grp[:, 1])\n</code></pre>\n\n<p>```</p>\n\n<p>this runs in 24 ms for me faster compared to</p>\n\n<p>```</p>\n\n<pre><code>grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].groupby('object_id')['mjd']\nmjddiff  = grp.max() - grp.min()\n</code></pre>\n\n<p>```\nwhich takes 26 ms</p>",
          "rawMarkdown": "```\n\n    @numba.jit(parallel=True)\n    def grp_max_min(a, b):\n        m = get_splits(a)\n        n = len(m)-1\n        c = np.empty((n, ), dtype=np.float32)\n        for i in prange(n):\n            x = b[m[i]:m[i+1]]\n            c[i] = np.max(x) - np.min(x)\n        return c\n\n    grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].values\n    mjddiff = grp_max_min(grp[:, 0], grp[:, 1])\n```\n\nthis runs in 24 ms for me faster compared to\n\n\n```\n\n    grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].groupby('object_id')['mjd']\n    mjddiff  = grp.max() - grp.min()\n\n```\nwhich takes 26 ms",
          "votes": 3
        },
        {
          "id": 421971,
          "postDate": "2018-11-15T16:08:48.220Z",
          "content": "<p>Numba has native implementations for max, min (and many other numpy functions) so above is 3x faster than calling np.ptp. Similarly, pandas has cython optimized functions for max and min which makes them 30x faster than a custom function in apply. </p>\n\n<p>In cases, where optimized functions are not available, using numba with numpy can still be a lot faster for than using pandas (10x in above example), ofcourse at the cost of writing more lines of code</p>",
          "rawMarkdown": "Numba has native implementations for max, min (and many other numpy functions) so above is 3x faster than calling np.ptp. Similarly, pandas has cython optimized functions for max and min which makes them 30x faster than a custom function in apply. \n\nIn cases, where optimized functions are not available, using numba with numpy can still be a lot faster for than using pandas (10x in above example), ofcourse at the cost of writing more lines of code",
          "votes": 1
        },
        {
          "id": 422003,
          "postDate": "2018-11-15T16:56:45.513Z",
          "content": "<p>Right, you get about 10% faster with numba.  The nice side effect is that your groupby code is also way faster than what I was using!  And faster than what S D is suggesting.</p>\n\n<p>On my machine, what I am using takes 31 ms, what S D proposed takes 19.6 ms, your groupby takes 13.5 ms, and your numba takes 11.6 ms.  </p>\n\n<p>Thanks, I've learned yet another thing today.</p>",
          "rawMarkdown": "Right, you get about 10% faster with numba.  The nice side effect is that your groupby code is also way faster than what I was using!  And faster than what S D is suggesting.\n\nOn my machine, what I am using takes 31 ms, what S D proposed takes 19.6 ms, your groupby takes 13.5 ms, and your numba takes 11.6 ms.  \n\nThanks, I've learned yet another thing today."
        },
        {
          "id": 422173,
          "postDate": "2018-11-15T21:51:33.247Z",
          "content": "<p>@Mohsin, How about something like:</p>\n\n<pre><code>c[i] = x[len(x)-1] - x[0]\n</code></pre>\n\n<p>instead of:</p>\n\n<pre><code>c[i] = np.max(x) - np.min(x)\n</code></pre>",
          "rawMarkdown": "@Mohsin, How about something like:\n\n    c[i] = x[len(x)-1] - x[0]\n\ninstead of:\n\n    c[i] = np.max(x) - np.min(x)\n",
          "votes": 1
        },
        {
          "id": 422399,
          "postDate": "2018-11-16T07:04:59.883Z",
          "content": "<p>@Gregorz:  yes , Doing that in this particular case would be much faster :)</p>",
          "rawMarkdown": "@Gregorz:  yes , Doing that in this particular case would be much faster :)"
        },
        {
          "id": 422413,
          "postDate": "2018-11-16T07:26:13.223Z",
          "content": "<p>@CPMP :  Actually, i only started using numba after your kernel in instakart competition, so thanks to you!  Also, credit goes to diwakar (<a href=\"https://stackoverflow.com/users/3293881/divakar\">https://stackoverflow.com/users/3293881/divakar</a>) who has lot of good answers for doing stuff fast in numpy. </p>\n\n<p>P.S. - Numpy solution also takes lot less RAM, it is possible to compute basic features for test set under 16 GB budget</p>",
          "rawMarkdown": "@CPMP :  Actually, i only started using numba after your kernel in instakart competition, so thanks to you!  Also, credit goes to diwakar (https://stackoverflow.com/users/3293881/divakar) who has lot of good answers for doing stuff fast in numpy. \n\nP.S. - Numpy solution also takes lot less RAM, it is possible to compute basic features for test set under 16 GB budget",
          "votes": 2
        }
      ]
    },
    {
      "id": 420628,
      "postDate": "2018-11-13T23:04:43.693Z",
      "content": "<p>This article has some other tricks for speeding up calculations in pandas.\n<a href=\"https://realpython.com/fast-flexible-pandas/\">https://realpython.com/fast-flexible-pandas/</a></p>",
      "rawMarkdown": "This article has some other tricks for speeding up calculations in pandas.\nhttps://realpython.com/fast-flexible-pandas/\n\n",
      "votes": 2
    },
    {
      "id": 420334,
      "postDate": "2018-11-13T13:21:35.663Z",
      "content": "<p>Thanks for sharing! This is a crucial yet often overlooked point</p>\n\n<p>Kaggle novice here,  but can attest first-hand at how writing fast feat engineering code has dramatically improved my score by allowing me to \"fail fast, fail often\" and try out new feats</p>\n\n<p>Another possible approach is to precompute and store curve data once in a cesium-like format, thus avoiding groupbys altogether during feat engineering- computing ~120 feats for training set takes me around 15secs with python multiprocessing using this method</p>",
      "rawMarkdown": "Thanks for sharing! This is a crucial yet often overlooked point\n\nKaggle novice here,  but can attest first-hand at how writing fast feat engineering code has dramatically improved my score by allowing me to \"fail fast, fail often\" and try out new feats\n\nAnother possible approach is to precompute and store curve data once in a cesium-like format, thus avoiding groupbys altogether during feat engineering- computing ~120 feats for training set takes me around 15secs with python multiprocessing using this method",
      "votes": 2,
      "replies": [
        {
          "id": 421131,
          "postDate": "2018-11-14T16:16:48.140Z",
          "content": "<p>Absolutely right !</p>\n\n<p>This is the way I always process ;)</p>\n\n<p>Moreover, in this competition with millions of rows, it is insane to try to make it fit in one single kernel.</p>",
          "rawMarkdown": "Absolutely right !\n\nThis is the way I always process ;)\n\nMoreover, in this competition with millions of rows, it is insane to try to make it fit in one single kernel.",
          "votes": 1
        }
      ]
    },
    {
      "id": 465286,
      "postDate": "2019-02-02T18:23:26.153Z",
      "content": "<p>there is also _.apply(raw=True) pass as an ndarray to function which also might speed up in a few seconds</p>",
      "rawMarkdown": "there is also _.apply(raw=True) pass as an ndarray to function which also might speed up in a few seconds"
    },
    {
      "id": 441402,
      "postDate": "2018-12-18T16:09:52.780Z",
      "content": "<p>Not sure whether its apt discussion to ask a  query, but here it's</p>\n\n<pre><code>X[['col1', cc]].groupby(['col']).agg(lambda x: ' '.join(x)).reset_index()\n#(X 's shape (9853253, 16))\n</code></pre>\n\n<p>How to make this faster?</p>\n\n<p>Thanks :)</p>",
      "rawMarkdown": "Not sure whether its apt discussion to ask a  query, but here it's\n\n    X[['col1', cc]].groupby(['col']).agg(lambda x: ' '.join(x)).reset_index()\n    #(X 's shape (9853253, 16))\n\nHow to make this faster?\n\nThanks :)",
      "replies": [
        {
          "id": 441420,
          "postDate": "2018-12-18T16:38:31.237Z",
          "content": "<p>That's how I would write it too.  Let's see if others have better ways.  But did you try this (not sure it is supported):</p>\n\n<p><code>X[['col1', cc]].groupby(['col']).join().reset_index()</code></p>",
          "rawMarkdown": "That's how I would write it too.  Let's see if others have better ways.  But did you try this (not sure it is supported):\n\n`X[['col1', cc]].groupby(['col']).join().reset_index()`"
        },
        {
          "id": 441942,
          "postDate": "2018-12-19T09:18:10.977Z",
          "content": "<p>Didn't try this.. Will try and see if(if its supported) the speed improves</p>",
          "rawMarkdown": "Didn't try this.. Will try and see if(if its supported) the speed improves"
        }
      ]
    },
    {
      "id": 420384,
      "postDate": "2018-11-13T14:41:53.727Z",
      "content": "<p>Hi @CPMP,</p>\n\n<p>Thanks for your example and explanation.</p>\n\n<p>I was curious about the claim we can't use <code>ptp</code> in <code>transform</code> so I tried your example.</p>\n\n<p>You can use it as:\n<code>train['mjd_diff']  = gr_mjd.transform(np.ptp)</code></p>\n\n<p>but that's much slower than your <code>gr_mjd.transform('max') - gr_mjd.transform('min')</code> method (30x on my machine), which shows the importance of examining any performance bottlenecks, even when you think you're using the right method or technique.</p>",
      "rawMarkdown": "Hi @CPMP,\n\nThanks for your example and explanation.\n\nI was curious about the claim we can't use ```ptp``` in ```transform``` so I tried your example.\n\nYou can use it as:\n```train['mjd_diff']  = gr_mjd.transform(np.ptp)```\n\nbut that's much slower than your ```gr_mjd.transform('max') - gr_mjd.transform('min')``` method (30x on my machine), which shows the importance of examining any performance bottlenecks, even when you think you're using the right method or technique.",
      "replies": [
        {
          "id": 420387,
          "postDate": "2018-11-13T14:48:21.697Z",
          "content": "<p>I meant you cannot use <code>transform('ptp')</code>.  Of course, you can pass a function as you did, but then you are back to the slowness of <code>apply()</code>.</p>\n\n<p>The point is to avoid calling a function for each object_id.</p>",
          "rawMarkdown": "I meant you cannot use `transform('ptp')`.  Of course, you can pass a function as you did, but then you are back to the slowness of `apply()`.\n\nThe point is to avoid calling a function for each object_id.",
          "votes": 2
        },
        {
          "id": 420388,
          "postDate": "2018-11-13T14:49:41.090Z",
          "content": "<p>Okay, thanks for that clarification.</p>",
          "rawMarkdown": "Okay, thanks for that clarification.",
          "votes": 1
        },
        {
          "id": 420392,
          "postDate": "2018-11-13T14:59:45.743Z",
          "content": "<p>Thanks, I edited the post to make it clearer.</p>",
          "rawMarkdown": "Thanks, I edited the post to make it clearer."
        }
      ]
    },
    {
      "id": 422269,
      "postDate": "2018-11-16T02:05:36.433Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 420413,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2018-11-13T15:38:20.917000",
      "content": "<p>Great sage advice <a href=\"/cpmpml\">@cpmpml</a> .</p>\n\n<p>Some other tricks I use are:</p>\n\n<ul>\n<li>The <code>sort=False</code> flag on <code>.groupby()</code>; no need to burn cycles sorting by the group unless desired.</li>\n<li>Also <code>as_index=False</code> flag if the aggregate I'm creating needs to be further worked on in its own table (as opposed to <code>.reset_index()</code> or <code>.to_frame()</code>.</li>\n<li>Finally, <code>in_place=True</code> wherever supported to reduce memcopying; but be careful and make sure you really wanna do it in place, especially if you have further analysis that depends on the original data.</li>\n</ul>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 421544,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-11-15T05:36:49.823000",
      "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> for the trick.</p>\n\n<p>Just in case you want to go a little faster, no need to create the <code>mjd_detected</code> feature</p>\n\n<pre><code>grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].groupby('object_id')['mjd']\ndf['test']  = grp.transform('max') - grp.transform('min')\n</code></pre>\n\n<p>92ms on kaggle kernel instead of 120ms.</p>\n\n<p>For people using linux <a href=\"https://github.com/modin-project/modin\">modin</a> may help you a bit, I believe it provides parallel apply.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 421556,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-15T05:52:22.127000",
          "content": "<p>Thanks, This runs in 49 ms vs 66 ms on my machine.  I've learned something else today!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 422137,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-11-15T20:33:17.173000",
      "content": "<p>Code optimization reminds me my youth. \nRegarding calculating mjd_diff, may I suggest you an approach using the properties of the data? As far as I know, simulated measurements are sorted by objectid and mjd, so you do not have to find max and min values of mjd, but just take last and first one:</p>\n\n<pre><code>dt[detected==1,mjd_diff:=mjd[length(mjd)]-mjd[1],by=object_id]\n</code></pre>",
      "votes": 5,
      "replies": [
        {
          "id": 422373,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-16T06:14:03.843000",
          "content": "<p>Sure, but this works only for this specific feature.  The discussion is about writing efficient pandas code in general.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421430,
      "author_name": "S D",
      "author_url": "",
      "post_date": "2018-11-15T01:59:18.547000",
      "content": "<p>Thanks for the reminder CPMP.</p>\n\n<p>I'm addicted to pivot_table rather than group_by and my code is about as fast (faster?) as your optimized code. And more intuitive to me.  It's probably not optimally written.</p>\n\n<p>This is what I've been using:</p>\n\n<pre><code> result1 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=[min,max])\n\n result1.columns = ['min_mjd','max_mjd']\n\n result1 ['detected_duration'] = result1['max_mjd'] - result1['min_mjd']\n</code></pre>\n\n<p>Interestingly, both of the following are slow:</p>\n\n<pre><code>result2 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: x.ptp())\n\ndf['mjd_diff'] = df.groupby('object_id').mjd_detected.apply(np.max)\n</code></pre>\n\n<p>And then this code:</p>\n\n<pre><code> train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=np.max)\n</code></pre>\n\n<p>is a lot faster than this one:</p>\n\n<p>train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: np.max(x))</p>\n\n<p>Not sure why! I guess the conclusion is to find your bottlenecks and to test alternatives when needed.</p>\n\n<p>Also, is there a need to compute mjd_diff for each train row in your code? I think we would only need the value for each object_id.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 421547,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-15T05:42:10.790000",
          "content": "<p>Thanks. Pivot_table is significantly faster than groupby.  I learned something today!  </p>\n\n<p>&gt; Also, is there a need to compute mjddiff for each train row in your code?</p>\n\n<p>Not in my code, but the code I saw proposed was doing that with apply().  To your point, I only compute it once per object_id, with:</p>\n\n<pre><code>df['mjd_detected'] = np.NaN\ndf.loc[df.detected == 1, 'mjd_detected'] = df.loc[df.detected == 1, 'mjd']\nagg = df.groupby('object_id').agg({'mjd_detected':['min','max']})\nagg.columns = ['mjd_detected_min', 'mjd_detected_max']\nagg['mjd_diff'] = agg['mjd_detected_max'] - agg['mjd_detected_min']\ndel agg['mjd_detected_max'],  agg['mjd_detected_min']\n</code></pre>\n\n<p>The equivalent way with pivot_table is</p>\n\n<pre><code>result1 = df[df['detected']==1].pivot_table('mjd','object_id',aggfunc=[min,max])\nresult1.columns = ['min_mjd','max_mjd']\nresult1 ['detected_duration'] = result1['max_mjd'] - result1['min_mjd']\ndel result1['max_mjd'], result1['min_mjd']\n</code></pre>\n\n<p>The latter runs in 19.6 ms while the former runs in 33 ms.  </p>\n\n<p>numpy/pandas also has specific code for the most common aggregations.  It is why using <code>'max'</code>, or <code>np.max</code> is faster than <code>lambda(x: np.max(x))</code></p>\n\n<p>PS.  I don't know how you cut and pasted your code, but it seems underscore got transformed into something weird...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 421865,
          "author_name": "S D",
          "author_url": "",
          "post_date": "2018-11-15T14:11:32.610000",
          "content": "<p>(look like I'm having trouble with Markdown styling; my blockquotes aren't showing anymore,and underscores create Italic)</p>\n\n<p>Thanks for the feedback.</p>\n\n<p>Looks like you found yet another way to run this, without transform. </p>\n\n<p>I think pivot_table uses groupby under the hood. So you should be able to get the same performance.</p>\n\n<p>Try not creating that extra column (the 2nd line of your code is slow):</p>\n\n<pre><code> agg = df[df.detected == 1].groupby('object_id').agg({'mjd':['min','max']})\n\n agg.columns = ['mjd_detected_min', 'mjd_detected_max']\n\n agg['mjd_diff'] = agg['mjd_detected_max'] - agg['mjd_detected_min']\n\n del agg['mjd_detected_max'],  agg['mjd_detected_min']\n</code></pre>\n\n<p>Edit: Even though we save half the time here by removing that step, I wonder if that will translate to the same savings when we run the on the bigger test data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421926,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-15T15:19:21.653000",
          "content": "<p>Thanks, you found why my code was slower.</p>\n\n<p>All the great feedback I got here from you and others was worth the time spent in writing the post :)</p>\n\n<p>Re markdown, I need to type backquote twice to see it happen.  And to get code style I need to prepend 4 white spaces.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425036,
          "author_name": "Jason Duncan-Wilson",
          "author_url": "",
          "post_date": "2018-11-21T03:34:43.983000",
          "content": "<p>SD, I love your pivot table suggestion so I tried it out on the code in my <a href=\"https://www.kaggle.com/jasonduncanwilson/grouping-contiguous-time-series-data\">Grouping contiguous time series data kernel</a> in step 6. Surprisingly I saw that the groupby was actually faster than the pivot_table.</p>\n\n<p>The groupby averaged around 7ms for execution time.\n<code>\nelec_data_grouped = elec_data.groupby(['group_id','dayofweek','month'], as_index=False).agg({'IT_solar_generation':sum, 'utc_timestamp':\"count\"})\n</code></p>\n\n<p>The pivot_table averaged around 10 ms for execution time.\n<code>\nelec_data_grouped = elec_data.pivot_table(index=['group_id','dayofweek','month'], values=['IT_solar_generation', 'utc_timestamp'], aggfunc={'IT_solar_generation':np.sum, 'utc_timestamp':\"count\"})\n</code></p>\n\n<p>I know this is a different variant than the thread here but figured it was worth mentioning that I'm seeing a case where groupby still performs better.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425080,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-21T05:29:51.670000",
          "content": "<p>Thanks for shairng.  Can you try 'sum' instead of a function?  In your code you use funciton sum while with pivot you use np.sum.  That may explain the difference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425334,
          "author_name": "Jason Duncan-Wilson",
          "author_url": "",
          "post_date": "2018-11-21T13:07:34.300000",
          "content": "<p>Good suggestion CPMP. I ran it again with and replaced the np.sum with a straight sum in the pivot table example. It did speed up the execution time by roughly 1 ms but that was all the gain I saw. So the groupby is still outperforming the pivot table by about 2 ms.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 420427,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "2018-11-13T15:59:19.297000",
      "content": "<p>I guess another example of writing efficient code is the loss function. By setting the weights properly, we do not need to write new custom objective and evaluation functions during training. Please correct me if I'm wrong. Thanks!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 420430,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-13T16:11:14.353000",
          "content": "<p>You are right!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420450,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-13T16:38:04.317000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420452,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "2018-11-13T16:39:23.640000",
          "content": "<p>Thanks for your confirmation and great suggestions!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421910,
          "author_name": "S D",
          "author_url": "",
          "post_date": "2018-11-15T14:58:00.970000",
          "content": "<p>Do you mean using the sample_weight parameter for sklearn.metrics.log_loss ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422136,
          "author_name": "S D",
          "author_url": "",
          "post_date": "2018-11-15T20:32:03.857000",
          "content": "<p>I also just found out about the weight parameter when creating lgbm datasets:</p>\n\n<p>train_data = lgb.Dataset(data, label=label, weight=w)</p>\n\n<p>So I guess we could use 'metric': 'multi_logloss', add the weights in the dataset, and voila?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 430428,
      "author_name": "Johannes Nowak",
      "author_url": "",
      "post_date": "2018-11-30T11:17:02.820000",
      "content": "<p>Thanks a lot for these tipps. I have been a notorious <code>groupby()</code>-<code>apply()</code>-offender.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 424574,
      "author_name": "Murat Korkmaz",
      "author_url": "",
      "post_date": "2018-11-20T10:45:53.957000",
      "content": "<p>Thanks for the tips. I'm also curious about the total running time for a single model. What is your total running time for a single model in your own computer?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 424591,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-20T11:36:30.150000",
          "content": "<p>Few hours from scratch, much less is I reuse previously computed features.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 422013,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2018-11-15T17:09:05.030000",
      "content": "<p>This is a great thread; thanks for kicking it off!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 422066,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-15T18:29:05.403000",
          "content": "<p>Thanks, you're welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421950,
      "author_name": "Mohsin hasan",
      "author_url": "",
      "post_date": "2018-11-15T15:43:58.163000",
      "content": "<p>One other option is too use numpy along with numba. That opens door to usage of more numpy functions , e.g. np.ptp</p>\n\n<p>For example, this code </p>\n\n<p>```</p>\n\n<pre><code>@numba.jit\n\ndef get_splits(a):\n\n    m = np.concatenate([[True], a[1:] != a[:-1], [True]])\n\n    m = np.flatnonzero(m)\n\n    return m\n\n @numba.jit(parallel=True)\n\n def grp_ptp(a, b):\n\n     m = get_splits(a)\n\n     n = len(m)-1\n\n     c = np.empty((n, ), dtype=np.float32)\n\n    for i in prange(n):\n\n        x = b[m[i]:m[i+1]]\n\n        c[i] = np.ptp(x)\n\n    return c \n</code></pre>\n\n<p>```</p>\n\n<p>runs in 95.1 s ± 2.1 ms on kaggle kernel compared to 1.53 s ± 31.8 ms for df.groupby(\"object_id\")[\"mjd\"].apply(np.ptp)</p>\n\n<p>similar implementation of max-min using numpy + numba runs in same time as df.groupby(cola)[colb].max() - df.groupby(cola)[colb].min()</p>\n\n<p>On the flip side, one has to implement more code (one for each feature) and compile using numba :P</p>",
      "votes": 2,
      "replies": [
        {
          "id": 421952,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-15T15:49:20.247000",
          "content": "<p>From what you say compared to apply, then using groupby without apply as discussed elsewhere in this topic is probably way faster than your numba code.  But using numba when groupby isn't applicable is certainly the way to go.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421961,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2018-11-15T15:58:35.603000",
          "content": "<p>```</p>\n\n<pre><code>@numba.jit(parallel=True)\ndef grp_max_min(a, b):\n    m = get_splits(a)\n    n = len(m)-1\n    c = np.empty((n, ), dtype=np.float32)\n    for i in prange(n):\n        x = b[m[i]:m[i+1]]\n        c[i] = np.max(x) - np.min(x)\n    return c\n\ngrp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].values\nmjddiff = grp_max_min(grp[:, 0], grp[:, 1])\n</code></pre>\n\n<p>```</p>\n\n<p>this runs in 24 ms for me faster compared to</p>\n\n<p>```</p>\n\n<pre><code>grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].groupby('object_id')['mjd']\nmjddiff  = grp.max() - grp.min()\n</code></pre>\n\n<p>```\nwhich takes 26 ms</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 421971,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2018-11-15T16:08:48.220000",
          "content": "<p>Numba has native implementations for max, min (and many other numpy functions) so above is 3x faster than calling np.ptp. Similarly, pandas has cython optimized functions for max and min which makes them 30x faster than a custom function in apply. </p>\n\n<p>In cases, where optimized functions are not available, using numba with numpy can still be a lot faster for than using pandas (10x in above example), ofcourse at the cost of writing more lines of code</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422003,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-15T16:56:45.513000",
          "content": "<p>Right, you get about 10% faster with numba.  The nice side effect is that your groupby code is also way faster than what I was using!  And faster than what S D is suggesting.</p>\n\n<p>On my machine, what I am using takes 31 ms, what S D proposed takes 19.6 ms, your groupby takes 13.5 ms, and your numba takes 11.6 ms.  </p>\n\n<p>Thanks, I've learned yet another thing today.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422173,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-11-15T21:51:33.247000",
          "content": "<p>@Mohsin, How about something like:</p>\n\n<pre><code>c[i] = x[len(x)-1] - x[0]\n</code></pre>\n\n<p>instead of:</p>\n\n<pre><code>c[i] = np.max(x) - np.min(x)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422399,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2018-11-16T07:04:59.883000",
          "content": "<p>@Gregorz:  yes , Doing that in this particular case would be much faster :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422413,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2018-11-16T07:26:13.223000",
          "content": "<p>@CPMP :  Actually, i only started using numba after your kernel in instakart competition, so thanks to you!  Also, credit goes to diwakar (<a href=\"https://stackoverflow.com/users/3293881/divakar\">https://stackoverflow.com/users/3293881/divakar</a>) who has lot of good answers for doing stuff fast in numpy. </p>\n\n<p>P.S. - Numpy solution also takes lot less RAM, it is possible to compute basic features for test set under 16 GB budget</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 420628,
      "author_name": "Allen",
      "author_url": "",
      "post_date": "2018-11-13T23:04:43.693000",
      "content": "<p>This article has some other tricks for speeding up calculations in pandas.\n<a href=\"https://realpython.com/fast-flexible-pandas/\">https://realpython.com/fast-flexible-pandas/</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 420334,
      "author_name": "Ganfear",
      "author_url": "",
      "post_date": "2018-11-13T13:21:35.663000",
      "content": "<p>Thanks for sharing! This is a crucial yet often overlooked point</p>\n\n<p>Kaggle novice here,  but can attest first-hand at how writing fast feat engineering code has dramatically improved my score by allowing me to \"fail fast, fail often\" and try out new feats</p>\n\n<p>Another possible approach is to precompute and store curve data once in a cesium-like format, thus avoiding groupbys altogether during feat engineering- computing ~120 feats for training set takes me around 15secs with python multiprocessing using this method</p>",
      "votes": 2,
      "replies": [
        {
          "id": 421131,
          "author_name": "mezoganet",
          "author_url": "",
          "post_date": "2018-11-14T16:16:48.140000",
          "content": "<p>Absolutely right !</p>\n\n<p>This is the way I always process ;)</p>\n\n<p>Moreover, in this competition with millions of rows, it is insane to try to make it fit in one single kernel.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 465286,
      "author_name": "Iurii Cojocari",
      "author_url": "",
      "post_date": "2019-02-02T18:23:26.153000",
      "content": "<p>there is also _.apply(raw=True) pass as an ndarray to function which also might speed up in a few seconds</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441402,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2018-12-18T16:09:52.780000",
      "content": "<p>Not sure whether its apt discussion to ask a  query, but here it's</p>\n\n<pre><code>X[['col1', cc]].groupby(['col']).agg(lambda x: ' '.join(x)).reset_index()\n#(X 's shape (9853253, 16))\n</code></pre>\n\n<p>How to make this faster?</p>\n\n<p>Thanks :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 441420,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-18T16:38:31.237000",
          "content": "<p>That's how I would write it too.  Let's see if others have better ways.  But did you try this (not sure it is supported):</p>\n\n<p><code>X[['col1', cc]].groupby(['col']).join().reset_index()</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441942,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2018-12-19T09:18:10.977000",
          "content": "<p>Didn't try this.. Will try and see if(if its supported) the speed improves</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420384,
      "author_name": "Greg Behm",
      "author_url": "",
      "post_date": "2018-11-13T14:41:53.727000",
      "content": "<p>Hi @CPMP,</p>\n\n<p>Thanks for your example and explanation.</p>\n\n<p>I was curious about the claim we can't use <code>ptp</code> in <code>transform</code> so I tried your example.</p>\n\n<p>You can use it as:\n<code>train['mjd_diff']  = gr_mjd.transform(np.ptp)</code></p>\n\n<p>but that's much slower than your <code>gr_mjd.transform('max') - gr_mjd.transform('min')</code> method (30x on my machine), which shows the importance of examining any performance bottlenecks, even when you think you're using the right method or technique.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 420387,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-13T14:48:21.697000",
          "content": "<p>I meant you cannot use <code>transform('ptp')</code>.  Of course, you can pass a function as you did, but then you are back to the slowness of <code>apply()</code>.</p>\n\n<p>The point is to avoid calling a function for each object_id.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 420388,
          "author_name": "Greg Behm",
          "author_url": "",
          "post_date": "2018-11-13T14:49:41.090000",
          "content": "<p>Okay, thanks for that clarification.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 420392,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-13T14:59:45.743000",
          "content": "<p>Thanks, I edited the post to make it clearer.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422269,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-16T02:05:36.433000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "420268": "Feature engineering is the crux of how to solve the problem if you're not using deep learning.  Tuning feature engineering code for speed is key:\n- Feature engineering can dramatically increase your running time if you aren't tuning your code.  \n- Having a faster code lets you do more experiments within the time you can devote to the competition.  Therefore you can try more ideas.\n\nFor instance, my current best model takes about 12 seconds on my machine (i7 4.2 GHz) to generate about 110 features for the training data, from scratch.\n\nHere is an example of how to tune code.  An example that many did not see probably as it is [buried in another topic][1].  I reproduce it here.  \n\nThe feature we want to compute was disclosed by Grzegorz Sionkovski.  It is the difference between the latest and earliest time at which a flux is detected.  Here is how he stated it:\n\n    dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n\nLet's see how one can implement it with pandas.\n\nFirst thing is to create a feature to represent the times when the flux is detected:\n\n    df['mjd_detected'] = np.NaN\n    df.loc[df.detected == 1, 'mjd_detected'] = df.loc[df.detected == 1, 'mjd']\n\nThen we must compute the difference between its min and max values, for each object_id.\n\nA seemingly nice way to do it is to use `groupby()` and `apply()`:\n\n    train['mjd_diff'] = train.groupby('object_id').mjd_detected.apply(lambda x: x.ptp())\n\nwhere `ptp` stands for \"peak to peak\".\n\nIssue is that `apply()` is slow and should be avoided as plague. If you want a faster code, here is one:\n\n    gr_mjd = train.groupby('object_id').mjd_detected\n    train['mjd_diff']  = gr_mjd.transform('max') - gr_mjd.transform('min')\n\nThere are two optimizations here: use `transform()` instead of `apply()`, and compute `groupby('object_id').mjd_detected` only once.  \n\nOn my machine, the code with apply takes 941 ms while without it it takes 30 ms. A 30x speedup ;)\n\nUnfortunately we cannot use `transform('ptp')`. It would be even faster. Not sure why it isn't supported.\n\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538",
    "420413": "Great sage advice @cpmpml .\n\nSome other tricks I use are:\n\n- The `sort=False` flag on `.groupby()`; no need to burn cycles sorting by the group unless desired.\n- Also `as_index=False` flag if the aggregate I'm creating needs to be further worked on in its own table (as opposed to `.reset_index()` or `.to_frame()`.\n- Finally, `in_place=True` wherever supported to reduce memcopying; but be careful and make sure you really wanna do it in place, especially if you have further analysis that depends on the original data.",
    "421544": "Thanks @cpmpml for the trick.\n\nJust in case you want to go a little faster, no need to create the `mjd_detected` feature\n\n    grp = df.loc[df['detected'] == 1, ['object_id', 'mjd']].groupby('object_id')['mjd']\n    df['test']  = grp.transform('max') - grp.transform('min')\n\n\n92ms on kaggle kernel instead of 120ms.\n\nFor people using linux [modin](https://github.com/modin-project/modin) may help you a bit, I believe it provides parallel apply.",
    "422137": "Code optimization reminds me my youth. \nRegarding calculating mjd_diff, may I suggest you an approach using the properties of the data? As far as I know, simulated measurements are sorted by objectid and mjd, so you do not have to find max and min values of mjd, but just take last and first one:\n\n    dt[detected==1,mjd_diff:=mjd[length(mjd)]-mjd[1],by=object_id]",
    "421430": "Thanks for the reminder CPMP.\n\nI'm addicted to pivot_table rather than group_by and my code is about as fast (faster?) as your optimized code. And more intuitive to me.  It's probably not optimally written.\n\nThis is what I've been using:\n\n     result1 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=[min,max])\n    \n     result1.columns = ['min_mjd','max_mjd']\n    \n     result1 ['detected_duration'] = result1['max_mjd'] - result1['min_mjd']\n\nInterestingly, both of the following are slow:\n\n    result2 = train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: x.ptp())\n\n    df['mjd_diff'] = df.groupby('object_id').mjd_detected.apply(np.max)\n\nAnd then this code:\n\n     train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=np.max)\n\nis a lot faster than this one:\n\n train[train['detected']==1].pivot_table('mjd','object_id',aggfunc=lambda x: np.max(x))\n\n\nNot sure why! I guess the conclusion is to find your bottlenecks and to test alternatives when needed.\n\nAlso, is there a need to compute mjd_diff for each train row in your code? I think we would only need the value for each object_id.\n\n\n\n",
    "420427": "I guess another example of writing efficient code is the loss function. By setting the weights properly, we do not need to write new custom objective and evaluation functions during training. Please correct me if I'm wrong. Thanks!",
    "430428": "Thanks a lot for these tipps. I have been a notorious `groupby()`-`apply()`-offender.",
    "424574": "Thanks for the tips. I'm also curious about the total running time for a single model. What is your total running time for a single model in your own computer?",
    "422013": "This is a great thread; thanks for kicking it off!",
    "421950": "One other option is too use numpy along with numba. That opens door to usage of more numpy functions , e.g. np.ptp\n\nFor example, this code \n\n```\n\n    @numba.jit\n\n    def get_splits(a):\n\n        m = np.concatenate([[True], a[1:] != a[:-1], [True]])\n\n        m = np.flatnonzero(m)\n\n        return m\n\n     @numba.jit(parallel=True)\n\n     def grp_ptp(a, b):\n\n         m = get_splits(a)\n\n         n = len(m)-1\n\n         c = np.empty((n, ), dtype=np.float32)\n\n        for i in prange(n):\n\n            x = b[m[i]:m[i+1]]\n\n            c[i] = np.ptp(x)\n\n        return c \n\n```\n\nruns in 95.1 s ± 2.1 ms on kaggle kernel compared to 1.53 s ± 31.8 ms for df.groupby(\"object_id\")[\"mjd\"].apply(np.ptp)\n\n\nsimilar implementation of max-min using numpy + numba runs in same time as df.groupby(cola)[colb].max() - df.groupby(cola)[colb].min()\n\nOn the flip side, one has to implement more code (one for each feature) and compile using numba :P",
    "420628": "This article has some other tricks for speeding up calculations in pandas.\nhttps://realpython.com/fast-flexible-pandas/\n\n",
    "420334": "Thanks for sharing! This is a crucial yet often overlooked point\n\nKaggle novice here,  but can attest first-hand at how writing fast feat engineering code has dramatically improved my score by allowing me to \"fail fast, fail often\" and try out new feats\n\nAnother possible approach is to precompute and store curve data once in a cesium-like format, thus avoiding groupbys altogether during feat engineering- computing ~120 feats for training set takes me around 15secs with python multiprocessing using this method",
    "465286": "there is also _.apply(raw=True) pass as an ndarray to function which also might speed up in a few seconds",
    "441402": "Not sure whether its apt discussion to ask a  query, but here it's\n\n    X[['col1', cc]].groupby(['col']).agg(lambda x: ' '.join(x)).reset_index()\n    #(X 's shape (9853253, 16))\n\nHow to make this faster?\n\nThanks :)",
    "420384": "Hi @CPMP,\n\nThanks for your example and explanation.\n\nI was curious about the claim we can't use ```ptp``` in ```transform``` so I tried your example.\n\nYou can use it as:\n```train['mjd_diff']  = gr_mjd.transform(np.ptp)```\n\nbut that's much slower than your ```gr_mjd.transform('max') - gr_mjd.transform('min')``` method (30x on my machine), which shows the importance of examining any performance bottlenecks, even when you think you're using the right method or technique.",
    "422269": ""
  }
}