{
  "id": 388788,
  "title": " [LB 0.693 in 16min] How long does your most efficient submission take?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388788",
  "author_name": "Carno Zhao",
  "post_date": "2023-02-19T14:16:12.446000",
  "votes": 42,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Here is my current fastest submission:</p>\n<p><strong>Time: 16min = 2 + 12 + 2</strong><br>\n<strong>LB: 0.693</strong></p>\n<p>~2min: basic <code>env.iter_test()</code> loops, (from baseline empty 0.226 submission)</p>\n<p>~12min: feature engineering, 200 features per target, no feature cross-level-group reuse</p>\n<p>~2min: 1-fold tree model prediction</p>\n<p>I rewrite my feature engineering using <code>numba</code>, which runs extremely fast on numpy arrays. </p>\n<p>However, <code>numba</code> cannot process array of strings very well. The string operations take almost 90% of the total processing time, and the float operations take about <strong>1.5 millisecond</strong> per dataframe. I will dig further in <code>numba</code> to optimize the feature engineering part. I think it is possible to reach &lt;10min submission with the same LB score.</p>",
  "messages": [
    {
      "id": 2150713,
      "postDate": "2023-02-19T14:16:12.447Z",
      "content": "<p>Here is my current fastest submission:</p>\n<p><strong>Time: 16min = 2 + 12 + 2</strong><br>\n<strong>LB: 0.693</strong></p>\n<p>~2min: basic <code>env.iter_test()</code> loops, (from baseline empty 0.226 submission)</p>\n<p>~12min: feature engineering, 200 features per target, no feature cross-level-group reuse</p>\n<p>~2min: 1-fold tree model prediction</p>\n<p>I rewrite my feature engineering using <code>numba</code>, which runs extremely fast on numpy arrays. </p>\n<p>However, <code>numba</code> cannot process array of strings very well. The string operations take almost 90% of the total processing time, and the float operations take about <strong>1.5 millisecond</strong> per dataframe. I will dig further in <code>numba</code> to optimize the feature engineering part. I think it is possible to reach &lt;10min submission with the same LB score.</p>",
      "rawMarkdown": "Here is my current fastest submission:\n\n**Time: 16min = 2 + 12 + 2**\n**LB: 0.693**\n\n~2min: basic `env.iter_test()` loops, (from baseline empty 0.226 submission)\n\n~12min: feature engineering, 200 features per target, no feature cross-level-group reuse\n\n~2min: 1-fold tree model prediction\n\nI rewrite my feature engineering using `numba`, which runs extremely fast on numpy arrays. \n\nHowever, `numba` cannot process array of strings very well. The string operations take almost 90% of the total processing time, and the float operations take about **1.5 millisecond** per dataframe. I will dig further in `numba` to optimize the feature engineering part. I think it is possible to reach <10min submission with the same LB score.",
      "votes": 41
    },
    {
      "id": 2152665,
      "postDate": "2023-02-20T23:21:05.837Z",
      "content": "<p>I have inference times:</p>\n<ul>\n<li>0.671 in 4min</li>\n<li>0.684 in 20min</li>\n<li>0.693 in 39min. </li>\n</ul>\n<p>I still not using numba.</p>",
      "rawMarkdown": "I have inference times:\n- 0.671 in 4min\n- 0.684 in 20min\n- 0.693 in 39min. \n\nI still not using numba.",
      "votes": 1,
      "replies": [
        {
          "id": 2152681,
          "postDate": "2023-02-20T23:39:59.780Z",
          "content": "<p>Here are my updates:</p>\n<ul>\n<li>0.690 in 6min</li>\n<li>0.693 in 10min</li>\n<li>0.696 in 16min</li>\n</ul>",
          "rawMarkdown": "Here are my updates:\n\n- 0.690 in 6min\n- 0.693 in 10min\n- 0.696 in 16min",
          "votes": 7,
          "replies": [
            {
              "id": 2213002,
              "postDate": "2023-04-07T08:33:19.833Z",
              "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>  thanks for telling us the times. Do you mind sharing your latest inference times, given the data update we had 2 weeks ago?</p>",
              "rawMarkdown": "@titericz @carnozhao  thanks for telling us the times. Do you mind sharing your latest inference times, given the data update we had 2 weeks ago?"
            }
          ]
        }
      ]
    },
    {
      "id": 2161507,
      "postDate": "2023-02-27T14:59:11.750Z",
      "content": "<p>0.683 in 12 min</p>",
      "rawMarkdown": "0.683 in 12 min"
    },
    {
      "id": 2153794,
      "postDate": "2023-02-21T16:41:26.157Z",
      "content": "<p>Hi Zhao, I'm wondering how you rewrite your feature engineering code with numba. Could you give some baseline examples?</p>\n<p>For example, how to deal with pandas.groupby.agg('sum')? </p>",
      "rawMarkdown": "Hi Zhao, I'm wondering how you rewrite your feature engineering code with numba. Could you give some baseline examples?\n\nFor example, how to deal with pandas.groupby.agg('sum')? ",
      "replies": [
        {
          "id": 2154037,
          "postDate": "2023-02-21T19:27:45.080Z",
          "content": "<p>Zhao, do you have a better way than split/apply/concatenate?</p>\n<p>Moonlit, best generic way I know is like:</p>\n<pre><code>\n\ngroups = np.split(array, np.unique(array[:, ], return_index=)[][:])\nfinal_result = np.concatenate([feature_engg(grp)  grp  groups])\n</code></pre>\n<p>(See also <a href=\"https://stackoverflow.com/questions/38013778/is-there-any-numpy-group-by-function\" target=\"_blank\">is-there-any-numpy-group-by-function</a>)</p>\n<p>And assuming that the batch size from the env.iter_test() API is small, even tricks to avoid the time cost of the list comprehension aka for loop may(?) not help much. So probably(?) the above is pretty effective, but I haven't benchmarked it. (Well, except in a very different context <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">over here</a>.)</p>",
          "rawMarkdown": "Zhao, do you have a better way than split/apply/concatenate?\n\nMoonlit, best generic way I know is like:\n```python\n## Assumes column 0 is sorted and is the column to group on\n## To sort, use: a = a[a[:, 0].argsort()]\ngroups = np.split(array, np.unique(array[:, 0], return_index=True)[1][1:])\nfinal_result = np.concatenate([feature_engg(grp) for grp in groups])\n```\n(See also [is-there-any-numpy-group-by-function](https://stackoverflow.com/questions/38013778/is-there-any-numpy-group-by-function))\n\nAnd assuming that the batch size from the env.iter_test() API is small, even tricks to avoid the time cost of the list comprehension aka for loop may(?) not help much. So probably(?) the above is pretty effective, but I haven't benchmarked it. (Well, except in a very different context [over here](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars).)",
          "votes": 1,
          "replies": [
            {
              "id": 2154275,
              "postDate": "2023-02-21T23:18:48.047Z",
              "content": "<p>Ha, I haven't thought about using sorting to do groupbys. my groupby implementation is merely use a boolean mask array to save the condition result during iteration over the groupby columns. And then use the mask to slice the target column and put the statistical function after it.</p>",
              "rawMarkdown": "Ha, I haven't thought about using sorting to do groupbys. my groupby implementation is merely use a boolean mask array to save the condition result during iteration over the groupby columns. And then use the mask to slice the target column and put the statistical function after it.",
              "votes": 2
            },
            {
              "id": 2193794,
              "postDate": "2023-03-23T13:40:15.900Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2154467,
          "postDate": "2023-02-22T03:18:05.987Z",
          "content": "<p>I get it. It's quite helpful. Thank you all so much.</p>",
          "rawMarkdown": "I get it. It's quite helpful. Thank you all so much."
        }
      ]
    },
    {
      "id": 2152728,
      "postDate": "2023-02-21T01:23:00.570Z",
      "content": "<p>Are you doing 'real' string operations as part of feature engineering, things like filtering on df[col].str.contains(substr)? If not, can you get a big speed up by using a label encoder or similar - basically, by converting to numeric type asap?</p>",
      "rawMarkdown": "Are you doing 'real' string operations as part of feature engineering, things like filtering on df[col].str.contains(substr)? If not, can you get a big speed up by using a label encoder or similar - basically, by converting to numeric type asap?",
      "replies": [
        {
          "id": 2152742,
          "postDate": "2023-02-21T01:38:31.417Z",
          "content": "<p>That's what I am doing in my implementation. However, the time complexity is O(n), where n is the number of string rows. But in polars, the time complexity over string columns is O(1). I have no iead why polars can do string operations so fast and my current result is the balance of numeric speed boost and string speed lag.</p>",
          "rawMarkdown": "That's what I am doing in my implementation. However, the time complexity is O(n), where n is the number of string rows. But in polars, the time complexity over string columns is O(1). I have no iead why polars can do string operations so fast and my current result is the balance of numeric speed boost and string speed lag.",
          "replies": [
            {
              "id": 2152824,
              "postDate": "2023-02-21T03:36:45.200Z",
              "content": "<p>Interesting, good to know. </p>",
              "rawMarkdown": "Interesting, good to know. "
            },
            {
              "id": 2154056,
              "postDate": "2023-02-21T19:34:58.137Z",
              "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> if you wrote a label encoder from scratch there might be a way to process all columns at once, rather than sequentially? Avoiding one or both of multiple passes through the entire data, and/or multiple 'resizing data' costs?</p>\n<p>Or, even simpler, numpy defaults to row-major ordering, but google thinks it can be configured either which way. If there's not a big upfront hit for transforming it, maybe processing each column sequentially will be ~O(1) if using column-major ordering (or just array.T so that columns are 'rows')?</p>\n<p>If you try any of that, let me know how it goes!</p>",
              "rawMarkdown": "@carnozhao if you wrote a label encoder from scratch there might be a way to process all columns at once, rather than sequentially? Avoiding one or both of multiple passes through the entire data, and/or multiple 'resizing data' costs?\n\nOr, even simpler, numpy defaults to row-major ordering, but google thinks it can be configured either which way. If there's not a big upfront hit for transforming it, maybe processing each column sequentially will be ~O(1) if using column-major ordering (or just array.T so that columns are 'rows')?\n\nIf you try any of that, let me know how it goes!"
            },
            {
              "id": 2154265,
              "postDate": "2023-02-21T23:13:52.417Z",
              "content": "<p>Yes, what you mentioned might be a good way to process strings. However AFAIK, you should treat numba codes as pythonic C codes, that is every for loop will be compiled and run as fast as original numpy function. You probably think Python's numpy process columns at once, but actually it is just faster loops.</p>\n<p>To be more specific, numba's string and numpy's string is different when the strings are stored in arrays. Numpy use '&lt;Uxx' dtype to store strings making sure every string in this array is shorter than xx. However, numba use real python strings, the convertion from numpy to numba is slow.</p>",
              "rawMarkdown": "Yes, what you mentioned might be a good way to process strings. However AFAIK, you should treat numba codes as pythonic C codes, that is every for loop will be compiled and run as fast as original numpy function. You probably think Python's numpy process columns at once, but actually it is just faster loops.\n\nTo be more specific, numba's string and numpy's string is different when the strings are stored in arrays. Numpy use '<Uxx' dtype to store strings making sure every string in this array is shorter than xx. However, numba use real python strings, the convertion from numpy to numba is slow."
            },
            {
              "id": 2160969,
              "postDate": "2023-02-27T06:07:45.993Z",
              "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> <br>\nPolars is faster than NumPy in string operations because the string in numpy is python object. Refer to <a href=\"https://pola-rs.github.io/polars-book/user-guide/howcani/data/strings.html\" target=\"_blank\">this</a>. I'm a little bit curious since ploars is faster, what's the befinits of using numpy + numba?</p>",
              "rawMarkdown": "Thanks for sharing @carnozhao \nPolars is faster than NumPy in string operations because the string in numpy is python object. Refer to [this](https://pola-rs.github.io/polars-book/user-guide/howcani/data/strings.html). I'm a little bit curious since ploars is faster, what's the befinits of using numpy + numba?"
            },
            {
              "id": 2160977,
              "postDate": "2023-02-27T06:15:19.293Z",
              "content": "<p>Because the given input is pd.DataFrame format, I can process its raw data by its \".values\" attribute using numba. However, for polars, I need convert it to polars DataFrame and then convert it back to pandas, which is a little bit redundent. And moreover, numba can do much faster numeric operation than polars, considering mean/std/max/min with groupby/filter/etc.</p>",
              "rawMarkdown": "Because the given input is pd.DataFrame format, I can process its raw data by its \".values\" attribute using numba. However, for polars, I need convert it to polars DataFrame and then convert it back to pandas, which is a little bit redundent. And moreover, numba can do much faster numeric operation than polars, considering mean/std/max/min with groupby/filter/etc.",
              "votes": 1
            },
            {
              "id": 2160988,
              "postDate": "2023-02-27T06:26:10.150Z",
              "content": "<p>Got it, thanks for the detail explanation</p>",
              "rawMarkdown": "Got it, thanks for the detail explanation"
            }
          ]
        }
      ]
    },
    {
      "id": 2151588,
      "postDate": "2023-02-20T07:22:00.153Z",
      "content": "<p>Impressive!<br>\nI've never used <code>numba</code>! How does it scale with bigger dataset like the train set?</p>",
      "rawMarkdown": "Impressive!\nI've never used `numba`! How does it scale with bigger dataset like the train set?",
      "replies": [
        {
          "id": 2151602,
          "postDate": "2023-02-20T07:36:15.937Z",
          "content": "<p>It is not recommended to use numba on large numeric &amp; string datasets, polars or cudf is a better choice. But if you only need to do some large numeric computation, numba is what you want.</p>",
          "rawMarkdown": "It is not recommended to use numba on large numeric & string datasets, polars or cudf is a better choice. But if you only need to do some large numeric computation, numba is what you want.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2150726,
      "postDate": "2023-02-19T14:27:21.493Z",
      "content": "<p>Thanks for sharing, numba master. I benefits a lot from your OTTO numba notebook.</p>",
      "rawMarkdown": "Thanks for sharing, numba master. I benefits a lot from your OTTO numba notebook.",
      "replies": [
        {
          "id": 2150732,
          "postDate": "2023-02-19T14:32:47.290Z",
          "content": "<p>I'm glad it can help you😃</p>",
          "rawMarkdown": "I'm glad it can help you😃"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2152665,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2023-02-20T23:21:05.837000",
      "content": "<p>I have inference times:</p>\n<ul>\n<li>0.671 in 4min</li>\n<li>0.684 in 20min</li>\n<li>0.693 in 39min. </li>\n</ul>\n<p>I still not using numba.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2152681,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-02-20T23:39:59.780000",
          "content": "<p>Here are my updates:</p>\n<ul>\n<li>0.690 in 6min</li>\n<li>0.693 in 10min</li>\n<li>0.696 in 16min</li>\n</ul>",
          "votes": 7,
          "replies": [
            {
              "id": 2213002,
              "author_name": "durvorezbariq",
              "author_url": "",
              "post_date": "2023-04-07T08:33:19.833000",
              "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>  thanks for telling us the times. Do you mind sharing your latest inference times, given the data update we had 2 weeks ago?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2161507,
      "author_name": "ln",
      "author_url": "",
      "post_date": "2023-02-27T14:59:11.750000",
      "content": "<p>0.683 in 12 min</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2153794,
      "author_name": "Moonlit",
      "author_url": "",
      "post_date": "2023-02-21T16:41:26.157000",
      "content": "<p>Hi Zhao, I'm wondering how you rewrite your feature engineering code with numba. Could you give some baseline examples?</p>\n<p>For example, how to deal with pandas.groupby.agg('sum')? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2154037,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2023-02-21T19:27:45.080000",
          "content": "<p>Zhao, do you have a better way than split/apply/concatenate?</p>\n<p>Moonlit, best generic way I know is like:</p>\n<pre><code>\n\ngroups = np.split(array, np.unique(array[:, ], return_index=)[][:])\nfinal_result = np.concatenate([feature_engg(grp)  grp  groups])\n</code></pre>\n<p>(See also <a href=\"https://stackoverflow.com/questions/38013778/is-there-any-numpy-group-by-function\" target=\"_blank\">is-there-any-numpy-group-by-function</a>)</p>\n<p>And assuming that the batch size from the env.iter_test() API is small, even tricks to avoid the time cost of the list comprehension aka for loop may(?) not help much. So probably(?) the above is pretty effective, but I haven't benchmarked it. (Well, except in a very different context <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">over here</a>.)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2154275,
              "author_name": "Carno Zhao",
              "author_url": "",
              "post_date": "2023-02-21T23:18:48.047000",
              "content": "<p>Ha, I haven't thought about using sorting to do groupbys. my groupby implementation is merely use a boolean mask array to save the condition result during iteration over the groupby columns. And then use the mask to slice the target column and put the statistical function after it.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2193794,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-23T13:40:15.900000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2154467,
          "author_name": "Moonlit",
          "author_url": "",
          "post_date": "2023-02-22T03:18:05.987000",
          "content": "<p>I get it. It's quite helpful. Thank you all so much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2152728,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2023-02-21T01:23:00.570000",
      "content": "<p>Are you doing 'real' string operations as part of feature engineering, things like filtering on df[col].str.contains(substr)? If not, can you get a big speed up by using a label encoder or similar - basically, by converting to numeric type asap?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2152742,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-02-21T01:38:31.417000",
          "content": "<p>That's what I am doing in my implementation. However, the time complexity is O(n), where n is the number of string rows. But in polars, the time complexity over string columns is O(1). I have no iead why polars can do string operations so fast and my current result is the balance of numeric speed boost and string speed lag.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2152824,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-21T03:36:45.200000",
              "content": "<p>Interesting, good to know. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2154056,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-21T19:34:58.137000",
              "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> if you wrote a label encoder from scratch there might be a way to process all columns at once, rather than sequentially? Avoiding one or both of multiple passes through the entire data, and/or multiple 'resizing data' costs?</p>\n<p>Or, even simpler, numpy defaults to row-major ordering, but google thinks it can be configured either which way. If there's not a big upfront hit for transforming it, maybe processing each column sequentially will be ~O(1) if using column-major ordering (or just array.T so that columns are 'rows')?</p>\n<p>If you try any of that, let me know how it goes!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2154265,
              "author_name": "Carno Zhao",
              "author_url": "",
              "post_date": "2023-02-21T23:13:52.417000",
              "content": "<p>Yes, what you mentioned might be a good way to process strings. However AFAIK, you should treat numba codes as pythonic C codes, that is every for loop will be compiled and run as fast as original numpy function. You probably think Python's numpy process columns at once, but actually it is just faster loops.</p>\n<p>To be more specific, numba's string and numpy's string is different when the strings are stored in arrays. Numpy use '&lt;Uxx' dtype to store strings making sure every string in this array is shorter than xx. However, numba use real python strings, the convertion from numpy to numba is slow.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2160969,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-02-27T06:07:45.993000",
              "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> <br>\nPolars is faster than NumPy in string operations because the string in numpy is python object. Refer to <a href=\"https://pola-rs.github.io/polars-book/user-guide/howcani/data/strings.html\" target=\"_blank\">this</a>. I'm a little bit curious since ploars is faster, what's the befinits of using numpy + numba?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2160977,
              "author_name": "Carno Zhao",
              "author_url": "",
              "post_date": "2023-02-27T06:15:19.293000",
              "content": "<p>Because the given input is pd.DataFrame format, I can process its raw data by its \".values\" attribute using numba. However, for polars, I need convert it to polars DataFrame and then convert it back to pandas, which is a little bit redundent. And moreover, numba can do much faster numeric operation than polars, considering mean/std/max/min with groupby/filter/etc.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2160988,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-02-27T06:26:10.150000",
              "content": "<p>Got it, thanks for the detail explanation</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2151588,
      "author_name": "Hoang Nguyen",
      "author_url": "",
      "post_date": "2023-02-20T07:22:00.153000",
      "content": "<p>Impressive!<br>\nI've never used <code>numba</code>! How does it scale with bigger dataset like the train set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2151602,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-02-20T07:36:15.937000",
          "content": "<p>It is not recommended to use numba on large numeric &amp; string datasets, polars or cudf is a better choice. But if you only need to do some large numeric computation, numba is what you want.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2150726,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2023-02-19T14:27:21.493000",
      "content": "<p>Thanks for sharing, numba master. I benefits a lot from your OTTO numba notebook.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2150732,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-02-19T14:32:47.290000",
          "content": "<p>I'm glad it can help you😃</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2150713": "Here is my current fastest submission:\n\n**Time: 16min = 2 + 12 + 2**\n**LB: 0.693**\n\n~2min: basic `env.iter_test()` loops, (from baseline empty 0.226 submission)\n\n~12min: feature engineering, 200 features per target, no feature cross-level-group reuse\n\n~2min: 1-fold tree model prediction\n\nI rewrite my feature engineering using `numba`, which runs extremely fast on numpy arrays. \n\nHowever, `numba` cannot process array of strings very well. The string operations take almost 90% of the total processing time, and the float operations take about **1.5 millisecond** per dataframe. I will dig further in `numba` to optimize the feature engineering part. I think it is possible to reach <10min submission with the same LB score.",
    "2152665": "I have inference times:\n- 0.671 in 4min\n- 0.684 in 20min\n- 0.693 in 39min. \n\nI still not using numba.",
    "2161507": "0.683 in 12 min",
    "2153794": "Hi Zhao, I'm wondering how you rewrite your feature engineering code with numba. Could you give some baseline examples?\n\nFor example, how to deal with pandas.groupby.agg('sum')? ",
    "2152728": "Are you doing 'real' string operations as part of feature engineering, things like filtering on df[col].str.contains(substr)? If not, can you get a big speed up by using a label encoder or similar - basically, by converting to numeric type asap?",
    "2151588": "Impressive!\nI've never used `numba`! How does it scale with bigger dataset like the train set?",
    "2150726": "Thanks for sharing, numba master. I benefits a lot from your OTTO numba notebook."
  }
}