{
  "id": 91077,
  "title": "Simple performance improvement",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91077",
  "author_name": "",
  "post_date": "2019-04-30T18:04:23.020469800Z",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<p>My feature generation procedure took more time than I could expect. \nThe reason was quite simple, in my case 'acoustic_data' was stored as np.int16 and every time a function was applied the data is converted into np.float before calling the function (this is my assumption). \nSolution is to convert data chunk to np.float32 once and then apply the functions.</p>\n\n<p>So code before was something like this:</p>\n\n<p>```\nfeature_desc = {\n    'mean': np.mean, \n    'std': np.std,\n    ...\n}</p>\n\n<p>signal = df['acoustic_data'].values\nchunk = signal[:150_000]\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)\n```</p>\n\n<p>new code is the following:</p>\n\n<p>```\nsignal = df['acoustic_data'].values\nchunk = signal[:150_000].astype(np.float32)\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)</p>\n\n<p>```\nThis simple trick saved me a lot of time. For this sample code, optimized version works ~1.5 times faster. \nUpdate: actually 3 times faster for simple functions like np.mean and chunk size 150000.</p>",
  "messages": [
    {
      "id": "525325",
      "postDate": "04/30/2019 18:04:23",
      "content": "<p>My feature generation procedure took more time than I could expect. \nThe reason was quite simple, in my case 'acoustic_data' was stored as np.int16 and every time a function was applied the data is converted into np.float before calling the function (this is my assumption). \nSolution is to convert data chunk to np.float32 once and then apply the functions.</p>\n\n<p>So code before was something like this:</p>\n\n<p>```\nfeature_desc = {\n    'mean': np.mean, \n    'std': np.std,\n    ...\n}</p>\n\n<p>signal = df['acoustic_data'].values\nchunk = signal[:150_000]\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)\n```</p>\n\n<p>new code is the following:</p>\n\n<p>```\nsignal = df['acoustic_data'].values\nchunk = signal[:150_000].astype(np.float32)\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)</p>\n\n<p>```\nThis simple trick saved me a lot of time. For this sample code, optimized version works ~1.5 times faster. \nUpdate: actually 3 times faster for simple functions like np.mean and chunk size 150000.</p>",
      "rawMarkdown": "My feature generation procedure took more time than I could expect. \nThe reason was quite simple, in my case 'acoustic_data' was stored as np.int16 and every time a function was applied the data is converted into np.float before calling the function (this is my assumption). \nSolution is to convert data chunk to np.float32 once and then apply the functions.\n\nSo code before was something like this:\n\n```\nfeature_desc = {\n    'mean': np.mean, \n    'std': np.std,\n    ...\n}\n\nsignal = df['acoustic_data'].values\nchunk = signal[:150_000]\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)\n```\n\n\nnew code is the following:\n\n```\nsignal = df['acoustic_data'].values\nchunk = signal[:150_000].astype(np.float32)\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)\n\n```\nThis simple trick saved me a lot of time. For this sample code, optimized version works ~1.5 times faster. \nUpdate: actually 3 times faster for simple functions like np.mean and chunk size 150000.",
      "votes": null
    },
    {
      "id": "525388",
      "postDate": "04/30/2019 21:54:50",
      "content": "<p>Thanks! This is cool. I didn't know that. I read through the numpy documentation and it seems that part of the speed you are gaining is because you convert them to <code>float32</code> while numpy will convert inputs to <code>float64</code> if passed as integers.\n<a href=\"https://docs.scipy.org/doc/numpy/reference/generated/numpy.mean.html\">https://docs.scipy.org/doc/numpy/reference/generated/numpy.mean.html</a></p>",
      "rawMarkdown": "Thanks! This is cool. I didn't know that. I read through the numpy documentation and it seems that part of the speed you are gaining is because you convert them to `float32` while numpy will convert inputs to `float64` if passed as integers.\nhttps://docs.scipy.org/doc/numpy/reference/generated/numpy.mean.html",
      "votes": null
    },
    {
      "id": "525390",
      "postDate": "04/30/2019 21:55:31",
      "content": "<p>Also please \"Note that for floating-point input, the mean is computed using the same precision the input has. Depending on the input data, this can cause the results to be inaccurate, especially for float32\"</p>",
      "rawMarkdown": "Also please \"Note that for floating-point input, the mean is computed using the same precision the input has. Depending on the input data, this can cause the results to be inaccurate, especially for float32\"",
      "votes": null
    },
    {
      "id": "525549",
      "postDate": "05/01/2019 09:04:26",
      "content": "<p><a href=\"/mhviraf\">@mhviraf</a>, thank for information. \nConverting to np.float64 is 2 times slower than np.float32 but still faster than not doing casting at all.</p>",
      "rawMarkdown": "mhviraf, thank for information. \nConverting to np.float64 is 2 times slower than np.float32 but still faster than not doing casting at all.",
      "votes": null
    },
    {
      "id": "525657",
      "postDate": "05/01/2019 13:25:43",
      "content": "<p>Another useful trick is to use <code>bottleneck</code> instead of NumPy: <a href=\"https://github.com/kwgoodman/bottleneck\">https://github.com/kwgoodman/bottleneck</a></p>\n\n<p>It's a library of common NumPy functions that are implemeted in C for extra speed and efficiency. As long as your dtype is 'int32', 'int64', 'float32' or 'float64' you'll see improvements, and it's especially good for rolling mean/variance, important in this competition. You can save a lot of time on your feature engineering.</p>",
      "rawMarkdown": "Another useful trick is to use `bottleneck` instead of NumPy: https://github.com/kwgoodman/bottleneck\n\nIt's a library of common NumPy functions that are implemeted in C for extra speed and efficiency. As long as your dtype is 'int32', 'int64', 'float32' or 'float64' you'll see improvements, and it's especially good for rolling mean/variance, important in this competition. You can save a lot of time on your feature engineering.",
      "votes": null
    },
    {
      "id": "525717",
      "postDate": "05/01/2019 15:38:58",
      "content": "<p>great</p>",
      "rawMarkdown": "great",
      "votes": null
    },
    {
      "id": "526181",
      "postDate": "05/02/2019 14:12:03",
      "content": "<p>Thanks, didn't know that one. Can you explain why it's so much faster than numpy? I just looked at their benchmark table <a href=\"https://kwgoodman.github.io/bottleneck-doc/intro.html\"></a> and the improvement is huge for some functions. I thought most of numpy was written in C (CPython), so i'm a little surprised.</p>",
      "rawMarkdown": "Thanks, didn't know that one. Can you explain why it's so much faster than numpy? I just looked at their benchmark table [](https://kwgoodman.github.io/bottleneck-doc/intro.html) and the improvement is huge for some functions. I thought most of numpy was written in C (CPython), so i'm a little surprised.",
      "votes": null
    },
    {
      "id": "526187",
      "postDate": "05/02/2019 14:26:26",
      "content": "<p>I'm afraid I haven't had time to delve into the source code of either so I can' t make direct comparisons. C can be fast as lightning if you're willing to play fast and loose with memory so I assume it has something to do with that. The <code>bottleneck</code> code does contain quite a few fixes for memory leakage which makes me think my suspicion is correct! </p>",
      "rawMarkdown": "I'm afraid I haven't had time to delve into the source code of either so I can' t make direct comparisons. C can be fast as lightning if you're willing to play fast and loose with memory so I assume it has something to do with that. The `bottleneck` code does contain quite a few fixes for memory leakage which makes me think my suspicion is correct!",
      "votes": null
    },
    {
      "id": "526510",
      "postDate": "05/03/2019 07:32:48",
      "content": "<p>Thanks for the informative write up.</p>",
      "rawMarkdown": "Thanks for the informative write up.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 525388,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "04/30/2019 21:54:50",
      "content": "<p>Thanks! This is cool. I didn't know that. I read through the numpy documentation and it seems that part of the speed you are gaining is because you convert them to <code>float32</code> while numpy will convert inputs to <code>float64</code> if passed as integers.\n<a href=\"https://docs.scipy.org/doc/numpy/reference/generated/numpy.mean.html\">https://docs.scipy.org/doc/numpy/reference/generated/numpy.mean.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 525390,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "04/30/2019 21:55:31",
      "content": "<p>Also please \"Note that for floating-point input, the mean is computed using the same precision the input has. Depending on the input data, this can cause the results to be inaccurate, especially for float32\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 525549,
      "author_name": "alexfir",
      "author_url": "",
      "post_date": "05/01/2019 09:04:26",
      "content": "<p><a href=\"/mhviraf\">@mhviraf</a>, thank for information. \nConverting to np.float64 is 2 times slower than np.float32 but still faster than not doing casting at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 525657,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "05/01/2019 13:25:43",
      "content": "<p>Another useful trick is to use <code>bottleneck</code> instead of NumPy: <a href=\"https://github.com/kwgoodman/bottleneck\">https://github.com/kwgoodman/bottleneck</a></p>\n\n<p>It's a library of common NumPy functions that are implemeted in C for extra speed and efficiency. As long as your dtype is 'int32', 'int64', 'float32' or 'float64' you'll see improvements, and it's especially good for rolling mean/variance, important in this competition. You can save a lot of time on your feature engineering.</p>",
      "votes": null,
      "replies": [
        {
          "id": 526181,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "05/02/2019 14:12:03",
          "content": "<p>Thanks, didn't know that one. Can you explain why it's so much faster than numpy? I just looked at their benchmark table <a href=\"https://kwgoodman.github.io/bottleneck-doc/intro.html\"></a> and the improvement is huge for some functions. I thought most of numpy was written in C (CPython), so i'm a little surprised.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526187,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/02/2019 14:26:26",
          "content": "<p>I'm afraid I haven't had time to delve into the source code of either so I can' t make direct comparisons. C can be fast as lightning if you're willing to play fast and loose with memory so I assume it has something to do with that. The <code>bottleneck</code> code does contain quite a few fixes for memory leakage which makes me think my suspicion is correct! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525717,
      "author_name": "daftintrovert",
      "author_url": "",
      "post_date": "05/01/2019 15:38:58",
      "content": "<p>great</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526510,
      "author_name": "farukcse",
      "author_url": "",
      "post_date": "05/03/2019 07:32:48",
      "content": "<p>Thanks for the informative write up.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "525325": "My feature generation procedure took more time than I could expect. \nThe reason was quite simple, in my case 'acoustic_data' was stored as np.int16 and every time a function was applied the data is converted into np.float before calling the function (this is my assumption). \nSolution is to convert data chunk to np.float32 once and then apply the functions.\n\nSo code before was something like this:\n\n```\nfeature_desc = {\n    'mean': np.mean, \n    'std': np.std,\n    ...\n}\n\nsignal = df['acoustic_data'].values\nchunk = signal[:150_000]\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)\n```\n\n\nnew code is the following:\n\n```\nsignal = df['acoustic_data'].values\nchunk = signal[:150_000].astype(np.float32)\nfor func_name, func in feature_desc.items():\n    features[func_name] = func(chunk)\n\n```\nThis simple trick saved me a lot of time. For this sample code, optimized version works ~1.5 times faster. \nUpdate: actually 3 times faster for simple functions like np.mean and chunk size 150000.",
    "525388": "Thanks! This is cool. I didn't know that. I read through the numpy documentation and it seems that part of the speed you are gaining is because you convert them to `float32` while numpy will convert inputs to `float64` if passed as integers.\nhttps://docs.scipy.org/doc/numpy/reference/generated/numpy.mean.html",
    "525390": "Also please \"Note that for floating-point input, the mean is computed using the same precision the input has. Depending on the input data, this can cause the results to be inaccurate, especially for float32\"",
    "525549": "mhviraf, thank for information. \nConverting to np.float64 is 2 times slower than np.float32 but still faster than not doing casting at all.",
    "525657": "Another useful trick is to use `bottleneck` instead of NumPy: https://github.com/kwgoodman/bottleneck\n\nIt's a library of common NumPy functions that are implemeted in C for extra speed and efficiency. As long as your dtype is 'int32', 'int64', 'float32' or 'float64' you'll see improvements, and it's especially good for rolling mean/variance, important in this competition. You can save a lot of time on your feature engineering.",
    "525717": "great",
    "526181": "Thanks, didn't know that one. Can you explain why it's so much faster than numpy? I just looked at their benchmark table [](https://kwgoodman.github.io/bottleneck-doc/intro.html) and the improvement is huge for some functions. I thought most of numpy was written in C (CPython), so i'm a little surprised.",
    "526187": "I'm afraid I haven't had time to delve into the source code of either so I can' t make direct comparisons. C can be fast as lightning if you're willing to play fast and loose with memory so I assume it has something to do with that. The `bottleneck` code does contain quite a few fixes for memory leakage which makes me think my suspicion is correct!",
    "526510": "Thanks for the informative write up."
  },
  "source": "meta"
}