{
  "id": 63210,
  "title": "Language of choice",
  "url": "/competitions/trackml-particle-identification/discussion/63210",
  "author_name": "",
  "post_date": "2018-08-13T16:08:14.148331400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,\nHas anyone used a static typed language (or cython) to speed up the calculations? I've been using Python/R and I think it slowed me down after the +- 0.5.\nEven if suggesting creating a DBScan algo in a different programming language takes time the post processing of the tracks would be done a lot quicker.</p>\n\n<p>What do you think?</p>",
  "messages": [
    {
      "id": "369688",
      "postDate": "08/13/2018 16:08:14",
      "content": "<p>Hi,\nHas anyone used a static typed language (or cython) to speed up the calculations? I've been using Python/R and I think it slowed me down after the +- 0.5.\nEven if suggesting creating a DBScan algo in a different programming language takes time the post processing of the tracks would be done a lot quicker.</p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "Hi,\nHas anyone used a static typed language (or cython) to speed up the calculations? I've been using Python/R and I think it slowed me down after the +- 0.5.\nEven if suggesting creating a DBScan algo in a different programming language takes time the post processing of the tracks would be done a lot quicker.\n\nWhat do you think?",
      "votes": null
    },
    {
      "id": "371287",
      "postDate": "08/16/2018 13:09:50",
      "content": "<p>Hi,\nThe 1st place solution has been written in C++, but I'm sure that was not the decisive factor. I used Python only and my solution is one of the fastest ones near the top.</p>\n\n<p>The key is vectorization of the calculations. My code has only a few \"for\" loops in about 3800 lines of code. All the rest is numpy vector arithmetic. In the end this means that the code spends practically all of its time in optimized, cache-friendly inner loops inside numpy, pandas, and scikit libraries.</p>\n\n<p>I was somewhat performance-conscious throughout the development for the simple reason that I did everything on my notebook and runtime was one of the limiting factors for how quickly I could test improvements. I also used hyperopt a lot, and there it makes a big difference how quickly hyperopt can try another set of hyper-parameters.</p>\n\n<p>P.S.: I also used Python's profiling features once, and that's a good way to find out which operations are inefficient. For example, I learned to avoid pandas \"groupby\" operation in cases where the number of groups would be large. Replacing such cases by custom code using numpy directly bought me a lot of time.</p>",
      "rawMarkdown": "Hi,\nThe 1st place solution has been written in C++, but I'm sure that was not the decisive factor. I used Python only and my solution is one of the fastest ones near the top.\n\nThe key is vectorization of the calculations. My code has only a few \"for\" loops in about 3800 lines of code. All the rest is numpy vector arithmetic. In the end this means that the code spends practically all of its time in optimized, cache-friendly inner loops inside numpy, pandas, and scikit libraries.\n\nI was somewhat performance-conscious throughout the development for the simple reason that I did everything on my notebook and runtime was one of the limiting factors for how quickly I could test improvements. I also used hyperopt a lot, and there it makes a big difference how quickly hyperopt can try another set of hyper-parameters.\n\nP.S.: I also used Python's profiling features once, and that's a good way to find out which operations are inefficient. For example, I learned to avoid pandas \"groupby\" operation in cases where the number of groups would be large. Replacing such cases by custom code using numpy directly bought me a lot of time.",
      "votes": null
    },
    {
      "id": "371292",
      "postDate": "08/16/2018 13:15:12",
      "content": "<p>I second that. And when one cannot use numpy one can use numba to compile Python code very effectively, see for instance:   <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/A_Comparison_Of_C_Julia_Python_Numba_Cython_Scipy_and_BLAS_on_LU_Factorization?lang=en\">A Speed Comparison Of C, Julia, Python, Numba, and Cython on LU Factorization</a> and <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Python_Meets_Julia_Micro_Performance?lang=en\">How To Make Julia Run As Fast As Python</a></p>",
      "rawMarkdown": "I second that. And when one cannot use numpy one can use numba to compile Python code very effectively, see for instance:   [A Speed Comparison Of C, Julia, Python, Numba, and Cython on LU Factorization][1] and [How To Make Julia Run As Fast As Python][2]\n\n\n  [1]: https://www.ibm.com/developerworks/community/blogs/jfp/entry/A_Comparison_Of_C_Julia_Python_Numba_Cython_Scipy_and_BLAS_on_LU_Factorization?lang=en\n  [2]: https://www.ibm.com/developerworks/community/blogs/jfp/entry/Python_Meets_Julia_Micro_Performance?lang=en",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 371287,
      "author_name": "edwinst",
      "author_url": "",
      "post_date": "08/16/2018 13:09:50",
      "content": "<p>Hi,\nThe 1st place solution has been written in C++, but I'm sure that was not the decisive factor. I used Python only and my solution is one of the fastest ones near the top.</p>\n\n<p>The key is vectorization of the calculations. My code has only a few \"for\" loops in about 3800 lines of code. All the rest is numpy vector arithmetic. In the end this means that the code spends practically all of its time in optimized, cache-friendly inner loops inside numpy, pandas, and scikit libraries.</p>\n\n<p>I was somewhat performance-conscious throughout the development for the simple reason that I did everything on my notebook and runtime was one of the limiting factors for how quickly I could test improvements. I also used hyperopt a lot, and there it makes a big difference how quickly hyperopt can try another set of hyper-parameters.</p>\n\n<p>P.S.: I also used Python's profiling features once, and that's a good way to find out which operations are inefficient. For example, I learned to avoid pandas \"groupby\" operation in cases where the number of groups would be large. Replacing such cases by custom code using numpy directly bought me a lot of time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 371292,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/16/2018 13:15:12",
          "content": "<p>I second that. And when one cannot use numpy one can use numba to compile Python code very effectively, see for instance:   <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/A_Comparison_Of_C_Julia_Python_Numba_Cython_Scipy_and_BLAS_on_LU_Factorization?lang=en\">A Speed Comparison Of C, Julia, Python, Numba, and Cython on LU Factorization</a> and <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Python_Meets_Julia_Micro_Performance?lang=en\">How To Make Julia Run As Fast As Python</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "369688": "Hi,\nHas anyone used a static typed language (or cython) to speed up the calculations? I've been using Python/R and I think it slowed me down after the +- 0.5.\nEven if suggesting creating a DBScan algo in a different programming language takes time the post processing of the tracks would be done a lot quicker.\n\nWhat do you think?",
    "371287": "Hi,\nThe 1st place solution has been written in C++, but I'm sure that was not the decisive factor. I used Python only and my solution is one of the fastest ones near the top.\n\nThe key is vectorization of the calculations. My code has only a few \"for\" loops in about 3800 lines of code. All the rest is numpy vector arithmetic. In the end this means that the code spends practically all of its time in optimized, cache-friendly inner loops inside numpy, pandas, and scikit libraries.\n\nI was somewhat performance-conscious throughout the development for the simple reason that I did everything on my notebook and runtime was one of the limiting factors for how quickly I could test improvements. I also used hyperopt a lot, and there it makes a big difference how quickly hyperopt can try another set of hyper-parameters.\n\nP.S.: I also used Python's profiling features once, and that's a good way to find out which operations are inefficient. For example, I learned to avoid pandas \"groupby\" operation in cases where the number of groups would be large. Replacing such cases by custom code using numpy directly bought me a lot of time.",
    "371292": "I second that. And when one cannot use numpy one can use numba to compile Python code very effectively, see for instance:   [A Speed Comparison Of C, Julia, Python, Numba, and Cython on LU Factorization][1] and [How To Make Julia Run As Fast As Python][2]\n\n\n  [1]: https://www.ibm.com/developerworks/community/blogs/jfp/entry/A_Comparison_Of_C_Julia_Python_Numba_Cython_Scipy_and_BLAS_on_LU_Factorization?lang=en\n  [2]: https://www.ibm.com/developerworks/community/blogs/jfp/entry/Python_Meets_Julia_Micro_Performance?lang=en"
  },
  "source": "meta"
}