{
  "id": 398608,
  "title": "Fast preprocessing with Polars and Numba",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/398608",
  "author_name": "",
  "post_date": "2023-03-30T21:42:45.840382200Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Fast preprocessing with Polars and Numba</h1>\n<p>If you're using a LSTM like method for this competition such as <a href=\"https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker\" target=\"_blank\">LSTM Baseline by ZhaounBooty</a>. Every time you want to try out a different preprocessing idea, you have to regenerate a large set of files. This is both time-consuming and takes up a lot of space if you're using your own machine.</p>\n<p>I wanted a preprocessing setup with which I can quickly try out different preprocessing ideas without having to pre-cache the data. </p>\n<h3>Timings for preprocessing one batch:</h3>\n<ul>\n<li>Kaggle - 7 seconds</li>\n<li>i5 12400 Processor - 1.5 seconds (probably because Kaggle disk read is slow)</li>\n</ul>\n<h3>Key Insights</h3>\n<ol>\n<li>Use <code>polars</code> for loading parquet files, way faster than <code>pandas</code></li>\n<li>For padding, pre-assign the <code>np.zeros</code> array instead of using <code>np.pad</code></li>\n<li>For loop based operations operations, such as sampling points or filtering based on particular data, use <code>numba</code>. </li>\n</ol>\n<p>Here's the code, hope it helps. <br>\n<a href=\"https://www.kaggle.com/code/dipamc77/fast-preprocessing-with-polars-and-numba-1-5-s/notebook\" target=\"_blank\">https://www.kaggle.com/code/dipamc77/fast-preprocessing-with-polars-and-numba-1-5-s/notebook</a></p>",
  "messages": [
    {
      "id": "2203520",
      "postDate": "03/30/2023 21:42:45",
      "content": "<h1>Fast preprocessing with Polars and Numba</h1>\n<p>If you're using a LSTM like method for this competition such as <a href=\"https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker\" target=\"_blank\">LSTM Baseline by ZhaounBooty</a>. Every time you want to try out a different preprocessing idea, you have to regenerate a large set of files. This is both time-consuming and takes up a lot of space if you're using your own machine.</p>\n<p>I wanted a preprocessing setup with which I can quickly try out different preprocessing ideas without having to pre-cache the data. </p>\n<h3>Timings for preprocessing one batch:</h3>\n<ul>\n<li>Kaggle - 7 seconds</li>\n<li>i5 12400 Processor - 1.5 seconds (probably because Kaggle disk read is slow)</li>\n</ul>\n<h3>Key Insights</h3>\n<ol>\n<li>Use <code>polars</code> for loading parquet files, way faster than <code>pandas</code></li>\n<li>For padding, pre-assign the <code>np.zeros</code> array instead of using <code>np.pad</code></li>\n<li>For loop based operations operations, such as sampling points or filtering based on particular data, use <code>numba</code>. </li>\n</ol>\n<p>Here's the code, hope it helps. <br>\n<a href=\"https://www.kaggle.com/code/dipamc77/fast-preprocessing-with-polars-and-numba-1-5-s/notebook\" target=\"_blank\">https://www.kaggle.com/code/dipamc77/fast-preprocessing-with-polars-and-numba-1-5-s/notebook</a></p>",
      "rawMarkdown": "# Fast preprocessing with Polars and Numba\n If you're using a LSTM like method for this competition such as [LSTM Baseline by ZhaounBooty](https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker). Every time you want to try out a different preprocessing idea, you have to regenerate a large set of files. This is both time-consuming and takes up a lot of space if you're using your own machine.\n\nI wanted a preprocessing setup with which I can quickly try out different preprocessing ideas without having to pre-cache the data. \n\n### Timings for preprocessing one batch: \n- Kaggle - 7 seconds\n- i5 12400 Processor - 1.5 seconds (probably because Kaggle disk read is slow)\n\n### Key Insights\n1. Use `polars` for loading parquet files, way faster than `pandas`\n2. For padding, pre-assign the `np.zeros` array instead of using `np.pad`\n3. For loop based operations operations, such as sampling points or filtering based on particular data, use `numba`. \n\nHere's the code, hope it helps. \nhttps://www.kaggle.com/code/dipamc77/fast-preprocessing-with-polars-and-numba-1-5-s/notebook",
      "votes": null
    },
    {
      "id": "2204248",
      "postDate": "03/31/2023 12:45:51",
      "content": "<p>very good notebook! I like notebooks which are about pipeline optimization. This is important topic. Definitely good job <a href=\"https://www.kaggle.com/dipamc77\" target=\"_blank\">@dipamc77</a> 👍</p>",
      "rawMarkdown": "very good notebook! I like notebooks which are about pipeline optimization. This is important topic. Definitely good job @dipamc77 👍",
      "votes": null
    },
    {
      "id": "2204358",
      "postDate": "03/31/2023 13:59:23",
      "content": "<p>Thanks! </p>\n<p>I guess the incentive to optimize is high when one is limited by compute. 😅 </p>",
      "rawMarkdown": "Thanks! \n\nI guess the incentive to optimize is high when one is limited by compute. 😅",
      "votes": null
    },
    {
      "id": "2205309",
      "postDate": "04/01/2023 12:20:33",
      "content": "<p>Numba is amazing. If you don't do that already, try to parallelize your Numba loops with nb.prange (remember to add parallel=True in the jit decorator) for even faster processing!  </p>",
      "rawMarkdown": "Numba is amazing. If you don't do that already, try to parallelize your Numba loops with nb.prange (remember to add parallel=True in the jit decorator) for even faster processing!",
      "votes": null
    },
    {
      "id": "2205310",
      "postDate": "04/01/2023 12:22:01",
      "content": "<p>Oh, I didn't know this, I check it out. Thanks a lot.</p>",
      "rawMarkdown": "Oh, I didn't know this, I check it out. Thanks a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2204248,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "03/31/2023 12:45:51",
      "content": "<p>very good notebook! I like notebooks which are about pipeline optimization. This is important topic. Definitely good job <a href=\"https://www.kaggle.com/dipamc77\" target=\"_blank\">@dipamc77</a> 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2204358,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "03/31/2023 13:59:23",
          "content": "<p>Thanks! </p>\n<p>I guess the incentive to optimize is high when one is limited by compute. 😅 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2205309,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "04/01/2023 12:20:33",
      "content": "<p>Numba is amazing. If you don't do that already, try to parallelize your Numba loops with nb.prange (remember to add parallel=True in the jit decorator) for even faster processing!  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2205310,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/01/2023 12:22:01",
          "content": "<p>Oh, I didn't know this, I check it out. Thanks a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2203520": "# Fast preprocessing with Polars and Numba\n If you're using a LSTM like method for this competition such as [LSTM Baseline by ZhaounBooty](https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker). Every time you want to try out a different preprocessing idea, you have to regenerate a large set of files. This is both time-consuming and takes up a lot of space if you're using your own machine.\n\nI wanted a preprocessing setup with which I can quickly try out different preprocessing ideas without having to pre-cache the data. \n\n### Timings for preprocessing one batch: \n- Kaggle - 7 seconds\n- i5 12400 Processor - 1.5 seconds (probably because Kaggle disk read is slow)\n\n### Key Insights\n1. Use `polars` for loading parquet files, way faster than `pandas`\n2. For padding, pre-assign the `np.zeros` array instead of using `np.pad`\n3. For loop based operations operations, such as sampling points or filtering based on particular data, use `numba`. \n\nHere's the code, hope it helps. \nhttps://www.kaggle.com/code/dipamc77/fast-preprocessing-with-polars-and-numba-1-5-s/notebook",
    "2204248": "very good notebook! I like notebooks which are about pipeline optimization. This is important topic. Definitely good job @dipamc77 👍",
    "2204358": "Thanks! \n\nI guess the incentive to optimize is high when one is limited by compute. 😅",
    "2205309": "Numba is amazing. If you don't do that already, try to parallelize your Numba loops with nb.prange (remember to add parallel=True in the jit decorator) for even faster processing!",
    "2205310": "Oh, I didn't know this, I check it out. Thanks a lot."
  },
  "source": "meta"
}