{
  "id": 365873,
  "title": "Say NO to GPU: Nx Faster Co-Visitation Matrices using single CPU!",
  "url": "/competitions/otto-recommender-system/discussion/365873",
  "author_name": "",
  "post_date": "2022-11-13T15:38:11.076840200Z",
  "votes": 43,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I published a notebook <a href=\"https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline\" target=\"_blank\">here</a> that uses <code>numba.jit</code> to compute Nx faster co-visitation matrices. </p>\n<h3>Running Time</h3>\n<p>The training time of one co-visitation matrix is <strong>7~8min</strong> using Kaggle's CPU Kernel, even faster than <a href=\"https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline\" target=\"_blank\">Chris's cuDF implementation</a>. The inference time is <strong>1min</strong>. The total running time is only <strong>20min</strong>! Although I used <code>numba</code>'s parallel in the notebook, however, it didn't give much time boost. So I believe one single CPU with 30GB RAM can easily benefit from this great speed up.</p>\n<h3>Logic</h3>\n<p>I simplified the logic of <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\" target=\"_blank\">Chris's notebook</a> and get <strong>0.574</strong> on public leaderboard.</p>",
  "messages": [
    {
      "id": "2028170",
      "postDate": "11/13/2022 15:38:11",
      "content": "<p>I published a notebook <a href=\"https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline\" target=\"_blank\">here</a> that uses <code>numba.jit</code> to compute Nx faster co-visitation matrices. </p>\n<h3>Running Time</h3>\n<p>The training time of one co-visitation matrix is <strong>7~8min</strong> using Kaggle's CPU Kernel, even faster than <a href=\"https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline\" target=\"_blank\">Chris's cuDF implementation</a>. The inference time is <strong>1min</strong>. The total running time is only <strong>20min</strong>! Although I used <code>numba</code>'s parallel in the notebook, however, it didn't give much time boost. So I believe one single CPU with 30GB RAM can easily benefit from this great speed up.</p>\n<h3>Logic</h3>\n<p>I simplified the logic of <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\" target=\"_blank\">Chris's notebook</a> and get <strong>0.574</strong> on public leaderboard.</p>",
      "rawMarkdown": "I published a notebook [here](https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline) that uses `numba.jit` to compute Nx faster co-visitation matrices. \n\n### Running Time\n\nThe training time of one co-visitation matrix is **7~8min** using Kaggle's CPU Kernel, even faster than [Chris's cuDF implementation](https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline). The inference time is **1min**. The total running time is only **20min**! Although I used `numba`'s parallel in the notebook, however, it didn't give much time boost. So I believe one single CPU with 30GB RAM can easily benefit from this great speed up.\n\n### Logic\n\nI simplified the logic of [Chris's notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573) and get **0.574** on public leaderboard.",
      "votes": null
    },
    {
      "id": "2028180",
      "postDate": "11/13/2022 15:45:17",
      "content": "<p>this is a nice work</p>",
      "rawMarkdown": "this is a nice work",
      "votes": null
    },
    {
      "id": "2028204",
      "postDate": "11/13/2022 15:58:06",
      "content": "<p>This seems like a very useful piece of information and can benefit a lot of users greatly. Sincere thanks and regards for the share <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> </p>",
      "rawMarkdown": "This seems like a very useful piece of information and can benefit a lot of users greatly. Sincere thanks and regards for the share @carnozhao",
      "votes": null
    },
    {
      "id": "2028335",
      "postDate": "11/13/2022 19:59:34",
      "content": "<p>Hey carno. This looks nice and extremly fast. Could you also publish the numba files for the local validation data so we can make Experiments using your base notebook. Tnx in advance :))<br>\nPs: what would be even greater would be a small notebook on how to create the numba Arrays. Then people less with experience like me learn from it :)))</p>",
      "rawMarkdown": "Hey carno. This looks nice and extremly fast. Could you also publish the numba files for the local validation data so we can make Experiments using your base notebook. Tnx in advance :))\nPs: what would be even greater would be a small notebook on how to create the numba Arrays. Then people less with experience like me learn from it :)))",
      "votes": null
    },
    {
      "id": "2028398",
      "postDate": "11/13/2022 21:36:12",
      "content": "<p>Great code Carno! Your code is very fast and it's <strong>almost as fast</strong> as Chris' cuDF notebook 😄 My latest version creates co-visitation matrices in <strong>6min</strong>. (And offline on 1xV100 in <strong>1.5min</strong>)</p>",
      "rawMarkdown": "Great code Carno! Your code is very fast and it's **almost as fast** as Chris' cuDF notebook 😄 My latest version creates co-visitation matrices in **6min**. (And offline on 1xV100 in **1.5min**)",
      "votes": null
    },
    {
      "id": "2028471",
      "postDate": "11/14/2022 00:56:58",
      "content": "<p>Hi Simon, maybe you should try Chris's newest version whose performance is better than mine. If I can further boost my notebook in the future, I will release my data preprocessing and local validation codes.</p>",
      "rawMarkdown": "Hi Simon, maybe you should try Chris's newest version whose performance is better than mine. If I can further boost my notebook in the future, I will release my data preprocessing and local validation codes.",
      "votes": null
    },
    {
      "id": "2028474",
      "postDate": "11/14/2022 01:01:42",
      "content": "<p>Finally… beaten by cuDF😂<br>\nIn my code, the bottleneck is how to merge multiple counters. For me, I can only use single-thread algorithm to iterate one by one. cuDF might have better implementation of <code>Seris.add()</code> on GPU. </p>",
      "rawMarkdown": "Finally... beaten by cuDF😂\nIn my code, the bottleneck is how to merge multiple counters. For me, I can only use single-thread algorithm to iterate one by one. cuDF might have better implementation of `Seris.add()` on GPU.",
      "votes": null
    },
    {
      "id": "2028689",
      "postDate": "11/14/2022 07:07:33",
      "content": "<p>Good idea! Tnx</p>",
      "rawMarkdown": "Good idea! Tnx",
      "votes": null
    },
    {
      "id": "2032864",
      "postDate": "11/16/2022 20:58:24",
      "content": "<p>GPU cuDF is fast! 😄 Perhaps you can overcome your bottle neck with 2 stage combination. Here is an example with 2 CPU. One CPU can combine all counters with first key between 0-1e6 while a second CPU is combining all counters with first key 1e6-2e6. Then when stage 1 is done, combine the two counters together to get your result.</p>\n<p>My notebook bottleneck is disk I/O. My cuDF notebook uses a public Kaggle dataset where the data is 146 separate files and the data still includes original strings and int64. </p>\n<p>I just updated my notebook with improve disk I/O to compute co-visitation matrices in <strong>3min</strong> each using Kaggle notebook 1xP100 GPU! 🔥</p>",
      "rawMarkdown": "GPU cuDF is fast! 😄 Perhaps you can overcome your bottle neck with 2 stage combination. Here is an example with 2 CPU. One CPU can combine all counters with first key between 0-1e6 while a second CPU is combining all counters with first key 1e6-2e6. Then when stage 1 is done, combine the two counters together to get your result.\n\nMy notebook bottleneck is disk I/O. My cuDF notebook uses a public Kaggle dataset where the data is 146 separate files and the data still includes original strings and int64. \n\nI just updated my notebook with improve disk I/O to compute co-visitation matrices in **3min** each using Kaggle notebook 1xP100 GPU! 🔥",
      "votes": null
    },
    {
      "id": "2065746",
      "postDate": "12/15/2022 03:20:56",
      "content": "<p>Are you using the full data with this notebook</p>",
      "rawMarkdown": "Are you using the full data with this notebook",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2028180,
      "author_name": "catchbat",
      "author_url": "",
      "post_date": "11/13/2022 15:45:17",
      "content": "<p>this is a nice work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2028204,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/13/2022 15:58:06",
      "content": "<p>This seems like a very useful piece of information and can benefit a lot of users greatly. Sincere thanks and regards for the share <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2028335,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "11/13/2022 19:59:34",
      "content": "<p>Hey carno. This looks nice and extremly fast. Could you also publish the numba files for the local validation data so we can make Experiments using your base notebook. Tnx in advance :))<br>\nPs: what would be even greater would be a small notebook on how to create the numba Arrays. Then people less with experience like me learn from it :)))</p>",
      "votes": null,
      "replies": [
        {
          "id": 2028471,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "11/14/2022 00:56:58",
          "content": "<p>Hi Simon, maybe you should try Chris's newest version whose performance is better than mine. If I can further boost my notebook in the future, I will release my data preprocessing and local validation codes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2028689,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "11/14/2022 07:07:33",
          "content": "<p>Good idea! Tnx</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2028398,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/13/2022 21:36:12",
      "content": "<p>Great code Carno! Your code is very fast and it's <strong>almost as fast</strong> as Chris' cuDF notebook 😄 My latest version creates co-visitation matrices in <strong>6min</strong>. (And offline on 1xV100 in <strong>1.5min</strong>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2028474,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "11/14/2022 01:01:42",
          "content": "<p>Finally… beaten by cuDF😂<br>\nIn my code, the bottleneck is how to merge multiple counters. For me, I can only use single-thread algorithm to iterate one by one. cuDF might have better implementation of <code>Seris.add()</code> on GPU. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2032864,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/16/2022 20:58:24",
          "content": "<p>GPU cuDF is fast! 😄 Perhaps you can overcome your bottle neck with 2 stage combination. Here is an example with 2 CPU. One CPU can combine all counters with first key between 0-1e6 while a second CPU is combining all counters with first key 1e6-2e6. Then when stage 1 is done, combine the two counters together to get your result.</p>\n<p>My notebook bottleneck is disk I/O. My cuDF notebook uses a public Kaggle dataset where the data is 146 separate files and the data still includes original strings and int64. </p>\n<p>I just updated my notebook with improve disk I/O to compute co-visitation matrices in <strong>3min</strong> each using Kaggle notebook 1xP100 GPU! 🔥</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2065746,
      "author_name": "luyi564",
      "author_url": "",
      "post_date": "12/15/2022 03:20:56",
      "content": "<p>Are you using the full data with this notebook</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2028170": "I published a notebook [here](https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline) that uses `numba.jit` to compute Nx faster co-visitation matrices. \n\n### Running Time\n\nThe training time of one co-visitation matrix is **7~8min** using Kaggle's CPU Kernel, even faster than [Chris's cuDF implementation](https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline). The inference time is **1min**. The total running time is only **20min**! Although I used `numba`'s parallel in the notebook, however, it didn't give much time boost. So I believe one single CPU with 30GB RAM can easily benefit from this great speed up.\n\n### Logic\n\nI simplified the logic of [Chris's notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573) and get **0.574** on public leaderboard.",
    "2028180": "this is a nice work",
    "2028204": "This seems like a very useful piece of information and can benefit a lot of users greatly. Sincere thanks and regards for the share @carnozhao",
    "2028335": "Hey carno. This looks nice and extremly fast. Could you also publish the numba files for the local validation data so we can make Experiments using your base notebook. Tnx in advance :))\nPs: what would be even greater would be a small notebook on how to create the numba Arrays. Then people less with experience like me learn from it :)))",
    "2028398": "Great code Carno! Your code is very fast and it's **almost as fast** as Chris' cuDF notebook 😄 My latest version creates co-visitation matrices in **6min**. (And offline on 1xV100 in **1.5min**)",
    "2028471": "Hi Simon, maybe you should try Chris's newest version whose performance is better than mine. If I can further boost my notebook in the future, I will release my data preprocessing and local validation codes.",
    "2028474": "Finally... beaten by cuDF😂\nIn my code, the bottleneck is how to merge multiple counters. For me, I can only use single-thread algorithm to iterate one by one. cuDF might have better implementation of `Seris.add()` on GPU.",
    "2028689": "Good idea! Tnx",
    "2032864": "GPU cuDF is fast! 😄 Perhaps you can overcome your bottle neck with 2 stage combination. Here is an example with 2 CPU. One CPU can combine all counters with first key between 0-1e6 while a second CPU is combining all counters with first key 1e6-2e6. Then when stage 1 is done, combine the two counters together to get your result.\n\nMy notebook bottleneck is disk I/O. My cuDF notebook uses a public Kaggle dataset where the data is 146 separate files and the data still includes original strings and int64. \n\nI just updated my notebook with improve disk I/O to compute co-visitation matrices in **3min** each using Kaggle notebook 1xP100 GPU! 🔥",
    "2065746": "Are you using the full data with this notebook"
  },
  "source": "meta"
}