{
  "id": 381672,
  "title": "X15 Faster Inference : Using the TPU VM -3-8 vs CPU",
  "url": "/competitions/otto-recommender-system/discussion/381672",
  "author_name": "",
  "post_date": "2023-01-27T17:43:57.923559500Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello community, I wanted to share this experiments with you.</p>\n<p>I struggled for a long time with the Ranker inference, especially when using multiples rankers through GroupKFold or any  other k-fold techniques.</p>\n<p>Here is my results:</p>\n<ul>\n<li>I'm using <strong>50 variables</strong> ( it's important to note that ).</li>\n<li>Those results concern 1/10 of the validation set, because I'm using 10 CHUNKS, just multiply the <strong>time per chunk</strong> by 10 to get the overall time elapsed .</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Number of models</th>\n<th>Number of estimators</th>\n<th>time per chunk</th>\n<th>overall time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3</td>\n<td>50</td>\n<td>1min12</td>\n<td>11 min</td>\n</tr>\n<tr>\n<td>3</td>\n<td>100</td>\n<td>2 min</td>\n<td>20 min</td>\n</tr>\n<tr>\n<td>3</td>\n<td>300</td>\n<td>5min23 sec</td>\n<td>51 min</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1000</td>\n<td>22 min.</td>\n<td>3h 43min</td>\n</tr>\n</tbody>\n</table>\n<p>As you can see, using the CPU kaggle kernel is  not enough, as pointed out by <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>   <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381469\" target=\"_blank\">here</a>. </p>\n<p>However, it's probably ideal to use GPU but I got some limitations and memory errors in my pipeline with kaggle kernels, and my ranker is already trained in CPU. </p>\n<p>Using the TPU VM -3-8,  the inference time is reduced by about 15 times :</p>\n<table>\n<thead>\n<tr>\n<th>Number of models</th>\n<th>Number of estimators</th>\n<th>time per chunk</th>\n<th>overall time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3</td>\n<td>1000</td>\n<td>1 min36.</td>\n<td>~15 min</td>\n</tr>\n<tr>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2Fcf068fe379cfb39f940338d2900b8983%2FCapture%20decran%202023-01-27%20a%206.38.11%20PM.png?generation=1674841108941183&amp;alt=media\" alt=\"\"></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>The only constraint is that I couldn't use Polars, so I convert it to Pandas  and you'll have to wait in order to use the TPU.</p>\n<p>Summary:</p>\n<table>\n<thead>\n<tr>\n<th>Case</th>\n<th>Disadvantages</th>\n<th>Advantages</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CPU</td>\n<td>Too slow</td>\n<td>immediat use / you don't have to wait</td>\n</tr>\n<tr>\n<td>TPU</td>\n<td>Fast / don't need to chunk</td>\n<td>Queue / Be aware of the lib you're using</td>\n</tr>\n</tbody>\n</table>\n<p><em>Another essential point before executing your code in TPU, please be really careful, test your code before execution, because if you made a little typo, you'll have to wait again.</em></p>",
  "messages": [
    {
      "id": "2118003",
      "postDate": "01/27/2023 17:43:57",
      "content": "<p>Hello community, I wanted to share this experiments with you.</p>\n<p>I struggled for a long time with the Ranker inference, especially when using multiples rankers through GroupKFold or any  other k-fold techniques.</p>\n<p>Here is my results:</p>\n<ul>\n<li>I'm using <strong>50 variables</strong> ( it's important to note that ).</li>\n<li>Those results concern 1/10 of the validation set, because I'm using 10 CHUNKS, just multiply the <strong>time per chunk</strong> by 10 to get the overall time elapsed .</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Number of models</th>\n<th>Number of estimators</th>\n<th>time per chunk</th>\n<th>overall time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3</td>\n<td>50</td>\n<td>1min12</td>\n<td>11 min</td>\n</tr>\n<tr>\n<td>3</td>\n<td>100</td>\n<td>2 min</td>\n<td>20 min</td>\n</tr>\n<tr>\n<td>3</td>\n<td>300</td>\n<td>5min23 sec</td>\n<td>51 min</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1000</td>\n<td>22 min.</td>\n<td>3h 43min</td>\n</tr>\n</tbody>\n</table>\n<p>As you can see, using the CPU kaggle kernel is  not enough, as pointed out by <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>   <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381469\" target=\"_blank\">here</a>. </p>\n<p>However, it's probably ideal to use GPU but I got some limitations and memory errors in my pipeline with kaggle kernels, and my ranker is already trained in CPU. </p>\n<p>Using the TPU VM -3-8,  the inference time is reduced by about 15 times :</p>\n<table>\n<thead>\n<tr>\n<th>Number of models</th>\n<th>Number of estimators</th>\n<th>time per chunk</th>\n<th>overall time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3</td>\n<td>1000</td>\n<td>1 min36.</td>\n<td>~15 min</td>\n</tr>\n<tr>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2Fcf068fe379cfb39f940338d2900b8983%2FCapture%20decran%202023-01-27%20a%206.38.11%20PM.png?generation=1674841108941183&amp;alt=media\" alt=\"\"></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>The only constraint is that I couldn't use Polars, so I convert it to Pandas  and you'll have to wait in order to use the TPU.</p>\n<p>Summary:</p>\n<table>\n<thead>\n<tr>\n<th>Case</th>\n<th>Disadvantages</th>\n<th>Advantages</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CPU</td>\n<td>Too slow</td>\n<td>immediat use / you don't have to wait</td>\n</tr>\n<tr>\n<td>TPU</td>\n<td>Fast / don't need to chunk</td>\n<td>Queue / Be aware of the lib you're using</td>\n</tr>\n</tbody>\n</table>\n<p><em>Another essential point before executing your code in TPU, please be really careful, test your code before execution, because if you made a little typo, you'll have to wait again.</em></p>",
      "rawMarkdown": "Hello community, I wanted to share this experiments with you.\n\nI struggled for a long time with the Ranker inference, especially when using multiples rankers through GroupKFold or any  other k-fold techniques.\n\nHere is my results:\n- I'm using **50 variables** ( it's important to note that ).\n- Those results concern 1/10 of the validation set, because I'm using 10 CHUNKS, just multiply the **time per chunk** by 10 to get the overall time elapsed .\n\n| Number of models | Number of estimators  | time per chunk  | overall time\n|---|---|---| ---|\n| 3 | 50 | 1min12 | 11 min\n| 3 | 100  | 2 min | 20 min\n| 3 | 300  |  5min23 sec | 51 min\n| 3 | 1000  | 22 min. |   3h 43min\n\n\nAs you can see, using the CPU kaggle kernel is  not enough, as pointed out by @buumoo   [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381469). \n\nHowever, it's probably ideal to use GPU but I got some limitations and memory errors in my pipeline with kaggle kernels, and my ranker is already trained in CPU. \n\n\nUsing the TPU VM -3-8,  the inference time is reduced by about 15 times :\n| Number of models | Number of estimators  | time per chunk  | overall time\n|---|---|---| ---|\n| 3 | 1000  | 1 min36. |   ~15 min\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2Fcf068fe379cfb39f940338d2900b8983%2FCapture%20decran%202023-01-27%20a%206.38.11%20PM.png?generation=1674841108941183&alt=media)\n\nThe only constraint is that I couldn't use Polars, so I convert it to Pandas  and you'll have to wait in order to use the TPU.\n\nSummary:\n| Case |  Disadvantages | Advantages\n| --- | --- |\n|  CPU | Too slow | immediat use / you don't have to wait\n|  TPU | Fast / don't need to chunk | Queue / Be aware of the lib you're using \n\n*Another essential point before executing your code in TPU, please be really careful, test your code before execution, because if you made a little typo, you'll have to wait again.*",
      "votes": null
    },
    {
      "id": "2118073",
      "postDate": "01/27/2023 18:39:34",
      "content": "<p>You may want to use cuml.ForestInference to accelerate the inference speed for tree-based models, no matter it is trained by GPU or CPU.</p>",
      "rawMarkdown": "You may want to use cuml.ForestInference to accelerate the inference speed for tree-based models, no matter it is trained by GPU or CPU.",
      "votes": null
    },
    {
      "id": "2119094",
      "postDate": "01/28/2023 13:49:17",
      "content": "<p>is it faster than treelite ? If run on CPU ?</p>",
      "rawMarkdown": "is it faster than treelite ? If run on CPU ?",
      "votes": null
    },
    {
      "id": "2119158",
      "postDate": "01/28/2023 14:21:35",
      "content": "<p>If I am right, I think ForestInference is a GPU-only lib, I am not very sure.<br>\nIf runing it on GPU, it is supposed to be much faster than treelite.<br>\nI found a comparision figure for your reference.<br>\nanyway, the syntax is simple, just give it a shot, maybe it will surprise you.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F8dd3e9d24a135ed92197153baeffe7b5%2F0_KAOl-8S4gIzFyhly.png?generation=1674915594600121&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "If I am right, I think ForestInference is a GPU-only lib, I am not very sure.\nIf runing it on GPU, it is supposed to be much faster than treelite.\nI found a comparision figure for your reference.\nanyway, the syntax is simple, just give it a shot, maybe it will surprise you.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F8dd3e9d24a135ed92197153baeffe7b5%2F0_KAOl-8S4gIzFyhly.png?generation=1674915594600121&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2118073,
      "author_name": "buumoo",
      "author_url": "",
      "post_date": "01/27/2023 18:39:34",
      "content": "<p>You may want to use cuml.ForestInference to accelerate the inference speed for tree-based models, no matter it is trained by GPU or CPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2119094,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "01/28/2023 13:49:17",
          "content": "<p>is it faster than treelite ? If run on CPU ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2119158,
              "author_name": "buumoo",
              "author_url": "",
              "post_date": "01/28/2023 14:21:35",
              "content": "<p>If I am right, I think ForestInference is a GPU-only lib, I am not very sure.<br>\nIf runing it on GPU, it is supposed to be much faster than treelite.<br>\nI found a comparision figure for your reference.<br>\nanyway, the syntax is simple, just give it a shot, maybe it will surprise you.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F8dd3e9d24a135ed92197153baeffe7b5%2F0_KAOl-8S4gIzFyhly.png?generation=1674915594600121&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2118003": "Hello community, I wanted to share this experiments with you.\n\nI struggled for a long time with the Ranker inference, especially when using multiples rankers through GroupKFold or any  other k-fold techniques.\n\nHere is my results:\n- I'm using **50 variables** ( it's important to note that ).\n- Those results concern 1/10 of the validation set, because I'm using 10 CHUNKS, just multiply the **time per chunk** by 10 to get the overall time elapsed .\n\n| Number of models | Number of estimators  | time per chunk  | overall time\n|---|---|---| ---|\n| 3 | 50 | 1min12 | 11 min\n| 3 | 100  | 2 min | 20 min\n| 3 | 300  |  5min23 sec | 51 min\n| 3 | 1000  | 22 min. |   3h 43min\n\n\nAs you can see, using the CPU kaggle kernel is  not enough, as pointed out by @buumoo   [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381469). \n\nHowever, it's probably ideal to use GPU but I got some limitations and memory errors in my pipeline with kaggle kernels, and my ranker is already trained in CPU. \n\n\nUsing the TPU VM -3-8,  the inference time is reduced by about 15 times :\n| Number of models | Number of estimators  | time per chunk  | overall time\n|---|---|---| ---|\n| 3 | 1000  | 1 min36. |   ~15 min\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2Fcf068fe379cfb39f940338d2900b8983%2FCapture%20decran%202023-01-27%20a%206.38.11%20PM.png?generation=1674841108941183&alt=media)\n\nThe only constraint is that I couldn't use Polars, so I convert it to Pandas  and you'll have to wait in order to use the TPU.\n\nSummary:\n| Case |  Disadvantages | Advantages\n| --- | --- |\n|  CPU | Too slow | immediat use / you don't have to wait\n|  TPU | Fast / don't need to chunk | Queue / Be aware of the lib you're using \n\n*Another essential point before executing your code in TPU, please be really careful, test your code before execution, because if you made a little typo, you'll have to wait again.*",
    "2118073": "You may want to use cuml.ForestInference to accelerate the inference speed for tree-based models, no matter it is trained by GPU or CPU.",
    "2119094": "is it faster than treelite ? If run on CPU ?",
    "2119158": "If I am right, I think ForestInference is a GPU-only lib, I am not very sure.\nIf runing it on GPU, it is supposed to be much faster than treelite.\nI found a comparision figure for your reference.\nanyway, the syntax is simple, just give it a shot, maybe it will surprise you.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F8dd3e9d24a135ed92197153baeffe7b5%2F0_KAOl-8S4gIzFyhly.png?generation=1674915594600121&alt=media)"
  },
  "source": "meta"
}