{
  "id": 180359,
  "title": "OpenGL image rendering on GPU",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/180359",
  "author_name": "Peter",
  "post_date": "2020-09-04T18:03:46.875000",
  "votes": 55,
  "comment_count": 29,
  "views": 0,
  "content": "<p>As you probably already know, the rasterization process is the bottleneck in this competition. At first, I tried to optimize the l5kit, but I only achieved minor speedup. With pre-generated images (and the optimized l5kit), I achieved ~ 4-5x  it/sec. Unfortunately, the GPU utilization was still ~40-45% (I have an RTX-2080ti I used 10 cores; 20 CPU threads). The obvious solution is to render the images on the GPU. As <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> mentioned on their <a href=\"https://github.com/lyft/l5kit/issues/136#issuecomment-686370259\" target=\"_blank\">github discussion</a> that it is not an easy task with the current implementation.</p>\n<p>I implemented an OpenGLRasterizer, and it looks very promising. Here is my baseline benchmark.</p>\n<h2>Results</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F85e632b2df7be4c66926a84160078834%2Frasterizer_benchmark.png?generation=1599242281108744&amp;alt=media\" alt=\"\"></p>\n<h2>Experiments:</h2>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Size</th>\n<th>History</th>\n<th># of samples</th>\n<th>Running time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CPU rasterizer</td>\n<td>650px</td>\n<td>0</td>\n<td>1,000</td>\n<td>1:14</td>\n<td>13.41</td>\n</tr>\n<tr>\n<td>OpenGL GPU</td>\n<td>650px</td>\n<td>0</td>\n<td>10,000</td>\n<td>1:16</td>\n<td>130.49</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Notes</strong></p>\n<ul>\n<li>I haven't implemented the crosswalks and the traffic lights yet. (CPU rasterizer generates both)</li>\n<li>It is hard to generate the history (in the current output form) on the GPU (at least need some CPU post-process). I am looking for alternatives.</li>\n<li>l5kit optimization is in progress (see <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">their github</a>; both methods will improve.</li>\n<li>Data loading is still a bottleneck. OpenGL can generate at a ~800 it/sec rate if I use the same preloaded frame's data. (I used a busy frame with lots of agents)</li>\n<li>I used 1 CPU core for both methods.</li>\n<li>GPU memory usage is low: ~35Mb.</li>\n</ul>\n<p>I have to solve/fix a few things before I can publish the code. I don't want to make any promises because I am busy with other projects, but I expect to be ready next week.</p>\n<hr>\n<p><strong>Update</strong><br>\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).</p>\n<p>You can find more details and the source code <a href=\"https://github.com/pestipeti/LyftOpenGLRasterizer\" target=\"_blank\">on my github</a>.</p>",
  "messages": [
    {
      "id": 998423,
      "postDate": "2020-09-04T18:03:46.877Z",
      "content": "<p>As you probably already know, the rasterization process is the bottleneck in this competition. At first, I tried to optimize the l5kit, but I only achieved minor speedup. With pre-generated images (and the optimized l5kit), I achieved ~ 4-5x  it/sec. Unfortunately, the GPU utilization was still ~40-45% (I have an RTX-2080ti I used 10 cores; 20 CPU threads). The obvious solution is to render the images on the GPU. As <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> mentioned on their <a href=\"https://github.com/lyft/l5kit/issues/136#issuecomment-686370259\" target=\"_blank\">github discussion</a> that it is not an easy task with the current implementation.</p>\n<p>I implemented an OpenGLRasterizer, and it looks very promising. Here is my baseline benchmark.</p>\n<h2>Results</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F85e632b2df7be4c66926a84160078834%2Frasterizer_benchmark.png?generation=1599242281108744&amp;alt=media\" alt=\"\"></p>\n<h2>Experiments:</h2>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Size</th>\n<th>History</th>\n<th># of samples</th>\n<th>Running time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CPU rasterizer</td>\n<td>650px</td>\n<td>0</td>\n<td>1,000</td>\n<td>1:14</td>\n<td>13.41</td>\n</tr>\n<tr>\n<td>OpenGL GPU</td>\n<td>650px</td>\n<td>0</td>\n<td>10,000</td>\n<td>1:16</td>\n<td>130.49</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Notes</strong></p>\n<ul>\n<li>I haven't implemented the crosswalks and the traffic lights yet. (CPU rasterizer generates both)</li>\n<li>It is hard to generate the history (in the current output form) on the GPU (at least need some CPU post-process). I am looking for alternatives.</li>\n<li>l5kit optimization is in progress (see <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">their github</a>; both methods will improve.</li>\n<li>Data loading is still a bottleneck. OpenGL can generate at a ~800 it/sec rate if I use the same preloaded frame's data. (I used a busy frame with lots of agents)</li>\n<li>I used 1 CPU core for both methods.</li>\n<li>GPU memory usage is low: ~35Mb.</li>\n</ul>\n<p>I have to solve/fix a few things before I can publish the code. I don't want to make any promises because I am busy with other projects, but I expect to be ready next week.</p>\n<hr>\n<p><strong>Update</strong><br>\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).</p>\n<p>You can find more details and the source code <a href=\"https://github.com/pestipeti/LyftOpenGLRasterizer\" target=\"_blank\">on my github</a>.</p>",
      "rawMarkdown": "As you probably already know, the rasterization process is the bottleneck in this competition. At first, I tried to optimize the l5kit, but I only achieved minor speedup. With pre-generated images (and the optimized l5kit), I achieved ~ 4-5x  it/sec. Unfortunately, the GPU utilization was still ~40-45% (I have an RTX-2080ti I used 10 cores; 20 CPU threads). The obvious solution is to render the images on the GPU. As @lucabergamini mentioned on their [github discussion](https://github.com/lyft/l5kit/issues/136#issuecomment-686370259) that it is not an easy task with the current implementation.\n\nI implemented an OpenGLRasterizer, and it looks very promising. Here is my baseline benchmark.\n\n## Results\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F85e632b2df7be4c66926a84160078834%2Frasterizer_benchmark.png?generation=1599242281108744&alt=media)\n\n## Experiments:\n| Method             | Size | History | \\# of samples | Running time | it/sec |\n| --------------- | ---- | ------- | ------------ | ------------- | ------ |\n| CPU rasterizer | 650px | 0           | 1,000           | 1:14                | 13.41   |\n| OpenGL GPU   | 650px | 0           | 10,000         | 1:16                | 130.49   |\n\n**Notes**\n- I haven't implemented the crosswalks and the traffic lights yet. (CPU rasterizer generates both)\n- It is hard to generate the history (in the current output form) on the GPU (at least need some CPU post-process). I am looking for alternatives.\n- l5kit optimization is in progress (see [their github](https://github.com/lyft/l5kit); both methods will improve.\n- Data loading is still a bottleneck. OpenGL can generate at a ~800 it/sec rate if I use the same preloaded frame's data. (I used a busy frame with lots of agents)\n- I used 1 CPU core for both methods.\n- GPU memory usage is low: ~35Mb.\n\nI have to solve/fix a few things before I can publish the code. I don't want to make any promises because I am busy with other projects, but I expect to be ready next week.\n\n--------------\n\n**Update**\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).\n\nYou can find more details and the source code [on my github](https://github.com/pestipeti/LyftOpenGLRasterizer).\n",
      "votes": 55
    },
    {
      "id": 1003526,
      "postDate": "2020-09-09T04:08:19.133Z",
      "content": "<p><strong>Update</strong><br>\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).</p>\n<p>You can find more details and the source code <a href=\"https://github.com/pestipeti/LyftOpenGLRasterizer\" target=\"_blank\">on my github</a>.</p>",
      "rawMarkdown": "**Update**\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).\n\nYou can find more details and the source code [on my github](https://github.com/pestipeti/LyftOpenGLRasterizer).\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 1012015,
          "postDate": "2020-09-15T21:05:03.117Z",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> Thanks for sharing your progress/work. Given the difficulties going back/forth between CPU GPU, GL/PyTorch wonder if you've looked at fully Torch tensor based options?</p>\n<p>After seeing how bad the bottleneck is, I looked for some native PyTorch 2D drawing/rasterization but didn't find anything. Kornia has some discussion in issue tracker but no implemented functionality. I then looked for 3d options and ran into two libs I'd looked at before but never tried.</p>\n<p>Both NVIDIA Kaolin (<a href=\"https://github.com/NVIDIAGameWorks/kaolin\" target=\"_blank\">https://github.com/NVIDIAGameWorks/kaolin</a>) and Facebook PyTorch3d (<a href=\"https://github.com/facebookresearch/pytorch3d\" target=\"_blank\">https://github.com/facebookresearch/pytorch3d</a>) have 3d rasterizers that may have enough functionality. They are full differentiable, which is not needed here, but does mean they should work with PyTorch cuda tensors.</p>",
          "rawMarkdown": "@pestipeti Thanks for sharing your progress/work. Given the difficulties going back/forth between CPU GPU, GL/PyTorch wonder if you've looked at fully Torch tensor based options?\n\nAfter seeing how bad the bottleneck is, I looked for some native PyTorch 2D drawing/rasterization but didn't find anything. Kornia has some discussion in issue tracker but no implemented functionality. I then looked for 3d options and ran into two libs I'd looked at before but never tried.\n\nBoth NVIDIA Kaolin (https://github.com/NVIDIAGameWorks/kaolin) and Facebook PyTorch3d (https://github.com/facebookresearch/pytorch3d) have 3d rasterizers that may have enough functionality. They are full differentiable, which is not needed here, but does mean they should work with PyTorch cuda tensors.",
          "votes": 4
        },
        {
          "id": 1012654,
          "postDate": "2020-09-16T07:58:23.917Z",
          "content": "<p>I looked for PyTorch 2D drawing, too; the same result. 3D tools did not come to my mind. I'll take a look at both. Thank you for the links.</p>",
          "rawMarkdown": "I looked for PyTorch 2D drawing, too; the same result. 3D tools did not come to my mind. I'll take a look at both. Thank you for the links.",
          "votes": 1
        },
        {
          "id": 1021480,
          "postDate": "2020-09-21T22:11:19.987Z",
          "content": "<p>I received impressions on your code, so I've written a rendering process in PyTorch3d. But I have yet to see any speed improvements. I'm rendering in Torch on the GPU, can you make this code better?<br>\n<a href=\"https://www.kaggle.com/nmygle/pytorch3d-rendering\" target=\"_blank\">https://www.kaggle.com/nmygle/pytorch3d-rendering</a></p>",
          "rawMarkdown": "I received impressions on your code, so I've written a rendering process in PyTorch3d. But I have yet to see any speed improvements. I'm rendering in Torch on the GPU, can you make this code better?\nhttps://www.kaggle.com/nmygle/pytorch3d-rendering"
        }
      ]
    },
    {
      "id": 1003062,
      "postDate": "2020-09-08T16:29:05.083Z",
      "content": "<p>this is <strong>amazing</strong> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>! I can't help you on the openGL side, but If you start a PR in L5Kit I can surely assist with integration. The rasterisation in Python is a huge bottleneck, so everything that can help on that side is more than welcomed!</p>",
      "rawMarkdown": "this is **amazing** @pestipeti! I can't help you on the openGL side, but If you start a PR in L5Kit I can surely assist with integration. The rasterisation in Python is a huge bottleneck, so everything that can help on that side is more than welcomed!",
      "votes": 1,
      "replies": [
        {
          "id": 1003075,
          "postDate": "2020-09-08T16:36:57.913Z",
          "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> </p>\n<p>Unfortunately, it won't work. The problem is the same that you mentioned on the github discussion. The data transfer between OpenGL gpu mem -&gt; ram -&gt; pytorch gpu tensor kills the entire idea. OpenGL is extremely fast until I want to \"download\" the data. There is a (possible) solution in C++, but that won't help us in python.</p>\n<p>I'll publish the code later, maybe someone will come up with a solution.</p>",
          "rawMarkdown": "@lucabergamini \n\nUnfortunately, it won't work. The problem is the same that you mentioned on the github discussion. The data transfer between OpenGL gpu mem -> ram -> pytorch gpu tensor kills the entire idea. OpenGL is extremely fast until I want to \"download\" the data. There is a (possible) solution in C++, but that won't help us in python.\n\n I'll publish the code later, maybe someone will come up with a solution.",
          "votes": 1
        },
        {
          "id": 1003086,
          "postDate": "2020-09-08T16:45:12.670Z",
          "content": "<p>Yeah, that's problematic. A possible solution would be a C++ implementation wrapped in Python (that alone would probably speed things up quite a lot), but sadly this is again outside my expertise :( </p>",
          "rawMarkdown": "Yeah, that's problematic. A possible solution would be a C++ implementation wrapped in Python (that alone would probably speed things up quite a lot), but sadly this is again outside my expertise :( ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1002197,
      "postDate": "2020-09-07T23:02:17.570Z",
      "content": "<p>Did you try a context.save / context.restore type method for the background data? - it would probably just be the lane info, but you should be able to reuse for some number of frames before drawing the next</p>\n<p>Added: there's a stray assert in an is_lane()?, but also the map api cache is set quite low at 90kb?</p>",
      "rawMarkdown": "Did you try a context.save / context.restore type method for the background data? - it would probably just be the lane info, but you should be able to reuse for some number of frames before drawing the next\n\nAdded: there's a stray assert in an is_lane()?, but also the map api cache is set quite low at 90kb?",
      "votes": 1,
      "replies": [
        {
          "id": 1002362,
          "postDate": "2020-09-08T04:38:49.490Z",
          "content": "<p>I stored the semantic map in vbo/vao (only ~30mb gpu mem). It is fast, no need for context. The real problem is moving the image/pixels from opengl -&gt; cpu -&gt; pytorch cuda tensor. </p>",
          "rawMarkdown": "I stored the semantic map in vbo/vao (only ~30mb gpu mem). It is fast, no need for context. The real problem is moving the image/pixels from opengl -> cpu -> pytorch cuda tensor. ",
          "votes": 1
        },
        {
          "id": 1002442,
          "postDate": "2020-09-08T06:08:32.153Z",
          "content": "<p>Maybe passing a shared memory reference would work, but it's just a guess</p>",
          "rawMarkdown": "Maybe passing a shared memory reference would work, but it's just a guess"
        },
        {
          "id": 1002480,
          "postDate": "2020-09-08T06:56:40.807Z",
          "content": "<p>Unfortunately,  I don't think we have that level of control over opengl in python (or I don't know how to do it). I'll share the code later today or tomorrow, maybe someone will come up with a solution.</p>",
          "rawMarkdown": "Unfortunately,  I don't think we have that level of control over opengl in python (or I don't know how to do it). I'll share the code later today or tomorrow, maybe someone will come up with a solution."
        },
        {
          "id": 1003395,
          "postDate": "2020-09-08T23:26:28.063Z",
          "content": "<p>Oldish article with c++ <a href=\"https://www.3dgep.com/opengl-interoperability-with-cuda/\" target=\"_blank\">https://www.3dgep.com/opengl-interoperability-with-cuda/</a></p>",
          "rawMarkdown": "Oldish article with c++ https://www.3dgep.com/opengl-interoperability-with-cuda/"
        }
      ]
    },
    {
      "id": 3300065,
      "postDate": "2025-10-09T12:24:40.057Z",
      "content": "<p>Thanks for the deep dive into OpenGL rasterization! That 10x speed boost is huge, especially since CPU rasterization is the main slow point.</p>\n<p>Your benchmark of 130 it/sec vs 13 it/sec is impressive, even without history rendering. It makes sense that history generation on GPU is tricky due to how l5kit works.</p>\n<p>Looks like data loading is now the bottleneck, showing how much faster GPU rasterization is. Pre-caching images might still be easier for training, but your method could be great for inference or live use.</p>",
      "rawMarkdown": "Thanks for the deep dive into OpenGL rasterization! That 10x speed boost is huge, especially since CPU rasterization is the main slow point.\n\nYour benchmark of 130 it/sec vs 13 it/sec is impressive, even without history rendering. It makes sense that history generation on GPU is tricky due to how l5kit works.\n\nLooks like data loading is now the bottleneck, showing how much faster GPU rasterization is. Pre-caching images might still be easier for training, but your method could be great for inference or live use."
    },
    {
      "id": 1061559,
      "postDate": "2020-10-27T04:53:19.547Z",
      "content": "<ul>\n<li>Is this notebook works on new L5kit?</li>\n<li>Do they have any improvement?</li>\n<li>Your table show it has but on the same time you are saying OpenGl GPU -&gt; CPU -&gt; Pytorch Tensor. Can you please clarify this?</li>\n</ul>\n<blockquote>\n  <p>I haven't implemented the crosswalks and the traffic lights yet.&lt;</p>\n</blockquote>\n<ul>\n<li>In your Github, Have you included them in the Github</li>\n<li>Will it improves if we improve all the function as a torch function?</li>\n<li>or by <code>torch.multiprocessing</code> will that works?</li>\n<li>What time it takes to complete the <code>train.zarr</code> file with your fastest method you drew for this competition to become 8th place in Public. If you don't mind Can you share how you done them?</li>\n</ul>",
      "rawMarkdown": "- Is this notebook works on new L5kit?\n- Do they have any improvement?\n- Your table show it has but on the same time you are saying OpenGl GPU -> CPU -> Pytorch Tensor. Can you please clarify this?\n> I haven't implemented the crosswalks and the traffic lights yet.<\n\n\n- In your Github, Have you included them in the Github\n- Will it improves if we improve all the function as a torch function?\n- or by `torch.multiprocessing` will that works?\n- What time it takes to complete the `train.zarr` file with your fastest method you drew for this competition to become 8th place in Public. If you don't mind Can you share how you done them?",
      "replies": [
        {
          "id": 1061616,
          "postDate": "2020-10-27T06:24:06.117Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> </p>\n<ul>\n<li>I did not try with the new version of l5kit</li>\n<li>No improvements. I gave up on this idea.</li>\n<li>The problem is the same. OpenGL is fast, but copying the data from opengl's memory to a tensor is slow</li>\n<li>No improvements, no crosswalks/traffic lights/no history</li>\n<li>If you can find a solution on how to pass memory reference between opengl and torch tensor it would be usable</li>\n<li>no, opengl is single-threaded. </li>\n<li>for 8th place, I used l5kit v1.0.6 and I optimized it to run faster. The training for 8th place ran ~8-9days</li>\n</ul>",
          "rawMarkdown": "@morizin \n\n- I did not try with the new version of l5kit\n- No improvements. I gave up on this idea.\n- The problem is the same. OpenGL is fast, but copying the data from opengl's memory to a tensor is slow\n- No improvements, no crosswalks/traffic lights/no history\n- If you can find a solution on how to pass memory reference between opengl and torch tensor it would be usable\n- no, opengl is single-threaded. \n- for 8th place, I used l5kit v1.0.6 and I optimized it to run faster. The training for 8th place ran ~8-9days",
          "votes": 1
        },
        {
          "id": 1061645,
          "postDate": "2020-10-27T06:59:36.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> </p>\n<ul>\n<li>I tried you Github, I had an error, but don't care.</li>\n<li>But you project out a table in the discussion that it's faster. I don't get you much right now</li>\n<li>I might think making all the function as Cython function will sometimes works don't know.</li>\n<li>will making all the functions as numba function will work (still don't know)</li>\n<li>I maybe will go on the way to recode all the code to C++ (if trying I want to dedicate 5 days for it)</li>\n<li>But don't know how <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> in this discussion made his training to 2.5 days</li>\n<li>when trained its 22 days for train.zarr (without optimizing l5kit)</li>\n<li>This is Open topic for this competition</li>\n</ul>\n<h5>Confused All about</h5>",
          "rawMarkdown": "@pestipeti \n- I tried you Github, I had an error, but don't care.\n- But you project out a table in the discussion that it's faster. I don't get you much right now\n- I might think making all the function as Cython function will sometimes works don't know.\n- will making all the functions as numba function will work (still don't know)\n- I maybe will go on the way to recode all the code to C++ (if trying I want to dedicate 5 days for it)\n- But don't know how @ilu000 in this discussion made his training to 2.5 days\n- when trained its 22 days for train.zarr (without optimizing l5kit)\n- This is Open topic for this competition\n##### Confused All about"
        },
        {
          "id": 1061666,
          "postDate": "2020-10-27T07:31:17.577Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> <br>\nSorry for the confusion. That table shows stats about one image generated on CPU (one core) vs GPU. When you want to generate a batch of images, you can use multiple CPU cores, but OpenGL is still working on one CPU thread + GPU. The main issue is that you have to move somehow the generated images from OpenGL's memory to a PyTorch Tensor. The simplest way is OpenGL (GPU mem) -&gt; CPU (RAM) -&gt; PyTorch/CUDA (GPU mem). Unfortunately, this data transfer is too slow and it kills the whole idea.</p>",
          "rawMarkdown": "@morizin \nSorry for the confusion. That table shows stats about one image generated on CPU (one core) vs GPU. When you want to generate a batch of images, you can use multiple CPU cores, but OpenGL is still working on one CPU thread + GPU. The main issue is that you have to move somehow the generated images from OpenGL's memory to a PyTorch Tensor. The simplest way is OpenGL (GPU mem) -> CPU (RAM) -> PyTorch/CUDA (GPU mem). Unfortunately, this data transfer is too slow and it kills the whole idea.\n\n\n",
          "votes": 1
        },
        {
          "id": 1061707,
          "postDate": "2020-10-27T08:34:12.003Z",
          "content": "<ul>\n<li>Have you tried OpenCV Cuda implementation or even OpenCV Cuda is also 1 thread.</li>\n<li>I cloned your fork of l5kit. Is that is our implementation. want to make some more modification of my own in them to faster them up</li>\n<li>Do you know where mixed-precision happening in the code</li>\n</ul>\n<blockquote>\n  <p>I looked for PyTorch 2D drawing, too; the same result. 3D tools did not come to my mind. I'll take a look at both. Thank you for the links.</p>\n</blockquote>\n<ul>\n<li><p>Have you done this? did they improved</p></li>\n<li><p>I ran this <a href=\"https://www.kaggle.com/nmygle/pytorch3d-rendering\" target=\"_blank\">notebook</a> I felt a improvement of  7.5%</p></li>\n</ul>",
          "rawMarkdown": "- Have you tried OpenCV Cuda implementation or even OpenCV Cuda is also 1 thread.\n- I cloned your fork of l5kit. Is that is our implementation. want to make some more modification of my own in them to faster them up\n- Do you know where mixed-precision happening in the code\n\n> I looked for PyTorch 2D drawing, too; the same result. 3D tools did not come to my mind. I'll take a look at both. Thank you for the links.\n\n- Have you done this? did they improved\n\n- I ran this [notebook](https://www.kaggle.com/nmygle/pytorch3d-rendering) I felt a improvement of  7.5%"
        },
        {
          "id": 1061716,
          "postDate": "2020-10-27T08:50:17.677Z",
          "content": "<p>I haven't tried it. I am not sure they have implemented the drawing API on GPU/Cuda. Maybe in the c++ version, I don't know.</p>",
          "rawMarkdown": "I haven't tried it. I am not sure they have implemented the drawing API on GPU/Cuda. Maybe in the c++ version, I don't know.",
          "votes": 1
        },
        {
          "id": 1061724,
          "postDate": "2020-10-27T09:03:40.413Z",
          "content": "<p>Thank for your multiple replies<br>\nI am trying recoding slow functions in C++<br>\nI have connected you in LinkedIn</p>\n<blockquote>\n  <table>\n  <thead>\n  <tr>\n  <th>Method</th>\n  <th>Size</th>\n  <th>History</th>\n  <th># of samples</th>\n  <th>Storage size</th>\n  <th>Training time</th>\n  <th>it/sec</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>Rasterizer</td>\n  <td>300</td>\n  <td>10</td>\n  <td>32000</td>\n  <td>0</td>\n  <td>4:48</td>\n  <td>3.47</td>\n  </tr>\n  <tr>\n  <td>Rasterizer</td>\n  <td>300</td>\n  <td>5</td>\n  <td>32000</td>\n  <td>0</td>\n  <td>3:46</td>\n  <td>4.42</td>\n  </tr>\n  <tr>\n  <td>Pre-generated data</td>\n  <td>300</td>\n  <td>10</td>\n  <td>32000</td>\n  <td>1.9 Gb</td>\n  <td>3:26</td>\n  <td>4.85</td>\n  </tr>\n  <tr>\n  <td>Pre-generated data</td>\n  <td>300</td>\n  <td>5</td>\n  <td>32000</td>\n  <td>1.8 Gb</td>\n  <td>2:56</td>\n  <td>5.68</td>\n  </tr>\n  </tbody>\n  </table>\n</blockquote>\n<p>When I choosed 16 batch and 224 rastersize it takes 32 min for 16*2000 iters<br>\nand you said only 4:48min Can you say how you did it?<br>\nhow many channels you inputted into network. no. of history frames?</p>\n<p>I checked out <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this</a> but got an error of <code>pbz file not found</code> to write</p>",
          "rawMarkdown": "Thank for your multiple replies\nI am trying recoding slow functions in C++\nI have connected you in LinkedIn\n\n> | Method             | Size | History | \\# of samples | Storage size | Training time | it/sec |\n> | ------------------ | ---- | ------- | ------------- | ------------ | ------------- | ------ |\n> | Rasterizer         | 300  | 10      | 32000         | 0            | 4:48          | 3.47   |\n> | Rasterizer         | 300  | 5       | 32000         | 0            | 3:46          | 4.42   |\n> | Pre-generated data | 300  | 10      | 32000         | 1.9 Gb       | 3:26          | 4.85   |\n> | Pre-generated data | 300  | 5       | 32000         | 1.8 Gb       | 2:56          | 5.68   |\n\nWhen I choosed 16 batch and 224 rastersize it takes 32 min for 16*2000 iters\nand you said only 4:48min Can you say how you did it?\nhow many channels you inputted into network. no. of history frames?\n\nI checked out [this](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637) but got an error of `pbz file not found` to write"
        },
        {
          "id": 1061772,
          "postDate": "2020-10-27T10:12:41.403Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2Fd867cf8f6a2501c63f994be4dbf879fd%2FIMG_20201027_154829.jpg?generation=1603793951042126&amp;alt=media\" alt=\"\">But even OpenGL is slow it takes only 3 days for 22million samples( over count in pic)<br>\nI have error in your LyftOpenGLRasterizer. Is the rasterized image is same as py semantic<br>\nCould you please show the rasterized image<br>\nCould we generate pbz files from this</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2Fd867cf8f6a2501c63f994be4dbf879fd%2FIMG_20201027_154829.jpg?generation=1603793951042126&alt=media)But even OpenGL is slow it takes only 3 days for 22million samples( over count in pic)\nI have error in your LyftOpenGLRasterizer. Is the rasterized image is same as py semantic\nCould you please show the rasterized image\nCould we generate pbz files from this"
        },
        {
          "id": 1062390,
          "postDate": "2020-10-27T19:28:37.713Z",
          "content": "<p>I can't remember the stats, but I did not finish the OpenGL rasterizer. My tests did not have historical frames or traffic lights, etc. </p>",
          "rawMarkdown": "I can't remember the stats, but I did not finish the OpenGL rasterizer. My tests did not have historical frames or traffic lights, etc. "
        },
        {
          "id": 1062683,
          "postDate": "2020-10-28T05:11:00.640Z",
          "content": "<p>I believe there is actually a way to have OpenGL render to texture and then have CUDA use the texture and thus skip the CPU transfer.  (Here's an example of the concept, I'm not sure if it still works with latest CUDA etc but I know it was possible with older CUDA <a href=\"https://stackoverflow.com/questions/19244191/cuda-opengl-interop-draw-to-opengl-texture-with-cuda\" target=\"_blank\">https://stackoverflow.com/questions/19244191/cuda-opengl-interop-draw-to-opengl-texture-with-cuda</a> ).  </p>\n<p>One nice feature of this OpenGL implementation is that you could use XvFB / llvmpipe / Mesa software rasterizer to do CPU rendering in the case you don't have a GPU.  I know that's not the focus on this work, but if you happened to have GCloud credits, you could use a 96-core instance and run 96 rendering threads in parallel.  It's also handy to use in cases where you need CI and CI doesn't have GPUs.  So OpenGL impl has some nice value.  </p>\n<p>Looking at the code, I'm probably missing something but I don't quite see the showstopper.  In any case this is a really nice effort!!  Comparable libraries in differential rendering, like pytorch3d or DIRT ( <a href=\"https://github.com/pmh47/dirt\" target=\"_blank\">https://github.com/pmh47/dirt</a> ) are still kinda slow, especially versus OpenGL for simple stuff.  There's actually probably a lot of value in trying to pick through these perf bottlenecks…  There's also Kaolin but that project looks a bit hectic right now :S</p>",
          "rawMarkdown": "I believe there is actually a way to have OpenGL render to texture and then have CUDA use the texture and thus skip the CPU transfer.  (Here's an example of the concept, I'm not sure if it still works with latest CUDA etc but I know it was possible with older CUDA https://stackoverflow.com/questions/19244191/cuda-opengl-interop-draw-to-opengl-texture-with-cuda ).  \n\nOne nice feature of this OpenGL implementation is that you could use XvFB / llvmpipe / Mesa software rasterizer to do CPU rendering in the case you don't have a GPU.  I know that's not the focus on this work, but if you happened to have GCloud credits, you could use a 96-core instance and run 96 rendering threads in parallel.  It's also handy to use in cases where you need CI and CI doesn't have GPUs.  So OpenGL impl has some nice value.  \n\nLooking at the code, I'm probably missing something but I don't quite see the showstopper.  In any case this is a really nice effort!!  Comparable libraries in differential rendering, like pytorch3d or DIRT ( https://github.com/pmh47/dirt ) are still kinda slow, especially versus OpenGL for simple stuff.  There's actually probably a lot of value in trying to pick through these perf bottlenecks...  There's also Kaolin but that project looks a bit hectic right now :S"
        },
        {
          "id": 1062794,
          "postDate": "2020-10-28T07:27:38.510Z",
          "content": "<p>is we able to get 96 core CPU + GPU in $300 GCP Credits</p>",
          "rawMarkdown": "is we able to get 96 core CPU + GPU in $300 GCP Credits"
        },
        {
          "id": 1062811,
          "postDate": "2020-10-28T08:05:37.437Z",
          "content": "<p>you can get the quota for sure… I think 96 cores is about $5/hr not including GPU.  so you'd have to be pretty efficient with your time…   aside: in the original pytorch + TPU demo, google wanted you to use a 96-core box because the TPU doesn't support a lot of ops you need for I/O</p>",
          "rawMarkdown": "you can get the quota for sure... I think 96 cores is about $5/hr not including GPU.  so you'd have to be pretty efficient with your time...   aside: in the original pytorch + TPU demo, google wanted you to use a 96-core box because the TPU doesn't support a lot of ops you need for I/O"
        },
        {
          "id": 1062815,
          "postDate": "2020-10-28T08:08:54.877Z",
          "content": "<p>What is 1vCPU = how many cores</p>",
          "rawMarkdown": "What is 1vCPU = how many cores"
        },
        {
          "id": 1069666,
          "postDate": "2020-11-04T18:23:44.713Z",
          "content": "<p><a href=\"https://www.kaggle.com/oarphme\" target=\"_blank\">@oarphme</a> I have a request to you. will you help me fix this issue which I faced during while increasing GCP quota? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2F0347f31470f6f821c5e149dffcdf717a%2FScreenshot%202020-11-02%20210247.jpg?generation=1604514217660408&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@oarphme I have a request to you. will you help me fix this issue which I faced during while increasing GCP quota? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2F0347f31470f6f821c5e149dffcdf717a%2FScreenshot%202020-11-02%20210247.jpg?generation=1604514217660408&alt=media)"
        }
      ]
    },
    {
      "id": 998693,
      "postDate": "2020-09-05T01:11:41.297Z",
      "content": "<p>Wow, great work Peter! This looks very promising.</p>",
      "rawMarkdown": "Wow, great work Peter! This looks very promising."
    },
    {
      "id": 1047727,
      "postDate": "2020-10-12T21:59:19.500Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1003526,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2020-09-09T04:08:19.133000",
      "content": "<p><strong>Update</strong><br>\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).</p>\n<p>You can find more details and the source code <a href=\"https://github.com/pestipeti/LyftOpenGLRasterizer\" target=\"_blank\">on my github</a>.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1012015,
          "author_name": "RossWightman",
          "author_url": "",
          "post_date": "2020-09-15T21:05:03.117000",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> Thanks for sharing your progress/work. Given the difficulties going back/forth between CPU GPU, GL/PyTorch wonder if you've looked at fully Torch tensor based options?</p>\n<p>After seeing how bad the bottleneck is, I looked for some native PyTorch 2D drawing/rasterization but didn't find anything. Kornia has some discussion in issue tracker but no implemented functionality. I then looked for 3d options and ran into two libs I'd looked at before but never tried.</p>\n<p>Both NVIDIA Kaolin (<a href=\"https://github.com/NVIDIAGameWorks/kaolin\" target=\"_blank\">https://github.com/NVIDIAGameWorks/kaolin</a>) and Facebook PyTorch3d (<a href=\"https://github.com/facebookresearch/pytorch3d\" target=\"_blank\">https://github.com/facebookresearch/pytorch3d</a>) have 3d rasterizers that may have enough functionality. They are full differentiable, which is not needed here, but does mean they should work with PyTorch cuda tensors.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1012654,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-16T07:58:23.917000",
          "content": "<p>I looked for PyTorch 2D drawing, too; the same result. 3D tools did not come to my mind. I'll take a look at both. Thank you for the links.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1021480,
          "author_name": "nmygle",
          "author_url": "",
          "post_date": "2020-09-21T22:11:19.987000",
          "content": "<p>I received impressions on your code, so I've written a rendering process in PyTorch3d. But I have yet to see any speed improvements. I'm rendering in Torch on the GPU, can you make this code better?<br>\n<a href=\"https://www.kaggle.com/nmygle/pytorch3d-rendering\" target=\"_blank\">https://www.kaggle.com/nmygle/pytorch3d-rendering</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1003062,
      "author_name": "Luca Bergamini",
      "author_url": "",
      "post_date": "2020-09-08T16:29:05.083000",
      "content": "<p>this is <strong>amazing</strong> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>! I can't help you on the openGL side, but If you start a PR in L5Kit I can surely assist with integration. The rasterisation in Python is a huge bottleneck, so everything that can help on that side is more than welcomed!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1003075,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-08T16:36:57.913000",
          "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> </p>\n<p>Unfortunately, it won't work. The problem is the same that you mentioned on the github discussion. The data transfer between OpenGL gpu mem -&gt; ram -&gt; pytorch gpu tensor kills the entire idea. OpenGL is extremely fast until I want to \"download\" the data. There is a (possible) solution in C++, but that won't help us in python.</p>\n<p>I'll publish the code later, maybe someone will come up with a solution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1003086,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-08T16:45:12.670000",
          "content": "<p>Yeah, that's problematic. A possible solution would be a C++ implementation wrapped in Python (that alone would probably speed things up quite a lot), but sadly this is again outside my expertise :( </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1002197,
      "author_name": "n3n7i",
      "author_url": "",
      "post_date": "2020-09-07T23:02:17.570000",
      "content": "<p>Did you try a context.save / context.restore type method for the background data? - it would probably just be the lane info, but you should be able to reuse for some number of frames before drawing the next</p>\n<p>Added: there's a stray assert in an is_lane()?, but also the map api cache is set quite low at 90kb?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1002362,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-08T04:38:49.490000",
          "content": "<p>I stored the semantic map in vbo/vao (only ~30mb gpu mem). It is fast, no need for context. The real problem is moving the image/pixels from opengl -&gt; cpu -&gt; pytorch cuda tensor. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1002442,
          "author_name": "n3n7i",
          "author_url": "",
          "post_date": "2020-09-08T06:08:32.153000",
          "content": "<p>Maybe passing a shared memory reference would work, but it's just a guess</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1002480,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-08T06:56:40.807000",
          "content": "<p>Unfortunately,  I don't think we have that level of control over opengl in python (or I don't know how to do it). I'll share the code later today or tomorrow, maybe someone will come up with a solution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1003395,
          "author_name": "n3n7i",
          "author_url": "",
          "post_date": "2020-09-08T23:26:28.063000",
          "content": "<p>Oldish article with c++ <a href=\"https://www.3dgep.com/opengl-interoperability-with-cuda/\" target=\"_blank\">https://www.3dgep.com/opengl-interoperability-with-cuda/</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3300065,
      "author_name": "daya.s",
      "author_url": "",
      "post_date": "2025-10-09T12:24:40.057000",
      "content": "<p>Thanks for the deep dive into OpenGL rasterization! That 10x speed boost is huge, especially since CPU rasterization is the main slow point.</p>\n<p>Your benchmark of 130 it/sec vs 13 it/sec is impressive, even without history rendering. It makes sense that history generation on GPU is tricky due to how l5kit works.</p>\n<p>Looks like data loading is now the bottleneck, showing how much faster GPU rasterization is. Pre-caching images might still be easier for training, but your method could be great for inference or live use.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1061559,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-27T04:53:19.547000",
      "content": "<ul>\n<li>Is this notebook works on new L5kit?</li>\n<li>Do they have any improvement?</li>\n<li>Your table show it has but on the same time you are saying OpenGl GPU -&gt; CPU -&gt; Pytorch Tensor. Can you please clarify this?</li>\n</ul>\n<blockquote>\n  <p>I haven't implemented the crosswalks and the traffic lights yet.&lt;</p>\n</blockquote>\n<ul>\n<li>In your Github, Have you included them in the Github</li>\n<li>Will it improves if we improve all the function as a torch function?</li>\n<li>or by <code>torch.multiprocessing</code> will that works?</li>\n<li>What time it takes to complete the <code>train.zarr</code> file with your fastest method you drew for this competition to become 8th place in Public. If you don't mind Can you share how you done them?</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 1061616,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-10-27T06:24:06.117000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> </p>\n<ul>\n<li>I did not try with the new version of l5kit</li>\n<li>No improvements. I gave up on this idea.</li>\n<li>The problem is the same. OpenGL is fast, but copying the data from opengl's memory to a tensor is slow</li>\n<li>No improvements, no crosswalks/traffic lights/no history</li>\n<li>If you can find a solution on how to pass memory reference between opengl and torch tensor it would be usable</li>\n<li>no, opengl is single-threaded. </li>\n<li>for 8th place, I used l5kit v1.0.6 and I optimized it to run faster. The training for 8th place ran ~8-9days</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1061645,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-27T06:59:36.350000",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> </p>\n<ul>\n<li>I tried you Github, I had an error, but don't care.</li>\n<li>But you project out a table in the discussion that it's faster. I don't get you much right now</li>\n<li>I might think making all the function as Cython function will sometimes works don't know.</li>\n<li>will making all the functions as numba function will work (still don't know)</li>\n<li>I maybe will go on the way to recode all the code to C++ (if trying I want to dedicate 5 days for it)</li>\n<li>But don't know how <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> in this discussion made his training to 2.5 days</li>\n<li>when trained its 22 days for train.zarr (without optimizing l5kit)</li>\n<li>This is Open topic for this competition</li>\n</ul>\n<h5>Confused All about</h5>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1061666,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-10-27T07:31:17.577000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> <br>\nSorry for the confusion. That table shows stats about one image generated on CPU (one core) vs GPU. When you want to generate a batch of images, you can use multiple CPU cores, but OpenGL is still working on one CPU thread + GPU. The main issue is that you have to move somehow the generated images from OpenGL's memory to a PyTorch Tensor. The simplest way is OpenGL (GPU mem) -&gt; CPU (RAM) -&gt; PyTorch/CUDA (GPU mem). Unfortunately, this data transfer is too slow and it kills the whole idea.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1061707,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-27T08:34:12.003000",
          "content": "<ul>\n<li>Have you tried OpenCV Cuda implementation or even OpenCV Cuda is also 1 thread.</li>\n<li>I cloned your fork of l5kit. Is that is our implementation. want to make some more modification of my own in them to faster them up</li>\n<li>Do you know where mixed-precision happening in the code</li>\n</ul>\n<blockquote>\n  <p>I looked for PyTorch 2D drawing, too; the same result. 3D tools did not come to my mind. I'll take a look at both. Thank you for the links.</p>\n</blockquote>\n<ul>\n<li><p>Have you done this? did they improved</p></li>\n<li><p>I ran this <a href=\"https://www.kaggle.com/nmygle/pytorch3d-rendering\" target=\"_blank\">notebook</a> I felt a improvement of  7.5%</p></li>\n</ul>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1061716,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-10-27T08:50:17.677000",
          "content": "<p>I haven't tried it. I am not sure they have implemented the drawing API on GPU/Cuda. Maybe in the c++ version, I don't know.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1061724,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-27T09:03:40.413000",
          "content": "<p>Thank for your multiple replies<br>\nI am trying recoding slow functions in C++<br>\nI have connected you in LinkedIn</p>\n<blockquote>\n  <table>\n  <thead>\n  <tr>\n  <th>Method</th>\n  <th>Size</th>\n  <th>History</th>\n  <th># of samples</th>\n  <th>Storage size</th>\n  <th>Training time</th>\n  <th>it/sec</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>Rasterizer</td>\n  <td>300</td>\n  <td>10</td>\n  <td>32000</td>\n  <td>0</td>\n  <td>4:48</td>\n  <td>3.47</td>\n  </tr>\n  <tr>\n  <td>Rasterizer</td>\n  <td>300</td>\n  <td>5</td>\n  <td>32000</td>\n  <td>0</td>\n  <td>3:46</td>\n  <td>4.42</td>\n  </tr>\n  <tr>\n  <td>Pre-generated data</td>\n  <td>300</td>\n  <td>10</td>\n  <td>32000</td>\n  <td>1.9 Gb</td>\n  <td>3:26</td>\n  <td>4.85</td>\n  </tr>\n  <tr>\n  <td>Pre-generated data</td>\n  <td>300</td>\n  <td>5</td>\n  <td>32000</td>\n  <td>1.8 Gb</td>\n  <td>2:56</td>\n  <td>5.68</td>\n  </tr>\n  </tbody>\n  </table>\n</blockquote>\n<p>When I choosed 16 batch and 224 rastersize it takes 32 min for 16*2000 iters<br>\nand you said only 4:48min Can you say how you did it?<br>\nhow many channels you inputted into network. no. of history frames?</p>\n<p>I checked out <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this</a> but got an error of <code>pbz file not found</code> to write</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1061772,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-27T10:12:41.403000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2Fd867cf8f6a2501c63f994be4dbf879fd%2FIMG_20201027_154829.jpg?generation=1603793951042126&amp;alt=media\" alt=\"\">But even OpenGL is slow it takes only 3 days for 22million samples( over count in pic)<br>\nI have error in your LyftOpenGLRasterizer. Is the rasterized image is same as py semantic<br>\nCould you please show the rasterized image<br>\nCould we generate pbz files from this</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062390,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-10-27T19:28:37.713000",
          "content": "<p>I can't remember the stats, but I did not finish the OpenGL rasterizer. My tests did not have historical frames or traffic lights, etc. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062683,
          "author_name": "oarph",
          "author_url": "",
          "post_date": "2020-10-28T05:11:00.640000",
          "content": "<p>I believe there is actually a way to have OpenGL render to texture and then have CUDA use the texture and thus skip the CPU transfer.  (Here's an example of the concept, I'm not sure if it still works with latest CUDA etc but I know it was possible with older CUDA <a href=\"https://stackoverflow.com/questions/19244191/cuda-opengl-interop-draw-to-opengl-texture-with-cuda\" target=\"_blank\">https://stackoverflow.com/questions/19244191/cuda-opengl-interop-draw-to-opengl-texture-with-cuda</a> ).  </p>\n<p>One nice feature of this OpenGL implementation is that you could use XvFB / llvmpipe / Mesa software rasterizer to do CPU rendering in the case you don't have a GPU.  I know that's not the focus on this work, but if you happened to have GCloud credits, you could use a 96-core instance and run 96 rendering threads in parallel.  It's also handy to use in cases where you need CI and CI doesn't have GPUs.  So OpenGL impl has some nice value.  </p>\n<p>Looking at the code, I'm probably missing something but I don't quite see the showstopper.  In any case this is a really nice effort!!  Comparable libraries in differential rendering, like pytorch3d or DIRT ( <a href=\"https://github.com/pmh47/dirt\" target=\"_blank\">https://github.com/pmh47/dirt</a> ) are still kinda slow, especially versus OpenGL for simple stuff.  There's actually probably a lot of value in trying to pick through these perf bottlenecks…  There's also Kaolin but that project looks a bit hectic right now :S</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062794,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-28T07:27:38.510000",
          "content": "<p>is we able to get 96 core CPU + GPU in $300 GCP Credits</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062811,
          "author_name": "oarph",
          "author_url": "",
          "post_date": "2020-10-28T08:05:37.437000",
          "content": "<p>you can get the quota for sure… I think 96 cores is about $5/hr not including GPU.  so you'd have to be pretty efficient with your time…   aside: in the original pytorch + TPU demo, google wanted you to use a 96-core box because the TPU doesn't support a lot of ops you need for I/O</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062815,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-28T08:08:54.877000",
          "content": "<p>What is 1vCPU = how many cores</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1069666,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-11-04T18:23:44.713000",
          "content": "<p><a href=\"https://www.kaggle.com/oarphme\" target=\"_blank\">@oarphme</a> I have a request to you. will you help me fix this issue which I faced during while increasing GCP quota? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2F0347f31470f6f821c5e149dffcdf717a%2FScreenshot%202020-11-02%20210247.jpg?generation=1604514217660408&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 998693,
      "author_name": "Khushal B",
      "author_url": "",
      "post_date": "2020-09-05T01:11:41.297000",
      "content": "<p>Wow, great work Peter! This looks very promising.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1047727,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-12T21:59:19.500000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "998423": "As you probably already know, the rasterization process is the bottleneck in this competition. At first, I tried to optimize the l5kit, but I only achieved minor speedup. With pre-generated images (and the optimized l5kit), I achieved ~ 4-5x  it/sec. Unfortunately, the GPU utilization was still ~40-45% (I have an RTX-2080ti I used 10 cores; 20 CPU threads). The obvious solution is to render the images on the GPU. As @lucabergamini mentioned on their [github discussion](https://github.com/lyft/l5kit/issues/136#issuecomment-686370259) that it is not an easy task with the current implementation.\n\nI implemented an OpenGLRasterizer, and it looks very promising. Here is my baseline benchmark.\n\n## Results\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F85e632b2df7be4c66926a84160078834%2Frasterizer_benchmark.png?generation=1599242281108744&alt=media)\n\n## Experiments:\n| Method             | Size | History | \\# of samples | Running time | it/sec |\n| --------------- | ---- | ------- | ------------ | ------------- | ------ |\n| CPU rasterizer | 650px | 0           | 1,000           | 1:14                | 13.41   |\n| OpenGL GPU   | 650px | 0           | 10,000         | 1:16                | 130.49   |\n\n**Notes**\n- I haven't implemented the crosswalks and the traffic lights yet. (CPU rasterizer generates both)\n- It is hard to generate the history (in the current output form) on the GPU (at least need some CPU post-process). I am looking for alternatives.\n- l5kit optimization is in progress (see [their github](https://github.com/lyft/l5kit); both methods will improve.\n- Data loading is still a bottleneck. OpenGL can generate at a ~800 it/sec rate if I use the same preloaded frame's data. (I used a busy frame with lots of agents)\n- I used 1 CPU core for both methods.\n- GPU memory usage is low: ~35Mb.\n\nI have to solve/fix a few things before I can publish the code. I don't want to make any promises because I am busy with other projects, but I expect to be ready next week.\n\n--------------\n\n**Update**\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).\n\nYou can find more details and the source code [on my github](https://github.com/pestipeti/LyftOpenGLRasterizer).\n",
    "1003526": "**Update**\nI was a bit optimistic when I posted this thread. Unfortunately, my idea is not working. It could give some improvement if you have only 1-2 CPU cores, or you have multiple GPUs (I haven't tested).\n\nYou can find more details and the source code [on my github](https://github.com/pestipeti/LyftOpenGLRasterizer).\n\n",
    "1003062": "this is **amazing** @pestipeti! I can't help you on the openGL side, but If you start a PR in L5Kit I can surely assist with integration. The rasterisation in Python is a huge bottleneck, so everything that can help on that side is more than welcomed!",
    "1002197": "Did you try a context.save / context.restore type method for the background data? - it would probably just be the lane info, but you should be able to reuse for some number of frames before drawing the next\n\nAdded: there's a stray assert in an is_lane()?, but also the map api cache is set quite low at 90kb?",
    "3300065": "Thanks for the deep dive into OpenGL rasterization! That 10x speed boost is huge, especially since CPU rasterization is the main slow point.\n\nYour benchmark of 130 it/sec vs 13 it/sec is impressive, even without history rendering. It makes sense that history generation on GPU is tricky due to how l5kit works.\n\nLooks like data loading is now the bottleneck, showing how much faster GPU rasterization is. Pre-caching images might still be easier for training, but your method could be great for inference or live use.",
    "1061559": "- Is this notebook works on new L5kit?\n- Do they have any improvement?\n- Your table show it has but on the same time you are saying OpenGl GPU -> CPU -> Pytorch Tensor. Can you please clarify this?\n> I haven't implemented the crosswalks and the traffic lights yet.<\n\n\n- In your Github, Have you included them in the Github\n- Will it improves if we improve all the function as a torch function?\n- or by `torch.multiprocessing` will that works?\n- What time it takes to complete the `train.zarr` file with your fastest method you drew for this competition to become 8th place in Public. If you don't mind Can you share how you done them?",
    "998693": "Wow, great work Peter! This looks very promising.",
    "1047727": ""
  }
}