{
  "id": 199244,
  "title": "Faster draw_boxes function for l5kit and some comment on the l5kit package",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199244",
  "author_name": "",
  "post_date": "2020-11-25T01:42:23.976310Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Sorry, this is probably too late to share but I just found this yesterday since I am pretty late to the competition. <br>\nWe know that the bottleneck of the competition is the speed that l5kit generates the training set. And, the most computation intense part is where it try to the render image of all the agents in each frames in the <code>box_rasterizer.py</code>. One of the reason why it so slow is because the l5kit code is not vectorize but instead simply use a lot of python <code>for</code> loop to iterate through each agent in each frames. So for each sample in the <code>AgentDataset</code>, it loop through (11 frames) * (N agents in each frame). If one can vectorize this part, it will make thing faster. Here I found the loop in the <code>draw_boxes</code> (link to <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/rasterization/box_rasterizer.py#L33-L72\" target=\"_blank\">original code</a>) can be vectorize as</p>\n<pre><code>_corners_base_coords = (np.asarray([[-1, -1], [-1, 1], [1, 1], [1, -1]]) / 2)[None, :, :]\n\ndef draw_boxes(\n    raster_size: Tuple[int, int],\n    raster_from_world: np.ndarray,\n    agents: np.ndarray,\n    color: Union[int, Tuple[int, int, int]],\n) -&gt; np.ndarray:\n    if isinstance(color, int):\n        im = np.zeros((raster_size[1], raster_size[0]), dtype=np.uint8)\n    else:\n        im = np.zeros((raster_size[1], raster_size[0], 3), dtype=np.uint8)\n\n    corners_m = _corners_base_coords * agents[\"extent\"][:, None, :2]  # corners in zero\n    s = np.sin(agents['yaw'])\n    c = np.cos(agents['yaw'])\n    rotation_m = np.moveaxis(np.array([[c, -s], [s, c]]), 2, 0)\n    box_world_coords = np.einsum('bti,bji-&gt;btj', corners_m, rotation_m) + agents['centroid'][:, None, :2]\n    box_raster_coords = transform_points(box_world_coords.reshape((-1, 2)), raster_from_world)\n\n    # fillPoly wants polys in a sequence with points inside as (x,y)\n    box_raster_coords = cv2_subpixel(box_raster_coords.reshape((-1, 4, 2)))\n    cv2.fillPoly(im, box_raster_coords, color=color, lineType=cv2.LINE_AA, shift=CV2_SHIFT)\n    return im\n</code></pre>\n<p>This makes the rendering about 40% faster on my machine. There are also several other places where one can vectorize the computation instead of using a slow python <code>for</code> loop. Though the original code has better readability, and that is probably why lyft didn't write in this way.</p>\n<p>Also, note that in several cases, a simple comprehension <code>a = np.array([f(i) for i in range(n)])</code> in python is still much (2x - 10x) faster than a for loop</p>\n<pre><code>a = np.zeros(n)\nfor i in range(n):\n    a[i] = f(i)\n</code></pre>\n<p>This is because the way python optimize the comprehension in C. But it can't do it for a general for loop. So perhaps, one can consider adding this.</p>\n<p>Any comments are welcome. Or, you have other tricks better than this. We probably won't have time to test them though ;)</p>",
  "messages": [
    {
      "id": "1090017",
      "postDate": "11/25/2020 01:42:23",
      "content": "<p>Sorry, this is probably too late to share but I just found this yesterday since I am pretty late to the competition. <br>\nWe know that the bottleneck of the competition is the speed that l5kit generates the training set. And, the most computation intense part is where it try to the render image of all the agents in each frames in the <code>box_rasterizer.py</code>. One of the reason why it so slow is because the l5kit code is not vectorize but instead simply use a lot of python <code>for</code> loop to iterate through each agent in each frames. So for each sample in the <code>AgentDataset</code>, it loop through (11 frames) * (N agents in each frame). If one can vectorize this part, it will make thing faster. Here I found the loop in the <code>draw_boxes</code> (link to <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/rasterization/box_rasterizer.py#L33-L72\" target=\"_blank\">original code</a>) can be vectorize as</p>\n<pre><code>_corners_base_coords = (np.asarray([[-1, -1], [-1, 1], [1, 1], [1, -1]]) / 2)[None, :, :]\n\ndef draw_boxes(\n    raster_size: Tuple[int, int],\n    raster_from_world: np.ndarray,\n    agents: np.ndarray,\n    color: Union[int, Tuple[int, int, int]],\n) -&gt; np.ndarray:\n    if isinstance(color, int):\n        im = np.zeros((raster_size[1], raster_size[0]), dtype=np.uint8)\n    else:\n        im = np.zeros((raster_size[1], raster_size[0], 3), dtype=np.uint8)\n\n    corners_m = _corners_base_coords * agents[\"extent\"][:, None, :2]  # corners in zero\n    s = np.sin(agents['yaw'])\n    c = np.cos(agents['yaw'])\n    rotation_m = np.moveaxis(np.array([[c, -s], [s, c]]), 2, 0)\n    box_world_coords = np.einsum('bti,bji-&gt;btj', corners_m, rotation_m) + agents['centroid'][:, None, :2]\n    box_raster_coords = transform_points(box_world_coords.reshape((-1, 2)), raster_from_world)\n\n    # fillPoly wants polys in a sequence with points inside as (x,y)\n    box_raster_coords = cv2_subpixel(box_raster_coords.reshape((-1, 4, 2)))\n    cv2.fillPoly(im, box_raster_coords, color=color, lineType=cv2.LINE_AA, shift=CV2_SHIFT)\n    return im\n</code></pre>\n<p>This makes the rendering about 40% faster on my machine. There are also several other places where one can vectorize the computation instead of using a slow python <code>for</code> loop. Though the original code has better readability, and that is probably why lyft didn't write in this way.</p>\n<p>Also, note that in several cases, a simple comprehension <code>a = np.array([f(i) for i in range(n)])</code> in python is still much (2x - 10x) faster than a for loop</p>\n<pre><code>a = np.zeros(n)\nfor i in range(n):\n    a[i] = f(i)\n</code></pre>\n<p>This is because the way python optimize the comprehension in C. But it can't do it for a general for loop. So perhaps, one can consider adding this.</p>\n<p>Any comments are welcome. Or, you have other tricks better than this. We probably won't have time to test them though ;)</p>",
      "rawMarkdown": "Sorry, this is probably too late to share but I just found this yesterday since I am pretty late to the competition. \nWe know that the bottleneck of the competition is the speed that l5kit generates the training set. And, the most computation intense part is where it try to the render image of all the agents in each frames in the `box_rasterizer.py`. One of the reason why it so slow is because the l5kit code is not vectorize but instead simply use a lot of python `for` loop to iterate through each agent in each frames. So for each sample in the `AgentDataset`, it loop through (11 frames) * (N agents in each frame). If one can vectorize this part, it will make thing faster. Here I found the loop in the `draw_boxes` (link to [original code](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/rasterization/box_rasterizer.py#L33-L72)) can be vectorize as\n```\n_corners_base_coords = (np.asarray([[-1, -1], [-1, 1], [1, 1], [1, -1]]) / 2)[None, :, :]\n\ndef draw_boxes(\n    raster_size: Tuple[int, int],\n    raster_from_world: np.ndarray,\n    agents: np.ndarray,\n    color: Union[int, Tuple[int, int, int]],\n) -> np.ndarray:\n    if isinstance(color, int):\n        im = np.zeros((raster_size[1], raster_size[0]), dtype=np.uint8)\n    else:\n        im = np.zeros((raster_size[1], raster_size[0], 3), dtype=np.uint8)\n\n    corners_m = _corners_base_coords * agents[\"extent\"][:, None, :2]  # corners in zero\n    s = np.sin(agents['yaw'])\n    c = np.cos(agents['yaw'])\n    rotation_m = np.moveaxis(np.array([[c, -s], [s, c]]), 2, 0)\n    box_world_coords = np.einsum('bti,bji->btj', corners_m, rotation_m) + agents['centroid'][:, None, :2]\n    box_raster_coords = transform_points(box_world_coords.reshape((-1, 2)), raster_from_world)\n\n    # fillPoly wants polys in a sequence with points inside as (x,y)\n    box_raster_coords = cv2_subpixel(box_raster_coords.reshape((-1, 4, 2)))\n    cv2.fillPoly(im, box_raster_coords, color=color, lineType=cv2.LINE_AA, shift=CV2_SHIFT)\n    return im\n```\nThis makes the rendering about 40% faster on my machine. There are also several other places where one can vectorize the computation instead of using a slow python `for` loop. Though the original code has better readability, and that is probably why lyft didn't write in this way.\n\nAlso, note that in several cases, a simple comprehension `a = np.array([f(i) for i in range(n)])` in python is still much (2x - 10x) faster than a for loop\n```\na = np.zeros(n)\nfor i in range(n):\n    a[i] = f(i)\n```\nThis is because the way python optimize the comprehension in C. But it can't do it for a general for loop. So perhaps, one can consider adding this.\n\nAny comments are welcome. Or, you have other tricks better than this. We probably won't have time to test them though ;)",
      "votes": null
    },
    {
      "id": "1090258",
      "postDate": "11/25/2020 07:51:06",
      "content": "<p>I am surprised to hear that list comprehensions are actually any faster than just standard for loops but looks like that is truly the case in some scenarios. </p>\n<p><a href=\"https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list\" target=\"_blank\">https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list</a></p>",
      "rawMarkdown": "I am surprised to hear that list comprehensions are actually any faster than just standard for loops but looks like that is truly the case in some scenarios. \n\n[https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list](https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list)",
      "votes": null
    },
    {
      "id": "1090375",
      "postDate": "11/25/2020 09:47:49",
      "content": "<p>Thank you. I hope this is enough to allow me to get a sub in, lol.</p>",
      "rawMarkdown": "Thank you. I hope this is enough to allow me to get a sub in, lol.",
      "votes": null
    },
    {
      "id": "1090425",
      "postDate": "11/25/2020 10:19:10",
      "content": "<blockquote>\n  <p>Though the original code has better readability, and that is probably why lyft didn't write in this way.</p>\n</blockquote>\n<p>I think your implementation is quite clear tbh.  Einstein summation convention is not something I have in my background, so I was not aware you could implement it this way :)</p>\n<p>Do you think you can start a PR with this in L5Kit?</p>",
      "rawMarkdown": "> Though the original code has better readability, and that is probably why lyft didn't write in this way.\n\nI think your implementation is quite clear tbh.  Einstein summation convention is not something I have in my background, so I was not aware you could implement it this way :)\n\nDo you think you can start a PR with this in L5Kit?",
      "votes": null
    },
    {
      "id": "1090955",
      "postDate": "11/25/2020 17:31:25",
      "content": "<p>Sure! Will give it a try!</p>",
      "rawMarkdown": "Sure! Will give it a try!",
      "votes": null
    },
    {
      "id": "1091379",
      "postDate": "11/26/2020 01:10:40",
      "content": "<p>not sure if this would help. there is specialised package for einsum<br>\n(maybe in GPU also?)</p>\n<p><a href=\"https://github.com/dgasmith/opt_einsum\" target=\"_blank\">https://github.com/dgasmith/opt_einsum</a></p>",
      "rawMarkdown": "not sure if this would help. there is specialised package for einsum\n(maybe in GPU also?)\n\nhttps://github.com/dgasmith/opt_einsum",
      "votes": null
    },
    {
      "id": "1093642",
      "postDate": "11/27/2020 21:50:09",
      "content": "<p>The PR is hear <a href=\"https://github.com/lyft/l5kit/pull/209\" target=\"_blank\">https://github.com/lyft/l5kit/pull/209</a></p>",
      "rawMarkdown": "The PR is hear https://github.com/lyft/l5kit/pull/209",
      "votes": null
    },
    {
      "id": "1093645",
      "postDate": "11/27/2020 21:56:51",
      "content": "<p>awesome! 👀</p>",
      "rawMarkdown": "awesome! 👀",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1090258,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/25/2020 07:51:06",
      "content": "<p>I am surprised to hear that list comprehensions are actually any faster than just standard for loops but looks like that is truly the case in some scenarios. </p>\n<p><a href=\"https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list\" target=\"_blank\">https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1090375,
      "author_name": "authman",
      "author_url": "",
      "post_date": "11/25/2020 09:47:49",
      "content": "<p>Thank you. I hope this is enough to allow me to get a sub in, lol.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1090425,
      "author_name": "lucabergamini",
      "author_url": "",
      "post_date": "11/25/2020 10:19:10",
      "content": "<blockquote>\n  <p>Though the original code has better readability, and that is probably why lyft didn't write in this way.</p>\n</blockquote>\n<p>I think your implementation is quite clear tbh.  Einstein summation convention is not something I have in my background, so I was not aware you could implement it this way :)</p>\n<p>Do you think you can start a PR with this in L5Kit?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1090955,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/25/2020 17:31:25",
          "content": "<p>Sure! Will give it a try!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091379,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/26/2020 01:10:40",
          "content": "<p>not sure if this would help. there is specialised package for einsum<br>\n(maybe in GPU also?)</p>\n<p><a href=\"https://github.com/dgasmith/opt_einsum\" target=\"_blank\">https://github.com/dgasmith/opt_einsum</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093642,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/27/2020 21:50:09",
          "content": "<p>The PR is hear <a href=\"https://github.com/lyft/l5kit/pull/209\" target=\"_blank\">https://github.com/lyft/l5kit/pull/209</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093645,
          "author_name": "lucabergamini",
          "author_url": "",
          "post_date": "11/27/2020 21:56:51",
          "content": "<p>awesome! 👀</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1090017": "Sorry, this is probably too late to share but I just found this yesterday since I am pretty late to the competition. \nWe know that the bottleneck of the competition is the speed that l5kit generates the training set. And, the most computation intense part is where it try to the render image of all the agents in each frames in the `box_rasterizer.py`. One of the reason why it so slow is because the l5kit code is not vectorize but instead simply use a lot of python `for` loop to iterate through each agent in each frames. So for each sample in the `AgentDataset`, it loop through (11 frames) * (N agents in each frame). If one can vectorize this part, it will make thing faster. Here I found the loop in the `draw_boxes` (link to [original code](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/rasterization/box_rasterizer.py#L33-L72)) can be vectorize as\n```\n_corners_base_coords = (np.asarray([[-1, -1], [-1, 1], [1, 1], [1, -1]]) / 2)[None, :, :]\n\ndef draw_boxes(\n    raster_size: Tuple[int, int],\n    raster_from_world: np.ndarray,\n    agents: np.ndarray,\n    color: Union[int, Tuple[int, int, int]],\n) -> np.ndarray:\n    if isinstance(color, int):\n        im = np.zeros((raster_size[1], raster_size[0]), dtype=np.uint8)\n    else:\n        im = np.zeros((raster_size[1], raster_size[0], 3), dtype=np.uint8)\n\n    corners_m = _corners_base_coords * agents[\"extent\"][:, None, :2]  # corners in zero\n    s = np.sin(agents['yaw'])\n    c = np.cos(agents['yaw'])\n    rotation_m = np.moveaxis(np.array([[c, -s], [s, c]]), 2, 0)\n    box_world_coords = np.einsum('bti,bji->btj', corners_m, rotation_m) + agents['centroid'][:, None, :2]\n    box_raster_coords = transform_points(box_world_coords.reshape((-1, 2)), raster_from_world)\n\n    # fillPoly wants polys in a sequence with points inside as (x,y)\n    box_raster_coords = cv2_subpixel(box_raster_coords.reshape((-1, 4, 2)))\n    cv2.fillPoly(im, box_raster_coords, color=color, lineType=cv2.LINE_AA, shift=CV2_SHIFT)\n    return im\n```\nThis makes the rendering about 40% faster on my machine. There are also several other places where one can vectorize the computation instead of using a slow python `for` loop. Though the original code has better readability, and that is probably why lyft didn't write in this way.\n\nAlso, note that in several cases, a simple comprehension `a = np.array([f(i) for i in range(n)])` in python is still much (2x - 10x) faster than a for loop\n```\na = np.zeros(n)\nfor i in range(n):\n    a[i] = f(i)\n```\nThis is because the way python optimize the comprehension in C. But it can't do it for a general for loop. So perhaps, one can consider adding this.\n\nAny comments are welcome. Or, you have other tricks better than this. We probably won't have time to test them though ;)",
    "1090258": "I am surprised to hear that list comprehensions are actually any faster than just standard for loops but looks like that is truly the case in some scenarios. \n\n[https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list](https://stackoverflow.com/questions/30245397/why-is-a-list-comprehension-so-much-faster-than-appending-to-a-list)",
    "1090375": "Thank you. I hope this is enough to allow me to get a sub in, lol.",
    "1090425": "> Though the original code has better readability, and that is probably why lyft didn't write in this way.\n\nI think your implementation is quite clear tbh.  Einstein summation convention is not something I have in my background, so I was not aware you could implement it this way :)\n\nDo you think you can start a PR with this in L5Kit?",
    "1090955": "Sure! Will give it a try!",
    "1091379": "not sure if this would help. there is specialised package for einsum\n(maybe in GPU also?)\n\nhttps://github.com/dgasmith/opt_einsum",
    "1093642": "The PR is hear https://github.com/lyft/l5kit/pull/209",
    "1093645": "awesome! 👀"
  },
  "source": "meta"
}