{
  "id": 193858,
  "title": "Redundancy of the `transform_points` in the l5kit package for prediction?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/193858",
  "author_name": "",
  "post_date": "2020-10-29T08:03:36.320241300Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am trying to understand what the final transformation in the prediction was doing. From popular notebook like <a href=\"https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline\" target=\"_blank\">this</a>, we have</p>\n<pre><code>for idx in range(len(preds)):\n    for mode in range(3):\n        preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]\n</code></pre>\n<p>And, when looks into what the <code>transform_points()</code> is doing <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L73-L99\" target=\"_blank\">github link</a>, which can be simplified as (for our 2d points):</p>\n<pre><code>def transform_points(points, transf_matrix)\n    return (points @ transf_matrix[:2, :2].T) + transf_matrix[:2, 2]\n</code></pre>\n<p>One can simplify this entire transformation as simple matrix multiplication:</p>\n<pre><code>preds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n + world_from_agents[idx, :2, 2] - centroids[idx, :2]\n</code></pre>\n<p>The funny thing is that when you check the <code>world_from_agents matrix[:, :2, 2]</code> matrix and <code>centroids[:, :2]</code> matrix. They are exactly the same….</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fff211c6780b02630d460dc0ec38e8787%2Fworld_from_agent.png?generation=1603958261780537&amp;alt=media\" alt=\"world from agents\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F19cd582ada25e21469f524908af1603d%2Fcentroid.png?generation=1603958284550144&amp;alt=media\" alt=\"centroids\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fcd56524370ed3785d8398685f5ae950d%2Fdiff.png?generation=1603958304404184&amp;alt=media\" alt=\"diff\"></p>\n<p>So I guess we can just simplify the transform as </p>\n<pre><code>preds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n</code></pre>\n<p>And, I believe with some numpy trick, you can easily do this for all the idx and mode without the for loop. <br>\nIs my assumption that the last column of <code>world_from _agents</code> is the same as the <code>centroid</code> correct?<br>\nI also noticed that the centroid is in real world coordinate so the value tend to be very large like ~500 compared to the agents local coordinate. So adding then subtracting such a large number can slightly harm the precision of prediction in general.</p>\n<p>Edit:<br>\nI found that you can simplify the last transformation as</p>\n<pre><code>preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n</code></pre>\n<p>I verify this transformation give the same score as the original formula on Kaggle in this notebook  (v12)<br>\n<a href=\"https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline/notebook#Prediction\" target=\"_blank\">https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline/notebook#Prediction</a></p>",
  "messages": [
    {
      "id": "1063691",
      "postDate": "10/29/2020 08:03:36",
      "content": "<p>I am trying to understand what the final transformation in the prediction was doing. From popular notebook like <a href=\"https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline\" target=\"_blank\">this</a>, we have</p>\n<pre><code>for idx in range(len(preds)):\n    for mode in range(3):\n        preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]\n</code></pre>\n<p>And, when looks into what the <code>transform_points()</code> is doing <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L73-L99\" target=\"_blank\">github link</a>, which can be simplified as (for our 2d points):</p>\n<pre><code>def transform_points(points, transf_matrix)\n    return (points @ transf_matrix[:2, :2].T) + transf_matrix[:2, 2]\n</code></pre>\n<p>One can simplify this entire transformation as simple matrix multiplication:</p>\n<pre><code>preds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n + world_from_agents[idx, :2, 2] - centroids[idx, :2]\n</code></pre>\n<p>The funny thing is that when you check the <code>world_from_agents matrix[:, :2, 2]</code> matrix and <code>centroids[:, :2]</code> matrix. They are exactly the same….</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fff211c6780b02630d460dc0ec38e8787%2Fworld_from_agent.png?generation=1603958261780537&amp;alt=media\" alt=\"world from agents\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F19cd582ada25e21469f524908af1603d%2Fcentroid.png?generation=1603958284550144&amp;alt=media\" alt=\"centroids\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fcd56524370ed3785d8398685f5ae950d%2Fdiff.png?generation=1603958304404184&amp;alt=media\" alt=\"diff\"></p>\n<p>So I guess we can just simplify the transform as </p>\n<pre><code>preds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n</code></pre>\n<p>And, I believe with some numpy trick, you can easily do this for all the idx and mode without the for loop. <br>\nIs my assumption that the last column of <code>world_from _agents</code> is the same as the <code>centroid</code> correct?<br>\nI also noticed that the centroid is in real world coordinate so the value tend to be very large like ~500 compared to the agents local coordinate. So adding then subtracting such a large number can slightly harm the precision of prediction in general.</p>\n<p>Edit:<br>\nI found that you can simplify the last transformation as</p>\n<pre><code>preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n</code></pre>\n<p>I verify this transformation give the same score as the original formula on Kaggle in this notebook  (v12)<br>\n<a href=\"https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline/notebook#Prediction\" target=\"_blank\">https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline/notebook#Prediction</a></p>",
      "rawMarkdown": "I am trying to understand what the final transformation in the prediction was doing. From popular notebook like [this](https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline), we have\n```python\nfor idx in range(len(preds)):\n    for mode in range(3):\n        preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]\n```\nAnd, when looks into what the `transform_points()` is doing [github link](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L73-L99), which can be simplified as (for our 2d points):\n```\ndef transform_points(points, transf_matrix)\n    return (points @ transf_matrix[:2, :2].T) + transf_matrix[:2, 2]\n```\nOne can simplify this entire transformation as simple matrix multiplication:\n```\npreds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n + world_from_agents[idx, :2, 2] - centroids[idx, :2]\n```\nThe funny thing is that when you check the `world_from_agents matrix[:, :2, 2]` matrix and `centroids[:, :2]` matrix. They are exactly the same....\n\n![world from agents](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fff211c6780b02630d460dc0ec38e8787%2Fworld_from_agent.png?generation=1603958261780537&alt=media)\n\n![centroids](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F19cd582ada25e21469f524908af1603d%2Fcentroid.png?generation=1603958284550144&alt=media)\n\n![diff](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fcd56524370ed3785d8398685f5ae950d%2Fdiff.png?generation=1603958304404184&alt=media)\n\nSo I guess we can just simplify the transform as \n```\npreds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n```\nAnd, I believe with some numpy trick, you can easily do this for all the idx and mode without the for loop. \nIs my assumption that the last column of `world_from _agents` is the same as the `centroid` correct?\nI also noticed that the centroid is in real world coordinate so the value tend to be very large like ~500 compared to the agents local coordinate. So adding then subtracting such a large number can slightly harm the precision of prediction in general.\n\nEdit:\nI found that you can simplify the last transformation as\n```\npreds = torch.einsum('bmti,bji->bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n```\nI verify this transformation give the same score as the original formula on Kaggle in this notebook  (v12)\nhttps://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline/notebook#Prediction",
      "votes": null
    },
    {
      "id": "1063717",
      "postDate": "10/29/2020 08:23:17",
      "content": "<p>OK, look at the code, <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L8-L25\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L8-L25</a><br>\nThe last column is indeed copy from <code>agent_centroid_m</code>.</p>",
      "rawMarkdown": "OK, look at the code, https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L8-L25\nThe last column is indeed copy from `agent_centroid_m`.",
      "votes": null
    },
    {
      "id": "1063724",
      "postDate": "10/29/2020 08:30:43",
      "content": "<p>Look at more than one sample from the dataset. You've chosen a sample where the AV is the agent. These transformations were introduced in l5kit 1.1.0 to fix an issue where the targets for non-AV vehicles were not properly rotated so while properly oriented for the AV they were randomly oriented for other vehicles. The l5kit rasterizer was updated to properly transform targets but they chose not to modify the competition ground truths given the issues this would cause (invalidating all old submissions). This transformation corrects predictions to match the mistaken ground truths. So it has no affect on AVs but does affect other vehicles.</p>",
      "rawMarkdown": "Look at more than one sample from the dataset. You've chosen a sample where the AV is the agent. These transformations were introduced in l5kit 1.1.0 to fix an issue where the targets for non-AV vehicles were not properly rotated so while properly oriented for the AV they were randomly oriented for other vehicles. The l5kit rasterizer was updated to properly transform targets but they chose not to modify the competition ground truths given the issues this would cause (invalidating all old submissions). This transformation corrects predictions to match the mistaken ground truths. So it has no affect on AVs but does affect other vehicles.",
      "votes": null
    },
    {
      "id": "1065067",
      "postDate": "10/30/2020 20:06:10",
      "content": "<p>I have look into the code. The last column of the transform matrix is just copy from centroid. So they are indeed identical. My test code also get the same scores between two approaches.</p>",
      "rawMarkdown": "I have look into the code. The last column of the transform matrix is just copy from centroid. So they are indeed identical. My test code also get the same scores between two approaches.",
      "votes": null
    },
    {
      "id": "1066064",
      "postDate": "11/01/2020 09:18:05",
      "content": "<p>I found that you can simplify the last transformation as</p>\n<pre><code>preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n</code></pre>\n<p>and doing all of them on GPU using pytorch instead of CPU. I verify this transformation give the same score as the original formula on Kaggle.</p>",
      "rawMarkdown": "I found that you can simplify the last transformation as\n```\npreds = torch.einsum('bmti,bji->bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n```\nand doing all of them on GPU using pytorch instead of CPU. I verify this transformation give the same score as the original formula on Kaggle.",
      "votes": null
    },
    {
      "id": "1069370",
      "postDate": "11/04/2020 11:13:12",
      "content": "<p>Or, perhaps, one can do everything on CPU. Not sure which one is faster for inference.</p>",
      "rawMarkdown": "Or, perhaps, one can do everything on CPU. Not sure which one is faster for inference.",
      "votes": null
    },
    {
      "id": "1069475",
      "postDate": "11/04/2020 13:57:32",
      "content": "<p>Yes, my mistake. The <code>world_from_agent</code> is doing a post rotation translation. So the last column is just adding the centroid. So your simplification is entirely valid.</p>",
      "rawMarkdown": "Yes, my mistake. The `world_from_agent` is doing a post rotation translation. So the last column is just adding the centroid. So your simplification is entirely valid.",
      "votes": null
    },
    {
      "id": "1071150",
      "postDate": "11/06/2020 15:01:11",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a> for sharing this. Looking at your kernel, I am suprised this did not improve the inference time by more than a few seconds. Is that right?</p>",
      "rawMarkdown": "Thanks @louis925 for sharing this. Looking at your kernel, I am suprised this did not improve the inference time by more than a few seconds. Is that right?",
      "votes": null
    },
    {
      "id": "1071354",
      "postDate": "11/06/2020 19:04:12",
      "content": "<p>Right, it doesn't seem to improve on Kaggle notebook. However, it seem to improve on my local machine. Maybe someone can double check that.</p>",
      "rawMarkdown": "Right, it doesn't seem to improve on Kaggle notebook. However, it seem to improve on my local machine. Maybe someone can double check that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1063717,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "10/29/2020 08:23:17",
      "content": "<p>OK, look at the code, <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L8-L25\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L8-L25</a><br>\nThe last column is indeed copy from <code>agent_centroid_m</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1063724,
      "author_name": "thomasbrandon",
      "author_url": "",
      "post_date": "10/29/2020 08:30:43",
      "content": "<p>Look at more than one sample from the dataset. You've chosen a sample where the AV is the agent. These transformations were introduced in l5kit 1.1.0 to fix an issue where the targets for non-AV vehicles were not properly rotated so while properly oriented for the AV they were randomly oriented for other vehicles. The l5kit rasterizer was updated to properly transform targets but they chose not to modify the competition ground truths given the issues this would cause (invalidating all old submissions). This transformation corrects predictions to match the mistaken ground truths. So it has no affect on AVs but does affect other vehicles.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1065067,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "10/30/2020 20:06:10",
          "content": "<p>I have look into the code. The last column of the transform matrix is just copy from centroid. So they are indeed identical. My test code also get the same scores between two approaches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069475,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "11/04/2020 13:57:32",
          "content": "<p>Yes, my mistake. The <code>world_from_agent</code> is doing a post rotation translation. So the last column is just adding the centroid. So your simplification is entirely valid.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1066064,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/01/2020 09:18:05",
      "content": "<p>I found that you can simplify the last transformation as</p>\n<pre><code>preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n</code></pre>\n<p>and doing all of them on GPU using pytorch instead of CPU. I verify this transformation give the same score as the original formula on Kaggle.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1069370,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/04/2020 11:13:12",
          "content": "<p>Or, perhaps, one can do everything on CPU. Not sure which one is faster for inference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1071150,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/06/2020 15:01:11",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a> for sharing this. Looking at your kernel, I am suprised this did not improve the inference time by more than a few seconds. Is that right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1071354,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/06/2020 19:04:12",
          "content": "<p>Right, it doesn't seem to improve on Kaggle notebook. However, it seem to improve on my local machine. Maybe someone can double check that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1063691": "I am trying to understand what the final transformation in the prediction was doing. From popular notebook like [this](https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline), we have\n```python\nfor idx in range(len(preds)):\n    for mode in range(3):\n        preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]\n```\nAnd, when looks into what the `transform_points()` is doing [github link](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L73-L99), which can be simplified as (for our 2d points):\n```\ndef transform_points(points, transf_matrix)\n    return (points @ transf_matrix[:2, :2].T) + transf_matrix[:2, 2]\n```\nOne can simplify this entire transformation as simple matrix multiplication:\n```\npreds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n + world_from_agents[idx, :2, 2] - centroids[idx, :2]\n```\nThe funny thing is that when you check the `world_from_agents matrix[:, :2, 2]` matrix and `centroids[:, :2]` matrix. They are exactly the same....\n\n![world from agents](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fff211c6780b02630d460dc0ec38e8787%2Fworld_from_agent.png?generation=1603958261780537&alt=media)\n\n![centroids](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F19cd582ada25e21469f524908af1603d%2Fcentroid.png?generation=1603958284550144&alt=media)\n\n![diff](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2Fcd56524370ed3785d8398685f5ae950d%2Fdiff.png?generation=1603958304404184&alt=media)\n\nSo I guess we can just simplify the transform as \n```\npreds[idx, mode, :, :] = np.matmul(preds[idx, mode, :, :], world_from_agents[idx, :2, :2].T)\n```\nAnd, I believe with some numpy trick, you can easily do this for all the idx and mode without the for loop. \nIs my assumption that the last column of `world_from _agents` is the same as the `centroid` correct?\nI also noticed that the centroid is in real world coordinate so the value tend to be very large like ~500 compared to the agents local coordinate. So adding then subtracting such a large number can slightly harm the precision of prediction in general.\n\nEdit:\nI found that you can simplify the last transformation as\n```\npreds = torch.einsum('bmti,bji->bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n```\nI verify this transformation give the same score as the original formula on Kaggle in this notebook  (v12)\nhttps://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline/notebook#Prediction",
    "1063717": "OK, look at the code, https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/geometry/transform.py#L8-L25\nThe last column is indeed copy from `agent_centroid_m`.",
    "1063724": "Look at more than one sample from the dataset. You've chosen a sample where the AV is the agent. These transformations were introduced in l5kit 1.1.0 to fix an issue where the targets for non-AV vehicles were not properly rotated so while properly oriented for the AV they were randomly oriented for other vehicles. The l5kit rasterizer was updated to properly transform targets but they chose not to modify the competition ground truths given the issues this would cause (invalidating all old submissions). This transformation corrects predictions to match the mistaken ground truths. So it has no affect on AVs but does affect other vehicles.",
    "1065067": "I have look into the code. The last column of the transform matrix is just copy from centroid. So they are indeed identical. My test code also get the same scores between two approaches.",
    "1066064": "I found that you can simplify the last transformation as\n```\npreds = torch.einsum('bmti,bji->bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n```\nand doing all of them on GPU using pytorch instead of CPU. I verify this transformation give the same score as the original formula on Kaggle.",
    "1069370": "Or, perhaps, one can do everything on CPU. Not sure which one is faster for inference.",
    "1069475": "Yes, my mistake. The `world_from_agent` is doing a post rotation translation. So the last column is just adding the centroid. So your simplification is entirely valid.",
    "1071150": "Thanks @louis925 for sharing this. Looking at your kernel, I am suprised this did not improve the inference time by more than a few seconds. Is that right?",
    "1071354": "Right, it doesn't seem to improve on Kaggle notebook. However, it seem to improve on my local machine. Maybe someone can double check that."
  },
  "source": "meta"
}