{
  "id": 186492,
  "title": "We did it all wrong",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/186492",
  "author_name": "nosound",
  "post_date": "2020-09-24T17:03:46.365000",
  "votes": 117,
  "comment_count": 107,
  "views": 0,
  "content": "<p>Hi Kagglers, I came across something important that I am excited to share. Despite my LB score improving from <code>32.210</code> to <code>25.524</code> after that fix with using <code>3x</code> less training epochs to get there, I am still not 100% sure that I do not miss something simple and obvious to everyone but me. But here it is for your judgement.</p>\n<h1>Statement</h1>\n<p>We have input images of the following kind, aligned for our agent vehicle to look right and at <code>'ego_center': [0.25, 0.5]</code>. It is the only input to our simple NNs that everyone uses.</p>\n<p><img src=\"https://i.imgur.com/uQYCT7u.png\" alt=\"typical input image\"></p>\n<p>But then we predict targets which are not aligned, at least if you follow <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">the example notebook</a> code. But the public notebooks that I saw use the same approach, if I draw several hundreds of target trajectories that we try to predict then it looks like this:</p>\n<p><img src=\"https://i.imgur.com/gxPjOAn.png\" alt=\"targets in world coordinates\"></p>\n<p>It is in all different directions, as expected. Then I spent a day trying to understand how is that possible to predict correct directions from direction-less input. To no avail, I still do not understand it. But the predictions are so accurate, they are actually very good! The direction is somehow predicted correctly from a direction-less input.</p>\n<p>After failing to understand why everyone does it wrong and why it even works, I tried to implement it correctly, that is, <strong>transform the targets to image coordinates</strong> before training. I attach the pytorch code with the transformations below so you can try it yourself fast. If you try and get an improvement please report it here.</p>\n<p>Transformed target trajectories look like this, this is 160 randomly sampled targets</p>\n<p><img src=\"https://i.imgur.com/NZVGGkT.png\" alt=\"targets in image coordinates\"></p>\n<p>And it worked like a charm. With only 15 epochs, 10k batches each, batch size 16, which is about 5h training on my machine, I improved my training score from <code>30.87</code> to <code>23.79</code> and public LB <code>32.210</code> to <code>25.524</code>. </p>\n<p>So how is the old approach able to predict the direction? I have two theories but no prove:</p>\n<ol>\n<li>Number of streets is very limited, and the network is able to learn what street has what orientation and then apply it to predictions.</li>\n<li>The image does somehow contain the rotation transformation in itself, maybe some subtle pixel level traces of it are there.</li>\n</ol>\n<p>What do you think? I have a lingering feeling that I am missing something obvious!</p>\n<h1>Code</h1>\n<p>Below boolean flag <code>cfg['train_params']['image_coords']</code> controls if I use the new approach. I apply the needed transformations on targets and then do inverse transformation on predictions in reverse order. Note that my predictions are multi-scenario there, you need to get rid of this additional dimensions if you use single scenario.</p>\n<pre><code>def forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    matrix = data[\"world_to_image\"].to(device)\n    centroid = data[\"centroid\"].to(device)[:,None,:].to(torch.float)\n\n    # Forward pass\n    outputs = model(inputs)\n\n    bs,tl,_ = targets.shape\n    assert tl == cfg[\"model_params\"][\"future_num_frames\"]\n\n    if cfg['train_params']['image_coords']:\n        targets = targets + centroid\n        targets = torch.cat([targets,torch.ones((bs,tl,1)).to(device)], dim=2)\n        targets = torch.matmul(matrix.to(torch.float), targets.transpose(1,2))\n        targets = targets.transpose(1,2)[:,:,:2]\n        bias = torch.tensor([56.25, 112.5])[None,None,:].to(device)\n        targets = targets - bias\n\n    confidences, pred = outputs[:,:3], outputs[:,3:]\n    pred = pred.view(bs, 3, tl, 2)\n    assert confidences.shape == (bs, 3)\n    confidences = torch.softmax(confidences, dim=1)\n\n    loss = criterion(targets, pred, confidences, target_availabilities)\n    loss = torch.mean(loss)\n\n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n\n    return loss, pred, confidences\n</code></pre>\n<p>and in your loss function you may want to divide the coordinates by 2:</p>\n<pre><code>    ...\n    error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    if cfg['train_params']['image_coords']:\n        error = error / 4\n    ...\n</code></pre>",
  "messages": [
    {
      "id": 1025600,
      "postDate": "2020-09-24T17:03:46.367Z",
      "content": "<p>Hi Kagglers, I came across something important that I am excited to share. Despite my LB score improving from <code>32.210</code> to <code>25.524</code> after that fix with using <code>3x</code> less training epochs to get there, I am still not 100% sure that I do not miss something simple and obvious to everyone but me. But here it is for your judgement.</p>\n<h1>Statement</h1>\n<p>We have input images of the following kind, aligned for our agent vehicle to look right and at <code>'ego_center': [0.25, 0.5]</code>. It is the only input to our simple NNs that everyone uses.</p>\n<p><img src=\"https://i.imgur.com/uQYCT7u.png\" alt=\"typical input image\"></p>\n<p>But then we predict targets which are not aligned, at least if you follow <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">the example notebook</a> code. But the public notebooks that I saw use the same approach, if I draw several hundreds of target trajectories that we try to predict then it looks like this:</p>\n<p><img src=\"https://i.imgur.com/gxPjOAn.png\" alt=\"targets in world coordinates\"></p>\n<p>It is in all different directions, as expected. Then I spent a day trying to understand how is that possible to predict correct directions from direction-less input. To no avail, I still do not understand it. But the predictions are so accurate, they are actually very good! The direction is somehow predicted correctly from a direction-less input.</p>\n<p>After failing to understand why everyone does it wrong and why it even works, I tried to implement it correctly, that is, <strong>transform the targets to image coordinates</strong> before training. I attach the pytorch code with the transformations below so you can try it yourself fast. If you try and get an improvement please report it here.</p>\n<p>Transformed target trajectories look like this, this is 160 randomly sampled targets</p>\n<p><img src=\"https://i.imgur.com/NZVGGkT.png\" alt=\"targets in image coordinates\"></p>\n<p>And it worked like a charm. With only 15 epochs, 10k batches each, batch size 16, which is about 5h training on my machine, I improved my training score from <code>30.87</code> to <code>23.79</code> and public LB <code>32.210</code> to <code>25.524</code>. </p>\n<p>So how is the old approach able to predict the direction? I have two theories but no prove:</p>\n<ol>\n<li>Number of streets is very limited, and the network is able to learn what street has what orientation and then apply it to predictions.</li>\n<li>The image does somehow contain the rotation transformation in itself, maybe some subtle pixel level traces of it are there.</li>\n</ol>\n<p>What do you think? I have a lingering feeling that I am missing something obvious!</p>\n<h1>Code</h1>\n<p>Below boolean flag <code>cfg['train_params']['image_coords']</code> controls if I use the new approach. I apply the needed transformations on targets and then do inverse transformation on predictions in reverse order. Note that my predictions are multi-scenario there, you need to get rid of this additional dimensions if you use single scenario.</p>\n<pre><code>def forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    matrix = data[\"world_to_image\"].to(device)\n    centroid = data[\"centroid\"].to(device)[:,None,:].to(torch.float)\n\n    # Forward pass\n    outputs = model(inputs)\n\n    bs,tl,_ = targets.shape\n    assert tl == cfg[\"model_params\"][\"future_num_frames\"]\n\n    if cfg['train_params']['image_coords']:\n        targets = targets + centroid\n        targets = torch.cat([targets,torch.ones((bs,tl,1)).to(device)], dim=2)\n        targets = torch.matmul(matrix.to(torch.float), targets.transpose(1,2))\n        targets = targets.transpose(1,2)[:,:,:2]\n        bias = torch.tensor([56.25, 112.5])[None,None,:].to(device)\n        targets = targets - bias\n\n    confidences, pred = outputs[:,:3], outputs[:,3:]\n    pred = pred.view(bs, 3, tl, 2)\n    assert confidences.shape == (bs, 3)\n    confidences = torch.softmax(confidences, dim=1)\n\n    loss = criterion(targets, pred, confidences, target_availabilities)\n    loss = torch.mean(loss)\n\n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n\n    return loss, pred, confidences\n</code></pre>\n<p>and in your loss function you may want to divide the coordinates by 2:</p>\n<pre><code>    ...\n    error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    if cfg['train_params']['image_coords']:\n        error = error / 4\n    ...\n</code></pre>",
      "rawMarkdown": "Hi Kagglers, I came across something important that I am excited to share. Despite my LB score improving from `32.210` to `25.524` after that fix with using `3x` less training epochs to get there, I am still not 100% sure that I do not miss something simple and obvious to everyone but me. But here it is for your judgement.\n\n# Statement\n\nWe have input images of the following kind, aligned for our agent vehicle to look right and at `'ego_center': [0.25, 0.5]`. It is the only input to our simple NNs that everyone uses.\n\n![typical input image](https://i.imgur.com/uQYCT7u.png)\n\nBut then we predict targets which are not aligned, at least if you follow [the example notebook](https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb) code. But the public notebooks that I saw use the same approach, if I draw several hundreds of target trajectories that we try to predict then it looks like this:\n\n![targets in world coordinates](https://i.imgur.com/gxPjOAn.png)\n\nIt is in all different directions, as expected. Then I spent a day trying to understand how is that possible to predict correct directions from direction-less input. To no avail, I still do not understand it. But the predictions are so accurate, they are actually very good! The direction is somehow predicted correctly from a direction-less input.\n\nAfter failing to understand why everyone does it wrong and why it even works, I tried to implement it correctly, that is, **transform the targets to image coordinates** before training. I attach the pytorch code with the transformations below so you can try it yourself fast. If you try and get an improvement please report it here.\n\nTransformed target trajectories look like this, this is 160 randomly sampled targets\n\n![targets in image coordinates](https://i.imgur.com/NZVGGkT.png)\n\nAnd it worked like a charm. With only 15 epochs, 10k batches each, batch size 16, which is about 5h training on my machine, I improved my training score from `30.87` to `23.79` and public LB `32.210` to `25.524`. \n\nSo how is the old approach able to predict the direction? I have two theories but no prove:\n1. Number of streets is very limited, and the network is able to learn what street has what orientation and then apply it to predictions.\n2. The image does somehow contain the rotation transformation in itself, maybe some subtle pixel level traces of it are there.\n\nWhat do you think? I have a lingering feeling that I am missing something obvious!\n\n# Code\n\nBelow boolean flag `cfg['train_params']['image_coords']` controls if I use the new approach. I apply the needed transformations on targets and then do inverse transformation on predictions in reverse order. Note that my predictions are multi-scenario there, you need to get rid of this additional dimensions if you use single scenario.\n\n```\ndef forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    matrix = data[\"world_to_image\"].to(device)\n    centroid = data[\"centroid\"].to(device)[:,None,:].to(torch.float)\n    \n    # Forward pass\n    outputs = model(inputs)\n    \n    bs,tl,_ = targets.shape\n    assert tl == cfg[\"model_params\"][\"future_num_frames\"]\n    \n    if cfg['train_params']['image_coords']:\n        targets = targets + centroid\n        targets = torch.cat([targets,torch.ones((bs,tl,1)).to(device)], dim=2)\n        targets = torch.matmul(matrix.to(torch.float), targets.transpose(1,2))\n        targets = targets.transpose(1,2)[:,:,:2]\n        bias = torch.tensor([56.25, 112.5])[None,None,:].to(device)\n        targets = targets - bias\n    \n    confidences, pred = outputs[:,:3], outputs[:,3:]\n    pred = pred.view(bs, 3, tl, 2)\n    assert confidences.shape == (bs, 3)\n    confidences = torch.softmax(confidences, dim=1)\n    \n    loss = criterion(targets, pred, confidences, target_availabilities)\n    loss = torch.mean(loss)\n    \n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n    \n    return loss, pred, confidences\n```\n\nand in your loss function you may want to divide the coordinates by 2:\n\n```\n    ...\n    error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    if cfg['train_params']['image_coords']:\n        error = error / 4\n    ...\n```",
      "votes": 117
    },
    {
      "id": 1026361,
      "postDate": "2020-09-25T08:49:31.527Z",
      "content": "<p>Hey Everyone,</p>\n<blockquote>\n  <p>What do you think? I have a lingering feeling that I am missing something obvious!</p>\n</blockquote>\n<p>On the contrary, you're spot on :) This is something that slipped through our tests and that we only recently became aware of.</p>\n<p>Using pixel coordinates is absolutely legit here, but <strong>we're also working on a long term fix in L5Kit</strong> we hope we can ship next week. Once that is thoroughly checked we will sync with Kaggle to update the default package.</p>\n<p>I'll start a post with more info about this issue and how we plan to fix it to increase visibility and awareness.</p>",
      "rawMarkdown": "Hey Everyone,\n> What do you think? I have a lingering feeling that I am missing something obvious!\n\nOn the contrary, you're spot on :) This is something that slipped through our tests and that we only recently became aware of.\n\nUsing pixel coordinates is absolutely legit here, but **we're also working on a long term fix in L5Kit** we hope we can ship next week. Once that is thoroughly checked we will sync with Kaggle to update the default package.\n\nI'll start a post with more info about this issue and how we plan to fix it to increase visibility and awareness.",
      "votes": 10,
      "replies": [
        {
          "id": 1030060,
          "postDate": "2020-09-28T11:39:01.493Z",
          "content": "<p>any updates?</p>",
          "rawMarkdown": "any updates?"
        },
        {
          "id": 1030368,
          "postDate": "2020-09-28T15:58:08.520Z",
          "content": "<p>almost ready to merge <a href=\"https://github.com/lyft/l5kit/pull/150\" target=\"_blank\">https://github.com/lyft/l5kit/pull/150</a> :) After that we will cut a new release and update Kaggle, doc and baseline</p>",
          "rawMarkdown": "almost ready to merge https://github.com/lyft/l5kit/pull/150 :) After that we will cut a new release and update Kaggle, doc and baseline",
          "votes": 2
        },
        {
          "id": 1030406,
          "postDate": "2020-09-28T16:41:34.090Z",
          "content": "<p>I compared that PR to the fix posted here and looks like you're keeping the target positions in metres whereas above code converts to pixels, otherwise they match (barring slight inaccuracy likely due to floating point). That's the intended final result correct? (No issue, I think metres is better, just checking)</p>\n<p>I gather also that you'll be updating the Kaggle ground truths so predictions will be in the new coordinates. Is that correct?</p>",
          "rawMarkdown": "I compared that PR to the fix posted here and looks like you're keeping the target positions in metres whereas above code converts to pixels, otherwise they match (barring slight inaccuracy likely due to floating point). That's the intended final result correct? (No issue, I think metres is better, just checking)\n\nI gather also that you'll be updating the Kaggle ground truths so predictions will be in the new coordinates. Is that correct?",
          "votes": 2
        },
        {
          "id": 1030418,
          "postDate": "2020-09-28T16:52:25.093Z",
          "content": "<p>I think pixels are better, - in pixels they can be directly connected in creative ways to the images.</p>",
          "rawMarkdown": "I think pixels are better, - in pixels they can be directly connected in creative ways to the images.",
          "votes": 1
        },
        {
          "id": 1030428,
          "postDate": "2020-09-28T16:58:47.383Z",
          "content": "<blockquote>\n  <p>That's the intended final result correct?</p>\n</blockquote>\n<p>yes, we're keeping metres (basically the only difference is a 2D rotation) :)</p>\n<blockquote>\n  <p>I gather also that you'll be updating the Kaggle ground truths so predictions will be in the new coordinates. Is that correct?</p>\n</blockquote>\n<p>We're not, this means you will have to convert them back into offsets in world coordinates (but we'll update the baseline notebook to show how to do that and we're adding a lot of matrices to make it easier). Updating the ground truth would have required a lot of additional testing, this was the best compromise :( </p>",
          "rawMarkdown": "> That's the intended final result correct?\n\nyes, we're keeping metres (basically the only difference is a 2D rotation) :)\n\n> I gather also that you'll be updating the Kaggle ground truths so predictions will be in the new coordinates. Is that correct?\n\nWe're not, this means you will have to convert them back into offsets in world coordinates (but we'll update the baseline notebook to show how to do that and we're adding a lot of matrices to make it easier). Updating the ground truth would have required a lot of additional testing, this was the best compromise :( ",
          "votes": 1
        },
        {
          "id": 1030433,
          "postDate": "2020-09-28T17:02:57.873Z",
          "content": "<blockquote>\n  <p>I think pixels are better</p>\n</blockquote>\n<p>How so? Either way you can just multiply/divide by pixel size right? I think if you want that you probably want to do it within a single model and then convert back to metres at the end. Otherwise if doing something like combining outputs from multiple models with different pixel sizes you'd have to be tracking the pixel sizes of all models.<br>\nWhat if you aren't even using a pixel based model, why convert to some arbitrary pizel size?</p>",
          "rawMarkdown": "> I think pixels are better\n\nHow so? Either way you can just multiply/divide by pixel size right? I think if you want that you probably want to do it within a single model and then convert back to metres at the end. Otherwise if doing something like combining outputs from multiple models with different pixel sizes you'd have to be tracking the pixel sizes of all models.\nWhat if you aren't even using a pixel based model, why convert to some arbitrary pizel size?",
          "votes": 1
        },
        {
          "id": 1030441,
          "postDate": "2020-09-28T17:09:37.190Z",
          "content": "<blockquote>\n  <p>We're not, this means you will have to convert them back into offsets in world coordinates (but we'll update the baseline notebook to show how to do that and we're adding a lot of matrices to make it easier). Updating the ground truth would have required a lot of additional testing, this was the best compromise :( </p>\n</blockquote>\n<p>Ah OK, I misunderstood the changes to the GT.csv code as meaning they were being updated to new coords when I guess they're converting new coords back to the old ones. Presumably they will match expected outputs on Kaggle.<br>\nI understand the compromise, obviously tricky either way.</p>",
          "rawMarkdown": "> We're not, this means you will have to convert them back into offsets in world coordinates (but we'll update the baseline notebook to show how to do that and we're adding a lot of matrices to make it easier). Updating the ground truth would have required a lot of additional testing, this was the best compromise :( \n\nAh OK, I misunderstood the changes to the GT.csv code as meaning they were being updated to new coords when I guess they're converting new coords back to the old ones. Presumably they will match expected outputs on Kaggle.\nI understand the compromise, obviously tricky either way.",
          "votes": 1
        },
        {
          "id": 1030447,
          "postDate": "2020-09-28T17:13:58.037Z",
          "content": "<blockquote>\n  <p>I guess they're converting new coords back to the old ones. Presumably they will match expected outputs on Kaggle</p>\n</blockquote>\n<p>That's is exactly what's happening there :) I've checked and they match the old ones (also I've submitted a new baseline trained using the latest master and results are much better so I'm positive we didn't screw everything up completely)</p>",
          "rawMarkdown": "> I guess they're converting new coords back to the old ones. Presumably they will match expected outputs on Kaggle\n\nThat's is exactly what's happening there :) I've checked and they match the old ones (also I've submitted a new baseline trained using the latest master and results are much better so I'm positive we didn't screw everything up completely)",
          "votes": 2
        },
        {
          "id": 1030458,
          "postDate": "2020-09-28T17:20:15.487Z",
          "content": "<p>Probably belongs elsewhere but in relation to the GCP credit prize for those beating the baseline does that mean they need to beat new baseline? Think the first deadline for that was tomorrow so not a lot of time to update models.</p>",
          "rawMarkdown": "Probably belongs elsewhere but in relation to the GCP credit prize for those beating the baseline does that mean they need to beat new baseline? Think the first deadline for that was tomorrow so not a lot of time to update models.",
          "votes": 2
        },
        {
          "id": 1030480,
          "postDate": "2020-09-28T17:29:02.487Z",
          "content": "<p>Where is the new baseline result? I could not find it. Or if someone can just report the new score here? </p>",
          "rawMarkdown": "Where is the new baseline result? I could not find it. Or if someone can just report the new score here? "
        },
        {
          "id": 1030498,
          "postDate": "2020-09-28T17:35:05.717Z",
          "content": "<p>82.246<br>\n<a href=\"https://www.kaggle.com/lucabergamini/lyft-baseline-09-02\" target=\"_blank\">https://www.kaggle.com/lucabergamini/lyft-baseline-09-02</a></p>",
          "rawMarkdown": "82.246\nhttps://www.kaggle.com/lucabergamini/lyft-baseline-09-02",
          "votes": 1
        },
        {
          "id": 1030499,
          "postDate": "2020-09-28T17:35:15.993Z",
          "content": "<blockquote>\n  <p>How so? Either way you can just multiply/divide by pixel size right? I think if you want that you probably want to do it within a single model and then convert back to metres at the end. Otherwise if doing something like combining outputs from multiple models with different pixel sizes you'd have to be tracking the pixel sizes of all models.<br>\n  What if you aren't even using a pixel based model, why convert to some arbitrary pizel size?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> , you have good points, hard to argue with what you say. In defense of pixels, - we get <code>data</code> object in which I would naturally expect everything to be in the same coordinate system. There is the image object, with its own coordinate system, so everything needs to match. My second argument is that if everything is in pixels then we can easily draw on the image without any additional transformations, rasterize those pixels into something etc.</p>",
          "rawMarkdown": "> How so? Either way you can just multiply/divide by pixel size right? I think if you want that you probably want to do it within a single model and then convert back to metres at the end. Otherwise if doing something like combining outputs from multiple models with different pixel sizes you'd have to be tracking the pixel sizes of all models.\nWhat if you aren't even using a pixel based model, why convert to some arbitrary pizel size?\n\n@thomasbrandon , you have good points, hard to argue with what you say. In defense of pixels, - we get `data` object in which I would naturally expect everything to be in the same coordinate system. There is the image object, with its own coordinate system, so everything needs to match. My second argument is that if everything is in pixels then we can easily draw on the image without any additional transformations, rasterize those pixels into something etc."
        }
      ]
    },
    {
      "id": 1025660,
      "postDate": "2020-09-24T17:35:20.237Z",
      "content": "<p>Did you calculate <code>bias = torch.tensor([56.25, 112.5])</code> from the <code>ego_center</code> config? So the input size is 225x225px?</p>",
      "rawMarkdown": "Did you calculate `bias = torch.tensor([56.25, 112.5])` from the `ego_center` config? So the input size is 225x225px?",
      "votes": 5,
      "replies": [
        {
          "id": 1025677,
          "postDate": "2020-09-24T17:47:58.253Z",
          "content": "<p>That was my question as well. I assume this will not be the correct alignment at other resolutions. </p>",
          "rawMarkdown": "That was my question as well. I assume this will not be the correct alignment at other resolutions. ",
          "votes": 1
        },
        {
          "id": 1025694,
          "postDate": "2020-09-24T17:58:03.900Z",
          "content": "<p>Maybe this..</p>\n<pre><code>rs = cfg[\"raster_params\"][\"raster_size\"]\nec = cfg[\"raster_params\"][\"ego_center\"]\n\nbias = torch.tensor([rs[0] * ec[0], rs[1] * ec[1]])[None, None, :].to(device)\n</code></pre>",
          "rawMarkdown": "Maybe this..\n\n```\nrs = cfg[\"raster_params\"][\"raster_size\"]\nec = cfg[\"raster_params\"][\"ego_center\"]\n\nbias = torch.tensor([rs[0] * ec[0], rs[1] * ec[1]])[None, None, :].to(device)\n```",
          "votes": 5
        },
        {
          "id": 1025730,
          "postDate": "2020-09-24T18:21:17.327Z",
          "content": "<p>Exactly right, thank you! The point of this is of course to have non-moving target at zero, as it was before. </p>",
          "rawMarkdown": "Exactly right, thank you! The point of this is of course to have non-moving target at zero, as it was before. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1029906,
      "postDate": "2020-09-28T08:37:46.400Z",
      "content": "<p>Thank you for sharing this. We were doing it completely wrong, and most people would have been oblivious to that fact had you not so generously highlighted it. The epitome of community - kudos.</p>",
      "rawMarkdown": "Thank you for sharing this. We were doing it completely wrong, and most people would have been oblivious to that fact had you not so generously highlighted it. The epitome of community - kudos.",
      "votes": 6
    },
    {
      "id": 1027976,
      "postDate": "2020-09-26T13:47:09.237Z",
      "content": "<p>Thank you for sharing. <br>\nWith this approach, in my case, the model(45.0) trained on all of the datasets was improved to the model(37.7) trained on 1/4 of the dataset.<br>\nHowever, there is a gap between validation metric (13.1) and public LB. </p>",
      "rawMarkdown": "Thank you for sharing. \nWith this approach, in my case, the model(45.0) trained on all of the datasets was improved to the model(37.7) trained on 1/4 of the dataset.\nHowever, there is a gap between validation metric (13.1) and public LB. ",
      "votes": 3,
      "replies": [
        {
          "id": 1028012,
          "postDate": "2020-09-26T14:27:01.327Z",
          "content": "<p>Your validation metric is too small, this can be because you divide the errors by 2 <strong>after</strong> the predictions are transformed. Basically check that the errors are divided by 2 only when needed.</p>",
          "rawMarkdown": "Your validation metric is too small, this can be because you divide the errors by 2 **after** the predictions are transformed. Basically check that the errors are divided by 2 only when needed.",
          "votes": 1
        },
        {
          "id": 1028488,
          "postDate": "2020-09-26T22:29:26.957Z",
          "content": "<p>Sorry. When I checked it, I used the square error, so I divided it by 4.</p>",
          "rawMarkdown": "Sorry. When I checked it, I used the square error, so I divided it by 4."
        }
      ]
    },
    {
      "id": 1026441,
      "postDate": "2020-09-25T10:22:50.063Z",
      "content": "<p>Good approach <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> . Congrats !<br>\nCan I ask you why do you choose to divide the error by 4 ?</p>",
      "rawMarkdown": "Good approach @zaharch . Congrats !\nCan I ask you why do you choose to divide the error by 4 ?",
      "votes": 3,
      "replies": [
        {
          "id": 1026516,
          "postDate": "2020-09-25T11:29:26.090Z",
          "content": "<p>In case one uses default configuration there is a scaling factor <code>'pixel_size': [0.5, 0.5]</code>. So there are two pixels (new coordinates) per meter (old coordinates). To get the error back into the old scale I need to divide it by 2, or by 4 in case of squared error. </p>\n<p>Depending on your model it may not be necessary to scale, you will just have 4x bigger loss. But I train a multi-scenario output, and there is a subtle inter-play between the errors and the confidences. Therefore I want to keep it at the old scale, which is the correct one for the loss.</p>",
          "rawMarkdown": "In case one uses default configuration there is a scaling factor `'pixel_size': [0.5, 0.5]`. So there are two pixels (new coordinates) per meter (old coordinates). To get the error back into the old scale I need to divide it by 2, or by 4 in case of squared error. \n\nDepending on your model it may not be necessary to scale, you will just have 4x bigger loss. But I train a multi-scenario output, and there is a subtle inter-play between the errors and the confidences. Therefore I want to keep it at the old scale, which is the correct one for the loss.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1025888,
      "postDate": "2020-09-24T21:06:16.230Z",
      "content": "<p>We really did it all wrong! Thanks for sharing and pointing it out. </p>",
      "rawMarkdown": "We really did it all wrong! Thanks for sharing and pointing it out. ",
      "votes": 3
    },
    {
      "id": 1031260,
      "postDate": "2020-09-29T10:54:47.223Z",
      "content": "<p>I'm really new to this community and this helped me a lot. Thanks!</p>",
      "rawMarkdown": "I'm really new to this community and this helped me a lot. Thanks!",
      "votes": 1
    },
    {
      "id": 1028934,
      "postDate": "2020-09-27T10:36:53.357Z",
      "content": "<p>Hi, Great insight! Can i know how you obtained those graphs?</p>",
      "rawMarkdown": "Hi, Great insight! Can i know how you obtained those graphs?",
      "votes": 1,
      "replies": [
        {
          "id": 1028945,
          "postDate": "2020-09-27T10:44:40.177Z",
          "content": "<p>Sure, the first graph is this code:</p>\n<pre><code>fig = plt.figure(figsize = (12,12))\nfor k in range(500):\n    data = valid_dataset[4001+1000*k]\n    tgt = data[\"target_positions\"]\n    plt.plot(tgt[:,0],tgt[:,1])\n</code></pre>\n<p>the second is similar but with the transformation, depends on how you shape your code.</p>",
          "rawMarkdown": "Sure, the first graph is this code:\n\n```\nfig = plt.figure(figsize = (12,12))\nfor k in range(500):\n    data = valid_dataset[4001+1000*k]\n    tgt = data[\"target_positions\"]\n    plt.plot(tgt[:,0],tgt[:,1])\n```\n\nthe second is similar but with the transformation, depends on how you shape your code.",
          "votes": 5
        },
        {
          "id": 1028966,
          "postDate": "2020-09-27T11:04:52.797Z",
          "content": "<p>Thanks a Lot! By the looks of it… the valid_dataset suggests you are using an evaluation dataset… why? just curious…</p>",
          "rawMarkdown": "Thanks a Lot! By the looks of it... the valid_dataset suggests you are using an evaluation dataset... why? just curious..."
        },
        {
          "id": 1029308,
          "postDate": "2020-09-27T16:34:59.013Z",
          "content": "<p>No reason, it doesn't really matter because it is just targets from the data. I don't think there is any meaningful difference between the train and the validation.</p>",
          "rawMarkdown": "No reason, it doesn't really matter because it is just targets from the data. I don't think there is any meaningful difference between the train and the validation.",
          "votes": 1
        },
        {
          "id": 1029763,
          "postDate": "2020-09-28T05:48:28.323Z",
          "content": "<p>Understood! thanks again😀</p>",
          "rawMarkdown": "Understood! thanks again😀\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1027575,
      "postDate": "2020-09-26T08:43:53.657Z",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> does this mean while inference we have to transform the predicted trajectories using the code block you shared</p>\n<pre><code>    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n</code></pre>\n<p>Trying to understand if I am interpreting the targets correctly</p>",
      "rawMarkdown": "@zaharch does this mean while inference we have to transform the predicted trajectories using the code block you shared\n```\n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n```\nTrying to understand if I am interpreting the targets correctly",
      "votes": 1,
      "replies": [
        {
          "id": 1027577,
          "postDate": "2020-09-26T08:44:51.447Z",
          "content": "<p>Yes, in inference you will have to use the same code block.</p>",
          "rawMarkdown": "Yes, in inference you will have to use the same code block.",
          "votes": 2
        },
        {
          "id": 1027613,
          "postDate": "2020-09-26T08:56:53.187Z",
          "content": "<p>Yes, that is right, same for inference.</p>",
          "rawMarkdown": "Yes, that is right, same for inference.",
          "votes": 1
        },
        {
          "id": 1027628,
          "postDate": "2020-09-26T09:06:41.810Z",
          "content": "<p>Interesting, I trained model for 6 epochs 10k steps each my train loss was around ~60.x but the inference kernel with transformation gave really bad result. Anything obvious I might be missing?</p>\n<p>Note that I am using raster rize 224x224 so I changed the bias value.</p>\n<p>Inference code for reference</p>\n<pre><code>        inputs = data[\"image\"].to(device)\n        #target_availabilities = data[\"target_availabilities\"].unsqueeze(-1).to(device)\n        target_availabilities = data[\"target_availabilities\"].to(device)\n        targets = data[\"target_positions\"].to(device)\n        matrix = data[\"world_to_image\"].to(device)\n        centroid = data[\"centroid\"].to(device)[:,None,:].to(torch.float)\n\n        bs,tl,_ = targets.shape\n        rs = cfg[\"raster_params\"][\"raster_size\"]\n        ec = cfg[\"raster_params\"][\"ego_center\"]\n        bias = torch.tensor([rs[0] * ec[0], rs[1] * ec[1]])[None, None, :].to(device)\n\n\n        #outputs = model(inputs).reshape(targets.shape)\n        pred, confidences = model(inputs)\n\n        if cfg['test_params']['image_coords']:\n            matrix_inv = torch.inverse(matrix)\n            pred = pred + bias[:,None,:,:]\n            pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n            pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                                for i in range(3)], dim=1)\n            pred = pred.transpose(2,3)[:,:,:,:2]\n            pred = pred - centroid[:,None,:,:]\n</code></pre>",
          "rawMarkdown": "Interesting, I trained model for 6 epochs 10k steps each my train loss was around ~60.x but the inference kernel with transformation gave really bad result. Anything obvious I might be missing?\n\nNote that I am using raster rize 224x224 so I changed the bias value.\n\nInference code for reference\n```\n        inputs = data[\"image\"].to(device)\n        #target_availabilities = data[\"target_availabilities\"].unsqueeze(-1).to(device)\n        target_availabilities = data[\"target_availabilities\"].to(device)\n        targets = data[\"target_positions\"].to(device)\n        matrix = data[\"world_to_image\"].to(device)\n        centroid = data[\"centroid\"].to(device)[:,None,:].to(torch.float)\n        \n        bs,tl,_ = targets.shape\n        rs = cfg[\"raster_params\"][\"raster_size\"]\n        ec = cfg[\"raster_params\"][\"ego_center\"]\n        bias = torch.tensor([rs[0] * ec[0], rs[1] * ec[1]])[None, None, :].to(device)\n\n\n        #outputs = model(inputs).reshape(targets.shape)\n        pred, confidences = model(inputs)\n        \n        if cfg['test_params']['image_coords']:\n            matrix_inv = torch.inverse(matrix)\n            pred = pred + bias[:,None,:,:]\n            pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n            pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                                for i in range(3)], dim=1)\n            pred = pred.transpose(2,3)[:,:,:,:2]\n            pred = pred - centroid[:,None,:,:]\n        \n```",
          "votes": 1
        },
        {
          "id": 1027636,
          "postDate": "2020-09-26T09:13:08.097Z",
          "content": "<p>Have you changed the bias to 56,112? How do you judge that the result is bad? Do prediction start from close to zero? I guess it is some kind of a technical error.</p>\n<p>Do you do softmax for the confidences and pred shape transformations inside the model?</p>",
          "rawMarkdown": "Have you changed the bias to 56,112? How do you judge that the result is bad? Do prediction start from close to zero? I guess it is some kind of a technical error.\n\nDo you do softmax for the confidences and pred shape transformations inside the model?"
        },
        {
          "id": 1027654,
          "postDate": "2020-09-26T09:31:26.883Z",
          "content": "<p><code>Have you changed the bias to 56,112?</code><br>\nYes,<br>\n<code>Do you do softmax for the confidences and pred shape transformations inside the model?</code><br>\nYes<br>\n<code>Do prediction start from close to zero?</code><br>\nYes<br>\n<code>How do you judge that the result is bad?</code><br>\nI am comparing model performance against my old model without transformation after n epochs. Model without transformations seemed to have converged faster in my case. I guess I'll dig deeper into this by comparing my results</p>",
          "rawMarkdown": "`Have you changed the bias to 56,112?`\nYes,\n` Do you do softmax for the confidences and pred shape transformations inside the model?`\nYes\n` Do prediction start from close to zero?`\nYes\n`How do you judge that the result is bad?`\nI am comparing model performance against my old model without transformation after n epochs. Model without transformations seemed to have converged faster in my case. I guess I'll dig deeper into this by comparing my results",
          "votes": 1
        }
      ]
    },
    {
      "id": 1025678,
      "postDate": "2020-09-24T17:48:07.827Z",
      "content": "<p>Awesome point. Thank you for sharing this</p>",
      "rawMarkdown": "Awesome point. Thank you for sharing this",
      "votes": 1
    },
    {
      "id": 1064052,
      "postDate": "2020-10-29T16:30:18.540Z",
      "content": "<p>important!!!! </p>\n<ol>\n<li>check if you can get the same values if convert from world to image coordinates, and then back again.</li>\n<li>i think you need a 64-bit float (double) to do this</li>\n</ol>",
      "rawMarkdown": "important!!!! \n1.  check if you can get the same values if convert from world to image coordinates, and then back again.\n2. i think you need a 64-bit float (double) to do this",
      "votes": 2
    },
    {
      "id": 1037369,
      "postDate": "2020-10-04T23:21:41.780Z",
      "content": "<p>So I tried to modify the <em>forward</em> function above for the new l5kit package (version 1.1.0) following the notebook <a href=\"https://www.kaggle.com/lucabergamini/lyft-baseline-09-02?scriptVersionId=43751473\" target=\"_blank\">Lyft-baseline-09-02</a></p>\n<pre><code>def forward(data, model, device, criterion = pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n\n    # Forward pass\n    preds, confidences = model(inputs)    \n\n\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n\n    world_from_agents = data[\"world_from_agent\"].numpy()\n    centroids = data[\"centroid\"].numpy()\n\n    #convert into world coordinates and compute offsets\n    for idx in range(len(preds)):\n        preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n\n    return loss, preds, confidences\n</code></pre>\n<p>So as I understand the target in the train set is fixed, however the test set is not so we still need to convert it? However when I ran the above code I encountered the following error:</p>\n<pre><code>---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n&lt;ipython-input-24-9205635f7dc5&gt; in &lt;module&gt;\n     12 for data in progress_bar:\n     13     inputs = data['image'].to(device)\n---&gt; 14     _, preds, confidences  = forward(data, model, device)\n     15 \n     16 #     # convert agent coordinates into world offsets\n\n&lt;ipython-input-22-276920502171&gt; in forward(data, model, device, criterion)\n     40     #convert into world coordinates and compute offsets\n     41     for idx in range(len(preds)):\n---&gt; 42         preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n     43 \n     44     return loss, preds, confidences\n\n/kaggle/usr/lib/lyft_l5kit_unofficial_fix/l5kit/geometry/transform.py in transform_points(points, transf_matrix)\n     87         np.ndarray: array of shape (N,2) for 2D input points, or (N,3) points for 3D input points\n     88     \"\"\"\n---&gt; 89     assert len(points.shape) == len(transf_matrix.shape) == 2\n     90     assert transf_matrix.shape[0] == transf_matrix.shape[1]\n     91 \n\nAssertionError: \n</code></pre>\n<p>Does anyone know what is going on? Also I assume that with the fix I do not have to divide the error by 4 anymore? Any insight is appreciated. </p>\n<p>Thanks a lot for your help!</p>",
      "rawMarkdown": "So I tried to modify the *forward* function above for the new l5kit package (version 1.1.0) following the notebook [Lyft-baseline-09-02](https://www.kaggle.com/lucabergamini/lyft-baseline-09-02?scriptVersionId=43751473)\n```\ndef forward(data, model, device, criterion = pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    \n    # Forward pass\n    preds, confidences = model(inputs)    \n    \n    \n    loss = criterion(targets, preds, confidences, target_availabilities)\n    \n  \n    world_from_agents = data[\"world_from_agent\"].numpy()\n    centroids = data[\"centroid\"].numpy()\n    \n    #convert into world coordinates and compute offsets\n    for idx in range(len(preds)):\n        preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n    \n    return loss, preds, confidences\n```\nSo as I understand the target in the train set is fixed, however the test set is not so we still need to convert it? However when I ran the above code I encountered the following error:\n\n```\n---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n<ipython-input-24-9205635f7dc5> in <module>\n     12 for data in progress_bar:\n     13     inputs = data['image'].to(device)\n---> 14     _, preds, confidences  = forward(data, model, device)\n     15 \n     16 #     # convert agent coordinates into world offsets\n\n<ipython-input-22-276920502171> in forward(data, model, device, criterion)\n     40     #convert into world coordinates and compute offsets\n     41     for idx in range(len(preds)):\n---> 42         preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n     43 \n     44     return loss, preds, confidences\n\n/kaggle/usr/lib/lyft_l5kit_unofficial_fix/l5kit/geometry/transform.py in transform_points(points, transf_matrix)\n     87         np.ndarray: array of shape (N,2) for 2D input points, or (N,3) points for 3D input points\n     88     \"\"\"\n---> 89     assert len(points.shape) == len(transf_matrix.shape) == 2\n     90     assert transf_matrix.shape[0] == transf_matrix.shape[1]\n     91 \n\nAssertionError: \n```\nDoes anyone know what is going on? Also I assume that with the fix I do not have to divide the error by 4 anymore? Any insight is appreciated. \n\nThanks a lot for your help!",
      "votes": 2,
      "replies": [
        {
          "id": 1037423,
          "postDate": "2020-10-05T02:35:19.387Z",
          "content": "<p>The official example is only for single-mode prediction. For multi-mode I am doing something like this (when making predictions):</p>\n<pre><code>outputs = outputs.cpu().numpy()\nworld_from_agents = data[\"world_from_agent\"].numpy()\ncentroids = data[\"centroid\"].numpy()\ncoords_offset = []\n\nfor agent_coords, world_from_agent, centroid in zip(outputs, world_from_agents, centroids):\n    for i in range(3):\n        agent_coords[i, :, :] = transform_points(agent_coords[i, :, :], world_from_agent) - centroid[:2]\n    coords_offset.append(agent_coords)\n</code></pre>\n<p>You can use some matrix manipulation to avoid the loops so it could be a liitle bit faster but this seems to work fine for me.</p>",
          "rawMarkdown": "The official example is only for single-mode prediction. For multi-mode I am doing something like this (when making predictions):\n\n```\noutputs = outputs.cpu().numpy()\nworld_from_agents = data[\"world_from_agent\"].numpy()\ncentroids = data[\"centroid\"].numpy()\ncoords_offset = []\n\nfor agent_coords, world_from_agent, centroid in zip(outputs, world_from_agents, centroids):\n    for i in range(3):\n        agent_coords[i, :, :] = transform_points(agent_coords[i, :, :], world_from_agent) - centroid[:2]\n    coords_offset.append(agent_coords)\n```\nYou can use some matrix manipulation to avoid the loops so it could be a liitle bit faster but this seems to work fine for me.",
          "votes": 3
        },
        {
          "id": 1037644,
          "postDate": "2020-10-05T07:40:22.533Z",
          "content": "<p>We're considering to support batch matrix multiplication for that function in the future release. It's not anything difficult to do (numpy can handle it without any troubles ), we just didn't have enough time to implement and test it :( </p>",
          "rawMarkdown": "We're considering to support batch matrix multiplication for that function in the future release. It's not anything difficult to do (numpy can handle it without any troubles ), we just didn't have enough time to implement and test it :( ",
          "votes": 3,
          "replies": [
            {
              "id": 1038958,
              "postDate": "2020-10-06T07:42:41.653Z",
              "content": "<blockquote>\n  <p>We're considering to support batch matrix multiplication for that function in the future release. It's not anything difficult to do (numpy can handle it without any troubles ), we just didn't have enough time to implement and test it :(</p>\n</blockquote>\n<p>worthy of trying <a href=\"https://github.com/lyft/l5kit/pull/166\" target=\"_blank\">https://github.com/lyft/l5kit/pull/166</a></p>",
              "rawMarkdown": "> We're considering to support batch matrix multiplication for that function in the future release. It's not anything difficult to do (numpy can handle it without any troubles ), we just didn't have enough time to implement and test it :(\n\nworthy of trying https://github.com/lyft/l5kit/pull/166"
            }
          ]
        },
        {
          "id": 1037713,
          "postDate": "2020-10-05T09:01:43.140Z",
          "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> Is <code>future_coords_offsets_pd.append()</code> in your code <code>coords_offset.append(agent_coords)</code> or do you stack it at the end?</p>\n<p><strong>EDIT</strong>: seems to need <code>future_coords_offsets_pd.append(np.stack(coords_offset))</code></p>\n<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> I have tried the same yesterday, it is like mentioned because of single model. They also use different shapes e.g.<code>output = model(inputs).reshape(targets.shape)</code> </p>\n<p>I tried to change everything accordingly but I ended up using the mentioned solution from <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a><br>\nI will try a couple of things out and report back.</p>\n<p><strong>UPDATE</strong>: <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> Solution works best for me.</p>",
          "rawMarkdown": "@frankpanxj Is ` future_coords_offsets_pd.append()` in your code `coords_offset.append(agent_coords)` or do you stack it at the end?\n\n**EDIT**: seems to need `future_coords_offsets_pd.append(np.stack(coords_offset))`\n\n@huanvo I have tried the same yesterday, it is like mentioned because of single model. They also use different shapes e.g.`output = model(inputs).reshape(targets.shape)` \n\nI tried to change everything accordingly but I ended up using the mentioned solution from @zaharch\nI will try a couple of things out and report back.\n\n**UPDATE**: @frankpanxj Solution works best for me.",
          "votes": 1
        },
        {
          "id": 1037772,
          "postDate": "2020-10-05T10:19:25.590Z",
          "content": "<p>Thanks for the input <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> . Is coords_offset the same shape as the original agents_coords ? It can replace it when we pass it to the loss function to get the metric of the training ?</p>",
          "rawMarkdown": "Thanks for the input @frankpanxj . Is coords_offset the same shape as the original agents_coords ? It can replace it when we pass it to the loss function to get the metric of the training ?",
          "votes": 1
        },
        {
          "id": 1037822,
          "postDate": "2020-10-05T11:10:27.450Z",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> You are right. You need to stack the reults.</p>",
          "rawMarkdown": "@aliabdin1 You are right. You need to stack the reults."
        },
        {
          "id": 1037828,
          "postDate": "2020-10-05T11:14:48.693Z",
          "content": "<p><a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> the shape of outputs is (batch size)x(modes)x(time)x(2D coords), which is needed for the loss function by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> at <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence</a>.<br>\ncoords_offset here is just a list of the preds for every sample, after stacking it you will get a np array with the same shape as outputs. (I changed some the variable names above to make this clearer.)</p>",
          "rawMarkdown": "@vladvdv the shape of outputs is (batch size)x(modes)x(time)x(2D coords), which is needed for the loss function by @corochann at https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence.\ncoords_offset here is just a list of the preds for every sample, after stacking it you will get a np array with the same shape as outputs. (I changed some the variable names above to make this clearer.)",
          "votes": 2
        },
        {
          "id": 1038042,
          "postDate": "2020-10-05T13:57:27.737Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  this seems to resolve my error. However my prediction error got much worse in the new l5kit environment (old env 18.x vs new env 4xxx).<br>\nMaybe I misunderstood something here. I was using the fix by <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> before in the old environment and it seemed to work fine. </p>",
          "rawMarkdown": "Thanks @frankpanxj  this seems to resolve my error. However my prediction error got much worse in the new l5kit environment (old env 18.x vs new env 4xxx).\nMaybe I misunderstood something here. I was using the fix by @zaharch before in the old environment and it seemed to work fine. "
        },
        {
          "id": 1038140,
          "postDate": "2020-10-05T15:25:57.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> have similar issues. Did you find the problem ?<br>\nlater edit: I was doing the transform at training time also. From what I understand now, the transformation should be done only on inference !</p>",
          "rawMarkdown": "@huanvo have similar issues. Did you find the problem ?\nlater edit: I was doing the transform at training time also. From what I understand now, the transformation should be done only on inference !"
        },
        {
          "id": 1038180,
          "postDate": "2020-10-05T15:59:46.360Z",
          "content": "<p>The fix posted here is different to the fix in l5kit v1.1.0. In the fix here it not only corrects for the rotation error but also converts from metres to pixels. The fix in v1.1 only corrects the rotation, leaving things in metres. So models trained with the fix here aren't compatible with the fix in v1.1.<br>\nTo use a model trained with the fix here in v1.1 you'd need to convert the targets from metres to pixels and then transform the predictions from pixels to metres before applying the rotation fix from the l5kit docs.</p>",
          "rawMarkdown": "The fix posted here is different to the fix in l5kit v1.1.0. In the fix here it not only corrects for the rotation error but also converts from metres to pixels. The fix in v1.1 only corrects the rotation, leaving things in metres. So models trained with the fix here aren't compatible with the fix in v1.1.\nTo use a model trained with the fix here in v1.1 you'd need to convert the targets from metres to pixels and then transform the predictions from pixels to metres before applying the rotation fix from the l5kit docs.",
          "votes": 2
        },
        {
          "id": 1038496,
          "postDate": "2020-10-05T20:37:22.327Z",
          "content": "<p>I used the above <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> 's code to inference a model trained with L5kit ver1.1.0 but my submission file created only up to coordinates 0-29 i.e. coord_x29, coord_y29. </p>\n<p>Does anybody experience the same thing?</p>\n<p>I could not see what I am doing wrong in my code right now.</p>",
          "rawMarkdown": "I used the above @frankpanxj 's code to inference a model trained with L5kit ver1.1.0 but my submission file created only up to coordinates 0-29 i.e. coord_x29, coord_y29. \n\nDoes anybody experience the same thing?\n\n I could not see what I am doing wrong in my code right now."
        },
        {
          "id": 1038634,
          "postDate": "2020-10-05T23:20:19.687Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> , I tried zaharch's approach and got similar results as using v1.1. However, my score was never that good (my validation loss was always much higher than my training loss, did you experience the same?) so there certainly is something I am missing.</p>\n<p><a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> very interesting results. I did not have the same problem. There are coords beyond 29 in my results.</p>",
          "rawMarkdown": "Hi @huanvo , I tried zaharch's approach and got similar results as using v1.1. However, my score was never that good (my validation loss was always much higher than my training loss, did you experience the same?) so there certainly is something I am missing.\n\n@sheriytm very interesting results. I did not have the same problem. There are coords beyond 29 in my results.",
          "votes": 1
        },
        {
          "id": 1040128,
          "postDate": "2020-10-07T01:48:21.617Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>. I just found the bug that caused cuts out the rest of the coordinates during prediction.</p>",
          "rawMarkdown": "Thanks @frankpanxj. I just found the bug that caused cuts out the rest of the coordinates during prediction.",
          "votes": 1
        },
        {
          "id": 1041478,
          "postDate": "2020-10-07T19:04:16.717Z",
          "content": "<p>I'm a bit confused. My results seem to be completely non-sensical with the new version and the example code provided. The issue seems to be stemming from subtracting the centroid. My centroid is values in the thousands typically and shifting my predictions way too much. Anyone having similar experience or an explanation of what I might be doing wrong?</p>",
          "rawMarkdown": "I'm a bit confused. My results seem to be completely non-sensical with the new version and the example code provided. The issue seems to be stemming from subtracting the centroid. My centroid is values in the thousands typically and shifting my predictions way too much. Anyone having similar experience or an explanation of what I might be doing wrong?"
        },
        {
          "id": 1041480,
          "postDate": "2020-10-07T19:05:52.430Z",
          "content": "<p>Did you remove the here mentioned training snippet? Its not needed anymore.</p>",
          "rawMarkdown": "Did you remove the here mentioned training snippet? Its not needed anymore."
        },
        {
          "id": 1041493,
          "postDate": "2020-10-07T19:17:16.053Z",
          "content": "<p>Yes I am training on the raw outputs from the dataloader on the newest version of l5kit. No transforms anymore. Training loss looks fine but validation is messed up. in the hundreds of thousands using roughly the code that was shared here. My predictions should not be in the thousands, right?</p>",
          "rawMarkdown": "Yes I am training on the raw outputs from the dataloader on the newest version of l5kit. No transforms anymore. Training loss looks fine but validation is messed up. in the hundreds of thousands using roughly the code that was shared here. My predictions should not be in the thousands, right?"
        },
        {
          "id": 1041494,
          "postDate": "2020-10-07T19:18:10.640Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> , the new version works fine for me, I haven't noticed anything interesting changed after the switch. Please share the code.</p>",
          "rawMarkdown": "Hi @ryches , the new version works fine for me, I haven't noticed anything interesting changed after the switch. Please share the code."
        },
        {
          "id": 1041514,
          "postDate": "2020-10-07T19:34:51.987Z",
          "content": "<p>My outputs from my model are roughly what they were prior to the l5kit upgrade. Previously I was doing the transform and then inverse transform and it was working fine. </p>\n<p>Now with my model trained on the new l5kit targets without any transformations its max output on x, y coordinates is ~77 which seems normal, but centroid is showing up as 785, -2024. Validation predictions used to look like <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F38339d77fe0592dbb4b92bed316dfff3%2FScreenshot%20from%202020-10-07%2012-33-02.png?generation=1602099206020101&amp;alt=media\" alt=\"\"></p>\n<p>but now they look like this and the big difference seems to be these huge centroid outputs <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fb7d99f675f22689d08b7ad4a37d43907%2FScreenshot%20from%202020-10-07%2012-33-48.png?generation=1602099263121087&amp;alt=media\" alt=\"\"></p>\n<p>Is the l5kit version you are using from pip or directly from github?</p>",
          "rawMarkdown": "My outputs from my model are roughly what they were prior to the l5kit upgrade. Previously I was doing the transform and then inverse transform and it was working fine. \n\nNow with my model trained on the new l5kit targets without any transformations its max output on x, y coordinates is ~77 which seems normal, but centroid is showing up as 785, -2024. Validation predictions used to look like ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F38339d77fe0592dbb4b92bed316dfff3%2FScreenshot%20from%202020-10-07%2012-33-02.png?generation=1602099206020101&alt=media)\n\nbut now they look like this and the big difference seems to be these huge centroid outputs ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fb7d99f675f22689d08b7ad4a37d43907%2FScreenshot%20from%202020-10-07%2012-33-48.png?generation=1602099263121087&alt=media)\n\nIs the l5kit version you are using from pip or directly from github?"
        },
        {
          "id": 1041527,
          "postDate": "2020-10-07T19:42:52.290Z",
          "content": "<p>The code I have been poking at to try to understand. If I dont subtract the huge centroids then I still get questionable results </p>\n<pre><code>torch.set_grad_enabled(False)\nwith torch.no_grad():\n    model.eval()\n    # store information for evaluation\n    future_coords_offsets_pd = []\n    timestamps = []\n    agent_ids = []\n    confidences = []\n    progress_bar = tqdm(eval_dataloader)\n    for data in progress_bar:\n        inputs = data\n        targets = data[\"target_positions\"]\n        timestamps.append(data[\"timestamp\"].cpu().numpy())\n        agent_ids.append(data[\"track_id\"].cpu().numpy())\n        outputs, confs = model(inputs)\n#             outputs = inverse_tranform_predictions(outputs, data[\"world_to_image\"], data[\"centroid\"])\n#             future_coords_offsets_pd.append(outputs.cpu().numpy())\n        for agent_coords, world_from_agent, centroid in zip(outputs.cpu().numpy(), \n                                                            data[\"world_to_image\"].cpu().numpy(),\n                                                            data[\"centroid\"].cpu().numpy()):\n            for mode in range(3):\n                agent_coords[mode, :, :] = transform_points(agent_coords[mode, :, :], world_from_agent) - centroid[:2]\n            future_coords_offsets_pd.append(agent_coords)\n        confidences.append(confs.cpu().numpy())\n    pred_path = f\"{gettempdir()}/validation_preds.csv\"\n    write_pred_csv(pred_path,\n               timestamps=np.concatenate(timestamps),\n               track_ids=np.concatenate(agent_ids),\n               coords=np.stack(future_coords_offsets_pd),\n               confs = np.concatenate(confidences)\n              )\n    del confidences, future_coords_offsets_pd\n    metrics = compute_metrics_csv(eval_gt_path, pred_path, [neg_multi_log_likelihood, time_displace])\n    for metric_name, metric_mean in metrics.items():\n        print(metric_name, metric_mean)\n</code></pre>",
          "rawMarkdown": "The code I have been poking at to try to understand. If I dont subtract the huge centroids then I still get questionable results \n\n```python\ntorch.set_grad_enabled(False)\nwith torch.no_grad():\n    model.eval()\n    # store information for evaluation\n    future_coords_offsets_pd = []\n    timestamps = []\n    agent_ids = []\n    confidences = []\n    progress_bar = tqdm(eval_dataloader)\n    for data in progress_bar:\n        inputs = data\n        targets = data[\"target_positions\"]\n        timestamps.append(data[\"timestamp\"].cpu().numpy())\n        agent_ids.append(data[\"track_id\"].cpu().numpy())\n        outputs, confs = model(inputs)\n#             outputs = inverse_tranform_predictions(outputs, data[\"world_to_image\"], data[\"centroid\"])\n#             future_coords_offsets_pd.append(outputs.cpu().numpy())\n        for agent_coords, world_from_agent, centroid in zip(outputs.cpu().numpy(), \n                                                            data[\"world_to_image\"].cpu().numpy(),\n                                                            data[\"centroid\"].cpu().numpy()):\n            for mode in range(3):\n                agent_coords[mode, :, :] = transform_points(agent_coords[mode, :, :], world_from_agent) - centroid[:2]\n            future_coords_offsets_pd.append(agent_coords)\n        confidences.append(confs.cpu().numpy())\n    pred_path = f\"{gettempdir()}/validation_preds.csv\"\n    write_pred_csv(pred_path,\n               timestamps=np.concatenate(timestamps),\n               track_ids=np.concatenate(agent_ids),\n               coords=np.stack(future_coords_offsets_pd),\n               confs = np.concatenate(confidences)\n              )\n    del confidences, future_coords_offsets_pd\n    metrics = compute_metrics_csv(eval_gt_path, pred_path, [neg_multi_log_likelihood, time_displace])\n    for metric_name, metric_mean in metrics.items():\n        print(metric_name, metric_mean)\n```\n\n",
          "votes": 1
        },
        {
          "id": 1041528,
          "postDate": "2020-10-07T19:43:19.807Z",
          "content": "<p>this is how you need to transform for multimode:</p>\n<pre><code>    # convert agent coordinates into world offsets\n    pred = pred.cpu().numpy()\n    world_from_agents = data[\"world_from_agent\"].numpy()\n    centroids = data[\"centroid\"].numpy()\n    coords_offset = []\n\n    # convert into world coordinates and compute offsets\n    for idx in range(len(pred)):\n        for mode in range(3):\n            pred[idx, mode, :, :] = transform_points(pred[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]\n\n    confidences_list.append(confidences.cpu().numpy().copy())\n    pred_coords_list.append(pred.copy())\n    timestamps.append(data[\"timestamp\"].numpy().copy())\n    agent_ids.append(data[\"track_id\"].numpy().copy())\n</code></pre>\n<p>I see that you used a batched approach <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> , that could be the problem. I am not sure if that works for multi-mode. I am actually pretty sure that it doesn't work as the code for transform_points is expecting points in (Nx2) or (Nx3).</p>",
          "rawMarkdown": "this is how you need to transform for multimode:\n\n```\n    # convert agent coordinates into world offsets\n    pred = pred.cpu().numpy()\n    world_from_agents = data[\"world_from_agent\"].numpy()\n    centroids = data[\"centroid\"].numpy()\n    coords_offset = []\n\n    # convert into world coordinates and compute offsets\n    for idx in range(len(pred)):\n        for mode in range(3):\n            pred[idx, mode, :, :] = transform_points(pred[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]\n\n    confidences_list.append(confidences.cpu().numpy().copy())\n    pred_coords_list.append(pred.copy())\n    timestamps.append(data[\"timestamp\"].numpy().copy())\n    agent_ids.append(data[\"track_id\"].numpy().copy())\n```\n\nI see that you used a batched approach @ryches , that could be the problem. I am not sure if that works for multi-mode. I am actually pretty sure that it doesn't work as the code for transform_points is expecting points in (Nx2) or (Nx3).",
          "votes": 7
        },
        {
          "id": 1041539,
          "postDate": "2020-10-07T19:50:27.217Z",
          "content": "<p>Thanks, I will look at that and try it out</p>",
          "rawMarkdown": "Thanks, I will look at that and try it out"
        },
        {
          "id": 1041548,
          "postDate": "2020-10-07T20:03:06.530Z",
          "content": "<p>Maybe the problem is that you are using <code>data[\"world_to_image\"]</code>, instead of <code>world_from_agent</code></p>",
          "rawMarkdown": "Maybe the problem is that you are using `data[\"world_to_image\"]`, instead of `world_from_agent`",
          "votes": 1
        },
        {
          "id": 1041550,
          "postDate": "2020-10-07T20:03:50.860Z",
          "content": "<p>Looks like that works. Thank you so much. Devil is in the details</p>",
          "rawMarkdown": "Looks like that works. Thank you so much. Devil is in the details"
        },
        {
          "id": 1041559,
          "postDate": "2020-10-07T20:08:28.657Z",
          "content": "<p>Yeah on second review I think the key was using the right transformation matrix. I was using the same method of zipping across the three inputs for the batches as shown in the example notebook and then iterating 3 times for the various modes as others have mentioned in this thread. I think my dimensions and operations were correct but I was using the wrong matrix. </p>",
          "rawMarkdown": "Yeah on second review I think the key was using the right transformation matrix. I was using the same method of zipping across the three inputs for the batches as shown in the example notebook and then iterating 3 times for the various modes as others have mentioned in this thread. I think my dimensions and operations were correct but I was using the wrong matrix. ",
          "votes": 2
        },
        {
          "id": 1041564,
          "postDate": "2020-10-07T20:12:05.090Z",
          "content": "<p>Great to hear, i'll switch to batched transformation, too if that works ;)</p>",
          "rawMarkdown": "Great to hear, i'll switch to batched transformation, too if that works ;)"
        },
        {
          "id": 1041590,
          "postDate": "2020-10-07T20:36:18.320Z",
          "content": "<p>Probably not worth switching. I did not see any appreciable difference in speed</p>",
          "rawMarkdown": "Probably not worth switching. I did not see any appreciable difference in speed"
        }
      ]
    },
    {
      "id": 1030216,
      "postDate": "2020-09-28T13:34:14.877Z",
      "content": "<p>First <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> helped the heard think &amp; helped us and now you ( <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> ) are doing the same. I like to thank you both. </p>",
      "rawMarkdown": "First @corochann helped the heard think & helped us and now you ( @zaharch ) are doing the same. I like to thank you both. ",
      "votes": 2,
      "replies": [
        {
          "id": 1030748,
          "postDate": "2020-09-28T23:22:36.097Z",
          "content": "<p>Thanks for feedback &amp; mention :)<br>\nI couldn't notice this before <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> pointed out, even I go through the EDA in several notebooks…</p>",
          "rawMarkdown": "Thanks for feedback & mention :)\nI couldn't notice this before @zaharch pointed out, even I go through the EDA in several notebooks...",
          "votes": 2
        }
      ]
    },
    {
      "id": 1030059,
      "postDate": "2020-09-28T11:38:05.003Z",
      "content": "<p>hey  <a href=\"https://www.kaggle.com/nosound\" target=\"_blank\">@nosound</a>, thanks for sharing this code and these insights. I'm still noob at ML and DS, but I've been \"studying\" some of the newest codes at current competitions and i feel I can get what you mean by directionless input's. don't know how to approach directly to that matter but If i get something I'll share it =) thanks for it.</p>",
      "rawMarkdown": "hey  @nosound, thanks for sharing this code and these insights. I'm still noob at ML and DS, but I've been \"studying\" some of the newest codes at current competitions and i feel I can get what you mean by directionless input's. don't know how to approach directly to that matter but If i get something I'll share it =) thanks for it.",
      "votes": 2
    },
    {
      "id": 1028503,
      "postDate": "2020-09-26T23:18:52.250Z",
      "content": "<p>The fact the original model performed well without any knowledge about the agent orientation raises the question about how generalisable the models are, even with the orientation fix applied. Would be easy to fix by the position specific train/validation split.</p>\n<p>Would be great if hosts allocated the test set in the different geographic position (and not available until the models freeze) to check how well solutions perform in more realistic situation of roads model have not seen during training, but I'd expect it's too much to ask at this stage.</p>",
      "rawMarkdown": "The fact the original model performed well without any knowledge about the agent orientation raises the question about how generalisable the models are, even with the orientation fix applied. Would be easy to fix by the position specific train/validation split.\n\nWould be great if hosts allocated the test set in the different geographic position (and not available until the models freeze) to check how well solutions perform in more realistic situation of roads model have not seen during training, but I'd expect it's too much to ask at this stage.",
      "votes": 2,
      "replies": [
        {
          "id": 1028864,
          "postDate": "2020-09-27T09:08:52.447Z",
          "content": "<p>On the other hand if the interesting zone is limited to what we have and AVs do not drive out of this zone anyway, then it is OK to overfit on it. But you have a good point.</p>\n<p>If direction is possible to learn then the models potentially can learn average speeds on those roads and common routes taken by cars, like rare vs common turns.</p>",
          "rawMarkdown": "On the other hand if the interesting zone is limited to what we have and AVs do not drive out of this zone anyway, then it is OK to overfit on it. But you have a good point.\n\nIf direction is possible to learn then the models potentially can learn average speeds on those roads and common routes taken by cars, like rare vs common turns.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1027116,
      "postDate": "2020-09-25T21:04:05.090Z",
      "content": "<p>Fantastic post, thanks for sharing! Will implement and report results :)</p>",
      "rawMarkdown": "Fantastic post, thanks for sharing! Will implement and report results :)",
      "votes": 2,
      "replies": [
        {
          "id": 1029481,
          "postDate": "2020-09-27T19:16:51.673Z",
          "content": "<p>Can confirm this provided a great boost to models from 3x.x to 2x.x, and current score 18.x using this approach too. Thanks again to nosound for sharing. Super interested to better understand how we were doing so well before, very odd!</p>",
          "rawMarkdown": "Can confirm this provided a great boost to models from 3x.x to 2x.x, and current score 18.x using this approach too. Thanks again to nosound for sharing. Super interested to better understand how we were doing so well before, very odd!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1026626,
      "postDate": "2020-09-25T13:07:50.927Z",
      "content": "<p>Oh wow thanks for sharing this! It took me forever to bring the error down, and it is not even close to 23. On the other hand, it is still surprising how the model can learn without any directions. </p>",
      "rawMarkdown": "Oh wow thanks for sharing this! It took me forever to bring the error down, and it is not even close to 23. On the other hand, it is still surprising how the model can learn without any directions. ",
      "votes": 2,
      "replies": [
        {
          "id": 1051298,
          "postDate": "2020-10-16T11:09:36.493Z",
          "content": "<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> I don't understand about what are you talking about<br>\nbut do have a question, please<br>\nRecently you have posted a high scoring kernel.<br>\nis the forward function which mentioned by nosound on the discussion topic in that kernel</p>",
          "rawMarkdown": "@huanvo I don't understand about what are you talking about\nbut do have a question, please\nRecently you have posted a high scoring kernel.\nis the forward function which mentioned by nosound on the discussion topic in that kernel\n"
        }
      ]
    },
    {
      "id": 1025830,
      "postDate": "2020-09-24T19:38:37.040Z",
      "content": "<p>Would it be possible for you to create a plot of the transformed paths like you did for the originals? I'd be interested to see what that looks like</p>",
      "rawMarkdown": "Would it be possible for you to create a plot of the transformed paths like you did for the originals? I'd be interested to see what that looks like",
      "votes": 2,
      "replies": [
        {
          "id": 1025860,
          "postDate": "2020-09-24T20:04:23.183Z",
          "content": "<p>Will do when I get home.</p>\n<p>Update: added</p>",
          "rawMarkdown": "Will do when I get home.\n\nUpdate: added",
          "votes": 4
        }
      ]
    },
    {
      "id": 1051279,
      "postDate": "2020-10-16T10:35:24.620Z",
      "content": "<p>I have a few questions after a long time of discussion of different people around the world</p>\n<p><strong>Question</strong></p>\n<ol>\n<li><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> does this discussion topic code works without any error also compatible for L5Kit 1.1.0 (do I want to divide by 4 the error</li>\n<li>Also you have given the visualization code for visualizing the discussion topic plot</li>\n<li><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> <a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> all the solution you produced out here in this discussion topic</li>\n<li>the actual hidden solution for this competition is </li>\n</ol>\n<ul>\n<li>model </li>\n<li>hyperparameter(i feel so)</li>\n<li>preprocess of input (transformation (do feels a little)</li>\n</ul>",
      "rawMarkdown": "I have a few questions after a long time of discussion of different people around the world\n\n**Question**\n1. @zaharch does this discussion topic code works without any error also compatible for L5Kit 1.1.0 (do I want to divide by 4 the error\n2. Also you have given the visualization code for visualizing the discussion topic plot\n3. @ryches @huanvo all the solution you produced out here in this discussion topic\n4. the actual hidden solution for this competition is \n- model \n- hyperparameter(i feel so)\n- preprocess of input (transformation (do feels a little)\n ",
      "replies": [
        {
          "id": 1051753,
          "postDate": "2020-10-16T19:56:11.640Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> , I am glad you join the discussion!</p>\n<ol>\n<li>No, the coordinates have already been rotated in 1.1.0, this code is not needed. And no, no need to divide by 4 in the new version, because they do not work in pixels.</li>\n</ol>\n<p>not sure what to comment on your other points.</p>",
          "rawMarkdown": "Hi, @morizin , I am glad you join the discussion!\n\n1. No, the coordinates have already been rotated in 1.1.0, this code is not needed. And no, no need to divide by 4 in the new version, because they do not work in pixels.\n\nnot sure what to comment on your other points."
        }
      ]
    },
    {
      "id": 1039767,
      "postDate": "2020-10-06T19:21:11.993Z",
      "content": "<p>I'm really new to this community and this helped me a lot. Thanks!</p>",
      "rawMarkdown": "I'm really new to this community and this helped me a lot. Thanks!"
    },
    {
      "id": 1037514,
      "postDate": "2020-10-05T05:00:42.687Z",
      "content": "<p>Thanks for sharing! This was extremely useful.</p>\n<p>Wondering if the following achieves the same transformation:</p>\n<p><code>\ntransform_points(data['target_positions'], data['raster_from_agent'])\n</code></p>",
      "rawMarkdown": "Thanks for sharing! This was extremely useful.\n\nWondering if the following achieves the same transformation:\n\n`\ntransform_points(data['target_positions'], data['raster_from_agent'])\n`"
    },
    {
      "id": 1034040,
      "postDate": "2020-10-01T12:57:59.357Z",
      "content": "<p>Why is centroid added to target_positions?  (The centroid is the initial position of the agent in world coordinates?)</p>",
      "rawMarkdown": "Why is centroid added to target_positions?  (The centroid is the initial position of the agent in world coordinates?)",
      "replies": [
        {
          "id": 1034058,
          "postDate": "2020-10-01T13:11:20.757Z",
          "content": "<p>Instead of understanding this I recommend to move to the new l5kit version, it is different there. Go to the github page and read the coordinate system description and see the example notebooks.</p>",
          "rawMarkdown": "Instead of understanding this I recommend to move to the new l5kit version, it is different there. Go to the github page and read the coordinate system description and see the example notebooks.",
          "votes": 2
        },
        {
          "id": 1034119,
          "postDate": "2020-10-01T14:11:38.867Z",
          "content": "<p>Yeah, the process is different in the l5kit implementation but you do still need to add/subtract the centroid in some places. In both approaches this is because the targets are relative positions from the initial agent position while some of the transformation matrices need absolute positions. So adding the centroid to relative world coordinates converts to absolute world coordinates (and vice-versa subtracting it). Though as <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>  says the places where it's needed differ from the code here as some of the new transformation matrices include this translation.</p>\n<p>So with l5 kit 1.1 you see that:</p>\n<pre><code>&gt; transform_point(np.array([0,0]), data['world_from_agent'])\narray([  671.78094482, -2196.83374023])\n</code></pre>\n<p>is adding the centroid in and converting the rotation. So you then subtract the centroid from the result to get back to the initial relative world targets.</p>\n<p>To go back the other way from old world relative target coordinates to agent relative coordinates you add the centroid and then transform. So:</p>\n<pre><code>&gt;target_in_agent = np.array([0.1, 0.2])\n&gt;target_in_world = transform_point(target_in_agent, data['world_from_agent']) - data['centroid']\n&gt;print(\"Relative World coords: \", target_in_world)\n&gt;print(\"Recovered: \", transform_point(target_in_world + data['centroid'], data['agent_from_world']))\nWorld coords:  [-0.11271524  0.19311985]\nRecovered:  [0.1 0.2]\n</code></pre>\n<p>Note the centroid is that same in each case as the change simply involves how targets are rotated around this centroid. Previously they were rotated with respect to the Autonomous Vehicle (/ego), now they are rotated with respect to the actual agent vehicle you're predicting (which may or may not be the ego). In both cases they're relative to centroid.</p>\n<p>(I think, obviously all a little complex and certainly encourage you to check my work)</p>",
          "rawMarkdown": "Yeah, the process is different in the l5kit implementation but you do still need to add/subtract the centroid in some places. In both approaches this is because the targets are relative positions from the initial agent position while some of the transformation matrices need absolute positions. So adding the centroid to relative world coordinates converts to absolute world coordinates (and vice-versa subtracting it). Though as @zaharch  says the places where it's needed differ from the code here as some of the new transformation matrices include this translation.\n\nSo with l5 kit 1.1 you see that:\n```\n> transform_point(np.array([0,0]), data['world_from_agent'])\narray([  671.78094482, -2196.83374023])\n```\nis adding the centroid in and converting the rotation. So you then subtract the centroid from the result to get back to the initial relative world targets.\n\nTo go back the other way from old world relative target coordinates to agent relative coordinates you add the centroid and then transform. So:\n```\n>target_in_agent = np.array([0.1, 0.2])\n>target_in_world = transform_point(target_in_agent, data['world_from_agent']) - data['centroid']\n>print(\"Relative World coords: \", target_in_world)\n>print(\"Recovered: \", transform_point(target_in_world + data['centroid'], data['agent_from_world']))\nWorld coords:  [-0.11271524  0.19311985]\nRecovered:  [0.1 0.2]\n```\n\nNote the centroid is that same in each case as the change simply involves how targets are rotated around this centroid. Previously they were rotated with respect to the Autonomous Vehicle (/ego), now they are rotated with respect to the actual agent vehicle you're predicting (which may or may not be the ego). In both cases they're relative to centroid.\n\n(I think, obviously all a little complex and certainly encourage you to check my work)",
          "votes": 3
        }
      ]
    },
    {
      "id": 1031817,
      "postDate": "2020-09-29T17:44:04.043Z",
      "content": "<p>Thank you for sharing, It is very helpful  👍</p>",
      "rawMarkdown": "Thank you for sharing, It is very helpful  👍"
    },
    {
      "id": 1031758,
      "postDate": "2020-09-29T17:09:28.250Z",
      "content": "<p>Thank you for sharing your findings with the community! It helps the people who are new to the competition! :)</p>",
      "rawMarkdown": "Thank you for sharing your findings with the community! It helps the people who are new to the competition! :)"
    },
    {
      "id": 1031163,
      "postDate": "2020-09-29T09:08:12.123Z",
      "content": "<p>good job ! good</p>",
      "rawMarkdown": "good job ! good"
    },
    {
      "id": 1031127,
      "postDate": "2020-09-29T09:05:10.613Z",
      "content": "<p>Cool :)))))))</p>",
      "rawMarkdown": "Cool :)))))))"
    },
    {
      "id": 1030450,
      "postDate": "2020-09-28T17:16:27.573Z",
      "content": "<p>Great job I am here to win</p>",
      "rawMarkdown": "Great job I am here to win"
    },
    {
      "id": 1029934,
      "postDate": "2020-09-28T09:24:30.040Z",
      "content": "<p>really helpful</p>",
      "rawMarkdown": "really helpful"
    },
    {
      "id": 1029493,
      "postDate": "2020-09-27T19:26:48.257Z",
      "content": "<p>Are you using a custom loss function?</p>",
      "rawMarkdown": "Are you using a custom loss function?",
      "replies": [
        {
          "id": 1029545,
          "postDate": "2020-09-27T20:47:08.440Z",
          "content": "<p>I am using a <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">loss function from here</a>, take a look. </p>",
          "rawMarkdown": "I am using a [loss function from here](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence), take a look. ",
          "votes": 2
        },
        {
          "id": 1030032,
          "postDate": "2020-09-28T11:15:47.660Z",
          "content": "<p>Thank you, I was wondering first why you are passing 4 arguments to the loss function :D</p>",
          "rawMarkdown": "Thank you, I was wondering first why you are passing 4 arguments to the loss function :D",
          "votes": 1
        },
        {
          "id": 1030642,
          "postDate": "2020-09-28T19:39:50.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> May I ask, why do you calculate the loss first and afterwards you transform the pred?</p>\n<pre><code>loss = criterion(targets, pred, confidences, target_availabilities)\nloss = torch.mean(loss)\n\n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n</code></pre>\n<p>What do you do with the returned pred, do you use it for validation?</p>",
          "rawMarkdown": "@zaharch May I ask, why do you calculate the loss first and afterwards you transform the pred?\n\n```\nloss = criterion(targets, pred, confidences, target_availabilities)\nloss = torch.mean(loss)\n\n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n```\n\nWhat do you do with the returned pred, do you use it for validation?",
          "votes": 1
        },
        {
          "id": 1030653,
          "postDate": "2020-09-28T20:05:39.573Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> , yes you are right, specifically I am using them for independent loss calculation and also calculating other metrics, like rmse. Want to do it in the original coordinates. Afterwards this transformation is actually needed for the test, as we submit in original coordinates.</p>",
          "rawMarkdown": "Hi @aliabdin1 , yes you are right, specifically I am using them for independent loss calculation and also calculating other metrics, like rmse. Want to do it in the original coordinates. Afterwards this transformation is actually needed for the test, as we submit in original coordinates.",
          "votes": 1
        },
        {
          "id": 1030657,
          "postDate": "2020-09-28T20:11:16.337Z",
          "content": "<p>Thank you :)</p>",
          "rawMarkdown": "Thank you :)"
        }
      ]
    },
    {
      "id": 1028755,
      "postDate": "2020-09-27T07:07:17.407Z",
      "content": "<p>sure, i hope it will work</p>",
      "rawMarkdown": "sure, i hope it will work"
    },
    {
      "id": 1028499,
      "postDate": "2020-09-26T22:57:25.237Z",
      "content": "<p>It might be worth trying to anonymize the vehicles, Maybe draw them all with the same extent?</p>\n<p>Probably be faster to try on inference, it should still show how much the vehicle size contributes to accuracy</p>",
      "rawMarkdown": "It might be worth trying to anonymize the vehicles, Maybe draw them all with the same extent?\n\nProbably be faster to try on inference, it should still show how much the vehicle size contributes to accuracy"
    },
    {
      "id": 1028438,
      "postDate": "2020-09-26T21:04:58.093Z",
      "content": "<p>Big thank you for sharing this huge and relevant information!</p>",
      "rawMarkdown": "Big thank you for sharing this huge and relevant information!"
    },
    {
      "id": 1027963,
      "postDate": "2020-09-26T13:42:47.897Z",
      "content": "<p>Good job, continue like this ! BRAVO </p>",
      "rawMarkdown": "Good job, continue like this ! BRAVO "
    },
    {
      "id": 1026606,
      "postDate": "2020-09-25T12:46:48.603Z",
      "content": "<p>I am wondering if drawing the images always with x axis pointing in the direction of travel will put too much bias on the predictions.</p>",
      "rawMarkdown": "I am wondering if drawing the images always with x axis pointing in the direction of travel will put too much bias on the predictions."
    },
    {
      "id": 1032440,
      "postDate": "2020-09-30T07:50:47.030Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1026099,
      "postDate": "2020-09-25T04:42:40.040Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1039704,
      "postDate": "2020-10-06T18:26:12.170Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 1035802,
      "postDate": "2020-10-03T05:35:21.447Z",
      "content": "<p>thanks for sharing this information</p>",
      "rawMarkdown": "thanks for sharing this information"
    },
    {
      "id": 1030352,
      "postDate": "2020-09-28T15:37:31.927Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1029508,
      "postDate": "2020-09-27T19:39:32.697Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1028482,
      "postDate": "2020-09-26T22:21:22.550Z",
      "content": "<p>Great job guys, thanks for sharing!</p>",
      "rawMarkdown": "Great job guys, thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 1026361,
      "author_name": "Luca Bergamini",
      "author_url": "",
      "post_date": "2020-09-25T08:49:31.527000",
      "content": "<p>Hey Everyone,</p>\n<blockquote>\n  <p>What do you think? I have a lingering feeling that I am missing something obvious!</p>\n</blockquote>\n<p>On the contrary, you're spot on :) This is something that slipped through our tests and that we only recently became aware of.</p>\n<p>Using pixel coordinates is absolutely legit here, but <strong>we're also working on a long term fix in L5Kit</strong> we hope we can ship next week. Once that is thoroughly checked we will sync with Kaggle to update the default package.</p>\n<p>I'll start a post with more info about this issue and how we plan to fix it to increase visibility and awareness.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1030060,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-09-28T11:39:01.493000",
          "content": "<p>any updates?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1030368,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-28T15:58:08.520000",
          "content": "<p>almost ready to merge <a href=\"https://github.com/lyft/l5kit/pull/150\" target=\"_blank\">https://github.com/lyft/l5kit/pull/150</a> :) After that we will cut a new release and update Kaggle, doc and baseline</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1030406,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-09-28T16:41:34.090000",
          "content": "<p>I compared that PR to the fix posted here and looks like you're keeping the target positions in metres whereas above code converts to pixels, otherwise they match (barring slight inaccuracy likely due to floating point). That's the intended final result correct? (No issue, I think metres is better, just checking)</p>\n<p>I gather also that you'll be updating the Kaggle ground truths so predictions will be in the new coordinates. Is that correct?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1030418,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-28T16:52:25.093000",
          "content": "<p>I think pixels are better, - in pixels they can be directly connected in creative ways to the images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030428,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-28T16:58:47.383000",
          "content": "<blockquote>\n  <p>That's the intended final result correct?</p>\n</blockquote>\n<p>yes, we're keeping metres (basically the only difference is a 2D rotation) :)</p>\n<blockquote>\n  <p>I gather also that you'll be updating the Kaggle ground truths so predictions will be in the new coordinates. Is that correct?</p>\n</blockquote>\n<p>We're not, this means you will have to convert them back into offsets in world coordinates (but we'll update the baseline notebook to show how to do that and we're adding a lot of matrices to make it easier). Updating the ground truth would have required a lot of additional testing, this was the best compromise :( </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030433,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-09-28T17:02:57.873000",
          "content": "<blockquote>\n  <p>I think pixels are better</p>\n</blockquote>\n<p>How so? Either way you can just multiply/divide by pixel size right? I think if you want that you probably want to do it within a single model and then convert back to metres at the end. Otherwise if doing something like combining outputs from multiple models with different pixel sizes you'd have to be tracking the pixel sizes of all models.<br>\nWhat if you aren't even using a pixel based model, why convert to some arbitrary pizel size?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030441,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-09-28T17:09:37.190000",
          "content": "<blockquote>\n  <p>We're not, this means you will have to convert them back into offsets in world coordinates (but we'll update the baseline notebook to show how to do that and we're adding a lot of matrices to make it easier). Updating the ground truth would have required a lot of additional testing, this was the best compromise :( </p>\n</blockquote>\n<p>Ah OK, I misunderstood the changes to the GT.csv code as meaning they were being updated to new coords when I guess they're converting new coords back to the old ones. Presumably they will match expected outputs on Kaggle.<br>\nI understand the compromise, obviously tricky either way.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030447,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-28T17:13:58.037000",
          "content": "<blockquote>\n  <p>I guess they're converting new coords back to the old ones. Presumably they will match expected outputs on Kaggle</p>\n</blockquote>\n<p>That's is exactly what's happening there :) I've checked and they match the old ones (also I've submitted a new baseline trained using the latest master and results are much better so I'm positive we didn't screw everything up completely)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1030458,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-09-28T17:20:15.487000",
          "content": "<p>Probably belongs elsewhere but in relation to the GCP credit prize for those beating the baseline does that mean they need to beat new baseline? Think the first deadline for that was tomorrow so not a lot of time to update models.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1030480,
          "author_name": "Huan Vo",
          "author_url": "",
          "post_date": "2020-09-28T17:29:02.487000",
          "content": "<p>Where is the new baseline result? I could not find it. Or if someone can just report the new score here? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1030498,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-09-28T17:35:05.717000",
          "content": "<p>82.246<br>\n<a href=\"https://www.kaggle.com/lucabergamini/lyft-baseline-09-02\" target=\"_blank\">https://www.kaggle.com/lucabergamini/lyft-baseline-09-02</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030499,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-28T17:35:15.993000",
          "content": "<blockquote>\n  <p>How so? Either way you can just multiply/divide by pixel size right? I think if you want that you probably want to do it within a single model and then convert back to metres at the end. Otherwise if doing something like combining outputs from multiple models with different pixel sizes you'd have to be tracking the pixel sizes of all models.<br>\n  What if you aren't even using a pixel based model, why convert to some arbitrary pizel size?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> , you have good points, hard to argue with what you say. In defense of pixels, - we get <code>data</code> object in which I would naturally expect everything to be in the same coordinate system. There is the image object, with its own coordinate system, so everything needs to match. My second argument is that if everything is in pixels then we can easily draw on the image without any additional transformations, rasterize those pixels into something etc.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1025660,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2020-09-24T17:35:20.237000",
      "content": "<p>Did you calculate <code>bias = torch.tensor([56.25, 112.5])</code> from the <code>ego_center</code> config? So the input size is 225x225px?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1025677,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-24T17:47:58.253000",
          "content": "<p>That was my question as well. I assume this will not be the correct alignment at other resolutions. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1025694,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-24T17:58:03.900000",
          "content": "<p>Maybe this..</p>\n<pre><code>rs = cfg[\"raster_params\"][\"raster_size\"]\nec = cfg[\"raster_params\"][\"ego_center\"]\n\nbias = torch.tensor([rs[0] * ec[0], rs[1] * ec[1]])[None, None, :].to(device)\n</code></pre>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1025730,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-24T18:21:17.327000",
          "content": "<p>Exactly right, thank you! The point of this is of course to have non-moving target at zero, as it was before. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1029906,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "2020-09-28T08:37:46.400000",
      "content": "<p>Thank you for sharing this. We were doing it completely wrong, and most people would have been oblivious to that fact had you not so generously highlighted it. The epitome of community - kudos.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1027976,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2020-09-26T13:47:09.237000",
      "content": "<p>Thank you for sharing. <br>\nWith this approach, in my case, the model(45.0) trained on all of the datasets was improved to the model(37.7) trained on 1/4 of the dataset.<br>\nHowever, there is a gap between validation metric (13.1) and public LB. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1028012,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-26T14:27:01.327000",
          "content": "<p>Your validation metric is too small, this can be because you divide the errors by 2 <strong>after</strong> the predictions are transformed. Basically check that the errors are divided by 2 only when needed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1028488,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2020-09-26T22:29:26.957000",
          "content": "<p>Sorry. When I checked it, I used the square error, so I divided it by 4.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1026441,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-09-25T10:22:50.063000",
      "content": "<p>Good approach <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> . Congrats !<br>\nCan I ask you why do you choose to divide the error by 4 ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1026516,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-25T11:29:26.090000",
          "content": "<p>In case one uses default configuration there is a scaling factor <code>'pixel_size': [0.5, 0.5]</code>. So there are two pixels (new coordinates) per meter (old coordinates). To get the error back into the old scale I need to divide it by 2, or by 4 in case of squared error. </p>\n<p>Depending on your model it may not be necessary to scale, you will just have 4x bigger loss. But I train a multi-scenario output, and there is a subtle inter-play between the errors and the confidences. Therefore I want to keep it at the old scale, which is the correct one for the loss.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1025888,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2020-09-24T21:06:16.230000",
      "content": "<p>We really did it all wrong! Thanks for sharing and pointing it out. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1031260,
      "author_name": "M. Esad Tutar",
      "author_url": "",
      "post_date": "2020-09-29T10:54:47.223000",
      "content": "<p>I'm really new to this community and this helped me a lot. Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1028934,
      "author_name": "Shashank Madan",
      "author_url": "",
      "post_date": "2020-09-27T10:36:53.357000",
      "content": "<p>Hi, Great insight! Can i know how you obtained those graphs?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1028945,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-27T10:44:40.177000",
          "content": "<p>Sure, the first graph is this code:</p>\n<pre><code>fig = plt.figure(figsize = (12,12))\nfor k in range(500):\n    data = valid_dataset[4001+1000*k]\n    tgt = data[\"target_positions\"]\n    plt.plot(tgt[:,0],tgt[:,1])\n</code></pre>\n<p>the second is similar but with the transformation, depends on how you shape your code.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1028966,
          "author_name": "Shashank Madan",
          "author_url": "",
          "post_date": "2020-09-27T11:04:52.797000",
          "content": "<p>Thanks a Lot! By the looks of it… the valid_dataset suggests you are using an evaluation dataset… why? just curious…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1029308,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-27T16:34:59.013000",
          "content": "<p>No reason, it doesn't really matter because it is just targets from the data. I don't think there is any meaningful difference between the train and the validation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1029763,
          "author_name": "Shashank Madan",
          "author_url": "",
          "post_date": "2020-09-28T05:48:28.323000",
          "content": "<p>Understood! thanks again😀</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1027575,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2020-09-26T08:43:53.657000",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> does this mean while inference we have to transform the predicted trajectories using the code block you shared</p>\n<pre><code>    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n</code></pre>\n<p>Trying to understand if I am interpreting the targets correctly</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1027577,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-26T08:44:51.447000",
          "content": "<p>Yes, in inference you will have to use the same code block.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1027613,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-26T08:56:53.187000",
          "content": "<p>Yes, that is right, same for inference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1027628,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-09-26T09:06:41.810000",
          "content": "<p>Interesting, I trained model for 6 epochs 10k steps each my train loss was around ~60.x but the inference kernel with transformation gave really bad result. Anything obvious I might be missing?</p>\n<p>Note that I am using raster rize 224x224 so I changed the bias value.</p>\n<p>Inference code for reference</p>\n<pre><code>        inputs = data[\"image\"].to(device)\n        #target_availabilities = data[\"target_availabilities\"].unsqueeze(-1).to(device)\n        target_availabilities = data[\"target_availabilities\"].to(device)\n        targets = data[\"target_positions\"].to(device)\n        matrix = data[\"world_to_image\"].to(device)\n        centroid = data[\"centroid\"].to(device)[:,None,:].to(torch.float)\n\n        bs,tl,_ = targets.shape\n        rs = cfg[\"raster_params\"][\"raster_size\"]\n        ec = cfg[\"raster_params\"][\"ego_center\"]\n        bias = torch.tensor([rs[0] * ec[0], rs[1] * ec[1]])[None, None, :].to(device)\n\n\n        #outputs = model(inputs).reshape(targets.shape)\n        pred, confidences = model(inputs)\n\n        if cfg['test_params']['image_coords']:\n            matrix_inv = torch.inverse(matrix)\n            pred = pred + bias[:,None,:,:]\n            pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n            pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                                for i in range(3)], dim=1)\n            pred = pred.transpose(2,3)[:,:,:,:2]\n            pred = pred - centroid[:,None,:,:]\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1027636,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-26T09:13:08.097000",
          "content": "<p>Have you changed the bias to 56,112? How do you judge that the result is bad? Do prediction start from close to zero? I guess it is some kind of a technical error.</p>\n<p>Do you do softmax for the confidences and pred shape transformations inside the model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1027654,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-09-26T09:31:26.883000",
          "content": "<p><code>Have you changed the bias to 56,112?</code><br>\nYes,<br>\n<code>Do you do softmax for the confidences and pred shape transformations inside the model?</code><br>\nYes<br>\n<code>Do prediction start from close to zero?</code><br>\nYes<br>\n<code>How do you judge that the result is bad?</code><br>\nI am comparing model performance against my old model without transformation after n epochs. Model without transformations seemed to have converged faster in my case. I guess I'll dig deeper into this by comparing my results</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1025678,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-09-24T17:48:07.827000",
      "content": "<p>Awesome point. Thank you for sharing this</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1064052,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-10-29T16:30:18.540000",
      "content": "<p>important!!!! </p>\n<ol>\n<li>check if you can get the same values if convert from world to image coordinates, and then back again.</li>\n<li>i think you need a 64-bit float (double) to do this</li>\n</ol>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1037369,
      "author_name": "Huan Vo",
      "author_url": "",
      "post_date": "2020-10-04T23:21:41.780000",
      "content": "<p>So I tried to modify the <em>forward</em> function above for the new l5kit package (version 1.1.0) following the notebook <a href=\"https://www.kaggle.com/lucabergamini/lyft-baseline-09-02?scriptVersionId=43751473\" target=\"_blank\">Lyft-baseline-09-02</a></p>\n<pre><code>def forward(data, model, device, criterion = pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n\n    # Forward pass\n    preds, confidences = model(inputs)    \n\n\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n\n    world_from_agents = data[\"world_from_agent\"].numpy()\n    centroids = data[\"centroid\"].numpy()\n\n    #convert into world coordinates and compute offsets\n    for idx in range(len(preds)):\n        preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n\n    return loss, preds, confidences\n</code></pre>\n<p>So as I understand the target in the train set is fixed, however the test set is not so we still need to convert it? However when I ran the above code I encountered the following error:</p>\n<pre><code>---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n&lt;ipython-input-24-9205635f7dc5&gt; in &lt;module&gt;\n     12 for data in progress_bar:\n     13     inputs = data['image'].to(device)\n---&gt; 14     _, preds, confidences  = forward(data, model, device)\n     15 \n     16 #     # convert agent coordinates into world offsets\n\n&lt;ipython-input-22-276920502171&gt; in forward(data, model, device, criterion)\n     40     #convert into world coordinates and compute offsets\n     41     for idx in range(len(preds)):\n---&gt; 42         preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n     43 \n     44     return loss, preds, confidences\n\n/kaggle/usr/lib/lyft_l5kit_unofficial_fix/l5kit/geometry/transform.py in transform_points(points, transf_matrix)\n     87         np.ndarray: array of shape (N,2) for 2D input points, or (N,3) points for 3D input points\n     88     \"\"\"\n---&gt; 89     assert len(points.shape) == len(transf_matrix.shape) == 2\n     90     assert transf_matrix.shape[0] == transf_matrix.shape[1]\n     91 \n\nAssertionError: \n</code></pre>\n<p>Does anyone know what is going on? Also I assume that with the fix I do not have to divide the error by 4 anymore? Any insight is appreciated. </p>\n<p>Thanks a lot for your help!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1037423,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-10-05T02:35:19.387000",
          "content": "<p>The official example is only for single-mode prediction. For multi-mode I am doing something like this (when making predictions):</p>\n<pre><code>outputs = outputs.cpu().numpy()\nworld_from_agents = data[\"world_from_agent\"].numpy()\ncentroids = data[\"centroid\"].numpy()\ncoords_offset = []\n\nfor agent_coords, world_from_agent, centroid in zip(outputs, world_from_agents, centroids):\n    for i in range(3):\n        agent_coords[i, :, :] = transform_points(agent_coords[i, :, :], world_from_agent) - centroid[:2]\n    coords_offset.append(agent_coords)\n</code></pre>\n<p>You can use some matrix manipulation to avoid the loops so it could be a liitle bit faster but this seems to work fine for me.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1037644,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-10-05T07:40:22.533000",
          "content": "<p>We're considering to support batch matrix multiplication for that function in the future release. It's not anything difficult to do (numpy can handle it without any troubles ), we just didn't have enough time to implement and test it :( </p>",
          "votes": 3,
          "replies": [
            {
              "id": 1038958,
              "author_name": "Jiayu",
              "author_url": "",
              "post_date": "2020-10-06T07:42:41.653000",
              "content": "<blockquote>\n  <p>We're considering to support batch matrix multiplication for that function in the future release. It's not anything difficult to do (numpy can handle it without any troubles ), we just didn't have enough time to implement and test it :(</p>\n</blockquote>\n<p>worthy of trying <a href=\"https://github.com/lyft/l5kit/pull/166\" target=\"_blank\">https://github.com/lyft/l5kit/pull/166</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1037713,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-05T09:01:43.140000",
          "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> Is <code>future_coords_offsets_pd.append()</code> in your code <code>coords_offset.append(agent_coords)</code> or do you stack it at the end?</p>\n<p><strong>EDIT</strong>: seems to need <code>future_coords_offsets_pd.append(np.stack(coords_offset))</code></p>\n<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> I have tried the same yesterday, it is like mentioned because of single model. They also use different shapes e.g.<code>output = model(inputs).reshape(targets.shape)</code> </p>\n<p>I tried to change everything accordingly but I ended up using the mentioned solution from <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a><br>\nI will try a couple of things out and report back.</p>\n<p><strong>UPDATE</strong>: <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> Solution works best for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1037772,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-10-05T10:19:25.590000",
          "content": "<p>Thanks for the input <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> . Is coords_offset the same shape as the original agents_coords ? It can replace it when we pass it to the loss function to get the metric of the training ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1037822,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-10-05T11:10:27.450000",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> You are right. You need to stack the reults.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1037828,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-10-05T11:14:48.693000",
          "content": "<p><a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> the shape of outputs is (batch size)x(modes)x(time)x(2D coords), which is needed for the loss function by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> at <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence</a>.<br>\ncoords_offset here is just a list of the preds for every sample, after stacking it you will get a np array with the same shape as outputs. (I changed some the variable names above to make this clearer.)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1038042,
          "author_name": "Huan Vo",
          "author_url": "",
          "post_date": "2020-10-05T13:57:27.737000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  this seems to resolve my error. However my prediction error got much worse in the new l5kit environment (old env 18.x vs new env 4xxx).<br>\nMaybe I misunderstood something here. I was using the fix by <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> before in the old environment and it seemed to work fine. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1038140,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-10-05T15:25:57.460000",
          "content": "<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> have similar issues. Did you find the problem ?<br>\nlater edit: I was doing the transform at training time also. From what I understand now, the transformation should be done only on inference !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1038180,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-10-05T15:59:46.360000",
          "content": "<p>The fix posted here is different to the fix in l5kit v1.1.0. In the fix here it not only corrects for the rotation error but also converts from metres to pixels. The fix in v1.1 only corrects the rotation, leaving things in metres. So models trained with the fix here aren't compatible with the fix in v1.1.<br>\nTo use a model trained with the fix here in v1.1 you'd need to convert the targets from metres to pixels and then transform the predictions from pixels to metres before applying the rotation fix from the l5kit docs.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1038496,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-10-05T20:37:22.327000",
          "content": "<p>I used the above <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> 's code to inference a model trained with L5kit ver1.1.0 but my submission file created only up to coordinates 0-29 i.e. coord_x29, coord_y29. </p>\n<p>Does anybody experience the same thing?</p>\n<p>I could not see what I am doing wrong in my code right now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1038634,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-10-05T23:20:19.687000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> , I tried zaharch's approach and got similar results as using v1.1. However, my score was never that good (my validation loss was always much higher than my training loss, did you experience the same?) so there certainly is something I am missing.</p>\n<p><a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> very interesting results. I did not have the same problem. There are coords beyond 29 in my results.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1040128,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-10-07T01:48:21.617000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>. I just found the bug that caused cuts out the rest of the coordinates during prediction.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1041478,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T19:04:16.717000",
          "content": "<p>I'm a bit confused. My results seem to be completely non-sensical with the new version and the example code provided. The issue seems to be stemming from subtracting the centroid. My centroid is values in the thousands typically and shifting my predictions way too much. Anyone having similar experience or an explanation of what I might be doing wrong?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041480,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-07T19:05:52.430000",
          "content": "<p>Did you remove the here mentioned training snippet? Its not needed anymore.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041493,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T19:17:16.053000",
          "content": "<p>Yes I am training on the raw outputs from the dataloader on the newest version of l5kit. No transforms anymore. Training loss looks fine but validation is messed up. in the hundreds of thousands using roughly the code that was shared here. My predictions should not be in the thousands, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041494,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-10-07T19:18:10.640000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> , the new version works fine for me, I haven't noticed anything interesting changed after the switch. Please share the code.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041514,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T19:34:51.987000",
          "content": "<p>My outputs from my model are roughly what they were prior to the l5kit upgrade. Previously I was doing the transform and then inverse transform and it was working fine. </p>\n<p>Now with my model trained on the new l5kit targets without any transformations its max output on x, y coordinates is ~77 which seems normal, but centroid is showing up as 785, -2024. Validation predictions used to look like <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F38339d77fe0592dbb4b92bed316dfff3%2FScreenshot%20from%202020-10-07%2012-33-02.png?generation=1602099206020101&amp;alt=media\" alt=\"\"></p>\n<p>but now they look like this and the big difference seems to be these huge centroid outputs <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fb7d99f675f22689d08b7ad4a37d43907%2FScreenshot%20from%202020-10-07%2012-33-48.png?generation=1602099263121087&amp;alt=media\" alt=\"\"></p>\n<p>Is the l5kit version you are using from pip or directly from github?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041527,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T19:42:52.290000",
          "content": "<p>The code I have been poking at to try to understand. If I dont subtract the huge centroids then I still get questionable results </p>\n<pre><code>torch.set_grad_enabled(False)\nwith torch.no_grad():\n    model.eval()\n    # store information for evaluation\n    future_coords_offsets_pd = []\n    timestamps = []\n    agent_ids = []\n    confidences = []\n    progress_bar = tqdm(eval_dataloader)\n    for data in progress_bar:\n        inputs = data\n        targets = data[\"target_positions\"]\n        timestamps.append(data[\"timestamp\"].cpu().numpy())\n        agent_ids.append(data[\"track_id\"].cpu().numpy())\n        outputs, confs = model(inputs)\n#             outputs = inverse_tranform_predictions(outputs, data[\"world_to_image\"], data[\"centroid\"])\n#             future_coords_offsets_pd.append(outputs.cpu().numpy())\n        for agent_coords, world_from_agent, centroid in zip(outputs.cpu().numpy(), \n                                                            data[\"world_to_image\"].cpu().numpy(),\n                                                            data[\"centroid\"].cpu().numpy()):\n            for mode in range(3):\n                agent_coords[mode, :, :] = transform_points(agent_coords[mode, :, :], world_from_agent) - centroid[:2]\n            future_coords_offsets_pd.append(agent_coords)\n        confidences.append(confs.cpu().numpy())\n    pred_path = f\"{gettempdir()}/validation_preds.csv\"\n    write_pred_csv(pred_path,\n               timestamps=np.concatenate(timestamps),\n               track_ids=np.concatenate(agent_ids),\n               coords=np.stack(future_coords_offsets_pd),\n               confs = np.concatenate(confidences)\n              )\n    del confidences, future_coords_offsets_pd\n    metrics = compute_metrics_csv(eval_gt_path, pred_path, [neg_multi_log_likelihood, time_displace])\n    for metric_name, metric_mean in metrics.items():\n        print(metric_name, metric_mean)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1041528,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-07T19:43:19.807000",
          "content": "<p>this is how you need to transform for multimode:</p>\n<pre><code>    # convert agent coordinates into world offsets\n    pred = pred.cpu().numpy()\n    world_from_agents = data[\"world_from_agent\"].numpy()\n    centroids = data[\"centroid\"].numpy()\n    coords_offset = []\n\n    # convert into world coordinates and compute offsets\n    for idx in range(len(pred)):\n        for mode in range(3):\n            pred[idx, mode, :, :] = transform_points(pred[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]\n\n    confidences_list.append(confidences.cpu().numpy().copy())\n    pred_coords_list.append(pred.copy())\n    timestamps.append(data[\"timestamp\"].numpy().copy())\n    agent_ids.append(data[\"track_id\"].numpy().copy())\n</code></pre>\n<p>I see that you used a batched approach <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> , that could be the problem. I am not sure if that works for multi-mode. I am actually pretty sure that it doesn't work as the code for transform_points is expecting points in (Nx2) or (Nx3).</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1041539,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T19:50:27.217000",
          "content": "<p>Thanks, I will look at that and try it out</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041548,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-10-07T20:03:06.530000",
          "content": "<p>Maybe the problem is that you are using <code>data[\"world_to_image\"]</code>, instead of <code>world_from_agent</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1041550,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T20:03:50.860000",
          "content": "<p>Looks like that works. Thank you so much. Devil is in the details</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041559,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T20:08:28.657000",
          "content": "<p>Yeah on second review I think the key was using the right transformation matrix. I was using the same method of zipping across the three inputs for the batches as shown in the example notebook and then iterating 3 times for the various modes as others have mentioned in this thread. I think my dimensions and operations were correct but I was using the wrong matrix. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1041564,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-07T20:12:05.090000",
          "content": "<p>Great to hear, i'll switch to batched transformation, too if that works ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041590,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-07T20:36:18.320000",
          "content": "<p>Probably not worth switching. I did not see any appreciable difference in speed</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1030216,
      "author_name": "The Brown Iceman",
      "author_url": "",
      "post_date": "2020-09-28T13:34:14.877000",
      "content": "<p>First <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> helped the heard think &amp; helped us and now you ( <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> ) are doing the same. I like to thank you both. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1030748,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-09-28T23:22:36.097000",
          "content": "<p>Thanks for feedback &amp; mention :)<br>\nI couldn't notice this before <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> pointed out, even I go through the EDA in several notebooks…</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1030059,
      "author_name": "Luis Ramírez",
      "author_url": "",
      "post_date": "2020-09-28T11:38:05.003000",
      "content": "<p>hey  <a href=\"https://www.kaggle.com/nosound\" target=\"_blank\">@nosound</a>, thanks for sharing this code and these insights. I'm still noob at ML and DS, but I've been \"studying\" some of the newest codes at current competitions and i feel I can get what you mean by directionless input's. don't know how to approach directly to that matter but If i get something I'll share it =) thanks for it.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1028503,
      "author_name": "Dmytro Poplavskiy",
      "author_url": "",
      "post_date": "2020-09-26T23:18:52.250000",
      "content": "<p>The fact the original model performed well without any knowledge about the agent orientation raises the question about how generalisable the models are, even with the orientation fix applied. Would be easy to fix by the position specific train/validation split.</p>\n<p>Would be great if hosts allocated the test set in the different geographic position (and not available until the models freeze) to check how well solutions perform in more realistic situation of roads model have not seen during training, but I'd expect it's too much to ask at this stage.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1028864,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-27T09:08:52.447000",
          "content": "<p>On the other hand if the interesting zone is limited to what we have and AVs do not drive out of this zone anyway, then it is OK to overfit on it. But you have a good point.</p>\n<p>If direction is possible to learn then the models potentially can learn average speeds on those roads and common routes taken by cars, like rare vs common turns.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1027116,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2020-09-25T21:04:05.090000",
      "content": "<p>Fantastic post, thanks for sharing! Will implement and report results :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1029481,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2020-09-27T19:16:51.673000",
          "content": "<p>Can confirm this provided a great boost to models from 3x.x to 2x.x, and current score 18.x using this approach too. Thanks again to nosound for sharing. Super interested to better understand how we were doing so well before, very odd!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1026626,
      "author_name": "Huan Vo",
      "author_url": "",
      "post_date": "2020-09-25T13:07:50.927000",
      "content": "<p>Oh wow thanks for sharing this! It took me forever to bring the error down, and it is not even close to 23. On the other hand, it is still surprising how the model can learn without any directions. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1051298,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-16T11:09:36.493000",
          "content": "<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> I don't understand about what are you talking about<br>\nbut do have a question, please<br>\nRecently you have posted a high scoring kernel.<br>\nis the forward function which mentioned by nosound on the discussion topic in that kernel</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1025830,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-09-24T19:38:37.040000",
      "content": "<p>Would it be possible for you to create a plot of the transformed paths like you did for the originals? I'd be interested to see what that looks like</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1025860,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-24T20:04:23.183000",
          "content": "<p>Will do when I get home.</p>\n<p>Update: added</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1051279,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-16T10:35:24.620000",
      "content": "<p>I have a few questions after a long time of discussion of different people around the world</p>\n<p><strong>Question</strong></p>\n<ol>\n<li><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> does this discussion topic code works without any error also compatible for L5Kit 1.1.0 (do I want to divide by 4 the error</li>\n<li>Also you have given the visualization code for visualizing the discussion topic plot</li>\n<li><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> <a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> all the solution you produced out here in this discussion topic</li>\n<li>the actual hidden solution for this competition is </li>\n</ol>\n<ul>\n<li>model </li>\n<li>hyperparameter(i feel so)</li>\n<li>preprocess of input (transformation (do feels a little)</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 1051753,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-10-16T19:56:11.640000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> , I am glad you join the discussion!</p>\n<ol>\n<li>No, the coordinates have already been rotated in 1.1.0, this code is not needed. And no, no need to divide by 4 in the new version, because they do not work in pixels.</li>\n</ol>\n<p>not sure what to comment on your other points.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1039767,
      "author_name": "Rahul Kumar",
      "author_url": "",
      "post_date": "2020-10-06T19:21:11.993000",
      "content": "<p>I'm really new to this community and this helped me a lot. Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1037514,
      "author_name": "Srihari Radhakrishna",
      "author_url": "",
      "post_date": "2020-10-05T05:00:42.687000",
      "content": "<p>Thanks for sharing! This was extremely useful.</p>\n<p>Wondering if the following achieves the same transformation:</p>\n<p><code>\ntransform_points(data['target_positions'], data['raster_from_agent'])\n</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1034040,
      "author_name": "Sidhant",
      "author_url": "",
      "post_date": "2020-10-01T12:57:59.357000",
      "content": "<p>Why is centroid added to target_positions?  (The centroid is the initial position of the agent in world coordinates?)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1034058,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-10-01T13:11:20.757000",
          "content": "<p>Instead of understanding this I recommend to move to the new l5kit version, it is different there. Go to the github page and read the coordinate system description and see the example notebooks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1034119,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-10-01T14:11:38.867000",
          "content": "<p>Yeah, the process is different in the l5kit implementation but you do still need to add/subtract the centroid in some places. In both approaches this is because the targets are relative positions from the initial agent position while some of the transformation matrices need absolute positions. So adding the centroid to relative world coordinates converts to absolute world coordinates (and vice-versa subtracting it). Though as <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>  says the places where it's needed differ from the code here as some of the new transformation matrices include this translation.</p>\n<p>So with l5 kit 1.1 you see that:</p>\n<pre><code>&gt; transform_point(np.array([0,0]), data['world_from_agent'])\narray([  671.78094482, -2196.83374023])\n</code></pre>\n<p>is adding the centroid in and converting the rotation. So you then subtract the centroid from the result to get back to the initial relative world targets.</p>\n<p>To go back the other way from old world relative target coordinates to agent relative coordinates you add the centroid and then transform. So:</p>\n<pre><code>&gt;target_in_agent = np.array([0.1, 0.2])\n&gt;target_in_world = transform_point(target_in_agent, data['world_from_agent']) - data['centroid']\n&gt;print(\"Relative World coords: \", target_in_world)\n&gt;print(\"Recovered: \", transform_point(target_in_world + data['centroid'], data['agent_from_world']))\nWorld coords:  [-0.11271524  0.19311985]\nRecovered:  [0.1 0.2]\n</code></pre>\n<p>Note the centroid is that same in each case as the change simply involves how targets are rotated around this centroid. Previously they were rotated with respect to the Autonomous Vehicle (/ego), now they are rotated with respect to the actual agent vehicle you're predicting (which may or may not be the ego). In both cases they're relative to centroid.</p>\n<p>(I think, obviously all a little complex and certainly encourage you to check my work)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1031817,
      "author_name": "sandeep",
      "author_url": "",
      "post_date": "2020-09-29T17:44:04.043000",
      "content": "<p>Thank you for sharing, It is very helpful  👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1031758,
      "author_name": "Tarun Paparaju",
      "author_url": "",
      "post_date": "2020-09-29T17:09:28.250000",
      "content": "<p>Thank you for sharing your findings with the community! It helps the people who are new to the competition! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1031163,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2020-09-29T09:08:12.123000",
      "content": "<p>good job ! good</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1031127,
      "author_name": "Saeed Zou",
      "author_url": "",
      "post_date": "2020-09-29T09:05:10.613000",
      "content": "<p>Cool :)))))))</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1030450,
      "author_name": "Akhil Bhalerao",
      "author_url": "",
      "post_date": "2020-09-28T17:16:27.573000",
      "content": "<p>Great job I am here to win</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1029934,
      "author_name": "Vinayak M S",
      "author_url": "",
      "post_date": "2020-09-28T09:24:30.040000",
      "content": "<p>really helpful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1029493,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-09-27T19:26:48.257000",
      "content": "<p>Are you using a custom loss function?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1029545,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-27T20:47:08.440000",
          "content": "<p>I am using a <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">loss function from here</a>, take a look. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1030032,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-09-28T11:15:47.660000",
          "content": "<p>Thank you, I was wondering first why you are passing 4 arguments to the loss function :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030642,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-09-28T19:39:50.927000",
          "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> May I ask, why do you calculate the loss first and afterwards you transform the pred?</p>\n<pre><code>loss = criterion(targets, pred, confidences, target_availabilities)\nloss = torch.mean(loss)\n\n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n</code></pre>\n<p>What do you do with the returned pred, do you use it for validation?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030653,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-28T20:05:39.573000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> , yes you are right, specifically I am using them for independent loss calculation and also calculating other metrics, like rmse. Want to do it in the original coordinates. Afterwards this transformation is actually needed for the test, as we submit in original coordinates.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030657,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-09-28T20:11:16.337000",
          "content": "<p>Thank you :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1028755,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-27T07:07:17.407000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1028499,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-26T22:57:25.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1028438,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-26T21:04:58.093000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1027963,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-26T13:42:47.897000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1026606,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-25T12:46:48.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1032440,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-30T07:50:47.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1026099,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-25T04:42:40.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1039704,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-06T18:26:12.170000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1035802,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-03T05:35:21.447000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1030352,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-28T15:37:31.927000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1029508,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-27T19:39:32.697000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1028482,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-26T22:21:22.550000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1025600": "Hi Kagglers, I came across something important that I am excited to share. Despite my LB score improving from `32.210` to `25.524` after that fix with using `3x` less training epochs to get there, I am still not 100% sure that I do not miss something simple and obvious to everyone but me. But here it is for your judgement.\n\n# Statement\n\nWe have input images of the following kind, aligned for our agent vehicle to look right and at `'ego_center': [0.25, 0.5]`. It is the only input to our simple NNs that everyone uses.\n\n![typical input image](https://i.imgur.com/uQYCT7u.png)\n\nBut then we predict targets which are not aligned, at least if you follow [the example notebook](https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb) code. But the public notebooks that I saw use the same approach, if I draw several hundreds of target trajectories that we try to predict then it looks like this:\n\n![targets in world coordinates](https://i.imgur.com/gxPjOAn.png)\n\nIt is in all different directions, as expected. Then I spent a day trying to understand how is that possible to predict correct directions from direction-less input. To no avail, I still do not understand it. But the predictions are so accurate, they are actually very good! The direction is somehow predicted correctly from a direction-less input.\n\nAfter failing to understand why everyone does it wrong and why it even works, I tried to implement it correctly, that is, **transform the targets to image coordinates** before training. I attach the pytorch code with the transformations below so you can try it yourself fast. If you try and get an improvement please report it here.\n\nTransformed target trajectories look like this, this is 160 randomly sampled targets\n\n![targets in image coordinates](https://i.imgur.com/NZVGGkT.png)\n\nAnd it worked like a charm. With only 15 epochs, 10k batches each, batch size 16, which is about 5h training on my machine, I improved my training score from `30.87` to `23.79` and public LB `32.210` to `25.524`. \n\nSo how is the old approach able to predict the direction? I have two theories but no prove:\n1. Number of streets is very limited, and the network is able to learn what street has what orientation and then apply it to predictions.\n2. The image does somehow contain the rotation transformation in itself, maybe some subtle pixel level traces of it are there.\n\nWhat do you think? I have a lingering feeling that I am missing something obvious!\n\n# Code\n\nBelow boolean flag `cfg['train_params']['image_coords']` controls if I use the new approach. I apply the needed transformations on targets and then do inverse transformation on predictions in reverse order. Note that my predictions are multi-scenario there, you need to get rid of this additional dimensions if you use single scenario.\n\n```\ndef forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    matrix = data[\"world_to_image\"].to(device)\n    centroid = data[\"centroid\"].to(device)[:,None,:].to(torch.float)\n    \n    # Forward pass\n    outputs = model(inputs)\n    \n    bs,tl,_ = targets.shape\n    assert tl == cfg[\"model_params\"][\"future_num_frames\"]\n    \n    if cfg['train_params']['image_coords']:\n        targets = targets + centroid\n        targets = torch.cat([targets,torch.ones((bs,tl,1)).to(device)], dim=2)\n        targets = torch.matmul(matrix.to(torch.float), targets.transpose(1,2))\n        targets = targets.transpose(1,2)[:,:,:2]\n        bias = torch.tensor([56.25, 112.5])[None,None,:].to(device)\n        targets = targets - bias\n    \n    confidences, pred = outputs[:,:3], outputs[:,3:]\n    pred = pred.view(bs, 3, tl, 2)\n    assert confidences.shape == (bs, 3)\n    confidences = torch.softmax(confidences, dim=1)\n    \n    loss = criterion(targets, pred, confidences, target_availabilities)\n    loss = torch.mean(loss)\n    \n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n    \n    return loss, pred, confidences\n```\n\nand in your loss function you may want to divide the coordinates by 2:\n\n```\n    ...\n    error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    if cfg['train_params']['image_coords']:\n        error = error / 4\n    ...\n```",
    "1026361": "Hey Everyone,\n> What do you think? I have a lingering feeling that I am missing something obvious!\n\nOn the contrary, you're spot on :) This is something that slipped through our tests and that we only recently became aware of.\n\nUsing pixel coordinates is absolutely legit here, but **we're also working on a long term fix in L5Kit** we hope we can ship next week. Once that is thoroughly checked we will sync with Kaggle to update the default package.\n\nI'll start a post with more info about this issue and how we plan to fix it to increase visibility and awareness.",
    "1025660": "Did you calculate `bias = torch.tensor([56.25, 112.5])` from the `ego_center` config? So the input size is 225x225px?",
    "1029906": "Thank you for sharing this. We were doing it completely wrong, and most people would have been oblivious to that fact had you not so generously highlighted it. The epitome of community - kudos.",
    "1027976": "Thank you for sharing. \nWith this approach, in my case, the model(45.0) trained on all of the datasets was improved to the model(37.7) trained on 1/4 of the dataset.\nHowever, there is a gap between validation metric (13.1) and public LB. ",
    "1026441": "Good approach @zaharch . Congrats !\nCan I ask you why do you choose to divide the error by 4 ?",
    "1025888": "We really did it all wrong! Thanks for sharing and pointing it out. ",
    "1031260": "I'm really new to this community and this helped me a lot. Thanks!",
    "1028934": "Hi, Great insight! Can i know how you obtained those graphs?",
    "1027575": "@zaharch does this mean while inference we have to transform the predicted trajectories using the code block you shared\n```\n    if cfg['train_params']['image_coords']:\n        matrix_inv = torch.inverse(matrix)\n        pred = pred + bias[:,None,:,:]\n        pred = torch.cat([pred,torch.ones((bs,3,tl,1)).to(device)], dim=3)\n        pred = torch.stack([torch.matmul(matrix_inv.to(torch.float), pred[:,i].transpose(1,2)) \n                            for i in range(3)], dim=1)\n        pred = pred.transpose(2,3)[:,:,:,:2]\n        pred = pred - centroid[:,None,:,:]\n```\nTrying to understand if I am interpreting the targets correctly",
    "1025678": "Awesome point. Thank you for sharing this",
    "1064052": "important!!!! \n1.  check if you can get the same values if convert from world to image coordinates, and then back again.\n2. i think you need a 64-bit float (double) to do this",
    "1037369": "So I tried to modify the *forward* function above for the new l5kit package (version 1.1.0) following the notebook [Lyft-baseline-09-02](https://www.kaggle.com/lucabergamini/lyft-baseline-09-02?scriptVersionId=43751473)\n```\ndef forward(data, model, device, criterion = pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    \n    # Forward pass\n    preds, confidences = model(inputs)    \n    \n    \n    loss = criterion(targets, preds, confidences, target_availabilities)\n    \n  \n    world_from_agents = data[\"world_from_agent\"].numpy()\n    centroids = data[\"centroid\"].numpy()\n    \n    #convert into world coordinates and compute offsets\n    for idx in range(len(preds)):\n        preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n    \n    return loss, preds, confidences\n```\nSo as I understand the target in the train set is fixed, however the test set is not so we still need to convert it? However when I ran the above code I encountered the following error:\n\n```\n---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n<ipython-input-24-9205635f7dc5> in <module>\n     12 for data in progress_bar:\n     13     inputs = data['image'].to(device)\n---> 14     _, preds, confidences  = forward(data, model, device)\n     15 \n     16 #     # convert agent coordinates into world offsets\n\n<ipython-input-22-276920502171> in forward(data, model, device, criterion)\n     40     #convert into world coordinates and compute offsets\n     41     for idx in range(len(preds)):\n---> 42         preds[idx] = transform_points(preds[idx], world_from_agents[idx]) - centroids[idx]\n     43 \n     44     return loss, preds, confidences\n\n/kaggle/usr/lib/lyft_l5kit_unofficial_fix/l5kit/geometry/transform.py in transform_points(points, transf_matrix)\n     87         np.ndarray: array of shape (N,2) for 2D input points, or (N,3) points for 3D input points\n     88     \"\"\"\n---> 89     assert len(points.shape) == len(transf_matrix.shape) == 2\n     90     assert transf_matrix.shape[0] == transf_matrix.shape[1]\n     91 \n\nAssertionError: \n```\nDoes anyone know what is going on? Also I assume that with the fix I do not have to divide the error by 4 anymore? Any insight is appreciated. \n\nThanks a lot for your help!",
    "1030216": "First @corochann helped the heard think & helped us and now you ( @zaharch ) are doing the same. I like to thank you both. ",
    "1030059": "hey  @nosound, thanks for sharing this code and these insights. I'm still noob at ML and DS, but I've been \"studying\" some of the newest codes at current competitions and i feel I can get what you mean by directionless input's. don't know how to approach directly to that matter but If i get something I'll share it =) thanks for it.",
    "1028503": "The fact the original model performed well without any knowledge about the agent orientation raises the question about how generalisable the models are, even with the orientation fix applied. Would be easy to fix by the position specific train/validation split.\n\nWould be great if hosts allocated the test set in the different geographic position (and not available until the models freeze) to check how well solutions perform in more realistic situation of roads model have not seen during training, but I'd expect it's too much to ask at this stage.",
    "1027116": "Fantastic post, thanks for sharing! Will implement and report results :)",
    "1026626": "Oh wow thanks for sharing this! It took me forever to bring the error down, and it is not even close to 23. On the other hand, it is still surprising how the model can learn without any directions. ",
    "1025830": "Would it be possible for you to create a plot of the transformed paths like you did for the originals? I'd be interested to see what that looks like",
    "1051279": "I have a few questions after a long time of discussion of different people around the world\n\n**Question**\n1. @zaharch does this discussion topic code works without any error also compatible for L5Kit 1.1.0 (do I want to divide by 4 the error\n2. Also you have given the visualization code for visualizing the discussion topic plot\n3. @ryches @huanvo all the solution you produced out here in this discussion topic\n4. the actual hidden solution for this competition is \n- model \n- hyperparameter(i feel so)\n- preprocess of input (transformation (do feels a little)\n ",
    "1039767": "I'm really new to this community and this helped me a lot. Thanks!",
    "1037514": "Thanks for sharing! This was extremely useful.\n\nWondering if the following achieves the same transformation:\n\n`\ntransform_points(data['target_positions'], data['raster_from_agent'])\n`",
    "1034040": "Why is centroid added to target_positions?  (The centroid is the initial position of the agent in world coordinates?)",
    "1031817": "Thank you for sharing, It is very helpful  👍",
    "1031758": "Thank you for sharing your findings with the community! It helps the people who are new to the competition! :)",
    "1031163": "good job ! good",
    "1031127": "Cool :)))))))",
    "1030450": "Great job I am here to win",
    "1029934": "really helpful",
    "1029493": "Are you using a custom loss function?",
    "1028755": "sure, i hope it will work",
    "1028499": "It might be worth trying to anonymize the vehicles, Maybe draw them all with the same extent?\n\nProbably be faster to try on inference, it should still show how much the vehicle size contributes to accuracy",
    "1028438": "Big thank you for sharing this huge and relevant information!",
    "1027963": "Good job, continue like this ! BRAVO ",
    "1026606": "I am wondering if drawing the images always with x axis pointing in the direction of travel will put too much bias on the predictions.",
    "1032440": "",
    "1026099": "",
    "1039704": "Thank you!",
    "1035802": "thanks for sharing this information",
    "1030352": "Thanks for sharing",
    "1029508": "Thanks for sharing!",
    "1028482": "Great job guys, thanks for sharing!"
  }
}