{
  "id": 197186,
  "title": "Calculate loss in image coordinates",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/197186",
  "author_name": "",
  "post_date": "2020-11-14T21:52:19.179373800Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I read the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492\" target=\"_blank\">We did it all wrong</a> topic, and my thought was, that my network could give more accurate results if I calculate the loss in the image coordinates. I saw that the code was written for an older version of l5kit, and now initially the <code>data[\"target_positions\"]</code> has a better format ([1]), so I am not able to use it the latest l5kit 1.1.0.<br>\nAlthough drawing the default [1] and the image coordinates [2] trajectories, I saw, that they have the same shape, but the second has a larger resolution.</p>\n<p>Based on this evidence, I tried to modify the forward function. It is obvious, that in the normal case, I can calculate the position transformation with the <code>transform_points(data['target_positions'], data['raster_from_agent'])</code> function. Although this function doesn't work in batch mode, so I can't use it for the training.</p>\n<p>This is my pseudo-code for the forward function:</p>\n<pre><code>def forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    # Forward pass\n    preds, confidences = model(inputs)\n\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = transform_points_BATCH(data['target_positions'], data['raster_from_agent']).to(device)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n    matrix_inv = torch.inverse(data['raster_from_agent'])\n    preds = transform_points_BATCH(preds, matrix_inv).to(device)\n\n    return loss, preds, confidences\n</code></pre>\n<p>I saw the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193858\" target=\"_blank\">Redundancy of the <code>transform_points</code> in the l5kit package for prediction?</a> post as well, where it solved the pred -&gt; word position transformation with this code, but I was not able to fit for my need.</p>\n<pre><code>preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n</code></pre>\n<p>My questions are:</p>\n<ol>\n<li>How should the <code>transform_points_BATCH</code> function look like?</li>\n<li>Is it works if I only calculate the inverse <code>data['raster_from_agent']</code> matrix to transform back the predicitions?</li>\n</ol>\n<p>[1] and [2] are in the attached image.</p>",
  "messages": [
    {
      "id": "1078498",
      "postDate": "11/14/2020 21:52:19",
      "content": "<p>I read the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492\" target=\"_blank\">We did it all wrong</a> topic, and my thought was, that my network could give more accurate results if I calculate the loss in the image coordinates. I saw that the code was written for an older version of l5kit, and now initially the <code>data[\"target_positions\"]</code> has a better format ([1]), so I am not able to use it the latest l5kit 1.1.0.<br>\nAlthough drawing the default [1] and the image coordinates [2] trajectories, I saw, that they have the same shape, but the second has a larger resolution.</p>\n<p>Based on this evidence, I tried to modify the forward function. It is obvious, that in the normal case, I can calculate the position transformation with the <code>transform_points(data['target_positions'], data['raster_from_agent'])</code> function. Although this function doesn't work in batch mode, so I can't use it for the training.</p>\n<p>This is my pseudo-code for the forward function:</p>\n<pre><code>def forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    # Forward pass\n    preds, confidences = model(inputs)\n\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = transform_points_BATCH(data['target_positions'], data['raster_from_agent']).to(device)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n    matrix_inv = torch.inverse(data['raster_from_agent'])\n    preds = transform_points_BATCH(preds, matrix_inv).to(device)\n\n    return loss, preds, confidences\n</code></pre>\n<p>I saw the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193858\" target=\"_blank\">Redundancy of the <code>transform_points</code> in the l5kit package for prediction?</a> post as well, where it solved the pred -&gt; word position transformation with this code, but I was not able to fit for my need.</p>\n<pre><code>preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n</code></pre>\n<p>My questions are:</p>\n<ol>\n<li>How should the <code>transform_points_BATCH</code> function look like?</li>\n<li>Is it works if I only calculate the inverse <code>data['raster_from_agent']</code> matrix to transform back the predicitions?</li>\n</ol>\n<p>[1] and [2] are in the attached image.</p>",
      "rawMarkdown": "I read the [We did it all wrong](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492) topic, and my thought was, that my network could give more accurate results if I calculate the loss in the image coordinates. I saw that the code was written for an older version of l5kit, and now initially the `data[\"target_positions\"]` has a better format ([1]), so I am not able to use it the latest l5kit 1.1.0.\nAlthough drawing the default [1] and the image coordinates [2] trajectories, I saw, that they have the same shape, but the second has a larger resolution.\n\nBased on this evidence, I tried to modify the forward function. It is obvious, that in the normal case, I can calculate the position transformation with the `transform_points(data['target_positions'], data['raster_from_agent'])` function. Although this function doesn't work in batch mode, so I can't use it for the training.\n\nThis is my pseudo-code for the forward function:\n```python\ndef forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    # Forward pass\n    preds, confidences = model(inputs)\n    \n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = transform_points_BATCH(data['target_positions'], data['raster_from_agent']).to(device)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n    matrix_inv = torch.inverse(data['raster_from_agent'])\n    preds = transform_points_BATCH(preds, matrix_inv).to(device)\n\n    return loss, preds, confidences\n```\n\nI saw the [Redundancy of the `transform_points` in the l5kit package for prediction?](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193858) post as well, where it solved the pred -> word position transformation with this code, but I was not able to fit for my need.\n\n```\npreds = torch.einsum('bmti,bji->bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n```\n\nMy questions are:\n1. How should the `transform_points_BATCH` function look like?\n2. Is it works if I only calculate the inverse `data['raster_from_agent']` matrix to transform back the predicitions?\n\n[1] and [2] are in the attached image.",
      "votes": null
    },
    {
      "id": "1078502",
      "postDate": "11/14/2020 22:14:57",
      "content": "<p>There is no rotation in <code>data['raster_from_agent']</code>, just scaling. Specifically multiplying both axes by 2 by default. Why wouldn't you do it?</p>",
      "rawMarkdown": "There is no rotation in `data['raster_from_agent']`, just scaling. Specifically multiplying both axes by 2 by default. Why wouldn't you do it?",
      "votes": null
    },
    {
      "id": "1081067",
      "postDate": "11/16/2020 19:16:25",
      "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>! <br>\nThis my solution right now. Is this over overcomplicated?<br>\nFrom scratch with 224x224 images, it goes fast down until the score LB score 24.x (around 400k iteration), but after that gets stuck there. Do you have any idea what can it cause?</p>\n<pre><code>cfg['img_coord_en'] = True\n\ndef forward(data, model, device, criterion=pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n\n    if cfg['img_coord_en']:\n        points = targets.to(torch.float64)\n        transf_matrix = (data['raster_from_agent']).to(torch.float64)\n\n        transf_matrix_T = transf_matrix.transpose(1,2)[:,:-1, :-1].to(device).to(torch.float64)\n        targets = torch.matmul(points, transf_matrix_T).to(torch.float64)\n\n        inv_transf_matrix =  torch.inverse(transf_matrix).to(torch.float64)\n\n    preds, confidences = model(inputs)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n    if cfg['img_coord_en']:\n        preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n            inv_transf_matrix.to(device)[:, :2, :2])\n\n    return loss, preds, confidences\n</code></pre>",
      "rawMarkdown": "Thanks for the answer @zaharch! \nThis my solution right now. Is this over overcomplicated?\nFrom scratch with 224x224 images, it goes fast down until the score LB score 24.x (around 400k iteration), but after that gets stuck there. Do you have any idea what can it cause?\n\n```python\ncfg['img_coord_en'] = True\n\ndef forward(data, model, device, criterion=pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    \n    if cfg['img_coord_en']:\n        points = targets.to(torch.float64)\n        transf_matrix = (data['raster_from_agent']).to(torch.float64)\n\n        transf_matrix_T = transf_matrix.transpose(1,2)[:,:-1, :-1].to(device).to(torch.float64)\n        targets = torch.matmul(points, transf_matrix_T).to(torch.float64)\n    \n        inv_transf_matrix =  torch.inverse(transf_matrix).to(torch.float64)\n\n    preds, confidences = model(inputs)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n    \n    if cfg['img_coord_en']:\n        preds = torch.einsum('bmti,bji->bmtj', preds.double(), \n            inv_transf_matrix.to(device)[:, :2, :2])\n\n    return loss, preds, confidences\n```",
      "votes": null
    },
    {
      "id": "1081077",
      "postDate": "11/16/2020 19:28:35",
      "content": "<p>It is over-complicated, as it is easier to multiply/divide by 2. But should work, good that you found the torch batch matrix multiplication. Further improvements are definitely outside of this code.</p>",
      "rawMarkdown": "It is over-complicated, as it is easier to multiply/divide by 2. But should work, good that you found the torch batch matrix multiplication. Further improvements are definitely outside of this code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1078502,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "11/14/2020 22:14:57",
      "content": "<p>There is no rotation in <code>data['raster_from_agent']</code>, just scaling. Specifically multiplying both axes by 2 by default. Why wouldn't you do it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1081067,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "11/16/2020 19:16:25",
          "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>! <br>\nThis my solution right now. Is this over overcomplicated?<br>\nFrom scratch with 224x224 images, it goes fast down until the score LB score 24.x (around 400k iteration), but after that gets stuck there. Do you have any idea what can it cause?</p>\n<pre><code>cfg['img_coord_en'] = True\n\ndef forward(data, model, device, criterion=pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n\n    if cfg['img_coord_en']:\n        points = targets.to(torch.float64)\n        transf_matrix = (data['raster_from_agent']).to(torch.float64)\n\n        transf_matrix_T = transf_matrix.transpose(1,2)[:,:-1, :-1].to(device).to(torch.float64)\n        targets = torch.matmul(points, transf_matrix_T).to(torch.float64)\n\n        inv_transf_matrix =  torch.inverse(transf_matrix).to(torch.float64)\n\n    preds, confidences = model(inputs)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n    if cfg['img_coord_en']:\n        preds = torch.einsum('bmti,bji-&gt;bmtj', preds.double(), \n            inv_transf_matrix.to(device)[:, :2, :2])\n\n    return loss, preds, confidences\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1081077,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/16/2020 19:28:35",
          "content": "<p>It is over-complicated, as it is easier to multiply/divide by 2. But should work, good that you found the torch batch matrix multiplication. Further improvements are definitely outside of this code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1078498": "I read the [We did it all wrong](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492) topic, and my thought was, that my network could give more accurate results if I calculate the loss in the image coordinates. I saw that the code was written for an older version of l5kit, and now initially the `data[\"target_positions\"]` has a better format ([1]), so I am not able to use it the latest l5kit 1.1.0.\nAlthough drawing the default [1] and the image coordinates [2] trajectories, I saw, that they have the same shape, but the second has a larger resolution.\n\nBased on this evidence, I tried to modify the forward function. It is obvious, that in the normal case, I can calculate the position transformation with the `transform_points(data['target_positions'], data['raster_from_agent'])` function. Although this function doesn't work in batch mode, so I can't use it for the training.\n\nThis is my pseudo-code for the forward function:\n```python\ndef forward(data, model, device, criterion):\n    inputs = data[\"image\"].to(device)\n    # Forward pass\n    preds, confidences = model(inputs)\n    \n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = transform_points_BATCH(data['target_positions'], data['raster_from_agent']).to(device)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n\n    matrix_inv = torch.inverse(data['raster_from_agent'])\n    preds = transform_points_BATCH(preds, matrix_inv).to(device)\n\n    return loss, preds, confidences\n```\n\nI saw the [Redundancy of the `transform_points` in the l5kit package for prediction?](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193858) post as well, where it solved the pred -> word position transformation with this code, but I was not able to fit for my need.\n\n```\npreds = torch.einsum('bmti,bji->bmtj', preds.double(), \n        data[\"world_from_agent\"].to(device)[:, :2, :2]).cpu().numpy()\n```\n\nMy questions are:\n1. How should the `transform_points_BATCH` function look like?\n2. Is it works if I only calculate the inverse `data['raster_from_agent']` matrix to transform back the predicitions?\n\n[1] and [2] are in the attached image.",
    "1078502": "There is no rotation in `data['raster_from_agent']`, just scaling. Specifically multiplying both axes by 2 by default. Why wouldn't you do it?",
    "1081067": "Thanks for the answer @zaharch! \nThis my solution right now. Is this over overcomplicated?\nFrom scratch with 224x224 images, it goes fast down until the score LB score 24.x (around 400k iteration), but after that gets stuck there. Do you have any idea what can it cause?\n\n```python\ncfg['img_coord_en'] = True\n\ndef forward(data, model, device, criterion=pytorch_neg_multi_log_likelihood_batch):\n    inputs = data[\"image\"].to(device)\n    target_availabilities = data[\"target_availabilities\"].to(device)\n    targets = data[\"target_positions\"].to(device)\n    \n    if cfg['img_coord_en']:\n        points = targets.to(torch.float64)\n        transf_matrix = (data['raster_from_agent']).to(torch.float64)\n\n        transf_matrix_T = transf_matrix.transpose(1,2)[:,:-1, :-1].to(device).to(torch.float64)\n        targets = torch.matmul(points, transf_matrix_T).to(torch.float64)\n    \n        inv_transf_matrix =  torch.inverse(transf_matrix).to(torch.float64)\n\n    preds, confidences = model(inputs)\n    loss = criterion(targets, preds, confidences, target_availabilities)\n    \n    if cfg['img_coord_en']:\n        preds = torch.einsum('bmti,bji->bmtj', preds.double(), \n            inv_transf_matrix.to(device)[:, :2, :2])\n\n    return loss, preds, confidences\n```",
    "1081077": "It is over-complicated, as it is easier to multiply/divide by 2. But should work, good that you found the torch batch matrix multiplication. Further improvements are definitely outside of this code."
  },
  "source": "meta"
}