{
  "id": 186767,
  "title": "Expected Error: Inaccuracy from Rasterization",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/186767",
  "author_name": "",
  "post_date": "2020-09-25T20:05:49.553952100Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>As we are using a rasterizer in most public Notebooks right now, I was wondering how large the expected error from the inaccuracy from rasterization is and went ahead to do some calculations. </p>\n<p>Each history position, each lane, each other agent is encoded into a pixels and our net is only able to predict the next positions on the map with pixel level accuracy. <br>\nIn many notebooks, the raster has a size of 0.50 m per pixel (hyperparameter). Thus, the expected mean error will be a 0.50 / 4 for each direction for each predicted position.</p>\n<p><img src=\"https://i.imgur.com/RXmUxBd.png\" alt=\"error\"><br>\n<img src=\"https://i.imgur.com/4zPjaFO.png\" alt=\"legend\"></p>\n<p>Additional expected error for pixel_size = 0.50 (nll metric): <strong>0.78125</strong>  <br>\nAdditional expected error for pixel_size = 0.25 (nll metric): <strong>0.1953125</strong>  </p>\n<p>Seems like we have to expect, that 0.78125 of our metric is actually just the inaccuracy from the rasterizer. We can reduce that value to 0.1953125 by creating a smaller grid (e.g. pixel_size = 0.25). Keep in mind, that you would need to double the raster_size to see the same region of the map.</p>\n<p>Please point out any errors that I might have in my assumptions and/or calculations. </p>",
  "messages": [
    {
      "id": "1027085",
      "postDate": "09/25/2020 20:05:49",
      "content": "<p>As we are using a rasterizer in most public Notebooks right now, I was wondering how large the expected error from the inaccuracy from rasterization is and went ahead to do some calculations. </p>\n<p>Each history position, each lane, each other agent is encoded into a pixels and our net is only able to predict the next positions on the map with pixel level accuracy. <br>\nIn many notebooks, the raster has a size of 0.50 m per pixel (hyperparameter). Thus, the expected mean error will be a 0.50 / 4 for each direction for each predicted position.</p>\n<p><img src=\"https://i.imgur.com/RXmUxBd.png\" alt=\"error\"><br>\n<img src=\"https://i.imgur.com/4zPjaFO.png\" alt=\"legend\"></p>\n<p>Additional expected error for pixel_size = 0.50 (nll metric): <strong>0.78125</strong>  <br>\nAdditional expected error for pixel_size = 0.25 (nll metric): <strong>0.1953125</strong>  </p>\n<p>Seems like we have to expect, that 0.78125 of our metric is actually just the inaccuracy from the rasterizer. We can reduce that value to 0.1953125 by creating a smaller grid (e.g. pixel_size = 0.25). Keep in mind, that you would need to double the raster_size to see the same region of the map.</p>\n<p>Please point out any errors that I might have in my assumptions and/or calculations. </p>",
      "rawMarkdown": "As we are using a rasterizer in most public Notebooks right now, I was wondering how large the expected error from the inaccuracy from rasterization is and went ahead to do some calculations. \n\nEach history position, each lane, each other agent is encoded into a pixels and our net is only able to predict the next positions on the map with pixel level accuracy. \nIn many notebooks, the raster has a size of 0.50 m per pixel (hyperparameter). Thus, the expected mean error will be a 0.50 / 4 for each direction for each predicted position.\n\n![error](https://i.imgur.com/RXmUxBd.png)\n![legend](https://i.imgur.com/4zPjaFO.png)\n\nAdditional expected error for pixel_size = 0.50 (nll metric): **0.78125**  \nAdditional expected error for pixel_size = 0.25 (nll metric): **0.1953125**  \n\nSeems like we have to expect, that 0.78125 of our metric is actually just the inaccuracy from the rasterizer. We can reduce that value to 0.1953125 by creating a smaller grid (e.g. pixel_size = 0.25). Keep in mind, that you would need to double the raster_size to see the same region of the map.\n\nPlease point out any errors that I might have in my assumptions and/or calculations.",
      "votes": null
    },
    {
      "id": "1027172",
      "postDate": "09/25/2020 23:07:56",
      "content": "<p>I'm not sure that the net is limited to the pixel level accuracy for its output predictions, but it definitely does apply to the input accuracy. The training targets are still floats / long decimals?</p>\n<p>There is also the angular precision, so you often have a number of single pixel accuracies combined ? / eg one pixel forward at probably at least  degree per step accuracy in a 90 degree sweep? -- ah, likely less than 90 steps for smaller rectangles.. should be around [ pi/2 x radius ] corner pixels </p>\n<p>Add: You can see it illustrated in your picture, if you try to shift the line perpendicular (eg to the top right a little) the pixels will step slightly</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19411%2F5ee0a535fefbd2f69fae77989b695c8a%2Fpixels.png?generation=1601082264304852&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'm not sure that the net is limited to the pixel level accuracy for its output predictions, but it definitely does apply to the input accuracy. The training targets are still floats / long decimals?\n\nThere is also the angular precision, so you often have a number of single pixel accuracies combined ? / eg one pixel forward at probably at least ~~1~~ degree per step accuracy in a 90 degree sweep? -- ah, likely less than 90 steps for smaller rectangles.. should be around [ pi/2 x radius ] corner pixels \n\nAdd: You can see it illustrated in your picture, if you try to shift the line perpendicular (eg to the top right a little) the pixels will step slightly\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19411%2F5ee0a535fefbd2f69fae77989b695c8a%2Fpixels.png?generation=1601082264304852&alt=media)",
      "votes": null
    },
    {
      "id": "1027241",
      "postDate": "09/26/2020 01:19:13",
      "content": "<p>There is some ambiguity though, i guess the vehicle corners might be more important than the sides</p>",
      "rawMarkdown": "There is some ambiguity though, i guess the vehicle corners might be more important than the sides",
      "votes": null
    },
    {
      "id": "1027522",
      "postDate": "09/26/2020 08:03:52",
      "content": "<p>Yeah, as <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a>  notes the outputs aren't rounded to pixels.<br>\nAlso, the semantic map is actually rendered with sub-pixel accuracy <a href=\"https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/semantic_rasterizer.py#L39\" target=\"_blank\">(implemented here</a>). Positions for points are specified to 1/256th of a pixel (though the opencv rendering may not distinguish to this level). <br>\nHere's lines at various subpixel locations as used in <code>SemanticRasterizer</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3156659%2F6017ed830018427314874e2451ffbc30%2FSubPixel_x.png?generation=1601106811234301&amp;alt=media\" alt=\"\">.</p>\n<p>For some reason the <code>BoxRasterizer</code> for vehicles doesn't do this (in fact it's just simple truncation to int not even rounding to nearest pixel, <a href=\"https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/box_rasterizer.py#L68\" target=\"_blank\">implemented here</a>). So actually up to essentially a full pixel of error (well whatever the nearest float to 1 is, 0.999…). Not sure why it differs.</p>\n<p>So there is some representation of subpixel values. Obviously not clear how much useful information a NN will get from this but it is there. The fact this is only present for the semantic map is also an issue as I'd have thought accuracy of vehicle positions would probably be more important.</p>",
      "rawMarkdown": "Yeah, as @n3n77i  notes the outputs aren't rounded to pixels.\nAlso, the semantic map is actually rendered with sub-pixel accuracy [(implemented here](https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/semantic_rasterizer.py#L39)). Positions for points are specified to 1/256th of a pixel (though the opencv rendering may not distinguish to this level). \nHere's lines at various subpixel locations as used in `SemanticRasterizer`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3156659%2F6017ed830018427314874e2451ffbc30%2FSubPixel_x.png?generation=1601106811234301&alt=media).\n\nFor some reason the `BoxRasterizer` for vehicles doesn't do this (in fact it's just simple truncation to int not even rounding to nearest pixel, [implemented here](https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/box_rasterizer.py#L68)). So actually up to essentially a full pixel of error (well whatever the nearest float to 1 is, 0.999...). Not sure why it differs.\n\nSo there is some representation of subpixel values. Obviously not clear how much useful information a NN will get from this but it is there. The fact this is only present for the semantic map is also an issue as I'd have thought accuracy of vehicle positions would probably be more important.",
      "votes": null
    },
    {
      "id": "1027566",
      "postDate": "09/26/2020 08:40:36",
      "content": "<p>Ah, thanks for pointing out. That is good news, though, we probably also want the positions to have sub pixel accuracy. </p>\n<p>I know, the targets are not rounded, but the NN would have a hard time predicting those with inaccurate input, that was my concern. </p>",
      "rawMarkdown": "Ah, thanks for pointing out. That is good news, though, we probably also want the positions to have sub pixel accuracy. \n\nI know, the targets are not rounded, but the NN would have a hard time predicting those with inaccurate input, that was my concern.",
      "votes": null
    },
    {
      "id": "1027573",
      "postDate": "09/26/2020 08:43:25",
      "content": "<p>Good point, I haven't thought about the time relationship and rotation. That has additional value in terms of accuracy. </p>\n<p>Seems like overall, the 0.50 pixel_size is not such a bad value and decreasing that value does not too much impact on the score. </p>",
      "rawMarkdown": "Good point, I haven't thought about the time relationship and rotation. That has additional value in terms of accuracy. \n\nSeems like overall, the 0.50 pixel_size is not such a bad value and decreasing that value does not too much impact on the score.",
      "votes": null
    },
    {
      "id": "1027688",
      "postDate": "09/26/2020 09:49:10",
      "content": "<p>Yeah, would obviously be limited ability to generate good predictions without the subpixel rasterizing and still higher pixel density might help (but also slow things down given the rasterizing bottleneck). I haven't yet run any experiments on that. Obviously also the issue then of running NNs with non-optimal input sizes in terms of pre-training and also architecture. Testing on efficientnet would be a good way to check given the variants for different input sizes but I gather those who've tried using haven't got good results to start with (though I haven't seen much code there so not sure if they're implementing things ideally).</p>\n<p>Not sure why the <code>BoxRasterizer</code> doesn't use subpixel. I did check and subpixel works for single channel images as used for positions so should be easy enough to implement. I did ask about this when discussing another issue on the l5kit github but might raise it again separately.</p>",
      "rawMarkdown": "Yeah, would obviously be limited ability to generate good predictions without the subpixel rasterizing and still higher pixel density might help (but also slow things down given the rasterizing bottleneck). I haven't yet run any experiments on that. Obviously also the issue then of running NNs with non-optimal input sizes in terms of pre-training and also architecture. Testing on efficientnet would be a good way to check given the variants for different input sizes but I gather those who've tried using haven't got good results to start with (though I haven't seen much code there so not sure if they're implementing things ideally).\n\nNot sure why the `BoxRasterizer` doesn't use subpixel. I did check and subpixel works for single channel images as used for positions so should be easy enough to implement. I did ask about this when discussing another issue on the l5kit github but might raise it again separately.",
      "votes": null
    },
    {
      "id": "1028506",
      "postDate": "09/26/2020 23:39:04",
      "content": "<p>There are models that would suffer from that kind of output inaccuracy though, maybe gans or auto-encoders that try to return rasterized predictions. They could be fast though, the rasterization bottleneck wouldn't really exist?</p>",
      "rawMarkdown": "There are models that would suffer from that kind of output inaccuracy though, maybe gans or auto-encoders that try to return rasterized predictions. They could be fast though, the rasterization bottleneck wouldn't really exist?",
      "votes": null
    },
    {
      "id": "1030354",
      "postDate": "09/28/2020 15:42:55",
      "content": "<p>sub-pixel precision was introduced only recently in L5Kit so I still haven't piggybacked it to the <code>BoxRasterizer</code>. I'm working on it with now though :) </p>",
      "rawMarkdown": "sub-pixel precision was introduced only recently in L5Kit so I still haven't piggybacked it to the `BoxRasterizer`. I'm working on it with now though :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1027172,
      "author_name": "n3n77i",
      "author_url": "",
      "post_date": "09/25/2020 23:07:56",
      "content": "<p>I'm not sure that the net is limited to the pixel level accuracy for its output predictions, but it definitely does apply to the input accuracy. The training targets are still floats / long decimals?</p>\n<p>There is also the angular precision, so you often have a number of single pixel accuracies combined ? / eg one pixel forward at probably at least  degree per step accuracy in a 90 degree sweep? -- ah, likely less than 90 steps for smaller rectangles.. should be around [ pi/2 x radius ] corner pixels </p>\n<p>Add: You can see it illustrated in your picture, if you try to shift the line perpendicular (eg to the top right a little) the pixels will step slightly</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19411%2F5ee0a535fefbd2f69fae77989b695c8a%2Fpixels.png?generation=1601082264304852&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1027241,
          "author_name": "n3n77i",
          "author_url": "",
          "post_date": "09/26/2020 01:19:13",
          "content": "<p>There is some ambiguity though, i guess the vehicle corners might be more important than the sides</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1027573,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "09/26/2020 08:43:25",
          "content": "<p>Good point, I haven't thought about the time relationship and rotation. That has additional value in terms of accuracy. </p>\n<p>Seems like overall, the 0.50 pixel_size is not such a bad value and decreasing that value does not too much impact on the score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1028506,
          "author_name": "n3n77i",
          "author_url": "",
          "post_date": "09/26/2020 23:39:04",
          "content": "<p>There are models that would suffer from that kind of output inaccuracy though, maybe gans or auto-encoders that try to return rasterized predictions. They could be fast though, the rasterization bottleneck wouldn't really exist?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1027522,
      "author_name": "thomasbrandon",
      "author_url": "",
      "post_date": "09/26/2020 08:03:52",
      "content": "<p>Yeah, as <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a>  notes the outputs aren't rounded to pixels.<br>\nAlso, the semantic map is actually rendered with sub-pixel accuracy <a href=\"https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/semantic_rasterizer.py#L39\" target=\"_blank\">(implemented here</a>). Positions for points are specified to 1/256th of a pixel (though the opencv rendering may not distinguish to this level). <br>\nHere's lines at various subpixel locations as used in <code>SemanticRasterizer</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3156659%2F6017ed830018427314874e2451ffbc30%2FSubPixel_x.png?generation=1601106811234301&amp;alt=media\" alt=\"\">.</p>\n<p>For some reason the <code>BoxRasterizer</code> for vehicles doesn't do this (in fact it's just simple truncation to int not even rounding to nearest pixel, <a href=\"https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/box_rasterizer.py#L68\" target=\"_blank\">implemented here</a>). So actually up to essentially a full pixel of error (well whatever the nearest float to 1 is, 0.999…). Not sure why it differs.</p>\n<p>So there is some representation of subpixel values. Obviously not clear how much useful information a NN will get from this but it is there. The fact this is only present for the semantic map is also an issue as I'd have thought accuracy of vehicle positions would probably be more important.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1027566,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "09/26/2020 08:40:36",
          "content": "<p>Ah, thanks for pointing out. That is good news, though, we probably also want the positions to have sub pixel accuracy. </p>\n<p>I know, the targets are not rounded, but the NN would have a hard time predicting those with inaccurate input, that was my concern. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1027688,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "09/26/2020 09:49:10",
          "content": "<p>Yeah, would obviously be limited ability to generate good predictions without the subpixel rasterizing and still higher pixel density might help (but also slow things down given the rasterizing bottleneck). I haven't yet run any experiments on that. Obviously also the issue then of running NNs with non-optimal input sizes in terms of pre-training and also architecture. Testing on efficientnet would be a good way to check given the variants for different input sizes but I gather those who've tried using haven't got good results to start with (though I haven't seen much code there so not sure if they're implementing things ideally).</p>\n<p>Not sure why the <code>BoxRasterizer</code> doesn't use subpixel. I did check and subpixel works for single channel images as used for positions so should be easy enough to implement. I did ask about this when discussing another issue on the l5kit github but might raise it again separately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1030354,
          "author_name": "lucabergamini",
          "author_url": "",
          "post_date": "09/28/2020 15:42:55",
          "content": "<p>sub-pixel precision was introduced only recently in L5Kit so I still haven't piggybacked it to the <code>BoxRasterizer</code>. I'm working on it with now though :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1027085": "As we are using a rasterizer in most public Notebooks right now, I was wondering how large the expected error from the inaccuracy from rasterization is and went ahead to do some calculations. \n\nEach history position, each lane, each other agent is encoded into a pixels and our net is only able to predict the next positions on the map with pixel level accuracy. \nIn many notebooks, the raster has a size of 0.50 m per pixel (hyperparameter). Thus, the expected mean error will be a 0.50 / 4 for each direction for each predicted position.\n\n![error](https://i.imgur.com/RXmUxBd.png)\n![legend](https://i.imgur.com/4zPjaFO.png)\n\nAdditional expected error for pixel_size = 0.50 (nll metric): **0.78125**  \nAdditional expected error for pixel_size = 0.25 (nll metric): **0.1953125**  \n\nSeems like we have to expect, that 0.78125 of our metric is actually just the inaccuracy from the rasterizer. We can reduce that value to 0.1953125 by creating a smaller grid (e.g. pixel_size = 0.25). Keep in mind, that you would need to double the raster_size to see the same region of the map.\n\nPlease point out any errors that I might have in my assumptions and/or calculations.",
    "1027172": "I'm not sure that the net is limited to the pixel level accuracy for its output predictions, but it definitely does apply to the input accuracy. The training targets are still floats / long decimals?\n\nThere is also the angular precision, so you often have a number of single pixel accuracies combined ? / eg one pixel forward at probably at least ~~1~~ degree per step accuracy in a 90 degree sweep? -- ah, likely less than 90 steps for smaller rectangles.. should be around [ pi/2 x radius ] corner pixels \n\nAdd: You can see it illustrated in your picture, if you try to shift the line perpendicular (eg to the top right a little) the pixels will step slightly\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19411%2F5ee0a535fefbd2f69fae77989b695c8a%2Fpixels.png?generation=1601082264304852&alt=media)",
    "1027241": "There is some ambiguity though, i guess the vehicle corners might be more important than the sides",
    "1027522": "Yeah, as @n3n77i  notes the outputs aren't rounded to pixels.\nAlso, the semantic map is actually rendered with sub-pixel accuracy [(implemented here](https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/semantic_rasterizer.py#L39)). Positions for points are specified to 1/256th of a pixel (though the opencv rendering may not distinguish to this level). \nHere's lines at various subpixel locations as used in `SemanticRasterizer`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3156659%2F6017ed830018427314874e2451ffbc30%2FSubPixel_x.png?generation=1601106811234301&alt=media).\n\nFor some reason the `BoxRasterizer` for vehicles doesn't do this (in fact it's just simple truncation to int not even rounding to nearest pixel, [implemented here](https://github.com/lyft/l5kit/blob/41f3bbaf4bcf9d2396460dac57b68106df88ed68/l5kit/l5kit/rasterization/box_rasterizer.py#L68)). So actually up to essentially a full pixel of error (well whatever the nearest float to 1 is, 0.999...). Not sure why it differs.\n\nSo there is some representation of subpixel values. Obviously not clear how much useful information a NN will get from this but it is there. The fact this is only present for the semantic map is also an issue as I'd have thought accuracy of vehicle positions would probably be more important.",
    "1027566": "Ah, thanks for pointing out. That is good news, though, we probably also want the positions to have sub pixel accuracy. \n\nI know, the targets are not rounded, but the NN would have a hard time predicting those with inaccurate input, that was my concern.",
    "1027573": "Good point, I haven't thought about the time relationship and rotation. That has additional value in terms of accuracy. \n\nSeems like overall, the 0.50 pixel_size is not such a bad value and decreasing that value does not too much impact on the score.",
    "1027688": "Yeah, would obviously be limited ability to generate good predictions without the subpixel rasterizing and still higher pixel density might help (but also slow things down given the rasterizing bottleneck). I haven't yet run any experiments on that. Obviously also the issue then of running NNs with non-optimal input sizes in terms of pre-training and also architecture. Testing on efficientnet would be a good way to check given the variants for different input sizes but I gather those who've tried using haven't got good results to start with (though I haven't seen much code there so not sure if they're implementing things ideally).\n\nNot sure why the `BoxRasterizer` doesn't use subpixel. I did check and subpixel works for single channel images as used for positions so should be easy enough to implement. I did ask about this when discussing another issue on the l5kit github but might raise it again separately.",
    "1028506": "There are models that would suffer from that kind of output inaccuracy though, maybe gans or auto-encoders that try to return rasterized predictions. They could be fast though, the rasterization bottleneck wouldn't really exist?",
    "1030354": "sub-pixel precision was introduced only recently in L5Kit so I still haven't piggybacked it to the `BoxRasterizer`. I'm working on it with now though :)"
  },
  "source": "meta"
}