{
  "id": 188368,
  "title": "Augmentation considerations",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/188368",
  "author_name": "",
  "post_date": "2020-10-02T21:53:00.363726700Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We already have a pretty large pool of data, but as always augmentation will likely help with generalization. We don’t want our model memorizing the scenarios it sees in the training set and we likely have quite a bit of overlap because of the nature of the way the data was gathered. Because of this I was looking for methods of augmentation but found some snags.</p>\n<p>To some these may have been obvious but it took a little bit of thinking for me to configure something I think makes sense. Here is the path I went through in order to come up with an augmentation I think makes sense.</p>\n<p>Here is the original representation of the target vehicle's last known point and then the plot of the future points when using code from <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492\" target=\"_blank\">here</a> to transform targets to image units. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fea57225ded9effce4165285cceb1ea00%2Ffirst_image.png?generation=1601675352257839&amp;alt=media\" alt=\"\"></p>\n<p>The default configuration aligns the vehicle halfway down the image and a quarter of the image from the left. This comes from the ego_center param. It looks a little bit weird because the path is not aligned with the vehicle even after we do the transform to image units process. This is kind of nice because the path originates from 0,0, but might cause complications later on when applying augmentations. One major issue is that if the path happened to veer up instead of down then it would go outside of the plot. Some augmentation methods might struggle with this</p>\n<p>An alternate approach is to have the path not centered on the origin but on the car's location. This is a fairly simple constant addition to align with the vehicle or simply removing the bias term from the transform and inverse transform step. This would look something like this: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa3bb23c2a275d4a6984047cd319360e1%2Fsecond_image.png?generation=1601675392167194&amp;alt=media\" alt=\"\"></p>\n<p>This looks a bit more natural to have the two aligned together but still has some issues with various common augmentations. For example, imagine a simple rotation or flip and how that affects the output of the model's predictions. In the original unaugmented images, the starting point will always be height/2, width/4, but in the rotated or flipped image, the starting point can be anywhere in the image. This means rather than just trying to extrapolate a path given an origin it must also do the work to locate the origin. Another minor issue is that in the original setup with the trajectory starting at 0,0 the model is already reasonably well-calibrated, but with this alternate origin starting in the left side middle of the image a bias term needs to be learned to calibrate predictions. Not the end of the world but something to consider.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9793a67cbaffd1175ed6002a8ac82c81%2Fthird_image.png?generation=1601675446503051&amp;alt=media\" alt=\"\"></p>\n<p>In this example the model will then need to output ~25, 60 where it is more accustomed to having the car start at a fixed point. This is especially harsh when doing horizontal flips because the origin point is rather far from the augmented origin point.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa4fcf5b343d7862071d55879536ba278%2Ffourth_image.png?generation=1601675482895699&amp;alt=media\" alt=\"\"></p>\n<p>Notice that the image has been rotated, but the origin point is the same so the model just needs to learn the direction of travel and extrapolate from there.The major disadvantage with centering the starting point is it gives significantly less forward context. In the default settings 75% of the pixels are focused on the front of the vehicle, but when centered we lose 25% of that. It may require a higher resolution image to get adequate context. </p>\n<p><em>Extra notes about implementation</em>:</p>\n<ol>\n<li>Augmenting images + keypoints/coordinate outputs<br>\nWith <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> I knew that I was able to do lots of different augmentation techniques and also augment masks corresponding with those images, but I was not sure exactly how things would work given that we have x, y coordinates and not raw images. Of course we could just plot these points on an image and then augment those images and then extract the points back out, but albumentations already has the capability to augment coordinates with various different operations like flips and rotations</li>\n<li>Keypoints/coordinates are off the image<br>\nAnother nice thing that I discovered out of albumentations was that it is able to do these operations even if the coordinates fall off the image. This is important because it is potentially possible for our car to drive outside of the range of our image if we are using low resolution or have arranged the data such that the path starts at 0, 0 and can either go positive or negative like so:</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F4657666388cc94500ed0178954a8a88c%2Ffifth_image.png?generation=1601675495304757&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1035628",
      "postDate": "10/02/2020 21:53:00",
      "content": "<p>We already have a pretty large pool of data, but as always augmentation will likely help with generalization. We don’t want our model memorizing the scenarios it sees in the training set and we likely have quite a bit of overlap because of the nature of the way the data was gathered. Because of this I was looking for methods of augmentation but found some snags.</p>\n<p>To some these may have been obvious but it took a little bit of thinking for me to configure something I think makes sense. Here is the path I went through in order to come up with an augmentation I think makes sense.</p>\n<p>Here is the original representation of the target vehicle's last known point and then the plot of the future points when using code from <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492\" target=\"_blank\">here</a> to transform targets to image units. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fea57225ded9effce4165285cceb1ea00%2Ffirst_image.png?generation=1601675352257839&amp;alt=media\" alt=\"\"></p>\n<p>The default configuration aligns the vehicle halfway down the image and a quarter of the image from the left. This comes from the ego_center param. It looks a little bit weird because the path is not aligned with the vehicle even after we do the transform to image units process. This is kind of nice because the path originates from 0,0, but might cause complications later on when applying augmentations. One major issue is that if the path happened to veer up instead of down then it would go outside of the plot. Some augmentation methods might struggle with this</p>\n<p>An alternate approach is to have the path not centered on the origin but on the car's location. This is a fairly simple constant addition to align with the vehicle or simply removing the bias term from the transform and inverse transform step. This would look something like this: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa3bb23c2a275d4a6984047cd319360e1%2Fsecond_image.png?generation=1601675392167194&amp;alt=media\" alt=\"\"></p>\n<p>This looks a bit more natural to have the two aligned together but still has some issues with various common augmentations. For example, imagine a simple rotation or flip and how that affects the output of the model's predictions. In the original unaugmented images, the starting point will always be height/2, width/4, but in the rotated or flipped image, the starting point can be anywhere in the image. This means rather than just trying to extrapolate a path given an origin it must also do the work to locate the origin. Another minor issue is that in the original setup with the trajectory starting at 0,0 the model is already reasonably well-calibrated, but with this alternate origin starting in the left side middle of the image a bias term needs to be learned to calibrate predictions. Not the end of the world but something to consider.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9793a67cbaffd1175ed6002a8ac82c81%2Fthird_image.png?generation=1601675446503051&amp;alt=media\" alt=\"\"></p>\n<p>In this example the model will then need to output ~25, 60 where it is more accustomed to having the car start at a fixed point. This is especially harsh when doing horizontal flips because the origin point is rather far from the augmented origin point.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa4fcf5b343d7862071d55879536ba278%2Ffourth_image.png?generation=1601675482895699&amp;alt=media\" alt=\"\"></p>\n<p>Notice that the image has been rotated, but the origin point is the same so the model just needs to learn the direction of travel and extrapolate from there.The major disadvantage with centering the starting point is it gives significantly less forward context. In the default settings 75% of the pixels are focused on the front of the vehicle, but when centered we lose 25% of that. It may require a higher resolution image to get adequate context. </p>\n<p><em>Extra notes about implementation</em>:</p>\n<ol>\n<li>Augmenting images + keypoints/coordinate outputs<br>\nWith <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> I knew that I was able to do lots of different augmentation techniques and also augment masks corresponding with those images, but I was not sure exactly how things would work given that we have x, y coordinates and not raw images. Of course we could just plot these points on an image and then augment those images and then extract the points back out, but albumentations already has the capability to augment coordinates with various different operations like flips and rotations</li>\n<li>Keypoints/coordinates are off the image<br>\nAnother nice thing that I discovered out of albumentations was that it is able to do these operations even if the coordinates fall off the image. This is important because it is potentially possible for our car to drive outside of the range of our image if we are using low resolution or have arranged the data such that the path starts at 0, 0 and can either go positive or negative like so:</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F4657666388cc94500ed0178954a8a88c%2Ffifth_image.png?generation=1601675495304757&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "We already have a pretty large pool of data, but as always augmentation will likely help with generalization. We don’t want our model memorizing the scenarios it sees in the training set and we likely have quite a bit of overlap because of the nature of the way the data was gathered. Because of this I was looking for methods of augmentation but found some snags.\n\nTo some these may have been obvious but it took a little bit of thinking for me to configure something I think makes sense. Here is the path I went through in order to come up with an augmentation I think makes sense.\n\nHere is the original representation of the target vehicle's last known point and then the plot of the future points when using code from [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492) to transform targets to image units. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fea57225ded9effce4165285cceb1ea00%2Ffirst_image.png?generation=1601675352257839&alt=media)\n\nThe default configuration aligns the vehicle halfway down the image and a quarter of the image from the left. This comes from the ego_center param. It looks a little bit weird because the path is not aligned with the vehicle even after we do the transform to image units process. This is kind of nice because the path originates from 0,0, but might cause complications later on when applying augmentations. One major issue is that if the path happened to veer up instead of down then it would go outside of the plot. Some augmentation methods might struggle with this\n\nAn alternate approach is to have the path not centered on the origin but on the car's location. This is a fairly simple constant addition to align with the vehicle or simply removing the bias term from the transform and inverse transform step. This would look something like this: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa3bb23c2a275d4a6984047cd319360e1%2Fsecond_image.png?generation=1601675392167194&alt=media)\n\nThis looks a bit more natural to have the two aligned together but still has some issues with various common augmentations. For example, imagine a simple rotation or flip and how that affects the output of the model's predictions. In the original unaugmented images, the starting point will always be height/2, width/4, but in the rotated or flipped image, the starting point can be anywhere in the image. This means rather than just trying to extrapolate a path given an origin it must also do the work to locate the origin. Another minor issue is that in the original setup with the trajectory starting at 0,0 the model is already reasonably well-calibrated, but with this alternate origin starting in the left side middle of the image a bias term needs to be learned to calibrate predictions. Not the end of the world but something to consider.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9793a67cbaffd1175ed6002a8ac82c81%2Fthird_image.png?generation=1601675446503051&alt=media)\n\n\nIn this example the model will then need to output ~25, 60 where it is more accustomed to having the car start at a fixed point. This is especially harsh when doing horizontal flips because the origin point is rather far from the augmented origin point.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa4fcf5b343d7862071d55879536ba278%2Ffourth_image.png?generation=1601675482895699&alt=media)\n\nNotice that the image has been rotated, but the origin point is the same so the model just needs to learn the direction of travel and extrapolate from there.The major disadvantage with centering the starting point is it gives significantly less forward context. In the default settings 75% of the pixels are focused on the front of the vehicle, but when centered we lose 25% of that. It may require a higher resolution image to get adequate context. \n\n*Extra notes about implementation*:\n\n1. Augmenting images + keypoints/coordinate outputs\n    With [albumentations](https://github.com/albumentations-team/albumentations) I knew that I was able to do lots of different augmentation techniques and also augment masks corresponding with those images, but I was not sure exactly how things would work given that we have x, y coordinates and not raw images. Of course we could just plot these points on an image and then augment those images and then extract the points back out, but albumentations already has the capability to augment coordinates with various different operations like flips and rotations\n2. Keypoints/coordinates are off the image\n    Another nice thing that I discovered out of albumentations was that it is able to do these operations even if the coordinates fall off the image. This is important because it is potentially possible for our car to drive outside of the range of our image if we are using low resolution or have arranged the data such that the path starts at 0, 0 and can either go positive or negative like so:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F4657666388cc94500ed0178954a8a88c%2Ffifth_image.png?generation=1601675495304757&alt=media)",
      "votes": null
    },
    {
      "id": "1039919",
      "postDate": "10/06/2020 22:08:12",
      "content": "<p>If you are planning to use this kind of augmentation, would probebly be better to adjust the world_to_image matrix (or what is used by AgentDataset) before rasterization so everything is consistent.</p>",
      "rawMarkdown": "If you are planning to use this kind of augmentation, would probebly be better to adjust the world_to_image matrix (or what is used by AgentDataset) before rasterization so everything is consistent.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1039919,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "10/06/2020 22:08:12",
      "content": "<p>If you are planning to use this kind of augmentation, would probebly be better to adjust the world_to_image matrix (or what is used by AgentDataset) before rasterization so everything is consistent.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1035628": "We already have a pretty large pool of data, but as always augmentation will likely help with generalization. We don’t want our model memorizing the scenarios it sees in the training set and we likely have quite a bit of overlap because of the nature of the way the data was gathered. Because of this I was looking for methods of augmentation but found some snags.\n\nTo some these may have been obvious but it took a little bit of thinking for me to configure something I think makes sense. Here is the path I went through in order to come up with an augmentation I think makes sense.\n\nHere is the original representation of the target vehicle's last known point and then the plot of the future points when using code from [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492) to transform targets to image units. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fea57225ded9effce4165285cceb1ea00%2Ffirst_image.png?generation=1601675352257839&alt=media)\n\nThe default configuration aligns the vehicle halfway down the image and a quarter of the image from the left. This comes from the ego_center param. It looks a little bit weird because the path is not aligned with the vehicle even after we do the transform to image units process. This is kind of nice because the path originates from 0,0, but might cause complications later on when applying augmentations. One major issue is that if the path happened to veer up instead of down then it would go outside of the plot. Some augmentation methods might struggle with this\n\nAn alternate approach is to have the path not centered on the origin but on the car's location. This is a fairly simple constant addition to align with the vehicle or simply removing the bias term from the transform and inverse transform step. This would look something like this: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa3bb23c2a275d4a6984047cd319360e1%2Fsecond_image.png?generation=1601675392167194&alt=media)\n\nThis looks a bit more natural to have the two aligned together but still has some issues with various common augmentations. For example, imagine a simple rotation or flip and how that affects the output of the model's predictions. In the original unaugmented images, the starting point will always be height/2, width/4, but in the rotated or flipped image, the starting point can be anywhere in the image. This means rather than just trying to extrapolate a path given an origin it must also do the work to locate the origin. Another minor issue is that in the original setup with the trajectory starting at 0,0 the model is already reasonably well-calibrated, but with this alternate origin starting in the left side middle of the image a bias term needs to be learned to calibrate predictions. Not the end of the world but something to consider.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9793a67cbaffd1175ed6002a8ac82c81%2Fthird_image.png?generation=1601675446503051&alt=media)\n\n\nIn this example the model will then need to output ~25, 60 where it is more accustomed to having the car start at a fixed point. This is especially harsh when doing horizontal flips because the origin point is rather far from the augmented origin point.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fa4fcf5b343d7862071d55879536ba278%2Ffourth_image.png?generation=1601675482895699&alt=media)\n\nNotice that the image has been rotated, but the origin point is the same so the model just needs to learn the direction of travel and extrapolate from there.The major disadvantage with centering the starting point is it gives significantly less forward context. In the default settings 75% of the pixels are focused on the front of the vehicle, but when centered we lose 25% of that. It may require a higher resolution image to get adequate context. \n\n*Extra notes about implementation*:\n\n1. Augmenting images + keypoints/coordinate outputs\n    With [albumentations](https://github.com/albumentations-team/albumentations) I knew that I was able to do lots of different augmentation techniques and also augment masks corresponding with those images, but I was not sure exactly how things would work given that we have x, y coordinates and not raw images. Of course we could just plot these points on an image and then augment those images and then extract the points back out, but albumentations already has the capability to augment coordinates with various different operations like flips and rotations\n2. Keypoints/coordinates are off the image\n    Another nice thing that I discovered out of albumentations was that it is able to do these operations even if the coordinates fall off the image. This is important because it is potentially possible for our car to drive outside of the range of our image if we are using low resolution or have arranged the data such that the path starts at 0, 0 and can either go positive or negative like so:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F4657666388cc94500ed0178954a8a88c%2Ffifth_image.png?generation=1601675495304757&alt=media)",
    "1039919": "If you are planning to use this kind of augmentation, would probebly be better to adjust the world_to_image matrix (or what is used by AgentDataset) before rasterization so everything is consistent."
  },
  "source": "meta"
}