{
  "id": 201143,
  "title": "14th place solution. Custom Mask and LSTM encoder/decoder",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/matsuoiv-14th-place-solution-custom-mask-and-lstm-",
  "author_name": "",
  "post_date": "2020-12-03T11:26:41.101723500Z",
  "votes": 23,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Great thanks to Lyft and kaggle team for hosting this competition. Congrats to all the winners. And a special thanks to all my teammates for teaming up!! The time I spent discussing together was very insightful and most enjoyable.</p>\n<h1>Create Mask</h1>\n<p>Our solution’s point is that we made the custom mask.<br>\nAgent data one frame away from the same scene is very similar. So, in order to learn efficiently, we created a custom mask so that the data is loaded every N frames. For example, by sampling data every 5 frames, the training data is reduced by a factor of 5, allowing for faster learning.<br>\nTo make it more efficient, we balanced the data according to the target distance. About 65% of the data had a target distance of less than 6 meters, and they were easy tasks. So in order to efficiently train the difficult data, we downsampled the data below 6 m. This halved the data. In the end, the number of data was roughly 80,000 iterations x 64 batch size.</p>\n<h1>Model architecture</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F98319e9931caf1a99cbfdfdbf0c78237%2Ffinal_model.jpg?generation=1606994527055759&amp;alt=media\" alt=\"\"></p>\n<p>In addition to CNN, we used LSTM encoder and LSTM decoder.</p>\n<ul>\n<li>Image size: 300x300</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: Onecyclelr</li>\n<li>Num history: 10</li>\n<li>Target: positions + cosine/sine of yaw</li>\n</ul>\n<p>First, we trained the model with the target distance balanced mask. Additional learning with the non-balanced mask. We used log10-scaled loss when additional training, because train/validate loss goes up and down with a large margin repeatedly. For two days of model training by using this method, we got a public score under 13, and a private score about 12.</p>\n<h1>Didn’t work</h1>\n<p>history_positions more than 10 (Public score around 15, slow convergence)<br>\nmetalabeling<br>\nSeq2seq like LSTM decoder</p>\n<hr>\n<h1>Discussion</h1>\n<h2>Stacking</h2>\n<p>As discussed in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199531#1091632\" target=\"_blank\">this thread</a>, we found that stacking boosted the score, but had no time left integrating the best model to stacking…</p>\n<h2>Satellite image</h2>\n<p>Until halfway through, I had mistakenly used satellite_debug as the 3 channel map information instead of semantic_map. Interestingly, this model got a public score of 13.3.<br>\nI didn’t pursue it further because I realized the mistake, but it seems that satellite maps can score somewhat well in this task. Try it if you like it!!!</p>\n<h2>3D-CNN(R2+1D)</h2>\n<p>Rasterized image data in the form of RGB×T×H×W, then fed them into R2+1D.<br>\nValidation score was equivalent to resnet18, but the process of rasterizing and the 3D-CNN model were so heavy that the training required twice ~ triple as much time.</p>\n<h2>Only sequence model</h2>\n<p>We trained Seq2seq model not only with the target agent’s history position, but also with the surrounding vehicle’s history position. Despite no map information, this model got a validation score of 24. We integrated this trained seq2seq model into the trained resnet18 model, and fine-tuned the concat model. Sadly, improvement of the score was marginal.</p>",
  "messages": [
    {
      "id": "1100817",
      "postDate": "12/03/2020 11:26:41",
      "content": "<p>Great thanks to Lyft and kaggle team for hosting this competition. Congrats to all the winners. And a special thanks to all my teammates for teaming up!! The time I spent discussing together was very insightful and most enjoyable.</p>\n<h1>Create Mask</h1>\n<p>Our solution’s point is that we made the custom mask.<br>\nAgent data one frame away from the same scene is very similar. So, in order to learn efficiently, we created a custom mask so that the data is loaded every N frames. For example, by sampling data every 5 frames, the training data is reduced by a factor of 5, allowing for faster learning.<br>\nTo make it more efficient, we balanced the data according to the target distance. About 65% of the data had a target distance of less than 6 meters, and they were easy tasks. So in order to efficiently train the difficult data, we downsampled the data below 6 m. This halved the data. In the end, the number of data was roughly 80,000 iterations x 64 batch size.</p>\n<h1>Model architecture</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F98319e9931caf1a99cbfdfdbf0c78237%2Ffinal_model.jpg?generation=1606994527055759&amp;alt=media\" alt=\"\"></p>\n<p>In addition to CNN, we used LSTM encoder and LSTM decoder.</p>\n<ul>\n<li>Image size: 300x300</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: Onecyclelr</li>\n<li>Num history: 10</li>\n<li>Target: positions + cosine/sine of yaw</li>\n</ul>\n<p>First, we trained the model with the target distance balanced mask. Additional learning with the non-balanced mask. We used log10-scaled loss when additional training, because train/validate loss goes up and down with a large margin repeatedly. For two days of model training by using this method, we got a public score under 13, and a private score about 12.</p>\n<h1>Didn’t work</h1>\n<p>history_positions more than 10 (Public score around 15, slow convergence)<br>\nmetalabeling<br>\nSeq2seq like LSTM decoder</p>\n<hr>\n<h1>Discussion</h1>\n<h2>Stacking</h2>\n<p>As discussed in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199531#1091632\" target=\"_blank\">this thread</a>, we found that stacking boosted the score, but had no time left integrating the best model to stacking…</p>\n<h2>Satellite image</h2>\n<p>Until halfway through, I had mistakenly used satellite_debug as the 3 channel map information instead of semantic_map. Interestingly, this model got a public score of 13.3.<br>\nI didn’t pursue it further because I realized the mistake, but it seems that satellite maps can score somewhat well in this task. Try it if you like it!!!</p>\n<h2>3D-CNN(R2+1D)</h2>\n<p>Rasterized image data in the form of RGB×T×H×W, then fed them into R2+1D.<br>\nValidation score was equivalent to resnet18, but the process of rasterizing and the 3D-CNN model were so heavy that the training required twice ~ triple as much time.</p>\n<h2>Only sequence model</h2>\n<p>We trained Seq2seq model not only with the target agent’s history position, but also with the surrounding vehicle’s history position. Despite no map information, this model got a validation score of 24. We integrated this trained seq2seq model into the trained resnet18 model, and fine-tuned the concat model. Sadly, improvement of the score was marginal.</p>",
      "rawMarkdown": "Great thanks to Lyft and kaggle team for hosting this competition. Congrats to all the winners. And a special thanks to all my teammates for teaming up!! The time I spent discussing together was very insightful and most enjoyable.\n\n# Create Mask\n\nOur solution’s point is that we made the custom mask.\nAgent data one frame away from the same scene is very similar. So, in order to learn efficiently, we created a custom mask so that the data is loaded every N frames. For example, by sampling data every 5 frames, the training data is reduced by a factor of 5, allowing for faster learning.\nTo make it more efficient, we balanced the data according to the target distance. About 65% of the data had a target distance of less than 6 meters, and they were easy tasks. So in order to efficiently train the difficult data, we downsampled the data below 6 m. This halved the data. In the end, the number of data was roughly 80,000 iterations x 64 batch size.\n\n# Model architecture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F98319e9931caf1a99cbfdfdbf0c78237%2Ffinal_model.jpg?generation=1606994527055759&alt=media)\n\nIn addition to CNN, we used LSTM encoder and LSTM decoder.\n- Image size: 300x300\n- Optimizer: Adam\n- Scheduler: Onecyclelr\n- Num history: 10\n- Target: positions + cosine/sine of yaw\n\nFirst, we trained the model with the target distance balanced mask. Additional learning with the non-balanced mask. We used log10-scaled loss when additional training, because train/validate loss goes up and down with a large margin repeatedly. For two days of model training by using this method, we got a public score under 13, and a private score about 12.\n\n\n# Didn’t work\nhistory_positions more than 10 (Public score around 15, slow convergence)\nmetalabeling\nSeq2seq like LSTM decoder\n\n----------\n# Discussion\n\n## Stacking\n\nAs discussed in [this thread](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199531#1091632), we found that stacking boosted the score, but had no time left integrating the best model to stacking...\n\n## Satellite image\n\nUntil halfway through, I had mistakenly used satellite_debug as the 3 channel map information instead of semantic_map. Interestingly, this model got a public score of 13.3.\nI didn’t pursue it further because I realized the mistake, but it seems that satellite maps can score somewhat well in this task. Try it if you like it!!!\n\n## 3D-CNN(R2+1D)\n\nRasterized image data in the form of RGB×T×H×W, then fed them into R2+1D.\nValidation score was equivalent to resnet18, but the process of rasterizing and the 3D-CNN model were so heavy that the training required twice ~ triple as much time.\n\n## Only sequence model\n\nWe trained Seq2seq model not only with the target agent’s history position, but also with the surrounding vehicle’s history position. Despite no map information, this model got a validation score of 24. We integrated this trained seq2seq model into the trained resnet18 model, and fine-tuned the concat model. Sadly, improvement of the score was marginal.",
      "votes": null
    },
    {
      "id": "1103471",
      "postDate": "12/05/2020 23:45:57",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/Inoichan\" target=\"_blank\">@Inoichan</a> with the summarized figure!</p>",
      "rawMarkdown": "Thanks for sharing @Inoichan with the summarized figure!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1103471,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "12/05/2020 23:45:57",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/Inoichan\" target=\"_blank\">@Inoichan</a> with the summarized figure!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1100817": "Great thanks to Lyft and kaggle team for hosting this competition. Congrats to all the winners. And a special thanks to all my teammates for teaming up!! The time I spent discussing together was very insightful and most enjoyable.\n\n# Create Mask\n\nOur solution’s point is that we made the custom mask.\nAgent data one frame away from the same scene is very similar. So, in order to learn efficiently, we created a custom mask so that the data is loaded every N frames. For example, by sampling data every 5 frames, the training data is reduced by a factor of 5, allowing for faster learning.\nTo make it more efficient, we balanced the data according to the target distance. About 65% of the data had a target distance of less than 6 meters, and they were easy tasks. So in order to efficiently train the difficult data, we downsampled the data below 6 m. This halved the data. In the end, the number of data was roughly 80,000 iterations x 64 batch size.\n\n# Model architecture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F98319e9931caf1a99cbfdfdbf0c78237%2Ffinal_model.jpg?generation=1606994527055759&alt=media)\n\nIn addition to CNN, we used LSTM encoder and LSTM decoder.\n- Image size: 300x300\n- Optimizer: Adam\n- Scheduler: Onecyclelr\n- Num history: 10\n- Target: positions + cosine/sine of yaw\n\nFirst, we trained the model with the target distance balanced mask. Additional learning with the non-balanced mask. We used log10-scaled loss when additional training, because train/validate loss goes up and down with a large margin repeatedly. For two days of model training by using this method, we got a public score under 13, and a private score about 12.\n\n\n# Didn’t work\nhistory_positions more than 10 (Public score around 15, slow convergence)\nmetalabeling\nSeq2seq like LSTM decoder\n\n----------\n# Discussion\n\n## Stacking\n\nAs discussed in [this thread](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199531#1091632), we found that stacking boosted the score, but had no time left integrating the best model to stacking...\n\n## Satellite image\n\nUntil halfway through, I had mistakenly used satellite_debug as the 3 channel map information instead of semantic_map. Interestingly, this model got a public score of 13.3.\nI didn’t pursue it further because I realized the mistake, but it seems that satellite maps can score somewhat well in this task. Try it if you like it!!!\n\n## 3D-CNN(R2+1D)\n\nRasterized image data in the form of RGB×T×H×W, then fed them into R2+1D.\nValidation score was equivalent to resnet18, but the process of rasterizing and the 3D-CNN model were so heavy that the training required twice ~ triple as much time.\n\n## Only sequence model\n\nWe trained Seq2seq model not only with the target agent’s history position, but also with the surrounding vehicle’s history position. Despite no map information, this model got a validation score of 24. We integrated this trained seq2seq model into the trained resnet18 model, and fine-tuned the concat model. Sadly, improvement of the score was marginal.",
    "1103471": "Thanks for sharing @Inoichan with the summarized figure!"
  },
  "source": "meta"
}