{
  "id": 199494,
  "title": "22nd place journey : a completely different motion prediction approach",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199494",
  "author_name": "kkiller",
  "post_date": "2020-11-26T00:11:55.561000",
  "votes": 49,
  "comment_count": 19,
  "views": 0,
  "content": "<p>First of all, many thanks to the Kaggle team and Lyft team for hosting this competition, and congrats to all winners! Thanks to my teammates too, especially <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> for his hard work and commitment during all those 3 laborious months.</p>\n<h1>Preditcting agent motion with PointNet architecture …</h1>\n<p>Our ideas are mainly based on the pointnet architechture. We totally forgot about L5kit package and deal with the raw data which was transformed into a 4D tensors of shape (num_mini_scenes, max_agents_on_frame, num_backward_frames, num_features). Those tensors are concatenated over n_batches scenes and feed to the pointnet architecture. We was able to reach a score of 21.xx  by using a single pointnet model but things become harder and we got stuck.</p>\n<h1>Breakthrough ideas</h1>\n<p>To push our model performance a little bit, we manage to stack many pointnet models. This was doable since pointnet is very lightweight. Each model will look at a limited time step back to the agent history and output some embeddings of the scenes, the embeddging are then projected into a lower dimension space and concatenated. A simple full connected head is responsible for outputing the final 300+3 predictions.<br>\nOur best model is composed of 4 stacked models which resp. look at 10, 5, 3, and 1 frame back into the agent history. We use a simple zero padding for agents with no enough frames. With that giant pointnet model (~40 M params), we was able to reach a score of 13.353 on the public LB and 12.912 on the private.</p>\n<h1>Custom loss and training</h1>\n<p>We implement a custom version of the competition loss in which agent's loss is ignored when it leaves the scene. This custom loss allows us to try things like sample weight, penalization … We use <strong>pytorch-lightening</strong> to ease things. We maingly train on Colab, which is just owesome given the huge size of the competiion dataset. The optimizer is the classical Adam with a step learning rate scheduler, nothing fancy over there.</p>\n<h1>Things that doesn't work</h1>\n<ul>\n<li>Sample weight</li>\n<li>Bagging: we try a lot of ideas, and they  all fail :(</li>\n<li>RNN : we try a LSTM over time-stacked models without any success</li>\n<li>Longer history : augementing the history step leads to worse results (likely because of the many zeros coming from  our zero padding stratedy)</li>\n</ul>\n<h1>Advantages of our model</h1>\n<ul>\n<li>Very fast model, whole inference last 14 min's with single model and  less than 30 minutes with tens of stacked models</li>\n<li>Training on whole train_full  takes just 30' with a light pointnet model and 1h30' with 4 stacked models on a Colab Tesla-V100 single GPU</li>\n</ul>\n<h1>Cons of our model</h1>\n<ul>\n<li>Our pointnet implementation is completely road lanes blinded, even if we manage to incorporate light faces info in some extents</li>\n</ul>\n<h1>Things we may like to try</h1>\n<ul>\n<li>Moving from pointnet to other point-cloud models or voxel based models (pointCNN, point-Voxel, ShapeNet, …)</li>\n<li>Stacking &amp; transfer learning:  use our best models as embedders and train a simple model on top of them</li>\n<li>Combining our model with other image raster models (this one could make the pointnet road lanes aware)</li>\n<li>3D convolutions</li>\n</ul>\n<p>PS:  Our inference code by <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> is available <a href=\"https://www.kaggle.com/doanquanvietnamca/22st-solution-kkiller\" target=\"_blank\">here</a> .</p>",
  "messages": [
    {
      "id": 1091329,
      "postDate": "2020-11-26T00:11:55.560Z",
      "content": "<p>First of all, many thanks to the Kaggle team and Lyft team for hosting this competition, and congrats to all winners! Thanks to my teammates too, especially <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> for his hard work and commitment during all those 3 laborious months.</p>\n<h1>Preditcting agent motion with PointNet architecture …</h1>\n<p>Our ideas are mainly based on the pointnet architechture. We totally forgot about L5kit package and deal with the raw data which was transformed into a 4D tensors of shape (num_mini_scenes, max_agents_on_frame, num_backward_frames, num_features). Those tensors are concatenated over n_batches scenes and feed to the pointnet architecture. We was able to reach a score of 21.xx  by using a single pointnet model but things become harder and we got stuck.</p>\n<h1>Breakthrough ideas</h1>\n<p>To push our model performance a little bit, we manage to stack many pointnet models. This was doable since pointnet is very lightweight. Each model will look at a limited time step back to the agent history and output some embeddings of the scenes, the embeddging are then projected into a lower dimension space and concatenated. A simple full connected head is responsible for outputing the final 300+3 predictions.<br>\nOur best model is composed of 4 stacked models which resp. look at 10, 5, 3, and 1 frame back into the agent history. We use a simple zero padding for agents with no enough frames. With that giant pointnet model (~40 M params), we was able to reach a score of 13.353 on the public LB and 12.912 on the private.</p>\n<h1>Custom loss and training</h1>\n<p>We implement a custom version of the competition loss in which agent's loss is ignored when it leaves the scene. This custom loss allows us to try things like sample weight, penalization … We use <strong>pytorch-lightening</strong> to ease things. We maingly train on Colab, which is just owesome given the huge size of the competiion dataset. The optimizer is the classical Adam with a step learning rate scheduler, nothing fancy over there.</p>\n<h1>Things that doesn't work</h1>\n<ul>\n<li>Sample weight</li>\n<li>Bagging: we try a lot of ideas, and they  all fail :(</li>\n<li>RNN : we try a LSTM over time-stacked models without any success</li>\n<li>Longer history : augementing the history step leads to worse results (likely because of the many zeros coming from  our zero padding stratedy)</li>\n</ul>\n<h1>Advantages of our model</h1>\n<ul>\n<li>Very fast model, whole inference last 14 min's with single model and  less than 30 minutes with tens of stacked models</li>\n<li>Training on whole train_full  takes just 30' with a light pointnet model and 1h30' with 4 stacked models on a Colab Tesla-V100 single GPU</li>\n</ul>\n<h1>Cons of our model</h1>\n<ul>\n<li>Our pointnet implementation is completely road lanes blinded, even if we manage to incorporate light faces info in some extents</li>\n</ul>\n<h1>Things we may like to try</h1>\n<ul>\n<li>Moving from pointnet to other point-cloud models or voxel based models (pointCNN, point-Voxel, ShapeNet, …)</li>\n<li>Stacking &amp; transfer learning:  use our best models as embedders and train a simple model on top of them</li>\n<li>Combining our model with other image raster models (this one could make the pointnet road lanes aware)</li>\n<li>3D convolutions</li>\n</ul>\n<p>PS:  Our inference code by <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> is available <a href=\"https://www.kaggle.com/doanquanvietnamca/22st-solution-kkiller\" target=\"_blank\">here</a> .</p>",
      "rawMarkdown": "First of all, many thanks to the Kaggle team and Lyft team for hosting this competition, and congrats to all winners! Thanks to my teammates too, especially @doanquanvietnamca for his hard work and commitment during all those 3 laborious months.\n\n# Preditcting agent motion with PointNet architecture ...\nOur ideas are mainly based on the pointnet architechture. We totally forgot about L5kit package and deal with the raw data which was transformed into a 4D tensors of shape (num_mini_scenes, max_agents_on_frame, num_backward_frames, num_features). Those tensors are concatenated over n_batches scenes and feed to the pointnet architecture. We was able to reach a score of 21.xx  by using a single pointnet model but things become harder and we got stuck.\n\n\n# Breakthrough ideas\nTo push our model performance a little bit, we manage to stack many pointnet models. This was doable since pointnet is very lightweight. Each model will look at a limited time step back to the agent history and output some embeddings of the scenes, the embeddging are then projected into a lower dimension space and concatenated. A simple full connected head is responsible for outputing the final 300+3 predictions.\nOur best model is composed of 4 stacked models which resp. look at 10, 5, 3, and 1 frame back into the agent history. We use a simple zero padding for agents with no enough frames. With that giant pointnet model (~40 M params), we was able to reach a score of 13.353 on the public LB and 12.912 on the private.\n\n# Custom loss and training\nWe implement a custom version of the competition loss in which agent's loss is ignored when it leaves the scene. This custom loss allows us to try things like sample weight, penalization ... We use **pytorch-lightening** to ease things. We maingly train on Colab, which is just owesome given the huge size of the competiion dataset. The optimizer is the classical Adam with a step learning rate scheduler, nothing fancy over there.\n\n# Things that doesn't work\n* Sample weight\n* Bagging: we try a lot of ideas, and they  all fail :(\n* RNN : we try a LSTM over time-stacked models without any success\n* Longer history : augementing the history step leads to worse results (likely because of the many zeros coming from  our zero padding stratedy)\n\n\n# Advantages of our model\n* Very fast model, whole inference last 14 min's with single model and  less than 30 minutes with tens of stacked models\n* Training on whole train_full  takes just 30' with a light pointnet model and 1h30' with 4 stacked models on a Colab Tesla-V100 single GPU\n\n# Cons of our model\n* Our pointnet implementation is completely road lanes blinded, even if we manage to incorporate light faces info in some extents\n\n\n# Things we may like to try\n* Moving from pointnet to other point-cloud models or voxel based models (pointCNN, point-Voxel, ShapeNet, ...)\n* Stacking & transfer learning:  use our best models as embedders and train a simple model on top of them\n* Combining our model with other image raster models (this one could make the pointnet road lanes aware)\n* 3D convolutions\n\nPS:  Our inference code by @doanquanvietnamca is available [here](https://www.kaggle.com/doanquanvietnamca/22st-solution-kkiller) .",
      "votes": 49
    },
    {
      "id": 1091436,
      "postDate": "2020-11-26T02:51:44.510Z",
      "content": "<p>Good job. Base on your idea, I trying the PointVoxelNet, which outperformance the PointNet and PointNet++. Something wants to ask you but you not belong to my team 😄.</p>\n<p>I'm sharing here: <a href=\"https://www.kaggle.com/truonghoang/training-motion-prediction-with-pointvoxelnet\" target=\"_blank\">https://www.kaggle.com/truonghoang/training-motion-prediction-with-pointvoxelnet</a>.</p>\n<p>Congrats.</p>",
      "rawMarkdown": "Good job. Base on your idea, I trying the PointVoxelNet, which outperformance the PointNet and PointNet++. Something wants to ask you but you not belong to my team 😄.\n\nI'm sharing here: https://www.kaggle.com/truonghoang/training-motion-prediction-with-pointvoxelnet.\n\nCongrats.",
      "votes": 3,
      "replies": [
        {
          "id": 1091795,
          "postDate": "2020-11-26T09:38:20.993Z",
          "content": "<p>Nice extrapolation of our ideas over there :) </p>",
          "rawMarkdown": "Nice extrapolation of our ideas over there :) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1091375,
      "postDate": "2020-11-26T01:04:24.480Z",
      "content": "<p>Congrats on the results <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> and <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> </p>",
      "rawMarkdown": "Congrats on the results @kneroma and @doanquanvietnamca ",
      "votes": 3,
      "replies": [
        {
          "id": 1091796,
          "postDate": "2020-11-26T09:38:43.510Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
          "rawMarkdown": "Thanks @duykhanh99 "
        }
      ]
    },
    {
      "id": 1091350,
      "postDate": "2020-11-26T00:36:27.313Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> </p>",
      "rawMarkdown": "Great work @kneroma ",
      "votes": 3,
      "replies": [
        {
          "id": 1091361,
          "postDate": "2020-11-26T00:49:46Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> </p>",
          "rawMarkdown": "Thanks @ulrich07 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1091337,
      "postDate": "2020-11-26T00:23:48.307Z",
      "content": "<p>Wow…. stunning approach</p>",
      "rawMarkdown": "Wow.... stunning approach",
      "votes": 3,
      "replies": [
        {
          "id": 1091362,
          "postDate": "2020-11-26T00:50:01.753Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/seriousran\" target=\"_blank\">@seriousran</a> </p>",
          "rawMarkdown": "Thanks @seriousran ",
          "votes": 3
        }
      ]
    },
    {
      "id": 1091336,
      "postDate": "2020-11-26T00:21:50.740Z",
      "content": "<p>I want to point some thing: We use train_full dataset and get 1 epochs in 30mins with light version and 1h30min with stack version. That's so efficient model. After 5 epochs we can get 20x LB.</p>",
      "rawMarkdown": "I want to point some thing: We use train_full dataset and get 1 epochs in 30mins with light version and 1h30min with stack version. That's so efficient model. After 5 epochs we can get 20x LB.",
      "votes": 4,
      "replies": [
        {
          "id": 1091363,
          "postDate": "2020-11-26T00:51:42.387Z",
          "content": "<p>Wow! How do you even render the complete train_full in 30mins?</p>",
          "rawMarkdown": "Wow! How do you even render the complete train_full in 30mins?"
        },
        {
          "id": 1092484,
          "postDate": "2020-11-26T21:36:37.647Z",
          "content": "<p>we don't read every scene. and PointNet is very simple, don't need more compute calculation.</p>",
          "rawMarkdown": "we don't read every scene. and PointNet is very simple, don't need more compute calculation.\n"
        }
      ]
    },
    {
      "id": 1091330,
      "postDate": "2020-11-26T00:14:53.540Z",
      "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  hard work and genius teammate. I learned many thing when worked with him.</p>",
      "rawMarkdown": "@kneroma  hard work and genius teammate. I learned many thing when worked with him.",
      "votes": 4,
      "replies": [
        {
          "id": 1091335,
          "postDate": "2020-11-26T00:21:21.020Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> , I learn't a lot from you as well :) </p>",
          "rawMarkdown": "Thanks @doanquanvietnamca , I learn't a lot from you as well :) ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1091333,
      "postDate": "2020-11-26T00:19:27.787Z",
      "content": "<p>We tried time-related model like resnet3d, but the result was way worse than we expected…</p>",
      "rawMarkdown": "We tried time-related model like resnet3d, but the result was way worse than we expected...",
      "votes": 1
    },
    {
      "id": 1091331,
      "postDate": "2020-11-26T00:15:17.360Z",
      "content": "<p>Really interesting idea with PointNet architecture, Thank you for sharing and congratulation 👍</p>",
      "rawMarkdown": "Really interesting idea with PointNet architecture, Thank you for sharing and congratulation 👍",
      "votes": 1
    },
    {
      "id": 1091922,
      "postDate": "2020-11-26T11:34:22.740Z",
      "content": "<p>Fair play for taking such a different approach! Iterating through train_full in 30 mins!! Wow. Serious potential for innovation there. Well done.</p>",
      "rawMarkdown": "Fair play for taking such a different approach! Iterating through train_full in 30 mins!! Wow. Serious potential for innovation there. Well done.",
      "votes": 2
    },
    {
      "id": 1278947,
      "postDate": "2021-04-20T12:46:37.873Z",
      "content": "<p>Why are you using 3d based network with this type of dataset? </p>",
      "rawMarkdown": "Why are you using 3d based network with this type of dataset? "
    },
    {
      "id": 1103466,
      "postDate": "2020-12-05T23:43:49.343Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> , Non-image approach is quite interesting and may be beneficial for practically due to its fast training &amp; inference!</p>",
      "rawMarkdown": "Thanks for sharing @kneroma , Non-image approach is quite interesting and may be beneficial for practically due to its fast training & inference!"
    },
    {
      "id": 1091340,
      "postDate": "2020-11-26T00:29:02.467Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1091436,
      "author_name": "( ͡° ͜ʖ ͡°)",
      "author_url": "",
      "post_date": "2020-11-26T02:51:44.510000",
      "content": "<p>Good job. Base on your idea, I trying the PointVoxelNet, which outperformance the PointNet and PointNet++. Something wants to ask you but you not belong to my team 😄.</p>\n<p>I'm sharing here: <a href=\"https://www.kaggle.com/truonghoang/training-motion-prediction-with-pointvoxelnet\" target=\"_blank\">https://www.kaggle.com/truonghoang/training-motion-prediction-with-pointvoxelnet</a>.</p>\n<p>Congrats.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1091795,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-11-26T09:38:20.993000",
          "content": "<p>Nice extrapolation of our ideas over there :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1091375,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-11-26T01:04:24.480000",
      "content": "<p>Congrats on the results <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> and <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1091796,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-11-26T09:38:43.510000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091350,
      "author_name": "Ulrich G.",
      "author_url": "",
      "post_date": "2020-11-26T00:36:27.313000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1091361,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-11-26T00:49:46",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1091337,
      "author_name": "Chanran Kim",
      "author_url": "",
      "post_date": "2020-11-26T00:23:48.307000",
      "content": "<p>Wow…. stunning approach</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1091362,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-11-26T00:50:01.753000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/seriousran\" target=\"_blank\">@seriousran</a> </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1091336,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-11-26T00:21:50.740000",
      "content": "<p>I want to point some thing: We use train_full dataset and get 1 epochs in 30mins with light version and 1h30min with stack version. That's so efficient model. After 5 epochs we can get 20x LB.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1091363,
          "author_name": "Louis Yang",
          "author_url": "",
          "post_date": "2020-11-26T00:51:42.387000",
          "content": "<p>Wow! How do you even render the complete train_full in 30mins?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1092484,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2020-11-26T21:36:37.647000",
          "content": "<p>we don't read every scene. and PointNet is very simple, don't need more compute calculation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091330,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-11-26T00:14:53.540000",
      "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  hard work and genius teammate. I learned many thing when worked with him.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1091335,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-11-26T00:21:21.020000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> , I learn't a lot from you as well :) </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1091333,
      "author_name": "Yannik",
      "author_url": "",
      "post_date": "2020-11-26T00:19:27.787000",
      "content": "<p>We tried time-related model like resnet3d, but the result was way worse than we expected…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1091331,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-26T00:15:17.360000",
      "content": "<p>Really interesting idea with PointNet architecture, Thank you for sharing and congratulation 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1091922,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "2020-11-26T11:34:22.740000",
      "content": "<p>Fair play for taking such a different approach! Iterating through train_full in 30 mins!! Wow. Serious potential for innovation there. Well done.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1278947,
      "author_name": "FraMan",
      "author_url": "",
      "post_date": "2021-04-20T12:46:37.873000",
      "content": "<p>Why are you using 3d based network with this type of dataset? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1103466,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-12-05T23:43:49.343000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> , Non-image approach is quite interesting and may be beneficial for practically due to its fast training &amp; inference!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1091340,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-26T00:29:02.467000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091329": "First of all, many thanks to the Kaggle team and Lyft team for hosting this competition, and congrats to all winners! Thanks to my teammates too, especially @doanquanvietnamca for his hard work and commitment during all those 3 laborious months.\n\n# Preditcting agent motion with PointNet architecture ...\nOur ideas are mainly based on the pointnet architechture. We totally forgot about L5kit package and deal with the raw data which was transformed into a 4D tensors of shape (num_mini_scenes, max_agents_on_frame, num_backward_frames, num_features). Those tensors are concatenated over n_batches scenes and feed to the pointnet architecture. We was able to reach a score of 21.xx  by using a single pointnet model but things become harder and we got stuck.\n\n\n# Breakthrough ideas\nTo push our model performance a little bit, we manage to stack many pointnet models. This was doable since pointnet is very lightweight. Each model will look at a limited time step back to the agent history and output some embeddings of the scenes, the embeddging are then projected into a lower dimension space and concatenated. A simple full connected head is responsible for outputing the final 300+3 predictions.\nOur best model is composed of 4 stacked models which resp. look at 10, 5, 3, and 1 frame back into the agent history. We use a simple zero padding for agents with no enough frames. With that giant pointnet model (~40 M params), we was able to reach a score of 13.353 on the public LB and 12.912 on the private.\n\n# Custom loss and training\nWe implement a custom version of the competition loss in which agent's loss is ignored when it leaves the scene. This custom loss allows us to try things like sample weight, penalization ... We use **pytorch-lightening** to ease things. We maingly train on Colab, which is just owesome given the huge size of the competiion dataset. The optimizer is the classical Adam with a step learning rate scheduler, nothing fancy over there.\n\n# Things that doesn't work\n* Sample weight\n* Bagging: we try a lot of ideas, and they  all fail :(\n* RNN : we try a LSTM over time-stacked models without any success\n* Longer history : augementing the history step leads to worse results (likely because of the many zeros coming from  our zero padding stratedy)\n\n\n# Advantages of our model\n* Very fast model, whole inference last 14 min's with single model and  less than 30 minutes with tens of stacked models\n* Training on whole train_full  takes just 30' with a light pointnet model and 1h30' with 4 stacked models on a Colab Tesla-V100 single GPU\n\n# Cons of our model\n* Our pointnet implementation is completely road lanes blinded, even if we manage to incorporate light faces info in some extents\n\n\n# Things we may like to try\n* Moving from pointnet to other point-cloud models or voxel based models (pointCNN, point-Voxel, ShapeNet, ...)\n* Stacking & transfer learning:  use our best models as embedders and train a simple model on top of them\n* Combining our model with other image raster models (this one could make the pointnet road lanes aware)\n* 3D convolutions\n\nPS:  Our inference code by @doanquanvietnamca is available [here](https://www.kaggle.com/doanquanvietnamca/22st-solution-kkiller) .",
    "1091436": "Good job. Base on your idea, I trying the PointVoxelNet, which outperformance the PointNet and PointNet++. Something wants to ask you but you not belong to my team 😄.\n\nI'm sharing here: https://www.kaggle.com/truonghoang/training-motion-prediction-with-pointvoxelnet.\n\nCongrats.",
    "1091375": "Congrats on the results @kneroma and @doanquanvietnamca ",
    "1091350": "Great work @kneroma ",
    "1091337": "Wow.... stunning approach",
    "1091336": "I want to point some thing: We use train_full dataset and get 1 epochs in 30mins with light version and 1h30min with stack version. That's so efficient model. After 5 epochs we can get 20x LB.",
    "1091330": "@kneroma  hard work and genius teammate. I learned many thing when worked with him.",
    "1091333": "We tried time-related model like resnet3d, but the result was way worse than we expected...",
    "1091331": "Really interesting idea with PointNet architecture, Thank you for sharing and congratulation 👍",
    "1091922": "Fair play for taking such a different approach! Iterating through train_full in 30 mins!! Wow. Serious potential for innovation there. Well done.",
    "1278947": "Why are you using 3d based network with this type of dataset? ",
    "1103466": "Thanks for sharing @kneroma , Non-image approach is quite interesting and may be beneficial for practically due to its fast training & inference!",
    "1091340": ""
  }
}