{
  "id": 199588,
  "title": "6th place:  Micro-inputs, Lots of Data + Distance Order to Ensemble!",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/cyr-6th-place-micro-inputs-lots-of-data-distance-o",
  "author_name": "",
  "post_date": "2020-11-27T12:16:35.550Z",
  "votes": 31,
  "comment_count": 18,
  "views": 0,
  "content": "<p>First up: great competition! The data was about as stable as I’ve ever encountered on Kaggle – kudos to the organisers. Congratulations to all of the winners - and to everyone who worked hard and completed the competition. </p>\n<p>It was a pleasure to work with <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> and <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> on this – thank you both! </p>\n<p>I’m sure that my teammates will post separately, so I’ll stick with some observations from my own workflow here. Very interested to hear how everyone else approached the problem, as there was a lot of scope for different approaches! </p>\n<p>I ended up with two small input sizes in the interest of speed: 128 + 5 channels, 196 + 5 channels. </p>\n<ul>\n<li>The 128+5 model got to 11.9 after 68.8M samples. </li>\n<li>The 196+5 model reached 11.15 after 71.4M samples. </li>\n</ul>\n<p>Reading others’ comments and solutions it seems that some of the items that I thought were key to doing well were not, in fact! I’m sure cherry picking ideas that worked from different teams could lead to something very interesting.</p>\n<p><strong>Data, data, data:</strong></p>\n<p>There seemed to be one major key to success in this competition: how much data you could access, and how you sampled it. </p>\n<p>Early on it became clear that grouping scenes while training gave a significant uplift: when a single agent from each scene was selected before moving to the next iteration, scores improved.</p>\n<p>Taking this idea and following it through to train_full.zarr gave the next breakthrough: training on a chopped version of train_full.zarr yielded further improvements. </p>\n<p>The next jump in performance came from training on multiple chops of train_full.zarr. (You can adapt the l5kit code to create lightweight chops comprising just a small number of historic frames. This allowed for big savings on RAM/disk space.)</p>\n<p>We debated why there were such performance gains from training using chopped versions of train_full.zarr rather than randomly accessing indices. From early experiments it seemed like ensuring scene diversity was important: in the same way that we might enforce class balance while training, enforcing scene diversity (and perhaps more fundamentally, driver diversity?) seemed to matter. </p>\n<p><strong>Raster size v data coverage tradeoff:</strong></p>\n<p>Covering as much data as possible mattered, and training for as long as possible mattered. A single chop of train_full.zarr (approx 825K samples) could be shown to the model 12 times before it stopped learning. Obviously if you could show the model different samples you would do a lot better! But this seemed to be the hard limit.</p>\n<p>Keeping the model as small as possible meant it could iterate through these samples much more quickly. The inputs for the models that I ended up using were raster size 128 and 196, history_num_frames = 5, condensed into five input channels: </p>\n<p>sum(agent_history), agent_current, sum(ego_history), ego_current, sum(semantic_map)</p>\n<p>I was originally using Resnet18, but then switched to Resnest50 on <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a>'s recommendation and it gave an improvement of about -1 in nll. (Incidentally, <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> had the best performing single model in our group).</p>\n<p>Summing the semantic map meant that red and green traffic signals were treated the same. This seemed to work fine, surprisingly. Perhaps because only yellow gave additional information not already contained in the traffic movement.</p>\n<p><strong>Acceleration:</strong></p>\n<p>Eyeballing the predicted trajectories, it became clear that the models were ultimately making a bet on acceleration: typically, modes 0, 1, 2 represented trajectories arising from different agent speeds. </p>\n<p>Once it became clear that this was key it was possible to look at sampling the data such that we balanced these cases. </p>\n<p>A scene containing lots of agents was indicative of traffic. Gridlocked traffic obviously does not move much and leads to a lot of duplication in inputs. The models implemented a sampling scheme whereby scenes with a large number of agents were undersampled. The sampling proportion was: min(1, 7/agent_count). Thus, scenes containing 14 agents had 50% of those agents selected for each training iteration, etc.</p>\n<p><strong>Ensembling:</strong></p>\n<p>This was initially a tricky one: averaging models based on confidence values didn’t work. However, when the importance of acceleration was taken into account, the route to ensembling made sense: order the model modes by distance covered, then average the results. When this was implemented all models could be ensembled very quickly, with positive results. We also looked at incorporating curvature here, but it didn’t make any difference: distance was the key.</p>\n<p>The final ensemble optimized weights on the validation set. The weighting scheme incorporated distance and confidence values. </p>\n<p><strong>Code:</strong></p>\n<p>Code for contribution to our team solution can be found <a href=\"https://github.com/ciararogerson/Kaggle_Lyft\" target=\"_blank\">here</a></p>\n<p><strong>Ideas that didn’t work: many! Here are a few…</strong></p>\n<p><em>Traffic lights:</em></p>\n<p>I couldn’t get additional traffic light information to add anything: many different angles were tried! The most promising was probably traffic light persistence: we included an additional channel in the model where instead of traffic light lane lines, we drew lines containing the number of frames since the traffic light had turned to its current colour (up to a maximum of 100). Ultimately this didn’t add anything.</p>\n<p><em>Day/hour:</em></p>\n<p>Adding channels for day/hour values proved better than concatenating them directly before the dense layers of the model, but it still didn’t help much.</p>\n<p><em>Interpolation:</em></p>\n<p>Having the penultimate model layer output 25 sets of (x, y) points, followed by an interpolation between these points to make the final 50 point trajectory did not work.</p>\n<p><em>Weighted loss function:</em></p>\n<p>Given the importance of capturing acceleration, we tried a version of the loss function that weighted  the last 10 points of the trajectory equal to the first 40. This didn’t help.</p>",
  "messages": [
    {
      "id": "1091863",
      "postDate": "11/26/2020 10:47:37",
      "content": "<p>First up: great competition! The data was about as stable as I’ve ever encountered on Kaggle – kudos to the organisers. Congratulations to all of the winners - and to everyone who worked hard and completed the competition. </p>\n<p>It was a pleasure to work with <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> and <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> on this – thank you both! </p>\n<p>I’m sure that my teammates will post separately, so I’ll stick with some observations from my own workflow here. Very interested to hear how everyone else approached the problem, as there was a lot of scope for different approaches! </p>\n<p>I ended up with two small input sizes in the interest of speed: 128 + 5 channels, 196 + 5 channels. </p>\n<ul>\n<li>The 128+5 model got to 11.9 after 68.8M samples. </li>\n<li>The 196+5 model reached 11.15 after 71.4M samples. </li>\n</ul>\n<p>Reading others’ comments and solutions it seems that some of the items that I thought were key to doing well were not, in fact! I’m sure cherry picking ideas that worked from different teams could lead to something very interesting.</p>\n<p><strong>Data, data, data:</strong></p>\n<p>There seemed to be one major key to success in this competition: how much data you could access, and how you sampled it. </p>\n<p>Early on it became clear that grouping scenes while training gave a significant uplift: when a single agent from each scene was selected before moving to the next iteration, scores improved.</p>\n<p>Taking this idea and following it through to train_full.zarr gave the next breakthrough: training on a chopped version of train_full.zarr yielded further improvements. </p>\n<p>The next jump in performance came from training on multiple chops of train_full.zarr. (You can adapt the l5kit code to create lightweight chops comprising just a small number of historic frames. This allowed for big savings on RAM/disk space.)</p>\n<p>We debated why there were such performance gains from training using chopped versions of train_full.zarr rather than randomly accessing indices. From early experiments it seemed like ensuring scene diversity was important: in the same way that we might enforce class balance while training, enforcing scene diversity (and perhaps more fundamentally, driver diversity?) seemed to matter. </p>\n<p><strong>Raster size v data coverage tradeoff:</strong></p>\n<p>Covering as much data as possible mattered, and training for as long as possible mattered. A single chop of train_full.zarr (approx 825K samples) could be shown to the model 12 times before it stopped learning. Obviously if you could show the model different samples you would do a lot better! But this seemed to be the hard limit.</p>\n<p>Keeping the model as small as possible meant it could iterate through these samples much more quickly. The inputs for the models that I ended up using were raster size 128 and 196, history_num_frames = 5, condensed into five input channels: </p>\n<p>sum(agent_history), agent_current, sum(ego_history), ego_current, sum(semantic_map)</p>\n<p>I was originally using Resnet18, but then switched to Resnest50 on <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a>'s recommendation and it gave an improvement of about -1 in nll. (Incidentally, <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> had the best performing single model in our group).</p>\n<p>Summing the semantic map meant that red and green traffic signals were treated the same. This seemed to work fine, surprisingly. Perhaps because only yellow gave additional information not already contained in the traffic movement.</p>\n<p><strong>Acceleration:</strong></p>\n<p>Eyeballing the predicted trajectories, it became clear that the models were ultimately making a bet on acceleration: typically, modes 0, 1, 2 represented trajectories arising from different agent speeds. </p>\n<p>Once it became clear that this was key it was possible to look at sampling the data such that we balanced these cases. </p>\n<p>A scene containing lots of agents was indicative of traffic. Gridlocked traffic obviously does not move much and leads to a lot of duplication in inputs. The models implemented a sampling scheme whereby scenes with a large number of agents were undersampled. The sampling proportion was: min(1, 7/agent_count). Thus, scenes containing 14 agents had 50% of those agents selected for each training iteration, etc.</p>\n<p><strong>Ensembling:</strong></p>\n<p>This was initially a tricky one: averaging models based on confidence values didn’t work. However, when the importance of acceleration was taken into account, the route to ensembling made sense: order the model modes by distance covered, then average the results. When this was implemented all models could be ensembled very quickly, with positive results. We also looked at incorporating curvature here, but it didn’t make any difference: distance was the key.</p>\n<p>The final ensemble optimized weights on the validation set. The weighting scheme incorporated distance and confidence values. </p>\n<p><strong>Code:</strong></p>\n<p>Code for contribution to our team solution can be found <a href=\"https://github.com/ciararogerson/Kaggle_Lyft\" target=\"_blank\">here</a></p>\n<p><strong>Ideas that didn’t work: many! Here are a few…</strong></p>\n<p><em>Traffic lights:</em></p>\n<p>I couldn’t get additional traffic light information to add anything: many different angles were tried! The most promising was probably traffic light persistence: we included an additional channel in the model where instead of traffic light lane lines, we drew lines containing the number of frames since the traffic light had turned to its current colour (up to a maximum of 100). Ultimately this didn’t add anything.</p>\n<p><em>Day/hour:</em></p>\n<p>Adding channels for day/hour values proved better than concatenating them directly before the dense layers of the model, but it still didn’t help much.</p>\n<p><em>Interpolation:</em></p>\n<p>Having the penultimate model layer output 25 sets of (x, y) points, followed by an interpolation between these points to make the final 50 point trajectory did not work.</p>\n<p><em>Weighted loss function:</em></p>\n<p>Given the importance of capturing acceleration, we tried a version of the loss function that weighted  the last 10 points of the trajectory equal to the first 40. This didn’t help.</p>",
      "rawMarkdown": "First up: great competition! The data was about as stable as I’ve ever encountered on Kaggle – kudos to the organisers. Congratulations to all of the winners - and to everyone who worked hard and completed the competition. \n\nIt was a pleasure to work with @rytisva88 and @sheriytm on this – thank you both! \n\nI’m sure that my teammates will post separately, so I’ll stick with some observations from my own workflow here. Very interested to hear how everyone else approached the problem, as there was a lot of scope for different approaches! \n\nI ended up with two small input sizes in the interest of speed: 128 + 5 channels, 196 + 5 channels. \n- The 128+5 model got to 11.9 after 68.8M samples. \n- The 196+5 model reached 11.15 after 71.4M samples. \n\nReading others’ comments and solutions it seems that some of the items that I thought were key to doing well were not, in fact! I’m sure cherry picking ideas that worked from different teams could lead to something very interesting.\n\n\n**Data, data, data:**\n\nThere seemed to be one major key to success in this competition: how much data you could access, and how you sampled it. \n\nEarly on it became clear that grouping scenes while training gave a significant uplift: when a single agent from each scene was selected before moving to the next iteration, scores improved.\n\nTaking this idea and following it through to train_full.zarr gave the next breakthrough: training on a chopped version of train_full.zarr yielded further improvements. \n\nThe next jump in performance came from training on multiple chops of train_full.zarr. (You can adapt the l5kit code to create lightweight chops comprising just a small number of historic frames. This allowed for big savings on RAM/disk space.)\n\nWe debated why there were such performance gains from training using chopped versions of train_full.zarr rather than randomly accessing indices. From early experiments it seemed like ensuring scene diversity was important: in the same way that we might enforce class balance while training, enforcing scene diversity (and perhaps more fundamentally, driver diversity?) seemed to matter. \n\n\n**Raster size v data coverage tradeoff:**\n\nCovering as much data as possible mattered, and training for as long as possible mattered. A single chop of train_full.zarr (approx 825K samples) could be shown to the model 12 times before it stopped learning. Obviously if you could show the model different samples you would do a lot better! But this seemed to be the hard limit.\n\nKeeping the model as small as possible meant it could iterate through these samples much more quickly. The inputs for the models that I ended up using were raster size 128 and 196, history_num_frames = 5, condensed into five input channels: \n\nsum(agent_history), agent_current, sum(ego_history), ego_current, sum(semantic_map)\n\nI was originally using Resnet18, but then switched to Resnest50 on @rytisva88's recommendation and it gave an improvement of about -1 in nll. (Incidentally, @rytisva88 had the best performing single model in our group).\n\nSumming the semantic map meant that red and green traffic signals were treated the same. This seemed to work fine, surprisingly. Perhaps because only yellow gave additional information not already contained in the traffic movement.\n\n\n**Acceleration:**\n\nEyeballing the predicted trajectories, it became clear that the models were ultimately making a bet on acceleration: typically, modes 0, 1, 2 represented trajectories arising from different agent speeds. \n\nOnce it became clear that this was key it was possible to look at sampling the data such that we balanced these cases. \n\nA scene containing lots of agents was indicative of traffic. Gridlocked traffic obviously does not move much and leads to a lot of duplication in inputs. The models implemented a sampling scheme whereby scenes with a large number of agents were undersampled. The sampling proportion was: min(1, 7/agent_count). Thus, scenes containing 14 agents had 50% of those agents selected for each training iteration, etc.\n\n\n**Ensembling:**\n\nThis was initially a tricky one: averaging models based on confidence values didn’t work. However, when the importance of acceleration was taken into account, the route to ensembling made sense: order the model modes by distance covered, then average the results. When this was implemented all models could be ensembled very quickly, with positive results. We also looked at incorporating curvature here, but it didn’t make any difference: distance was the key.\n\nThe final ensemble optimized weights on the validation set. The weighting scheme incorporated distance and confidence values. \n\n\n**Code:**\n\nCode for contribution to our team solution can be found [here](https://github.com/ciararogerson/Kaggle_Lyft)\n\n\n**Ideas that didn’t work: many! Here are a few...**\n\n_Traffic lights:_\n\nI couldn’t get additional traffic light information to add anything: many different angles were tried! The most promising was probably traffic light persistence: we included an additional channel in the model where instead of traffic light lane lines, we drew lines containing the number of frames since the traffic light had turned to its current colour (up to a maximum of 100). Ultimately this didn’t add anything.\n\n_Day/hour:_\n\nAdding channels for day/hour values proved better than concatenating them directly before the dense layers of the model, but it still didn’t help much.\n\n_Interpolation:_\n\nHaving the penultimate model layer output 25 sets of (x, y) points, followed by an interpolation between these points to make the final 50 point trajectory did not work.\n\n_Weighted loss function:_\n\nGiven the importance of capturing acceleration, we tried a version of the loss function that weighted  the last 10 points of the trajectory equal to the first 40. This didn’t help.",
      "votes": null
    },
    {
      "id": "1091866",
      "postDate": "11/26/2020 10:51:08",
      "content": "<p>WHat was your hardware</p>",
      "rawMarkdown": "WHat was your hardware",
      "votes": null
    },
    {
      "id": "1091876",
      "postDate": "11/26/2020 10:55:00",
      "content": "<p>3x1080ti. 128GB RAM, 10 cores. Could have done with more!!</p>",
      "rawMarkdown": "3x1080ti. 128GB RAM, 10 cores. Could have done with more!!",
      "votes": null
    },
    {
      "id": "1091884",
      "postDate": "11/26/2020 10:57:32",
      "content": "<p>Great writeup, look forward to other members responses. I also chopped train_full, initially at frame 100 then slowly expanded to include additional frames, and saw consistent improvement. After chopping did you continue to use rasterization on the fly while training or did you store rasters on disk? How long did training take and what hardware did you use?</p>\n<p>Congratulations :)</p>",
      "rawMarkdown": "Great writeup, look forward to other members responses. I also chopped train_full, initially at frame 100 then slowly expanded to include additional frames, and saw consistent improvement. After chopping did you continue to use rasterization on the fly while training or did you store rasters on disk? How long did training take and what hardware did you use?\n\nCongratulations :)",
      "votes": null
    },
    {
      "id": "1091886",
      "postDate": "11/26/2020 11:00:15",
      "content": "<p><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588#1091876\" target=\"_blank\">Here they mention about their hardware</a></p>",
      "rawMarkdown": "[Here they mention about their hardware](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588#1091876)",
      "votes": null
    },
    {
      "id": "1091887",
      "postDate": "11/26/2020 11:01:04",
      "content": "<p>Yes, I kept the rasterization throughout as I was so often playing around with different configs that it didn't make sense to freeze it. Also, I wasn't sure whether I'd have the disk space… </p>",
      "rawMarkdown": "Yes, I kept the rasterization throughout as I was so often playing around with different configs that it didn't make sense to freeze it. Also, I wasn't sure whether I'd have the disk space...",
      "votes": null
    },
    {
      "id": "1091903",
      "postDate": "11/26/2020 11:12:59",
      "content": "<p>Thanks for the writeup! And congratulations!<br>\nSome smart ideas you had like using the daytime in the model and the summing trick</p>\n<p>I see that you experienced large performance gains while using a subsampled train_full over a random sampled train_full. Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?<br>\nIt makes sense, that subsampling most diverse samples can get the validation score down quicker, but I would be very interested if it actually makes a difference how you shuffle after a full epoch.<br>\nWhy did you use the same subsample again for a second epoch? Why not e.g. epoch0_frame +1?</p>",
      "rawMarkdown": "Thanks for the writeup! And congratulations!\nSome smart ideas you had like using the daytime in the model and the summing trick\n\nI see that you experienced large performance gains while using a subsampled train_full over a random sampled train_full. Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?\nIt makes sense, that subsampling most diverse samples can get the validation score down quicker, but I would be very interested if it actually makes a difference how you shuffle after a full epoch.\nWhy did you use the same subsample again for a second epoch? Why not e.g. epoch0_frame +1?",
      "votes": null
    },
    {
      "id": "1091912",
      "postDate": "11/26/2020 11:21:45",
      "content": "<p>Thanks! Same to you!</p>\n<p>No, I never trained train_full for a complete epoch, I think I'd be here until Christmas if I did! <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a>  had been training on a random shuffle of train_full before switching to the chopped version, so he will have a better gauge on what the impact was. <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> also went from using train.zarr to a chop of train_full.zarr so will have an idea of how much that was worth.</p>\n<p>We ended up reusing chopped datasets (i.e. showing the same sample multiple times). It would have been better if we didn't, but we were running out of time and were resource constrained. </p>\n<p>If I were starting over I would have done it differently: rather than creating complete chops I would just store off the indices in train_full.zarr that those chops refer to, thereby allowing us to create sequences of diverse samples without having to show samples multiple times or run out of resources. <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> managed resources better and he was able to use 20 chopped datasets (but could have gone higher). I was stuck using 16.</p>",
      "rawMarkdown": "Thanks! Same to you!\n\nNo, I never trained train_full for a complete epoch, I think I'd be here until Christmas if I did! @rytisva88  had been training on a random shuffle of train_full before switching to the chopped version, so he will have a better gauge on what the impact was. @sheriytm also went from using train.zarr to a chop of train_full.zarr so will have an idea of how much that was worth.\n\nWe ended up reusing chopped datasets (i.e. showing the same sample multiple times). It would have been better if we didn't, but we were running out of time and were resource constrained. \n\nIf I were starting over I would have done it differently: rather than creating complete chops I would just store off the indices in train_full.zarr that those chops refer to, thereby allowing us to create sequences of diverse samples without having to show samples multiple times or run out of resources. @rytisva88 managed resources better and he was able to use 20 chopped datasets (but could have gone higher). I was stuck using 16.",
      "votes": null
    },
    {
      "id": "1091967",
      "postDate": "11/26/2020 12:28:44",
      "content": "<p>Great result with 3 1080ti !</p>",
      "rawMarkdown": "Great result with 3 1080ti !",
      "votes": null
    },
    {
      "id": "1092011",
      "postDate": "11/26/2020 13:26:59",
      "content": "<p>Wow. Great job and Thanks for sharing</p>",
      "rawMarkdown": "Wow. Great job and Thanks for sharing",
      "votes": null
    },
    {
      "id": "1092047",
      "postDate": "11/26/2020 13:55:10",
      "content": "<p>Congratulations! Could I ask that how much improvement you got using the ensembling scheme you has mentioned?</p>",
      "rawMarkdown": "Congratulations! Could I ask that how much improvement you got using the ensembling scheme you has mentioned?",
      "votes": null
    },
    {
      "id": "1092049",
      "postDate": "11/26/2020 13:55:12",
      "content": "<p>Very thorough and impressive.</p>",
      "rawMarkdown": "Very thorough and impressive.",
      "votes": null
    },
    {
      "id": "1092203",
      "postDate": "11/26/2020 15:48:17",
      "content": "<p>Did you ever try something like applying k-means to ensemble your trajectories? In theory similar to your distance covered ensembling, but potentially more automatic and you can also apply sample weights directly in sklearn based on the confidence output. </p>",
      "rawMarkdown": "Did you ever try something like applying k-means to ensemble your trajectories? In theory similar to your distance covered ensembling, but potentially more automatic and you can also apply sample weights directly in sklearn based on the confidence output.",
      "votes": null
    },
    {
      "id": "1092206",
      "postDate": "11/26/2020 15:53:56",
      "content": "<p>I did try kmeans early on, but it wasn't working for me, perhaps because of how I implemented it: if there were three models, say, I was clustering 9 modes into three clusters. Then taking the max conf mode for each cluster or averaging within each cluster (I tried both). But obviously had something wrong somewhere…</p>",
      "rawMarkdown": "I did try kmeans early on, but it wasn't working for me, perhaps because of how I implemented it: if there were three models, say, I was clustering 9 modes into three clusters. Then taking the max conf mode for each cluster or averaging within each cluster (I tried both). But obviously had something wrong somewhere...",
      "votes": null
    },
    {
      "id": "1092211",
      "postDate": "11/26/2020 15:57:45",
      "content": "<p>The top score was an ensemble of the two best models (slightly lower score was an ensemble of four). For that one, 11.15 + 10.52 -&gt; 10.32</p>",
      "rawMarkdown": "The top score was an ensemble of the two best models (slightly lower score was an ensemble of four). For that one, 11.15 + 10.52 -> 10.32",
      "votes": null
    },
    {
      "id": "1095464",
      "postDate": "11/29/2020 16:08:50",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> ! Very Insightful 👍</p>",
      "rawMarkdown": "Thanks for sharing @fergusoci ! Very Insightful 👍",
      "votes": null
    },
    {
      "id": "1095737",
      "postDate": "11/29/2020 22:26:31",
      "content": "<p>Exhausted from the competition, my day job and Thanksgiving holiday stuff this had to wait and I apologize for the delayed post.</p>\n<p>My first thanks is to the competition organizers and the Kaggle team for this really interesting contest. Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> a.k.a.  team NIPD for a well deserved win.</p>\n<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> have pretty much presented the team efforts so well I am just going to add a few points here instead of a separate post. First I want to thank my team mates <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> and <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> for making this competition experience a pleasure.</p>\n<p>Since much far useful solutions from the top teams are already posted, I will give a brief on some of my training procedures:-</p>\n<p>I have limited GPU resource hence could not use train_full.zarr before teaming up. The GCP credit was unusable as GPU is not available where I am on my travels at this time. My best model before merger was a Resnet34 (224+5) trained on train.zarr to a reasonable LB score of 21.403 before it started to overfit. Training this model on the chopped data made improvements. </p>\n<p>Once we merged looking at all the teams models and various efforts after brain storming, I turned my attention to training a different type of model to add diversity. After a few (i.e. PointNet, EfficientNet, etc) trials, settled on training a Resnext50_32x4d (224+5) with train.zarr and chopped train_full.zarr to help add diversity to our ensemble. Basically, after training on 2.5M samples of train.zarr, I alternated by training between train.zarr and chopped 100 train_full.zarr. Each cycle trains on 256000 samples of the data. That made the scores go down significantly but got slower after reaching the sub 20 LB scores. At this point, I continued training on the chopped data only, which reached sub 14 scores after going through the data 2.5 times at competition close. I have continued training this model to see if it can get to the scores of our best models and make a late submission.</p>\n<p>And thanks everyone who contributed in many ways via kernels and discussions.</p>",
      "rawMarkdown": "Exhausted from the competition, my day job and Thanksgiving holiday stuff this had to wait and I apologize for the delayed post.\n\nMy first thanks is to the competition organizers and the Kaggle team for this really interesting contest. Congratulations @philippsinger, @christofhenkel, @ilu000, @nvnnghia a.k.a.  team NIPD for a well deserved win.\n\n@fergusoci have pretty much presented the team efforts so well I am just going to add a few points here instead of a separate post. First I want to thank my team mates @fergusoci and @rytisva88 for making this competition experience a pleasure.\n\nSince much far useful solutions from the top teams are already posted, I will give a brief on some of my training procedures:-\n\nI have limited GPU resource hence could not use train_full.zarr before teaming up. The GCP credit was unusable as GPU is not available where I am on my travels at this time. My best model before merger was a Resnet34 (224+5) trained on train.zarr to a reasonable LB score of 21.403 before it started to overfit. Training this model on the chopped data made improvements. \n\nOnce we merged looking at all the teams models and various efforts after brain storming, I turned my attention to training a different type of model to add diversity. After a few (i.e. PointNet, EfficientNet, etc) trials, settled on training a Resnext50_32x4d (224+5) with train.zarr and chopped train_full.zarr to help add diversity to our ensemble. Basically, after training on 2.5M samples of train.zarr, I alternated by training between train.zarr and chopped 100 train_full.zarr. Each cycle trains on 256000 samples of the data. That made the scores go down significantly but got slower after reaching the sub 20 LB scores. At this point, I continued training on the chopped data only, which reached sub 14 scores after going through the data 2.5 times at competition close. I have continued training this model to see if it can get to the scores of our best models and make a late submission.\n\nAnd thanks everyone who contributed in many ways via kernels and discussions.",
      "votes": null
    },
    {
      "id": "1095739",
      "postDate": "11/29/2020 22:30:33",
      "content": "<blockquote>\n  <p>Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?</p>\n</blockquote>\n<p>No <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, in my case that was not an option due GPU challenged. See my post in the comment section here for how I went about training my models.</p>",
      "rawMarkdown": "> Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?\n\nNo @ilu000, in my case that was not an option due GPU challenged. See my post in the comment section here for how I went about training my models.",
      "votes": null
    },
    {
      "id": "1103461",
      "postDate": "12/05/2020 23:37:01",
      "content": "<p>Thanks for sharing detailed write-up <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> , congrats! Ensemble method is interesting.</p>",
      "rawMarkdown": "Thanks for sharing detailed write-up @fergusoci , congrats! Ensemble method is interesting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091866,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "11/26/2020 10:51:08",
      "content": "<p>WHat was your hardware</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091876,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/26/2020 10:55:00",
          "content": "<p>3x1080ti. 128GB RAM, 10 cores. Could have done with more!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091967,
          "author_name": "dmytropoplavskiy",
          "author_url": "",
          "post_date": "11/26/2020 12:28:44",
          "content": "<p>Great result with 3 1080ti !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092011,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "11/26/2020 13:26:59",
          "content": "<p>Wow. Great job and Thanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091884,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "11/26/2020 10:57:32",
      "content": "<p>Great writeup, look forward to other members responses. I also chopped train_full, initially at frame 100 then slowly expanded to include additional frames, and saw consistent improvement. After chopping did you continue to use rasterization on the fly while training or did you store rasters on disk? How long did training take and what hardware did you use?</p>\n<p>Congratulations :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091886,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "11/26/2020 11:00:15",
          "content": "<p><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588#1091876\" target=\"_blank\">Here they mention about their hardware</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091887,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/26/2020 11:01:04",
          "content": "<p>Yes, I kept the rasterization throughout as I was so often playing around with different configs that it didn't make sense to freeze it. Also, I wasn't sure whether I'd have the disk space… </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091903,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "11/26/2020 11:12:59",
      "content": "<p>Thanks for the writeup! And congratulations!<br>\nSome smart ideas you had like using the daytime in the model and the summing trick</p>\n<p>I see that you experienced large performance gains while using a subsampled train_full over a random sampled train_full. Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?<br>\nIt makes sense, that subsampling most diverse samples can get the validation score down quicker, but I would be very interested if it actually makes a difference how you shuffle after a full epoch.<br>\nWhy did you use the same subsample again for a second epoch? Why not e.g. epoch0_frame +1?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091912,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/26/2020 11:21:45",
          "content": "<p>Thanks! Same to you!</p>\n<p>No, I never trained train_full for a complete epoch, I think I'd be here until Christmas if I did! <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a>  had been training on a random shuffle of train_full before switching to the chopped version, so he will have a better gauge on what the impact was. <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> also went from using train.zarr to a chop of train_full.zarr so will have an idea of how much that was worth.</p>\n<p>We ended up reusing chopped datasets (i.e. showing the same sample multiple times). It would have been better if we didn't, but we were running out of time and were resource constrained. </p>\n<p>If I were starting over I would have done it differently: rather than creating complete chops I would just store off the indices in train_full.zarr that those chops refer to, thereby allowing us to create sequences of diverse samples without having to show samples multiple times or run out of resources. <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> managed resources better and he was able to use 20 chopped datasets (but could have gone higher). I was stuck using 16.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095739,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "11/29/2020 22:30:33",
          "content": "<blockquote>\n  <p>Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?</p>\n</blockquote>\n<p>No <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, in my case that was not an option due GPU challenged. See my post in the comment section here for how I went about training my models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092047,
      "author_name": "hardworkingkaggler",
      "author_url": "",
      "post_date": "11/26/2020 13:55:10",
      "content": "<p>Congratulations! Could I ask that how much improvement you got using the ensembling scheme you has mentioned?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092211,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/26/2020 15:57:45",
          "content": "<p>The top score was an ensemble of the two best models (slightly lower score was an ensemble of four). For that one, 11.15 + 10.52 -&gt; 10.32</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092049,
      "author_name": "authman",
      "author_url": "",
      "post_date": "11/26/2020 13:55:12",
      "content": "<p>Very thorough and impressive.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092203,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/26/2020 15:48:17",
      "content": "<p>Did you ever try something like applying k-means to ensemble your trajectories? In theory similar to your distance covered ensembling, but potentially more automatic and you can also apply sample weights directly in sklearn based on the confidence output. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1092206,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/26/2020 15:53:56",
          "content": "<p>I did try kmeans early on, but it wasn't working for me, perhaps because of how I implemented it: if there were three models, say, I was clustering 9 modes into three clusters. Then taking the max conf mode for each cluster or averaging within each cluster (I tried both). But obviously had something wrong somewhere…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1095464,
      "author_name": "datajameson",
      "author_url": "",
      "post_date": "11/29/2020 16:08:50",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> ! Very Insightful 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1095737,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/29/2020 22:26:31",
      "content": "<p>Exhausted from the competition, my day job and Thanksgiving holiday stuff this had to wait and I apologize for the delayed post.</p>\n<p>My first thanks is to the competition organizers and the Kaggle team for this really interesting contest. Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> a.k.a.  team NIPD for a well deserved win.</p>\n<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> have pretty much presented the team efforts so well I am just going to add a few points here instead of a separate post. First I want to thank my team mates <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> and <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a> for making this competition experience a pleasure.</p>\n<p>Since much far useful solutions from the top teams are already posted, I will give a brief on some of my training procedures:-</p>\n<p>I have limited GPU resource hence could not use train_full.zarr before teaming up. The GCP credit was unusable as GPU is not available where I am on my travels at this time. My best model before merger was a Resnet34 (224+5) trained on train.zarr to a reasonable LB score of 21.403 before it started to overfit. Training this model on the chopped data made improvements. </p>\n<p>Once we merged looking at all the teams models and various efforts after brain storming, I turned my attention to training a different type of model to add diversity. After a few (i.e. PointNet, EfficientNet, etc) trials, settled on training a Resnext50_32x4d (224+5) with train.zarr and chopped train_full.zarr to help add diversity to our ensemble. Basically, after training on 2.5M samples of train.zarr, I alternated by training between train.zarr and chopped 100 train_full.zarr. Each cycle trains on 256000 samples of the data. That made the scores go down significantly but got slower after reaching the sub 20 LB scores. At this point, I continued training on the chopped data only, which reached sub 14 scores after going through the data 2.5 times at competition close. I have continued training this model to see if it can get to the scores of our best models and make a late submission.</p>\n<p>And thanks everyone who contributed in many ways via kernels and discussions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103461,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "12/05/2020 23:37:01",
      "content": "<p>Thanks for sharing detailed write-up <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> , congrats! Ensemble method is interesting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091863": "First up: great competition! The data was about as stable as I’ve ever encountered on Kaggle – kudos to the organisers. Congratulations to all of the winners - and to everyone who worked hard and completed the competition. \n\nIt was a pleasure to work with @rytisva88 and @sheriytm on this – thank you both! \n\nI’m sure that my teammates will post separately, so I’ll stick with some observations from my own workflow here. Very interested to hear how everyone else approached the problem, as there was a lot of scope for different approaches! \n\nI ended up with two small input sizes in the interest of speed: 128 + 5 channels, 196 + 5 channels. \n- The 128+5 model got to 11.9 after 68.8M samples. \n- The 196+5 model reached 11.15 after 71.4M samples. \n\nReading others’ comments and solutions it seems that some of the items that I thought were key to doing well were not, in fact! I’m sure cherry picking ideas that worked from different teams could lead to something very interesting.\n\n\n**Data, data, data:**\n\nThere seemed to be one major key to success in this competition: how much data you could access, and how you sampled it. \n\nEarly on it became clear that grouping scenes while training gave a significant uplift: when a single agent from each scene was selected before moving to the next iteration, scores improved.\n\nTaking this idea and following it through to train_full.zarr gave the next breakthrough: training on a chopped version of train_full.zarr yielded further improvements. \n\nThe next jump in performance came from training on multiple chops of train_full.zarr. (You can adapt the l5kit code to create lightweight chops comprising just a small number of historic frames. This allowed for big savings on RAM/disk space.)\n\nWe debated why there were such performance gains from training using chopped versions of train_full.zarr rather than randomly accessing indices. From early experiments it seemed like ensuring scene diversity was important: in the same way that we might enforce class balance while training, enforcing scene diversity (and perhaps more fundamentally, driver diversity?) seemed to matter. \n\n\n**Raster size v data coverage tradeoff:**\n\nCovering as much data as possible mattered, and training for as long as possible mattered. A single chop of train_full.zarr (approx 825K samples) could be shown to the model 12 times before it stopped learning. Obviously if you could show the model different samples you would do a lot better! But this seemed to be the hard limit.\n\nKeeping the model as small as possible meant it could iterate through these samples much more quickly. The inputs for the models that I ended up using were raster size 128 and 196, history_num_frames = 5, condensed into five input channels: \n\nsum(agent_history), agent_current, sum(ego_history), ego_current, sum(semantic_map)\n\nI was originally using Resnet18, but then switched to Resnest50 on @rytisva88's recommendation and it gave an improvement of about -1 in nll. (Incidentally, @rytisva88 had the best performing single model in our group).\n\nSumming the semantic map meant that red and green traffic signals were treated the same. This seemed to work fine, surprisingly. Perhaps because only yellow gave additional information not already contained in the traffic movement.\n\n\n**Acceleration:**\n\nEyeballing the predicted trajectories, it became clear that the models were ultimately making a bet on acceleration: typically, modes 0, 1, 2 represented trajectories arising from different agent speeds. \n\nOnce it became clear that this was key it was possible to look at sampling the data such that we balanced these cases. \n\nA scene containing lots of agents was indicative of traffic. Gridlocked traffic obviously does not move much and leads to a lot of duplication in inputs. The models implemented a sampling scheme whereby scenes with a large number of agents were undersampled. The sampling proportion was: min(1, 7/agent_count). Thus, scenes containing 14 agents had 50% of those agents selected for each training iteration, etc.\n\n\n**Ensembling:**\n\nThis was initially a tricky one: averaging models based on confidence values didn’t work. However, when the importance of acceleration was taken into account, the route to ensembling made sense: order the model modes by distance covered, then average the results. When this was implemented all models could be ensembled very quickly, with positive results. We also looked at incorporating curvature here, but it didn’t make any difference: distance was the key.\n\nThe final ensemble optimized weights on the validation set. The weighting scheme incorporated distance and confidence values. \n\n\n**Code:**\n\nCode for contribution to our team solution can be found [here](https://github.com/ciararogerson/Kaggle_Lyft)\n\n\n**Ideas that didn’t work: many! Here are a few...**\n\n_Traffic lights:_\n\nI couldn’t get additional traffic light information to add anything: many different angles were tried! The most promising was probably traffic light persistence: we included an additional channel in the model where instead of traffic light lane lines, we drew lines containing the number of frames since the traffic light had turned to its current colour (up to a maximum of 100). Ultimately this didn’t add anything.\n\n_Day/hour:_\n\nAdding channels for day/hour values proved better than concatenating them directly before the dense layers of the model, but it still didn’t help much.\n\n_Interpolation:_\n\nHaving the penultimate model layer output 25 sets of (x, y) points, followed by an interpolation between these points to make the final 50 point trajectory did not work.\n\n_Weighted loss function:_\n\nGiven the importance of capturing acceleration, we tried a version of the loss function that weighted  the last 10 points of the trajectory equal to the first 40. This didn’t help.",
    "1091866": "WHat was your hardware",
    "1091876": "3x1080ti. 128GB RAM, 10 cores. Could have done with more!!",
    "1091884": "Great writeup, look forward to other members responses. I also chopped train_full, initially at frame 100 then slowly expanded to include additional frames, and saw consistent improvement. After chopping did you continue to use rasterization on the fly while training or did you store rasters on disk? How long did training take and what hardware did you use?\n\nCongratulations :)",
    "1091886": "[Here they mention about their hardware](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588#1091876)",
    "1091887": "Yes, I kept the rasterization throughout as I was so often playing around with different configs that it didn't make sense to freeze it. Also, I wasn't sure whether I'd have the disk space...",
    "1091903": "Thanks for the writeup! And congratulations!\nSome smart ideas you had like using the daytime in the model and the summing trick\n\nI see that you experienced large performance gains while using a subsampled train_full over a random sampled train_full. Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?\nIt makes sense, that subsampling most diverse samples can get the validation score down quicker, but I would be very interested if it actually makes a difference how you shuffle after a full epoch.\nWhy did you use the same subsample again for a second epoch? Why not e.g. epoch0_frame +1?",
    "1091912": "Thanks! Same to you!\n\nNo, I never trained train_full for a complete epoch, I think I'd be here until Christmas if I did! @rytisva88  had been training on a random shuffle of train_full before switching to the chopped version, so he will have a better gauge on what the impact was. @sheriytm also went from using train.zarr to a chop of train_full.zarr so will have an idea of how much that was worth.\n\nWe ended up reusing chopped datasets (i.e. showing the same sample multiple times). It would have been better if we didn't, but we were running out of time and were resource constrained. \n\nIf I were starting over I would have done it differently: rather than creating complete chops I would just store off the indices in train_full.zarr that those chops refer to, thereby allowing us to create sequences of diverse samples without having to show samples multiple times or run out of resources. @rytisva88 managed resources better and he was able to use 20 chopped datasets (but could have gone higher). I was stuck using 16.",
    "1091967": "Great result with 3 1080ti !",
    "1092011": "Wow. Great job and Thanks for sharing",
    "1092047": "Congratulations! Could I ask that how much improvement you got using the ensembling scheme you has mentioned?",
    "1092049": "Very thorough and impressive.",
    "1092203": "Did you ever try something like applying k-means to ensemble your trajectories? In theory similar to your distance covered ensembling, but potentially more automatic and you can also apply sample weights directly in sklearn based on the confidence output.",
    "1092206": "I did try kmeans early on, but it wasn't working for me, perhaps because of how I implemented it: if there were three models, say, I was clustering 9 modes into three clusters. Then taking the max conf mode for each cluster or averaging within each cluster (I tried both). But obviously had something wrong somewhere...",
    "1092211": "The top score was an ensemble of the two best models (slightly lower score was an ensemble of four). For that one, 11.15 + 10.52 -> 10.32",
    "1095464": "Thanks for sharing @fergusoci ! Very Insightful 👍",
    "1095737": "Exhausted from the competition, my day job and Thanksgiving holiday stuff this had to wait and I apologize for the delayed post.\n\nMy first thanks is to the competition organizers and the Kaggle team for this really interesting contest. Congratulations @philippsinger, @christofhenkel, @ilu000, @nvnnghia a.k.a.  team NIPD for a well deserved win.\n\n@fergusoci have pretty much presented the team efforts so well I am just going to add a few points here instead of a separate post. First I want to thank my team mates @fergusoci and @rytisva88 for making this competition experience a pleasure.\n\nSince much far useful solutions from the top teams are already posted, I will give a brief on some of my training procedures:-\n\nI have limited GPU resource hence could not use train_full.zarr before teaming up. The GCP credit was unusable as GPU is not available where I am on my travels at this time. My best model before merger was a Resnet34 (224+5) trained on train.zarr to a reasonable LB score of 21.403 before it started to overfit. Training this model on the chopped data made improvements. \n\nOnce we merged looking at all the teams models and various efforts after brain storming, I turned my attention to training a different type of model to add diversity. After a few (i.e. PointNet, EfficientNet, etc) trials, settled on training a Resnext50_32x4d (224+5) with train.zarr and chopped train_full.zarr to help add diversity to our ensemble. Basically, after training on 2.5M samples of train.zarr, I alternated by training between train.zarr and chopped 100 train_full.zarr. Each cycle trains on 256000 samples of the data. That made the scores go down significantly but got slower after reaching the sub 20 LB scores. At this point, I continued training on the chopped data only, which reached sub 14 scores after going through the data 2.5 times at competition close. I have continued training this model to see if it can get to the scores of our best models and make a late submission.\n\nAnd thanks everyone who contributed in many ways via kernels and discussions.",
    "1095739": "> Have you ever trained the random shuffled train_full for a complete epoch? Was the performance still lower?\n\nNo @ilu000, in my case that was not an option due GPU challenged. See my post in the comment section here for how I went about training my models.",
    "1103461": "Thanks for sharing detailed write-up @fergusoci , congrats! Ensemble method is interesting."
  },
  "source": "meta"
}