{
  "id": 199531,
  "title": "What ensemble method did you use?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199531",
  "author_name": "",
  "post_date": "2020-11-26T04:52:37.471141700Z",
  "votes": 12,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I was having a hard to ensembling different models. Here are some naive methods I tried:</p>\n<p>1) Match the most confident projection for each id (every row), and weight them with confidences<br>\n2) Match the most confident projection for each column, and weight them with column-wise average confidences</p>\n<p>And I found non of the above is as good as simply averaging the trajectories in natural order directly (presumably because they are based on the same pertained model). So I thought of </p>\n<p>3) Utilize the <a href=\"https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min\" target=\"_blank\">chopped validation method</a> to iteratively test the best weights for each model. One ad hoc (quick and dirty) way to do that is inflating the inverse of test score, power(1/score, x), as weights. And I got the following. Basically X in power(1/score, x) and y for validation score. In this case the optimal was around 12 and reflects a weight of 0.7 for my best single model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Fe70e2756eaa1efa4727c036962efb7f1%2Fcurve.png?generation=1606365333528192&amp;alt=media\" alt=\"\"></p>\n<p>Because my case is simply ensembling the same model with different training data and parameters, so the combination of them as I understand it is simply smoothing out some noise. But it helped a single model from 18,6 to 17,8.</p>\n<p>Please let me know if this is overfitting or feel free the enlighten me on other ensemble methods.</p>",
  "messages": [
    {
      "id": "1091524",
      "postDate": "11/26/2020 04:52:37",
      "content": "<p>I was having a hard to ensembling different models. Here are some naive methods I tried:</p>\n<p>1) Match the most confident projection for each id (every row), and weight them with confidences<br>\n2) Match the most confident projection for each column, and weight them with column-wise average confidences</p>\n<p>And I found non of the above is as good as simply averaging the trajectories in natural order directly (presumably because they are based on the same pertained model). So I thought of </p>\n<p>3) Utilize the <a href=\"https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min\" target=\"_blank\">chopped validation method</a> to iteratively test the best weights for each model. One ad hoc (quick and dirty) way to do that is inflating the inverse of test score, power(1/score, x), as weights. And I got the following. Basically X in power(1/score, x) and y for validation score. In this case the optimal was around 12 and reflects a weight of 0.7 for my best single model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Fe70e2756eaa1efa4727c036962efb7f1%2Fcurve.png?generation=1606365333528192&amp;alt=media\" alt=\"\"></p>\n<p>Because my case is simply ensembling the same model with different training data and parameters, so the combination of them as I understand it is simply smoothing out some noise. But it helped a single model from 18,6 to 17,8.</p>\n<p>Please let me know if this is overfitting or feel free the enlighten me on other ensemble methods.</p>",
      "rawMarkdown": "I was having a hard to ensembling different models. Here are some naive methods I tried:\n\n1) Match the most confident projection for each id (every row), and weight them with confidences\n2) Match the most confident projection for each column, and weight them with column-wise average confidences\n\nAnd I found non of the above is as good as simply averaging the trajectories in natural order directly (presumably because they are based on the same pertained model). So I thought of \n\n3) Utilize the [chopped validation method](https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min) to iteratively test the best weights for each model. One ad hoc (quick and dirty) way to do that is inflating the inverse of test score, power(1/score, x), as weights. And I got the following. Basically X in power(1/score, x) and y for validation score. In this case the optimal was around 12 and reflects a weight of 0.7 for my best single model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Fe70e2756eaa1efa4727c036962efb7f1%2Fcurve.png?generation=1606365333528192&alt=media)\n\nBecause my case is simply ensembling the same model with different training data and parameters, so the combination of them as I understand it is simply smoothing out some noise. But it helped a single model from 18,6 to 17,8.\n\nPlease let me know if this is overfitting or feel free the enlighten me on other ensemble methods.",
      "votes": null
    },
    {
      "id": "1091530",
      "postDate": "11/26/2020 05:03:19",
      "content": "<p>Here are some results.<br>\nFig 1 for my best single model with dropout.<br>\nFig 2 for simply averaging trajectories for three similar models<br>\nFig 3 for fine tuning the weights with the naive method described<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Faa13d0c2ba0e30c5db8b010e4936e00a%2Fscore0.png?generation=1606366859617095&amp;alt=media\" alt=\"Fig 1\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2F3ab15c2815866e410ad90f2fbdc1f5ae%2Fscore1.png?generation=1606366873228927&amp;alt=media\" alt=\"Fig 2\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Feed3197e2c993ebe0f7c66b7f9890266%2Fscore2.png?generation=1606366960257055&amp;alt=media\" alt=\"Fig 3\"></p>",
      "rawMarkdown": "Here are some results.\nFig 1 for my best single model with dropout.\nFig 2 for simply averaging trajectories for three similar models\nFig 3 for fine tuning the weights with the naive method described\n![Fig 1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Faa13d0c2ba0e30c5db8b010e4936e00a%2Fscore0.png?generation=1606366859617095&alt=media)\n![Fig 2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2F3ab15c2815866e410ad90f2fbdc1f5ae%2Fscore1.png?generation=1606366873228927&alt=media)\n![Fig 3](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Feed3197e2c993ebe0f7c66b7f9890266%2Fscore2.png?generation=1606366960257055&alt=media)",
      "votes": null
    },
    {
      "id": "1091534",
      "postDate": "11/26/2020 05:09:05",
      "content": "<p>The most successful methods I was able to come up with were applying k-means to each set of predictions for each sample and then returning the cluster centers. And then also had similar results from the method <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> mentioned where I combined the embedding layers of multiple models and then trained a new prediction head on top of the frozen combination of the multiple models. </p>",
      "rawMarkdown": "The most successful methods I was able to come up with were applying k-means to each set of predictions for each sample and then returning the cluster centers. And then also had similar results from the method @hengck23 mentioned where I combined the embedding layers of multiple models and then trained a new prediction head on top of the frozen combination of the multiple models.",
      "votes": null
    },
    {
      "id": "1091548",
      "postDate": "11/26/2020 05:23:33",
      "content": "<p>\"And then also had similar results from the method <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> mentioned where I combined the embedding layers \"</p>\n<p>the gain of this method is about 1 to 1.5 lb. however, there is a catch. because the layers are freezed, performance is limited by the base models.</p>\n<p>i have another idea but I didn't try:<br>\n1) say you have N models, then you would have M=N*3 predictions (alternatively, you can have one models with M modes. M is larger than 3)<br>\n2) we can call these M prediction proposals (like those used on detection).<br>\n3) the problem then become choosing 3 out of M proposals, i.e. classification problems with bum of class = combination(M,3)<br>\n4) because it is classification, it is easier to the ensemble. just add up and average class probabability.</p>",
      "rawMarkdown": "\"And then also had similar results from the method @hengck23 mentioned where I combined the embedding layers \"\n\nthe gain of this method is about 1 to 1.5 lb. however, there is a catch. because the layers are freezed, performance is limited by the base models.\n\ni have another idea but I didn't try:\n1) say you have N models, then you would have M=N*3 predictions (alternatively, you can have one models with M modes. M is larger than 3)\n2) we can call these M prediction proposals (like those used on detection).\n3) the problem then become choosing 3 out of M proposals, i.e. classification problems with bum of class = combination(M,3)\n4) because it is classification, it is easier to the ensemble. just add up and average class probabability.",
      "votes": null
    },
    {
      "id": "1091585",
      "postDate": "11/26/2020 05:56:44",
      "content": "<p>I tried applying classification methods. Training a 10 mode model with extremely low loss and then using an entirely separate classifier model with various combinations of inputs. just the regular input from the data generator, inputs plus trajectories, inputs plus trajectories, inputs plus trajectories plus embeddings from the trajectory proposal model. My models always really struggled to classify correctly and yield anything comparable to directly training a 3 mode model. </p>",
      "rawMarkdown": "I tried applying classification methods. Training a 10 mode model with extremely low loss and then using an entirely separate classifier model with various combinations of inputs. just the regular input from the data generator, inputs plus trajectories, inputs plus trajectories, inputs plus trajectories plus embeddings from the trajectory proposal model. My models always really struggled to classify correctly and yield anything comparable to directly training a 3 mode model.",
      "votes": null
    },
    {
      "id": "1091632",
      "postDate": "11/26/2020 06:37:56",
      "content": "<p>Stacking worked fine.</p>\n<p>Overview:</p>\n<p>model1: val_chopped 13.2<br>\nmodel2: val_chopped 14.3</p>\n<p>First, concatenate flattened prediction of these models, <code>dim=(50 future frame * 2 xy * 3 modes + 3 conf.) * 2 models=606</code>. Next, linear(606, 4096) -&gt; relu. Concatenate this 4096 vector, model1 mid layer(4096), and model2 mid layer(4096), linear(4096*3, 4096) -&gt; relu -&gt; linear(4096, 303).</p>\n<p>This stacking model got 12.8 (val_chopped).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F96dfeb41c3a9ca7c7606989c33a4ab66%2FStacking.png?generation=1606372432970337&amp;alt=media\" alt=\"\"></p>\n<p>I came up with this stacking method 2 days before competition end, so I couldn't integrate the best model(val_chopped 12.3) to stacking…<br>\nWe will post the other parts of our main solution later.</p>",
      "rawMarkdown": "Stacking worked fine.\n\nOverview:\n\nmodel1: val_chopped 13.2\nmodel2: val_chopped 14.3\n\nFirst, concatenate flattened prediction of these models, `dim=(50 future frame * 2 xy * 3 modes + 3 conf.) * 2 models=606`. Next, linear(606, 4096) -> relu. Concatenate this 4096 vector, model1 mid layer(4096), and model2 mid layer(4096), linear(4096*3, 4096) -> relu -> linear(4096, 303).\n\nThis stacking model got 12.8 (val_chopped).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F96dfeb41c3a9ca7c7606989c33a4ab66%2FStacking.png?generation=1606372432970337&alt=media)\n\nI came up with this stacking method 2 days before competition end, so I couldn't integrate the best model(val_chopped 12.3) to stacking...\nWe will post the other parts of our main solution later.",
      "votes": null
    },
    {
      "id": "1091760",
      "postDate": "11/26/2020 08:59:13",
      "content": "<p>Yeah, stacking is the only thing that worked for us.</p>",
      "rawMarkdown": "Yeah, stacking is the only thing that worked for us.",
      "votes": null
    },
    {
      "id": "1091784",
      "postDate": "11/26/2020 09:27:33",
      "content": "<p>In the end I opted for training a single model with more data and for longer, and for an ensemble just used the average of checkpoints (also tried optimizing ensemble weights, but end performance was effectively the same). This provided a boost of around ~0.9 over the best single checkpoint.</p>",
      "rawMarkdown": "In the end I opted for training a single model with more data and for longer, and for an ensemble just used the average of checkpoints (also tried optimizing ensemble weights, but end performance was effectively the same). This provided a boost of around ~0.9 over the best single checkpoint.",
      "votes": null
    },
    {
      "id": "1091874",
      "postDate": "11/26/2020 10:54:11",
      "content": "<p>For diverse models, the key was distance ordering. (Well, stacking obviously works too, but you need to model this).</p>\n<p>The model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. </p>\n<p>I also looked at whether incorporating curvature was of any use, but it didn't seem to be…</p>",
      "rawMarkdown": "For diverse models, the key was distance ordering. (Well, stacking obviously works too, but you need to model this).\n\nThe model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. \n\nI also looked at whether incorporating curvature was of any use, but it didn't seem to be...",
      "votes": null
    },
    {
      "id": "1091896",
      "postDate": "11/26/2020 11:07:48",
      "content": "<p>I really like this approach, what kind of improvements did you see in score from single model to ensemble?</p>",
      "rawMarkdown": "I really like this approach, what kind of improvements did you see in score from single model to ensemble?",
      "votes": null
    },
    {
      "id": "1091902",
      "postDate": "11/26/2020 11:12:57",
      "content": "<p>It depended on the model set - early on when then models weren't great there was pretty big improvement, up to 1pt or so over the best model. As the models got better the magnitude of improvement was tighter. It depended a lot on how close the scores were. For e.g. if you had a model at 11.3 and a second at 12.3 it would only get to 11 or so.</p>",
      "rawMarkdown": "It depended on the model set - early on when then models weren't great there was pretty big improvement, up to 1pt or so over the best model. As the models got better the magnitude of improvement was tighter. It depended a lot on how close the scores were. For e.g. if you had a model at 11.3 and a second at 12.3 it would only get to 11 or so.",
      "votes": null
    },
    {
      "id": "1091969",
      "postDate": "11/26/2020 12:30:52",
      "content": "<p>Thank you for your reply. Your method enlightens me a lot. Just wondering how did you deal with the deviation in angles of different modes? That is, maybe the model is predicting a minor but different case. In that sense, would it make sense to first filter out the modes that deviate too much (beyond a level), and then average then?</p>",
      "rawMarkdown": "Thank you for your reply. Your method enlightens me a lot. Just wondering how did you deal with the deviation in angles of different modes? That is, maybe the model is predicting a minor but different case. In that sense, would it make sense to first filter out the modes that deviate too much (beyond a level), and then average then?",
      "votes": null
    },
    {
      "id": "1091972",
      "postDate": "11/26/2020 12:36:27",
      "content": "<p>I originally thought that the curvature (is that what you mean by angle?) would make a difference, but when tested it didn't. i.e. it worked fine to just straight average across models once you had ordered the modes by distance from first to last point. When visualising different samples it also didn't seem to throw anything way off. The averages seemed reasonable, even for cases where the trajectory is not straight.</p>",
      "rawMarkdown": "I originally thought that the curvature (is that what you mean by angle?) would make a difference, but when tested it didn't. i.e. it worked fine to just straight average across models once you had ordered the modes by distance from first to last point. When visualising different samples it also didn't seem to throw anything way off. The averages seemed reasonable, even for cases where the trajectory is not straight.",
      "votes": null
    },
    {
      "id": "1092051",
      "postDate": "11/26/2020 13:57:05",
      "content": "<p>Thank you for your reply. Yes I meant curvature.</p>",
      "rawMarkdown": "Thank you for your reply. Yes I meant curvature.",
      "votes": null
    },
    {
      "id": "1092057",
      "postDate": "11/26/2020 14:01:55",
      "content": "<blockquote>\n  <p>The model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. </p>\n</blockquote>\n<p>What an astute observation! I had three weak models and did a search w/o replacement to find the mode, per sample, with the nearest location and then weighted blended. This N*M^2 operation could have been simplified considerably by simply taking the dist mags first and then sorting on it. Well done!</p>",
      "rawMarkdown": "> The model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. \n\nWhat an astute observation! I had three weak models and did a search w/o replacement to find the mode, per sample, with the nearest location and then weighted blended. This N*M^2 operation could have been simplified considerably by simply taking the dist mags first and then sorting on it. Well done!",
      "votes": null
    },
    {
      "id": "1092758",
      "postDate": "11/27/2020 06:28:19",
      "content": "<p>We used Gaussian Mixture Model, as written in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657\" target=\"_blank\">here</a>. <br>\nMaybe the behavior is almost same with k-means clustering. It was very effective.</p>",
      "rawMarkdown": "We used Gaussian Mixture Model, as written in [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657). \nMaybe the behavior is almost same with k-means clustering. It was very effective.",
      "votes": null
    },
    {
      "id": "1092776",
      "postDate": "11/27/2020 06:45:17",
      "content": "<p>I tried both GMM and k-means and found k-means was significantly faster even with a much higher value of n_inits and yielded a marginally better score. At least in my implementation of the GMM where I was just using the cluster centers still and discarding the covariance info. I wasn't sure exactly how to use that in some useful way. </p>",
      "rawMarkdown": "I tried both GMM and k-means and found k-means was significantly faster even with a much higher value of n_inits and yielded a marginally better score. At least in my implementation of the GMM where I was just using the cluster centers still and discarding the covariance info. I wasn't sure exactly how to use that in some useful way.",
      "votes": null
    },
    {
      "id": "1092779",
      "postDate": "11/27/2020 06:47:51",
      "content": "<p>Thank you for your reply. You wrote a great discussion thread and congrats! <br>\nIt's interesting to see different ensemble methods and how important they are for this competition.</p>",
      "rawMarkdown": "Thank you for your reply. You wrote a great discussion thread and congrats! \nIt's interesting to see different ensemble methods and how important they are for this competition.",
      "votes": null
    },
    {
      "id": "1092793",
      "postDate": "11/27/2020 07:10:00",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> <br>\nThanks for comment.<br>\nDid you set <code>covariance_type=”spherical”</code>? I guess <code>covariance_type=\"full\"</code> takes time to fit covariance parameters.</p>",
      "rawMarkdown": "ryches \nThanks for comment.\nDid you set `covariance_type=”spherical”`? I guess `covariance_type=\"full\"` takes time to fit covariance parameters.",
      "votes": null
    },
    {
      "id": "1092795",
      "postDate": "11/27/2020 07:11:12",
      "content": "<p><a href=\"https://www.kaggle.com/ryanchun\" target=\"_blank\">@ryanchun</a> Thank you for comment. Yeah the ensemble method was not trivial in this competition and I also surprised that every team uses different approaches!</p>",
      "rawMarkdown": "ryanchun Thank you for comment. Yeah the ensemble method was not trivial in this competition and I also surprised that every team uses different approaches!",
      "votes": null
    },
    {
      "id": "1092866",
      "postDate": "11/27/2020 08:24:56",
      "content": "<p>No, I did not set it to spherical? Does that affect the speed? I did not dig too deeply into the GMM details. </p>",
      "rawMarkdown": "No, I did not set it to spherical? Does that affect the speed? I did not dig too deeply into the GMM details.",
      "votes": null
    },
    {
      "id": "1094892",
      "postDate": "11/29/2020 04:13:10",
      "content": "<p>Yeah it does affect to speed a lot.<br>\nDefault setting fits covariance matrix, which is the heavy part.</p>\n<p>But this competition's metric is used with fixed <code>sigma=1</code>, so we can/should skip fitting full covariance matrix.</p>",
      "rawMarkdown": "Yeah it does affect to speed a lot.\nDefault setting fits covariance matrix, which is the heavy part.\n\nBut this competition's metric is used with fixed `sigma=1`, so we can/should skip fitting full covariance matrix.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091530,
      "author_name": "ryanchun",
      "author_url": "",
      "post_date": "11/26/2020 05:03:19",
      "content": "<p>Here are some results.<br>\nFig 1 for my best single model with dropout.<br>\nFig 2 for simply averaging trajectories for three similar models<br>\nFig 3 for fine tuning the weights with the naive method described<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Faa13d0c2ba0e30c5db8b010e4936e00a%2Fscore0.png?generation=1606366859617095&amp;alt=media\" alt=\"Fig 1\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2F3ab15c2815866e410ad90f2fbdc1f5ae%2Fscore1.png?generation=1606366873228927&amp;alt=media\" alt=\"Fig 2\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Feed3197e2c993ebe0f7c66b7f9890266%2Fscore2.png?generation=1606366960257055&amp;alt=media\" alt=\"Fig 3\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091534,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/26/2020 05:09:05",
      "content": "<p>The most successful methods I was able to come up with were applying k-means to each set of predictions for each sample and then returning the cluster centers. And then also had similar results from the method <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> mentioned where I combined the embedding layers of multiple models and then trained a new prediction head on top of the frozen combination of the multiple models. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1091548,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/26/2020 05:23:33",
          "content": "<p>\"And then also had similar results from the method <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> mentioned where I combined the embedding layers \"</p>\n<p>the gain of this method is about 1 to 1.5 lb. however, there is a catch. because the layers are freezed, performance is limited by the base models.</p>\n<p>i have another idea but I didn't try:<br>\n1) say you have N models, then you would have M=N*3 predictions (alternatively, you can have one models with M modes. M is larger than 3)<br>\n2) we can call these M prediction proposals (like those used on detection).<br>\n3) the problem then become choosing 3 out of M proposals, i.e. classification problems with bum of class = combination(M,3)<br>\n4) because it is classification, it is easier to the ensemble. just add up and average class probabability.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091585,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/26/2020 05:56:44",
          "content": "<p>I tried applying classification methods. Training a 10 mode model with extremely low loss and then using an entirely separate classifier model with various combinations of inputs. just the regular input from the data generator, inputs plus trajectories, inputs plus trajectories, inputs plus trajectories plus embeddings from the trajectory proposal model. My models always really struggled to classify correctly and yield anything comparable to directly training a 3 mode model. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091632,
      "author_name": "inoueu1",
      "author_url": "",
      "post_date": "11/26/2020 06:37:56",
      "content": "<p>Stacking worked fine.</p>\n<p>Overview:</p>\n<p>model1: val_chopped 13.2<br>\nmodel2: val_chopped 14.3</p>\n<p>First, concatenate flattened prediction of these models, <code>dim=(50 future frame * 2 xy * 3 modes + 3 conf.) * 2 models=606</code>. Next, linear(606, 4096) -&gt; relu. Concatenate this 4096 vector, model1 mid layer(4096), and model2 mid layer(4096), linear(4096*3, 4096) -&gt; relu -&gt; linear(4096, 303).</p>\n<p>This stacking model got 12.8 (val_chopped).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F96dfeb41c3a9ca7c7606989c33a4ab66%2FStacking.png?generation=1606372432970337&amp;alt=media\" alt=\"\"></p>\n<p>I came up with this stacking method 2 days before competition end, so I couldn't integrate the best model(val_chopped 12.3) to stacking…<br>\nWe will post the other parts of our main solution later.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091760,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/26/2020 08:59:13",
          "content": "<p>Yeah, stacking is the only thing that worked for us.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091784,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "11/26/2020 09:27:33",
      "content": "<p>In the end I opted for training a single model with more data and for longer, and for an ensemble just used the average of checkpoints (also tried optimizing ensemble weights, but end performance was effectively the same). This provided a boost of around ~0.9 over the best single checkpoint.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091874,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "11/26/2020 10:54:11",
      "content": "<p>For diverse models, the key was distance ordering. (Well, stacking obviously works too, but you need to model this).</p>\n<p>The model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. </p>\n<p>I also looked at whether incorporating curvature was of any use, but it didn't seem to be…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091896,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "11/26/2020 11:07:48",
          "content": "<p>I really like this approach, what kind of improvements did you see in score from single model to ensemble?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091902,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/26/2020 11:12:57",
          "content": "<p>It depended on the model set - early on when then models weren't great there was pretty big improvement, up to 1pt or so over the best model. As the models got better the magnitude of improvement was tighter. It depended a lot on how close the scores were. For e.g. if you had a model at 11.3 and a second at 12.3 it would only get to 11 or so.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091969,
          "author_name": "ryanchun",
          "author_url": "",
          "post_date": "11/26/2020 12:30:52",
          "content": "<p>Thank you for your reply. Your method enlightens me a lot. Just wondering how did you deal with the deviation in angles of different modes? That is, maybe the model is predicting a minor but different case. In that sense, would it make sense to first filter out the modes that deviate too much (beyond a level), and then average then?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091972,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/26/2020 12:36:27",
          "content": "<p>I originally thought that the curvature (is that what you mean by angle?) would make a difference, but when tested it didn't. i.e. it worked fine to just straight average across models once you had ordered the modes by distance from first to last point. When visualising different samples it also didn't seem to throw anything way off. The averages seemed reasonable, even for cases where the trajectory is not straight.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092051,
          "author_name": "ryanchun",
          "author_url": "",
          "post_date": "11/26/2020 13:57:05",
          "content": "<p>Thank you for your reply. Yes I meant curvature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092057,
          "author_name": "authman",
          "author_url": "",
          "post_date": "11/26/2020 14:01:55",
          "content": "<blockquote>\n  <p>The model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. </p>\n</blockquote>\n<p>What an astute observation! I had three weak models and did a search w/o replacement to find the mode, per sample, with the nearest location and then weighted blended. This N*M^2 operation could have been simplified considerably by simply taking the dist mags first and then sorting on it. Well done!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092758,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "11/27/2020 06:28:19",
      "content": "<p>We used Gaussian Mixture Model, as written in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657\" target=\"_blank\">here</a>. <br>\nMaybe the behavior is almost same with k-means clustering. It was very effective.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092776,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/27/2020 06:45:17",
          "content": "<p>I tried both GMM and k-means and found k-means was significantly faster even with a much higher value of n_inits and yielded a marginally better score. At least in my implementation of the GMM where I was just using the cluster centers still and discarding the covariance info. I wasn't sure exactly how to use that in some useful way. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092779,
          "author_name": "ryanchun",
          "author_url": "",
          "post_date": "11/27/2020 06:47:51",
          "content": "<p>Thank you for your reply. You wrote a great discussion thread and congrats! <br>\nIt's interesting to see different ensemble methods and how important they are for this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092793,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "11/27/2020 07:10:00",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> <br>\nThanks for comment.<br>\nDid you set <code>covariance_type=”spherical”</code>? I guess <code>covariance_type=\"full\"</code> takes time to fit covariance parameters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092795,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "11/27/2020 07:11:12",
          "content": "<p><a href=\"https://www.kaggle.com/ryanchun\" target=\"_blank\">@ryanchun</a> Thank you for comment. Yeah the ensemble method was not trivial in this competition and I also surprised that every team uses different approaches!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092866,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/27/2020 08:24:56",
          "content": "<p>No, I did not set it to spherical? Does that affect the speed? I did not dig too deeply into the GMM details. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094892,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "11/29/2020 04:13:10",
          "content": "<p>Yeah it does affect to speed a lot.<br>\nDefault setting fits covariance matrix, which is the heavy part.</p>\n<p>But this competition's metric is used with fixed <code>sigma=1</code>, so we can/should skip fitting full covariance matrix.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1091524": "I was having a hard to ensembling different models. Here are some naive methods I tried:\n\n1) Match the most confident projection for each id (every row), and weight them with confidences\n2) Match the most confident projection for each column, and weight them with column-wise average confidences\n\nAnd I found non of the above is as good as simply averaging the trajectories in natural order directly (presumably because they are based on the same pertained model). So I thought of \n\n3) Utilize the [chopped validation method](https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min) to iteratively test the best weights for each model. One ad hoc (quick and dirty) way to do that is inflating the inverse of test score, power(1/score, x), as weights. And I got the following. Basically X in power(1/score, x) and y for validation score. In this case the optimal was around 12 and reflects a weight of 0.7 for my best single model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Fe70e2756eaa1efa4727c036962efb7f1%2Fcurve.png?generation=1606365333528192&alt=media)\n\nBecause my case is simply ensembling the same model with different training data and parameters, so the combination of them as I understand it is simply smoothing out some noise. But it helped a single model from 18,6 to 17,8.\n\nPlease let me know if this is overfitting or feel free the enlighten me on other ensemble methods.",
    "1091530": "Here are some results.\nFig 1 for my best single model with dropout.\nFig 2 for simply averaging trajectories for three similar models\nFig 3 for fine tuning the weights with the naive method described\n![Fig 1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Faa13d0c2ba0e30c5db8b010e4936e00a%2Fscore0.png?generation=1606366859617095&alt=media)\n![Fig 2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2F3ab15c2815866e410ad90f2fbdc1f5ae%2Fscore1.png?generation=1606366873228927&alt=media)\n![Fig 3](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2Feed3197e2c993ebe0f7c66b7f9890266%2Fscore2.png?generation=1606366960257055&alt=media)",
    "1091534": "The most successful methods I was able to come up with were applying k-means to each set of predictions for each sample and then returning the cluster centers. And then also had similar results from the method @hengck23 mentioned where I combined the embedding layers of multiple models and then trained a new prediction head on top of the frozen combination of the multiple models.",
    "1091548": "\"And then also had similar results from the method @hengck23 mentioned where I combined the embedding layers \"\n\nthe gain of this method is about 1 to 1.5 lb. however, there is a catch. because the layers are freezed, performance is limited by the base models.\n\ni have another idea but I didn't try:\n1) say you have N models, then you would have M=N*3 predictions (alternatively, you can have one models with M modes. M is larger than 3)\n2) we can call these M prediction proposals (like those used on detection).\n3) the problem then become choosing 3 out of M proposals, i.e. classification problems with bum of class = combination(M,3)\n4) because it is classification, it is easier to the ensemble. just add up and average class probabability.",
    "1091585": "I tried applying classification methods. Training a 10 mode model with extremely low loss and then using an entirely separate classifier model with various combinations of inputs. just the regular input from the data generator, inputs plus trajectories, inputs plus trajectories, inputs plus trajectories plus embeddings from the trajectory proposal model. My models always really struggled to classify correctly and yield anything comparable to directly training a 3 mode model.",
    "1091632": "Stacking worked fine.\n\nOverview:\n\nmodel1: val_chopped 13.2\nmodel2: val_chopped 14.3\n\nFirst, concatenate flattened prediction of these models, `dim=(50 future frame * 2 xy * 3 modes + 3 conf.) * 2 models=606`. Next, linear(606, 4096) -> relu. Concatenate this 4096 vector, model1 mid layer(4096), and model2 mid layer(4096), linear(4096*3, 4096) -> relu -> linear(4096, 303).\n\nThis stacking model got 12.8 (val_chopped).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1057275%2F96dfeb41c3a9ca7c7606989c33a4ab66%2FStacking.png?generation=1606372432970337&alt=media)\n\nI came up with this stacking method 2 days before competition end, so I couldn't integrate the best model(val_chopped 12.3) to stacking...\nWe will post the other parts of our main solution later.",
    "1091760": "Yeah, stacking is the only thing that worked for us.",
    "1091784": "In the end I opted for training a single model with more data and for longer, and for an ensemble just used the average of checkpoints (also tried optimizing ensemble weights, but end performance was effectively the same). This provided a boost of around ~0.9 over the best single checkpoint.",
    "1091874": "For diverse models, the key was distance ordering. (Well, stacking obviously works too, but you need to model this).\n\nThe model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. \n\nI also looked at whether incorporating curvature was of any use, but it didn't seem to be...",
    "1091896": "I really like this approach, what kind of improvements did you see in score from single model to ensemble?",
    "1091902": "It depended on the model set - early on when then models weren't great there was pretty big improvement, up to 1pt or so over the best model. As the models got better the magnitude of improvement was tighter. It depended a lot on how close the scores were. For e.g. if you had a model at 11.3 and a second at 12.3 it would only get to 11 or so.",
    "1091969": "Thank you for your reply. Your method enlightens me a lot. Just wondering how did you deal with the deviation in angles of different modes? That is, maybe the model is predicting a minor but different case. In that sense, would it make sense to first filter out the modes that deviate too much (beyond a level), and then average then?",
    "1091972": "I originally thought that the curvature (is that what you mean by angle?) would make a difference, but when tested it didn't. i.e. it worked fine to just straight average across models once you had ordered the modes by distance from first to last point. When visualising different samples it also didn't seem to throw anything way off. The averages seemed reasonable, even for cases where the trajectory is not straight.",
    "1092051": "Thank you for your reply. Yes I meant curvature.",
    "1092057": "> The model modes are typically a bet on acceleration: how far along a trajectory will an agent get. If you order the modes by distance covered and then average them, it works effectively. \n\nWhat an astute observation! I had three weak models and did a search w/o replacement to find the mode, per sample, with the nearest location and then weighted blended. This N*M^2 operation could have been simplified considerably by simply taking the dist mags first and then sorting on it. Well done!",
    "1092758": "We used Gaussian Mixture Model, as written in [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657). \nMaybe the behavior is almost same with k-means clustering. It was very effective.",
    "1092776": "I tried both GMM and k-means and found k-means was significantly faster even with a much higher value of n_inits and yielded a marginally better score. At least in my implementation of the GMM where I was just using the cluster centers still and discarding the covariance info. I wasn't sure exactly how to use that in some useful way.",
    "1092779": "Thank you for your reply. You wrote a great discussion thread and congrats! \nIt's interesting to see different ensemble methods and how important they are for this competition.",
    "1092793": "ryches \nThanks for comment.\nDid you set `covariance_type=”spherical”`? I guess `covariance_type=\"full\"` takes time to fit covariance parameters.",
    "1092795": "ryanchun Thank you for comment. Yeah the ensemble method was not trivial in this competition and I also surprised that every team uses different approaches!",
    "1092866": "No, I did not set it to spherical? Does that affect the speed? I did not dig too deeply into the GMM details.",
    "1094892": "Yeah it does affect to speed a lot.\nDefault setting fits covariance matrix, which is the heavy part.\n\nBut this competition's metric is used with fixed `sigma=1`, so we can/should skip fitting full covariance matrix."
  },
  "source": "meta"
}