{
  "id": 193908,
  "title": "multi-mode trick",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/193908",
  "author_name": "hengck23",
  "post_date": "2020-10-29T14:28:19.672000",
  "votes": 39,
  "comment_count": 75,
  "views": 0,
  "content": "<p>first, why is there multi-mode?</p>\n<ul>\n<li>e.g. car motion is different if traffic light condition is different ( green lights or red lights )</li>\n<li>e.g. intention of the driver (maybe he wants to go destination A instead of B)</li>\n</ul>\n<p>so is three mode enough?</p>\n<ul>\n<li>for each similar map (e.g. via clustering), you can cluster the train trajectory. how many clusters can you find?</li>\n<li>you can actually have K modes (i.e. K subclass linear classifier) and treat this as a selecting top-3 problem</li>\n</ul>\n<p>I see that the public notebook are setting K=3, but K can be greater than 3. since \"regression loss\" seek for least square you will end up in a solution that is equal distance to each of the non-matching K cluster. If you have a matching K cluster, your regression loss will be low.</p>\n<p>related: top-K ranking loss, classifiation with subclass, trajectory clustering<br>\ne.g. : <a href=\"https://arxiv.org/abs/1909.05235\" target=\"_blank\">https://arxiv.org/abs/1909.05235</a> - subclass softmax<br>\n'Multiple Centers Now, we assume that each class has K centers. Then, the similarity between the example x ….'</p>",
  "messages": [
    {
      "id": 1063956,
      "postDate": "2020-10-29T14:28:19.673Z",
      "content": "<p>first, why is there multi-mode?</p>\n<ul>\n<li>e.g. car motion is different if traffic light condition is different ( green lights or red lights )</li>\n<li>e.g. intention of the driver (maybe he wants to go destination A instead of B)</li>\n</ul>\n<p>so is three mode enough?</p>\n<ul>\n<li>for each similar map (e.g. via clustering), you can cluster the train trajectory. how many clusters can you find?</li>\n<li>you can actually have K modes (i.e. K subclass linear classifier) and treat this as a selecting top-3 problem</li>\n</ul>\n<p>I see that the public notebook are setting K=3, but K can be greater than 3. since \"regression loss\" seek for least square you will end up in a solution that is equal distance to each of the non-matching K cluster. If you have a matching K cluster, your regression loss will be low.</p>\n<p>related: top-K ranking loss, classifiation with subclass, trajectory clustering<br>\ne.g. : <a href=\"https://arxiv.org/abs/1909.05235\" target=\"_blank\">https://arxiv.org/abs/1909.05235</a> - subclass softmax<br>\n'Multiple Centers Now, we assume that each class has K centers. Then, the similarity between the example x ….'</p>",
      "rawMarkdown": "first, why is there multi-mode?\n- e.g. car motion is different if traffic light condition is different ( green lights or red lights )\n- e.g. intention of the driver (maybe he wants to go destination A instead of B)\n\nso is three mode enough?\n- for each similar map (e.g. via clustering), you can cluster the train trajectory. how many clusters can you find?\n- you can actually have K modes (i.e. K subclass linear classifier) and treat this as a selecting top-3 problem\n\nI see that the public notebook are setting K=3, but K can be greater than 3. since \"regression loss\" seek for least square you will end up in a solution that is equal distance to each of the non-matching K cluster. If you have a matching K cluster, your regression loss will be low.\n\n \n\nrelated: top-K ranking loss, classifiation with subclass, trajectory clustering\ne.g. : https://arxiv.org/abs/1909.05235 - subclass softmax\n'Multiple Centers Now, we assume that each class has K centers. Then, the similarity between the example x ....'",
      "votes": 39
    },
    {
      "id": 1072417,
      "postDate": "2020-11-08T07:31:19.873Z",
      "content": "<p>this is what you get if you plot the target on world coords.</p>\n<ul>\n<li>identify the agent type from meta data definitely helps (e.g. vehicle, pedestrian, cyclist … traveling at different speed)</li>\n<li>you can have average speed/yaw of vehicle on road as input (very strong prior)</li>\n</ul>\n<p>there could be some experiment design flaw here. The test data should be of a completely different road scene from the training dataset to avoid bias.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb8df48f983e647e2e913a98f13f7ca6f%2FSelection_209.png?generation=1604820518323524&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe05ba536a3afc41526f9167fadf877b9%2FSelection_208.png?generation=1604820537854691&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03cabdbc4e906b59cfc51e63c093365%2FSelection_207.png?generation=1604820558413844&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "this is what you get if you plot the target on world coords.\n- identify the agent type from meta data definitely helps (e.g. vehicle, pedestrian, cyclist ... traveling at different speed)\n- you can have average speed/yaw of vehicle on road as input (very strong prior)\n\nthere could be some experiment design flaw here. The test data should be of a completely different road scene from the training dataset to avoid bias.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb8df48f983e647e2e913a98f13f7ca6f%2FSelection_209.png?generation=1604820518323524&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe05ba536a3afc41526f9167fadf877b9%2FSelection_208.png?generation=1604820537854691&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03cabdbc4e906b59cfc51e63c093365%2FSelection_207.png?generation=1604820558413844&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 1072434,
          "postDate": "2020-11-08T08:02:08.103Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8666565343532f56a65a79a324bf676c%2FSelection_218.png?generation=1604822508676863&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7d827e2b859272da2fef33dd56529976%2FSelection_219.png?generation=1604822525999439&amp;alt=media\" alt=\"\"><br>\ntypo error : there is no red …</p>",
          "rawMarkdown": " ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8666565343532f56a65a79a324bf676c%2FSelection_218.png?generation=1604822508676863&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7d827e2b859272da2fef33dd56529976%2FSelection_219.png?generation=1604822525999439&alt=media)\n\ntypo error : there is no red ..."
        },
        {
          "id": 1072872,
          "postDate": "2020-11-08T19:37:55.223Z",
          "content": "<p>Do you have a hint how to implement additional inputs? I guess you are talking about the model architecture itself. Do you have an example how to add non image data to a pretained resnet34?</p>",
          "rawMarkdown": "Do you have a hint how to implement additional inputs? I guess you are talking about the model architecture itself. Do you have an example how to add non image data to a pretained resnet34?"
        },
        {
          "id": 1072888,
          "postDate": "2020-11-08T20:14:48.760Z",
          "content": "<p>In the forward function, you can extract features from the backbone, then flatten it and concatenate with any additional features you want using torch.cat(). After that, you will have to adjust the number of in_features in your dense layer (just increase it by the length of your non-image input). </p>",
          "rawMarkdown": "In the forward function, you can extract features from the backbone, then flatten it and concatenate with any additional features you want using torch.cat(). After that, you will have to adjust the number of in_features in your dense layer (just increase it by the length of your non-image input). ",
          "votes": 2
        },
        {
          "id": 1072896,
          "postDate": "2020-11-08T20:43:51.733Z",
          "content": "<p>Thanks for the quick reply <a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> </p>\n<p>Do you have a code example? I thought about something like this:</p>\n<p>1# forward</p>\n<pre><code>    def forward(self, x, non-image-data):\n        x = self.backbone.conv1(x)\n        x = self.backbone.bn1(x)\n        x = self.backbone.relu(x)\n        x = self.backbone.maxpool(x)\n\n        x = self.backbone.layer1(x)\n        x = self.backbone.layer2(x)\n        x = self.backbone.layer3(x)\n        x = self.backbone.layer4(x)\n\n        x = self.backbone.avgpool(x)\n        x = torch.flatten(x, 1)\n        y = non-image-data\n        x = torch.cat((x, y), dim=1)\n\n        x = self.head(x)\n        x = self.logit(x)\n         [ ...... ]\n</code></pre>\n<p>2# init</p>\n<pre><code>def __init__(self, cfg: Dict, num_modes=3):\n       [ ...... ]\n        # You can add more layers here.\n        self.head = nn.Sequential(\n            # nn.Dropout(0.2),\n            nn.Linear(in_features=backbone_out_features+len(non-image-data), out_features=4096),\n        )\n        self.num_preds = num_targets * num_modes\n        self.num_modes = num_modes\n\n        self.logit = nn.Linear(4096, out_features=self.num_preds + num_modes)\n</code></pre>\n<p>What do you think?</p>",
          "rawMarkdown": "Thanks for the quick reply @szacho \n\nDo you have a code example? I thought about something like this:\n\n1# forward\n\n```\n    def forward(self, x, non-image-data):\n        x = self.backbone.conv1(x)\n        x = self.backbone.bn1(x)\n        x = self.backbone.relu(x)\n        x = self.backbone.maxpool(x)\n\n        x = self.backbone.layer1(x)\n        x = self.backbone.layer2(x)\n        x = self.backbone.layer3(x)\n        x = self.backbone.layer4(x)\n\n        x = self.backbone.avgpool(x)\n        x = torch.flatten(x, 1)\n        y = non-image-data\n        x = torch.cat((x, y), dim=1)\n\n        x = self.head(x)\n        x = self.logit(x)\n         [ ...... ]\n```\n\n2# init\n```\ndef __init__(self, cfg: Dict, num_modes=3):\n       [ ...... ]\n        # You can add more layers here.\n        self.head = nn.Sequential(\n            # nn.Dropout(0.2),\n            nn.Linear(in_features=backbone_out_features+len(non-image-data), out_features=4096),\n        )\n        self.num_preds = num_targets * num_modes\n        self.num_modes = num_modes\n\n        self.logit = nn.Linear(4096, out_features=self.num_preds + num_modes)\n```\n\nWhat do you think?",
          "votes": 1
        },
        {
          "id": 1073382,
          "postDate": "2020-11-09T13:25:13.987Z",
          "content": "<p><a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> <br>\nSeems fine, I implemented it in a very similar if not the same way. Good job!</p>",
          "rawMarkdown": "@benbla \nSeems fine, I implemented it in a very similar if not the same way. Good job!",
          "votes": 1
        },
        {
          "id": 1073631,
          "postDate": "2020-11-09T18:12:09.943Z",
          "content": "<p>here is the loss breakdown by label:</p>\n<pre><code>for non-chopped data:\nlocal cv loss: 9.346676  (lb = 15.034)\nbreakdown: \nPERCEPTION_LABEL_CAR         9.259042 \nPERCEPTION_LABEL_CYCLIST    17.221859 \nPERCEPTION_LABEL_PEDESTRIAN  4.827463\n</code></pre>",
          "rawMarkdown": "here is the loss breakdown by label:\n\n```\n\n\nfor non-chopped data:\nlocal cv loss: 9.346676  (lb = 15.034)\nbreakdown: \nPERCEPTION_LABEL_CAR         9.259042 \nPERCEPTION_LABEL_CYCLIST    17.221859 \nPERCEPTION_LABEL_PEDESTRIAN  4.827463\n \n\n\n\n```",
          "votes": 3
        },
        {
          "id": 1073880,
          "postDate": "2020-11-10T02:22:08.400Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffc4872f9d3751a914d92103db8769fc0%2FSelection_063.png?generation=1604974926304836&amp;alt=media\" alt=\"\"></p>\n<p>i wonder are these ground truth senor errors?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd29aed8a457cdcd270ca91cd9562457b%2FSelection_067.png?generation=1604975711807483&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffc4872f9d3751a914d92103db8769fc0%2FSelection_063.png?generation=1604974926304836&alt=media)\n\ni wonder are these ground truth senor errors?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd29aed8a457cdcd270ca91cd9562457b%2FSelection_067.png?generation=1604975711807483&alt=media)"
        },
        {
          "id": 1074374,
          "postDate": "2020-11-10T15:26:38.020Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a>. Did you feed \"history_positions\" and other data into the model? During training, is \"history_availabilities \" useful to imply the available agent in history frames?</p>",
          "rawMarkdown": "Hi, @szacho. Did you feed \"history_positions\" and other data into the model? During training, is \"history_availabilities \" useful to imply the available agent in history frames?"
        },
        {
          "id": 1074375,
          "postDate": "2020-11-10T15:29:06.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> <br>\nIntegrating \"history_positions\" leads to a huge overfit and a big derivation between loss and lb score to me.</p>",
          "rawMarkdown": "@spicychicken38 \nIntegrating \"history_positions\" leads to a huge overfit and a big derivation between loss and lb score to me.",
          "replies": [
            {
              "id": 1074776,
              "postDate": "2020-11-11T03:37:10.663Z",
              "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> <br>\n  Integrating \"history_positions\" leads to a huge overfit and a big derivation between loss and lb score to me.</p>\n</blockquote>\n<p>Hi, <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a>. I don't think this huge deviation is caused by overfitting. According to my experiment, there is a data mismatch that is caused by the differences between train.zarr and validation.zarr(or test.zarr), which can lead to this huge deviation. Because the \"history_availabilities\" in train.zarr are always 1, which means agents exist in every frame in the training sample that they belong to. However, this is totally different in the chopped validation set or the testing set. So I think this might be the main reason why we have such a huge deviation when we are adding \"history_positions\".</p>",
              "rawMarkdown": "> @spicychicken38 \n> Integrating \"history_positions\" leads to a huge overfit and a big derivation between loss and lb score to me.\n\nHi, @benbla. I don't think this huge deviation is caused by overfitting. According to my experiment, there is a data mismatch that is caused by the differences between train.zarr and validation.zarr(or test.zarr), which can lead to this huge deviation. Because the \"history_availabilities\" in train.zarr are always 1, which means agents exist in every frame in the training sample that they belong to. However, this is totally different in the chopped validation set or the testing set. So I think this might be the main reason why we have such a huge deviation when we are adding \"history_positions\".",
              "votes": 1
            }
          ]
        },
        {
          "id": 1074389,
          "postDate": "2020-11-10T15:44:36.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> In my case, there is also a huge deviation between validation score and training loss. But I can not certainly say that it is overfitting. I also think that it is also possible that the differences of \"history_availabilities\" between the training set and the chopped validation set to cause this huge deviation. But I can not verify it. As mention <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> also meet the same problem.</p>",
          "rawMarkdown": "@benbla In my case, there is also a huge deviation between validation score and training loss. But I can not certainly say that it is overfitting. I also think that it is also possible that the differences of \"history_availabilities\" between the training set and the chopped validation set to cause this huge deviation. But I can not verify it. As mention [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762), @frankpanxj also meet the same problem."
        },
        {
          "id": 1074393,
          "postDate": "2020-11-10T15:50:07.757Z",
          "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> No, I feed the agent state vector as described in this article: <a href=\"https://arxiv.org/abs/1809.10732\" target=\"_blank\">https://arxiv.org/abs/1809.10732</a></p>",
          "rawMarkdown": "@spicychicken38 No, I feed the agent state vector as described in this article: https://arxiv.org/abs/1809.10732"
        },
        {
          "id": 1074406,
          "postDate": "2020-11-10T16:06:30.183Z",
          "content": "<p><a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> But we don't have the velocity, acceleration mentioned in the article. How you implement this?</p>",
          "rawMarkdown": "@szacho But we don't have the velocity, acceleration mentioned in the article. How you implement this?"
        },
        {
          "id": 1074459,
          "postDate": "2020-11-10T17:07:33.470Z",
          "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> Velocity is available in zarr data, but you can get everything from data returned by AgentDataset - you have history_positions and history_yaws in time. Features like velocity, acceleration, and yaw change rate can be calculated using basic physics knowledge.</p>",
          "rawMarkdown": "@spicychicken38 Velocity is available in zarr data, but you can get everything from data returned by AgentDataset - you have history_positions and history_yaws in time. Features like velocity, acceleration, and yaw change rate can be calculated using basic physics knowledge.",
          "votes": 1
        },
        {
          "id": 1074777,
          "postDate": "2020-11-11T03:38:06.713Z",
          "content": "<p><a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> Thanks for your views. I will try to do some experiments on that later!</p>",
          "rawMarkdown": "@szacho Thanks for your views. I will try to do some experiments on that later!"
        },
        {
          "id": 1075234,
          "postDate": "2020-11-11T14:09:26.740Z",
          "content": "<p><a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> As you mentioned we can get the velocity from AgentDataset. But what about the agent's <strong>label_probabilities</strong>? seems <strong>agent_samples.py</strong> won't return them. If possible, can suggest some method to get them along with the \"data\" object?</p>",
          "rawMarkdown": "@szacho As you mentioned we can get the velocity from AgentDataset. But what about the agent's **label_probabilities**? seems **agent_samples.py** won't return them. If possible, can suggest some method to get them along with the \"data\" object?\n"
        },
        {
          "id": 1075254,
          "postDate": "2020-11-11T14:19:33.403Z",
          "content": "<p>for the file: <br>\n/l5kit/sampling/agent_sampling.py</p>\n<pre><code>def generate_agent_sample(...):\n     ...\n        \"extent\": agent_extent_m,\n        #\"agent\": selected_agent, # &lt;add here&gt;\n        \"label_probabilities\": selected_agent[\"label_probabilities\"], # &lt;add here&gt;\n</code></pre>\n<p>for the file: <br>\n/l5kit/dataset/ego.py</p>\n<pre><code>class EgoDataset(Dataset):\n...\ndef get_frame(...):\n\n            \"extent\": data[\"extent\"],\n            #\"agent\": data[\"agent\"], #&lt;add here&gt;\n            \"label_probabilities\": data[\"label_probabilities\"], #&lt;add here&gt;\n        }\n</code></pre>",
          "rawMarkdown": "for the file: \n/l5kit/sampling/agent_sampling.py\n```\ndef generate_agent_sample(...):\n     ...\n        \"extent\": agent_extent_m,\n        #\"agent\": selected_agent, # <add here>\n        \"label_probabilities\": selected_agent[\"label_probabilities\"], # <add here>\n```\n\nfor the file: \n/l5kit/dataset/ego.py\n```\nclass EgoDataset(Dataset):\n...\ndef get_frame(...):\n\n            \"extent\": data[\"extent\"],\n            #\"agent\": data[\"agent\"], #<add here>\n            \"label_probabilities\": data[\"label_probabilities\"], #<add here>\n        }\n```",
          "votes": 5,
          "replies": [
            {
              "id": 1075771,
              "postDate": "2020-11-12T00:21:30.830Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , Thanks for your insights. How to add this custom l5kit library to one's kaggle notebook? Instead I thought of calculating velocities in rudimentary way like this (I guess we have to calculate yaw rate and acceleration ourselves anyway):</p>\n<pre><code>    if data['history_availabilities'].to(device)shape[1] &gt; 1: \n        velocities = torch.sqrt(torch.sum((pose_data[:,1:,:] - pose_data[:,0:-1, :]) ** 2, 2)) / dt\n        avg_velocity = torch.mean(velocities, dim=1).view(bs,1)\n        #acceleration = torch.abs(velocities[:, 1:] - velocities[:, 0:-1])\n        avg_acceleration = torch.mean(torch.abs(velocities[:, 1:] - velocities[:, 0:-1]) / dt, dim=1).view(bs,1)\n        yaw_rate = torch.mean(torch.abs(yaw_data[:, 1:] - yaw_data[:, 0:-1]) /  dt, dim=1).view(bs,1)\n\n        avg_velocity = avg_velocity / MAX_VELOCITY # 20 m/s\n        avg_acceleration = avg_acceleration / MAX_ACCELERATION # 2m/s^2\n        yaw_rate = yaw_rate / MAX_YAW_RATE # 45 deg/s\n\n    else:\n        avg_velocity = 0.0;   avg_acceleration = 0.0;    yaw_rate = 0.0; \n</code></pre>\n<p>Is there anything I'm missing here? Any any exception handling required?</p>",
              "rawMarkdown": "Hi @hengck23 , Thanks for your insights. How to add this custom l5kit library to one's kaggle notebook? Instead I thought of calculating velocities in rudimentary way like this (I guess we have to calculate yaw rate and acceleration ourselves anyway):\n\n```\n    if data['history_availabilities'].to(device)shape[1] > 1: \n        velocities = torch.sqrt(torch.sum((pose_data[:,1:,:] - pose_data[:,0:-1, :]) ** 2, 2)) / dt\n        avg_velocity = torch.mean(velocities, dim=1).view(bs,1)\n        #acceleration = torch.abs(velocities[:, 1:] - velocities[:, 0:-1])\n        avg_acceleration = torch.mean(torch.abs(velocities[:, 1:] - velocities[:, 0:-1]) / dt, dim=1).view(bs,1)\n        yaw_rate = torch.mean(torch.abs(yaw_data[:, 1:] - yaw_data[:, 0:-1]) /  dt, dim=1).view(bs,1)\n        \n        avg_velocity = avg_velocity / MAX_VELOCITY # 20 m/s\n        avg_acceleration = avg_acceleration / MAX_ACCELERATION # 2m/s^2\n        yaw_rate = yaw_rate / MAX_YAW_RATE # 45 deg/s\n\n    else:\n        avg_velocity = 0.0;   avg_acceleration = 0.0;    yaw_rate = 0.0; \n```\n\nIs there anything I'm missing here? Any any exception handling required?",
              "votes": 1
            },
            {
              "id": 1082506,
              "postDate": "2020-11-18T00:59:28.477Z",
              "content": "<p>Adding these state inputs leads to large overfitting for me, anyone else seeing similar results ?</p>",
              "rawMarkdown": "Adding these state inputs leads to large overfitting for me, anyone else seeing similar results ?"
            }
          ]
        },
        {
          "id": 1075328,
          "postDate": "2020-11-11T15:35:08.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Oh you are using l5kit/ as part of your code. I am using it as a library. Sure I will do require changes. Thanks for your quick response.</p>",
          "rawMarkdown": "@hengck23  Oh you are using l5kit/ as part of your code. I am using it as a library. Sure I will do require changes. Thanks for your quick response."
        },
        {
          "id": 1078444,
          "postDate": "2020-11-14T19:50:00.310Z",
          "content": "<p>Speaking of velocities, agent_sampling.py and ego.py have been updated. You can now get velocity (and speed) through the <code>AgentDataset()</code> class.</p>",
          "rawMarkdown": "Speaking of velocities, agent_sampling.py and ego.py have been updated. You can now get velocity (and speed) through the `AgentDataset()` class.",
          "votes": 1
        },
        {
          "id": 1078452,
          "postDate": "2020-11-14T20:04:05.507Z",
          "content": "<p>Where can I find the updates ?</p>",
          "rawMarkdown": "Where can I find the updates ?"
        },
        {
          "id": 1078462,
          "postDate": "2020-11-14T20:17:03.597Z",
          "content": "<p>l5kit github repository, <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "l5kit github repository, [here](https://github.com/lyft/l5kit)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1076901,
      "postDate": "2020-11-13T02:52:24.260Z",
      "content": "<p>surprise, surprise, surprise …<br>\nI put the conclusion first and then the observations.</p>\n<p>conclusion:</p>\n<ul>\n<li>there are two types of trajectories: </li>\n</ul>\n<ol>\n<li>parametric curves, that obey kinematics. you can get very accurate results via curve fitting (and maybe using lstm of past history points)</li>\n<li>non-parametric and irregular curves</li>\n</ol>\n<p>we probably want to detect, predict, post or pre-processing them separately. they are easy to identify via past histories.</p>\n<hr>\n<p>observation:</p>\n<ul>\n<li>I study the metric and note that the effect of l2 loss is very much greater than the confidence values (because world coords is in meters?). i.e. in a prediction of mode M, as long as one of them has low l2 loss (even if the confidence is low) the kaggle metric will be low.</li>\n</ul>\n<p>see below simulation:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcadc4dbd60a83b50138a2a7daaf45c3b%2FSelection_149.png?generation=1605235825678435&amp;alt=media\" alt=\"\"></p>\n<p>hence I trained a mode-7 predictor from my current mode-3 predictor. i just freeze the backbone weights and change the head. it is very fast to train. on chopped validation data, kaggle metric reduces from 15+ to 11+ within hours.</p>\n<p>but when I inspect the results,  I was surprised</p>",
      "rawMarkdown": "surprise, surprise, surprise ...\nI put the conclusion first and then the observations.\n\nconclusion:\n- there are two types of trajectories: \n1. parametric curves, that obey kinematics. you can get very accurate results via curve fitting (and maybe using lstm of past history points)\n2. non-parametric and irregular curves\n\nwe probably want to detect, predict, post or pre-processing them separately. they are easy to identify via past histories.\n\n--- \n\nobservation:\n- I study the metric and note that the effect of l2 loss is very much greater than the confidence values (because world coords is in meters?). i.e. in a prediction of mode M, as long as one of them has low l2 loss (even if the confidence is low) the kaggle metric will be low.\n\nsee below simulation:\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcadc4dbd60a83b50138a2a7daaf45c3b%2FSelection_149.png?generation=1605235825678435&alt=media)\n\nhence I trained a mode-7 predictor from my current mode-3 predictor. i just freeze the backbone weights and change the head. it is very fast to train. on chopped validation data, kaggle metric reduces from 15+ to 11+ within hours.\n\nbut when I inspect the results,  I was surprised",
      "votes": 4,
      "replies": [
        {
          "id": 1076902,
          "postDate": "2020-11-13T02:53:07.980Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F28b6cea52808ab6ba5a725609a2e1da3%2FSelection_144.png?generation=1605235977467342&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3057c4f9087c723954509a2d25e0a2ce%2FSelection_157.png?generation=1605261672198016&amp;alt=media\" alt=\"\"><br>\nNote: if there are shakeup, it can be due to just a few outliers</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F28b6cea52808ab6ba5a725609a2e1da3%2FSelection_144.png?generation=1605235977467342&alt=media)\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3057c4f9087c723954509a2d25e0a2ce%2FSelection_157.png?generation=1605261672198016&alt=media)\nNote: if there are shakeup, it can be due to just a few outliers"
        },
        {
          "id": 1076905,
          "postDate": "2020-11-13T02:55:36.757Z",
          "content": "<p>so, what are the cases, mode-7 performs better?<br>\ntruth: disconnected black dots<br>\n(history is not shown)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F490491b0c3190dc69ef6a15d00b88b86%2FSelection_145.png?generation=1605236075073834&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F26cb2e998011236e99e4b3de568a11f2%2FSelection_146.png?generation=1605236092430499&amp;alt=media\" alt=\"\"></p>\n<p>you can see that we can gain a lot if we can estimate the \"straight trajectory\" more accurate … which seems to be an easy task </p>",
          "rawMarkdown": "so, what are the cases, mode-7 performs better?\ntruth: disconnected black dots\n(history is not shown)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F490491b0c3190dc69ef6a15d00b88b86%2FSelection_145.png?generation=1605236075073834&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F26cb2e998011236e99e4b3de568a11f2%2FSelection_146.png?generation=1605236092430499&alt=media)\n\n\nyou can see that we can gain a lot if we can estimate the \"straight trajectory\" more accurate ... which seems to be an easy task ",
          "votes": 1
        },
        {
          "id": 1077784,
          "postDate": "2020-11-13T23:19:28Z",
          "content": "<p>Nice catch but the competition metric is biased towards long trajectories. Sigma is fixed to 1 which is not real! It should increase as the predictions goes further in time.</p>",
          "rawMarkdown": "Nice catch but the competition metric is biased towards long trajectories. Sigma is fixed to 1 which is not real! It should increase as the predictions goes further in time.\n"
        },
        {
          "id": 1078686,
          "postDate": "2020-11-15T07:12:14.137Z",
          "content": "<p>as an example<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F004de876653e436c02f85469f7911ca4%2FSelection_185.png?generation=1605424264162784&amp;alt=media\" alt=\"\"></p>\n<p>you can add other labels for specific post-processing, e.g. di-biasing (the kaggle multiplier trick in regression problem, which is essentially manually adjusting the std of error)</p>",
          "rawMarkdown": "as an example\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F004de876653e436c02f85469f7911ca4%2FSelection_185.png?generation=1605424264162784&alt=media)\n\nyou can add other labels for specific post-processing, e.g. di-biasing (the kaggle multiplier trick in regression problem, which is essentially manually adjusting the std of error)",
          "votes": 1
        },
        {
          "id": 1078739,
          "postDate": "2020-11-15T08:49:06.343Z",
          "content": "<p>How do you set the <code>is_parametric</code> flag? It's not clear to me. Do you use a custom curve fitting function?</p>",
          "rawMarkdown": "How do you set the `is_parametric` flag? It's not clear to me. Do you use a custom curve fitting function?"
        },
        {
          "id": 1078773,
          "postDate": "2020-11-15T09:36:13.760Z",
          "content": "<p>e.g. for each train sample, you have x =[history, future]</p>\n<p>then you devise a curve function, predict_future = f(history). you will have to design this by hand, making some assumptions. you can google for some kinematics model.</p>\n<p>now each train sample will have x =[history, f(history), future].</p>\n<p>compute error of f(history), e.g. l2 loss (f(history), future). now if this error is small, you can label is as True (i.e. you can recover future values from f(history))</p>",
          "rawMarkdown": "e.g. for each train sample, you have x =[history, future]\n\nthen you devise a curve function, predict\\_future = f(history). you will have to design this by hand, making some assumptions. you can google for some kinematics model.\n\nnow each train sample will have x =[history, f(history), future].\n\ncompute error of f(history), e.g. l2 loss (f(history), future). now if this error is small, you can label is as True (i.e. you can recover future values from f(history))",
          "votes": 1
        },
        {
          "id": 1078778,
          "postDate": "2020-11-15T09:44:17.033Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, this helps :)</p>",
          "rawMarkdown": "Thanks @hengck23, this helps :)"
        }
      ]
    },
    {
      "id": 1072969,
      "postDate": "2020-11-09T00:44:23.703Z",
      "content": "<p>there is much more information in the zarr file (instead of the AgentDataset)</p>\n<p>e.g. you get velocity and label probability</p>\n<pre><code>AGENT_DTYPE = [\n    (\"centroid\", np.float64, (2,)),\n    (\"extent\", np.float32, (3,)),\n    (\"yaw\", np.float32),\n    (\"velocity\", np.float32, (2,)),\n    (\"track_id\", np.uint64),\n    (\"label_probabilities\", np.float32, (len(LABELS),)),\n]\n</code></pre>\n<p>here is an example of reading the raw zarr file. note that you can recover the whole track!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa3d90739319a1ec4bea51004002f61d3%2FSelection_230.png?generation=1604882661363275&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "there is much more information in the zarr file (instead of the AgentDataset)\n\ne.g. you get velocity and label probability\n\n```\nAGENT_DTYPE = [\n    (\"centroid\", np.float64, (2,)),\n    (\"extent\", np.float32, (3,)),\n    (\"yaw\", np.float32),\n    (\"velocity\", np.float32, (2,)),\n    (\"track_id\", np.uint64),\n    (\"label_probabilities\", np.float32, (len(LABELS),)),\n]\n\n\n```\n\nhere is an example of reading the raw zarr file. note that you can recover the whole track!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa3d90739319a1ec4bea51004002f61d3%2FSelection_230.png?generation=1604882661363275&alt=media)",
      "votes": 1
    },
    {
      "id": 1069034,
      "postDate": "2020-11-04T02:08:55.307Z",
      "content": "<p>you can submit the top-1, top-2 and top-3 results to prob the correctness of your prediction. then it may be possible to recalibrate  or adjust the temperate of the confidence.</p>\n<p>some of the  predictions are clearly only single-mode, e.g. moving in a straight lane or not moving.  in that case a single high confidence trajectory prediction may be better?</p>",
      "rawMarkdown": "you can submit the top-1, top-2 and top-3 results to prob the correctness of your prediction. then it may be possible to recalibrate  or adjust the temperate of the confidence.\n\nsome of the  predictions are clearly only single-mode, e.g. moving in a straight lane or not moving.  in that case a single high confidence trajectory prediction may be better?",
      "votes": 1
    },
    {
      "id": 1067368,
      "postDate": "2020-11-02T13:52:56.493Z",
      "content": "<p>self-supervised?</p>\n<p>for the test data, say we ask for number of history frame = 30.<br>\nwe use t-30 to t-20 as input and train an embedding to predict t-20 to t.</p>\n<p>would it work?</p>\n<p>how about :<br>\nfuture = predict (history) and then history1 = predict(future) … like to and from lanuage translation …<br>\nwould that be a kind of augmentation (i.e. history --&gt;history1)?</p>",
      "rawMarkdown": "self-supervised?\n\nfor the test data, say we ask for number of history frame = 30.\nwe use t-30 to t-20 as input and train an embedding to predict t-20 to t.\n\nwould it work?\n\nhow about :\nfuture = predict (history) and then history1 = predict(future) ... like to and from lanuage translation ...\nwould that be a kind of augmentation (i.e. history -->history1)?",
      "votes": 1
    },
    {
      "id": 1066792,
      "postDate": "2020-11-02T03:34:06.673Z",
      "content": "<p>MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction<br>\n<a href=\"https://openaccess.thecvf.com/content_CVPR_2020/papers/Marchetti_MANTRA_Memory_Augmented_Networks_for_Multiple_Trajectory_Prediction_CVPR_2020_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPR_2020/papers/Marchetti_MANTRA_Memory_Augmented_Networks_for_Multiple_Trajectory_Prediction_CVPR_2020_paper.pdf</a></p>\n<p><img src=\"https://images.deepai.org/converted-papers/2006.03340/x3.png\" alt=\"\"></p>",
      "rawMarkdown": "MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction\nhttps://openaccess.thecvf.com/content_CVPR_2020/papers/Marchetti_MANTRA_Memory_Augmented_Networks_for_Multiple_Trajectory_Prediction_CVPR_2020_paper.pdf\n\n\n![](https://images.deepai.org/converted-papers/2006.03340/x3.png)",
      "votes": 1
    },
    {
      "id": 1064120,
      "postDate": "2020-10-29T18:29:54.943Z",
      "content": "<p>In the original paper of the base model that lyft provided to us it seems to show that 3 is empirically the most performant number of modes <a href=\"https://arxiv.org/pdf/1809.10732.pdf\" target=\"_blank\">https://arxiv.org/pdf/1809.10732.pdf</a></p>\n<p>1, 2, and 4 modes do worse. This has been reproduced by a few other papers as well. </p>",
      "rawMarkdown": "In the original paper of the base model that lyft provided to us it seems to show that 3 is empirically the most performant number of modes https://arxiv.org/pdf/1809.10732.pdf\n\n1, 2, and 4 modes do worse. This has been reproduced by a few other papers as well. ",
      "votes": 1,
      "replies": [
        {
          "id": 1064288,
          "postDate": "2020-10-29T23:49:07.947Z",
          "content": "<p>thanks for the paper.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc0e2508c364917765fd2d65e8436fcc8%2FSelection_030.png?generation=1604014930131979&amp;alt=media\" alt=\"\"></p>\n<p>it seems that it is due to the nature road junction. in that case, we can just train for going straight, turn left and turn right. i need to check the distribution/frequency of turns of junctions. thanks!</p>",
          "rawMarkdown": "thanks for the paper.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc0e2508c364917765fd2d65e8436fcc8%2FSelection_030.png?generation=1604014930131979&alt=media)\n\nit seems that it is due to the nature road junction. in that case, we can just train for going straight, turn left and turn right. i need to check the distribution/frequency of turns of junctions. thanks!",
          "votes": 1
        },
        {
          "id": 1065049,
          "postDate": "2020-10-30T19:38:03.467Z",
          "content": "<p>Intresting, thank you for the paper.</p>\n<p>The relations are greatly explained and visualized.</p>",
          "rawMarkdown": "Intresting, thank you for the paper.\n\nThe relations are greatly explained and visualized."
        },
        {
          "id": 1065578,
          "postDate": "2020-10-31T13:13:12.893Z",
          "content": "<p>Does anyone know if the mentioned MTP loss is somewhere available as public code?</p>",
          "rawMarkdown": "Does anyone know if the mentioned MTP loss is somewhere available as public code?"
        },
        {
          "id": 1065608,
          "postDate": "2020-10-31T13:58:34.417Z",
          "content": "<p>Here, <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/prediction/models/mtp.py\" target=\"_blank\">https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/prediction/models/mtp.py</a> </p>",
          "rawMarkdown": "Here, https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/prediction/models/mtp.py ",
          "votes": 1
        },
        {
          "id": 1065732,
          "postDate": "2020-10-31T16:56:38.637Z",
          "content": "<p>Thank you :)</p>",
          "rawMarkdown": "Thank you :)"
        }
      ]
    },
    {
      "id": 1076954,
      "postDate": "2020-11-13T04:28:13.600Z",
      "content": "<p>to disable image rasterizer and just do prediction using history information:</p>\n<pre><code>    dm = LocalDataManager()\n    rasterizer = None #build_rasterizer(cfg, dm)\n\n    zarr = ChunkedDataset( dm.require('scenes/train.zarr')).open()\n    dataset = MyAgentDataset(cfg, zarr, rasterizer) \n\n\n\nin ego.py: (around line 99)\n\nclass EgoDataset(Dataset):\n\ndef get_frame(self, scene_index ...):\n\n        # 0,1,C -&gt; C,0,1\n        # &lt;hck&gt;\n        if data[\"image\"] is not None:\n            image = data[\"image\"].transpose(2, 0, 1)\n        else:\n            image = 0 #dummy\n</code></pre>\n<p>i am thinking of using:</p>\n<ul>\n<li>image : e.g. up to past 5 history frames</li>\n<li>locations : e.g. up to past 100 history frames</li>\n</ul>\n<p>so i need to create 2 dataset, one with rasterizer and one without (to speed up cpu data loading).</p>",
      "rawMarkdown": "to disable image rasterizer and just do prediction using history information:\n\n```\n    dm = LocalDataManager()\n    rasterizer = None #build_rasterizer(cfg, dm)\n\n    zarr = ChunkedDataset( dm.require('scenes/train.zarr')).open()\n    dataset = MyAgentDataset(cfg, zarr, rasterizer) \n\n\n\nin ego.py: (around line 99)\n\nclass EgoDataset(Dataset):\n\ndef get_frame(self, scene_index ...):\n\n        # 0,1,C -> C,0,1\n        # <hck>\n        if data[\"image\"] is not None:\n            image = data[\"image\"].transpose(2, 0, 1)\n        else:\n            image = 0 #dummy\n\n```\n\ni am thinking of using:\n- image : e.g. up to past 5 history frames\n- locations : e.g. up to past 100 history frames\n\nso i need to create 2 dataset, one with rasterizer and one without (to speed up cpu data loading).\n ",
      "votes": 2,
      "replies": [
        {
          "id": 1077782,
          "postDate": "2020-11-13T23:15:21.343Z",
          "content": "<p>I reported this as a bug long time ago on github and I guess it is now fixed</p>",
          "rawMarkdown": "I reported this as a bug long time ago on github and I guess it is now fixed"
        },
        {
          "id": 1077851,
          "postDate": "2020-11-14T02:13:08.017Z",
          "content": "<p>You mean that by doing disabling the rasterizer, we can use history information for improvement regardless of the differences of history_availabilities between train.zarr and test.zarr?</p>",
          "rawMarkdown": "You mean that by doing disabling the rasterizer, we can use history information for improvement regardless of the differences of history_availabilities between train.zarr and test.zarr?"
        },
        {
          "id": 1077896,
          "postDate": "2020-11-14T04:22:54.570Z",
          "content": "<p>rendering long history frames is very slow. hence I want to experiment using history positions only without image. (or few history images)</p>\n<p>this is non- fixed length sequence learning, e.g</p>\n<pre><code>input:\nx(t), x(t-1),x(t-2) ... &lt;end&gt;\n</code></pre>\n<p>you can treat it as LSTM seq with an \"end\" token. alternatively you can pad to equal length.</p>",
          "rawMarkdown": "rendering long history frames is very slow. hence I want to experiment using history positions only without image. (or few history images)\n\nthis is non- fixed length sequence learning, e.g\n\n```\n\ninput:\nx(t), x(t-1),x(t-2) ... <end>\n\n``` \n\nyou can treat it as LSTM seq with an \"end\" token. alternatively you can pad to equal length."
        },
        {
          "id": 1078104,
          "postDate": "2020-11-14T11:21:04.800Z",
          "content": "<p>You can use <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/rasterization/stub_rasterizer.py\" target=\"_blank\">StubRasterizer</a> to render only blank images</p>",
          "rawMarkdown": "You can use [StubRasterizer](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/rasterization/stub_rasterizer.py) to render only blank images"
        }
      ]
    },
    {
      "id": 1065510,
      "postDate": "2020-10-31T11:17:03.593Z",
      "content": "<p>i wonder did anyone plot and visualise the predicted trajectory?</p>\n<p>i don't think it would be a perfect parametric curve. it might look a little \"wiggling\" and not smooth. So there are be some post-processing required, or better still, should we  predict the parameters of the curve of trajectory instead?</p>\n<p>history trajectory can also be represented as a parametric curve</p>",
      "rawMarkdown": "i wonder did anyone plot and visualise the predicted trajectory?\n\ni don't think it would be a perfect parametric curve. it might look a little \"wiggling\" and not smooth. So there are be some post-processing required, or better still, should we  predict the parameters of the curve of trajectory instead?\n\nhistory trajectory can also be represented as a parametric curve",
      "votes": 2,
      "replies": [
        {
          "id": 1065524,
          "postDate": "2020-10-31T11:33:06.847Z",
          "content": "<p>There has been some discussion about maybe using a Kalman filter to smooth out predictions, although at the moment my predicted trajectories tend to be pretty smooth:  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Fe89a6a4e6fe1239bf3aa224d2fe83537%2Finbox_314368_2ae6a5e2632e97085e567e4951c51536_mask.png?generation=1604143906678498&amp;alt=media\" alt=\"\"></p>\n<p>The targets often have some amount of noise in them, so being smoother than the targets might not give that much improvement</p>",
          "rawMarkdown": "There has been some discussion about maybe using a Kalman filter to smooth out predictions, although at the moment my predicted trajectories tend to be pretty smooth:  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Fe89a6a4e6fe1239bf3aa224d2fe83537%2Finbox_314368_2ae6a5e2632e97085e567e4951c51536_mask.png?generation=1604143906678498&alt=media)\n\nThe targets often have some amount of noise in them, so being smoother than the targets might not give that much improvement",
          "votes": 5
        },
        {
          "id": 1066394,
          "postDate": "2020-11-01T17:54:15.257Z",
          "content": "<p>seems that predicting the end-point correctly is very important. </p>",
          "rawMarkdown": "seems that predicting the end-point correctly is very important. "
        }
      ]
    },
    {
      "id": 1077979,
      "postDate": "2020-11-14T07:13:11.497Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3093b436d5ee8a87c7594c04ed1fb923%2FSelection_164.png?generation=1605337987662608&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3093b436d5ee8a87c7594c04ed1fb923%2FSelection_164.png?generation=1605337987662608&alt=media)",
      "replies": [
        {
          "id": 1078092,
          "postDate": "2020-11-14T10:40:49.500Z",
          "content": "<p>I believe, that picture is rather misleading with different counts in each bin. You may subsample the 50 avails part to match the count for the others, or plot e.g. violin graphs to get a better view on the loss distribution.</p>",
          "rawMarkdown": "I believe, that picture is rather misleading with different counts in each bin. You may subsample the 50 avails part to match the count for the others, or plot e.g. violin graphs to get a better view on the loss distribution."
        },
        {
          "id": 1078093,
          "postDate": "2020-11-14T10:50:16.457Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F84f497214505a4fcaaac8589277f494c%2FSelection_165.png?generation=1605350990328844&amp;alt=media\" alt=\"\"></p>\n<p>count of samples vs samples' target availability</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F84f497214505a4fcaaac8589277f494c%2FSelection_165.png?generation=1605350990328844&alt=media)\n\ncount of samples vs samples' target availability",
          "votes": 1
        },
        {
          "id": 1078110,
          "postDate": "2020-11-14T11:31:02.323Z",
          "content": "<p>I know, that's why your first Image gives a misleading view on the loss distribution. It looks like the loss distribution for 50 avails is much broader than for example 49 avails, but in reality they are very similar.</p>",
          "rawMarkdown": "I know, that's why your first Image gives a misleading view on the loss distribution. It looks like the loss distribution for 50 avails is much broader than for example 49 avails, but in reality they are very similar."
        },
        {
          "id": 1078120,
          "postDate": "2020-11-14T11:43:59.517Z",
          "content": "<p>i think the loss for e.g. availability=10 would be small because only \"10 nearby points are counted\" in the evaluation metric.</p>\n<p>in the case where availability=50, errors can get accumulated for many predicted points. hence the loss will be higher.</p>",
          "rawMarkdown": "i think the loss for e.g. availability=10 would be small because only \"10 nearby points are counted\" in the evaluation metric.\n\nin the case where availability=50, errors can get accumulated for many predicted points. hence the loss will be higher."
        },
        {
          "id": 1078142,
          "postDate": "2020-11-14T12:17:25.673Z",
          "content": "<p>of course that is the case, you can see it in the l5kit <code>time_displace</code> error, which is exactly evaluating this. But naturally, 49 and 50 are similar, which is not what your first plot suggests, that's all I am saying. </p>",
          "rawMarkdown": "of course that is the case, you can see it in the l5kit `time_displace` error, which is exactly evaluating this. But naturally, 49 and 50 are similar, which is not what your first plot suggests, that's all I am saying. ",
          "votes": 1
        },
        {
          "id": 1078330,
          "postDate": "2020-11-14T16:32:17.310Z",
          "content": "<p>Yes, the plot is rather misleading because the frequency of availability=50 makes it difficult to evaluate  the distribution of losses. Because you are viewing thousands of samples rather than just a few it makes the spread look much more drastic even though it could just be a very small subset that makes the graph look like the spread is actually much greater. If you had thousands of samples for any other availability near 50 you might get a visualization that looks similar to availability=50</p>",
          "rawMarkdown": "Yes, the plot is rather misleading because the frequency of availability=50 makes it difficult to evaluate  the distribution of losses. Because you are viewing thousands of samples rather than just a few it makes the spread look much more drastic even though it could just be a very small subset that makes the graph look like the spread is actually much greater. If you had thousands of samples for any other availability near 50 you might get a visualization that looks similar to availability=50",
          "votes": 1
        }
      ]
    },
    {
      "id": 1077899,
      "postDate": "2020-11-14T04:29:24.903Z",
      "content": "<p>\"ensembling\" that works. i implemented it and it has some improvement</p>\n<pre><code>1. train 2 models:\nmodel1(x) : x --&gt;embed1(x)--&gt;head1(x)\nmodel2(x) : x --&gt;embed2(x)--&gt;head2(x)\n\n2. measure loss on validation set.\n- we plot loss1 vs loss2 to visually inspect correlation and diversity\n- we compute upper bound for improvement: best loss = average { min(loss1(x),loss2(x)) } \n\n3. make ensemble model\nmode: x --&gt;embed1(x),embed2(x) --&gt; concat --&gt; head(concat)\n\n- we freeze embed1,embed2 to prevent overfitting. you should see some improvement\n</code></pre>",
      "rawMarkdown": "\"ensembling\" that works. i implemented it and it has some improvement\n\n```\n1. train 2 models:\nmodel1(x) : x -->embed1(x)-->head1(x)\nmodel2(x) : x -->embed2(x)-->head2(x)\n\n2. measure loss on validation set.\n- we plot loss1 vs loss2 to visually inspect correlation and diversity\n- we compute upper bound for improvement: best loss = average { min(loss1(x),loss2(x)) } \n\n3. make ensemble model\nmode: x -->embed1(x),embed2(x) --> concat --> head(concat)\n\n- we freeze embed1,embed2 to prevent overfitting. you should see some improvement\n\n```\n"
    },
    {
      "id": 1076909,
      "postDate": "2020-11-13T03:04:44.093Z",
      "content": "<p>What I am curious about is how do mode7 make inference? Take the top 3 confidence?</p>",
      "rawMarkdown": "What I am curious about is how do mode7 make inference? Take the top 3 confidence?",
      "replies": [
        {
          "id": 1076928,
          "postDate": "2020-11-13T03:37:25.307Z",
          "content": "<p>this is local validation. i use all 7 mode</p>",
          "rawMarkdown": "this is local validation. i use all 7 mode"
        },
        {
          "id": 1076957,
          "postDate": "2020-11-13T04:32:26.447Z",
          "content": "<p>What about \"total confidences should sum to 1\"? Simply increase the top 1 confidence and make the total sum to 1?</p>",
          "rawMarkdown": "What about \"total confidences should sum to 1\"? Simply increase the top 1 confidence and make the total sum to 1?"
        },
        {
          "id": 1076964,
          "postDate": "2020-11-13T04:44:47.037Z",
          "content": "<p>i use all 7 mode, meaning that I have 7 confidence values and 7 set of x,y coords.<br>\nthis will not be able to make a submission, but it is for analyzing results and data.</p>\n<p>the evaluation function is not affected as the function does not assume 3 mode.</p>\n<hr>\n<p>but it does show that if you can make more modes, e.g. via more prediction from one model or concate predictions from multiple models, you may be able to get very low metric score. </p>\n<p>but of course, you will have another problem of selecting the correct 3 for submission. --- that is the problem of ensmbling</p>",
          "rawMarkdown": "i use all 7 mode, meaning that I have 7 confidence values and 7 set of x,y coords.\nthis will not be able to make a submission, but it is for analyzing results and data.\n\nthe evaluation function is not affected as the function does not assume 3 mode.\n\n---\n\nbut it does show that if you can make more modes, e.g. via more prediction from one model or concate predictions from multiple models, you may be able to get very low metric score. \n\n\nbut of course, you will have another problem of selecting the correct 3 for submission. --- that is the problem of ensmbling\n"
        },
        {
          "id": 1077620,
          "postDate": "2020-11-13T18:34:57.503Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1072972,
      "postDate": "2020-11-09T00:47:29.280Z",
      "content": "<p>you might wan to check:</p>\n<p>\"Cycling rules in the US<br>\nThere are differences in the cycling laws across the states, so please click here to check the bike laws based on your destination.</p>\n<p>There are, however, a few general rules applicable wherever within the USA:</p>\n<p>In the United States, everyone must drive on the right-hand side of the roadway. Never ride your bike against the traffic flow.\"</p>",
      "rawMarkdown": "you might wan to check:\n\n\"Cycling rules in the US\nThere are differences in the cycling laws across the states, so please click here to check the bike laws based on your destination.\n\nThere are, however, a few general rules applicable wherever within the USA:\n\nIn the United States, everyone must drive on the right-hand side of the roadway. Never ride your bike against the traffic flow.\""
    },
    {
      "id": 1066996,
      "postDate": "2020-11-02T09:32:20.790Z",
      "content": "<p>i was reading the paper introduced by <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> </p>\n<p>Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks<br>\n<a href=\"https://arxiv.org/pdf/1809.10732.pdf\" target=\"_blank\">https://arxiv.org/pdf/1809.10732.pdf</a></p>\n<p>\"We collected 240 hours of data by manually driving SDV in<br>\nPittsburgh, PA and Phoenix, AZ in various traffic conditions<br>\n(e.g., varying times of day, days of the week). Traffic actors<br>\nwere tracked using Unscented Kalman filter (UKF) \"</p>\n<p>Now I need to read the document of l5kit to see how the data is collected. i don't think anyone would annotate the location of the agent frame by frame. there must be some kind of processing data.</p>",
      "rawMarkdown": "i was reading the paper introduced by @ryches \n\nMultimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks\nhttps://arxiv.org/pdf/1809.10732.pdf\n\n\n\"We collected 240 hours of data by manually driving SDV in\nPittsburgh, PA and Phoenix, AZ in various traffic conditions\n(e.g., varying times of day, days of the week). Traffic actors\nwere tracked using Unscented Kalman filter (UKF) \"\n\nNow I need to read the document of l5kit to see how the data is collected. i don't think anyone would annotate the location of the agent frame by frame. there must be some kind of processing data."
    },
    {
      "id": 1064337,
      "postDate": "2020-10-30T02:32:16.760Z",
      "content": "<p>Using classification loss instead of regression for post trajectory selection makes sense. covernet uses the similar concept.<br>\n<a href=\"https://arxiv.org/pdf/1911.10298.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.10298.pdf</a></p>\n<p>for classification loss they first sample sets of trajectories from train data and then select smallest set of fix trajectories using NP hard set cover problem.<br>\nBut from my experiments, nnl loss performs better than covernets classification loss for this data.  Two reasons  that I could think of are first different types of agent like pedestrians and cyclist in this data which has very jittery trajectories and secondly covernet loss is usually better for higher number of modes.</p>",
      "rawMarkdown": "Using classification loss instead of regression for post trajectory selection makes sense. covernet uses the similar concept.\nhttps://arxiv.org/pdf/1911.10298.pdf\n\nfor classification loss they first sample sets of trajectories from train data and then select smallest set of fix trajectories using NP hard set cover problem.\nBut from my experiments, nnl loss performs better than covernets classification loss for this data.  Two reasons  that I could think of are first different types of agent like pedestrians and cyclist in this data which has very jittery trajectories and secondly covernet loss is usually better for higher number of modes.",
      "replies": [
        {
          "id": 1064389,
          "postDate": "2020-10-30T04:28:05.757Z",
          "content": "<p>The problem I had with covernet is the trajectories they provided are configured for 6 seconds at a .5s resolution and the ones they provided were not the hybrid or dynamic which performed significantly better. It would be much better if they had provided their methods for generating these</p>",
          "rawMarkdown": "The problem I had with covernet is the trajectories they provided are configured for 6 seconds at a .5s resolution and the ones they provided were not the hybrid or dynamic which performed significantly better. It would be much better if they had provided their methods for generating these",
          "votes": 1
        },
        {
          "id": 1064811,
          "postDate": "2020-10-30T14:51:50.403Z",
          "content": "<p>Methods used to generate sets of trajectories are described in the article. I created hybrid and dynamic sets and set up the CoverNet pipeline. However, the main problem with this approach is that you have a lower boundary for NLL.  I did a test to answer the question \"What would be the nll value if CoverNet predicted everything correctly (unimodal)?\" - and given a fixed trajectory set (~1400 trajectories, epsilon 2.0) it was about 9 to 12 nll score. Of course, you cannot assume 100% accuracy, so it's probably not worth further investigation. </p>\n<p>Although, those trajectories sets are useful for debugging.</p>\n<p>And this is how fixed trajectories set (epsilon 2) looks like in Lyft competition:  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3611454%2F128398b61a49bbb0c49dcaa4ddb4c084%2Fdownload.png?generation=1604069404715931&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Methods used to generate sets of trajectories are described in the article. I created hybrid and dynamic sets and set up the CoverNet pipeline. However, the main problem with this approach is that you have a lower boundary for NLL.  I did a test to answer the question \"What would be the nll value if CoverNet predicted everything correctly (unimodal)?\" - and given a fixed trajectory set (~1400 trajectories, epsilon 2.0) it was about 9 to 12 nll score. Of course, you cannot assume 100% accuracy, so it's probably not worth further investigation. \n\nAlthough, those trajectories sets are useful for debugging.\n\nAnd this is how fixed trajectories set (epsilon 2) looks like in Lyft competition:  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3611454%2F128398b61a49bbb0c49dcaa4ddb4c084%2Fdownload.png?generation=1604069404715931&alt=media)\n\n ",
          "votes": 5
        },
        {
          "id": 1065297,
          "postDate": "2020-10-31T06:07:53.987Z",
          "content": "<p>l believe having the correct way to train the mixture expert is a key to get good results and win the competition. And since the mixture number is only 3 and small, there are many tricks to get better results. (what's more, the car only travels on road, or nearly on road)</p>\n<p>i am quite busy at the moment and I haven't tried this out yet. but this is what one can do:<br>\n1) say we divide trajectectory into type = left, front and right .<br>\n2) we train 3 predictors for each type using :</p>\n<p>individual predictor(type, image) --&gt; function( cat [ embed (image), type ] )</p>\n<p>this is the upper bound of our performance if we have guess the mode type correctly. Note that embedding weights are shared.</p>\n<p>3) then, we have<br>\nfinal_predictor (image) -&gt; type = function( …), image --&gt; individual predictor</p>\n<p>this reminds me of the early of face detection and face point localization where estimation of face orientation first and then apply specific face detector to specific orientation.</p>\n<p>no sure how it will work here or comparing to end-to-end methods. in theory end-to-end should out-performance. and 2 stage method maybe a good baseline.</p>",
          "rawMarkdown": "l believe having the correct way to train the mixture expert is a key to get good results and win the competition. And since the mixture number is only 3 and small, there are many tricks to get better results. (what's more, the car only travels on road, or nearly on road)\n\ni am quite busy at the moment and I haven't tried this out yet. but this is what one can do:\n1) say we divide trajectectory into type = left, front and right .\n2) we train 3 predictors for each type using :\n\nindividual predictor(type, image) --> function( cat [ embed (image), type ] )\n\nthis is the upper bound of our performance if we have guess the mode type correctly. Note that embedding weights are shared.\n\n3) then, we have\nfinal\\_predictor (image) -> type = function( ...), image --> individual predictor\n\nthis reminds me of the early of face detection and face point localization where estimation of face orientation first and then apply specific face detector to specific orientation.\n\nno sure how it will work here or comparing to end-to-end methods. in theory end-to-end should out-performance. and 2 stage method maybe a good baseline.",
          "votes": 4
        },
        {
          "id": 1065847,
          "postDate": "2020-10-31T22:27:18.307Z",
          "content": "<p>if we predict turn in the first step before trajectories don't you think it can lead to mode collapse and since no one can estimate the direction of agent accurately all the time, for some of them it can be terribly wrong? . I think MTP (<a href=\"https://arxiv.org/pdf/1809.10732.pdf\" target=\"_blank\">https://arxiv.org/pdf/1809.10732.pdf</a>) follows similar type of approach.<br>\nFirst they find best possible mode using classification loss and then backprop only for that mode to estimate trajectories. I ran into model collapse problem with it though, and not so good results.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4573165%2F0a27d166e533fa8c2a7962ead42a7d4f%2Fdownload.png?generation=1604183925157810&amp;alt=media\" alt=\"\"></p>\n<p>So in this case all three modes are going straight. Loss is 1674 for this sample.</p>",
          "rawMarkdown": "if we predict turn in the first step before trajectories don't you think it can lead to mode collapse and since no one can estimate the direction of agent accurately all the time, for some of them it can be terribly wrong? . I think MTP (https://arxiv.org/pdf/1809.10732.pdf) follows similar type of approach.\nFirst they find best possible mode using classification loss and then backprop only for that mode to estimate trajectories. I ran into model collapse problem with it though, and not so good results.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4573165%2F0a27d166e533fa8c2a7962ead42a7d4f%2Fdownload.png?generation=1604183925157810&alt=media)\n\nSo in this case all three modes are going straight. Loss is 1674 for this sample."
        },
        {
          "id": 1066387,
          "postDate": "2020-11-01T17:35:26.197Z",
          "content": "<p>if you can find a way to label the type of trajectory, then you can do ensembling.<br>\nit is just like object detection framework:</p>\n<p>classification  --&gt; trajectory selection<br>\nbox regression --&gt;estimation of the trajectory</p>\n<hr>\n<p>but  of course, you still can do ensembling without classification framework. In this case, assume you train 3 model using resnet34, efficient-b0, resnet-50. the predictions are:</p>\n<ul>\n<li>ca0, ca1, ca2, ta0, ta1, ta2  (c for confidence, t for predicted trajectory)</li>\n<li>cb0, cb1, cb2, tb0, tb1, tb2 </li>\n<li>cc0, cc1, cc2, tc0, tc1, tc2 </li>\n</ul>\n<p>you would have to cluster all the t's into groups. (think of this as non-max suppresion in object detection). you may end up with group1,2,3,4..5 (e.g. 5 groups)</p>\n<p>for each group, you need to think of a way to find the center trajectory<br>\nand also a way to compute the group score. And for grouping, you need to away to measure trajectory distance (just like iou in non-max suppression)</p>\n<hr>\n<p>there are papers that use transformer or train a network for non-max suppression to replace heuristics non-max-suppresion. these methods can be applied here too.</p>\n<pre><code>def decide_to_ensemble_net(trajectory1, trajectory2):\n     ...\n     it is like 2 layer 2 stacking net\n</code></pre>\n<hr>\n<p>if you can do ensembling, you can out-performance one single model. this opens up for the further application of pesudo-label, semi-supervised, weak-supervised methods.</p>",
          "rawMarkdown": "if you can find a way to label the type of trajectory, then you can do ensembling.\nit is just like object detection framework:\n\nclassification  --> trajectory selection\nbox regression -->estimation of the trajectory\n\n---\n\nbut  of course, you still can do ensembling without classification framework. In this case, assume you train 3 model using resnet34, efficient-b0, resnet-50. the predictions are:\n- ca0, ca1, ca2, ta0, ta1, ta2  (c for confidence, t for predicted trajectory)\n- cb0, cb1, cb2, tb0, tb1, tb2 \n- cc0, cc1, cc2, tc0, tc1, tc2 \n\nyou would have to cluster all the t's into groups. (think of this as non-max suppresion in object detection). you may end up with group1,2,3,4..5 (e.g. 5 groups)\n\nfor each group, you need to think of a way to find the center trajectory\nand also a way to compute the group score. And for grouping, you need to away to measure trajectory distance (just like iou in non-max suppression)\n\n\n---\n\nthere are papers that use transformer or train a network for non-max suppression to replace heuristics non-max-suppresion. these methods can be applied here too.\n\n```\n\ndef decide_to_ensemble_net(trajectory1, trajectory2):\n     ...\n     it is like 2 layer 2 stacking net\n```\n\n---\n\nif you can do ensembling, you can out-performance one single model. this opens up for the further application of pesudo-label, semi-supervised, weak-supervised methods.\n\n",
          "votes": 4
        },
        {
          "id": 1067227,
          "postDate": "2020-11-02T12:04:14.710Z",
          "content": "<p>coverNet: <a href=\"https://www.youtube.com/watch?v=fTM6ZLwmp10\" target=\"_blank\">https://www.youtube.com/watch?v=fTM6ZLwmp10</a><br>\nclassifiy trajectory</p>\n<p>TPNet: <a href=\"https://www.youtube.com/watch?v=Bfvy7qby1fg\" target=\"_blank\">https://www.youtube.com/watch?v=Bfvy7qby1fg</a><br>\npropose trajectory</p>",
          "rawMarkdown": "coverNet: https://www.youtube.com/watch?v=fTM6ZLwmp10\nclassifiy trajectory\n\nTPNet: https://www.youtube.com/watch?v=Bfvy7qby1fg\npropose trajectory"
        },
        {
          "id": 1070452,
          "postDate": "2020-11-05T19:43:23.130Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Great insights, wondering if you have any code reference for classification head and regression head that could be adapted to this ??</p>",
          "rawMarkdown": "@hengck23 Great insights, wondering if you have any code reference for classification head and regression head that could be adapted to this ??",
          "votes": -1
        },
        {
          "id": 1071328,
          "postDate": "2020-11-06T18:28:14.923Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1064055,
      "postDate": "2020-10-29T16:31:51.690Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/engck23\" target=\"_blank\">@engck23</a>, cool to see that you are jumping in!</p>\n<p>Yeah it's something I thought of. I was playing around with a model which has a prediction head and classification head, so in principle, you could predict 6 trajectories, and then only take the top 3 based on the classification score from the other head. </p>\n<p>Not sure if it will help, but it makes sense to me that there could be more then 3 viable routes, and the additional exploration might help. <br>\nDon't have any concrete results as to whether or not it actually helps. </p>",
      "rawMarkdown": "Hey @engck23, cool to see that you are jumping in!\n\nYeah it's something I thought of. I was playing around with a model which has a prediction head and classification head, so in principle, you could predict 6 trajectories, and then only take the top 3 based on the classification score from the other head. \n\nNot sure if it will help, but it makes sense to me that there could be more then 3 viable routes, and the additional exploration might help. \nDon't have any concrete results as to whether or not it actually helps. \n\n"
    }
  ],
  "comments": [
    {
      "id": 1072417,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-08T07:31:19.873000",
      "content": "<p>this is what you get if you plot the target on world coords.</p>\n<ul>\n<li>identify the agent type from meta data definitely helps (e.g. vehicle, pedestrian, cyclist … traveling at different speed)</li>\n<li>you can have average speed/yaw of vehicle on road as input (very strong prior)</li>\n</ul>\n<p>there could be some experiment design flaw here. The test data should be of a completely different road scene from the training dataset to avoid bias.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb8df48f983e647e2e913a98f13f7ca6f%2FSelection_209.png?generation=1604820518323524&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe05ba536a3afc41526f9167fadf877b9%2FSelection_208.png?generation=1604820537854691&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03cabdbc4e906b59cfc51e63c093365%2FSelection_207.png?generation=1604820558413844&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1072434,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-08T08:02:08.103000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8666565343532f56a65a79a324bf676c%2FSelection_218.png?generation=1604822508676863&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7d827e2b859272da2fef33dd56529976%2FSelection_219.png?generation=1604822525999439&amp;alt=media\" alt=\"\"><br>\ntypo error : there is no red …</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1072872,
          "author_name": "Benedikt Droste",
          "author_url": "",
          "post_date": "2020-11-08T19:37:55.223000",
          "content": "<p>Do you have a hint how to implement additional inputs? I guess you are talking about the model architecture itself. Do you have an example how to add non image data to a pretained resnet34?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1072888,
          "author_name": "Michał Szachniewicz",
          "author_url": "",
          "post_date": "2020-11-08T20:14:48.760000",
          "content": "<p>In the forward function, you can extract features from the backbone, then flatten it and concatenate with any additional features you want using torch.cat(). After that, you will have to adjust the number of in_features in your dense layer (just increase it by the length of your non-image input). </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1072896,
          "author_name": "Benedikt Droste",
          "author_url": "",
          "post_date": "2020-11-08T20:43:51.733000",
          "content": "<p>Thanks for the quick reply <a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> </p>\n<p>Do you have a code example? I thought about something like this:</p>\n<p>1# forward</p>\n<pre><code>    def forward(self, x, non-image-data):\n        x = self.backbone.conv1(x)\n        x = self.backbone.bn1(x)\n        x = self.backbone.relu(x)\n        x = self.backbone.maxpool(x)\n\n        x = self.backbone.layer1(x)\n        x = self.backbone.layer2(x)\n        x = self.backbone.layer3(x)\n        x = self.backbone.layer4(x)\n\n        x = self.backbone.avgpool(x)\n        x = torch.flatten(x, 1)\n        y = non-image-data\n        x = torch.cat((x, y), dim=1)\n\n        x = self.head(x)\n        x = self.logit(x)\n         [ ...... ]\n</code></pre>\n<p>2# init</p>\n<pre><code>def __init__(self, cfg: Dict, num_modes=3):\n       [ ...... ]\n        # You can add more layers here.\n        self.head = nn.Sequential(\n            # nn.Dropout(0.2),\n            nn.Linear(in_features=backbone_out_features+len(non-image-data), out_features=4096),\n        )\n        self.num_preds = num_targets * num_modes\n        self.num_modes = num_modes\n\n        self.logit = nn.Linear(4096, out_features=self.num_preds + num_modes)\n</code></pre>\n<p>What do you think?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073382,
          "author_name": "Michał Szachniewicz",
          "author_url": "",
          "post_date": "2020-11-09T13:25:13.987000",
          "content": "<p><a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> <br>\nSeems fine, I implemented it in a very similar if not the same way. Good job!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073631,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-09T18:12:09.943000",
          "content": "<p>here is the loss breakdown by label:</p>\n<pre><code>for non-chopped data:\nlocal cv loss: 9.346676  (lb = 15.034)\nbreakdown: \nPERCEPTION_LABEL_CAR         9.259042 \nPERCEPTION_LABEL_CYCLIST    17.221859 \nPERCEPTION_LABEL_PEDESTRIAN  4.827463\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1073880,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-10T02:22:08.400000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffc4872f9d3751a914d92103db8769fc0%2FSelection_063.png?generation=1604974926304836&amp;alt=media\" alt=\"\"></p>\n<p>i wonder are these ground truth senor errors?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd29aed8a457cdcd270ca91cd9562457b%2FSelection_067.png?generation=1604975711807483&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1074374,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-11-10T15:26:38.020000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a>. Did you feed \"history_positions\" and other data into the model? During training, is \"history_availabilities \" useful to imply the available agent in history frames?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1074375,
          "author_name": "Benedikt Droste",
          "author_url": "",
          "post_date": "2020-11-10T15:29:06.357000",
          "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> <br>\nIntegrating \"history_positions\" leads to a huge overfit and a big derivation between loss and lb score to me.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1074776,
              "author_name": "Yannik",
              "author_url": "",
              "post_date": "2020-11-11T03:37:10.663000",
              "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> <br>\n  Integrating \"history_positions\" leads to a huge overfit and a big derivation between loss and lb score to me.</p>\n</blockquote>\n<p>Hi, <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a>. I don't think this huge deviation is caused by overfitting. According to my experiment, there is a data mismatch that is caused by the differences between train.zarr and validation.zarr(or test.zarr), which can lead to this huge deviation. Because the \"history_availabilities\" in train.zarr are always 1, which means agents exist in every frame in the training sample that they belong to. However, this is totally different in the chopped validation set or the testing set. So I think this might be the main reason why we have such a huge deviation when we are adding \"history_positions\".</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1074389,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-11-10T15:44:36.343000",
          "content": "<p><a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> In my case, there is also a huge deviation between validation score and training loss. But I can not certainly say that it is overfitting. I also think that it is also possible that the differences of \"history_availabilities\" between the training set and the chopped validation set to cause this huge deviation. But I can not verify it. As mention <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> also meet the same problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1074393,
          "author_name": "Michał Szachniewicz",
          "author_url": "",
          "post_date": "2020-11-10T15:50:07.757000",
          "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> No, I feed the agent state vector as described in this article: <a href=\"https://arxiv.org/abs/1809.10732\" target=\"_blank\">https://arxiv.org/abs/1809.10732</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1074406,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-11-10T16:06:30.183000",
          "content": "<p><a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> But we don't have the velocity, acceleration mentioned in the article. How you implement this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1074459,
          "author_name": "Michał Szachniewicz",
          "author_url": "",
          "post_date": "2020-11-10T17:07:33.470000",
          "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> Velocity is available in zarr data, but you can get everything from data returned by AgentDataset - you have history_positions and history_yaws in time. Features like velocity, acceleration, and yaw change rate can be calculated using basic physics knowledge.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1074777,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-11-11T03:38:06.713000",
          "content": "<p><a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> Thanks for your views. I will try to do some experiments on that later!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1075234,
          "author_name": "Pallavi Ramicetty",
          "author_url": "",
          "post_date": "2020-11-11T14:09:26.740000",
          "content": "<p><a href=\"https://www.kaggle.com/szacho\" target=\"_blank\">@szacho</a> As you mentioned we can get the velocity from AgentDataset. But what about the agent's <strong>label_probabilities</strong>? seems <strong>agent_samples.py</strong> won't return them. If possible, can suggest some method to get them along with the \"data\" object?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1075254,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-11T14:19:33.403000",
          "content": "<p>for the file: <br>\n/l5kit/sampling/agent_sampling.py</p>\n<pre><code>def generate_agent_sample(...):\n     ...\n        \"extent\": agent_extent_m,\n        #\"agent\": selected_agent, # &lt;add here&gt;\n        \"label_probabilities\": selected_agent[\"label_probabilities\"], # &lt;add here&gt;\n</code></pre>\n<p>for the file: <br>\n/l5kit/dataset/ego.py</p>\n<pre><code>class EgoDataset(Dataset):\n...\ndef get_frame(...):\n\n            \"extent\": data[\"extent\"],\n            #\"agent\": data[\"agent\"], #&lt;add here&gt;\n            \"label_probabilities\": data[\"label_probabilities\"], #&lt;add here&gt;\n        }\n</code></pre>",
          "votes": 5,
          "replies": [
            {
              "id": 1075771,
              "author_name": "SuryaJR_Rafl",
              "author_url": "",
              "post_date": "2020-11-12T00:21:30.830000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , Thanks for your insights. How to add this custom l5kit library to one's kaggle notebook? Instead I thought of calculating velocities in rudimentary way like this (I guess we have to calculate yaw rate and acceleration ourselves anyway):</p>\n<pre><code>    if data['history_availabilities'].to(device)shape[1] &gt; 1: \n        velocities = torch.sqrt(torch.sum((pose_data[:,1:,:] - pose_data[:,0:-1, :]) ** 2, 2)) / dt\n        avg_velocity = torch.mean(velocities, dim=1).view(bs,1)\n        #acceleration = torch.abs(velocities[:, 1:] - velocities[:, 0:-1])\n        avg_acceleration = torch.mean(torch.abs(velocities[:, 1:] - velocities[:, 0:-1]) / dt, dim=1).view(bs,1)\n        yaw_rate = torch.mean(torch.abs(yaw_data[:, 1:] - yaw_data[:, 0:-1]) /  dt, dim=1).view(bs,1)\n\n        avg_velocity = avg_velocity / MAX_VELOCITY # 20 m/s\n        avg_acceleration = avg_acceleration / MAX_ACCELERATION # 2m/s^2\n        yaw_rate = yaw_rate / MAX_YAW_RATE # 45 deg/s\n\n    else:\n        avg_velocity = 0.0;   avg_acceleration = 0.0;    yaw_rate = 0.0; \n</code></pre>\n<p>Is there anything I'm missing here? Any any exception handling required?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1082506,
              "author_name": "Saurabh7",
              "author_url": "",
              "post_date": "2020-11-18T00:59:28.477000",
              "content": "<p>Adding these state inputs leads to large overfitting for me, anyone else seeing similar results ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1075328,
          "author_name": "Pallavi Ramicetty",
          "author_url": "",
          "post_date": "2020-11-11T15:35:08.927000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Oh you are using l5kit/ as part of your code. I am using it as a library. Sure I will do require changes. Thanks for your quick response.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078444,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2020-11-14T19:50:00.310000",
          "content": "<p>Speaking of velocities, agent_sampling.py and ego.py have been updated. You can now get velocity (and speed) through the <code>AgentDataset()</code> class.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1078452,
          "author_name": "Saurabh7",
          "author_url": "",
          "post_date": "2020-11-14T20:04:05.507000",
          "content": "<p>Where can I find the updates ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078462,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2020-11-14T20:17:03.597000",
          "content": "<p>l5kit github repository, <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">here</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1076901,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-13T02:52:24.260000",
      "content": "<p>surprise, surprise, surprise …<br>\nI put the conclusion first and then the observations.</p>\n<p>conclusion:</p>\n<ul>\n<li>there are two types of trajectories: </li>\n</ul>\n<ol>\n<li>parametric curves, that obey kinematics. you can get very accurate results via curve fitting (and maybe using lstm of past history points)</li>\n<li>non-parametric and irregular curves</li>\n</ol>\n<p>we probably want to detect, predict, post or pre-processing them separately. they are easy to identify via past histories.</p>\n<hr>\n<p>observation:</p>\n<ul>\n<li>I study the metric and note that the effect of l2 loss is very much greater than the confidence values (because world coords is in meters?). i.e. in a prediction of mode M, as long as one of them has low l2 loss (even if the confidence is low) the kaggle metric will be low.</li>\n</ul>\n<p>see below simulation:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcadc4dbd60a83b50138a2a7daaf45c3b%2FSelection_149.png?generation=1605235825678435&amp;alt=media\" alt=\"\"></p>\n<p>hence I trained a mode-7 predictor from my current mode-3 predictor. i just freeze the backbone weights and change the head. it is very fast to train. on chopped validation data, kaggle metric reduces from 15+ to 11+ within hours.</p>\n<p>but when I inspect the results,  I was surprised</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1076902,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-13T02:53:07.980000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F28b6cea52808ab6ba5a725609a2e1da3%2FSelection_144.png?generation=1605235977467342&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3057c4f9087c723954509a2d25e0a2ce%2FSelection_157.png?generation=1605261672198016&amp;alt=media\" alt=\"\"><br>\nNote: if there are shakeup, it can be due to just a few outliers</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1076905,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-13T02:55:36.757000",
          "content": "<p>so, what are the cases, mode-7 performs better?<br>\ntruth: disconnected black dots<br>\n(history is not shown)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F490491b0c3190dc69ef6a15d00b88b86%2FSelection_145.png?generation=1605236075073834&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F26cb2e998011236e99e4b3de568a11f2%2FSelection_146.png?generation=1605236092430499&amp;alt=media\" alt=\"\"></p>\n<p>you can see that we can gain a lot if we can estimate the \"straight trajectory\" more accurate … which seems to be an easy task </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1077784,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-11-13T23:19:28",
          "content": "<p>Nice catch but the competition metric is biased towards long trajectories. Sigma is fixed to 1 which is not real! It should increase as the predictions goes further in time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078686,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-15T07:12:14.137000",
          "content": "<p>as an example<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F004de876653e436c02f85469f7911ca4%2FSelection_185.png?generation=1605424264162784&amp;alt=media\" alt=\"\"></p>\n<p>you can add other labels for specific post-processing, e.g. di-biasing (the kaggle multiplier trick in regression problem, which is essentially manually adjusting the std of error)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1078739,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-11-15T08:49:06.343000",
          "content": "<p>How do you set the <code>is_parametric</code> flag? It's not clear to me. Do you use a custom curve fitting function?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078773,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-15T09:36:13.760000",
          "content": "<p>e.g. for each train sample, you have x =[history, future]</p>\n<p>then you devise a curve function, predict_future = f(history). you will have to design this by hand, making some assumptions. you can google for some kinematics model.</p>\n<p>now each train sample will have x =[history, f(history), future].</p>\n<p>compute error of f(history), e.g. l2 loss (f(history), future). now if this error is small, you can label is as True (i.e. you can recover future values from f(history))</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1078778,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-11-15T09:44:17.033000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, this helps :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1072969,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-09T00:44:23.703000",
      "content": "<p>there is much more information in the zarr file (instead of the AgentDataset)</p>\n<p>e.g. you get velocity and label probability</p>\n<pre><code>AGENT_DTYPE = [\n    (\"centroid\", np.float64, (2,)),\n    (\"extent\", np.float32, (3,)),\n    (\"yaw\", np.float32),\n    (\"velocity\", np.float32, (2,)),\n    (\"track_id\", np.uint64),\n    (\"label_probabilities\", np.float32, (len(LABELS),)),\n]\n</code></pre>\n<p>here is an example of reading the raw zarr file. note that you can recover the whole track!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa3d90739319a1ec4bea51004002f61d3%2FSelection_230.png?generation=1604882661363275&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1069034,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-04T02:08:55.307000",
      "content": "<p>you can submit the top-1, top-2 and top-3 results to prob the correctness of your prediction. then it may be possible to recalibrate  or adjust the temperate of the confidence.</p>\n<p>some of the  predictions are clearly only single-mode, e.g. moving in a straight lane or not moving.  in that case a single high confidence trajectory prediction may be better?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1067368,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-02T13:52:56.493000",
      "content": "<p>self-supervised?</p>\n<p>for the test data, say we ask for number of history frame = 30.<br>\nwe use t-30 to t-20 as input and train an embedding to predict t-20 to t.</p>\n<p>would it work?</p>\n<p>how about :<br>\nfuture = predict (history) and then history1 = predict(future) … like to and from lanuage translation …<br>\nwould that be a kind of augmentation (i.e. history --&gt;history1)?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1066792,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-02T03:34:06.673000",
      "content": "<p>MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction<br>\n<a href=\"https://openaccess.thecvf.com/content_CVPR_2020/papers/Marchetti_MANTRA_Memory_Augmented_Networks_for_Multiple_Trajectory_Prediction_CVPR_2020_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPR_2020/papers/Marchetti_MANTRA_Memory_Augmented_Networks_for_Multiple_Trajectory_Prediction_CVPR_2020_paper.pdf</a></p>\n<p><img src=\"https://images.deepai.org/converted-papers/2006.03340/x3.png\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1064120,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-10-29T18:29:54.943000",
      "content": "<p>In the original paper of the base model that lyft provided to us it seems to show that 3 is empirically the most performant number of modes <a href=\"https://arxiv.org/pdf/1809.10732.pdf\" target=\"_blank\">https://arxiv.org/pdf/1809.10732.pdf</a></p>\n<p>1, 2, and 4 modes do worse. This has been reproduced by a few other papers as well. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1064288,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-29T23:49:07.947000",
          "content": "<p>thanks for the paper.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc0e2508c364917765fd2d65e8436fcc8%2FSelection_030.png?generation=1604014930131979&amp;alt=media\" alt=\"\"></p>\n<p>it seems that it is due to the nature road junction. in that case, we can just train for going straight, turn left and turn right. i need to check the distribution/frequency of turns of junctions. thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1065049,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-30T19:38:03.467000",
          "content": "<p>Intresting, thank you for the paper.</p>\n<p>The relations are greatly explained and visualized.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1065578,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-31T13:13:12.893000",
          "content": "<p>Does anyone know if the mentioned MTP loss is somewhere available as public code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1065608,
          "author_name": "Michał Szachniewicz",
          "author_url": "",
          "post_date": "2020-10-31T13:58:34.417000",
          "content": "<p>Here, <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/prediction/models/mtp.py\" target=\"_blank\">https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/prediction/models/mtp.py</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1065732,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-31T16:56:38.637000",
          "content": "<p>Thank you :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1076954,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-13T04:28:13.600000",
      "content": "<p>to disable image rasterizer and just do prediction using history information:</p>\n<pre><code>    dm = LocalDataManager()\n    rasterizer = None #build_rasterizer(cfg, dm)\n\n    zarr = ChunkedDataset( dm.require('scenes/train.zarr')).open()\n    dataset = MyAgentDataset(cfg, zarr, rasterizer) \n\n\n\nin ego.py: (around line 99)\n\nclass EgoDataset(Dataset):\n\ndef get_frame(self, scene_index ...):\n\n        # 0,1,C -&gt; C,0,1\n        # &lt;hck&gt;\n        if data[\"image\"] is not None:\n            image = data[\"image\"].transpose(2, 0, 1)\n        else:\n            image = 0 #dummy\n</code></pre>\n<p>i am thinking of using:</p>\n<ul>\n<li>image : e.g. up to past 5 history frames</li>\n<li>locations : e.g. up to past 100 history frames</li>\n</ul>\n<p>so i need to create 2 dataset, one with rasterizer and one without (to speed up cpu data loading).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1077782,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-11-13T23:15:21.343000",
          "content": "<p>I reported this as a bug long time ago on github and I guess it is now fixed</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1077851,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-11-14T02:13:08.017000",
          "content": "<p>You mean that by doing disabling the rasterizer, we can use history information for improvement regardless of the differences of history_availabilities between train.zarr and test.zarr?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1077896,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-14T04:22:54.570000",
          "content": "<p>rendering long history frames is very slow. hence I want to experiment using history positions only without image. (or few history images)</p>\n<p>this is non- fixed length sequence learning, e.g</p>\n<pre><code>input:\nx(t), x(t-1),x(t-2) ... &lt;end&gt;\n</code></pre>\n<p>you can treat it as LSTM seq with an \"end\" token. alternatively you can pad to equal length.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078104,
          "author_name": "Yosshi999",
          "author_url": "",
          "post_date": "2020-11-14T11:21:04.800000",
          "content": "<p>You can use <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/rasterization/stub_rasterizer.py\" target=\"_blank\">StubRasterizer</a> to render only blank images</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1065510,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-10-31T11:17:03.593000",
      "content": "<p>i wonder did anyone plot and visualise the predicted trajectory?</p>\n<p>i don't think it would be a perfect parametric curve. it might look a little \"wiggling\" and not smooth. So there are be some post-processing required, or better still, should we  predict the parameters of the curve of trajectory instead?</p>\n<p>history trajectory can also be represented as a parametric curve</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1065524,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "2020-10-31T11:33:06.847000",
          "content": "<p>There has been some discussion about maybe using a Kalman filter to smooth out predictions, although at the moment my predicted trajectories tend to be pretty smooth:  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Fe89a6a4e6fe1239bf3aa224d2fe83537%2Finbox_314368_2ae6a5e2632e97085e567e4951c51536_mask.png?generation=1604143906678498&amp;alt=media\" alt=\"\"></p>\n<p>The targets often have some amount of noise in them, so being smoother than the targets might not give that much improvement</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1066394,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-01T17:54:15.257000",
          "content": "<p>seems that predicting the end-point correctly is very important. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1077979,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-14T07:13:11.497000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3093b436d5ee8a87c7594c04ed1fb923%2FSelection_164.png?generation=1605337987662608&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1078092,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-11-14T10:40:49.500000",
          "content": "<p>I believe, that picture is rather misleading with different counts in each bin. You may subsample the 50 avails part to match the count for the others, or plot e.g. violin graphs to get a better view on the loss distribution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078093,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-14T10:50:16.457000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F84f497214505a4fcaaac8589277f494c%2FSelection_165.png?generation=1605350990328844&amp;alt=media\" alt=\"\"></p>\n<p>count of samples vs samples' target availability</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1078110,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-11-14T11:31:02.323000",
          "content": "<p>I know, that's why your first Image gives a misleading view on the loss distribution. It looks like the loss distribution for 50 avails is much broader than for example 49 avails, but in reality they are very similar.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078120,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-14T11:43:59.517000",
          "content": "<p>i think the loss for e.g. availability=10 would be small because only \"10 nearby points are counted\" in the evaluation metric.</p>\n<p>in the case where availability=50, errors can get accumulated for many predicted points. hence the loss will be higher.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078142,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-11-14T12:17:25.673000",
          "content": "<p>of course that is the case, you can see it in the l5kit <code>time_displace</code> error, which is exactly evaluating this. But naturally, 49 and 50 are similar, which is not what your first plot suggests, that's all I am saying. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1078330,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-11-14T16:32:17.310000",
          "content": "<p>Yes, the plot is rather misleading because the frequency of availability=50 makes it difficult to evaluate  the distribution of losses. Because you are viewing thousands of samples rather than just a few it makes the spread look much more drastic even though it could just be a very small subset that makes the graph look like the spread is actually much greater. If you had thousands of samples for any other availability near 50 you might get a visualization that looks similar to availability=50</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1077899,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-14T04:29:24.903000",
      "content": "<p>\"ensembling\" that works. i implemented it and it has some improvement</p>\n<pre><code>1. train 2 models:\nmodel1(x) : x --&gt;embed1(x)--&gt;head1(x)\nmodel2(x) : x --&gt;embed2(x)--&gt;head2(x)\n\n2. measure loss on validation set.\n- we plot loss1 vs loss2 to visually inspect correlation and diversity\n- we compute upper bound for improvement: best loss = average { min(loss1(x),loss2(x)) } \n\n3. make ensemble model\nmode: x --&gt;embed1(x),embed2(x) --&gt; concat --&gt; head(concat)\n\n- we freeze embed1,embed2 to prevent overfitting. you should see some improvement\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1076909,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2020-11-13T03:04:44.093000",
      "content": "<p>What I am curious about is how do mode7 make inference? Take the top 3 confidence?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1076928,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-13T03:37:25.307000",
          "content": "<p>this is local validation. i use all 7 mode</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1076957,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-11-13T04:32:26.447000",
          "content": "<p>What about \"total confidences should sum to 1\"? Simply increase the top 1 confidence and make the total sum to 1?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1076964,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-13T04:44:47.037000",
          "content": "<p>i use all 7 mode, meaning that I have 7 confidence values and 7 set of x,y coords.<br>\nthis will not be able to make a submission, but it is for analyzing results and data.</p>\n<p>the evaluation function is not affected as the function does not assume 3 mode.</p>\n<hr>\n<p>but it does show that if you can make more modes, e.g. via more prediction from one model or concate predictions from multiple models, you may be able to get very low metric score. </p>\n<p>but of course, you will have another problem of selecting the correct 3 for submission. --- that is the problem of ensmbling</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1077620,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-13T18:34:57.503000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1072972,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-09T00:47:29.280000",
      "content": "<p>you might wan to check:</p>\n<p>\"Cycling rules in the US<br>\nThere are differences in the cycling laws across the states, so please click here to check the bike laws based on your destination.</p>\n<p>There are, however, a few general rules applicable wherever within the USA:</p>\n<p>In the United States, everyone must drive on the right-hand side of the roadway. Never ride your bike against the traffic flow.\"</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1066996,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-02T09:32:20.790000",
      "content": "<p>i was reading the paper introduced by <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> </p>\n<p>Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks<br>\n<a href=\"https://arxiv.org/pdf/1809.10732.pdf\" target=\"_blank\">https://arxiv.org/pdf/1809.10732.pdf</a></p>\n<p>\"We collected 240 hours of data by manually driving SDV in<br>\nPittsburgh, PA and Phoenix, AZ in various traffic conditions<br>\n(e.g., varying times of day, days of the week). Traffic actors<br>\nwere tracked using Unscented Kalman filter (UKF) \"</p>\n<p>Now I need to read the document of l5kit to see how the data is collected. i don't think anyone would annotate the location of the agent frame by frame. there must be some kind of processing data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1064337,
      "author_name": "SujaydKhandekar",
      "author_url": "",
      "post_date": "2020-10-30T02:32:16.760000",
      "content": "<p>Using classification loss instead of regression for post trajectory selection makes sense. covernet uses the similar concept.<br>\n<a href=\"https://arxiv.org/pdf/1911.10298.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.10298.pdf</a></p>\n<p>for classification loss they first sample sets of trajectories from train data and then select smallest set of fix trajectories using NP hard set cover problem.<br>\nBut from my experiments, nnl loss performs better than covernets classification loss for this data.  Two reasons  that I could think of are first different types of agent like pedestrians and cyclist in this data which has very jittery trajectories and secondly covernet loss is usually better for higher number of modes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1064389,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-10-30T04:28:05.757000",
          "content": "<p>The problem I had with covernet is the trajectories they provided are configured for 6 seconds at a .5s resolution and the ones they provided were not the hybrid or dynamic which performed significantly better. It would be much better if they had provided their methods for generating these</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1064811,
          "author_name": "Michał Szachniewicz",
          "author_url": "",
          "post_date": "2020-10-30T14:51:50.403000",
          "content": "<p>Methods used to generate sets of trajectories are described in the article. I created hybrid and dynamic sets and set up the CoverNet pipeline. However, the main problem with this approach is that you have a lower boundary for NLL.  I did a test to answer the question \"What would be the nll value if CoverNet predicted everything correctly (unimodal)?\" - and given a fixed trajectory set (~1400 trajectories, epsilon 2.0) it was about 9 to 12 nll score. Of course, you cannot assume 100% accuracy, so it's probably not worth further investigation. </p>\n<p>Although, those trajectories sets are useful for debugging.</p>\n<p>And this is how fixed trajectories set (epsilon 2) looks like in Lyft competition:  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3611454%2F128398b61a49bbb0c49dcaa4ddb4c084%2Fdownload.png?generation=1604069404715931&amp;alt=media\" alt=\"\"></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1065297,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-31T06:07:53.987000",
          "content": "<p>l believe having the correct way to train the mixture expert is a key to get good results and win the competition. And since the mixture number is only 3 and small, there are many tricks to get better results. (what's more, the car only travels on road, or nearly on road)</p>\n<p>i am quite busy at the moment and I haven't tried this out yet. but this is what one can do:<br>\n1) say we divide trajectectory into type = left, front and right .<br>\n2) we train 3 predictors for each type using :</p>\n<p>individual predictor(type, image) --&gt; function( cat [ embed (image), type ] )</p>\n<p>this is the upper bound of our performance if we have guess the mode type correctly. Note that embedding weights are shared.</p>\n<p>3) then, we have<br>\nfinal_predictor (image) -&gt; type = function( …), image --&gt; individual predictor</p>\n<p>this reminds me of the early of face detection and face point localization where estimation of face orientation first and then apply specific face detector to specific orientation.</p>\n<p>no sure how it will work here or comparing to end-to-end methods. in theory end-to-end should out-performance. and 2 stage method maybe a good baseline.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1065847,
          "author_name": "SujaydKhandekar",
          "author_url": "",
          "post_date": "2020-10-31T22:27:18.307000",
          "content": "<p>if we predict turn in the first step before trajectories don't you think it can lead to mode collapse and since no one can estimate the direction of agent accurately all the time, for some of them it can be terribly wrong? . I think MTP (<a href=\"https://arxiv.org/pdf/1809.10732.pdf\" target=\"_blank\">https://arxiv.org/pdf/1809.10732.pdf</a>) follows similar type of approach.<br>\nFirst they find best possible mode using classification loss and then backprop only for that mode to estimate trajectories. I ran into model collapse problem with it though, and not so good results.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4573165%2F0a27d166e533fa8c2a7962ead42a7d4f%2Fdownload.png?generation=1604183925157810&amp;alt=media\" alt=\"\"></p>\n<p>So in this case all three modes are going straight. Loss is 1674 for this sample.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1066387,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-01T17:35:26.197000",
          "content": "<p>if you can find a way to label the type of trajectory, then you can do ensembling.<br>\nit is just like object detection framework:</p>\n<p>classification  --&gt; trajectory selection<br>\nbox regression --&gt;estimation of the trajectory</p>\n<hr>\n<p>but  of course, you still can do ensembling without classification framework. In this case, assume you train 3 model using resnet34, efficient-b0, resnet-50. the predictions are:</p>\n<ul>\n<li>ca0, ca1, ca2, ta0, ta1, ta2  (c for confidence, t for predicted trajectory)</li>\n<li>cb0, cb1, cb2, tb0, tb1, tb2 </li>\n<li>cc0, cc1, cc2, tc0, tc1, tc2 </li>\n</ul>\n<p>you would have to cluster all the t's into groups. (think of this as non-max suppresion in object detection). you may end up with group1,2,3,4..5 (e.g. 5 groups)</p>\n<p>for each group, you need to think of a way to find the center trajectory<br>\nand also a way to compute the group score. And for grouping, you need to away to measure trajectory distance (just like iou in non-max suppression)</p>\n<hr>\n<p>there are papers that use transformer or train a network for non-max suppression to replace heuristics non-max-suppresion. these methods can be applied here too.</p>\n<pre><code>def decide_to_ensemble_net(trajectory1, trajectory2):\n     ...\n     it is like 2 layer 2 stacking net\n</code></pre>\n<hr>\n<p>if you can do ensembling, you can out-performance one single model. this opens up for the further application of pesudo-label, semi-supervised, weak-supervised methods.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1067227,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-02T12:04:14.710000",
          "content": "<p>coverNet: <a href=\"https://www.youtube.com/watch?v=fTM6ZLwmp10\" target=\"_blank\">https://www.youtube.com/watch?v=fTM6ZLwmp10</a><br>\nclassifiy trajectory</p>\n<p>TPNet: <a href=\"https://www.youtube.com/watch?v=Bfvy7qby1fg\" target=\"_blank\">https://www.youtube.com/watch?v=Bfvy7qby1fg</a><br>\npropose trajectory</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1070452,
          "author_name": "Saurabh7",
          "author_url": "",
          "post_date": "2020-11-05T19:43:23.130000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Great insights, wondering if you have any code reference for classification head and regression head that could be adapted to this ??</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1071328,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-06T18:28:14.923000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1064055,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "2020-10-29T16:31:51.690000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/engck23\" target=\"_blank\">@engck23</a>, cool to see that you are jumping in!</p>\n<p>Yeah it's something I thought of. I was playing around with a model which has a prediction head and classification head, so in principle, you could predict 6 trajectories, and then only take the top 3 based on the classification score from the other head. </p>\n<p>Not sure if it will help, but it makes sense to me that there could be more then 3 viable routes, and the additional exploration might help. <br>\nDon't have any concrete results as to whether or not it actually helps. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1063956": "first, why is there multi-mode?\n- e.g. car motion is different if traffic light condition is different ( green lights or red lights )\n- e.g. intention of the driver (maybe he wants to go destination A instead of B)\n\nso is three mode enough?\n- for each similar map (e.g. via clustering), you can cluster the train trajectory. how many clusters can you find?\n- you can actually have K modes (i.e. K subclass linear classifier) and treat this as a selecting top-3 problem\n\nI see that the public notebook are setting K=3, but K can be greater than 3. since \"regression loss\" seek for least square you will end up in a solution that is equal distance to each of the non-matching K cluster. If you have a matching K cluster, your regression loss will be low.\n\n \n\nrelated: top-K ranking loss, classifiation with subclass, trajectory clustering\ne.g. : https://arxiv.org/abs/1909.05235 - subclass softmax\n'Multiple Centers Now, we assume that each class has K centers. Then, the similarity between the example x ....'",
    "1072417": "this is what you get if you plot the target on world coords.\n- identify the agent type from meta data definitely helps (e.g. vehicle, pedestrian, cyclist ... traveling at different speed)\n- you can have average speed/yaw of vehicle on road as input (very strong prior)\n\nthere could be some experiment design flaw here. The test data should be of a completely different road scene from the training dataset to avoid bias.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb8df48f983e647e2e913a98f13f7ca6f%2FSelection_209.png?generation=1604820518323524&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe05ba536a3afc41526f9167fadf877b9%2FSelection_208.png?generation=1604820537854691&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd03cabdbc4e906b59cfc51e63c093365%2FSelection_207.png?generation=1604820558413844&alt=media)",
    "1076901": "surprise, surprise, surprise ...\nI put the conclusion first and then the observations.\n\nconclusion:\n- there are two types of trajectories: \n1. parametric curves, that obey kinematics. you can get very accurate results via curve fitting (and maybe using lstm of past history points)\n2. non-parametric and irregular curves\n\nwe probably want to detect, predict, post or pre-processing them separately. they are easy to identify via past histories.\n\n--- \n\nobservation:\n- I study the metric and note that the effect of l2 loss is very much greater than the confidence values (because world coords is in meters?). i.e. in a prediction of mode M, as long as one of them has low l2 loss (even if the confidence is low) the kaggle metric will be low.\n\nsee below simulation:\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcadc4dbd60a83b50138a2a7daaf45c3b%2FSelection_149.png?generation=1605235825678435&alt=media)\n\nhence I trained a mode-7 predictor from my current mode-3 predictor. i just freeze the backbone weights and change the head. it is very fast to train. on chopped validation data, kaggle metric reduces from 15+ to 11+ within hours.\n\nbut when I inspect the results,  I was surprised",
    "1072969": "there is much more information in the zarr file (instead of the AgentDataset)\n\ne.g. you get velocity and label probability\n\n```\nAGENT_DTYPE = [\n    (\"centroid\", np.float64, (2,)),\n    (\"extent\", np.float32, (3,)),\n    (\"yaw\", np.float32),\n    (\"velocity\", np.float32, (2,)),\n    (\"track_id\", np.uint64),\n    (\"label_probabilities\", np.float32, (len(LABELS),)),\n]\n\n\n```\n\nhere is an example of reading the raw zarr file. note that you can recover the whole track!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa3d90739319a1ec4bea51004002f61d3%2FSelection_230.png?generation=1604882661363275&alt=media)",
    "1069034": "you can submit the top-1, top-2 and top-3 results to prob the correctness of your prediction. then it may be possible to recalibrate  or adjust the temperate of the confidence.\n\nsome of the  predictions are clearly only single-mode, e.g. moving in a straight lane or not moving.  in that case a single high confidence trajectory prediction may be better?",
    "1067368": "self-supervised?\n\nfor the test data, say we ask for number of history frame = 30.\nwe use t-30 to t-20 as input and train an embedding to predict t-20 to t.\n\nwould it work?\n\nhow about :\nfuture = predict (history) and then history1 = predict(future) ... like to and from lanuage translation ...\nwould that be a kind of augmentation (i.e. history -->history1)?",
    "1066792": "MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction\nhttps://openaccess.thecvf.com/content_CVPR_2020/papers/Marchetti_MANTRA_Memory_Augmented_Networks_for_Multiple_Trajectory_Prediction_CVPR_2020_paper.pdf\n\n\n![](https://images.deepai.org/converted-papers/2006.03340/x3.png)",
    "1064120": "In the original paper of the base model that lyft provided to us it seems to show that 3 is empirically the most performant number of modes https://arxiv.org/pdf/1809.10732.pdf\n\n1, 2, and 4 modes do worse. This has been reproduced by a few other papers as well. ",
    "1076954": "to disable image rasterizer and just do prediction using history information:\n\n```\n    dm = LocalDataManager()\n    rasterizer = None #build_rasterizer(cfg, dm)\n\n    zarr = ChunkedDataset( dm.require('scenes/train.zarr')).open()\n    dataset = MyAgentDataset(cfg, zarr, rasterizer) \n\n\n\nin ego.py: (around line 99)\n\nclass EgoDataset(Dataset):\n\ndef get_frame(self, scene_index ...):\n\n        # 0,1,C -> C,0,1\n        # <hck>\n        if data[\"image\"] is not None:\n            image = data[\"image\"].transpose(2, 0, 1)\n        else:\n            image = 0 #dummy\n\n```\n\ni am thinking of using:\n- image : e.g. up to past 5 history frames\n- locations : e.g. up to past 100 history frames\n\nso i need to create 2 dataset, one with rasterizer and one without (to speed up cpu data loading).\n ",
    "1065510": "i wonder did anyone plot and visualise the predicted trajectory?\n\ni don't think it would be a perfect parametric curve. it might look a little \"wiggling\" and not smooth. So there are be some post-processing required, or better still, should we  predict the parameters of the curve of trajectory instead?\n\nhistory trajectory can also be represented as a parametric curve",
    "1077979": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3093b436d5ee8a87c7594c04ed1fb923%2FSelection_164.png?generation=1605337987662608&alt=media)",
    "1077899": "\"ensembling\" that works. i implemented it and it has some improvement\n\n```\n1. train 2 models:\nmodel1(x) : x -->embed1(x)-->head1(x)\nmodel2(x) : x -->embed2(x)-->head2(x)\n\n2. measure loss on validation set.\n- we plot loss1 vs loss2 to visually inspect correlation and diversity\n- we compute upper bound for improvement: best loss = average { min(loss1(x),loss2(x)) } \n\n3. make ensemble model\nmode: x -->embed1(x),embed2(x) --> concat --> head(concat)\n\n- we freeze embed1,embed2 to prevent overfitting. you should see some improvement\n\n```\n",
    "1076909": "What I am curious about is how do mode7 make inference? Take the top 3 confidence?",
    "1072972": "you might wan to check:\n\n\"Cycling rules in the US\nThere are differences in the cycling laws across the states, so please click here to check the bike laws based on your destination.\n\nThere are, however, a few general rules applicable wherever within the USA:\n\nIn the United States, everyone must drive on the right-hand side of the roadway. Never ride your bike against the traffic flow.\"",
    "1066996": "i was reading the paper introduced by @ryches \n\nMultimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks\nhttps://arxiv.org/pdf/1809.10732.pdf\n\n\n\"We collected 240 hours of data by manually driving SDV in\nPittsburgh, PA and Phoenix, AZ in various traffic conditions\n(e.g., varying times of day, days of the week). Traffic actors\nwere tracked using Unscented Kalman filter (UKF) \"\n\nNow I need to read the document of l5kit to see how the data is collected. i don't think anyone would annotate the location of the agent frame by frame. there must be some kind of processing data.",
    "1064337": "Using classification loss instead of regression for post trajectory selection makes sense. covernet uses the similar concept.\nhttps://arxiv.org/pdf/1911.10298.pdf\n\nfor classification loss they first sample sets of trajectories from train data and then select smallest set of fix trajectories using NP hard set cover problem.\nBut from my experiments, nnl loss performs better than covernets classification loss for this data.  Two reasons  that I could think of are first different types of agent like pedestrians and cyclist in this data which has very jittery trajectories and secondly covernet loss is usually better for higher number of modes.",
    "1064055": "Hey @engck23, cool to see that you are jumping in!\n\nYeah it's something I thought of. I was playing around with a model which has a prediction head and classification head, so in principle, you could predict 6 trajectories, and then only take the top 3 based on the classification score from the other head. \n\nNot sure if it will help, but it makes sense to me that there could be more then 3 viable routes, and the additional exploration might help. \nDon't have any concrete results as to whether or not it actually helps. \n\n"
  }
}