{
  "id": 200035,
  "title": "23rd solution (single model based on Resnet18)",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/a-elsheikh-23rd-solution-single-model-based-on-res",
  "author_name": "",
  "post_date": "2020-11-28T13:45:32.488871400Z",
  "votes": 18,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This was a nice dataset to play with. It is like HPC simulations, one have to relax and let things run for long. To my surprise, one can go very far using simple concepts and published work. My solution is simply an implementation of the two papers from Uber<br>\n1- Uncertainty-aware Short-term Motion Prediction of Traffic Actors for Autonomous Driving, <a href=\"https://arxiv.org/abs/1808.05819\" target=\"_blank\">https://arxiv.org/abs/1808.05819</a><br>\n2- Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks, <a href=\"https://arxiv.org/abs/1809.10732\" target=\"_blank\">https://arxiv.org/abs/1809.10732</a></p>\n<p>Now to some details which I think contributed to my best performing model:</p>\n<p>1- Sample the data to break the serial correlation in time. I only trained on frames with 60+10+1 gaps. I wanted to avoid data leakages so I kept a gap of history length (10) + future prediction (50) + (1) extra. Also, if a frame is too crowded (i.e. too many agents), I sampled a sub-set of the agents (20 in my code below).</p>\n<pre><code>def get_dataset_masks(\n        zarr_path,\n        zarr_dt,\n        th_agent_prob: float,\n        min_frame_history: int,\n        min_frame_future: int,\n        scene_mask=None,\n        chop_data=False,\n        chop_idx_list=[100, 200],\n        chop_agents=False):\n\n    \"\"\"\n\n    Modified from create_chopped_dataset\n\n\n    Returns:\n        mask: numpy array of bool to pass to AgentDatasetExtended\n    \"\"\"\n\n    if scene_mask is None:\n        scene_mask = np.ones(len(zarr_dt.scenes), dtype=np.bool)\n    else:\n        assert len(zarr_dt.scenes) == len(scene_mask), \"mask should be equal length\"\n\n    agents_mask_path = Path(zarr_path) / f\"agents_mask/{th_agent_prob}\"\n\n    if not agents_mask_path.exists():  # don't check in root but check for the path\n        assert 0\n    agents_mask_original = np.asarray(convenience.load(str(agents_mask_path)))\n\n    agents_mask = np.zeros(len(zarr_dt.agents), dtype=np.bool)\n    # for scene in zarr_dt.scenes.get_mask_selection(scene_mask):\n    for scene_idx in range(len(zarr_dt.scenes)):\n        if scene_mask[scene_idx] == 0:\n            continue\n        scene = zarr_dt.scenes[scene_idx]\n        if chop_data:\n            for num_frame_to_copy in chop_idx_list:\n                kept_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0] + num_frame_to_copy - 1]\n                agents_slice = get_agents_slice_from_frames(kept_frame)\n                # In create_chopped_dataset: no mask for min_frame_history\n                mask = agents_mask_original[agents_slice][:, 1] &gt;= min_frame_future\n                num_agents_per_frame = mask.sum()\n                max_agents_per_frame = 20\n                if chop_agents and num_agents_per_frame &gt; max_agents_per_frame:  # or 10\n                    removed_indices = np.random.choice(\n                        np.where(mask)[0],\n                        num_agents_per_frame - max_agents_per_frame,\n                        replace=False)\n                    mask[removed_indices] = False\n                agents_mask[agents_slice] = mask.copy()\n\n        else:\n            first_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0]]\n            last_frame = zarr_dt.frames[scene[\"frame_index_interval\"][1] - 1]\n            agents_slice = get_agents_slice_from_frames(first_frame, last_frame)\n\n            past_mask = agents_mask_original[agents_slice][:, 0] &gt;= min_frame_history\n            future_mask = agents_mask_original[agents_slice][:, 1] &gt;= min_frame_future\n            mask = past_mask * future_mask\n            agents_mask[agents_slice] = mask.copy()\n\n    return agents_mask\n</code></pre>\n<p>2-  I had a clean implementation of the Multiple-Trajectory Prediction (MTP) loss. I trained on the MoN loss and validated on the nll loss</p>\n<pre><code>def uber_like_loss_new(gt, pred, confidences, avails):\n    batch_size, num_modes, future_len, num_coords = pred.shape\n\n    # ensure that your model outputs logits\n    gt = gt[:, None, :, :]  # add modes\n    avails = avails[:, None, :, None]  # add modes and cords\n    l2_error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    l2_error = torch.sum(l2_error, dim=-1)  # reduce future_len\n\n    best_mode_target = torch.argmin(l2_error, dim=1).detach()\n    classification_loss = torch.nn.functional.cross_entropy(confidences, best_mode_target, reduction='none')\n\n    alpha = 1.0\n    MoN_error = classification_loss + alpha * l2_error[torch.arange(batch_size), best_mode_target]\n    MoN_error = MoN_error.reshape(-1, 1)\n    error = torch.nn.functional.log_softmax(confidences, dim=1) - 0.5 * l2_error  # reduce future_len\n    max_value, _ = torch.max(error, dim=-1, keepdim=True)  # error are negative at this point, so max() gives the minimum one\n    nll_error = -torch.log(torch.sum(torch.exp(error - max_value), dim=-1, keepdim=True)) - max_value\n    return MoN_error, nll_error\n</code></pre>\n<p>3- I extracted some meta data from the AgentDataset. Here is a minimal implementation </p>\n<pre><code>class AgentDatasetExtended(AgentDataset):\n    def __init__(\n        self,\n        cfg,\n        zarr_dataset,\n        rasterizer,\n        perturbation,\n        agents_mask,\n        min_frame_history,\n        min_frame_future,\n        transform,\n        l5kit_version,\n    ):\n        assert perturbation is None, \"AgentDataset does not support perturbation (yet)\"\n        super(AgentDatasetExtended, self).__init__(\n            cfg, zarr_dataset, rasterizer, perturbation, agents_mask, min_frame_history, min_frame_future)\n        self.min_frame_future = min_frame_future\n        self.min_frame_history = min_frame_history\n        self.transform = transform\n        self.l5kit_version = l5kit_version\n\n    def __getitem__(self, index: int) -&gt; dict:\n        \"\"\"\n        Differs from parent returning the indices of the frame, agent and scene\n        \"\"\"\n        if index &lt; 0:\n            if -index &gt; len(self):\n                raise ValueError(\"absolute value of index should not exceed dataset length\")\n            index = len(self) + index\n\n        index = self.agents_indices[index]\n        track_id = self.dataset.agents[index][\"track_id\"]\n        frame_index = bisect.bisect_right(self.cumulative_sizes_agents, index)\n        scene_index = bisect.bisect_right(self.cumulative_sizes, frame_index)\n\n        if scene_index == 0:\n            state_index = frame_index\n        else:\n            state_index = frame_index - self.cumulative_sizes[scene_index - 1]\n        data_dic = self.get_frame(scene_index, state_index, track_id=track_id)\n\n        # track_id = self.dataset.agents[index][\"track_id\"]\n        # centroid = self.dataset.agents[index][\"centroid\"]\n        # yaw = self.dataset.agents[index][\"yaw\"]\n        velocity = self.dataset.agents[index][\"velocity\"]\n        label_probabilities = self.dataset.agents[index][\"label_probabilities\"]\n        # data_dic['track_id2'] = np.int64(track_id)\n        # data_dic['centroid2'] = centroid\n        # data_dic['yaw2'] = yaw\n        data_dic['velocity'] = velocity\n        data_dic['label_probabilities'] = label_probabilities\n\n        data_dic['ego_translation'] = self.dataset.frames[frame_index][\"ego_translation\"]\n        data_dic['ego_rotation'] = self.dataset.frames[frame_index][\"ego_rotation\"]  # matrix\n        data_dic['timestamp'] = self.dataset.frames[frame_index][\"timestamp\"]\n        data_dic['hour'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).hour\n        data_dic['weekday'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).weekday()\n\n        if self.transform:\n            data_dic = self.transform(data_dic)\n        return data_dic\n</code></pre>\n<p>4- I had a second head of 47 inputs (probably too much), which is fed all sort of metadata (label_probabilities, yaw, extent, acceleration, ego_translation, ego_centroid_diff, hour, weekday) and then passed through two fully connected layers before merging with the pooling layers of the backbone. This concatenated vector is then passed to dense layer, relu and then the output layer.</p>\n<p>5- I applied a <code>cumsum</code> function to NN output to force the NN to produce something like the differences, then I applied the image_to_world_matrix transformation before calling the loss function of the real coordinate. I wanted the output of my NN model to be in the image space (conceptually). </p>\n<p>There are many small additional details. I used lookahead optimizer wrapped around ADAM with starting learning rate of 1e-4 that and then decreased the LR by 0.99 every 100000 steps over batches of 32 samples (which takes about an hour on my machine RTX-Titan-X). Learning rate is reduced after 20 iterations of the 320k samples. Best solution is reached after 110 of these 320k samples (best model Private LB 13.580, Public LB 14.265). I then averaged 7 models based on checkpointing (Private LB 13.154, Public LB 13.722). My image size is 336x336 with a resolution of 0.25x0.25, trained on full data with the following frames <code>chop_idx_list_train = [11, 11 + 61, 11 + 2 * 61, 11 + 3 * 61]</code></p>\n<p>Probably, few simple modification could produce a bit better score.<br>\n(a) A bigger patch of 64 samples might perform better (diversity of the samples). <br>\n(b) Change the backbone to ResNet50 or MobileNet-V2</p>",
  "messages": [
    {
      "id": "1094304",
      "postDate": "11/28/2020 13:45:32",
      "content": "<p>This was a nice dataset to play with. It is like HPC simulations, one have to relax and let things run for long. To my surprise, one can go very far using simple concepts and published work. My solution is simply an implementation of the two papers from Uber<br>\n1- Uncertainty-aware Short-term Motion Prediction of Traffic Actors for Autonomous Driving, <a href=\"https://arxiv.org/abs/1808.05819\" target=\"_blank\">https://arxiv.org/abs/1808.05819</a><br>\n2- Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks, <a href=\"https://arxiv.org/abs/1809.10732\" target=\"_blank\">https://arxiv.org/abs/1809.10732</a></p>\n<p>Now to some details which I think contributed to my best performing model:</p>\n<p>1- Sample the data to break the serial correlation in time. I only trained on frames with 60+10+1 gaps. I wanted to avoid data leakages so I kept a gap of history length (10) + future prediction (50) + (1) extra. Also, if a frame is too crowded (i.e. too many agents), I sampled a sub-set of the agents (20 in my code below).</p>\n<pre><code>def get_dataset_masks(\n        zarr_path,\n        zarr_dt,\n        th_agent_prob: float,\n        min_frame_history: int,\n        min_frame_future: int,\n        scene_mask=None,\n        chop_data=False,\n        chop_idx_list=[100, 200],\n        chop_agents=False):\n\n    \"\"\"\n\n    Modified from create_chopped_dataset\n\n\n    Returns:\n        mask: numpy array of bool to pass to AgentDatasetExtended\n    \"\"\"\n\n    if scene_mask is None:\n        scene_mask = np.ones(len(zarr_dt.scenes), dtype=np.bool)\n    else:\n        assert len(zarr_dt.scenes) == len(scene_mask), \"mask should be equal length\"\n\n    agents_mask_path = Path(zarr_path) / f\"agents_mask/{th_agent_prob}\"\n\n    if not agents_mask_path.exists():  # don't check in root but check for the path\n        assert 0\n    agents_mask_original = np.asarray(convenience.load(str(agents_mask_path)))\n\n    agents_mask = np.zeros(len(zarr_dt.agents), dtype=np.bool)\n    # for scene in zarr_dt.scenes.get_mask_selection(scene_mask):\n    for scene_idx in range(len(zarr_dt.scenes)):\n        if scene_mask[scene_idx] == 0:\n            continue\n        scene = zarr_dt.scenes[scene_idx]\n        if chop_data:\n            for num_frame_to_copy in chop_idx_list:\n                kept_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0] + num_frame_to_copy - 1]\n                agents_slice = get_agents_slice_from_frames(kept_frame)\n                # In create_chopped_dataset: no mask for min_frame_history\n                mask = agents_mask_original[agents_slice][:, 1] &gt;= min_frame_future\n                num_agents_per_frame = mask.sum()\n                max_agents_per_frame = 20\n                if chop_agents and num_agents_per_frame &gt; max_agents_per_frame:  # or 10\n                    removed_indices = np.random.choice(\n                        np.where(mask)[0],\n                        num_agents_per_frame - max_agents_per_frame,\n                        replace=False)\n                    mask[removed_indices] = False\n                agents_mask[agents_slice] = mask.copy()\n\n        else:\n            first_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0]]\n            last_frame = zarr_dt.frames[scene[\"frame_index_interval\"][1] - 1]\n            agents_slice = get_agents_slice_from_frames(first_frame, last_frame)\n\n            past_mask = agents_mask_original[agents_slice][:, 0] &gt;= min_frame_history\n            future_mask = agents_mask_original[agents_slice][:, 1] &gt;= min_frame_future\n            mask = past_mask * future_mask\n            agents_mask[agents_slice] = mask.copy()\n\n    return agents_mask\n</code></pre>\n<p>2-  I had a clean implementation of the Multiple-Trajectory Prediction (MTP) loss. I trained on the MoN loss and validated on the nll loss</p>\n<pre><code>def uber_like_loss_new(gt, pred, confidences, avails):\n    batch_size, num_modes, future_len, num_coords = pred.shape\n\n    # ensure that your model outputs logits\n    gt = gt[:, None, :, :]  # add modes\n    avails = avails[:, None, :, None]  # add modes and cords\n    l2_error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    l2_error = torch.sum(l2_error, dim=-1)  # reduce future_len\n\n    best_mode_target = torch.argmin(l2_error, dim=1).detach()\n    classification_loss = torch.nn.functional.cross_entropy(confidences, best_mode_target, reduction='none')\n\n    alpha = 1.0\n    MoN_error = classification_loss + alpha * l2_error[torch.arange(batch_size), best_mode_target]\n    MoN_error = MoN_error.reshape(-1, 1)\n    error = torch.nn.functional.log_softmax(confidences, dim=1) - 0.5 * l2_error  # reduce future_len\n    max_value, _ = torch.max(error, dim=-1, keepdim=True)  # error are negative at this point, so max() gives the minimum one\n    nll_error = -torch.log(torch.sum(torch.exp(error - max_value), dim=-1, keepdim=True)) - max_value\n    return MoN_error, nll_error\n</code></pre>\n<p>3- I extracted some meta data from the AgentDataset. Here is a minimal implementation </p>\n<pre><code>class AgentDatasetExtended(AgentDataset):\n    def __init__(\n        self,\n        cfg,\n        zarr_dataset,\n        rasterizer,\n        perturbation,\n        agents_mask,\n        min_frame_history,\n        min_frame_future,\n        transform,\n        l5kit_version,\n    ):\n        assert perturbation is None, \"AgentDataset does not support perturbation (yet)\"\n        super(AgentDatasetExtended, self).__init__(\n            cfg, zarr_dataset, rasterizer, perturbation, agents_mask, min_frame_history, min_frame_future)\n        self.min_frame_future = min_frame_future\n        self.min_frame_history = min_frame_history\n        self.transform = transform\n        self.l5kit_version = l5kit_version\n\n    def __getitem__(self, index: int) -&gt; dict:\n        \"\"\"\n        Differs from parent returning the indices of the frame, agent and scene\n        \"\"\"\n        if index &lt; 0:\n            if -index &gt; len(self):\n                raise ValueError(\"absolute value of index should not exceed dataset length\")\n            index = len(self) + index\n\n        index = self.agents_indices[index]\n        track_id = self.dataset.agents[index][\"track_id\"]\n        frame_index = bisect.bisect_right(self.cumulative_sizes_agents, index)\n        scene_index = bisect.bisect_right(self.cumulative_sizes, frame_index)\n\n        if scene_index == 0:\n            state_index = frame_index\n        else:\n            state_index = frame_index - self.cumulative_sizes[scene_index - 1]\n        data_dic = self.get_frame(scene_index, state_index, track_id=track_id)\n\n        # track_id = self.dataset.agents[index][\"track_id\"]\n        # centroid = self.dataset.agents[index][\"centroid\"]\n        # yaw = self.dataset.agents[index][\"yaw\"]\n        velocity = self.dataset.agents[index][\"velocity\"]\n        label_probabilities = self.dataset.agents[index][\"label_probabilities\"]\n        # data_dic['track_id2'] = np.int64(track_id)\n        # data_dic['centroid2'] = centroid\n        # data_dic['yaw2'] = yaw\n        data_dic['velocity'] = velocity\n        data_dic['label_probabilities'] = label_probabilities\n\n        data_dic['ego_translation'] = self.dataset.frames[frame_index][\"ego_translation\"]\n        data_dic['ego_rotation'] = self.dataset.frames[frame_index][\"ego_rotation\"]  # matrix\n        data_dic['timestamp'] = self.dataset.frames[frame_index][\"timestamp\"]\n        data_dic['hour'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).hour\n        data_dic['weekday'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).weekday()\n\n        if self.transform:\n            data_dic = self.transform(data_dic)\n        return data_dic\n</code></pre>\n<p>4- I had a second head of 47 inputs (probably too much), which is fed all sort of metadata (label_probabilities, yaw, extent, acceleration, ego_translation, ego_centroid_diff, hour, weekday) and then passed through two fully connected layers before merging with the pooling layers of the backbone. This concatenated vector is then passed to dense layer, relu and then the output layer.</p>\n<p>5- I applied a <code>cumsum</code> function to NN output to force the NN to produce something like the differences, then I applied the image_to_world_matrix transformation before calling the loss function of the real coordinate. I wanted the output of my NN model to be in the image space (conceptually). </p>\n<p>There are many small additional details. I used lookahead optimizer wrapped around ADAM with starting learning rate of 1e-4 that and then decreased the LR by 0.99 every 100000 steps over batches of 32 samples (which takes about an hour on my machine RTX-Titan-X). Learning rate is reduced after 20 iterations of the 320k samples. Best solution is reached after 110 of these 320k samples (best model Private LB 13.580, Public LB 14.265). I then averaged 7 models based on checkpointing (Private LB 13.154, Public LB 13.722). My image size is 336x336 with a resolution of 0.25x0.25, trained on full data with the following frames <code>chop_idx_list_train = [11, 11 + 61, 11 + 2 * 61, 11 + 3 * 61]</code></p>\n<p>Probably, few simple modification could produce a bit better score.<br>\n(a) A bigger patch of 64 samples might perform better (diversity of the samples). <br>\n(b) Change the backbone to ResNet50 or MobileNet-V2</p>",
      "rawMarkdown": "This was a nice dataset to play with. It is like HPC simulations, one have to relax and let things run for long. To my surprise, one can go very far using simple concepts and published work. My solution is simply an implementation of the two papers from Uber\n1- Uncertainty-aware Short-term Motion Prediction of Traffic Actors for Autonomous Driving, https://arxiv.org/abs/1808.05819\n2- Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks, https://arxiv.org/abs/1809.10732\n\nNow to some details which I think contributed to my best performing model:\n\n1- Sample the data to break the serial correlation in time. I only trained on frames with 60+10+1 gaps. I wanted to avoid data leakages so I kept a gap of history length (10) + future prediction (50) + (1) extra. Also, if a frame is too crowded (i.e. too many agents), I sampled a sub-set of the agents (20 in my code below).\n\n```\ndef get_dataset_masks(\n        zarr_path,\n        zarr_dt,\n        th_agent_prob: float,\n        min_frame_history: int,\n        min_frame_future: int,\n        scene_mask=None,\n        chop_data=False,\n        chop_idx_list=[100, 200],\n        chop_agents=False):\n\n    \"\"\"\n\n    Modified from create_chopped_dataset\n\n\n    Returns:\n        mask: numpy array of bool to pass to AgentDatasetExtended\n    \"\"\"\n\n    if scene_mask is None:\n        scene_mask = np.ones(len(zarr_dt.scenes), dtype=np.bool)\n    else:\n        assert len(zarr_dt.scenes) == len(scene_mask), \"mask should be equal length\"\n\n    agents_mask_path = Path(zarr_path) / f\"agents_mask/{th_agent_prob}\"\n\n    if not agents_mask_path.exists():  # don't check in root but check for the path\n        assert 0\n    agents_mask_original = np.asarray(convenience.load(str(agents_mask_path)))\n\n    agents_mask = np.zeros(len(zarr_dt.agents), dtype=np.bool)\n    # for scene in zarr_dt.scenes.get_mask_selection(scene_mask):\n    for scene_idx in range(len(zarr_dt.scenes)):\n        if scene_mask[scene_idx] == 0:\n            continue\n        scene = zarr_dt.scenes[scene_idx]\n        if chop_data:\n            for num_frame_to_copy in chop_idx_list:\n                kept_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0] + num_frame_to_copy - 1]\n                agents_slice = get_agents_slice_from_frames(kept_frame)\n                # In create_chopped_dataset: no mask for min_frame_history\n                mask = agents_mask_original[agents_slice][:, 1] >= min_frame_future\n                num_agents_per_frame = mask.sum()\n                max_agents_per_frame = 20\n                if chop_agents and num_agents_per_frame > max_agents_per_frame:  # or 10\n                    removed_indices = np.random.choice(\n                        np.where(mask)[0],\n                        num_agents_per_frame - max_agents_per_frame,\n                        replace=False)\n                    mask[removed_indices] = False\n                agents_mask[agents_slice] = mask.copy()\n\n        else:\n            first_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0]]\n            last_frame = zarr_dt.frames[scene[\"frame_index_interval\"][1] - 1]\n            agents_slice = get_agents_slice_from_frames(first_frame, last_frame)\n\n            past_mask = agents_mask_original[agents_slice][:, 0] >= min_frame_history\n            future_mask = agents_mask_original[agents_slice][:, 1] >= min_frame_future\n            mask = past_mask * future_mask\n            agents_mask[agents_slice] = mask.copy()\n\n    return agents_mask\n```\n2-  I had a clean implementation of the Multiple-Trajectory Prediction (MTP) loss. I trained on the MoN loss and validated on the nll loss\n\n```\ndef uber_like_loss_new(gt, pred, confidences, avails):\n    batch_size, num_modes, future_len, num_coords = pred.shape\n\n    # ensure that your model outputs logits\n    gt = gt[:, None, :, :]  # add modes\n    avails = avails[:, None, :, None]  # add modes and cords\n    l2_error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    l2_error = torch.sum(l2_error, dim=-1)  # reduce future_len\n\n    best_mode_target = torch.argmin(l2_error, dim=1).detach()\n    classification_loss = torch.nn.functional.cross_entropy(confidences, best_mode_target, reduction='none')\n\n    alpha = 1.0\n    MoN_error = classification_loss + alpha * l2_error[torch.arange(batch_size), best_mode_target]\n    MoN_error = MoN_error.reshape(-1, 1)\n    error = torch.nn.functional.log_softmax(confidences, dim=1) - 0.5 * l2_error  # reduce future_len\n    max_value, _ = torch.max(error, dim=-1, keepdim=True)  # error are negative at this point, so max() gives the minimum one\n    nll_error = -torch.log(torch.sum(torch.exp(error - max_value), dim=-1, keepdim=True)) - max_value\n    return MoN_error, nll_error\n```\n\n3- I extracted some meta data from the AgentDataset. Here is a minimal implementation \n\n```\nclass AgentDatasetExtended(AgentDataset):\n    def __init__(\n        self,\n        cfg,\n        zarr_dataset,\n        rasterizer,\n        perturbation,\n        agents_mask,\n        min_frame_history,\n        min_frame_future,\n        transform,\n        l5kit_version,\n    ):\n        assert perturbation is None, \"AgentDataset does not support perturbation (yet)\"\n        super(AgentDatasetExtended, self).__init__(\n            cfg, zarr_dataset, rasterizer, perturbation, agents_mask, min_frame_history, min_frame_future)\n        self.min_frame_future = min_frame_future\n        self.min_frame_history = min_frame_history\n        self.transform = transform\n        self.l5kit_version = l5kit_version\n\n    def __getitem__(self, index: int) -> dict:\n        \"\"\"\n        Differs from parent returning the indices of the frame, agent and scene\n        \"\"\"\n        if index < 0:\n            if -index > len(self):\n                raise ValueError(\"absolute value of index should not exceed dataset length\")\n            index = len(self) + index\n\n        index = self.agents_indices[index]\n        track_id = self.dataset.agents[index][\"track_id\"]\n        frame_index = bisect.bisect_right(self.cumulative_sizes_agents, index)\n        scene_index = bisect.bisect_right(self.cumulative_sizes, frame_index)\n\n        if scene_index == 0:\n            state_index = frame_index\n        else:\n            state_index = frame_index - self.cumulative_sizes[scene_index - 1]\n        data_dic = self.get_frame(scene_index, state_index, track_id=track_id)\n\n        # track_id = self.dataset.agents[index][\"track_id\"]\n        # centroid = self.dataset.agents[index][\"centroid\"]\n        # yaw = self.dataset.agents[index][\"yaw\"]\n        velocity = self.dataset.agents[index][\"velocity\"]\n        label_probabilities = self.dataset.agents[index][\"label_probabilities\"]\n        # data_dic['track_id2'] = np.int64(track_id)\n        # data_dic['centroid2'] = centroid\n        # data_dic['yaw2'] = yaw\n        data_dic['velocity'] = velocity\n        data_dic['label_probabilities'] = label_probabilities\n\n        data_dic['ego_translation'] = self.dataset.frames[frame_index][\"ego_translation\"]\n        data_dic['ego_rotation'] = self.dataset.frames[frame_index][\"ego_rotation\"]  # matrix\n        data_dic['timestamp'] = self.dataset.frames[frame_index][\"timestamp\"]\n        data_dic['hour'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).hour\n        data_dic['weekday'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).weekday()\n\n        if self.transform:\n            data_dic = self.transform(data_dic)\n        return data_dic\n```\n\n4- I had a second head of 47 inputs (probably too much), which is fed all sort of metadata (label_probabilities, yaw, extent, acceleration, ego_translation, ego_centroid_diff, hour, weekday) and then passed through two fully connected layers before merging with the pooling layers of the backbone. This concatenated vector is then passed to dense layer, relu and then the output layer.\n\n5- I applied a `cumsum` function to NN output to force the NN to produce something like the differences, then I applied the image_to_world_matrix transformation before calling the loss function of the real coordinate. I wanted the output of my NN model to be in the image space (conceptually). \n\nThere are many small additional details. I used lookahead optimizer wrapped around ADAM with starting learning rate of 1e-4 that and then decreased the LR by 0.99 every 100000 steps over batches of 32 samples (which takes about an hour on my machine RTX-Titan-X). Learning rate is reduced after 20 iterations of the 320k samples. Best solution is reached after 110 of these 320k samples (best model Private LB 13.580, Public LB 14.265). I then averaged 7 models based on checkpointing (Private LB 13.154, Public LB 13.722). My image size is 336x336 with a resolution of 0.25x0.25, trained on full data with the following frames `chop_idx_list_train = [11, 11 + 61, 11 + 2 * 61, 11 + 3 * 61]`\n\nProbably, few simple modification could produce a bit better score.\n(a) A bigger patch of 64 samples might perform better (diversity of the samples). \n(b) Change the backbone to ResNet50 or MobileNet-V2",
      "votes": null
    },
    {
      "id": "1098871",
      "postDate": "12/01/2020 23:10:53",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ggrizzly\" target=\"_blank\">@ggrizzly</a> and thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats @ggrizzly and thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "1115233",
      "postDate": "12/16/2020 05:36:58",
      "content": "<p>multi-modal paper suggests to use angle based distance metric to choose best mode, did you try that?</p>",
      "rawMarkdown": "multi-modal paper suggests to use angle based distance metric to choose best mode, did you try that?",
      "votes": null
    },
    {
      "id": "1117287",
      "postDate": "12/17/2020 21:54:58",
      "content": "<p>I didn't test that and as reported in the paper, the differences interms of the evaluation metrics are not that big. However, I recently came across the diversity loss (see <a href=\"https://arxiv.org/abs/2011.15084\" target=\"_blank\">https://arxiv.org/abs/2011.15084</a>) by adding a L2 penalty on similar predictions to the NLL loss.</p>",
      "rawMarkdown": "I didn't test that and as reported in the paper, the differences interms of the evaluation metrics are not that big. However, I recently came across the diversity loss (see https://arxiv.org/abs/2011.15084) by adding a L2 penalty on similar predictions to the NLL loss.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1098871,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "12/01/2020 23:10:53",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ggrizzly\" target=\"_blank\">@ggrizzly</a> and thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115233,
      "author_name": "akshayraos",
      "author_url": "",
      "post_date": "12/16/2020 05:36:58",
      "content": "<p>multi-modal paper suggests to use angle based distance metric to choose best mode, did you try that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1117287,
          "author_name": "ggrizzly",
          "author_url": "",
          "post_date": "12/17/2020 21:54:58",
          "content": "<p>I didn't test that and as reported in the paper, the differences interms of the evaluation metrics are not that big. However, I recently came across the diversity loss (see <a href=\"https://arxiv.org/abs/2011.15084\" target=\"_blank\">https://arxiv.org/abs/2011.15084</a>) by adding a L2 penalty on similar predictions to the NLL loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1094304": "This was a nice dataset to play with. It is like HPC simulations, one have to relax and let things run for long. To my surprise, one can go very far using simple concepts and published work. My solution is simply an implementation of the two papers from Uber\n1- Uncertainty-aware Short-term Motion Prediction of Traffic Actors for Autonomous Driving, https://arxiv.org/abs/1808.05819\n2- Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks, https://arxiv.org/abs/1809.10732\n\nNow to some details which I think contributed to my best performing model:\n\n1- Sample the data to break the serial correlation in time. I only trained on frames with 60+10+1 gaps. I wanted to avoid data leakages so I kept a gap of history length (10) + future prediction (50) + (1) extra. Also, if a frame is too crowded (i.e. too many agents), I sampled a sub-set of the agents (20 in my code below).\n\n```\ndef get_dataset_masks(\n        zarr_path,\n        zarr_dt,\n        th_agent_prob: float,\n        min_frame_history: int,\n        min_frame_future: int,\n        scene_mask=None,\n        chop_data=False,\n        chop_idx_list=[100, 200],\n        chop_agents=False):\n\n    \"\"\"\n\n    Modified from create_chopped_dataset\n\n\n    Returns:\n        mask: numpy array of bool to pass to AgentDatasetExtended\n    \"\"\"\n\n    if scene_mask is None:\n        scene_mask = np.ones(len(zarr_dt.scenes), dtype=np.bool)\n    else:\n        assert len(zarr_dt.scenes) == len(scene_mask), \"mask should be equal length\"\n\n    agents_mask_path = Path(zarr_path) / f\"agents_mask/{th_agent_prob}\"\n\n    if not agents_mask_path.exists():  # don't check in root but check for the path\n        assert 0\n    agents_mask_original = np.asarray(convenience.load(str(agents_mask_path)))\n\n    agents_mask = np.zeros(len(zarr_dt.agents), dtype=np.bool)\n    # for scene in zarr_dt.scenes.get_mask_selection(scene_mask):\n    for scene_idx in range(len(zarr_dt.scenes)):\n        if scene_mask[scene_idx] == 0:\n            continue\n        scene = zarr_dt.scenes[scene_idx]\n        if chop_data:\n            for num_frame_to_copy in chop_idx_list:\n                kept_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0] + num_frame_to_copy - 1]\n                agents_slice = get_agents_slice_from_frames(kept_frame)\n                # In create_chopped_dataset: no mask for min_frame_history\n                mask = agents_mask_original[agents_slice][:, 1] >= min_frame_future\n                num_agents_per_frame = mask.sum()\n                max_agents_per_frame = 20\n                if chop_agents and num_agents_per_frame > max_agents_per_frame:  # or 10\n                    removed_indices = np.random.choice(\n                        np.where(mask)[0],\n                        num_agents_per_frame - max_agents_per_frame,\n                        replace=False)\n                    mask[removed_indices] = False\n                agents_mask[agents_slice] = mask.copy()\n\n        else:\n            first_frame = zarr_dt.frames[scene[\"frame_index_interval\"][0]]\n            last_frame = zarr_dt.frames[scene[\"frame_index_interval\"][1] - 1]\n            agents_slice = get_agents_slice_from_frames(first_frame, last_frame)\n\n            past_mask = agents_mask_original[agents_slice][:, 0] >= min_frame_history\n            future_mask = agents_mask_original[agents_slice][:, 1] >= min_frame_future\n            mask = past_mask * future_mask\n            agents_mask[agents_slice] = mask.copy()\n\n    return agents_mask\n```\n2-  I had a clean implementation of the Multiple-Trajectory Prediction (MTP) loss. I trained on the MoN loss and validated on the nll loss\n\n```\ndef uber_like_loss_new(gt, pred, confidences, avails):\n    batch_size, num_modes, future_len, num_coords = pred.shape\n\n    # ensure that your model outputs logits\n    gt = gt[:, None, :, :]  # add modes\n    avails = avails[:, None, :, None]  # add modes and cords\n    l2_error = torch.sum(((gt - pred) * avails) ** 2, dim=-1)  # reduce coords and use availability\n    l2_error = torch.sum(l2_error, dim=-1)  # reduce future_len\n\n    best_mode_target = torch.argmin(l2_error, dim=1).detach()\n    classification_loss = torch.nn.functional.cross_entropy(confidences, best_mode_target, reduction='none')\n\n    alpha = 1.0\n    MoN_error = classification_loss + alpha * l2_error[torch.arange(batch_size), best_mode_target]\n    MoN_error = MoN_error.reshape(-1, 1)\n    error = torch.nn.functional.log_softmax(confidences, dim=1) - 0.5 * l2_error  # reduce future_len\n    max_value, _ = torch.max(error, dim=-1, keepdim=True)  # error are negative at this point, so max() gives the minimum one\n    nll_error = -torch.log(torch.sum(torch.exp(error - max_value), dim=-1, keepdim=True)) - max_value\n    return MoN_error, nll_error\n```\n\n3- I extracted some meta data from the AgentDataset. Here is a minimal implementation \n\n```\nclass AgentDatasetExtended(AgentDataset):\n    def __init__(\n        self,\n        cfg,\n        zarr_dataset,\n        rasterizer,\n        perturbation,\n        agents_mask,\n        min_frame_history,\n        min_frame_future,\n        transform,\n        l5kit_version,\n    ):\n        assert perturbation is None, \"AgentDataset does not support perturbation (yet)\"\n        super(AgentDatasetExtended, self).__init__(\n            cfg, zarr_dataset, rasterizer, perturbation, agents_mask, min_frame_history, min_frame_future)\n        self.min_frame_future = min_frame_future\n        self.min_frame_history = min_frame_history\n        self.transform = transform\n        self.l5kit_version = l5kit_version\n\n    def __getitem__(self, index: int) -> dict:\n        \"\"\"\n        Differs from parent returning the indices of the frame, agent and scene\n        \"\"\"\n        if index < 0:\n            if -index > len(self):\n                raise ValueError(\"absolute value of index should not exceed dataset length\")\n            index = len(self) + index\n\n        index = self.agents_indices[index]\n        track_id = self.dataset.agents[index][\"track_id\"]\n        frame_index = bisect.bisect_right(self.cumulative_sizes_agents, index)\n        scene_index = bisect.bisect_right(self.cumulative_sizes, frame_index)\n\n        if scene_index == 0:\n            state_index = frame_index\n        else:\n            state_index = frame_index - self.cumulative_sizes[scene_index - 1]\n        data_dic = self.get_frame(scene_index, state_index, track_id=track_id)\n\n        # track_id = self.dataset.agents[index][\"track_id\"]\n        # centroid = self.dataset.agents[index][\"centroid\"]\n        # yaw = self.dataset.agents[index][\"yaw\"]\n        velocity = self.dataset.agents[index][\"velocity\"]\n        label_probabilities = self.dataset.agents[index][\"label_probabilities\"]\n        # data_dic['track_id2'] = np.int64(track_id)\n        # data_dic['centroid2'] = centroid\n        # data_dic['yaw2'] = yaw\n        data_dic['velocity'] = velocity\n        data_dic['label_probabilities'] = label_probabilities\n\n        data_dic['ego_translation'] = self.dataset.frames[frame_index][\"ego_translation\"]\n        data_dic['ego_rotation'] = self.dataset.frames[frame_index][\"ego_rotation\"]  # matrix\n        data_dic['timestamp'] = self.dataset.frames[frame_index][\"timestamp\"]\n        data_dic['hour'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).hour\n        data_dic['weekday'] = datetime.fromtimestamp(data_dic['timestamp'] / 1e9).weekday()\n\n        if self.transform:\n            data_dic = self.transform(data_dic)\n        return data_dic\n```\n\n4- I had a second head of 47 inputs (probably too much), which is fed all sort of metadata (label_probabilities, yaw, extent, acceleration, ego_translation, ego_centroid_diff, hour, weekday) and then passed through two fully connected layers before merging with the pooling layers of the backbone. This concatenated vector is then passed to dense layer, relu and then the output layer.\n\n5- I applied a `cumsum` function to NN output to force the NN to produce something like the differences, then I applied the image_to_world_matrix transformation before calling the loss function of the real coordinate. I wanted the output of my NN model to be in the image space (conceptually). \n\nThere are many small additional details. I used lookahead optimizer wrapped around ADAM with starting learning rate of 1e-4 that and then decreased the LR by 0.99 every 100000 steps over batches of 32 samples (which takes about an hour on my machine RTX-Titan-X). Learning rate is reduced after 20 iterations of the 320k samples. Best solution is reached after 110 of these 320k samples (best model Private LB 13.580, Public LB 14.265). I then averaged 7 models based on checkpointing (Private LB 13.154, Public LB 13.722). My image size is 336x336 with a resolution of 0.25x0.25, trained on full data with the following frames `chop_idx_list_train = [11, 11 + 61, 11 + 2 * 61, 11 + 3 * 61]`\n\nProbably, few simple modification could produce a bit better score.\n(a) A bigger patch of 64 samples might perform better (diversity of the samples). \n(b) Change the backbone to ResNet50 or MobileNet-V2",
    "1098871": "Congrats @ggrizzly and thanks for sharing your solution.",
    "1115233": "multi-modal paper suggests to use angle based distance metric to choose best mode, did you try that?",
    "1117287": "I didn't test that and as reported in the paper, the differences interms of the evaluation metrics are not that big. However, I recently came across the diversity loss (see https://arxiv.org/abs/2011.15084) by adding a L2 penalty on similar predictions to the NLL loss."
  },
  "source": "meta"
}