{
  "id": 178097,
  "title": "what is the data in each image channel?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/178097",
  "author_name": "",
  "post_date": "2020-08-28T15:30:26.999230100Z",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello fellows,</p>\n<p>I have spent the whole night just trying to understand the exact structure of the data generated by EgoDataset and AgentDataset. There's (hopefully) one last thing I haven't figured out.</p>\n<p>According to the official example, the number of channels of dataset image is calculated as:</p>\n<pre><code>num_history_channels = (cfg[\"model_params\"][\"history_num_frames\"] + 1) * 2\nnum_in_channels = 3 + num_history_channels\n</code></pre>\n<ol>\n<li>To my understanding, the reason for adding 1 in the first formula is that we need to add the present frame. But why multiplying by 2 here? Why we need two (not three) channels to express one moment?</li>\n<li>Why further adding a number of 3 in the second formula? What data goes into these three channels?</li>\n</ol>\n<p>I tried to read the source code, tracing all the way back to the function rasterize of module rasterizer. There is just a \"pass\" in the source code for function rasterize, which makes me even more confused.</p>\n<p>Any thoughts on this would be much appreciated.</p>\n<p>Thanks!<br>\nFrank</p>",
  "messages": [
    {
      "id": "989169",
      "postDate": "08/28/2020 15:30:27",
      "content": "<p>Hello fellows,</p>\n<p>I have spent the whole night just trying to understand the exact structure of the data generated by EgoDataset and AgentDataset. There's (hopefully) one last thing I haven't figured out.</p>\n<p>According to the official example, the number of channels of dataset image is calculated as:</p>\n<pre><code>num_history_channels = (cfg[\"model_params\"][\"history_num_frames\"] + 1) * 2\nnum_in_channels = 3 + num_history_channels\n</code></pre>\n<ol>\n<li>To my understanding, the reason for adding 1 in the first formula is that we need to add the present frame. But why multiplying by 2 here? Why we need two (not three) channels to express one moment?</li>\n<li>Why further adding a number of 3 in the second formula? What data goes into these three channels?</li>\n</ol>\n<p>I tried to read the source code, tracing all the way back to the function rasterize of module rasterizer. There is just a \"pass\" in the source code for function rasterize, which makes me even more confused.</p>\n<p>Any thoughts on this would be much appreciated.</p>\n<p>Thanks!<br>\nFrank</p>",
      "rawMarkdown": "Hello fellows,\n\nI have spent the whole night just trying to understand the exact structure of the data generated by EgoDataset and AgentDataset. There's (hopefully) one last thing I haven't figured out.\n\nAccording to the official example, the number of channels of dataset image is calculated as:\n\n    num_history_channels = (cfg[\"model_params\"][\"history_num_frames\"] + 1) * 2\n    num_in_channels = 3 + num_history_channels\n\n1. To my understanding, the reason for adding 1 in the first formula is that we need to add the present frame. But why multiplying by 2 here? Why we need two (not three) channels to express one moment?\n2. Why further adding a number of 3 in the second formula? What data goes into these three channels?\n\nI tried to read the source code, tracing all the way back to the function rasterize of module rasterizer. There is just a \"pass\" in the source code for function rasterize, which makes me even more confused.\n\nAny thoughts on this would be much appreciated.\n\nThanks!\nFrank",
      "votes": null
    },
    {
      "id": "989198",
      "postDate": "08/28/2020 15:58:43",
      "content": "<ul>\n<li>The history (+current) frames of the EGO and the agents are on different channels.</li>\n<li>3 extra (RGB) are for the semantic map (roads, lanes, crosswalks, etc)</li>\n</ul>",
      "rawMarkdown": "The history (+current) frames of the EGO and the agents are on different channels.\n- 3 extra (RGB) are for the semantic map (roads, lanes, crosswalks, etc)",
      "votes": null
    },
    {
      "id": "989646",
      "postDate": "08/29/2020 02:03:26",
      "content": "<p>Thank you so much Peter. That makes sense.</p>",
      "rawMarkdown": "Thank you so much Peter. That makes sense.",
      "votes": null
    },
    {
      "id": "990115",
      "postDate": "08/29/2020 10:48:28",
      "content": "<p>Can anyone confirm the order of these? </p>\n<p>Say we have history_num_frames = 2, when we load the 'image' dict entry does this comprise:</p>\n<p>agent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2,  semantic map R, G, B</p>\n<p>or is it</p>\n<p>agent t-2, agent t-1, agent t, ego t-2, ego t-1, ego t,  semantic map R, G, B</p>\n<p>??</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Can anyone confirm the order of these? \n\nSay we have history_num_frames = 2, when we load the 'image' dict entry does this comprise:\n\nagent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2,  semantic map R, G, B\n\nor is it\n\nagent t-2, agent t-1, agent t, ego t-2, ego t-1, ego t,  semantic map R, G, B\n\n??\n\nThanks!",
      "votes": null
    },
    {
      "id": "990296",
      "postDate": "08/29/2020 14:01:20",
      "content": "<p>Update:<br>\nI think it's agent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2, semantic map R, G, B<br>\nRef: <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> work <a href=\"https://www.kaggle.com/ryches/lyft-constant-velocity-extrapolation-baseline/notebook\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Update:\nI think it's agent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2, semantic map R, G, B\nRef: @ryches work [here](https://www.kaggle.com/ryches/lyft-constant-velocity-extrapolation-baseline/notebook)",
      "votes": null
    },
    {
      "id": "990375",
      "postDate": "08/29/2020 15:08:58",
      "content": "<p>Great! Another puzzle solved.</p>",
      "rawMarkdown": "Great! Another puzzle solved.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 989198,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "08/28/2020 15:58:43",
      "content": "<ul>\n<li>The history (+current) frames of the EGO and the agents are on different channels.</li>\n<li>3 extra (RGB) are for the semantic map (roads, lanes, crosswalks, etc)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 989646,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "08/29/2020 02:03:26",
          "content": "<p>Thank you so much Peter. That makes sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 990115,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "08/29/2020 10:48:28",
      "content": "<p>Can anyone confirm the order of these? </p>\n<p>Say we have history_num_frames = 2, when we load the 'image' dict entry does this comprise:</p>\n<p>agent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2,  semantic map R, G, B</p>\n<p>or is it</p>\n<p>agent t-2, agent t-1, agent t, ego t-2, ego t-1, ego t,  semantic map R, G, B</p>\n<p>??</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 990296,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "08/29/2020 14:01:20",
          "content": "<p>Update:<br>\nI think it's agent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2, semantic map R, G, B<br>\nRef: <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> work <a href=\"https://www.kaggle.com/ryches/lyft-constant-velocity-extrapolation-baseline/notebook\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990375,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "08/29/2020 15:08:58",
          "content": "<p>Great! Another puzzle solved.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "989169": "Hello fellows,\n\nI have spent the whole night just trying to understand the exact structure of the data generated by EgoDataset and AgentDataset. There's (hopefully) one last thing I haven't figured out.\n\nAccording to the official example, the number of channels of dataset image is calculated as:\n\n    num_history_channels = (cfg[\"model_params\"][\"history_num_frames\"] + 1) * 2\n    num_in_channels = 3 + num_history_channels\n\n1. To my understanding, the reason for adding 1 in the first formula is that we need to add the present frame. But why multiplying by 2 here? Why we need two (not three) channels to express one moment?\n2. Why further adding a number of 3 in the second formula? What data goes into these three channels?\n\nI tried to read the source code, tracing all the way back to the function rasterize of module rasterizer. There is just a \"pass\" in the source code for function rasterize, which makes me even more confused.\n\nAny thoughts on this would be much appreciated.\n\nThanks!\nFrank",
    "989198": "The history (+current) frames of the EGO and the agents are on different channels.\n- 3 extra (RGB) are for the semantic map (roads, lanes, crosswalks, etc)",
    "989646": "Thank you so much Peter. That makes sense.",
    "990115": "Can anyone confirm the order of these? \n\nSay we have history_num_frames = 2, when we load the 'image' dict entry does this comprise:\n\nagent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2,  semantic map R, G, B\n\nor is it\n\nagent t-2, agent t-1, agent t, ego t-2, ego t-1, ego t,  semantic map R, G, B\n\n??\n\nThanks!",
    "990296": "Update:\nI think it's agent t, agent t-1, agent t-2, ego t, ego t-1, ego t-2, semantic map R, G, B\nRef: @ryches work [here](https://www.kaggle.com/ryches/lyft-constant-velocity-extrapolation-baseline/notebook)",
    "990375": "Great! Another puzzle solved."
  },
  "source": "meta"
}