{
  "id": 194396,
  "title": "Difficulties to understand exact problem and data",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/194396",
  "author_name": "",
  "post_date": "2020-11-01T15:52:12.995665600Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This post is meant to be both a question as well as a feedback to the organizers. I really do appreciate posting this exciting problem, sharing the data and the following threads from the organizer:</p>\n<ul>\n<li><a href=\"https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format</a></li>\n<li><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/179246\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/179246</a></li>\n<li><a href=\"https://www.kaggle.com/lucabergamini/lyft-baseline-09-02\" target=\"_blank\">https://www.kaggle.com/lucabergamini/lyft-baseline-09-02</a></li>\n<li><a href=\"https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format</a></li>\n</ul>\n<p>Because of that I wanted to join this competition also at a late stage for the purpose of learning. However, after spending ~3-4 hours reading the links above and some example notebooks, I still have trouble understanding the data and the exact problem setting. Because of that, I will now probably focus on other competitions… Anyhow, those are my current most important questions:</p>\n<ul>\n<li><p>In the baseline model as well as in the notebook \"pytorch-baseline-train\" by <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>, the model input is an image with 25 \"channels\". What does each channel represent, in particular the last three? I believed it to be a boolean mask, but if I'm not mistaken it's somehow encoded uint8 values? Github only states that \"image\" in EgoDataset represents \"The BEV raster as a multi-channel tensor\". Since this seems to be the most important input, I find it strange to have such poor documentation, or am I missing something?</p></li>\n<li><p>In the Q&amp;A, it is stated that \"you will have to predict multiple agents in a given frame\". Yet, the baseline model only seems to output a 50x2 vector representing the trajectory of ONE vehicle?!</p></li>\n<li><p>Each item of the EgoDataset contains not only images, but also e.g. the history positions and yaws for the egovehicle. Can those be used as well? Or more generally: Which of the keys are meant to be known during test time? (On a similar note, the \"test.zarr\" also seems to contain non-empty values for the target positions, which I though we were meant to predict?)</p></li>\n</ul>\n<p>Again, thank you very much for hosting this event. Personally I just feel that the entrance hurdle seems quite big. Or I missed some crucial information pages… </p>",
  "messages": [
    {
      "id": "1066312",
      "postDate": "11/01/2020 15:52:12",
      "content": "<p>This post is meant to be both a question as well as a feedback to the organizers. I really do appreciate posting this exciting problem, sharing the data and the following threads from the organizer:</p>\n<ul>\n<li><a href=\"https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format</a></li>\n<li><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/179246\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/179246</a></li>\n<li><a href=\"https://www.kaggle.com/lucabergamini/lyft-baseline-09-02\" target=\"_blank\">https://www.kaggle.com/lucabergamini/lyft-baseline-09-02</a></li>\n<li><a href=\"https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format</a></li>\n</ul>\n<p>Because of that I wanted to join this competition also at a late stage for the purpose of learning. However, after spending ~3-4 hours reading the links above and some example notebooks, I still have trouble understanding the data and the exact problem setting. Because of that, I will now probably focus on other competitions… Anyhow, those are my current most important questions:</p>\n<ul>\n<li><p>In the baseline model as well as in the notebook \"pytorch-baseline-train\" by <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>, the model input is an image with 25 \"channels\". What does each channel represent, in particular the last three? I believed it to be a boolean mask, but if I'm not mistaken it's somehow encoded uint8 values? Github only states that \"image\" in EgoDataset represents \"The BEV raster as a multi-channel tensor\". Since this seems to be the most important input, I find it strange to have such poor documentation, or am I missing something?</p></li>\n<li><p>In the Q&amp;A, it is stated that \"you will have to predict multiple agents in a given frame\". Yet, the baseline model only seems to output a 50x2 vector representing the trajectory of ONE vehicle?!</p></li>\n<li><p>Each item of the EgoDataset contains not only images, but also e.g. the history positions and yaws for the egovehicle. Can those be used as well? Or more generally: Which of the keys are meant to be known during test time? (On a similar note, the \"test.zarr\" also seems to contain non-empty values for the target positions, which I though we were meant to predict?)</p></li>\n</ul>\n<p>Again, thank you very much for hosting this event. Personally I just feel that the entrance hurdle seems quite big. Or I missed some crucial information pages… </p>",
      "rawMarkdown": "This post is meant to be both a question as well as a feedback to the organizers. I really do appreciate posting this exciting problem, sharing the data and the following threads from the organizer:\n- https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\n- https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/179246\n- https://www.kaggle.com/lucabergamini/lyft-baseline-09-02\n- https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\n\nBecause of that I wanted to join this competition also at a late stage for the purpose of learning. However, after spending ~3-4 hours reading the links above and some example notebooks, I still have trouble understanding the data and the exact problem setting. Because of that, I will now probably focus on other competitions... Anyhow, those are my current most important questions:\n\n- In the baseline model as well as in the notebook \"pytorch-baseline-train\" by @pestipeti, the model input is an image with 25 \"channels\". What does each channel represent, in particular the last three? I believed it to be a boolean mask, but if I'm not mistaken it's somehow encoded uint8 values? Github only states that \"image\" in EgoDataset represents \"The BEV raster as a multi-channel tensor\". Since this seems to be the most important input, I find it strange to have such poor documentation, or am I missing something?\n\n- In the Q&A, it is stated that \"you will have to predict multiple agents in a given frame\". Yet, the baseline model only seems to output a 50x2 vector representing the trajectory of ONE vehicle?!\n\n- Each item of the EgoDataset contains not only images, but also e.g. the history positions and yaws for the egovehicle. Can those be used as well? Or more generally: Which of the keys are meant to be known during test time? (On a similar note, the \"test.zarr\" also seems to contain non-empty values for the target positions, which I though we were meant to predict?)\n\nAgain, thank you very much for hosting this event. Personally I just feel that the entrance hurdle seems quite big. Or I missed some crucial information pages...",
      "votes": null
    },
    {
      "id": "1066356",
      "postDate": "11/01/2020 16:59:54",
      "content": "<ol>\n<li>Take a look at <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187366#1030547\" target=\"_blank\">this pic</a> for an example of channels visualization.</li>\n<li>Yes, the prediction is for one agent at a time, but many of these agents share same frame. It can be about 5-10 predictions (different agents) that we need to make per frame. There are about 16k frames in the test set if I remember correctly, with about 70k agents.</li>\n<li>Target positions, target availability and target yaws are zero in the test set (a placeholder), target positions is what needs to be predicted. The rest you can use.</li>\n</ol>",
      "rawMarkdown": "1. Take a look at [this pic](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187366#1030547) for an example of channels visualization.\n2. Yes, the prediction is for one agent at a time, but many of these agents share same frame. It can be about 5-10 predictions (different agents) that we need to make per frame. There are about 16k frames in the test set if I remember correctly, with about 70k agents.\n3. Target positions, target availability and target yaws are zero in the test set (a placeholder), target positions is what needs to be predicted. The rest you can use.",
      "votes": null
    },
    {
      "id": "1066484",
      "postDate": "11/01/2020 20:20:56",
      "content": "<p>That image is really helpful <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>, thank you.</p>\n<p>It looks like there is a lot of redundant info  - especially for  the \"agent under consideration\". <br>\nDo you think we would lose much info by setting history_step_size =2, effectively making it 15 channels ?</p>",
      "rawMarkdown": "That image is really helpful @zaharch, thank you.\n\nIt looks like there is a lot of redundant info  - especially for  the \"agent under consideration\". \nDo you think we would lose much info by setting history_step_size =2, effectively making it 15 channels ?",
      "votes": null
    },
    {
      "id": "1066488",
      "postDate": "11/01/2020 20:30:18",
      "content": "<p>Probably not much loss, the frames are separated by only 0.1s. But I haven't tried it, so can't say for sure.</p>",
      "rawMarkdown": "Probably not much loss, the frames are separated by only 0.1s. But I haven't tried it, so can't say for sure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1066356,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "11/01/2020 16:59:54",
      "content": "<ol>\n<li>Take a look at <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187366#1030547\" target=\"_blank\">this pic</a> for an example of channels visualization.</li>\n<li>Yes, the prediction is for one agent at a time, but many of these agents share same frame. It can be about 5-10 predictions (different agents) that we need to make per frame. There are about 16k frames in the test set if I remember correctly, with about 70k agents.</li>\n<li>Target positions, target availability and target yaws are zero in the test set (a placeholder), target positions is what needs to be predicted. The rest you can use.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1066484,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "11/01/2020 20:20:56",
          "content": "<p>That image is really helpful <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>, thank you.</p>\n<p>It looks like there is a lot of redundant info  - especially for  the \"agent under consideration\". <br>\nDo you think we would lose much info by setting history_step_size =2, effectively making it 15 channels ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066488,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/01/2020 20:30:18",
          "content": "<p>Probably not much loss, the frames are separated by only 0.1s. But I haven't tried it, so can't say for sure.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1066312": "This post is meant to be both a question as well as a feedback to the organizers. I really do appreciate posting this exciting problem, sharing the data and the following threads from the organizer:\n- https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\n- https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/179246\n- https://www.kaggle.com/lucabergamini/lyft-baseline-09-02\n- https://github.com/lyft/l5kit/blob/master/data_format.md#2020-lyft-competition-dataset-format\n\nBecause of that I wanted to join this competition also at a late stage for the purpose of learning. However, after spending ~3-4 hours reading the links above and some example notebooks, I still have trouble understanding the data and the exact problem setting. Because of that, I will now probably focus on other competitions... Anyhow, those are my current most important questions:\n\n- In the baseline model as well as in the notebook \"pytorch-baseline-train\" by @pestipeti, the model input is an image with 25 \"channels\". What does each channel represent, in particular the last three? I believed it to be a boolean mask, but if I'm not mistaken it's somehow encoded uint8 values? Github only states that \"image\" in EgoDataset represents \"The BEV raster as a multi-channel tensor\". Since this seems to be the most important input, I find it strange to have such poor documentation, or am I missing something?\n\n- In the Q&A, it is stated that \"you will have to predict multiple agents in a given frame\". Yet, the baseline model only seems to output a 50x2 vector representing the trajectory of ONE vehicle?!\n\n- Each item of the EgoDataset contains not only images, but also e.g. the history positions and yaws for the egovehicle. Can those be used as well? Or more generally: Which of the keys are meant to be known during test time? (On a similar note, the \"test.zarr\" also seems to contain non-empty values for the target positions, which I though we were meant to predict?)\n\nAgain, thank you very much for hosting this event. Personally I just feel that the entrance hurdle seems quite big. Or I missed some crucial information pages...",
    "1066356": "1. Take a look at [this pic](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187366#1030547) for an example of channels visualization.\n2. Yes, the prediction is for one agent at a time, but many of these agents share same frame. It can be about 5-10 predictions (different agents) that we need to make per frame. There are about 16k frames in the test set if I remember correctly, with about 70k agents.\n3. Target positions, target availability and target yaws are zero in the test set (a placeholder), target positions is what needs to be predicted. The rest you can use.",
    "1066484": "That image is really helpful @zaharch, thank you.\n\nIt looks like there is a lot of redundant info  - especially for  the \"agent under consideration\". \nDo you think we would lose much info by setting history_step_size =2, effectively making it 15 channels ?",
    "1066488": "Probably not much loss, the frames are separated by only 0.1s. But I haven't tried it, so can't say for sure."
  },
  "source": "meta"
}