{
  "id": 183915,
  "title": "What are contained in 'agents_mask'? How to read it ?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/183915",
  "author_name": "",
  "post_date": "2020-09-18T14:30:33.585095800Z",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I think this file can help me to understand which columns should be predicted in the output, but I have no idea to deal with it.</p>",
  "messages": [
    {
      "id": "1015930",
      "postDate": "09/18/2020 14:30:33",
      "content": "<p>I think this file can help me to understand which columns should be predicted in the output, but I have no idea to deal with it.</p>",
      "rawMarkdown": "I think this file can help me to understand which columns should be predicted in the output, but I have no idea to deal with it.",
      "votes": null
    },
    {
      "id": "1015964",
      "postDate": "09/18/2020 15:10:08",
      "content": "<p><code>agents_mask</code> is a 1D boolean array with the same length of the agent array. If a value is True, that agent will be considered as an element of the dataset, otherwise it will be skipped.</p>",
      "rawMarkdown": "`agents_mask` is a 1D boolean array with the same length of the agent array. If a value is True, that agent will be considered as an element of the dataset, otherwise it will be skipped.",
      "votes": null
    },
    {
      "id": "1020584",
      "postDate": "09/21/2020 09:25:28",
      "content": "<p>I see. Thanks a lot!</p>",
      "rawMarkdown": "I see. Thanks a lot!",
      "votes": null
    },
    {
      "id": "1025165",
      "postDate": "09/24/2020 11:31:39",
      "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a><br>\nIsn't <code>agents_mask</code> a 2D array? In function <code>load_agents_mask</code> in <code>agent.py</code> file, it returns a <code>numpy.ndarray</code> with shape <code>(1893736, 2)</code> for <code>sample.zarr</code> dataset:</p>\n<pre><code>def load_agents_mask(self) -&gt; np.ndarray:\n    \"\"\"\n    Loads a boolean mask of the agent availability stored into the zarr. Performs some sanity check against cfg.\n    Returns: a boolean mask of the same length of the dataset agents\n    \"\"\"\n\n    ...\n    ...\n    agents_mask = convenience.load(str(agents_mask_path))  # This returns `(1893736, 2)` for sample.zarr\n    return agents_mask\n</code></pre>\n<p>Later on in the <code>__init__</code> function of <code>class AgentDataset</code>, the <code>agents_mask</code> is compared against <code>min_frame_history</code> and <code>min_frame_future</code> as follows:</p>\n<pre><code>past_mask = agents_mask[:, 0] &gt;= min_frame_history\nfuture_mask = agents_mask[:, 1] &gt;= min_frame_future\nagents_mask = past_mask * future_mask\n</code></pre>\n<p>I wonder what those two columns of <code>agents_mask</code> correspond to?</p>",
      "rawMarkdown": "lucabergamini\nIsn't `agents_mask` a 2D array? In function `load_agents_mask` in `agent.py` file, it returns a `numpy.ndarray` with shape `(1893736, 2)` for `sample.zarr` dataset:\n\n```python\ndef load_agents_mask(self) -> np.ndarray:\n    \"\"\"\n    Loads a boolean mask of the agent availability stored into the zarr. Performs some sanity check against cfg.\n    Returns: a boolean mask of the same length of the dataset agents\n    \"\"\"\n    \n    ...\n    ...\n    agents_mask = convenience.load(str(agents_mask_path))  # This returns `(1893736, 2)` for sample.zarr\n    return agents_mask\n```\n\nLater on in the `__init__` function of `class AgentDataset`, the `agents_mask` is compared against `min_frame_history` and `min_frame_future` as follows:\n\n```python\npast_mask = agents_mask[:, 0] >= min_frame_history\nfuture_mask = agents_mask[:, 1] >= min_frame_future\nagents_mask = past_mask * future_mask\n```\n\nI wonder what those two columns of `agents_mask` correspond to?",
      "votes": null
    },
    {
      "id": "1025240",
      "postDate": "09/24/2020 12:42:26",
      "content": "<p>You are correct that it is a little bit confusing, <code>agents_mask</code> is different things depending on its shape. If it is 1D boolean vector then what <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> said. If it is a 2D integer vector then it specifies in how many frames agent is relevant before and after the sample in question (\"relevant\" here is my word). Note that relevant agent frames is not the same as frames where agent is available, read <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762#1022114\" target=\"_blank\">this discussion</a> for explanation of the difference.</p>",
      "rawMarkdown": "You are correct that it is a little bit confusing, `agents_mask` is different things depending on its shape. If it is 1D boolean vector then what @lucabergamini said. If it is a 2D integer vector then it specifies in how many frames agent is relevant before and after the sample in question (\"relevant\" here is my word). Note that relevant agent frames is not the same as frames where agent is available, read [this discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762#1022114) for explanation of the difference.",
      "votes": null
    },
    {
      "id": "1025303",
      "postDate": "09/24/2020 13:16:43",
      "content": "<p>Yeah, there's two different stored masks. The <code>agents_mask</code> array in the zarr datasets which is the 2D format and is used internally in <code>AgentDataset</code> to filer frames (and create an internal 1D mask, affected by certain config options). You don't need to manually load this or generally deal with it.<br>\nThen there's the <code>mask.npz</code> file you load (with <code>numpy.load</code>) and pass in to <code>AgentDataset</code> which is 1D boolean mask with a pre-specified set of frames for which you must make test predictions.</p>",
      "rawMarkdown": "Yeah, there's two different stored masks. The `agents_mask` array in the zarr datasets which is the 2D format and is used internally in `AgentDataset` to filer frames (and create an internal 1D mask, affected by certain config options). You don't need to manually load this or generally deal with it.\nThen there's the `mask.npz` file you load (with `numpy.load`) and pass in to `AgentDataset` which is 1D boolean mask with a pre-specified set of frames for which you must make test predictions.",
      "votes": null
    },
    {
      "id": "1025331",
      "postDate": "09/24/2020 13:35:30",
      "content": "<p>Yes, sorry for the confusion. The mask stored on the disk has future and history availabilities (hence 2D) while the one used at runtime is a 1D boolean array. The first is converted into the second one at runtime <a href=\"https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L38\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Yes, sorry for the confusion. The mask stored on the disk has future and history availabilities (hence 2D) while the one used at runtime is a 1D boolean array. The first is converted into the second one at runtime [here](https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L38)",
      "votes": null
    },
    {
      "id": "1025388",
      "postDate": "09/24/2020 14:20:44",
      "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> , don't you introduce another confusion here by calling it \"availabilities\"? Isn't it availability + the additional checks?</p>",
      "rawMarkdown": "lucabergamini , don't you introduce another confusion here by calling it \"availabilities\"? Isn't it availability + the additional checks?",
      "votes": null
    },
    {
      "id": "1025392",
      "postDate": "09/24/2020 14:26:34",
      "content": "<p>depends what you mean with \"availabilities\" I guess :) Originally the array stored on disk was also 1D and there were no checks there..</p>",
      "rawMarkdown": "depends what you mean with \"availabilities\" I guess :) Originally the array stored on disk was also 1D and there were no checks there..",
      "votes": null
    },
    {
      "id": "1025436",
      "postDate": "09/24/2020 14:48:45",
      "content": "<p>The fields <code>target_availabilities</code> and <code>history_availabilities</code> which are generated by <code>_create_targets_for_deep_prediction</code> and based solely on presence in a frame in zarr arrays is what naturally we call \"availability\". </p>\n<p>On the other hand, mask values are subset of those. To avoid confusion I wouldn't call masks by name \"availability\", it is a reserved word already. Does it make sense? :)</p>",
      "rawMarkdown": "The fields `target_availabilities` and `history_availabilities` which are generated by `_create_targets_for_deep_prediction` and based solely on presence in a frame in zarr arrays is what naturally we call \"availability\". \n\nOn the other hand, mask values are subset of those. To avoid confusion I wouldn't call masks by name \"availability\", it is a reserved word already. Does it make sense? :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1015964,
      "author_name": "lucabergamini",
      "author_url": "",
      "post_date": "09/18/2020 15:10:08",
      "content": "<p><code>agents_mask</code> is a 1D boolean array with the same length of the agent array. If a value is True, that agent will be considered as an element of the dataset, otherwise it will be skipped.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1020584,
      "author_name": "pythonql93",
      "author_url": "",
      "post_date": "09/21/2020 09:25:28",
      "content": "<p>I see. Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1025165,
      "author_name": "walkpast",
      "author_url": "",
      "post_date": "09/24/2020 11:31:39",
      "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a><br>\nIsn't <code>agents_mask</code> a 2D array? In function <code>load_agents_mask</code> in <code>agent.py</code> file, it returns a <code>numpy.ndarray</code> with shape <code>(1893736, 2)</code> for <code>sample.zarr</code> dataset:</p>\n<pre><code>def load_agents_mask(self) -&gt; np.ndarray:\n    \"\"\"\n    Loads a boolean mask of the agent availability stored into the zarr. Performs some sanity check against cfg.\n    Returns: a boolean mask of the same length of the dataset agents\n    \"\"\"\n\n    ...\n    ...\n    agents_mask = convenience.load(str(agents_mask_path))  # This returns `(1893736, 2)` for sample.zarr\n    return agents_mask\n</code></pre>\n<p>Later on in the <code>__init__</code> function of <code>class AgentDataset</code>, the <code>agents_mask</code> is compared against <code>min_frame_history</code> and <code>min_frame_future</code> as follows:</p>\n<pre><code>past_mask = agents_mask[:, 0] &gt;= min_frame_history\nfuture_mask = agents_mask[:, 1] &gt;= min_frame_future\nagents_mask = past_mask * future_mask\n</code></pre>\n<p>I wonder what those two columns of <code>agents_mask</code> correspond to?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1025240,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/24/2020 12:42:26",
          "content": "<p>You are correct that it is a little bit confusing, <code>agents_mask</code> is different things depending on its shape. If it is 1D boolean vector then what <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> said. If it is a 2D integer vector then it specifies in how many frames agent is relevant before and after the sample in question (\"relevant\" here is my word). Note that relevant agent frames is not the same as frames where agent is available, read <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762#1022114\" target=\"_blank\">this discussion</a> for explanation of the difference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025303,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "09/24/2020 13:16:43",
          "content": "<p>Yeah, there's two different stored masks. The <code>agents_mask</code> array in the zarr datasets which is the 2D format and is used internally in <code>AgentDataset</code> to filer frames (and create an internal 1D mask, affected by certain config options). You don't need to manually load this or generally deal with it.<br>\nThen there's the <code>mask.npz</code> file you load (with <code>numpy.load</code>) and pass in to <code>AgentDataset</code> which is 1D boolean mask with a pre-specified set of frames for which you must make test predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025331,
          "author_name": "lucabergamini",
          "author_url": "",
          "post_date": "09/24/2020 13:35:30",
          "content": "<p>Yes, sorry for the confusion. The mask stored on the disk has future and history availabilities (hence 2D) while the one used at runtime is a 1D boolean array. The first is converted into the second one at runtime <a href=\"https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L38\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025388,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/24/2020 14:20:44",
          "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> , don't you introduce another confusion here by calling it \"availabilities\"? Isn't it availability + the additional checks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025392,
          "author_name": "lucabergamini",
          "author_url": "",
          "post_date": "09/24/2020 14:26:34",
          "content": "<p>depends what you mean with \"availabilities\" I guess :) Originally the array stored on disk was also 1D and there were no checks there..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025436,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/24/2020 14:48:45",
          "content": "<p>The fields <code>target_availabilities</code> and <code>history_availabilities</code> which are generated by <code>_create_targets_for_deep_prediction</code> and based solely on presence in a frame in zarr arrays is what naturally we call \"availability\". </p>\n<p>On the other hand, mask values are subset of those. To avoid confusion I wouldn't call masks by name \"availability\", it is a reserved word already. Does it make sense? :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1015930": "I think this file can help me to understand which columns should be predicted in the output, but I have no idea to deal with it.",
    "1015964": "`agents_mask` is a 1D boolean array with the same length of the agent array. If a value is True, that agent will be considered as an element of the dataset, otherwise it will be skipped.",
    "1020584": "I see. Thanks a lot!",
    "1025165": "lucabergamini\nIsn't `agents_mask` a 2D array? In function `load_agents_mask` in `agent.py` file, it returns a `numpy.ndarray` with shape `(1893736, 2)` for `sample.zarr` dataset:\n\n```python\ndef load_agents_mask(self) -> np.ndarray:\n    \"\"\"\n    Loads a boolean mask of the agent availability stored into the zarr. Performs some sanity check against cfg.\n    Returns: a boolean mask of the same length of the dataset agents\n    \"\"\"\n    \n    ...\n    ...\n    agents_mask = convenience.load(str(agents_mask_path))  # This returns `(1893736, 2)` for sample.zarr\n    return agents_mask\n```\n\nLater on in the `__init__` function of `class AgentDataset`, the `agents_mask` is compared against `min_frame_history` and `min_frame_future` as follows:\n\n```python\npast_mask = agents_mask[:, 0] >= min_frame_history\nfuture_mask = agents_mask[:, 1] >= min_frame_future\nagents_mask = past_mask * future_mask\n```\n\nI wonder what those two columns of `agents_mask` correspond to?",
    "1025240": "You are correct that it is a little bit confusing, `agents_mask` is different things depending on its shape. If it is 1D boolean vector then what @lucabergamini said. If it is a 2D integer vector then it specifies in how many frames agent is relevant before and after the sample in question (\"relevant\" here is my word). Note that relevant agent frames is not the same as frames where agent is available, read [this discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762#1022114) for explanation of the difference.",
    "1025303": "Yeah, there's two different stored masks. The `agents_mask` array in the zarr datasets which is the 2D format and is used internally in `AgentDataset` to filer frames (and create an internal 1D mask, affected by certain config options). You don't need to manually load this or generally deal with it.\nThen there's the `mask.npz` file you load (with `numpy.load`) and pass in to `AgentDataset` which is 1D boolean mask with a pre-specified set of frames for which you must make test predictions.",
    "1025331": "Yes, sorry for the confusion. The mask stored on the disk has future and history availabilities (hence 2D) while the one used at runtime is a 1D boolean array. The first is converted into the second one at runtime [here](https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L38)",
    "1025388": "lucabergamini , don't you introduce another confusion here by calling it \"availabilities\"? Isn't it availability + the additional checks?",
    "1025392": "depends what you mean with \"availabilities\" I guess :) Originally the array stored on disk was also 1D and there were no checks there..",
    "1025436": "The fields `target_availabilities` and `history_availabilities` which are generated by `_create_targets_for_deep_prediction` and based solely on presence in a frame in zarr arrays is what naturally we call \"availability\". \n\nOn the other hand, mask values are subset of those. To avoid confusion I wouldn't call masks by name \"availability\", it is a reserved word already. Does it make sense? :)"
  },
  "source": "meta"
}