{
  "id": 233734,
  "title": "Better Dataset Description",
  "url": "/competitions/cse151b-spring/discussion/233734",
  "author_name": "Justin Allen",
  "post_date": "2021-04-20T20:03:51.160000",
  "votes": 53,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The dataset descriptions for the Kaggle competition weren't very descriptive, so I figured I'd post what I've learned about the raw data to help any other students struggling with it. There aren't any statistics or anything showing how to build a good model, only descriptions on how the data is actually laid out in the pickle files, and how it interconnects.</p>\n<p>Note: this is all from my own experimentation that could very well be wrong in some cases, so I encourage you to check out the data yourself for a bit to confirm what I'm saying is true, and reply if there is anything wrong.</p>\n<p>Raw data description:</p>\n<p>The raw data is in two folders: one for training, and one for validation. Inside each folder is many pickle (.pkl) files. Each pickle file is a dictionary containing information on one individual scene.</p>\n<p>Each scene consists of the following information:<br>\n    - Positions of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.<br>\n    - Velocities of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.<br>\n    - An ID for each vehicle.<br>\n    - The ID of the vehicle that we are tracking (the 'agent' vehicle).<br>\n    - Information on where the lanes are. These are positions of the centers of the lanes, and what direction those lanes are pointing (the lane normals). There are an arbitrary number of lanes for each scene.</p>\n<p>Each pickle file contains a dictionary with the following fields:</p>\n<ul>\n<li><p>'city': a string for the city the data was taken from. Can either be 'PIT' or 'MIA'</p></li>\n<li><p>'scene_idx': a unique non-zero integer acting as an ID for the current scene (unique across the union of training and validation sets)</p></li>\n<li><p>'agent_id': a string of the form \"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits that identifies the 'agent' we want to track (predict future movement of) in this scene. This ID is always 36 characters long (32 digits if you don't include the -'s)</p></li>\n<li><p>'car_mask': a (60, 1) numpy array of float 1's and 0's. Each row corresponds to the row index of each car in the scene (IE: this matches with the track_id, p_in, p_out, v_in, and v_out first dimensions). There is a '1' in the row if the values in the first dimension of track_id, p_in, etc. correspond to an actual object in the scene, or are just empty because there were less than 60 objects in the scene. The total number of 1's in this mask is the total number of objects in the scene, including the agent car we are tracking. This allows us to quickly get all of the values in track_id, p_in, etc. that correspond to actual objects like so:</p>\n<pre><code>actual_objects = track_id[car_mask.reshape([-1]).astype(int)]\n\nNOTE: the values in the numpy array are floats, and thus need to be converted to integers before using them as a mask (hence the .astype(int))\n</code></pre>\n<p>The car_mask is laid out such that the only 1's occur at the beginning of the array. So, in the event you had 10 objects in the scene: car_mask[:10, 0] == 1, and car_mask[10:, 0] == 0. This means all of the actually tracked objects in each of track_id, p_in, etc come first, and the rest of the values are dummy values.</p></li>\n<li><p>'track_id': a (60, 30, 1) numpy array of strings. Each string is of the form <br>\n\"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits. This describes a unique id (unique to the scene, not all of the data) for each object in the scene. The whole array is 60 \"30 by 1 column vectors\" where every one of the 30 elements in the column vector are the same track_id string. Don't ask me why this is the case, I have no idea. They could have just as easily made a 1-D list of 60 track_ids and accomplished the same thing. Honestly, whoever made this was probably just too lazy to change it, and I don't blame them; I'd probably do the same thing.<br>\nThis array follows the same indexing as the 'car_mask'. So, the first n 1's in the car mask array describe the track_ids for all of the actually tracked objects in the scene. After the first n actually tracked objects, the track_id changes to 'dummyK' where K is the integer index of the dummy variable starting at 0 (IE: 'dummy0', 'dummy1', …, 'dummy[60-n-1]').</p></li>\n<li><p>'p_in': a (60, 19, 2) numpy array of floats. These are the x,y coordinates of each tracked object in the scene, for a max of 60 objects. What those coordinates are relative to I do not know (probably some longitude/latitude in or near whatever city the data is from), so you may not want to use the exact values in your project, and instead care only about the change in position at each time step. There are 19 samples per object meaning the data is probably only for 1.9 seconds, not 2 seconds exactly. This array follows the same indexing as car_mask, so the first n arrays in p_in correspond to the n tracked objects in the scene, and the following 60-n arrays are filled with all 0's.<br>\n    NOTE: some values may be 0 or negative even if we are tracking them, indicating the object is at the very edge of the coordinate system</p></li>\n<li><p>'v_in': same as 'p_in', but for tracking velocities instead of position. These values can be positive, negative, or 0. Again, I don't know what the units are. It might be on the Argoverse website somewhere?<br>\n    NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)</p></li>\n<li><p>'p_out': same as 'p_in', but is instead of shape (60, 30, 2). These are the positions your model should learn to predict.<br>\n    NOTE: the validation sets do not contain this key as this is what your model should predict.</p></li>\n<li><p>'v_out': same as 'v_in', but is instead of shape (60, 30, 2). You do not need to predict these for the final project. You might want to discard them, unless you want to instead have your model predict velocity instead of position, and you can calculate the final positions by adding some small multiple of the velocity?<br>\n    NOTE: the validation sets do not contain this key<br>\n    NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)</p></li>\n<li><p>'lane': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene. There can be a different number of lanes for each scene, so there is no guarantee on the size of 'k'. These describe the x,y,z coordinates of the center of lane nodes (where driving lanes are in the scene). For some reason, the z-coordinate is included, but is always 0, so you can ignore it and just look at the x,y. These seem to be in the same coordinate system as the p_in and p_out coordinates. Perhaps you want to change these to be relative to something else in the scene?<br>\n    NOTE: some x,y values may be 0 or negative indicating the center of the lane is at the very edge of the coordinate system</p></li>\n<li><p>'lane_norm': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene (same size as the 'k' in 'lane'). These describe the x,y,z vector direction of the corresponding lane normal (the direction the lane center with the same index is pointing). These values can be positive or negative, but never 0 (for x,y). Again, the z direction is always 0 and can be ignored, and the coordinate system seems to be the same as p_in.</p></li>\n</ul>",
  "messages": [
    {
      "id": 1279365,
      "postDate": "2021-04-20T20:03:51.160Z",
      "content": "<p>The dataset descriptions for the Kaggle competition weren't very descriptive, so I figured I'd post what I've learned about the raw data to help any other students struggling with it. There aren't any statistics or anything showing how to build a good model, only descriptions on how the data is actually laid out in the pickle files, and how it interconnects.</p>\n<p>Note: this is all from my own experimentation that could very well be wrong in some cases, so I encourage you to check out the data yourself for a bit to confirm what I'm saying is true, and reply if there is anything wrong.</p>\n<p>Raw data description:</p>\n<p>The raw data is in two folders: one for training, and one for validation. Inside each folder is many pickle (.pkl) files. Each pickle file is a dictionary containing information on one individual scene.</p>\n<p>Each scene consists of the following information:<br>\n    - Positions of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.<br>\n    - Velocities of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.<br>\n    - An ID for each vehicle.<br>\n    - The ID of the vehicle that we are tracking (the 'agent' vehicle).<br>\n    - Information on where the lanes are. These are positions of the centers of the lanes, and what direction those lanes are pointing (the lane normals). There are an arbitrary number of lanes for each scene.</p>\n<p>Each pickle file contains a dictionary with the following fields:</p>\n<ul>\n<li><p>'city': a string for the city the data was taken from. Can either be 'PIT' or 'MIA'</p></li>\n<li><p>'scene_idx': a unique non-zero integer acting as an ID for the current scene (unique across the union of training and validation sets)</p></li>\n<li><p>'agent_id': a string of the form \"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits that identifies the 'agent' we want to track (predict future movement of) in this scene. This ID is always 36 characters long (32 digits if you don't include the -'s)</p></li>\n<li><p>'car_mask': a (60, 1) numpy array of float 1's and 0's. Each row corresponds to the row index of each car in the scene (IE: this matches with the track_id, p_in, p_out, v_in, and v_out first dimensions). There is a '1' in the row if the values in the first dimension of track_id, p_in, etc. correspond to an actual object in the scene, or are just empty because there were less than 60 objects in the scene. The total number of 1's in this mask is the total number of objects in the scene, including the agent car we are tracking. This allows us to quickly get all of the values in track_id, p_in, etc. that correspond to actual objects like so:</p>\n<pre><code>actual_objects = track_id[car_mask.reshape([-1]).astype(int)]\n\nNOTE: the values in the numpy array are floats, and thus need to be converted to integers before using them as a mask (hence the .astype(int))\n</code></pre>\n<p>The car_mask is laid out such that the only 1's occur at the beginning of the array. So, in the event you had 10 objects in the scene: car_mask[:10, 0] == 1, and car_mask[10:, 0] == 0. This means all of the actually tracked objects in each of track_id, p_in, etc come first, and the rest of the values are dummy values.</p></li>\n<li><p>'track_id': a (60, 30, 1) numpy array of strings. Each string is of the form <br>\n\"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits. This describes a unique id (unique to the scene, not all of the data) for each object in the scene. The whole array is 60 \"30 by 1 column vectors\" where every one of the 30 elements in the column vector are the same track_id string. Don't ask me why this is the case, I have no idea. They could have just as easily made a 1-D list of 60 track_ids and accomplished the same thing. Honestly, whoever made this was probably just too lazy to change it, and I don't blame them; I'd probably do the same thing.<br>\nThis array follows the same indexing as the 'car_mask'. So, the first n 1's in the car mask array describe the track_ids for all of the actually tracked objects in the scene. After the first n actually tracked objects, the track_id changes to 'dummyK' where K is the integer index of the dummy variable starting at 0 (IE: 'dummy0', 'dummy1', …, 'dummy[60-n-1]').</p></li>\n<li><p>'p_in': a (60, 19, 2) numpy array of floats. These are the x,y coordinates of each tracked object in the scene, for a max of 60 objects. What those coordinates are relative to I do not know (probably some longitude/latitude in or near whatever city the data is from), so you may not want to use the exact values in your project, and instead care only about the change in position at each time step. There are 19 samples per object meaning the data is probably only for 1.9 seconds, not 2 seconds exactly. This array follows the same indexing as car_mask, so the first n arrays in p_in correspond to the n tracked objects in the scene, and the following 60-n arrays are filled with all 0's.<br>\n    NOTE: some values may be 0 or negative even if we are tracking them, indicating the object is at the very edge of the coordinate system</p></li>\n<li><p>'v_in': same as 'p_in', but for tracking velocities instead of position. These values can be positive, negative, or 0. Again, I don't know what the units are. It might be on the Argoverse website somewhere?<br>\n    NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)</p></li>\n<li><p>'p_out': same as 'p_in', but is instead of shape (60, 30, 2). These are the positions your model should learn to predict.<br>\n    NOTE: the validation sets do not contain this key as this is what your model should predict.</p></li>\n<li><p>'v_out': same as 'v_in', but is instead of shape (60, 30, 2). You do not need to predict these for the final project. You might want to discard them, unless you want to instead have your model predict velocity instead of position, and you can calculate the final positions by adding some small multiple of the velocity?<br>\n    NOTE: the validation sets do not contain this key<br>\n    NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)</p></li>\n<li><p>'lane': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene. There can be a different number of lanes for each scene, so there is no guarantee on the size of 'k'. These describe the x,y,z coordinates of the center of lane nodes (where driving lanes are in the scene). For some reason, the z-coordinate is included, but is always 0, so you can ignore it and just look at the x,y. These seem to be in the same coordinate system as the p_in and p_out coordinates. Perhaps you want to change these to be relative to something else in the scene?<br>\n    NOTE: some x,y values may be 0 or negative indicating the center of the lane is at the very edge of the coordinate system</p></li>\n<li><p>'lane_norm': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene (same size as the 'k' in 'lane'). These describe the x,y,z vector direction of the corresponding lane normal (the direction the lane center with the same index is pointing). These values can be positive or negative, but never 0 (for x,y). Again, the z direction is always 0 and can be ignored, and the coordinate system seems to be the same as p_in.</p></li>\n</ul>",
      "rawMarkdown": "The dataset descriptions for the Kaggle competition weren't very descriptive, so I figured I'd post what I've learned about the raw data to help any other students struggling with it. There aren't any statistics or anything showing how to build a good model, only descriptions on how the data is actually laid out in the pickle files, and how it interconnects.\n\nNote: this is all from my own experimentation that could very well be wrong in some cases, so I encourage you to check out the data yourself for a bit to confirm what I'm saying is true, and reply if there is anything wrong.\n\nRaw data description:\n\nThe raw data is in two folders: one for training, and one for validation. Inside each folder is many pickle (.pkl) files. Each pickle file is a dictionary containing information on one individual scene.\n\nEach scene consists of the following information:\n    - Positions of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.\n    - Velocities of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.\n    - An ID for each vehicle.\n    - The ID of the vehicle that we are tracking (the 'agent' vehicle).\n    - Information on where the lanes are. These are positions of the centers of the lanes, and what direction those lanes are pointing (the lane normals). There are an arbitrary number of lanes for each scene.\n\nEach pickle file contains a dictionary with the following fields:\n- 'city': a string for the city the data was taken from. Can either be 'PIT' or 'MIA'\n\n- 'scene_idx': a unique non-zero integer acting as an ID for the current scene (unique across the union of training and validation sets)\n\n- 'agent_id': a string of the form \"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits that identifies the 'agent' we want to track (predict future movement of) in this scene. This ID is always 36 characters long (32 digits if you don't include the -'s)\n\n- 'car_mask': a (60, 1) numpy array of float 1's and 0's. Each row corresponds to the row index of each car in the scene (IE: this matches with the track_id, p_in, p_out, v_in, and v_out first dimensions). There is a '1' in the row if the values in the first dimension of track_id, p_in, etc. correspond to an actual object in the scene, or are just empty because there were less than 60 objects in the scene. The total number of 1's in this mask is the total number of objects in the scene, including the agent car we are tracking. This allows us to quickly get all of the values in track_id, p_in, etc. that correspond to actual objects like so:\n\n        actual_objects = track_id[car_mask.reshape([-1]).astype(int)]\n\n        NOTE: the values in the numpy array are floats, and thus need to be converted to integers before using them as a mask (hence the .astype(int))\n\n    The car_mask is laid out such that the only 1's occur at the beginning of the array. So, in the event you had 10 objects in the scene: car_mask[:10, 0] == 1, and car_mask[10:, 0] == 0. This means all of the actually tracked objects in each of track_id, p_in, etc come first, and the rest of the values are dummy values.\n\n- 'track_id': a (60, 30, 1) numpy array of strings. Each string is of the form \n\"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits. This describes a unique id (unique to the scene, not all of the data) for each object in the scene. The whole array is 60 \"30 by 1 column vectors\" where every one of the 30 elements in the column vector are the same track_id string. Don't ask me why this is the case, I have no idea. They could have just as easily made a 1-D list of 60 track_ids and accomplished the same thing. Honestly, whoever made this was probably just too lazy to change it, and I don't blame them; I'd probably do the same thing.\n    This array follows the same indexing as the 'car_mask'. So, the first n 1's in the car mask array describe the track_ids for all of the actually tracked objects in the scene. After the first n actually tracked objects, the track_id changes to 'dummyK' where K is the integer index of the dummy variable starting at 0 (IE: 'dummy0', 'dummy1', ..., 'dummy[60-n-1]').\n\n- 'p_in': a (60, 19, 2) numpy array of floats. These are the x,y coordinates of each tracked object in the scene, for a max of 60 objects. What those coordinates are relative to I do not know (probably some longitude/latitude in or near whatever city the data is from), so you may not want to use the exact values in your project, and instead care only about the change in position at each time step. There are 19 samples per object meaning the data is probably only for 1.9 seconds, not 2 seconds exactly. This array follows the same indexing as car_mask, so the first n arrays in p_in correspond to the n tracked objects in the scene, and the following 60-n arrays are filled with all 0's.\n        NOTE: some values may be 0 or negative even if we are tracking them, indicating the object is at the very edge of the coordinate system\n\n- 'v_in': same as 'p_in', but for tracking velocities instead of position. These values can be positive, negative, or 0. Again, I don't know what the units are. It might be on the Argoverse website somewhere?\n        NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)\n\n- 'p_out': same as 'p_in', but is instead of shape (60, 30, 2). These are the positions your model should learn to predict.\n        NOTE: the validation sets do not contain this key as this is what your model should predict.\n\n- 'v_out': same as 'v_in', but is instead of shape (60, 30, 2). You do not need to predict these for the final project. You might want to discard them, unless you want to instead have your model predict velocity instead of position, and you can calculate the final positions by adding some small multiple of the velocity?\n        NOTE: the validation sets do not contain this key\n        NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)\n\n- 'lane': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene. There can be a different number of lanes for each scene, so there is no guarantee on the size of 'k'. These describe the x,y,z coordinates of the center of lane nodes (where driving lanes are in the scene). For some reason, the z-coordinate is included, but is always 0, so you can ignore it and just look at the x,y. These seem to be in the same coordinate system as the p_in and p_out coordinates. Perhaps you want to change these to be relative to something else in the scene?\n        NOTE: some x,y values may be 0 or negative indicating the center of the lane is at the very edge of the coordinate system\n\n- 'lane_norm': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene (same size as the 'k' in 'lane'). These describe the x,y,z vector direction of the corresponding lane normal (the direction the lane center with the same index is pointing). These values can be positive or negative, but never 0 (for x,y). Again, the z direction is always 0 and can be ignored, and the coordinate system seems to be the same as p_in.",
      "votes": 53
    },
    {
      "id": 1315241,
      "postDate": "2021-05-19T16:47:13.480Z",
      "content": "<p>Hi. Really great work. But yesterday I was messing around with certain ways of preprocessing my data and I discovered that using <code>car_mask[i].reshape([-1]).astype(int)</code> (where <code>i</code> is designating the current scene in the batch) as a mask still resulted in a normal sized array. I.e. <code>p_in[car_mask[i].reshape([-1]).astype(int)]</code> would still result in an array of size (60, 19, 2) regardless of car_mask's contents. I was pretty sure that this was not the case earlier, but I guess I was wrong. Anyways, I think someone might find it useful to use car_mask as the following:</p>\n<p><code>mask= np.count_nonzero(car_mask[i])</code><br>\n<code>p_in_relevant = p_in[i, :mask, :, :]</code></p>\n<p>That seemed to work better for me in regards to ignoring dummy values. Thanks for coming to my TED talk.</p>",
      "rawMarkdown": "Hi. Really great work. But yesterday I was messing around with certain ways of preprocessing my data and I discovered that using `car_mask[i].reshape([-1]).astype(int)` (where `i` is designating the current scene in the batch) as a mask still resulted in a normal sized array. I.e. `p_in[car_mask[i].reshape([-1]).astype(int)]` would still result in an array of size (60, 19, 2) regardless of car_mask's contents. I was pretty sure that this was not the case earlier, but I guess I was wrong. Anyways, I think someone might find it useful to use car_mask as the following:\n\n`mask= np.count_nonzero(car_mask[i])`\n`p_in_relevant = p_in[i, :mask, :, :]`\n\nThat seemed to work better for me in regards to ignoring dummy values. Thanks for coming to my TED talk."
    },
    {
      "id": 1310705,
      "postDate": "2021-05-16T21:26:51.620Z",
      "content": "<p>Did you mean <code>So, in the event you had 10 objects in the scene: car_mask[:10, 1] == 1, and **car_mask[10:, 1]** == 0?</code></p>",
      "rawMarkdown": "Did you mean `So, in the event you had 10 objects in the scene: car_mask[:10, 1] == 1, and **car_mask[10:, 1]** == 0?`",
      "replies": [
        {
          "id": 1310800,
          "postDate": "2021-05-17T01:16:41.637Z",
          "content": "<p>Actually, what I meant was \"car_mask[:10, 0] == 1, and car_mask[10:, 0] == 0\". Fixed now, thanks for pointing it out!</p>",
          "rawMarkdown": "Actually, what I meant was \"car_mask[:10, 0] == 1, and car_mask[10:, 0] == 0\". Fixed now, thanks for pointing it out!"
        }
      ]
    },
    {
      "id": 1284566,
      "postDate": "2021-04-26T05:05:24.313Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1315241,
      "author_name": "Yishai Silver",
      "author_url": "",
      "post_date": "2021-05-19T16:47:13.480000",
      "content": "<p>Hi. Really great work. But yesterday I was messing around with certain ways of preprocessing my data and I discovered that using <code>car_mask[i].reshape([-1]).astype(int)</code> (where <code>i</code> is designating the current scene in the batch) as a mask still resulted in a normal sized array. I.e. <code>p_in[car_mask[i].reshape([-1]).astype(int)]</code> would still result in an array of size (60, 19, 2) regardless of car_mask's contents. I was pretty sure that this was not the case earlier, but I guess I was wrong. Anyways, I think someone might find it useful to use car_mask as the following:</p>\n<p><code>mask= np.count_nonzero(car_mask[i])</code><br>\n<code>p_in_relevant = p_in[i, :mask, :, :]</code></p>\n<p>That seemed to work better for me in regards to ignoring dummy values. Thanks for coming to my TED talk.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1310705,
      "author_name": "Advitya Gemawat",
      "author_url": "",
      "post_date": "2021-05-16T21:26:51.620000",
      "content": "<p>Did you mean <code>So, in the event you had 10 objects in the scene: car_mask[:10, 1] == 1, and **car_mask[10:, 1]** == 0?</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1310800,
          "author_name": "Justin Allen",
          "author_url": "",
          "post_date": "2021-05-17T01:16:41.637000",
          "content": "<p>Actually, what I meant was \"car_mask[:10, 0] == 1, and car_mask[10:, 0] == 0\". Fixed now, thanks for pointing it out!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1284566,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-26T05:05:24.313000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1279365": "The dataset descriptions for the Kaggle competition weren't very descriptive, so I figured I'd post what I've learned about the raw data to help any other students struggling with it. There aren't any statistics or anything showing how to build a good model, only descriptions on how the data is actually laid out in the pickle files, and how it interconnects.\n\nNote: this is all from my own experimentation that could very well be wrong in some cases, so I encourage you to check out the data yourself for a bit to confirm what I'm saying is true, and reply if there is anything wrong.\n\nRaw data description:\n\nThe raw data is in two folders: one for training, and one for validation. Inside each folder is many pickle (.pkl) files. Each pickle file is a dictionary containing information on one individual scene.\n\nEach scene consists of the following information:\n    - Positions of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.\n    - Velocities of other objects in the scene (always 60 total, even if they aren't used. Unused values are filled with 0's). These are sampled at 10HZ.\n    - An ID for each vehicle.\n    - The ID of the vehicle that we are tracking (the 'agent' vehicle).\n    - Information on where the lanes are. These are positions of the centers of the lanes, and what direction those lanes are pointing (the lane normals). There are an arbitrary number of lanes for each scene.\n\nEach pickle file contains a dictionary with the following fields:\n- 'city': a string for the city the data was taken from. Can either be 'PIT' or 'MIA'\n\n- 'scene_idx': a unique non-zero integer acting as an ID for the current scene (unique across the union of training and validation sets)\n\n- 'agent_id': a string of the form \"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits that identifies the 'agent' we want to track (predict future movement of) in this scene. This ID is always 36 characters long (32 digits if you don't include the -'s)\n\n- 'car_mask': a (60, 1) numpy array of float 1's and 0's. Each row corresponds to the row index of each car in the scene (IE: this matches with the track_id, p_in, p_out, v_in, and v_out first dimensions). There is a '1' in the row if the values in the first dimension of track_id, p_in, etc. correspond to an actual object in the scene, or are just empty because there were less than 60 objects in the scene. The total number of 1's in this mask is the total number of objects in the scene, including the agent car we are tracking. This allows us to quickly get all of the values in track_id, p_in, etc. that correspond to actual objects like so:\n\n        actual_objects = track_id[car_mask.reshape([-1]).astype(int)]\n\n        NOTE: the values in the numpy array are floats, and thus need to be converted to integers before using them as a mask (hence the .astype(int))\n\n    The car_mask is laid out such that the only 1's occur at the beginning of the array. So, in the event you had 10 objects in the scene: car_mask[:10, 0] == 1, and car_mask[10:, 0] == 0. This means all of the actually tracked objects in each of track_id, p_in, etc come first, and the rest of the values are dummy values.\n\n- 'track_id': a (60, 30, 1) numpy array of strings. Each string is of the form \n\"00000000-0000-0000-0000-0000XXXXXXXX\" with the X's being digits. This describes a unique id (unique to the scene, not all of the data) for each object in the scene. The whole array is 60 \"30 by 1 column vectors\" where every one of the 30 elements in the column vector are the same track_id string. Don't ask me why this is the case, I have no idea. They could have just as easily made a 1-D list of 60 track_ids and accomplished the same thing. Honestly, whoever made this was probably just too lazy to change it, and I don't blame them; I'd probably do the same thing.\n    This array follows the same indexing as the 'car_mask'. So, the first n 1's in the car mask array describe the track_ids for all of the actually tracked objects in the scene. After the first n actually tracked objects, the track_id changes to 'dummyK' where K is the integer index of the dummy variable starting at 0 (IE: 'dummy0', 'dummy1', ..., 'dummy[60-n-1]').\n\n- 'p_in': a (60, 19, 2) numpy array of floats. These are the x,y coordinates of each tracked object in the scene, for a max of 60 objects. What those coordinates are relative to I do not know (probably some longitude/latitude in or near whatever city the data is from), so you may not want to use the exact values in your project, and instead care only about the change in position at each time step. There are 19 samples per object meaning the data is probably only for 1.9 seconds, not 2 seconds exactly. This array follows the same indexing as car_mask, so the first n arrays in p_in correspond to the n tracked objects in the scene, and the following 60-n arrays are filled with all 0's.\n        NOTE: some values may be 0 or negative even if we are tracking them, indicating the object is at the very edge of the coordinate system\n\n- 'v_in': same as 'p_in', but for tracking velocities instead of position. These values can be positive, negative, or 0. Again, I don't know what the units are. It might be on the Argoverse website somewhere?\n        NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)\n\n- 'p_out': same as 'p_in', but is instead of shape (60, 30, 2). These are the positions your model should learn to predict.\n        NOTE: the validation sets do not contain this key as this is what your model should predict.\n\n- 'v_out': same as 'v_in', but is instead of shape (60, 30, 2). You do not need to predict these for the final project. You might want to discard them, unless you want to instead have your model predict velocity instead of position, and you can calculate the final positions by adding some small multiple of the velocity?\n        NOTE: the validation sets do not contain this key\n        NOTE: some values may still be 0.0 for velocity even if it is for an object we are tracking (IE: the object is not moving)\n\n- 'lane': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene. There can be a different number of lanes for each scene, so there is no guarantee on the size of 'k'. These describe the x,y,z coordinates of the center of lane nodes (where driving lanes are in the scene). For some reason, the z-coordinate is included, but is always 0, so you can ignore it and just look at the x,y. These seem to be in the same coordinate system as the p_in and p_out coordinates. Perhaps you want to change these to be relative to something else in the scene?\n        NOTE: some x,y values may be 0 or negative indicating the center of the lane is at the very edge of the coordinate system\n\n- 'lane_norm': a (k, 3) numpy array of floats where 'k' is the number of lanes in the scene (same size as the 'k' in 'lane'). These describe the x,y,z vector direction of the corresponding lane normal (the direction the lane center with the same index is pointing). These values can be positive or negative, but never 0 (for x,y). Again, the z direction is always 0 and can be ignored, and the coordinate system seems to be the same as p_in.",
    "1315241": "Hi. Really great work. But yesterday I was messing around with certain ways of preprocessing my data and I discovered that using `car_mask[i].reshape([-1]).astype(int)` (where `i` is designating the current scene in the batch) as a mask still resulted in a normal sized array. I.e. `p_in[car_mask[i].reshape([-1]).astype(int)]` would still result in an array of size (60, 19, 2) regardless of car_mask's contents. I was pretty sure that this was not the case earlier, but I guess I was wrong. Anyways, I think someone might find it useful to use car_mask as the following:\n\n`mask= np.count_nonzero(car_mask[i])`\n`p_in_relevant = p_in[i, :mask, :, :]`\n\nThat seemed to work better for me in regards to ignoring dummy values. Thanks for coming to my TED talk.",
    "1310705": "Did you mean `So, in the event you had 10 objects in the scene: car_mask[:10, 1] == 1, and **car_mask[10:, 1]** == 0?`",
    "1284566": ""
  }
}