{
  "id": 183814,
  "title": "Improving convergence and sample efficiency",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/183814",
  "author_name": "ryches",
  "post_date": "2020-09-18T06:38:13.031000",
  "votes": 32,
  "comment_count": 26,
  "views": 0,
  "content": "<p>One of the battles many of us are fighting against this data is that the data loading is a serious bottleneck. Obviously it would be great to remove this bottleneck and improve training throughput but so far it doesn't seem like there has been a great way around the rasterization slow down. Rather than trying to load the data faster maybe it is prudent to load only the most useful data. </p>\n<p>In my preliminary experiments, I am seeing that the loss numbers per batch are significantly skewed.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F05a5209502afd0e81baacd9d2d0ef00e%2Faveragenll.png?generation=1600410251301531&amp;alt=media\" alt=\"Mean NLL per batch\"></p>\n<p>This histogram is showing the distribution of the mean NLL per batch across my last 200k training batches. There is a significant mass centered around 20 and then a long tail with a max reaching all the way to &gt;900. My thought is that it likely isnt particularly worthwhile to continue optimizing for samples that are already well predicted by our model. This seems to be a common theme in self-driving cars, there is a lot of data, but it is difficult to hone in on the important data and not waste time on the mundane and predictable. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ff191ffc2d0e491f7ab37f7934c322315%2Fdistribution_describe.png?generation=1600410606005697&amp;alt=media\" alt=\"Describe of last 200k training steps\"></p>\n<p>I looked at OHEM(Online Hard Example Mining), but this doesn't seem ideal to me as the focus is put on the hard samples but all data is still loaded. One alternate idea is to do some sort of weighted sampler that increases the prevalence of scenes based on their loss on previous epochs. Then on future epochs the harder samples will happen more frequently and unuseful training data will not be loaded. The issue with this is that running through all of the data even once is fairly time-consuming. I have begun logging these numbers because I can't imagine I will have time to do a continuous run but maybe if I capture the data early it will be useful to survey the mean across my various different experiments per training sample. </p>\n<p>In theory with the correctly chosen points we could likely converge a model much faster with much less processing from the dataloader. Does anyone else have any ideas in this regard?</p>",
  "messages": [
    {
      "id": 1015380,
      "postDate": "2020-09-18T06:38:13.030Z",
      "content": "<p>One of the battles many of us are fighting against this data is that the data loading is a serious bottleneck. Obviously it would be great to remove this bottleneck and improve training throughput but so far it doesn't seem like there has been a great way around the rasterization slow down. Rather than trying to load the data faster maybe it is prudent to load only the most useful data. </p>\n<p>In my preliminary experiments, I am seeing that the loss numbers per batch are significantly skewed.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F05a5209502afd0e81baacd9d2d0ef00e%2Faveragenll.png?generation=1600410251301531&amp;alt=media\" alt=\"Mean NLL per batch\"></p>\n<p>This histogram is showing the distribution of the mean NLL per batch across my last 200k training batches. There is a significant mass centered around 20 and then a long tail with a max reaching all the way to &gt;900. My thought is that it likely isnt particularly worthwhile to continue optimizing for samples that are already well predicted by our model. This seems to be a common theme in self-driving cars, there is a lot of data, but it is difficult to hone in on the important data and not waste time on the mundane and predictable. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ff191ffc2d0e491f7ab37f7934c322315%2Fdistribution_describe.png?generation=1600410606005697&amp;alt=media\" alt=\"Describe of last 200k training steps\"></p>\n<p>I looked at OHEM(Online Hard Example Mining), but this doesn't seem ideal to me as the focus is put on the hard samples but all data is still loaded. One alternate idea is to do some sort of weighted sampler that increases the prevalence of scenes based on their loss on previous epochs. Then on future epochs the harder samples will happen more frequently and unuseful training data will not be loaded. The issue with this is that running through all of the data even once is fairly time-consuming. I have begun logging these numbers because I can't imagine I will have time to do a continuous run but maybe if I capture the data early it will be useful to survey the mean across my various different experiments per training sample. </p>\n<p>In theory with the correctly chosen points we could likely converge a model much faster with much less processing from the dataloader. Does anyone else have any ideas in this regard?</p>",
      "rawMarkdown": "One of the battles many of us are fighting against this data is that the data loading is a serious bottleneck. Obviously it would be great to remove this bottleneck and improve training throughput but so far it doesn't seem like there has been a great way around the rasterization slow down. Rather than trying to load the data faster maybe it is prudent to load only the most useful data. \n\nIn my preliminary experiments, I am seeing that the loss numbers per batch are significantly skewed.\n\n![Mean NLL per batch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F05a5209502afd0e81baacd9d2d0ef00e%2Faveragenll.png?generation=1600410251301531&alt=media)\n\nThis histogram is showing the distribution of the mean NLL per batch across my last 200k training batches. There is a significant mass centered around 20 and then a long tail with a max reaching all the way to >900. My thought is that it likely isnt particularly worthwhile to continue optimizing for samples that are already well predicted by our model. This seems to be a common theme in self-driving cars, there is a lot of data, but it is difficult to hone in on the important data and not waste time on the mundane and predictable. \n\n![Describe of last 200k training steps](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ff191ffc2d0e491f7ab37f7934c322315%2Fdistribution_describe.png?generation=1600410606005697&alt=media)\n\nI looked at OHEM(Online Hard Example Mining), but this doesn't seem ideal to me as the focus is put on the hard samples but all data is still loaded. One alternate idea is to do some sort of weighted sampler that increases the prevalence of scenes based on their loss on previous epochs. Then on future epochs the harder samples will happen more frequently and unuseful training data will not be loaded. The issue with this is that running through all of the data even once is fairly time-consuming. I have begun logging these numbers because I can't imagine I will have time to do a continuous run but maybe if I capture the data early it will be useful to survey the mean across my various different experiments per training sample. \n\nIn theory with the correctly chosen points we could likely converge a model much faster with much less processing from the dataloader. Does anyone else have any ideas in this regard?",
      "votes": 32
    },
    {
      "id": 1016386,
      "postDate": "2020-09-18T22:50:20.873Z",
      "content": "<p>Can you try increasing the min_frame_history and min_frame_future and retrying the test? <br>\ni.e. &gt;<code>AgentDataset(cfg, sample_zarr, rasterizer, min_frame_history=20, min_frame_future=20)</code></p>\n<p>The defaults are set to 10 and 1 respectively. Changing these to a bigger number means you actually get agents with more of a future and a past, as having an agent for which you have to only predict 1 frame in the future won't be nearly as useful as having to 20 target values. </p>\n<p>I did a quick check on the sample dataset, and if you increase min_frame_history and min_frame_future to 20 the amount of agents go from 111634 to 52277. If you up it to 50 it drops down to 19560. </p>\n<p>I'm willing to bet doing this will increase your mean loss significantly.  </p>",
      "rawMarkdown": "Can you try increasing the min_frame_history and min_frame_future and retrying the test? \ni.e. >`AgentDataset(cfg, sample_zarr, rasterizer, min_frame_history=20, min_frame_future=20)`\n\nThe defaults are set to 10 and 1 respectively. Changing these to a bigger number means you actually get agents with more of a future and a past, as having an agent for which you have to only predict 1 frame in the future won't be nearly as useful as having to 20 target values. \n\nI did a quick check on the sample dataset, and if you increase min_frame_history and min_frame_future to 20 the amount of agents go from 111634 to 52277. If you up it to 50 it drops down to 19560. \n\nI'm willing to bet doing this will increase your mean loss significantly.  ",
      "votes": 7,
      "replies": [
        {
          "id": 1016400,
          "postDate": "2020-09-18T23:31:44.717Z",
          "content": "<p><a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> The number might go down. But its a very bad idea. You have to predict 50 frames into the future. You can reduce agent by making history to zero even though it compromises accuracy. But why would you want to decrease future frames? reducing future up to 30 makes little sense(50 should be out of the question). But I am sure test data is set in such a way to account for diverse futures. with up to 30 you will definitely wont meet optimal(its my thinking). This reduction also can only be used to test if test data is diverse enough but not to arrive at best solution.    </p>",
          "rawMarkdown": "@fnands The number might go down. But its a very bad idea. You have to predict 50 frames into the future. You can reduce agent by making history to zero even though it compromises accuracy. But why would you want to decrease future frames? reducing future up to 30 makes little sense(50 should be out of the question). But I am sure test data is set in such a way to account for diverse futures. with up to 30 you will definitely wont meet optimal(its my thinking). This reduction also can only be used to test if test data is diverse enough but not to arrive at best solution.    ",
          "votes": -3
        },
        {
          "id": 1016702,
          "postDate": "2020-09-19T07:11:05.693Z",
          "content": "<p>Changing these parameters will <strong>increase</strong> the average number of future frames, and I believe therefore <strong>increase</strong> the mean loss.  </p>\n<p>I think you are misunderstanding how the AgentDataset works, and I did too, as it's not clearly explained.   </p>\n<p>The AgentDataset doesn't iterate through each agent by track_id, it iterates through each agent by the agent_index, which can change per frame, even for the same \"track_id\", i.e. the same vehicle.</p>\n<p>At the moment the default minimum  <code>min_frame_future = 1</code>, meaning if you don't set it, you will get many images in your dataset with only 1 real future frame. It doesn't matter if you set <code>future_num_frames</code> to 50, as you will just end up with 1 frame with non-zero 'target_positions' and 49 frames of padded on zeroes.   </p>",
          "rawMarkdown": "Changing these parameters will **increase** the average number of future frames, and I believe therefore **increase** the mean loss.  \n   \nI think you are misunderstanding how the AgentDataset works, and I did too, as it's not clearly explained.   \n\nThe AgentDataset doesn't iterate through each agent by track_id, it iterates through each agent by the agent_index, which can change per frame, even for the same \"track_id\", i.e. the same vehicle.\n\nAt the moment the default minimum  `min_frame_future = 1`, meaning if you don't set it, you will get many images in your dataset with only 1 real future frame. It doesn't matter if you set `future_num_frames` to 50, as you will just end up with 1 frame with non-zero 'target_positions' and 49 frames of padded on zeroes.   ",
          "votes": 5
        },
        {
          "id": 1018412,
          "postDate": "2020-09-19T17:01:26.903Z",
          "content": "<p>Do we know how the test set was generated? i.e. is there a min_frame_future implemented implicitly there?</p>",
          "rawMarkdown": "Do we know how the test set was generated? i.e. is there a min_frame_future implemented implicitly there?",
          "votes": 5
        },
        {
          "id": 1018425,
          "postDate": "2020-09-19T17:14:36.917Z",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> great question, I just checked the history for the first couple test cases.<br>\n<code>test_dataset[34]['history_availabilities'].sum() = 3.0</code><br>\n<code>test_dataset[138]['history_availabilities'].sum() = 4.0</code></p>\n<p>so, there are test cases with only very few history frames. I would suspect the same for the future frames.</p>",
          "rawMarkdown": "@fergusoci great question, I just checked the history for the first couple test cases.\n`test_dataset[34]['history_availabilities'].sum() = 3.0`\n`test_dataset[138]['history_availabilities'].sum() = 4.0`\n\nso, there are test cases with only very few history frames. I would suspect the same for the future frames.",
          "votes": 6
        },
        {
          "id": 1018434,
          "postDate": "2020-09-19T17:19:47.780Z",
          "content": "<p>I believe that is 10. I believe that is what the create_chopped_dataset function does, find all the samples that satisfy the constraints. Looking in the evaluation portion of <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">this notebook</a> the create_chopped_dataset takes the MIN_FUTURE_STEPS argument and it is set to 10 in <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/evaluation/chop_dataset.py\" target=\"_blank\">this file</a></p>",
          "rawMarkdown": "I believe that is 10. I believe that is what the create_chopped_dataset function does, find all the samples that satisfy the constraints. Looking in the evaluation portion of [this notebook](https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb) the create_chopped_dataset takes the MIN_FUTURE_STEPS argument and it is set to 10 in [this file](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/evaluation/chop_dataset.py)",
          "votes": 5
        },
        {
          "id": 1018488,
          "postDate": "2020-09-19T17:51:13.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> , I did a quick test and it seems there are a few cases in the test set with very short histories (1 frame, even when using the test mask), although the vast majority give you plenty of history to work with. So I guess the cases where we have only 1 frame of history to go on can be pretty challenging.  </p>\n<p>To answer <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a>'s question I think <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> is right, as we know they used <code>create_chopped_dataset</code> to make the test set, and I guess we can assume the used <code>MIN_FUTURE_STEPS = 10</code>, which could maybe mean that setting <code>min_frame_future=10</code> will get you a better approximation of the test set. </p>",
          "rawMarkdown": "@ilu000 , I did a quick test and it seems there are a few cases in the test set with very short histories (1 frame, even when using the test mask), although the vast majority give you plenty of history to work with. So I guess the cases where we have only 1 frame of history to go on can be pretty challenging.  \n\nTo answer @fergusoci's question I think @ryches is right, as we know they used `create_chopped_dataset` to make the test set, and I guess we can assume the used `MIN_FUTURE_STEPS = 10`, which could maybe mean that setting `min_frame_future=10` will get you a better approximation of the test set. ",
          "votes": 3
        },
        {
          "id": 1018494,
          "postDate": "2020-09-19T17:56:17.517Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> good find for the <code>MIN_FUTURE_STEPS</code><br>\nWe also know that <code>num_frames_to_copy</code> is set to 100, but that doesn't carry over to a minimum history length. So, I assume we see anything from 1 to 100 for <code>history_availabilities</code> in the test set with a peak at 100. </p>",
          "rawMarkdown": "@ryches good find for the `MIN_FUTURE_STEPS`\nWe also know that `num_frames_to_copy` is set to 100, but that doesn't carry over to a minimum history length. So, I assume we see anything from 1 to 100 for `history_availabilities` in the test set with a peak at 100. ",
          "votes": 1
        },
        {
          "id": 1019322,
          "postDate": "2020-09-20T10:58:07.687Z",
          "content": "<p>I might be wrong but I dont agree with few mentioned statement like min history length =1 in some cases. If some of ye all want to look at Ground truth and target availability with sample.zaar. I will also do the same for eval data so everyone can directly see the groundtrutt.csv , use it eval and you dont have to generate that zarr. You can visit my dataset. <a href=\"https://www.kaggle.com/deepakrajpurushothaman/sample-chopped-100\" target=\"_blank\">https://www.kaggle.com/deepakrajpurushothaman/sample-chopped-100</a> . </p>",
          "rawMarkdown": "I might be wrong but I dont agree with few mentioned statement like min history length =1 in some cases. If some of ye all want to look at Ground truth and target availability with sample.zaar. I will also do the same for eval data so everyone can directly see the groundtrutt.csv , use it eval and you dont have to generate that zarr. You can visit my dataset. https://www.kaggle.com/deepakrajpurushothaman/sample-chopped-100 . ",
          "votes": -1
        },
        {
          "id": 1019408,
          "postDate": "2020-09-20T12:10:32.017Z",
          "content": "<p><a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a> i can't follow your thoughts. History is always 100 for the entire test set. But the availability (<code>history_availabilities</code>) is lower, sometimes as low as 1 as <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> said above. <br>\nFor the future frames we can only speculate, but given the hardcoded value for <code>MIN_FUTURE_STEPS = 10</code> in the repository this value is probably a good guess. </p>",
          "rawMarkdown": "@deepakrajpurushothaman i can't follow your thoughts. History is always 100 for the entire test set. But the availability (`history_availabilities`) is lower, sometimes as low as 1 as @fnands said above. \nFor the future frames we can only speculate, but given the hardcoded value for `MIN_FUTURE_STEPS = 10` in the repository this value is probably a good guess. ",
          "votes": 2
        },
        {
          "id": 1021347,
          "postDate": "2020-09-21T19:15:54.343Z",
          "content": "<p>Does anybody know how validate.zarr was created? Was that done using <code>create_chopped_dataset()</code> too? Dug around, but can't see any mention of it anywhere…</p>",
          "rawMarkdown": "Does anybody know how validate.zarr was created? Was that done using `create_chopped_dataset()` too? Dug around, but can't see any mention of it anywhere...",
          "votes": 1
        },
        {
          "id": 1021368,
          "postDate": "2020-09-21T19:37:13.563Z",
          "content": "<p>The validate file was not created with that function. The create_chopped_dataset function must be run against it to make data that is representative of the test set. I've found good correlation between val and test by doing this. The validate is raw like the train files</p>",
          "rawMarkdown": "The validate file was not created with that function. The create_chopped_dataset function must be run against it to make data that is representative of the test set. I've found good correlation between val and test by doing this. The validate is raw like the train files",
          "votes": 6
        },
        {
          "id": 1021371,
          "postDate": "2020-09-21T19:37:56.680Z",
          "content": "<p>The examples in the l5kit repo show this process</p>",
          "rawMarkdown": "The examples in the l5kit repo show this process",
          "votes": 1
        },
        {
          "id": 1021383,
          "postDate": "2020-09-21T19:44:49.900Z",
          "content": "<p>Great, thanks. </p>",
          "rawMarkdown": "Great, thanks. ",
          "votes": 1
        },
        {
          "id": 1045772,
          "postDate": "2020-10-11T02:51:11.557Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1045773,
          "postDate": "2020-10-11T02:52:03.320Z",
          "content": "<p>Should I follow the example in <a href=\"url\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb</a> to test the model on validation set?</p>",
          "rawMarkdown": "Should I follow the example in [https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb](url) to test the model on validation set?"
        }
      ]
    },
    {
      "id": 1015520,
      "postDate": "2020-09-18T08:31:58.563Z",
      "content": "<p>Good observation. Continuing your line of thought, at first we can load the data without rasterization, and analyze <code>history_positions</code> and <code>target_positions</code>. If the targets can be accurately predicted with a constant speed model, then skip or undersample these items.</p>",
      "rawMarkdown": "Good observation. Continuing your line of thought, at first we can load the data without rasterization, and analyze `history_positions` and `target_positions`. If the targets can be accurately predicted with a constant speed model, then skip or undersample these items.",
      "votes": 7,
      "replies": [
        {
          "id": 1015578,
          "postDate": "2020-09-18T09:13:34.667Z",
          "content": "<p>Exactly what I though. You can easily predict the motion for many agents just by extrapolating their initial positions/velocities, so these probably don't teach you much. </p>\n<p>As a first experiment you can probably just completely drop these from the training set. </p>",
          "rawMarkdown": "Exactly what I though. You can easily predict the motion for many agents just by extrapolating their initial positions/velocities, so these probably don't teach you much. \n\nAs a first experiment you can probably just completely drop these from the training set. ",
          "votes": 2
        },
        {
          "id": 1069586,
          "postDate": "2020-11-04T16:24:45.537Z",
          "content": "<p>Thanks for your insights <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>! What do you mean by constant speed model? Is it related to this notebook by you - <a href=\"https://www.kaggle.com/zaharch/kalman-filter-baseline\" target=\"_blank\">https://www.kaggle.com/zaharch/kalman-filter-baseline</a>?</p>",
          "rawMarkdown": "Thanks for your insights @zaharch! What do you mean by constant speed model? Is it related to this notebook by you - https://www.kaggle.com/zaharch/kalman-filter-baseline?"
        },
        {
          "id": 1069598,
          "postDate": "2020-11-04T16:55:50.703Z",
          "content": "<p>Not necessarily related, by constant speed I mean estimating speed by any simple method. For example just calculating it from the last two history position points.</p>",
          "rawMarkdown": "Not necessarily related, by constant speed I mean estimating speed by any simple method. For example just calculating it from the last two history position points.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1046803,
      "postDate": "2020-10-12T02:48:43.560Z",
      "content": "<p>great</p>",
      "rawMarkdown": "great"
    },
    {
      "id": 1015804,
      "postDate": "2020-09-18T12:46:28.713Z",
      "content": "<p>I  idea of 'weighted sampler' gets my attention. How do you go about doing that?</p>",
      "rawMarkdown": "I  idea of 'weighted sampler' gets my attention. How do you go about doing that?",
      "replies": [
        {
          "id": 1015866,
          "postDate": "2020-09-18T13:45:28.877Z",
          "content": "<p>In pytorch it isn't too difficult. There is a built in weighted sampler you can configure and then use with your dataloaders. <a href=\"https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html\" target=\"_blank\">https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html</a></p>\n<p>I tried working in that direction but ran into an issue with the sampling not being able to work when the dataset size is more than 2^24 samples. <a href=\"https://github.com/pytorch/pytorch/issues/2576\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/2576</a></p>",
          "rawMarkdown": "In pytorch it isn't too difficult. There is a built in weighted sampler you can configure and then use with your dataloaders. https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html\n\nI tried working in that direction but ran into an issue with the sampling not being able to work when the dataset size is more than 2^24 samples. https://github.com/pytorch/pytorch/issues/2576",
          "votes": 1
        },
        {
          "id": 1016304,
          "postDate": "2020-09-18T20:02:49.450Z",
          "content": "<p>Thank! I will have a look at it.</p>",
          "rawMarkdown": "Thank! I will have a look at it."
        },
        {
          "id": 1059838,
          "postDate": "2020-10-25T14:15:34.860Z",
          "content": "<p>Can we use a 50% subset of the training set when we are using weighted sampler? But I am confused that how we can use weighted sampler according to the label or the loss. If we use sampler refers to losses, should we go through all the data in the training set?</p>",
          "rawMarkdown": "Can we use a 50% subset of the training set when we are using weighted sampler? But I am confused that how we can use weighted sampler according to the label or the loss. If we use sampler refers to losses, should we go through all the data in the training set?"
        }
      ]
    },
    {
      "id": 1020035,
      "postDate": "2020-09-20T20:34:53.787Z",
      "content": "<p>thank you! </p>",
      "rawMarkdown": "thank you! ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1016386,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "2020-09-18T22:50:20.873000",
      "content": "<p>Can you try increasing the min_frame_history and min_frame_future and retrying the test? <br>\ni.e. &gt;<code>AgentDataset(cfg, sample_zarr, rasterizer, min_frame_history=20, min_frame_future=20)</code></p>\n<p>The defaults are set to 10 and 1 respectively. Changing these to a bigger number means you actually get agents with more of a future and a past, as having an agent for which you have to only predict 1 frame in the future won't be nearly as useful as having to 20 target values. </p>\n<p>I did a quick check on the sample dataset, and if you increase min_frame_history and min_frame_future to 20 the amount of agents go from 111634 to 52277. If you up it to 50 it drops down to 19560. </p>\n<p>I'm willing to bet doing this will increase your mean loss significantly.  </p>",
      "votes": 7,
      "replies": [
        {
          "id": 1016400,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-09-18T23:31:44.717000",
          "content": "<p><a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> The number might go down. But its a very bad idea. You have to predict 50 frames into the future. You can reduce agent by making history to zero even though it compromises accuracy. But why would you want to decrease future frames? reducing future up to 30 makes little sense(50 should be out of the question). But I am sure test data is set in such a way to account for diverse futures. with up to 30 you will definitely wont meet optimal(its my thinking). This reduction also can only be used to test if test data is diverse enough but not to arrive at best solution.    </p>",
          "votes": -3,
          "replies": []
        },
        {
          "id": 1016702,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "2020-09-19T07:11:05.693000",
          "content": "<p>Changing these parameters will <strong>increase</strong> the average number of future frames, and I believe therefore <strong>increase</strong> the mean loss.  </p>\n<p>I think you are misunderstanding how the AgentDataset works, and I did too, as it's not clearly explained.   </p>\n<p>The AgentDataset doesn't iterate through each agent by track_id, it iterates through each agent by the agent_index, which can change per frame, even for the same \"track_id\", i.e. the same vehicle.</p>\n<p>At the moment the default minimum  <code>min_frame_future = 1</code>, meaning if you don't set it, you will get many images in your dataset with only 1 real future frame. It doesn't matter if you set <code>future_num_frames</code> to 50, as you will just end up with 1 frame with non-zero 'target_positions' and 49 frames of padded on zeroes.   </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1018412,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-19T17:01:26.903000",
          "content": "<p>Do we know how the test set was generated? i.e. is there a min_frame_future implemented implicitly there?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1018425,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-19T17:14:36.917000",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> great question, I just checked the history for the first couple test cases.<br>\n<code>test_dataset[34]['history_availabilities'].sum() = 3.0</code><br>\n<code>test_dataset[138]['history_availabilities'].sum() = 4.0</code></p>\n<p>so, there are test cases with only very few history frames. I would suspect the same for the future frames.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1018434,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-19T17:19:47.780000",
          "content": "<p>I believe that is 10. I believe that is what the create_chopped_dataset function does, find all the samples that satisfy the constraints. Looking in the evaluation portion of <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">this notebook</a> the create_chopped_dataset takes the MIN_FUTURE_STEPS argument and it is set to 10 in <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/evaluation/chop_dataset.py\" target=\"_blank\">this file</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1018488,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "2020-09-19T17:51:13.080000",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> , I did a quick test and it seems there are a few cases in the test set with very short histories (1 frame, even when using the test mask), although the vast majority give you plenty of history to work with. So I guess the cases where we have only 1 frame of history to go on can be pretty challenging.  </p>\n<p>To answer <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a>'s question I think <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> is right, as we know they used <code>create_chopped_dataset</code> to make the test set, and I guess we can assume the used <code>MIN_FUTURE_STEPS = 10</code>, which could maybe mean that setting <code>min_frame_future=10</code> will get you a better approximation of the test set. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1018494,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-19T17:56:17.517000",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> good find for the <code>MIN_FUTURE_STEPS</code><br>\nWe also know that <code>num_frames_to_copy</code> is set to 100, but that doesn't carry over to a minimum history length. So, I assume we see anything from 1 to 100 for <code>history_availabilities</code> in the test set with a peak at 100. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1019322,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-09-20T10:58:07.687000",
          "content": "<p>I might be wrong but I dont agree with few mentioned statement like min history length =1 in some cases. If some of ye all want to look at Ground truth and target availability with sample.zaar. I will also do the same for eval data so everyone can directly see the groundtrutt.csv , use it eval and you dont have to generate that zarr. You can visit my dataset. <a href=\"https://www.kaggle.com/deepakrajpurushothaman/sample-chopped-100\" target=\"_blank\">https://www.kaggle.com/deepakrajpurushothaman/sample-chopped-100</a> . </p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1019408,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-20T12:10:32.017000",
          "content": "<p><a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a> i can't follow your thoughts. History is always 100 for the entire test set. But the availability (<code>history_availabilities</code>) is lower, sometimes as low as 1 as <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> said above. <br>\nFor the future frames we can only speculate, but given the hardcoded value for <code>MIN_FUTURE_STEPS = 10</code> in the repository this value is probably a good guess. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1021347,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-21T19:15:54.343000",
          "content": "<p>Does anybody know how validate.zarr was created? Was that done using <code>create_chopped_dataset()</code> too? Dug around, but can't see any mention of it anywhere…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1021368,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-21T19:37:13.563000",
          "content": "<p>The validate file was not created with that function. The create_chopped_dataset function must be run against it to make data that is representative of the test set. I've found good correlation between val and test by doing this. The validate is raw like the train files</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1021371,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-21T19:37:56.680000",
          "content": "<p>The examples in the l5kit repo show this process</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1021383,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-21T19:44:49.900000",
          "content": "<p>Great, thanks. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1045772,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-11T02:51:11.557000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1045773,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-10-11T02:52:03.320000",
          "content": "<p>Should I follow the example in <a href=\"url\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb</a> to test the model on validation set?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1015520,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2020-09-18T08:31:58.563000",
      "content": "<p>Good observation. Continuing your line of thought, at first we can load the data without rasterization, and analyze <code>history_positions</code> and <code>target_positions</code>. If the targets can be accurately predicted with a constant speed model, then skip or undersample these items.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1015578,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "2020-09-18T09:13:34.667000",
          "content": "<p>Exactly what I though. You can easily predict the motion for many agents just by extrapolating their initial positions/velocities, so these probably don't teach you much. </p>\n<p>As a first experiment you can probably just completely drop these from the training set. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1069586,
          "author_name": "Arpit",
          "author_url": "",
          "post_date": "2020-11-04T16:24:45.537000",
          "content": "<p>Thanks for your insights <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>! What do you mean by constant speed model? Is it related to this notebook by you - <a href=\"https://www.kaggle.com/zaharch/kalman-filter-baseline\" target=\"_blank\">https://www.kaggle.com/zaharch/kalman-filter-baseline</a>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1069598,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-04T16:55:50.703000",
          "content": "<p>Not necessarily related, by constant speed I mean estimating speed by any simple method. For example just calculating it from the last two history position points.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1046803,
      "author_name": "Rohit Dube",
      "author_url": "",
      "post_date": "2020-10-12T02:48:43.560000",
      "content": "<p>great</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1015804,
      "author_name": "The Brown Iceman",
      "author_url": "",
      "post_date": "2020-09-18T12:46:28.713000",
      "content": "<p>I  idea of 'weighted sampler' gets my attention. How do you go about doing that?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1015866,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-18T13:45:28.877000",
          "content": "<p>In pytorch it isn't too difficult. There is a built in weighted sampler you can configure and then use with your dataloaders. <a href=\"https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html\" target=\"_blank\">https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html</a></p>\n<p>I tried working in that direction but ran into an issue with the sampling not being able to work when the dataset size is more than 2^24 samples. <a href=\"https://github.com/pytorch/pytorch/issues/2576\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/2576</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1016304,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-09-18T20:02:49.450000",
          "content": "<p>Thank! I will have a look at it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059838,
          "author_name": "Yannik",
          "author_url": "",
          "post_date": "2020-10-25T14:15:34.860000",
          "content": "<p>Can we use a 50% subset of the training set when we are using weighted sampler? But I am confused that how we can use weighted sampler according to the label or the loss. If we use sampler refers to losses, should we go through all the data in the training set?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1020035,
      "author_name": "berkaycihan",
      "author_url": "",
      "post_date": "2020-09-20T20:34:53.787000",
      "content": "<p>thank you! </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1015380": "One of the battles many of us are fighting against this data is that the data loading is a serious bottleneck. Obviously it would be great to remove this bottleneck and improve training throughput but so far it doesn't seem like there has been a great way around the rasterization slow down. Rather than trying to load the data faster maybe it is prudent to load only the most useful data. \n\nIn my preliminary experiments, I am seeing that the loss numbers per batch are significantly skewed.\n\n![Mean NLL per batch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F05a5209502afd0e81baacd9d2d0ef00e%2Faveragenll.png?generation=1600410251301531&alt=media)\n\nThis histogram is showing the distribution of the mean NLL per batch across my last 200k training batches. There is a significant mass centered around 20 and then a long tail with a max reaching all the way to >900. My thought is that it likely isnt particularly worthwhile to continue optimizing for samples that are already well predicted by our model. This seems to be a common theme in self-driving cars, there is a lot of data, but it is difficult to hone in on the important data and not waste time on the mundane and predictable. \n\n![Describe of last 200k training steps](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ff191ffc2d0e491f7ab37f7934c322315%2Fdistribution_describe.png?generation=1600410606005697&alt=media)\n\nI looked at OHEM(Online Hard Example Mining), but this doesn't seem ideal to me as the focus is put on the hard samples but all data is still loaded. One alternate idea is to do some sort of weighted sampler that increases the prevalence of scenes based on their loss on previous epochs. Then on future epochs the harder samples will happen more frequently and unuseful training data will not be loaded. The issue with this is that running through all of the data even once is fairly time-consuming. I have begun logging these numbers because I can't imagine I will have time to do a continuous run but maybe if I capture the data early it will be useful to survey the mean across my various different experiments per training sample. \n\nIn theory with the correctly chosen points we could likely converge a model much faster with much less processing from the dataloader. Does anyone else have any ideas in this regard?",
    "1016386": "Can you try increasing the min_frame_history and min_frame_future and retrying the test? \ni.e. >`AgentDataset(cfg, sample_zarr, rasterizer, min_frame_history=20, min_frame_future=20)`\n\nThe defaults are set to 10 and 1 respectively. Changing these to a bigger number means you actually get agents with more of a future and a past, as having an agent for which you have to only predict 1 frame in the future won't be nearly as useful as having to 20 target values. \n\nI did a quick check on the sample dataset, and if you increase min_frame_history and min_frame_future to 20 the amount of agents go from 111634 to 52277. If you up it to 50 it drops down to 19560. \n\nI'm willing to bet doing this will increase your mean loss significantly.  ",
    "1015520": "Good observation. Continuing your line of thought, at first we can load the data without rasterization, and analyze `history_positions` and `target_positions`. If the targets can be accurately predicted with a constant speed model, then skip or undersample these items.",
    "1046803": "great",
    "1015804": "I  idea of 'weighted sampler' gets my attention. How do you go about doing that?",
    "1020035": "thank you! "
  }
}