{
  "id": 185762,
  "title": "Large deviation between training loss and evaluation loss",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/185762",
  "author_name": "",
  "post_date": "2020-09-22T03:25:39.686299700Z",
  "votes": 41,
  "comment_count": 66,
  "views": 0,
  "content": "<p>Hello fellows,</p>\n<p>The loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:</p>\n<ul>\n<li><p>Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.</p></li>\n<li><p>Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.</p></li>\n<li><p>Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.</p></li>\n</ul>\n<p>Did you guys have the same problem? Any thoughts?</p>\n<p>Thanks!<br>\nFrank</p>\n<p>Update on 14 Oct:<br>\nFor me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.<br>\nNow I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄</p>",
  "messages": [
    {
      "id": "1021627",
      "postDate": "09/22/2020 03:25:39",
      "content": "<p>Hello fellows,</p>\n<p>The loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:</p>\n<ul>\n<li><p>Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.</p></li>\n<li><p>Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.</p></li>\n<li><p>Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.</p></li>\n</ul>\n<p>Did you guys have the same problem? Any thoughts?</p>\n<p>Thanks!<br>\nFrank</p>\n<p>Update on 14 Oct:<br>\nFor me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.<br>\nNow I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄</p>",
      "rawMarkdown": "Hello fellows,\n\nThe loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:\n\n - Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.\n\n - Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.\n\n - Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.\n\nDid you guys have the same problem? Any thoughts?\n\nThanks!\nFrank\n\nUpdate on 14 Oct:\nFor me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.\nNow I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄",
      "votes": null
    },
    {
      "id": "1021743",
      "postDate": "09/22/2020 05:39:55",
      "content": "<p>Same for me.<br>\nI think the problem is information leaking. If you use the <code>train.zarr</code> or the <code>train_full.zar</code>, you have samples like these:</p>\n<ul>\n<li>Current: scene 1 - frame 14; history: frame 4-13; target: frame 15-64</li>\n<li>Current: scene 1 - frame 15; history: frame 5-14; target: frame: 16-65</li>\n</ul>\n<p>There is an overlap in the samples.</p>\n<p>What was your validation score for your current LB (63.481)?</p>",
      "rawMarkdown": "Same for me.\nI think the problem is information leaking. If you use the `train.zarr` or the `train_full.zar`, you have samples like these:\n\n- Current: scene 1 - frame 14; history: frame 4-13; target: frame 15-64\n- Current: scene 1 - frame 15; history: frame 5-14; target: frame: 16-65\n\nThere is an overlap in the samples.\n\nWhat was your validation score for your current LB (63.481)?",
      "votes": null
    },
    {
      "id": "1021789",
      "postDate": "09/22/2020 06:14:17",
      "content": "<p>Run <code>create_chopped_dataset</code> vs. the validation set and use that for validation. What's the loss (nll) there?</p>",
      "rawMarkdown": "Run `create_chopped_dataset` vs. the validation set and use that for validation. What's the loss (nll) there?",
      "votes": null
    },
    {
      "id": "1021800",
      "postDate": "09/22/2020 06:20:52",
      "content": "<p>Good point! I have never thought of this. Thanks Peter!</p>\n<p>So this is really related to my third point? In the evaluation dataset, we only sample agents from the 100th frame, while in the training dataset we sample angents from all frames, thus there are many overlaps.</p>\n<p>I am not completely convinced this is the main reason, but if it is, removing these overlaps might help improving our training greatly?</p>\n<p>Yes, my current LB score is 63.5, which was achieved by a training loss of 20. 😂 </p>",
      "rawMarkdown": "Good point! I have never thought of this. Thanks Peter!\n\nSo this is really related to my third point? In the evaluation dataset, we only sample agents from the 100th frame, while in the training dataset we sample angents from all frames, thus there are many overlaps.\n\nI am not completely convinced this is the main reason, but if it is, removing these overlaps might help improving our training greatly?\n\nYes, my current LB score is 63.5, which was achieved by a training loss of 20. 😂",
      "votes": null
    },
    {
      "id": "1021804",
      "postDate": "09/22/2020 06:22:54",
      "content": "<p>The same. Huge deviation.</p>",
      "rawMarkdown": "The same. Huge deviation.",
      "votes": null
    },
    {
      "id": "1021813",
      "postDate": "09/22/2020 06:29:26",
      "content": "<p>Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. <br>\nAre you shuffling the training set? If not, <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> has a point with the overlap. </p>",
      "rawMarkdown": "Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. \nAre you shuffling the training set? If not, @pestipeti has a point with the overlap.",
      "votes": null
    },
    {
      "id": "1021827",
      "postDate": "09/22/2020 06:37:25",
      "content": "<p>Thanks for the feedback, llu. This is helpful information.</p>",
      "rawMarkdown": "Thanks for the feedback, llu. This is helpful information.",
      "votes": null
    },
    {
      "id": "1021843",
      "postDate": "09/22/2020 06:54:42",
      "content": "<p>Have you tried cutout?</p>",
      "rawMarkdown": "Have you tried cutout?",
      "votes": null
    },
    {
      "id": "1021905",
      "postDate": "09/22/2020 07:38:27",
      "content": "<p>Hi Peter, what do you mean by \"cutout\"? I tried dropout, that seemed to get a training loss close to validation loss, but the result was not good.</p>",
      "rawMarkdown": "Hi Peter, what do you mean by \"cutout\"? I tried dropout, that seemed to get a training loss close to validation loss, but the result was not good.",
      "votes": null
    },
    {
      "id": "1021923",
      "postDate": "09/22/2020 08:01:27",
      "content": "<p><a href=\"https://arxiv.org/pdf/1708.04552v2.pdf\" target=\"_blank\">Cutout</a> is an augmentation technique. In short, you cut out random number/size squares (or rectangles) from the input image. You can make \"less similar\" images. In the image below the original images almost identical (frame 0, frame 1) but if you cut out random parts the result will be different. Of course, you have to find the right number/size of cuts, but I think it would help with the overlapping problem. (I haven't tried it)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3bc3975618543fd778719139878cb17%2Fcutout.png?generation=1600761651764733&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "[Cutout](https://arxiv.org/pdf/1708.04552v2.pdf) is an augmentation technique. In short, you cut out random number/size squares (or rectangles) from the input image. You can make \"less similar\" images. In the image below the original images almost identical (frame 0, frame 1) but if you cut out random parts the result will be different. Of course, you have to find the right number/size of cuts, but I think it would help with the overlapping problem. (I haven't tried it)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3bc3975618543fd778719139878cb17%2Fcutout.png?generation=1600761651764733&alt=media)",
      "votes": null
    },
    {
      "id": "1021974",
      "postDate": "09/22/2020 08:45:10",
      "content": "<p>Thanks for the good suggestions Peter. You have been so helpful since the beginning of this competition.</p>",
      "rawMarkdown": "Thanks for the good suggestions Peter. You have been so helpful since the beginning of this competition.",
      "votes": null
    },
    {
      "id": "1022114",
      "postDate": "09/22/2020 10:17:00",
      "content": "<p>I observed the same problem and tried to deep dive yesterday. I found out that the root of my issue was incorrect <code>min_frame_future</code>, I used either 1 or 50, but it probably should be 10, as mentioned by others. </p>\n<p>In more detail, <code>min_frame_future</code> and <code>min_frame_history</code> which are parameters for <code>AgentDataset</code> are responsible for cutting <code>agents_mask</code>. I was very surprised that decreasing <code>min_frame_future</code> actually increases the loss. Intuitively, it should be the opposite, because if we add samples with only short future available, we should have had increased accuracy for two reasons:</p>\n<ol>\n<li>Close-future points can be more accurately predicted. Further into the future - less accurate.</li>\n<li>Points with availability 0 are counted as error 0.</li>\n</ol>\n<p>And that does happen, but it is a smaller effect. The bigger effect stems from the fact that <code>availability</code> and <code>agents_mask</code> are very different things. I initially assumed that it was the same, but in fact <code>agents_mask</code> are all the points which are:</p>\n<ol>\n<li>available</li>\n<li>pass perception threshold on agents</li>\n<li>pass max absolute distance in degree</li>\n<li>pass max change in area allowed</li>\n<li>pass max distance from AV in meters</li>\n</ol>\n<p>The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing <code>min_frame_future</code> we introduce all these noisy (and irrelevant) agents which mess up with our loss. </p>\n<p>Summarizing, we probably want to have <code>min_frame_future</code> and <code>min_frame_history</code> at exactly same values as the test, which is probably 10 and 10. But also applying <code>create_chopped_dataset</code> should align them even more. </p>",
      "rawMarkdown": "I observed the same problem and tried to deep dive yesterday. I found out that the root of my issue was incorrect `min_frame_future`, I used either 1 or 50, but it probably should be 10, as mentioned by others. \n\nIn more detail, `min_frame_future` and `min_frame_history` which are parameters for `AgentDataset` are responsible for cutting `agents_mask`. I was very surprised that decreasing `min_frame_future` actually increases the loss. Intuitively, it should be the opposite, because if we add samples with only short future available, we should have had increased accuracy for two reasons:\n1. Close-future points can be more accurately predicted. Further into the future - less accurate.\n2. Points with availability 0 are counted as error 0.\n\nAnd that does happen, but it is a smaller effect. The bigger effect stems from the fact that `availability` and `agents_mask` are very different things. I initially assumed that it was the same, but in fact `agents_mask` are all the points which are:\n1. available\n2. pass perception threshold on agents\n3. pass max absolute distance in degree\n4. pass max change in area allowed\n5. pass max distance from AV in meters\n\nThe sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing `min_frame_future` we introduce all these noisy (and irrelevant) agents which mess up with our loss. \n\nSummarizing, we probably want to have `min_frame_future` and `min_frame_history` at exactly same values as the test, which is probably 10 and 10. But also applying `create_chopped_dataset` should align them even more.",
      "votes": null
    },
    {
      "id": "1022121",
      "postDate": "09/22/2020 10:24:27",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> , not sure you can call it leaking. It is just a correlation in the training set, it is allowed and perfectly fine. Conceptually, for the sake of arguments, one can look at it as data augmentation.</p>",
      "rawMarkdown": "pestipeti , not sure you can call it leaking. It is just a correlation in the training set, it is allowed and perfectly fine. Conceptually, for the sake of arguments, one can look at it as data augmentation.",
      "votes": null
    },
    {
      "id": "1022181",
      "postDate": "09/22/2020 11:13:06",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Yes, you are probably right. Leaking is not completely accurate. </p>",
      "rawMarkdown": "zaharch Yes, you are probably right. Leaking is not completely accurate.",
      "votes": null
    },
    {
      "id": "1022198",
      "postDate": "09/22/2020 11:25:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> great insights.</p>\n<p>But according to the finding of <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> here: <br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814</a><br>\nThe test dataset has lot of agents with less than 10 available history frames </p>",
      "rawMarkdown": "Hi @zaharch great insights.\n\nBut according to the finding of @ilu000 here: \nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\nThe test dataset has lot of agents with less than 10 available history frames",
      "votes": null
    },
    {
      "id": "1022251",
      "postDate": "09/22/2020 12:11:13",
      "content": "<p>Good point, <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> </p>\n<p>As <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> wrote<br>\n<code>test_dataset[34]['history_availabilities'].sum() = 3.0</code><br>\nI checked that if a mask is generated with default parameters for the test then it is not even 3 but goes down to 1. Does it mean that the test mask is generated with <code>min_frame_history = 1</code>?</p>",
      "rawMarkdown": "Good point, @frankpanxj \n\nAs @ilu000 wrote\n```test_dataset[34]['history_availabilities'].sum() = 3.0```\nI checked that if a mask is generated with default parameters for the test then it is not even 3 but goes down to 1. Does it mean that the test mask is generated with `min_frame_history = 1`?",
      "votes": null
    },
    {
      "id": "1022607",
      "postDate": "09/22/2020 16:28:20",
      "content": "<p>That was just an example. There are histories with len = 1 in the test set as well.<br>\nThe problem is, that the <code>create_chopped_dataset</code> function only applies a filter w.r.t. history frames in a scene (100 for this test set). That doesn't necessarily mean, that the agent was always visible. Some are only visible in the very last frame.</p>",
      "rawMarkdown": "That was just an example. There are histories with len = 1 in the test set as well.\nThe problem is, that the `create_chopped_dataset` function only applies a filter w.r.t. history frames in a scene (100 for this test set). That doesn't necessarily mean, that the agent was always visible. Some are only visible in the very last frame.",
      "votes": null
    },
    {
      "id": "1025285",
      "postDate": "09/24/2020 13:04:29",
      "content": "<blockquote>\n  <p>The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy (and irrelevant) agents which mess up with our loss. </p>\n</blockquote>\n<p>I don't think that's correct unless I'm misunderstanding the code. If you look at how <code>min_frame_future</code> is applied in <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/dataset/agent.py#L36\" target=\"_blank\"><code>AgentDataset.__init__</code></a> it's used to filter the <code>agents_mask</code> array. This is loaded from the dataset in <a href=\"https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L67\" target=\"_blank\"><code>load_agents_mask</code></a> and only recreated if you use a different <code>filter_agents_threshold</code> to the one in the dataset (0.5). In this case a new agents mask is created with <a href=\"https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/select_agents.py#L153\" target=\"_blank\"><code>select_agents</code></a> which applies all those other checks.<br>\nSo I don't think that varying the <code>min_frame_future</code> bypasses those checks. You always only get agents that pass all 5 checks, <code>min_frame_future</code> just re-calculates check 1 (and changing <code>filter_agents_threshold</code> would recalculate all the others).</p>",
      "rawMarkdown": "> The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy (and irrelevant) agents which mess up with our loss. \n\nI don't think that's correct unless I'm misunderstanding the code. If you look at how `min_frame_future` is applied in [`AgentDataset.__init__`](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/dataset/agent.py#L36) it's used to filter the `agents_mask` array. This is loaded from the dataset in [`load_agents_mask`](https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L67) and only recreated if you use a different `filter_agents_threshold` to the one in the dataset (0.5). In this case a new agents mask is created with [`select_agents`](https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/select_agents.py#L153) which applies all those other checks.\nSo I don't think that varying the `min_frame_future` bypasses those checks. You always only get agents that pass all 5 checks, `min_frame_future` just re-calculates check 1 (and changing `filter_agents_threshold` would recalculate all the others).",
      "votes": null
    },
    {
      "id": "1025408",
      "postDate": "09/24/2020 14:34:05",
      "content": "<p><a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> , first, I assume that all masks stored on disk can be reproduced exactly by calling <code>selected_agents</code> with default parameters. I haven't checked that. If true, it is just a cache for performance reasons. As a side note, of course the test mask can not be reproduced, because we don't have the future for the test.</p>\n<p>Regarding the second part of your statement. When I said </p>\n<blockquote>\n  <p>Therefore by decreasing <code>min_frame_future</code> we introduce all these noisy (and irrelevant) agents which mess up with our loss.</p>\n</blockquote>\n<p>I meant specifically examples where future availability is full 50, but future masks are now smaller because we decrease <code>min_frame_future</code>. Such examples are very noise, and there are many of them. </p>\n<p>I feel that I have not necessarily cleared your concern, if you still think I made a mistake in any sentence please continue your arguments.</p>",
      "rawMarkdown": "thomasbrandon , first, I assume that all masks stored on disk can be reproduced exactly by calling `selected_agents` with default parameters. I haven't checked that. If true, it is just a cache for performance reasons. As a side note, of course the test mask can not be reproduced, because we don't have the future for the test.\n\nRegarding the second part of your statement. When I said \n> Therefore by decreasing `min_frame_future` we introduce all these noisy (and irrelevant) agents which mess up with our loss.\n\nI meant specifically examples where future availability is full 50, but future masks are now smaller because we decrease `min_frame_future`. Such examples are very noise, and there are many of them. \n\nI feel that I have not necessarily cleared your concern, if you still think I made a mistake in any sentence please continue your arguments.",
      "votes": null
    },
    {
      "id": "1025473",
      "postDate": "09/24/2020 15:13:30",
      "content": "<p>Might be misunderstanding you, still not quite sure (and of course I may be misunderstanding the code).<br>\nI agree the <code>agents_mask</code> in the zarr datasets is just a performance thing (the <code>mask.npz</code> used in test is different, not talking about that).</p>\n<p>I thought when you said:</p>\n<blockquote>\n  <p>The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy …</p>\n</blockquote>\n<p>you were suggesting that if you pass a custom <code>min_frame_future</code>/<code>min_frame_history</code> then only check 1 (availability) will be be performed and you'll end up with lots of frames that don't pass the other checks. Whereas if you stick to the defaults then you get frames with all checks verified. That's what I was questioning. My understanding is that regardless of whether you pass a custom <code>min_frame_future</code>/<code>min_frame_history</code> you will only get frames with all checks passed. But of course with a lower <code>min_frame_future</code>/<code>min_frame_history</code> that will be a weaker test and may introduce issues.</p>\n<p>Though actually, diving into the logic of <code>select_agents</code>  and <code>get_valid_agents</code> I'm struggling to follow the logic, and may have been misunderstanding. Though not sure that affects my concern here. Will have to investigate further but that was my concern in posting.</p>",
      "rawMarkdown": "Might be misunderstanding you, still not quite sure (and of course I may be misunderstanding the code).\nI agree the `agents_mask` in the zarr datasets is just a performance thing (the `mask.npz` used in test is different, not talking about that).\n\nI thought when you said:\n> The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy ...\n\nyou were suggesting that if you pass a custom `min_frame_future`/`min_frame_history` then only check 1 (availability) will be be performed and you'll end up with lots of frames that don't pass the other checks. Whereas if you stick to the defaults then you get frames with all checks verified. That's what I was questioning. My understanding is that regardless of whether you pass a custom `min_frame_future`/`min_frame_history` you will only get frames with all checks passed. But of course with a lower `min_frame_future`/`min_frame_history` that will be a weaker test and may introduce issues.\n\nThough actually, diving into the logic of `select_agents`  and `get_valid_agents` I'm struggling to follow the logic, and may have been misunderstanding. Though not sure that affects my concern here. Will have to investigate further but that was my concern in posting.",
      "votes": null
    },
    {
      "id": "1025486",
      "postDate": "09/24/2020 15:24:25",
      "content": "<p>It is hard for me to pinpoint exactly where we diverge. So I will just share a few statements. First, <code>min_frame_future</code>/<code>min_frame_history</code> are not related to availability at all, they are thresholds on masks, which is availability plus additional checks, a subset but very different. Second, the problem of increased errors are not in the frame in question itself. The frame itself is of course passes both availability and the additional checks. The problem is what future follows it for the prediction purposes. And if what follows is both 50 masks and 50 availability, it is easy to predict. But if it is only 10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.</p>",
      "rawMarkdown": "It is hard for me to pinpoint exactly where we diverge. So I will just share a few statements. First, `min_frame_future`/`min_frame_history` are not related to availability at all, they are thresholds on masks, which is availability plus additional checks, a subset but very different. Second, the problem of increased errors are not in the frame in question itself. The frame itself is of course passes both availability and the additional checks. The problem is what future follows it for the prediction purposes. And if what follows is both 50 masks and 50 availability, it is easy to predict. But if it is only 10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.",
      "votes": null
    },
    {
      "id": "1025505",
      "postDate": "09/24/2020 15:42:32",
      "content": "<p>Ah, OK. I think I see the divergence. You're talking about noise in errors. I was talking about noise in data. So you say:</p>\n<blockquote>\n  <p>10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.</p>\n</blockquote>\n<p>A perfectly reasonable meaning for noise but not the one I had. I would have instead said that was 10 points with data and 40 empty points so not really noisy (not to say my definition is better, just to identify divergence).</p>\n<p>I thought you were saying that if you kept the default <code>min_future_frames=1</code> then you'd end up with one non-empty, non-noisy frame and 49 empty. But if you changed it to <code>min_future_frames=10</code> then it would stop performing check 2-5 and you'd get 1 non-noisy frame, 9 noisy frames (because of no additional checks) and then 40 empty frames.</p>",
      "rawMarkdown": "Ah, OK. I think I see the divergence. You're talking about noise in errors. I was talking about noise in data. So you say:\n\n> 10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.\n \nA perfectly reasonable meaning for noise but not the one I had. I would have instead said that was 10 points with data and 40 empty points so not really noisy (not to say my definition is better, just to identify divergence).\n\nI thought you were saying that if you kept the default `min_future_frames=1` then you'd end up with one non-empty, non-noisy frame and 49 empty. But if you changed it to `min_future_frames=10` then it would stop performing check 2-5 and you'd get 1 non-noisy frame, 9 noisy frames (because of no additional checks) and then 40 empty frames.",
      "votes": null
    },
    {
      "id": "1025535",
      "postDate": "09/24/2020 16:14:46",
      "content": "<p>I also think we might differ on the meaning of availability given the <code>target_availability</code> mask.<br>\nI think given a <code>target_availability</code> mask of size 50 where 10 values in it are 1 and then other 40 are 0 then you would say that 10 values are available (or maybe 40 if it was a negative mask but that's less likely). There aren't 50 available frames just because that's the size of the mask.</p>",
      "rawMarkdown": "I also think we might differ on the meaning of availability given the `target_availability` mask.\nI think given a `target_availability` mask of size 50 where 10 values in it are 1 and then other 40 are 0 then you would say that 10 values are available (or maybe 40 if it was a negative mask but that's less likely). There aren't 50 available frames just because that's the size of the mask.",
      "votes": null
    },
    {
      "id": "1035432",
      "postDate": "10/02/2020 17:50:12",
      "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> did you find any explanation for deviation other than the one mentioned here and something that doesnt involve creating custom mask? I have been training with 224*224 and 0.25. For 3M sample the loss is like 100+ and for 6M sample the loss goes up to 6000 nll</p>",
      "rawMarkdown": "frankpanxj did you find any explanation for deviation other than the one mentioned here and something that doesnt involve creating custom mask? I have been training with 224*224 and 0.25. For 3M sample the loss is like 100+ and for 6M sample the loss goes up to 6000 nll",
      "votes": null
    },
    {
      "id": "1035441",
      "postDate": "10/02/2020 17:58:09",
      "content": "<p>The sample.zarr can be easily overfitted. It's too small to actually train on. </p>",
      "rawMarkdown": "The sample.zarr can be easily overfitted. It's too small to actually train on.",
      "votes": null
    },
    {
      "id": "1035453",
      "postDate": "10/02/2020 18:06:42",
      "content": "<p>Well I am not using sample.Zarr. I have been on train_full</p>",
      "rawMarkdown": "Well I am not using sample.Zarr. I have been on train_full",
      "votes": null
    },
    {
      "id": "1035525",
      "postDate": "10/02/2020 19:06:37",
      "content": "<p>If you're getting a loss of 6000 then you likely have an error somewhere. That's the sort of loss you'd get with random predictions.<br>\nYou aren't trying to use a model trained for l5kit 1.0.6 with the newly released 1.1,0 are you? Models trained with the old version aren't compatible with the new version without modifications (and really they need to be re-trained on the new version as that will improve performance significantly).<br>\nOtherwise you'd need to debug to figure out why the model is going off the rails. Someone did report exploding gradients which would cause that sort of thing but I haven't experienced that (but I'm using my own re-implementation of the loss function and don't know what implementation they were using).</p>",
      "rawMarkdown": "If you're getting a loss of 6000 then you likely have an error somewhere. That's the sort of loss you'd get with random predictions.\nYou aren't trying to use a model trained for l5kit 1.0.6 with the newly released 1.1,0 are you? Models trained with the old version aren't compatible with the new version without modifications (and really they need to be re-trained on the new version as that will improve performance significantly).\nOtherwise you'd need to debug to figure out why the model is going off the rails. Someone did report exploding gradients which would cause that sort of thing but I haven't experienced that (but I'm using my own re-implementation of the loss function and don't know what implementation they were using).",
      "votes": null
    },
    {
      "id": "1035887",
      "postDate": "10/03/2020 07:33:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a>, no I did not find a good enough explanation and the problem is still not solved. I noticed that not all people experienced the same so there might be something I an missing.</p>",
      "rawMarkdown": "Hi @deepakrajpurushothaman, no I did not find a good enough explanation and the problem is still not solved. I noticed that not all people experienced the same so there might be something I an missing.",
      "votes": null
    },
    {
      "id": "1037421",
      "postDate": "10/05/2020 02:30:58",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> &gt; Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. &gt; </p>\n<p>How were you able to get those 3 to agree?</p>",
      "rawMarkdown": "ilu000 > Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. > \n\nHow were you able to get those 3 to agree?",
      "votes": null
    },
    {
      "id": "1038185",
      "postDate": "10/05/2020 16:02:44",
      "content": "<p>If you are using the latest l5kit (<a href=\"https://github.com/lyft/l5kit/releases/tag/v1.1.0\" target=\"_blank\">v1.1.0</a>), make sure you convert agent coordinates into world offsets. Refer to Evaluation block in <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">this example</a> or refer to <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187825\" target=\"_blank\">this notebook</a>. Otherwise, your loss on test set will be super huge.</p>",
      "rawMarkdown": "If you are using the latest l5kit ([v1.1.0](https://github.com/lyft/l5kit/releases/tag/v1.1.0)), make sure you convert agent coordinates into world offsets. Refer to Evaluation block in [this example](https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb) or refer to [this notebook](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187825). Otherwise, your loss on test set will be super huge.",
      "votes": null
    },
    {
      "id": "1038597",
      "postDate": "10/05/2020 22:56:21",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> , <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> and <a href=\"https://www.kaggle.com/arkaung\" target=\"_blank\">@arkaung</a> your inputs helps to think for sure . I will try to debug and try to find a reason. And about version, I am trying both l5kit 1.0.6 and 1.1.0 and taking care of the compatibility. the deviations exists both cases.  </p>",
      "rawMarkdown": "thanks @thomasbrandon , @frankpanxj and @arkaung your inputs helps to think for sure . I will try to debug and try to find a reason. And about version, I am trying both l5kit 1.0.6 and 1.1.0 and taking care of the compatibility. the deviations exists both cases.",
      "votes": null
    },
    {
      "id": "1039295",
      "postDate": "10/06/2020 13:28:11",
      "content": "<p>I have huge differences (between train loss and public LB) and trying to figure out what seems to be the problem too. </p>",
      "rawMarkdown": "I have huge differences (between train loss and public LB) and trying to figure out what seems to be the problem too.",
      "votes": null
    },
    {
      "id": "1043428",
      "postDate": "10/09/2020 01:25:47",
      "content": "<p>Similar experience for me! After <code>10, 000</code>loops, My training <code>NLL loss</code> is  about <code>100</code> which is very close to my <code>validate NLL loss</code> and the <code>LB</code>, but <code>the evaluate metic  score(neg_multi_NLL) is around 1, 000</code>, its a huge gap. </p>\n<p>My evaluate dataset is a subset (size is 1000*12) of the Chuck validate.zarr, obtained by torch.utill.data.Subset.</p>\n<p>I see other people got a very close result between evaulation and the LB. How you guys make it? Thanks</p>",
      "rawMarkdown": "Similar experience for me! After `10, 000 `loops, My training `NLL loss` is  about `100` which is very close to my `validate NLL loss` and the `LB`, but `the evaluate metic  score(neg_multi_NLL) is around 1, 000`, its a huge gap. \n\nMy evaluate dataset is a subset (size is 1000*12) of the Chuck validate.zarr, obtained by torch.utill.data.Subset.\n  \nI see other people got a very close result between evaulation and the LB. How you guys make it? Thanks",
      "votes": null
    },
    {
      "id": "1049045",
      "postDate": "10/14/2020 03:58:40",
      "content": "<p>I met the same problem too. </p>",
      "rawMarkdown": "I met the same problem too.",
      "votes": null
    },
    {
      "id": "1053287",
      "postDate": "10/18/2020 20:10:03",
      "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> did you make any progress regarding this topic? I still face the same problem… avg loss around ~25 and leaderboard around ~50</p>",
      "rawMarkdown": "frankpanxj did you make any progress regarding this topic? I still face the same problem... avg loss around ~25 and leaderboard around ~50",
      "votes": null
    },
    {
      "id": "1053636",
      "postDate": "10/19/2020 07:27:11",
      "content": "<p>Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation…</p>",
      "rawMarkdown": "Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation...",
      "votes": null
    },
    {
      "id": "1058086",
      "postDate": "10/23/2020 10:19:14",
      "content": "<p>update:Oct 23rd:<br>\ntrain cfg:<br>\n    'model_params': {<br>\n        'model_architecture': 'resnet34',<br>\n        'history_num_frames': 10,<br>\n        'history_step_size': 1,<br>\n…<br>\n    'raster_params': {<br>\n        'raster_size': [480, 360],<br>\n        'pixel_size': [0.5, 0.5],<br>\n        'ego_center': [0.25, 0.5],<br>\n…<br>\n    'val_data_loader': {<br>\n        'key': 'scenes/validate.zarr',<br>\n        'batch_size': 32,<br>\n        'shuffle': False,<br>\n        'num_workers': 16<br>\n…<br>\ntrained for 600k iterations, train NLL loss is around 10, evaluation NLL loss around 30, LB score 26~28.<br>\nevaluation method:<br>\ncreate_chopped_dataset<br>\nnum_frames_to_chop = 100<br>\nMIN_FUTURE_STEPS = 10<br>\nMIN_FRAME_FUTURE = 10<br>\nMIN_FRAME_HISTORY = 10.<br>\nI guess there are some difference between training and validation. still working on the problem….</p>",
      "rawMarkdown": "update:Oct 23rd:\ntrain cfg:\n    'model_params': {\n        'model_architecture': 'resnet34',\n        'history_num_frames': 10,\n        'history_step_size': 1,\n...\n    'raster_params': {\n        'raster_size': [480, 360],\n        'pixel_size': [0.5, 0.5],\n        'ego_center': [0.25, 0.5],\n...\n    'val_data_loader': {\n        'key': 'scenes/validate.zarr',\n        'batch_size': 32,\n        'shuffle': False,\n        'num_workers': 16\n...\ntrained for 600k iterations, train NLL loss is around 10, evaluation NLL loss around 30, LB score 26~28.\nevaluation method:\ncreate_chopped_dataset\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\nMIN_FRAME_FUTURE = 10\nMIN_FRAME_HISTORY = 10.\nI guess there are some difference between training and validation. still working on the problem....",
      "votes": null
    },
    {
      "id": "1058110",
      "postDate": "10/23/2020 10:58:05",
      "content": "<p>Could you post the full config? Are you using shuffle in your training loop? If not, this may explain the large difference. </p>",
      "rawMarkdown": "Could you post the full config? Are you using shuffle in your training loop? If not, this may explain the large difference.",
      "votes": null
    },
    {
      "id": "1058160",
      "postDate": "10/23/2020 12:16:08",
      "content": "<p><a href=\"https://www.kaggle.com/xiaoyaopeng\" target=\"_blank\">@xiaoyaopeng</a> same here. Training loss ~11 after 600K iterations but LB in upper 20's, my raster is smaller though 224x224.<br>\nshuffle = True</p>",
      "rawMarkdown": "xiaoyaopeng same here. Training loss ~11 after 600K iterations but LB in upper 20's, my raster is smaller though 224x224.\nshuffle = True",
      "votes": null
    },
    {
      "id": "1058166",
      "postDate": "10/23/2020 12:21:26",
      "content": "<p>600k on train or train_full? <br>\nfor train, you can already see overfitting/memorization at that point. </p>",
      "rawMarkdown": "600k on train or train_full? \nfor train, you can already see overfitting/memorization at that point.",
      "votes": null
    },
    {
      "id": "1058172",
      "postDate": "10/23/2020 12:27:49",
      "content": "<p>unfortunately, train.csv in my case</p>",
      "rawMarkdown": "unfortunately, train.csv in my case",
      "votes": null
    },
    {
      "id": "1058357",
      "postDate": "10/23/2020 15:51:51",
      "content": "<p>training config:    <br>\n'train_data_loader': {<br>\n        'key': 'scenes/train.zarr',<br>\n        'batch_size': 32,<br>\n        'shuffle': True,<br>\n        'num_workers': 16<br>\n    },</p>",
      "rawMarkdown": "training config:    \n'train_data_loader': {\n        'key': 'scenes/train.zarr',\n        'batch_size': 32,\n        'shuffle': True,\n        'num_workers': 16\n    },",
      "votes": null
    },
    {
      "id": "1058823",
      "postDate": "10/24/2020 10:29:52",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> a bit of a simple (and to some obvious) question here, but how would one correctly detect overfitting in this type of problem?  I would assume a continuous decrease in training loss with a sudden increase in evaluation loss, which can be found only by experimentation - early stopping the training a number of times.</p>\n<p>Since I constantly run experiments on minimal number of samples, I am at minimal risk of going into the overfitting territory, therefore not very familiar with it. Except during that one time I filtered way too many agents…</p>",
      "rawMarkdown": "ilu000 a bit of a simple (and to some obvious) question here, but how would one correctly detect overfitting in this type of problem?  I would assume a continuous decrease in training loss with a sudden increase in evaluation loss, which can be found only by experimentation - early stopping the training a number of times.\n\nSince I constantly run experiments on minimal number of samples, I am at minimal risk of going into the overfitting territory, therefore not very familiar with it. Except during that one time I filtered way too many agents...",
      "votes": null
    },
    {
      "id": "1059623",
      "postDate": "10/25/2020 09:39:35",
      "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> this is indeed the definition of overfitting. It doesn't need to be a sudden increase in evaluation loss as it can be a slow increase, too. The model fails to generalize to unseen data and starts to memorize training data almost perfectly. Thus, the training loss can reach very very low values while the evaluation score on unseen data begins to increase. You can tackle that problem with augmentation or lots of data (amongst other things). We have the insanly large dataset here, so we can make full use of it. If there is no way around the problem, early stopping is a valid technique.</p>",
      "rawMarkdown": "indswetrust this is indeed the definition of overfitting. It doesn't need to be a sudden increase in evaluation loss as it can be a slow increase, too. The model fails to generalize to unseen data and starts to memorize training data almost perfectly. Thus, the training loss can reach very very low values while the evaluation score on unseen data begins to increase. You can tackle that problem with augmentation or lots of data (amongst other things). We have the insanly large dataset here, so we can make full use of it. If there is no way around the problem, early stopping is a valid technique.",
      "votes": null
    },
    {
      "id": "1059810",
      "postDate": "10/25/2020 13:44:38",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> thanks for the clarification! I think augmentation could be useful for overfitting prevention in this problem if the original data is filtered from \"noisy\" inputs, eg filtering agents based on various criteria to avoid irrelevant training inputs. Since we will have a smaller number of inputs, memorization is more likely and augmentation might help in this case.</p>\n<p>Otherwise, if the input train dataset is relatively untouched, augmentation is probably not so important  due to the amount of data (as raised in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042\" target=\"_blank\">this discussion</a>) and early stopping may be sufficient.<br>\nBut as we know theory and practice can differ, so everything is up for experimentation;)</p>",
      "rawMarkdown": "ilu000 thanks for the clarification! I think augmentation could be useful for overfitting prevention in this problem if the original data is filtered from \"noisy\" inputs, eg filtering agents based on various criteria to avoid irrelevant training inputs. Since we will have a smaller number of inputs, memorization is more likely and augmentation might help in this case.\n\nOtherwise, if the input train dataset is relatively untouched, augmentation is probably not so important  due to the amount of data (as raised in [this discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042)) and early stopping may be sufficient.\nBut as we know theory and practice can differ, so everything is up for experimentation;)",
      "votes": null
    },
    {
      "id": "1065243",
      "postDate": "10/31/2020 04:22:35",
      "content": "<blockquote>\n  <p>Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation…</p>\n</blockquote>\n<p>In addition to the similar performance on training loss and LB score/validation loss, can it help to improve the LB score finally?</p>",
      "rawMarkdown": "> Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation...\n\nIn addition to the similar performance on training loss and LB score/validation loss, can it help to improve the LB score finally?",
      "votes": null
    },
    {
      "id": "1065244",
      "postDate": "10/31/2020 04:25:00",
      "content": "<p>Maybe you need to adjust the pixel size correspondingly.</p>",
      "rawMarkdown": "Maybe you need to adjust the pixel size correspondingly.",
      "votes": null
    },
    {
      "id": "1074009",
      "postDate": "11/10/2020 06:24:06",
      "content": "<p>Are you guys talking about the config parameter <code>future_num_frames</code>?</p>",
      "rawMarkdown": "Are you guys talking about the config parameter `future_num_frames`?",
      "votes": null
    },
    {
      "id": "1074023",
      "postDate": "11/10/2020 06:48:21",
      "content": "<p>Nope, min_frame_future, min_frame_history are AgentDataset class parameters. Look at the AgentDataset() class in the l5kit and you'll see whats going on.</p>",
      "rawMarkdown": "Nope, min_frame_future, min_frame_history are AgentDataset class parameters. Look at the AgentDataset() class in the l5kit and you'll see whats going on.",
      "votes": null
    },
    {
      "id": "1074320",
      "postDate": "11/10/2020 14:12:06",
      "content": "<blockquote>\n  <p>Hello fellows,</p>\n  <p>The loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:</p>\n  <ul>\n  <li><p>Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.</p></li>\n  <li><p>Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.</p></li>\n  <li><p>Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.</p></li>\n  </ul>\n  <p>Did you guys have the same problem? Any thoughts?</p>\n  <p>Thanks!<br>\n  Frank</p>\n  <p>Update on 14 Oct:<br>\n  For me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.<br>\n  Now I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄</p>\n</blockquote>\n<p>Hi, I also meet this problem. But before this, I try baseline only using images without \"history_positions\" and the validation value is fine and reasonable. After adding history information, I also have a huge deviation between train loss and validation value. How you fix this problem finally?</p>",
      "rawMarkdown": "> Hello fellows,\n> \n> The loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:\n> \n>  - Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.\n> \n>  - Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.\n> \n>  - Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.\n> \n> Did you guys have the same problem? Any thoughts?\n> \n> Thanks!\n> Frank\n> \n> Update on 14 Oct:\n> For me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.\n> Now I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄\n\nHi, I also meet this problem. But before this, I try baseline only using images without \"history_positions\" and the validation value is fine and reasonable. After adding history information, I also have a huge deviation between train loss and validation value. How you fix this problem finally?",
      "votes": null
    },
    {
      "id": "1075049",
      "postDate": "11/11/2020 11:04:33",
      "content": "<p>Hi, I am currently not using history_positions as input. I can think of two possible reasons why this does not work. First, there are many agents with very short history, as you can see from some of the earlier discussions here. So, the model cannot really rely on this data to predict. Second, to feed in this data, we need to add some non-linear layers on top of the standard Resnet, which, based on my limited experiments, does not do any good (if not harm).</p>\n<p>Please note that I did not do enough experiments to be sure that this won't work.</p>",
      "rawMarkdown": "Hi, I am currently not using history_positions as input. I can think of two possible reasons why this does not work. First, there are many agents with very short history, as you can see from some of the earlier discussions here. So, the model cannot really rely on this data to predict. Second, to feed in this data, we need to add some non-linear layers on top of the standard Resnet, which, based on my limited experiments, does not do any good (if not harm).\n\nPlease note that I did not do enough experiments to be sure that this won't work.",
      "votes": null
    },
    {
      "id": "1075122",
      "postDate": "11/11/2020 12:07:22",
      "content": "<p>Yeah. Thank you for your reply. I do some experiments and it does confirm my thought that the different \"history_availabilities\" in train.zarr and chopped validation set can cause some data mismatch problems which could cause a huge deviation between train loss and actual LB score/validation score.</p>",
      "rawMarkdown": "Yeah. Thank you for your reply. I do some experiments and it does confirm my thought that the different \"history_availabilities\" in train.zarr and chopped validation set can cause some data mismatch problems which could cause a huge deviation between train loss and actual LB score/validation score.",
      "votes": null
    },
    {
      "id": "1075150",
      "postDate": "11/11/2020 12:53:25",
      "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> hi there, after reading the post I could grasp a better understanding of the terms. Thanks for your help though!</p>",
      "rawMarkdown": "indswetrust hi there, after reading the post I could grasp a better understanding of the terms. Thanks for your help though!",
      "votes": null
    },
    {
      "id": "1075155",
      "postDate": "11/11/2020 12:56:51",
      "content": "<p>I once used both the original train.zarr &amp; validate.zarr Chunkdataset, my training loss &amp; validate loss matched each other pretty. 1. Check your loss function, may you're using different ones. 2. You may need more training steps/training longer.</p>",
      "rawMarkdown": "I once used both the original train.zarr & validate.zarr Chunkdataset, my training loss & validate loss matched each other pretty. 1. Check your loss function, may you're using different ones. 2. You may need more training steps/training longer.",
      "votes": null
    },
    {
      "id": "1075162",
      "postDate": "11/11/2020 13:03:37",
      "content": "<p>For example, how to adjust the pixel size? You mean adjust it w.r.t raster size?</p>",
      "rawMarkdown": "For example, how to adjust the pixel size? You mean adjust it w.r.t raster size?",
      "votes": null
    },
    {
      "id": "1075168",
      "postDate": "11/11/2020 13:09:31",
      "content": "<p>So could you share some observations? </p>",
      "rawMarkdown": "So could you share some observations?",
      "votes": null
    },
    {
      "id": "1075943",
      "postDate": "11/12/2020 05:34:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/fuckvenkatraman\" target=\"_blank\">@fuckvenkatraman</a> , raster size selection is already discussed in this post <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\" target=\"_blank\">Raster size selection</a>. Basically, idea is to increase your width and have lesser height. (since ego vehicle moves primarily in x-direction). Also adjust the pixel_size parameter (which corresponds to how much distance a pixel means, default 1 pixel equals 0.5m). By adjusting both these parameters, you could focus more on your region of interest.</p>",
      "rawMarkdown": "Hi @fuckvenkatraman , raster size selection is already discussed in this post [Raster size selection](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323). Basically, idea is to increase your width and have lesser height. (since ego vehicle moves primarily in x-direction). Also adjust the pixel_size parameter (which corresponds to how much distance a pixel means, default 1 pixel equals 0.5m). By adjusting both these parameters, you could focus more on your region of interest.",
      "votes": null
    },
    {
      "id": "1076101",
      "postDate": "11/12/2020 09:08:25",
      "content": "<p><a href=\"https://www.kaggle.com/fuckvenkatraman\" target=\"_blank\">@fuckvenkatraman</a> Are you asking for my observation?</p>",
      "rawMarkdown": "fuckvenkatraman Are you asking for my observation?",
      "votes": null
    },
    {
      "id": "1077089",
      "postDate": "11/13/2020 08:21:53",
      "content": "<p><a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a> Hi, How to set min_frame_future=10? Do I need to call create_chopped_dataset function to train dataset?</p>",
      "rawMarkdown": "yousof9 Hi, How to set min_frame_future=10? Do I need to call create_chopped_dataset function to train dataset?",
      "votes": null
    },
    {
      "id": "1077097",
      "postDate": "11/13/2020 08:32:44",
      "content": "<p>No, it can be passed as an argument to Agent dataset class. By default the value is 1. You can refer to l5kit github page for more information. </p>",
      "rawMarkdown": "No, it can be passed as an argument to Agent dataset class. By default the value is 1. You can refer to l5kit github page for more information.",
      "votes": null
    },
    {
      "id": "1081109",
      "postDate": "11/16/2020 20:18:23",
      "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> Yeah might be interesting in this case. Still struggling to figure out how to deal with the mismatches. Additional data didn't improve it for me.</p>",
      "rawMarkdown": "spicychicken38 Yeah might be interesting in this case. Still struggling to figure out how to deal with the mismatches. Additional data didn't improve it for me.",
      "votes": null
    },
    {
      "id": "1081110",
      "postDate": "11/16/2020 20:19:24",
      "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a></p>\n<p>If you do it the right way, I think yes. However, I didn't manage to improve my score with increasing raster sizes..</p>",
      "rawMarkdown": "spicychicken38\n\nIf you do it the right way, I think yes. However, I didn't manage to improve my score with increasing raster sizes..",
      "votes": null
    },
    {
      "id": "1084377",
      "postDate": "11/20/2020 01:26:32",
      "content": "<p><a href=\"https://www.kaggle.com/xiaoyaopeng\" target=\"_blank\">@xiaoyaopeng</a> were you also using history positions as input to the model when you encountered this problem? </p>",
      "rawMarkdown": "xiaoyaopeng were you also using history positions as input to the model when you encountered this problem?",
      "votes": null
    },
    {
      "id": "1091664",
      "postDate": "11/26/2020 07:24:28",
      "content": "<p>no, I just run the baseline notebook and made few changes.<br>\nI successfully made eval loss and training loss close by chop the training data.<br>\nI think <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> is right. there are overlaps in the training data.</p>",
      "rawMarkdown": "no, I just run the baseline notebook and made few changes.\nI successfully made eval loss and training loss close by chop the training data.\nI think @pestipeti is right. there are overlaps in the training data.",
      "votes": null
    },
    {
      "id": "1091752",
      "postDate": "11/26/2020 08:55:13",
      "content": "<p>From what we have seen, a large deviation between train and validation / test score originates from different sample distributions. To match test, for AgentDataset we used <br>\n<code>min_frame_history=1</code><br>\n<code>min_frame_future=10</code><br>\nin validation and in most training setups</p>\n<p>Not to be confused with <code>history_num_frames</code></p>",
      "rawMarkdown": "From what we have seen, a large deviation between train and validation / test score originates from different sample distributions. To match test, for AgentDataset we used \n`min_frame_history=1`\n`min_frame_future=10`\nin validation and in most training setups\n\nNot to be confused with `history_num_frames`",
      "votes": null
    },
    {
      "id": "1091832",
      "postDate": "11/26/2020 10:21:39",
      "content": "<p>Yeah, I firmly believe that such a huge deviation is caused by data mismatch, but I dealt with it with a more complicated method… I looked at the probability distribution of history_availailities in the test set and try to reconcile the probability distribution in the training set with that in the test set…</p>",
      "rawMarkdown": "Yeah, I firmly believe that such a huge deviation is caused by data mismatch, but I dealt with it with a more complicated method... I looked at the probability distribution of history_availailities in the test set and try to reconcile the probability distribution in the training set with that in the test set...",
      "votes": null
    },
    {
      "id": "1091959",
      "postDate": "11/26/2020 12:21:07",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> , I actually did that, I even set the min_frame_history=0, which surprisingly still could increase valid agents. So, for your final submission, what training loss and validation loss did you get?</p>",
      "rawMarkdown": "Thanks @ilu000 , I actually did that, I even set the min_frame_history=0, which surprisingly still could increase valid agents. So, for your final submission, what training loss and validation loss did you get?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1021743,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "09/22/2020 05:39:55",
      "content": "<p>Same for me.<br>\nI think the problem is information leaking. If you use the <code>train.zarr</code> or the <code>train_full.zar</code>, you have samples like these:</p>\n<ul>\n<li>Current: scene 1 - frame 14; history: frame 4-13; target: frame 15-64</li>\n<li>Current: scene 1 - frame 15; history: frame 5-14; target: frame: 16-65</li>\n</ul>\n<p>There is an overlap in the samples.</p>\n<p>What was your validation score for your current LB (63.481)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1021800,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "09/22/2020 06:20:52",
          "content": "<p>Good point! I have never thought of this. Thanks Peter!</p>\n<p>So this is really related to my third point? In the evaluation dataset, we only sample agents from the 100th frame, while in the training dataset we sample angents from all frames, thus there are many overlaps.</p>\n<p>I am not completely convinced this is the main reason, but if it is, removing these overlaps might help improving our training greatly?</p>\n<p>Yes, my current LB score is 63.5, which was achieved by a training loss of 20. 😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022121,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/22/2020 10:24:27",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> , not sure you can call it leaking. It is just a correlation in the training set, it is allowed and perfectly fine. Conceptually, for the sake of arguments, one can look at it as data augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022181,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "09/22/2020 11:13:06",
          "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Yes, you are probably right. Leaking is not completely accurate. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1021789,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "09/22/2020 06:14:17",
      "content": "<p>Run <code>create_chopped_dataset</code> vs. the validation set and use that for validation. What's the loss (nll) there?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1021804,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "09/22/2020 06:22:54",
          "content": "<p>The same. Huge deviation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1021813,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "09/22/2020 06:29:26",
          "content": "<p>Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. <br>\nAre you shuffling the training set? If not, <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> has a point with the overlap. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1021827,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "09/22/2020 06:37:25",
          "content": "<p>Thanks for the feedback, llu. This is helpful information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1021843,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "09/22/2020 06:54:42",
          "content": "<p>Have you tried cutout?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1021905,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "09/22/2020 07:38:27",
          "content": "<p>Hi Peter, what do you mean by \"cutout\"? I tried dropout, that seemed to get a training loss close to validation loss, but the result was not good.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1021923,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "09/22/2020 08:01:27",
          "content": "<p><a href=\"https://arxiv.org/pdf/1708.04552v2.pdf\" target=\"_blank\">Cutout</a> is an augmentation technique. In short, you cut out random number/size squares (or rectangles) from the input image. You can make \"less similar\" images. In the image below the original images almost identical (frame 0, frame 1) but if you cut out random parts the result will be different. Of course, you have to find the right number/size of cuts, but I think it would help with the overlapping problem. (I haven't tried it)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3bc3975618543fd778719139878cb17%2Fcutout.png?generation=1600761651764733&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1021974,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "09/22/2020 08:45:10",
          "content": "<p>Thanks for the good suggestions Peter. You have been so helpful since the beginning of this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1037421,
          "author_name": "yousof9",
          "author_url": "",
          "post_date": "10/05/2020 02:30:58",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> &gt; Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. &gt; </p>\n<p>How were you able to get those 3 to agree?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1074009,
          "author_name": "fuckvenkatraman",
          "author_url": "",
          "post_date": "11/10/2020 06:24:06",
          "content": "<p>Are you guys talking about the config parameter <code>future_num_frames</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1074023,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "11/10/2020 06:48:21",
          "content": "<p>Nope, min_frame_future, min_frame_history are AgentDataset class parameters. Look at the AgentDataset() class in the l5kit and you'll see whats going on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075150,
          "author_name": "fuckvenkatraman",
          "author_url": "",
          "post_date": "11/11/2020 12:53:25",
          "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> hi there, after reading the post I could grasp a better understanding of the terms. Thanks for your help though!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1077089,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "11/13/2020 08:21:53",
          "content": "<p><a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a> Hi, How to set min_frame_future=10? Do I need to call create_chopped_dataset function to train dataset?</p>",
          "votes": null,
          "replies": [
            {
              "id": 1077097,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "11/13/2020 08:32:44",
              "content": "<p>No, it can be passed as an argument to Agent dataset class. By default the value is 1. You can refer to l5kit github page for more information. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1022114,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "09/22/2020 10:17:00",
      "content": "<p>I observed the same problem and tried to deep dive yesterday. I found out that the root of my issue was incorrect <code>min_frame_future</code>, I used either 1 or 50, but it probably should be 10, as mentioned by others. </p>\n<p>In more detail, <code>min_frame_future</code> and <code>min_frame_history</code> which are parameters for <code>AgentDataset</code> are responsible for cutting <code>agents_mask</code>. I was very surprised that decreasing <code>min_frame_future</code> actually increases the loss. Intuitively, it should be the opposite, because if we add samples with only short future available, we should have had increased accuracy for two reasons:</p>\n<ol>\n<li>Close-future points can be more accurately predicted. Further into the future - less accurate.</li>\n<li>Points with availability 0 are counted as error 0.</li>\n</ol>\n<p>And that does happen, but it is a smaller effect. The bigger effect stems from the fact that <code>availability</code> and <code>agents_mask</code> are very different things. I initially assumed that it was the same, but in fact <code>agents_mask</code> are all the points which are:</p>\n<ol>\n<li>available</li>\n<li>pass perception threshold on agents</li>\n<li>pass max absolute distance in degree</li>\n<li>pass max change in area allowed</li>\n<li>pass max distance from AV in meters</li>\n</ol>\n<p>The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing <code>min_frame_future</code> we introduce all these noisy (and irrelevant) agents which mess up with our loss. </p>\n<p>Summarizing, we probably want to have <code>min_frame_future</code> and <code>min_frame_history</code> at exactly same values as the test, which is probably 10 and 10. But also applying <code>create_chopped_dataset</code> should align them even more. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1022198,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "09/22/2020 11:25:28",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> great insights.</p>\n<p>But according to the finding of <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> here: <br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814</a><br>\nThe test dataset has lot of agents with less than 10 available history frames </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022251,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/22/2020 12:11:13",
          "content": "<p>Good point, <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> </p>\n<p>As <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> wrote<br>\n<code>test_dataset[34]['history_availabilities'].sum() = 3.0</code><br>\nI checked that if a mask is generated with default parameters for the test then it is not even 3 but goes down to 1. Does it mean that the test mask is generated with <code>min_frame_history = 1</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022607,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "09/22/2020 16:28:20",
          "content": "<p>That was just an example. There are histories with len = 1 in the test set as well.<br>\nThe problem is, that the <code>create_chopped_dataset</code> function only applies a filter w.r.t. history frames in a scene (100 for this test set). That doesn't necessarily mean, that the agent was always visible. Some are only visible in the very last frame.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025285,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "09/24/2020 13:04:29",
          "content": "<blockquote>\n  <p>The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy (and irrelevant) agents which mess up with our loss. </p>\n</blockquote>\n<p>I don't think that's correct unless I'm misunderstanding the code. If you look at how <code>min_frame_future</code> is applied in <a href=\"https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/dataset/agent.py#L36\" target=\"_blank\"><code>AgentDataset.__init__</code></a> it's used to filter the <code>agents_mask</code> array. This is loaded from the dataset in <a href=\"https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L67\" target=\"_blank\"><code>load_agents_mask</code></a> and only recreated if you use a different <code>filter_agents_threshold</code> to the one in the dataset (0.5). In this case a new agents mask is created with <a href=\"https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/select_agents.py#L153\" target=\"_blank\"><code>select_agents</code></a> which applies all those other checks.<br>\nSo I don't think that varying the <code>min_frame_future</code> bypasses those checks. You always only get agents that pass all 5 checks, <code>min_frame_future</code> just re-calculates check 1 (and changing <code>filter_agents_threshold</code> would recalculate all the others).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025408,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/24/2020 14:34:05",
          "content": "<p><a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> , first, I assume that all masks stored on disk can be reproduced exactly by calling <code>selected_agents</code> with default parameters. I haven't checked that. If true, it is just a cache for performance reasons. As a side note, of course the test mask can not be reproduced, because we don't have the future for the test.</p>\n<p>Regarding the second part of your statement. When I said </p>\n<blockquote>\n  <p>Therefore by decreasing <code>min_frame_future</code> we introduce all these noisy (and irrelevant) agents which mess up with our loss.</p>\n</blockquote>\n<p>I meant specifically examples where future availability is full 50, but future masks are now smaller because we decrease <code>min_frame_future</code>. Such examples are very noise, and there are many of them. </p>\n<p>I feel that I have not necessarily cleared your concern, if you still think I made a mistake in any sentence please continue your arguments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025473,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "09/24/2020 15:13:30",
          "content": "<p>Might be misunderstanding you, still not quite sure (and of course I may be misunderstanding the code).<br>\nI agree the <code>agents_mask</code> in the zarr datasets is just a performance thing (the <code>mask.npz</code> used in test is different, not talking about that).</p>\n<p>I thought when you said:</p>\n<blockquote>\n  <p>The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy …</p>\n</blockquote>\n<p>you were suggesting that if you pass a custom <code>min_frame_future</code>/<code>min_frame_history</code> then only check 1 (availability) will be be performed and you'll end up with lots of frames that don't pass the other checks. Whereas if you stick to the defaults then you get frames with all checks verified. That's what I was questioning. My understanding is that regardless of whether you pass a custom <code>min_frame_future</code>/<code>min_frame_history</code> you will only get frames with all checks passed. But of course with a lower <code>min_frame_future</code>/<code>min_frame_history</code> that will be a weaker test and may introduce issues.</p>\n<p>Though actually, diving into the logic of <code>select_agents</code>  and <code>get_valid_agents</code> I'm struggling to follow the logic, and may have been misunderstanding. Though not sure that affects my concern here. Will have to investigate further but that was my concern in posting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025486,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/24/2020 15:24:25",
          "content": "<p>It is hard for me to pinpoint exactly where we diverge. So I will just share a few statements. First, <code>min_frame_future</code>/<code>min_frame_history</code> are not related to availability at all, they are thresholds on masks, which is availability plus additional checks, a subset but very different. Second, the problem of increased errors are not in the frame in question itself. The frame itself is of course passes both availability and the additional checks. The problem is what future follows it for the prediction purposes. And if what follows is both 50 masks and 50 availability, it is easy to predict. But if it is only 10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025505,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "09/24/2020 15:42:32",
          "content": "<p>Ah, OK. I think I see the divergence. You're talking about noise in errors. I was talking about noise in data. So you say:</p>\n<blockquote>\n  <p>10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.</p>\n</blockquote>\n<p>A perfectly reasonable meaning for noise but not the one I had. I would have instead said that was 10 points with data and 40 empty points so not really noisy (not to say my definition is better, just to identify divergence).</p>\n<p>I thought you were saying that if you kept the default <code>min_future_frames=1</code> then you'd end up with one non-empty, non-noisy frame and 49 empty. But if you changed it to <code>min_future_frames=10</code> then it would stop performing check 2-5 and you'd get 1 non-noisy frame, 9 noisy frames (because of no additional checks) and then 40 empty frames.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025535,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "09/24/2020 16:14:46",
          "content": "<p>I also think we might differ on the meaning of availability given the <code>target_availability</code> mask.<br>\nI think given a <code>target_availability</code> mask of size 50 where 10 values in it are 1 and then other 40 are 0 then you would say that 10 values are available (or maybe 40 if it was a negative mask but that's less likely). There aren't 50 available frames just because that's the size of the mask.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1035432,
      "author_name": "deepakrajpurushothaman",
      "author_url": "",
      "post_date": "10/02/2020 17:50:12",
      "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> did you find any explanation for deviation other than the one mentioned here and something that doesnt involve creating custom mask? I have been training with 224*224 and 0.25. For 3M sample the loss is like 100+ and for 6M sample the loss goes up to 6000 nll</p>",
      "votes": null,
      "replies": [
        {
          "id": 1035441,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/02/2020 17:58:09",
          "content": "<p>The sample.zarr can be easily overfitted. It's too small to actually train on. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1035453,
          "author_name": "deepakrajpurushothaman",
          "author_url": "",
          "post_date": "10/02/2020 18:06:42",
          "content": "<p>Well I am not using sample.Zarr. I have been on train_full</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1035525,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "10/02/2020 19:06:37",
          "content": "<p>If you're getting a loss of 6000 then you likely have an error somewhere. That's the sort of loss you'd get with random predictions.<br>\nYou aren't trying to use a model trained for l5kit 1.0.6 with the newly released 1.1,0 are you? Models trained with the old version aren't compatible with the new version without modifications (and really they need to be re-trained on the new version as that will improve performance significantly).<br>\nOtherwise you'd need to debug to figure out why the model is going off the rails. Someone did report exploding gradients which would cause that sort of thing but I haven't experienced that (but I'm using my own re-implementation of the loss function and don't know what implementation they were using).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1035887,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "10/03/2020 07:33:08",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a>, no I did not find a good enough explanation and the problem is still not solved. I noticed that not all people experienced the same so there might be something I an missing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1038185,
          "author_name": "arkaung",
          "author_url": "",
          "post_date": "10/05/2020 16:02:44",
          "content": "<p>If you are using the latest l5kit (<a href=\"https://github.com/lyft/l5kit/releases/tag/v1.1.0\" target=\"_blank\">v1.1.0</a>), make sure you convert agent coordinates into world offsets. Refer to Evaluation block in <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">this example</a> or refer to <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187825\" target=\"_blank\">this notebook</a>. Otherwise, your loss on test set will be super huge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1038597,
          "author_name": "deepakrajpurushothaman",
          "author_url": "",
          "post_date": "10/05/2020 22:56:21",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> , <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> and <a href=\"https://www.kaggle.com/arkaung\" target=\"_blank\">@arkaung</a> your inputs helps to think for sure . I will try to debug and try to find a reason. And about version, I am trying both l5kit 1.0.6 and 1.1.0 and taking care of the compatibility. the deviations exists both cases.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1039295,
          "author_name": "arkaung",
          "author_url": "",
          "post_date": "10/06/2020 13:28:11",
          "content": "<p>I have huge differences (between train loss and public LB) and trying to figure out what seems to be the problem too. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1043428,
      "author_name": "wantsu",
      "author_url": "",
      "post_date": "10/09/2020 01:25:47",
      "content": "<p>Similar experience for me! After <code>10, 000</code>loops, My training <code>NLL loss</code> is  about <code>100</code> which is very close to my <code>validate NLL loss</code> and the <code>LB</code>, but <code>the evaluate metic  score(neg_multi_NLL) is around 1, 000</code>, its a huge gap. </p>\n<p>My evaluate dataset is a subset (size is 1000*12) of the Chuck validate.zarr, obtained by torch.utill.data.Subset.</p>\n<p>I see other people got a very close result between evaulation and the LB. How you guys make it? Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075155,
          "author_name": "fuckvenkatraman",
          "author_url": "",
          "post_date": "11/11/2020 12:56:51",
          "content": "<p>I once used both the original train.zarr &amp; validate.zarr Chunkdataset, my training loss &amp; validate loss matched each other pretty. 1. Check your loss function, may you're using different ones. 2. You may need more training steps/training longer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1049045,
      "author_name": "xiaoyaopeng",
      "author_url": "",
      "post_date": "10/14/2020 03:58:40",
      "content": "<p>I met the same problem too. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1058086,
          "author_name": "xiaoyaopeng",
          "author_url": "",
          "post_date": "10/23/2020 10:19:14",
          "content": "<p>update:Oct 23rd:<br>\ntrain cfg:<br>\n    'model_params': {<br>\n        'model_architecture': 'resnet34',<br>\n        'history_num_frames': 10,<br>\n        'history_step_size': 1,<br>\n…<br>\n    'raster_params': {<br>\n        'raster_size': [480, 360],<br>\n        'pixel_size': [0.5, 0.5],<br>\n        'ego_center': [0.25, 0.5],<br>\n…<br>\n    'val_data_loader': {<br>\n        'key': 'scenes/validate.zarr',<br>\n        'batch_size': 32,<br>\n        'shuffle': False,<br>\n        'num_workers': 16<br>\n…<br>\ntrained for 600k iterations, train NLL loss is around 10, evaluation NLL loss around 30, LB score 26~28.<br>\nevaluation method:<br>\ncreate_chopped_dataset<br>\nnum_frames_to_chop = 100<br>\nMIN_FUTURE_STEPS = 10<br>\nMIN_FRAME_FUTURE = 10<br>\nMIN_FRAME_HISTORY = 10.<br>\nI guess there are some difference between training and validation. still working on the problem….</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058110,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/23/2020 10:58:05",
          "content": "<p>Could you post the full config? Are you using shuffle in your training loop? If not, this may explain the large difference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058160,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "10/23/2020 12:16:08",
          "content": "<p><a href=\"https://www.kaggle.com/xiaoyaopeng\" target=\"_blank\">@xiaoyaopeng</a> same here. Training loss ~11 after 600K iterations but LB in upper 20's, my raster is smaller though 224x224.<br>\nshuffle = True</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058166,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/23/2020 12:21:26",
          "content": "<p>600k on train or train_full? <br>\nfor train, you can already see overfitting/memorization at that point. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058172,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "10/23/2020 12:27:49",
          "content": "<p>unfortunately, train.csv in my case</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058357,
          "author_name": "xiaoyaopeng",
          "author_url": "",
          "post_date": "10/23/2020 15:51:51",
          "content": "<p>training config:    <br>\n'train_data_loader': {<br>\n        'key': 'scenes/train.zarr',<br>\n        'batch_size': 32,<br>\n        'shuffle': True,<br>\n        'num_workers': 16<br>\n    },</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058823,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "10/24/2020 10:29:52",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> a bit of a simple (and to some obvious) question here, but how would one correctly detect overfitting in this type of problem?  I would assume a continuous decrease in training loss with a sudden increase in evaluation loss, which can be found only by experimentation - early stopping the training a number of times.</p>\n<p>Since I constantly run experiments on minimal number of samples, I am at minimal risk of going into the overfitting territory, therefore not very familiar with it. Except during that one time I filtered way too many agents…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1059623,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/25/2020 09:39:35",
          "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> this is indeed the definition of overfitting. It doesn't need to be a sudden increase in evaluation loss as it can be a slow increase, too. The model fails to generalize to unseen data and starts to memorize training data almost perfectly. Thus, the training loss can reach very very low values while the evaluation score on unseen data begins to increase. You can tackle that problem with augmentation or lots of data (amongst other things). We have the insanly large dataset here, so we can make full use of it. If there is no way around the problem, early stopping is a valid technique.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1059810,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "10/25/2020 13:44:38",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> thanks for the clarification! I think augmentation could be useful for overfitting prevention in this problem if the original data is filtered from \"noisy\" inputs, eg filtering agents based on various criteria to avoid irrelevant training inputs. Since we will have a smaller number of inputs, memorization is more likely and augmentation might help in this case.</p>\n<p>Otherwise, if the input train dataset is relatively untouched, augmentation is probably not so important  due to the amount of data (as raised in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042\" target=\"_blank\">this discussion</a>) and early stopping may be sufficient.<br>\nBut as we know theory and practice can differ, so everything is up for experimentation;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065244,
          "author_name": "spicychicken38",
          "author_url": "",
          "post_date": "10/31/2020 04:25:00",
          "content": "<p>Maybe you need to adjust the pixel size correspondingly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075162,
          "author_name": "fuckvenkatraman",
          "author_url": "",
          "post_date": "11/11/2020 13:03:37",
          "content": "<p>For example, how to adjust the pixel size? You mean adjust it w.r.t raster size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084377,
          "author_name": "sujaydkhandekar",
          "author_url": "",
          "post_date": "11/20/2020 01:26:32",
          "content": "<p><a href=\"https://www.kaggle.com/xiaoyaopeng\" target=\"_blank\">@xiaoyaopeng</a> were you also using history positions as input to the model when you encountered this problem? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091664,
          "author_name": "xiaoyaopeng",
          "author_url": "",
          "post_date": "11/26/2020 07:24:28",
          "content": "<p>no, I just run the baseline notebook and made few changes.<br>\nI successfully made eval loss and training loss close by chop the training data.<br>\nI think <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> is right. there are overlaps in the training data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1053287,
      "author_name": "benbla",
      "author_url": "",
      "post_date": "10/18/2020 20:10:03",
      "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> did you make any progress regarding this topic? I still face the same problem… avg loss around ~25 and leaderboard around ~50</p>",
      "votes": null,
      "replies": [
        {
          "id": 1053636,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "10/19/2020 07:27:11",
          "content": "<p>Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation…</p>",
          "votes": null,
          "replies": [
            {
              "id": 1065243,
              "author_name": "spicychicken38",
              "author_url": "",
              "post_date": "10/31/2020 04:22:35",
              "content": "<blockquote>\n  <p>Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation…</p>\n</blockquote>\n<p>In addition to the similar performance on training loss and LB score/validation loss, can it help to improve the LB score finally?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1081110,
              "author_name": "benbla",
              "author_url": "",
              "post_date": "11/16/2020 20:19:24",
              "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a></p>\n<p>If you do it the right way, I think yes. However, I didn't manage to improve my score with increasing raster sizes..</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1074320,
      "author_name": "spicychicken38",
      "author_url": "",
      "post_date": "11/10/2020 14:12:06",
      "content": "<blockquote>\n  <p>Hello fellows,</p>\n  <p>The loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:</p>\n  <ul>\n  <li><p>Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.</p></li>\n  <li><p>Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.</p></li>\n  <li><p>Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.</p></li>\n  </ul>\n  <p>Did you guys have the same problem? Any thoughts?</p>\n  <p>Thanks!<br>\n  Frank</p>\n  <p>Update on 14 Oct:<br>\n  For me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.<br>\n  Now I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄</p>\n</blockquote>\n<p>Hi, I also meet this problem. But before this, I try baseline only using images without \"history_positions\" and the validation value is fine and reasonable. After adding history information, I also have a huge deviation between train loss and validation value. How you fix this problem finally?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075049,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "11/11/2020 11:04:33",
          "content": "<p>Hi, I am currently not using history_positions as input. I can think of two possible reasons why this does not work. First, there are many agents with very short history, as you can see from some of the earlier discussions here. So, the model cannot really rely on this data to predict. Second, to feed in this data, we need to add some non-linear layers on top of the standard Resnet, which, based on my limited experiments, does not do any good (if not harm).</p>\n<p>Please note that I did not do enough experiments to be sure that this won't work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075122,
          "author_name": "spicychicken38",
          "author_url": "",
          "post_date": "11/11/2020 12:07:22",
          "content": "<p>Yeah. Thank you for your reply. I do some experiments and it does confirm my thought that the different \"history_availabilities\" in train.zarr and chopped validation set can cause some data mismatch problems which could cause a huge deviation between train loss and actual LB score/validation score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075168,
          "author_name": "fuckvenkatraman",
          "author_url": "",
          "post_date": "11/11/2020 13:09:31",
          "content": "<p>So could you share some observations? </p>",
          "votes": null,
          "replies": [
            {
              "id": 1075943,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "11/12/2020 05:34:59",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/fuckvenkatraman\" target=\"_blank\">@fuckvenkatraman</a> , raster size selection is already discussed in this post <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\" target=\"_blank\">Raster size selection</a>. Basically, idea is to increase your width and have lesser height. (since ego vehicle moves primarily in x-direction). Also adjust the pixel_size parameter (which corresponds to how much distance a pixel means, default 1 pixel equals 0.5m). By adjusting both these parameters, you could focus more on your region of interest.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1076101,
          "author_name": "spicychicken38",
          "author_url": "",
          "post_date": "11/12/2020 09:08:25",
          "content": "<p><a href=\"https://www.kaggle.com/fuckvenkatraman\" target=\"_blank\">@fuckvenkatraman</a> Are you asking for my observation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1081109,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "11/16/2020 20:18:23",
          "content": "<p><a href=\"https://www.kaggle.com/spicychicken38\" target=\"_blank\">@spicychicken38</a> Yeah might be interesting in this case. Still struggling to figure out how to deal with the mismatches. Additional data didn't improve it for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091752,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "11/26/2020 08:55:13",
      "content": "<p>From what we have seen, a large deviation between train and validation / test score originates from different sample distributions. To match test, for AgentDataset we used <br>\n<code>min_frame_history=1</code><br>\n<code>min_frame_future=10</code><br>\nin validation and in most training setups</p>\n<p>Not to be confused with <code>history_num_frames</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1091832,
          "author_name": "spicychicken38",
          "author_url": "",
          "post_date": "11/26/2020 10:21:39",
          "content": "<p>Yeah, I firmly believe that such a huge deviation is caused by data mismatch, but I dealt with it with a more complicated method… I looked at the probability distribution of history_availailities in the test set and try to reconcile the probability distribution in the training set with that in the test set…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091959,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "11/26/2020 12:21:07",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> , I actually did that, I even set the min_frame_history=0, which surprisingly still could increase valid agents. So, for your final submission, what training loss and validation loss did you get?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1021627": "Hello fellows,\n\nThe loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:\n\n - Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.\n\n - Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.\n\n - Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.\n\nDid you guys have the same problem? Any thoughts?\n\nThanks!\nFrank\n\nUpdate on 14 Oct:\nFor me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.\nNow I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄",
    "1021743": "Same for me.\nI think the problem is information leaking. If you use the `train.zarr` or the `train_full.zar`, you have samples like these:\n\n- Current: scene 1 - frame 14; history: frame 4-13; target: frame 15-64\n- Current: scene 1 - frame 15; history: frame 5-14; target: frame: 16-65\n\nThere is an overlap in the samples.\n\nWhat was your validation score for your current LB (63.481)?",
    "1021789": "Run `create_chopped_dataset` vs. the validation set and use that for validation. What's the loss (nll) there?",
    "1021800": "Good point! I have never thought of this. Thanks Peter!\n\nSo this is really related to my third point? In the evaluation dataset, we only sample agents from the 100th frame, while in the training dataset we sample angents from all frames, thus there are many overlaps.\n\nI am not completely convinced this is the main reason, but if it is, removing these overlaps might help improving our training greatly?\n\nYes, my current LB score is 63.5, which was achieved by a training loss of 20. 😂",
    "1021804": "The same. Huge deviation.",
    "1021813": "Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. \nAre you shuffling the training set? If not, @pestipeti has a point with the overlap.",
    "1021827": "Thanks for the feedback, llu. This is helpful information.",
    "1021843": "Have you tried cutout?",
    "1021905": "Hi Peter, what do you mean by \"cutout\"? I tried dropout, that seemed to get a training loss close to validation loss, but the result was not good.",
    "1021923": "[Cutout](https://arxiv.org/pdf/1708.04552v2.pdf) is an augmentation technique. In short, you cut out random number/size squares (or rectangles) from the input image. You can make \"less similar\" images. In the image below the original images almost identical (frame 0, frame 1) but if you cut out random parts the result will be different. Of course, you have to find the right number/size of cuts, but I think it would help with the overlapping problem. (I haven't tried it)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fb3bc3975618543fd778719139878cb17%2Fcutout.png?generation=1600761651764733&alt=media)",
    "1021974": "Thanks for the good suggestions Peter. You have been so helpful since the beginning of this competition.",
    "1022114": "I observed the same problem and tried to deep dive yesterday. I found out that the root of my issue was incorrect `min_frame_future`, I used either 1 or 50, but it probably should be 10, as mentioned by others. \n\nIn more detail, `min_frame_future` and `min_frame_history` which are parameters for `AgentDataset` are responsible for cutting `agents_mask`. I was very surprised that decreasing `min_frame_future` actually increases the loss. Intuitively, it should be the opposite, because if we add samples with only short future available, we should have had increased accuracy for two reasons:\n1. Close-future points can be more accurately predicted. Further into the future - less accurate.\n2. Points with availability 0 are counted as error 0.\n\nAnd that does happen, but it is a smaller effect. The bigger effect stems from the fact that `availability` and `agents_mask` are very different things. I initially assumed that it was the same, but in fact `agents_mask` are all the points which are:\n1. available\n2. pass perception threshold on agents\n3. pass max absolute distance in degree\n4. pass max change in area allowed\n5. pass max distance from AV in meters\n\nThe sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing `min_frame_future` we introduce all these noisy (and irrelevant) agents which mess up with our loss. \n\nSummarizing, we probably want to have `min_frame_future` and `min_frame_history` at exactly same values as the test, which is probably 10 and 10. But also applying `create_chopped_dataset` should align them even more.",
    "1022121": "pestipeti , not sure you can call it leaking. It is just a correlation in the training set, it is allowed and perfectly fine. Conceptually, for the sake of arguments, one can look at it as data augmentation.",
    "1022181": "zaharch Yes, you are probably right. Leaking is not completely accurate.",
    "1022198": "Hi @zaharch great insights.\n\nBut according to the finding of @ilu000 here: \nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\nThe test dataset has lot of agents with less than 10 available history frames",
    "1022251": "Good point, @frankpanxj \n\nAs @ilu000 wrote\n```test_dataset[34]['history_availabilities'].sum() = 3.0```\nI checked that if a mask is generated with default parameters for the test then it is not even 3 but goes down to 1. Does it mean that the test mask is generated with `min_frame_history = 1`?",
    "1022607": "That was just an example. There are histories with len = 1 in the test set as well.\nThe problem is, that the `create_chopped_dataset` function only applies a filter w.r.t. history frames in a scene (100 for this test set). That doesn't necessarily mean, that the agent was always visible. Some are only visible in the very last frame.",
    "1025285": "> The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy (and irrelevant) agents which mess up with our loss. \n\nI don't think that's correct unless I'm misunderstanding the code. If you look at how `min_frame_future` is applied in [`AgentDataset.__init__`](https://github.com/lyft/l5kit/blob/master/l5kit/l5kit/dataset/agent.py#L36) it's used to filter the `agents_mask` array. This is loaded from the dataset in [`load_agents_mask`](https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/agent.py#L67) and only recreated if you use a different `filter_agents_threshold` to the one in the dataset (0.5). In this case a new agents mask is created with [`select_agents`](https://github.com/lyft/l5kit/blob/90a6109754c8a75199219188dff2e25d1b4489ab/l5kit/l5kit/dataset/select_agents.py#L153) which applies all those other checks.\nSo I don't think that varying the `min_frame_future` bypasses those checks. You always only get agents that pass all 5 checks, `min_frame_future` just re-calculates check 1 (and changing `filter_agents_threshold` would recalculate all the others).",
    "1025408": "thomasbrandon , first, I assume that all masks stored on disk can be reproduced exactly by calling `selected_agents` with default parameters. I haven't checked that. If true, it is just a cache for performance reasons. As a side note, of course the test mask can not be reproduced, because we don't have the future for the test.\n\nRegarding the second part of your statement. When I said \n> Therefore by decreasing `min_frame_future` we introduce all these noisy (and irrelevant) agents which mess up with our loss.\n\nI meant specifically examples where future availability is full 50, but future masks are now smaller because we decrease `min_frame_future`. Such examples are very noise, and there are many of them. \n\nI feel that I have not necessarily cleared your concern, if you still think I made a mistake in any sentence please continue your arguments.",
    "1025473": "Might be misunderstanding you, still not quite sure (and of course I may be misunderstanding the code).\nI agree the `agents_mask` in the zarr datasets is just a performance thing (the `mask.npz` used in test is different, not talking about that).\n\nI thought when you said:\n> The sample points which do not pass 2-5 checks but do pass 1 tend to be much noisier, harder to predict. Therefore by decreasing min_frame_future we introduce all these noisy ...\n\nyou were suggesting that if you pass a custom `min_frame_future`/`min_frame_history` then only check 1 (availability) will be be performed and you'll end up with lots of frames that don't pass the other checks. Whereas if you stick to the defaults then you get frames with all checks verified. That's what I was questioning. My understanding is that regardless of whether you pass a custom `min_frame_future`/`min_frame_history` you will only get frames with all checks passed. But of course with a lower `min_frame_future`/`min_frame_history` that will be a weaker test and may introduce issues.\n\nThough actually, diving into the logic of `select_agents`  and `get_valid_agents` I'm struggling to follow the logic, and may have been misunderstanding. Though not sure that affects my concern here. Will have to investigate further but that was my concern in posting.",
    "1025486": "It is hard for me to pinpoint exactly where we diverge. So I will just share a few statements. First, `min_frame_future`/`min_frame_history` are not related to availability at all, they are thresholds on masks, which is availability plus additional checks, a subset but very different. Second, the problem of increased errors are not in the frame in question itself. The frame itself is of course passes both availability and the additional checks. The problem is what future follows it for the prediction purposes. And if what follows is both 50 masks and 50 availability, it is easy to predict. But if it is only 10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.",
    "1025505": "Ah, OK. I think I see the divergence. You're talking about noise in errors. I was talking about noise in data. So you say:\n\n> 10 masks and 50 availability, the last 40 points are hard to predict. They are available but masked, - a lot of noise here.\n \nA perfectly reasonable meaning for noise but not the one I had. I would have instead said that was 10 points with data and 40 empty points so not really noisy (not to say my definition is better, just to identify divergence).\n\nI thought you were saying that if you kept the default `min_future_frames=1` then you'd end up with one non-empty, non-noisy frame and 49 empty. But if you changed it to `min_future_frames=10` then it would stop performing check 2-5 and you'd get 1 non-noisy frame, 9 noisy frames (because of no additional checks) and then 40 empty frames.",
    "1025535": "I also think we might differ on the meaning of availability given the `target_availability` mask.\nI think given a `target_availability` mask of size 50 where 10 values in it are 1 and then other 40 are 0 then you would say that 10 values are available (or maybe 40 if it was a negative mask but that's less likely). There aren't 50 available frames just because that's the size of the mask.",
    "1035432": "frankpanxj did you find any explanation for deviation other than the one mentioned here and something that doesnt involve creating custom mask? I have been training with 224*224 and 0.25. For 3M sample the loss is like 100+ and for 6M sample the loss goes up to 6000 nll",
    "1035441": "The sample.zarr can be easily overfitted. It's too small to actually train on.",
    "1035453": "Well I am not using sample.Zarr. I have been on train_full",
    "1035525": "If you're getting a loss of 6000 then you likely have an error somewhere. That's the sort of loss you'd get with random predictions.\nYou aren't trying to use a model trained for l5kit 1.0.6 with the newly released 1.1,0 are you? Models trained with the old version aren't compatible with the new version without modifications (and really they need to be re-trained on the new version as that will improve performance significantly).\nOtherwise you'd need to debug to figure out why the model is going off the rails. Someone did report exploding gradients which would cause that sort of thing but I haven't experienced that (but I'm using my own re-implementation of the loss function and don't know what implementation they were using).",
    "1035887": "Hi @deepakrajpurushothaman, no I did not find a good enough explanation and the problem is still not solved. I noticed that not all people experienced the same so there might be something I an missing.",
    "1037421": "ilu000 > Very strange behavior indeed. Setting min_frame_future=10 really got my training loss, validation loss, and LB score very close together. > \n\nHow were you able to get those 3 to agree?",
    "1038185": "If you are using the latest l5kit ([v1.1.0](https://github.com/lyft/l5kit/releases/tag/v1.1.0)), make sure you convert agent coordinates into world offsets. Refer to Evaluation block in [this example](https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb) or refer to [this notebook](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187825). Otherwise, your loss on test set will be super huge.",
    "1038597": "thanks @thomasbrandon , @frankpanxj and @arkaung your inputs helps to think for sure . I will try to debug and try to find a reason. And about version, I am trying both l5kit 1.0.6 and 1.1.0 and taking care of the compatibility. the deviations exists both cases.",
    "1039295": "I have huge differences (between train loss and public LB) and trying to figure out what seems to be the problem too.",
    "1043428": "Similar experience for me! After `10, 000 `loops, My training `NLL loss` is  about `100` which is very close to my `validate NLL loss` and the `LB`, but `the evaluate metic  score(neg_multi_NLL) is around 1, 000`, its a huge gap. \n\nMy evaluate dataset is a subset (size is 1000*12) of the Chuck validate.zarr, obtained by torch.utill.data.Subset.\n  \nI see other people got a very close result between evaulation and the LB. How you guys make it? Thanks",
    "1049045": "I met the same problem too.",
    "1053287": "frankpanxj did you make any progress regarding this topic? I still face the same problem... avg loss around ~25 and leaderboard around ~50",
    "1053636": "Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation...",
    "1058086": "update:Oct 23rd:\ntrain cfg:\n    'model_params': {\n        'model_architecture': 'resnet34',\n        'history_num_frames': 10,\n        'history_step_size': 1,\n...\n    'raster_params': {\n        'raster_size': [480, 360],\n        'pixel_size': [0.5, 0.5],\n        'ego_center': [0.25, 0.5],\n...\n    'val_data_loader': {\n        'key': 'scenes/validate.zarr',\n        'batch_size': 32,\n        'shuffle': False,\n        'num_workers': 16\n...\ntrained for 600k iterations, train NLL loss is around 10, evaluation NLL loss around 30, LB score 26~28.\nevaluation method:\ncreate_chopped_dataset\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\nMIN_FRAME_FUTURE = 10\nMIN_FRAME_HISTORY = 10.\nI guess there are some difference between training and validation. still working on the problem....",
    "1058110": "Could you post the full config? Are you using shuffle in your training loop? If not, this may explain the large difference.",
    "1058160": "xiaoyaopeng same here. Training loss ~11 after 600K iterations but LB in upper 20's, my raster is smaller though 224x224.\nshuffle = True",
    "1058166": "600k on train or train_full? \nfor train, you can already see overfitting/memorization at that point.",
    "1058172": "unfortunately, train.csv in my case",
    "1058357": "training config:    \n'train_data_loader': {\n        'key': 'scenes/train.zarr',\n        'batch_size': 32,\n        'shuffle': True,\n        'num_workers': 16\n    },",
    "1058823": "ilu000 a bit of a simple (and to some obvious) question here, but how would one correctly detect overfitting in this type of problem?  I would assume a continuous decrease in training loss with a sudden increase in evaluation loss, which can be found only by experimentation - early stopping the training a number of times.\n\nSince I constantly run experiments on minimal number of samples, I am at minimal risk of going into the overfitting territory, therefore not very familiar with it. Except during that one time I filtered way too many agents...",
    "1059623": "indswetrust this is indeed the definition of overfitting. It doesn't need to be a sudden increase in evaluation loss as it can be a slow increase, too. The model fails to generalize to unseen data and starts to memorize training data almost perfectly. Thus, the training loss can reach very very low values while the evaluation score on unseen data begins to increase. You can tackle that problem with augmentation or lots of data (amongst other things). We have the insanly large dataset here, so we can make full use of it. If there is no way around the problem, early stopping is a valid technique.",
    "1059810": "ilu000 thanks for the clarification! I think augmentation could be useful for overfitting prevention in this problem if the original data is filtered from \"noisy\" inputs, eg filtering agents based on various criteria to avoid irrelevant training inputs. Since we will have a smaller number of inputs, memorization is more likely and augmentation might help in this case.\n\nOtherwise, if the input train dataset is relatively untouched, augmentation is probably not so important  due to the amount of data (as raised in [this discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042)) and early stopping may be sufficient.\nBut as we know theory and practice can differ, so everything is up for experimentation;)",
    "1065243": "> Answering to myself this time ;). I noticed that if I increase the raster sizes to larger numbers (e.g. 400 x 400 instead of 240 x 240), the avg loss is more similar to the lb score. I dont know why yet but my newest tests show a stable correlation...\n\nIn addition to the similar performance on training loss and LB score/validation loss, can it help to improve the LB score finally?",
    "1065244": "Maybe you need to adjust the pixel size correspondingly.",
    "1074009": "Are you guys talking about the config parameter `future_num_frames`?",
    "1074023": "Nope, min_frame_future, min_frame_history are AgentDataset class parameters. Look at the AgentDataset() class in the l5kit and you'll see whats going on.",
    "1074320": "> Hello fellows,\n> \n> The loss value I got during training was always much lower than the evaluation loss. I got a loss of below 30 during training very soon (just too good to be true), but the evaluation loss was always above 60. I investigated the following possibilities:\n> \n>  - Overfitting. This is unlikely. Due to the limited computing power of my PC, I was never able to run a complete cycle of the train.zarr. My model has never seen the same sample twice. Moreover, this “overfitting” begins after just 10000 iterations.\n> \n>  - Different future frame numbers. I first thought this must be the reason. The longer the future frames available, the larger the loss. I checked the gt.csv generated by the create_chopped_dataset function. I can confirm that the minimum future frames in the evaluation dataset is indeed 10. However, after setting the min_frame_future=10 for train.zarr, which reduces the overall number of samples by about 25%, the problem is still there, not even improved.\n> \n>  - Different sampling of agents. The agents in the evaluation data is the “valid agents in the 100th frame” after cutting every scene to 100 frames. The agents loaded from train.zarr include all valid agents in all frames (I guess). I cannot see how this could make such a big difference.\n> \n> Did you guys have the same problem? Any thoughts?\n> \n> Thanks!\n> Frank\n> \n> Update on 14 Oct:\n> For me the problem seems to be partly related to the data I feed in the model. So, besides inputting the image, I also feed in the history positions, hoping to provide some more accurate information in case the large pixel size makes the position not so precise. After removing this data, there is still significant difference between training loss and validation loss but much better. So this is definitely part of the reason for me, although I do not understand why this happens.\n> Now I am basically back to square one. All the \"improvements\" I made to the baseline solution and public notebooks have failed.😄\n\nHi, I also meet this problem. But before this, I try baseline only using images without \"history_positions\" and the validation value is fine and reasonable. After adding history information, I also have a huge deviation between train loss and validation value. How you fix this problem finally?",
    "1075049": "Hi, I am currently not using history_positions as input. I can think of two possible reasons why this does not work. First, there are many agents with very short history, as you can see from some of the earlier discussions here. So, the model cannot really rely on this data to predict. Second, to feed in this data, we need to add some non-linear layers on top of the standard Resnet, which, based on my limited experiments, does not do any good (if not harm).\n\nPlease note that I did not do enough experiments to be sure that this won't work.",
    "1075122": "Yeah. Thank you for your reply. I do some experiments and it does confirm my thought that the different \"history_availabilities\" in train.zarr and chopped validation set can cause some data mismatch problems which could cause a huge deviation between train loss and actual LB score/validation score.",
    "1075150": "indswetrust hi there, after reading the post I could grasp a better understanding of the terms. Thanks for your help though!",
    "1075155": "I once used both the original train.zarr & validate.zarr Chunkdataset, my training loss & validate loss matched each other pretty. 1. Check your loss function, may you're using different ones. 2. You may need more training steps/training longer.",
    "1075162": "For example, how to adjust the pixel size? You mean adjust it w.r.t raster size?",
    "1075168": "So could you share some observations?",
    "1075943": "Hi @fuckvenkatraman , raster size selection is already discussed in this post [Raster size selection](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323). Basically, idea is to increase your width and have lesser height. (since ego vehicle moves primarily in x-direction). Also adjust the pixel_size parameter (which corresponds to how much distance a pixel means, default 1 pixel equals 0.5m). By adjusting both these parameters, you could focus more on your region of interest.",
    "1076101": "fuckvenkatraman Are you asking for my observation?",
    "1077089": "yousof9 Hi, How to set min_frame_future=10? Do I need to call create_chopped_dataset function to train dataset?",
    "1077097": "No, it can be passed as an argument to Agent dataset class. By default the value is 1. You can refer to l5kit github page for more information.",
    "1081109": "spicychicken38 Yeah might be interesting in this case. Still struggling to figure out how to deal with the mismatches. Additional data didn't improve it for me.",
    "1081110": "spicychicken38\n\nIf you do it the right way, I think yes. However, I didn't manage to improve my score with increasing raster sizes..",
    "1084377": "xiaoyaopeng were you also using history positions as input to the model when you encountered this problem?",
    "1091664": "no, I just run the baseline notebook and made few changes.\nI successfully made eval loss and training loss close by chop the training data.\nI think @pestipeti is right. there are overlaps in the training data.",
    "1091752": "From what we have seen, a large deviation between train and validation / test score originates from different sample distributions. To match test, for AgentDataset we used \n`min_frame_history=1`\n`min_frame_future=10`\nin validation and in most training setups\n\nNot to be confused with `history_num_frames`",
    "1091832": "Yeah, I firmly believe that such a huge deviation is caused by data mismatch, but I dealt with it with a more complicated method... I looked at the probability distribution of history_availailities in the test set and try to reconcile the probability distribution in the training set with that in the test set...",
    "1091959": "Thanks @ilu000 , I actually did that, I even set the min_frame_history=0, which surprisingly still could increase valid agents. So, for your final submission, what training loss and validation loss did you get?"
  },
  "source": "meta"
}