{
  "id": 199433,
  "title": "Trouble in Paradise",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199433",
  "author_name": "",
  "post_date": "2020-11-25T17:59:39.628848200Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>It's too late I know, but I wanted to try this last second hustle.</p>\n<p>My model's training has <em>seemingly</em> been going okay form yesterday. Haven't been validating in train loop because on forums people said models need ~24h to converge for real, so I figured just let the thing do it's thing:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Faef19867fd6892eccc9dcbb647f47399%2Fwhat.png?generation=1606326887130863&amp;alt=media\" alt=\"\"></p>\n<p>The pictured model is a resnet34 trained for 140000 steps of 64-sized batches from random samples of train_full.zarr.</p>\n<p>So the problem is, when I look at the most recent, say, 15k predictions <code>z.loss.iloc[-15000:].mean()</code>, I get a train loss of 17.9. That feels comfortable. When I <a href=\"https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min\" target=\"_blank\">try and validate</a> against a sample from the validation set, I get back values that look like this:</p>\n<pre><code>{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n</code></pre>\n<p>Yikes! I clearly jacked something up. I saw a few posts on people having similar issues but no definitive solution. Did I miss it / is it buried in the discussions somewhere? Do we have to scale our solutions by some multiplier or something?</p>",
  "messages": [
    {
      "id": "1090995",
      "postDate": "11/25/2020 17:59:39",
      "content": "<p>It's too late I know, but I wanted to try this last second hustle.</p>\n<p>My model's training has <em>seemingly</em> been going okay form yesterday. Haven't been validating in train loop because on forums people said models need ~24h to converge for real, so I figured just let the thing do it's thing:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Faef19867fd6892eccc9dcbb647f47399%2Fwhat.png?generation=1606326887130863&amp;alt=media\" alt=\"\"></p>\n<p>The pictured model is a resnet34 trained for 140000 steps of 64-sized batches from random samples of train_full.zarr.</p>\n<p>So the problem is, when I look at the most recent, say, 15k predictions <code>z.loss.iloc[-15000:].mean()</code>, I get a train loss of 17.9. That feels comfortable. When I <a href=\"https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min\" target=\"_blank\">try and validate</a> against a sample from the validation set, I get back values that look like this:</p>\n<pre><code>{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n</code></pre>\n<p>Yikes! I clearly jacked something up. I saw a few posts on people having similar issues but no definitive solution. Did I miss it / is it buried in the discussions somewhere? Do we have to scale our solutions by some multiplier or something?</p>",
      "rawMarkdown": "It's too late I know, but I wanted to try this last second hustle.\n\nMy model's training has *seemingly* been going okay form yesterday. Haven't been validating in train loop because on forums people said models need ~24h to converge for real, so I figured just let the thing do it's thing:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Faef19867fd6892eccc9dcbb647f47399%2Fwhat.png?generation=1606326887130863&alt=media)\n\nThe pictured model is a resnet34 trained for 140000 steps of 64-sized batches from random samples of train_full.zarr.\n\nSo the problem is, when I look at the most recent, say, 15k predictions `z.loss.iloc[-15000:].mean()`, I get a train loss of 17.9. That feels comfortable. When I [try and validate](https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min) against a sample from the validation set, I get back values that look like this:\n\n```\n{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n```\n\nYikes! I clearly jacked something up. I saw a few posts on people having similar issues but no definitive solution. Did I miss it / is it buried in the discussions somewhere? Do we have to scale our solutions by some multiplier or something?",
      "votes": null
    },
    {
      "id": "1091005",
      "postDate": "11/25/2020 18:05:01",
      "content": "<p>For reference, it's not like val loss is moving down but… it feels like everything just shifted or something?</p>\n<pre><code>model_resnet34_output_84000.pth\n{'neg_multi_log_likelihood': 82.68555224218878, 'time_displace': array([0.08216104, 0.14191299, 0.17862896, 0.23167751, 0.29029346,\n       0.34343956, 0.39779761, 0.44261665, 0.48665519, 0.53399251,\n       0.57908257, 0.61783541, 0.65807086, 0.69021504, 0.71906   ,\n       0.75315359, 0.78096231, 0.81671914, 0.84816372, 0.8765357 ,\n       0.9001293 , 0.92258684, 0.94231826, 0.95649926, 0.97533073,\n       0.99249186, 1.01404677, 1.0282244 , 1.03808593, 1.04595177,\n       1.05631953, 1.06164656, 1.07050866, 1.07312559, 1.07921475,\n       1.09664675, 1.1066373 , 1.11167966, 1.11038037, 1.12044869,\n       1.12766502, 1.14252446, 1.15611682, 1.16043487, 1.17211799,\n       1.18489064, 1.20385188, 1.22025188, 1.24206752, 1.26920105])}\n\nmodel_resnet34_output_102000.pth\n{'neg_multi_log_likelihood': 60.15995701082701, 'time_displace': array([0.0814472 , 0.13519718, 0.18767383, 0.23921354, 0.29237729,\n       0.33924203, 0.38126515, 0.41581896, 0.45892417, 0.49914023,\n       0.53566297, 0.57807686, 0.6126809 , 0.64138179, 0.67282553,\n       0.69441783, 0.71572199, 0.73372158, 0.7567145 , 0.77424734,\n       0.78061963, 0.80467917, 0.82018216, 0.83072158, 0.84229814,\n       0.85418034, 0.8649891 , 0.87012398, 0.87173393, 0.87519373,\n       0.86675252, 0.87186382, 0.88307607, 0.88930812, 0.90104548,\n       0.91990247, 0.93296282, 0.93163789, 0.93750535, 0.9523782 ,\n       0.96132791, 0.96947078, 0.98608108, 0.99482964, 1.02291647,\n       1.04211571, 1.070679  , 1.08794308, 1.10915816, 1.13385227])}\n\nmodel_resnet34_output_140000.pth\n{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n\nmodel_resnet34_output_194000.pth\n{'neg_multi_log_likelihood': 36.540700766245884, 'time_displace': array([0.06768786, 0.11239789, 0.15527047, 0.19707848, 0.23415925,\n       0.27042494, 0.30206143, 0.32945963, 0.3599874 , 0.38719589,\n       0.41793629, 0.44404189, 0.46830858, 0.49044816, 0.50436334,\n       0.52224877, 0.53952061, 0.55973396, 0.57723214, 0.59279727,\n       0.60702455, 0.6243174 , 0.63574038, 0.65546905, 0.67029015,\n       0.68104914, 0.69563094, 0.70341774, 0.71044856, 0.70976926,\n       0.71465194, 0.71927185, 0.72643562, 0.73100875, 0.7356024 ,\n       0.74746589, 0.76176941, 0.7696118 , 0.77722741, 0.79098686,\n       0.80050605, 0.8083918 , 0.82216129, 0.83489121, 0.85430071,\n       0.87171541, 0.88682221, 0.90260645, 0.92327949, 0.94319656])}\n</code></pre>\n<p>You can approximate where train loss was from the first picture at the corresponding time step.</p>",
      "rawMarkdown": "For reference, it's not like val loss is moving down but... it feels like everything just shifted or something?\n\n```\nmodel_resnet34_output_84000.pth\n{'neg_multi_log_likelihood': 82.68555224218878, 'time_displace': array([0.08216104, 0.14191299, 0.17862896, 0.23167751, 0.29029346,\n       0.34343956, 0.39779761, 0.44261665, 0.48665519, 0.53399251,\n       0.57908257, 0.61783541, 0.65807086, 0.69021504, 0.71906   ,\n       0.75315359, 0.78096231, 0.81671914, 0.84816372, 0.8765357 ,\n       0.9001293 , 0.92258684, 0.94231826, 0.95649926, 0.97533073,\n       0.99249186, 1.01404677, 1.0282244 , 1.03808593, 1.04595177,\n       1.05631953, 1.06164656, 1.07050866, 1.07312559, 1.07921475,\n       1.09664675, 1.1066373 , 1.11167966, 1.11038037, 1.12044869,\n       1.12766502, 1.14252446, 1.15611682, 1.16043487, 1.17211799,\n       1.18489064, 1.20385188, 1.22025188, 1.24206752, 1.26920105])}\n\nmodel_resnet34_output_102000.pth\n{'neg_multi_log_likelihood': 60.15995701082701, 'time_displace': array([0.0814472 , 0.13519718, 0.18767383, 0.23921354, 0.29237729,\n       0.33924203, 0.38126515, 0.41581896, 0.45892417, 0.49914023,\n       0.53566297, 0.57807686, 0.6126809 , 0.64138179, 0.67282553,\n       0.69441783, 0.71572199, 0.73372158, 0.7567145 , 0.77424734,\n       0.78061963, 0.80467917, 0.82018216, 0.83072158, 0.84229814,\n       0.85418034, 0.8649891 , 0.87012398, 0.87173393, 0.87519373,\n       0.86675252, 0.87186382, 0.88307607, 0.88930812, 0.90104548,\n       0.91990247, 0.93296282, 0.93163789, 0.93750535, 0.9523782 ,\n       0.96132791, 0.96947078, 0.98608108, 0.99482964, 1.02291647,\n       1.04211571, 1.070679  , 1.08794308, 1.10915816, 1.13385227])}\n\nmodel_resnet34_output_140000.pth\n{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n\nmodel_resnet34_output_194000.pth\n{'neg_multi_log_likelihood': 36.540700766245884, 'time_displace': array([0.06768786, 0.11239789, 0.15527047, 0.19707848, 0.23415925,\n       0.27042494, 0.30206143, 0.32945963, 0.3599874 , 0.38719589,\n       0.41793629, 0.44404189, 0.46830858, 0.49044816, 0.50436334,\n       0.52224877, 0.53952061, 0.55973396, 0.57723214, 0.59279727,\n       0.60702455, 0.6243174 , 0.63574038, 0.65546905, 0.67029015,\n       0.68104914, 0.69563094, 0.70341774, 0.71044856, 0.70976926,\n       0.71465194, 0.71927185, 0.72643562, 0.73100875, 0.7356024 ,\n       0.74746589, 0.76176941, 0.7696118 , 0.77722741, 0.79098686,\n       0.80050605, 0.8083918 , 0.82216129, 0.83489121, 0.85430071,\n       0.87171541, 0.88682221, 0.90260645, 0.92327949, 0.94319656])}\n```\n\nYou can approximate where train loss was from the first picture at the corresponding time step.",
      "votes": null
    },
    {
      "id": "1091009",
      "postDate": "11/25/2020 18:11:34",
      "content": "<p>Yeah validation loss can be twice as high as train loss. I once overfit the train set and got train loss = 17 but chopped validation set loss = 74, which is when I tried to do replay memory trick in one of the discussion.<br>\n(Note that you need to validate on the \"chopped\" validation set to get real score)</p>",
      "rawMarkdown": "Yeah validation loss can be twice as high as train loss. I once overfit the train set and got train loss = 17 but chopped validation set loss = 74, which is when I tried to do replay memory trick in one of the discussion.\n(Note that you need to validate on the \"chopped\" validation set to get real score)",
      "votes": null
    },
    {
      "id": "1091020",
      "postDate": "11/25/2020 18:20:59",
      "content": "<p>Thank for for your comments. Currently, I am validating w/ chopped but training with train_full.</p>\n<blockquote>\n  <p>Yeah validation loss can be twice as high as train loss</p>\n</blockquote>\n<p>I see. In my case it's looking closer to ~2.8x, but… if that's expected and if everyone is encountering this, I guess it's fine. In some discussions I've read though, people are able to get train + valid scores neck-n-neck. I guess they must have been training on chopped trained.</p>\n<p>Last night before I started running my model, I followed <a href=\"https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset\" target=\"_blank\">the steps here</a> in order to created a chopped train_full.zarr set. After about 2 and a half hours of processing, it was finally ready to train on. But when I attempted to train, train loss started and stayed at 0.0(?). It was already like 2am at that time, so I figured forget it.. I'll just train with train_full since some kernels scored 23.x just using the regular train set.</p>\n<blockquote>\n  <p>which is when I tried to do replay memory trick</p>\n</blockquote>\n<p>Eh? What's that</p>",
      "rawMarkdown": "Thank for for your comments. Currently, I am validating w/ chopped but training with train_full.\n\n> Yeah validation loss can be twice as high as train loss\n\nI see. In my case it's looking closer to ~2.8x, but... if that's expected and if everyone is encountering this, I guess it's fine. In some discussions I've read though, people are able to get train + valid scores neck-n-neck. I guess they must have been training on chopped trained.\n\nLast night before I started running my model, I followed [the steps here](https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset) in order to created a chopped train_full.zarr set. After about 2 and a half hours of processing, it was finally ready to train on. But when I attempted to train, train loss started and stayed at 0.0(?). It was already like 2am at that time, so I figured forget it.. I'll just train with train_full since some kernels scored 23.x just using the regular train set.\n\n> which is when I tried to do replay memory trick\n\nEh? What's that",
      "votes": null
    },
    {
      "id": "1091055",
      "postDate": "11/25/2020 18:43:01",
      "content": "<p>The chop dataset function provided in l5kit removes the future positions for all the agents so you can't use that to train the model directly.<br>\nBut I guess if you do it properly, this sounds a really good idea. I wish I have heard this earlier…<br>\nThe replay memory trick is to store the loaded training batches in memory and randomly retrain on those cached batches multiple time during training to speed up since most of the computation time of this competition is spent on the loading the data but not training. I saw this suggest by someone in the discussion. Though it lead to overfit when I tried it.</p>",
      "rawMarkdown": "The chop dataset function provided in l5kit removes the future positions for all the agents so you can't use that to train the model directly.\nBut I guess if you do it properly, this sounds a really good idea. I wish I have heard this earlier...\nThe replay memory trick is to store the loaded training batches in memory and randomly retrain on those cached batches multiple time during training to speed up since most of the computation time of this competition is spent on the loading the data but not training. I saw this suggest by someone in the discussion. Though it lead to overfit when I tried it.",
      "votes": null
    },
    {
      "id": "1091158",
      "postDate": "11/25/2020 19:59:47",
      "content": "<p>What are your args when you call AgentDataset on the training set?</p>",
      "rawMarkdown": "What are your args when you call AgentDataset on the training set?",
      "votes": null
    },
    {
      "id": "1091183",
      "postDate": "11/25/2020 20:42:01",
      "content": "<pre><code>cfg = {\n    'format_version': 4,\n    'data_path': \"../input/lyft-motion-prediction-autonomous-vehicles\",\n    'model_params': {\n        'model_architecture': 'resnet34',\n        'history_num_frames': 6,\n        'history_step_size': 1,\n        'history_delta_time': 0.1,\n        'future_num_frames': 50,\n        'future_step_size': 1,\n        'future_delta_time': 0.1,\n        'model_name': \"model_resnet34_output\",\n        'lr': 1e-3, #1e-3,        \n        'train': True,\n        'predict': True\n    },\n\n    'raster_params': {\n        'raster_size': [224*2, 224//2],\n        'pixel_size': [0.3, 0.3],\n        'ego_center': [0.2, 0.5], # [0.25, 0.5]\n        'map_type': 'py_semantic',\n        'satellite_map_key': 'aerial_map/aerial_map.png',\n        'semantic_map_key': 'semantic_map/semantic_map.pb',\n        'dataset_meta_key': 'meta.json',\n        'filter_agents_threshold': 0.5 \n    },\n\n    'train_data_loader': {\n        'key': 'scenes/train_full.zarr',\n        'batch_size': 64,\n        'shuffle': True,\n        'num_workers': 4 if KAGGLE else 11\n    },\n}\n</code></pre>",
      "rawMarkdown": "```\ncfg = {\n    'format_version': 4,\n    'data_path': \"../input/lyft-motion-prediction-autonomous-vehicles\",\n    'model_params': {\n        'model_architecture': 'resnet34',\n        'history_num_frames': 6,\n        'history_step_size': 1,\n        'history_delta_time': 0.1,\n        'future_num_frames': 50,\n        'future_step_size': 1,\n        'future_delta_time': 0.1,\n        'model_name': \"model_resnet34_output\",\n        'lr': 1e-3, #1e-3,        \n        'train': True,\n        'predict': True\n    },\n\n    'raster_params': {\n        'raster_size': [224*2, 224//2],\n        'pixel_size': [0.3, 0.3],\n        'ego_center': [0.2, 0.5], # [0.25, 0.5]\n        'map_type': 'py_semantic',\n        'satellite_map_key': 'aerial_map/aerial_map.png',\n        'semantic_map_key': 'semantic_map/semantic_map.pb',\n        'dataset_meta_key': 'meta.json',\n        'filter_agents_threshold': 0.5 \n    },\n\n    'train_data_loader': {\n        'key': 'scenes/train_full.zarr',\n        'batch_size': 64,\n        'shuffle': True,\n        'num_workers': 4 if KAGGLE else 11\n    },\n}\n```",
      "votes": null
    },
    {
      "id": "1091225",
      "postDate": "11/25/2020 21:54:27",
      "content": "<p>I actually mean the <code>AgentDataset(cfg, train_zarr ...)</code><br>\nThe cfg looks fine, though. <br>\nWhat did you chose for <code>min_frame_history</code> and <code>min_frame_future</code>?</p>",
      "rawMarkdown": "I actually mean the `AgentDataset(cfg, train_zarr ...)`\nThe cfg looks fine, though. \nWhat did you chose for `min_frame_history` and `min_frame_future`?",
      "votes": null
    },
    {
      "id": "1091227",
      "postDate": "11/25/2020 21:57:31",
      "content": "<p>Hmm… Shouldn't the <code>history_num_frames</code> be 10? Or you won't have the same history as the test set problem.</p>",
      "rawMarkdown": "Hmm... Shouldn't the `history_num_frames` be 10? Or you won't have the same history as the test set problem.",
      "votes": null
    },
    {
      "id": "1091233",
      "postDate": "11/25/2020 22:05:19",
      "content": "<p><code>history_num_frames</code> is just a hyperparameter of how many history frames you are using in your rasterizer. You can go up to 100 as this is the max len in test. </p>",
      "rawMarkdown": "`history_num_frames` is just a hyperparameter of how many history frames you are using in your rasterizer. You can go up to 100 as this is the max len in test.",
      "votes": null
    },
    {
      "id": "1091260",
      "postDate": "11/25/2020 22:32:53",
      "content": "<p>It's too late for me now 🙃</p>",
      "rawMarkdown": "It's too late for me now 🙃",
      "votes": null
    },
    {
      "id": "1091388",
      "postDate": "11/26/2020 01:16:07",
      "content": "<p>Oh shoo… I guess I completely misunderstood that..</p>",
      "rawMarkdown": "Oh shoo... I guess I completely misunderstood that..",
      "votes": null
    },
    {
      "id": "1091674",
      "postDate": "11/26/2020 07:32:20",
      "content": "<p>with min_frame_history = 1 and min_frame_future = 10 it resembles the test dataset best (and matches training loss with validation loss and test score)</p>",
      "rawMarkdown": "with min_frame_history = 1 and min_frame_future = 10 it resembles the test dataset best (and matches training loss with validation loss and test score)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091005,
      "author_name": "authman",
      "author_url": "",
      "post_date": "11/25/2020 18:05:01",
      "content": "<p>For reference, it's not like val loss is moving down but… it feels like everything just shifted or something?</p>\n<pre><code>model_resnet34_output_84000.pth\n{'neg_multi_log_likelihood': 82.68555224218878, 'time_displace': array([0.08216104, 0.14191299, 0.17862896, 0.23167751, 0.29029346,\n       0.34343956, 0.39779761, 0.44261665, 0.48665519, 0.53399251,\n       0.57908257, 0.61783541, 0.65807086, 0.69021504, 0.71906   ,\n       0.75315359, 0.78096231, 0.81671914, 0.84816372, 0.8765357 ,\n       0.9001293 , 0.92258684, 0.94231826, 0.95649926, 0.97533073,\n       0.99249186, 1.01404677, 1.0282244 , 1.03808593, 1.04595177,\n       1.05631953, 1.06164656, 1.07050866, 1.07312559, 1.07921475,\n       1.09664675, 1.1066373 , 1.11167966, 1.11038037, 1.12044869,\n       1.12766502, 1.14252446, 1.15611682, 1.16043487, 1.17211799,\n       1.18489064, 1.20385188, 1.22025188, 1.24206752, 1.26920105])}\n\nmodel_resnet34_output_102000.pth\n{'neg_multi_log_likelihood': 60.15995701082701, 'time_displace': array([0.0814472 , 0.13519718, 0.18767383, 0.23921354, 0.29237729,\n       0.33924203, 0.38126515, 0.41581896, 0.45892417, 0.49914023,\n       0.53566297, 0.57807686, 0.6126809 , 0.64138179, 0.67282553,\n       0.69441783, 0.71572199, 0.73372158, 0.7567145 , 0.77424734,\n       0.78061963, 0.80467917, 0.82018216, 0.83072158, 0.84229814,\n       0.85418034, 0.8649891 , 0.87012398, 0.87173393, 0.87519373,\n       0.86675252, 0.87186382, 0.88307607, 0.88930812, 0.90104548,\n       0.91990247, 0.93296282, 0.93163789, 0.93750535, 0.9523782 ,\n       0.96132791, 0.96947078, 0.98608108, 0.99482964, 1.02291647,\n       1.04211571, 1.070679  , 1.08794308, 1.10915816, 1.13385227])}\n\nmodel_resnet34_output_140000.pth\n{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n\nmodel_resnet34_output_194000.pth\n{'neg_multi_log_likelihood': 36.540700766245884, 'time_displace': array([0.06768786, 0.11239789, 0.15527047, 0.19707848, 0.23415925,\n       0.27042494, 0.30206143, 0.32945963, 0.3599874 , 0.38719589,\n       0.41793629, 0.44404189, 0.46830858, 0.49044816, 0.50436334,\n       0.52224877, 0.53952061, 0.55973396, 0.57723214, 0.59279727,\n       0.60702455, 0.6243174 , 0.63574038, 0.65546905, 0.67029015,\n       0.68104914, 0.69563094, 0.70341774, 0.71044856, 0.70976926,\n       0.71465194, 0.71927185, 0.72643562, 0.73100875, 0.7356024 ,\n       0.74746589, 0.76176941, 0.7696118 , 0.77722741, 0.79098686,\n       0.80050605, 0.8083918 , 0.82216129, 0.83489121, 0.85430071,\n       0.87171541, 0.88682221, 0.90260645, 0.92327949, 0.94319656])}\n</code></pre>\n<p>You can approximate where train loss was from the first picture at the corresponding time step.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091009,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/25/2020 18:11:34",
      "content": "<p>Yeah validation loss can be twice as high as train loss. I once overfit the train set and got train loss = 17 but chopped validation set loss = 74, which is when I tried to do replay memory trick in one of the discussion.<br>\n(Note that you need to validate on the \"chopped\" validation set to get real score)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091020,
          "author_name": "authman",
          "author_url": "",
          "post_date": "11/25/2020 18:20:59",
          "content": "<p>Thank for for your comments. Currently, I am validating w/ chopped but training with train_full.</p>\n<blockquote>\n  <p>Yeah validation loss can be twice as high as train loss</p>\n</blockquote>\n<p>I see. In my case it's looking closer to ~2.8x, but… if that's expected and if everyone is encountering this, I guess it's fine. In some discussions I've read though, people are able to get train + valid scores neck-n-neck. I guess they must have been training on chopped trained.</p>\n<p>Last night before I started running my model, I followed <a href=\"https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset\" target=\"_blank\">the steps here</a> in order to created a chopped train_full.zarr set. After about 2 and a half hours of processing, it was finally ready to train on. But when I attempted to train, train loss started and stayed at 0.0(?). It was already like 2am at that time, so I figured forget it.. I'll just train with train_full since some kernels scored 23.x just using the regular train set.</p>\n<blockquote>\n  <p>which is when I tried to do replay memory trick</p>\n</blockquote>\n<p>Eh? What's that</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091055,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/25/2020 18:43:01",
          "content": "<p>The chop dataset function provided in l5kit removes the future positions for all the agents so you can't use that to train the model directly.<br>\nBut I guess if you do it properly, this sounds a really good idea. I wish I have heard this earlier…<br>\nThe replay memory trick is to store the loaded training batches in memory and randomly retrain on those cached batches multiple time during training to speed up since most of the computation time of this competition is spent on the loading the data but not training. I saw this suggest by someone in the discussion. Though it lead to overfit when I tried it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091158,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "11/25/2020 19:59:47",
      "content": "<p>What are your args when you call AgentDataset on the training set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091183,
          "author_name": "authman",
          "author_url": "",
          "post_date": "11/25/2020 20:42:01",
          "content": "<pre><code>cfg = {\n    'format_version': 4,\n    'data_path': \"../input/lyft-motion-prediction-autonomous-vehicles\",\n    'model_params': {\n        'model_architecture': 'resnet34',\n        'history_num_frames': 6,\n        'history_step_size': 1,\n        'history_delta_time': 0.1,\n        'future_num_frames': 50,\n        'future_step_size': 1,\n        'future_delta_time': 0.1,\n        'model_name': \"model_resnet34_output\",\n        'lr': 1e-3, #1e-3,        \n        'train': True,\n        'predict': True\n    },\n\n    'raster_params': {\n        'raster_size': [224*2, 224//2],\n        'pixel_size': [0.3, 0.3],\n        'ego_center': [0.2, 0.5], # [0.25, 0.5]\n        'map_type': 'py_semantic',\n        'satellite_map_key': 'aerial_map/aerial_map.png',\n        'semantic_map_key': 'semantic_map/semantic_map.pb',\n        'dataset_meta_key': 'meta.json',\n        'filter_agents_threshold': 0.5 \n    },\n\n    'train_data_loader': {\n        'key': 'scenes/train_full.zarr',\n        'batch_size': 64,\n        'shuffle': True,\n        'num_workers': 4 if KAGGLE else 11\n    },\n}\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091225,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "11/25/2020 21:54:27",
          "content": "<p>I actually mean the <code>AgentDataset(cfg, train_zarr ...)</code><br>\nThe cfg looks fine, though. <br>\nWhat did you chose for <code>min_frame_history</code> and <code>min_frame_future</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091227,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/25/2020 21:57:31",
          "content": "<p>Hmm… Shouldn't the <code>history_num_frames</code> be 10? Or you won't have the same history as the test set problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091233,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "11/25/2020 22:05:19",
          "content": "<p><code>history_num_frames</code> is just a hyperparameter of how many history frames you are using in your rasterizer. You can go up to 100 as this is the max len in test. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091260,
          "author_name": "authman",
          "author_url": "",
          "post_date": "11/25/2020 22:32:53",
          "content": "<p>It's too late for me now 🙃</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091388,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/26/2020 01:16:07",
          "content": "<p>Oh shoo… I guess I completely misunderstood that..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091674,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "11/26/2020 07:32:20",
          "content": "<p>with min_frame_history = 1 and min_frame_future = 10 it resembles the test dataset best (and matches training loss with validation loss and test score)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1090995": "It's too late I know, but I wanted to try this last second hustle.\n\nMy model's training has *seemingly* been going okay form yesterday. Haven't been validating in train loop because on forums people said models need ~24h to converge for real, so I figured just let the thing do it's thing:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Faef19867fd6892eccc9dcbb647f47399%2Fwhat.png?generation=1606326887130863&alt=media)\n\nThe pictured model is a resnet34 trained for 140000 steps of 64-sized batches from random samples of train_full.zarr.\n\nSo the problem is, when I look at the most recent, say, 15k predictions `z.loss.iloc[-15000:].mean()`, I get a train loss of 17.9. That feels comfortable. When I [try and validate](https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min) against a sample from the validation set, I get back values that look like this:\n\n```\n{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n```\n\nYikes! I clearly jacked something up. I saw a few posts on people having similar issues but no definitive solution. Did I miss it / is it buried in the discussions somewhere? Do we have to scale our solutions by some multiplier or something?",
    "1091005": "For reference, it's not like val loss is moving down but... it feels like everything just shifted or something?\n\n```\nmodel_resnet34_output_84000.pth\n{'neg_multi_log_likelihood': 82.68555224218878, 'time_displace': array([0.08216104, 0.14191299, 0.17862896, 0.23167751, 0.29029346,\n       0.34343956, 0.39779761, 0.44261665, 0.48665519, 0.53399251,\n       0.57908257, 0.61783541, 0.65807086, 0.69021504, 0.71906   ,\n       0.75315359, 0.78096231, 0.81671914, 0.84816372, 0.8765357 ,\n       0.9001293 , 0.92258684, 0.94231826, 0.95649926, 0.97533073,\n       0.99249186, 1.01404677, 1.0282244 , 1.03808593, 1.04595177,\n       1.05631953, 1.06164656, 1.07050866, 1.07312559, 1.07921475,\n       1.09664675, 1.1066373 , 1.11167966, 1.11038037, 1.12044869,\n       1.12766502, 1.14252446, 1.15611682, 1.16043487, 1.17211799,\n       1.18489064, 1.20385188, 1.22025188, 1.24206752, 1.26920105])}\n\nmodel_resnet34_output_102000.pth\n{'neg_multi_log_likelihood': 60.15995701082701, 'time_displace': array([0.0814472 , 0.13519718, 0.18767383, 0.23921354, 0.29237729,\n       0.33924203, 0.38126515, 0.41581896, 0.45892417, 0.49914023,\n       0.53566297, 0.57807686, 0.6126809 , 0.64138179, 0.67282553,\n       0.69441783, 0.71572199, 0.73372158, 0.7567145 , 0.77424734,\n       0.78061963, 0.80467917, 0.82018216, 0.83072158, 0.84229814,\n       0.85418034, 0.8649891 , 0.87012398, 0.87173393, 0.87519373,\n       0.86675252, 0.87186382, 0.88307607, 0.88930812, 0.90104548,\n       0.91990247, 0.93296282, 0.93163789, 0.93750535, 0.9523782 ,\n       0.96132791, 0.96947078, 0.98608108, 0.99482964, 1.02291647,\n       1.04211571, 1.070679  , 1.08794308, 1.10915816, 1.13385227])}\n\nmodel_resnet34_output_140000.pth\n{'neg_multi_log_likelihood': 50.8512012161938, 'time_displace': array([0.07315565, 0.12655913, 0.17892002, 0.2330734 , 0.28420516,\n       0.32747597, 0.36595389, 0.40379383, 0.43888783, 0.47513634,\n       0.50971709, 0.54011187, 0.56920184, 0.59744106, 0.62130535,\n       0.65341146, 0.67258954, 0.70006751, 0.71884056, 0.73727427,\n       0.7496481 , 0.76607399, 0.78562313, 0.79520553, 0.80104267,\n       0.80774203, 0.81788417, 0.81726488, 0.81748707, 0.81580307,\n       0.81671028, 0.81795009, 0.81919481, 0.82293758, 0.83519629,\n       0.84076967, 0.85065792, 0.85731167, 0.85335456, 0.8624345 ,\n       0.86496159, 0.87687561, 0.89450577, 0.90489753, 0.92299398,\n       0.93986202, 0.96203589, 0.97899521, 1.00011826, 1.02050441])}\n\nmodel_resnet34_output_194000.pth\n{'neg_multi_log_likelihood': 36.540700766245884, 'time_displace': array([0.06768786, 0.11239789, 0.15527047, 0.19707848, 0.23415925,\n       0.27042494, 0.30206143, 0.32945963, 0.3599874 , 0.38719589,\n       0.41793629, 0.44404189, 0.46830858, 0.49044816, 0.50436334,\n       0.52224877, 0.53952061, 0.55973396, 0.57723214, 0.59279727,\n       0.60702455, 0.6243174 , 0.63574038, 0.65546905, 0.67029015,\n       0.68104914, 0.69563094, 0.70341774, 0.71044856, 0.70976926,\n       0.71465194, 0.71927185, 0.72643562, 0.73100875, 0.7356024 ,\n       0.74746589, 0.76176941, 0.7696118 , 0.77722741, 0.79098686,\n       0.80050605, 0.8083918 , 0.82216129, 0.83489121, 0.85430071,\n       0.87171541, 0.88682221, 0.90260645, 0.92327949, 0.94319656])}\n```\n\nYou can approximate where train loss was from the first picture at the corresponding time step.",
    "1091009": "Yeah validation loss can be twice as high as train loss. I once overfit the train set and got train loss = 17 but chopped validation set loss = 74, which is when I tried to do replay memory trick in one of the discussion.\n(Note that you need to validate on the \"chopped\" validation set to get real score)",
    "1091020": "Thank for for your comments. Currently, I am validating w/ chopped but training with train_full.\n\n> Yeah validation loss can be twice as high as train loss\n\nI see. In my case it's looking closer to ~2.8x, but... if that's expected and if everyone is encountering this, I guess it's fine. In some discussions I've read though, people are able to get train + valid scores neck-n-neck. I guess they must have been training on chopped trained.\n\nLast night before I started running my model, I followed [the steps here](https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset) in order to created a chopped train_full.zarr set. After about 2 and a half hours of processing, it was finally ready to train on. But when I attempted to train, train loss started and stayed at 0.0(?). It was already like 2am at that time, so I figured forget it.. I'll just train with train_full since some kernels scored 23.x just using the regular train set.\n\n> which is when I tried to do replay memory trick\n\nEh? What's that",
    "1091055": "The chop dataset function provided in l5kit removes the future positions for all the agents so you can't use that to train the model directly.\nBut I guess if you do it properly, this sounds a really good idea. I wish I have heard this earlier...\nThe replay memory trick is to store the loaded training batches in memory and randomly retrain on those cached batches multiple time during training to speed up since most of the computation time of this competition is spent on the loading the data but not training. I saw this suggest by someone in the discussion. Though it lead to overfit when I tried it.",
    "1091158": "What are your args when you call AgentDataset on the training set?",
    "1091183": "```\ncfg = {\n    'format_version': 4,\n    'data_path': \"../input/lyft-motion-prediction-autonomous-vehicles\",\n    'model_params': {\n        'model_architecture': 'resnet34',\n        'history_num_frames': 6,\n        'history_step_size': 1,\n        'history_delta_time': 0.1,\n        'future_num_frames': 50,\n        'future_step_size': 1,\n        'future_delta_time': 0.1,\n        'model_name': \"model_resnet34_output\",\n        'lr': 1e-3, #1e-3,        \n        'train': True,\n        'predict': True\n    },\n\n    'raster_params': {\n        'raster_size': [224*2, 224//2],\n        'pixel_size': [0.3, 0.3],\n        'ego_center': [0.2, 0.5], # [0.25, 0.5]\n        'map_type': 'py_semantic',\n        'satellite_map_key': 'aerial_map/aerial_map.png',\n        'semantic_map_key': 'semantic_map/semantic_map.pb',\n        'dataset_meta_key': 'meta.json',\n        'filter_agents_threshold': 0.5 \n    },\n\n    'train_data_loader': {\n        'key': 'scenes/train_full.zarr',\n        'batch_size': 64,\n        'shuffle': True,\n        'num_workers': 4 if KAGGLE else 11\n    },\n}\n```",
    "1091225": "I actually mean the `AgentDataset(cfg, train_zarr ...)`\nThe cfg looks fine, though. \nWhat did you chose for `min_frame_history` and `min_frame_future`?",
    "1091227": "Hmm... Shouldn't the `history_num_frames` be 10? Or you won't have the same history as the test set problem.",
    "1091233": "`history_num_frames` is just a hyperparameter of how many history frames you are using in your rasterizer. You can go up to 100 as this is the max len in test.",
    "1091260": "It's too late for me now 🙃",
    "1091388": "Oh shoo... I guess I completely misunderstood that..",
    "1091674": "with min_frame_history = 1 and min_frame_future = 10 it resembles the test dataset best (and matches training loss with validation loss and test score)"
  },
  "source": "meta"
}