{
  "id": 188695,
  "title": "Validation vs LB score",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/188695",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2020-10-04T18:11:28.717000",
  "votes": 45,
  "comment_count": 43,
  "views": 0,
  "content": "<p>I would like to get a little more light onto the validation strategy that everyone is using. </p>\n<p>The hosts are suggesting to use a chopped dataset for validation, as this was also used to create the test dataset.</p>\n<p>There have been some discussions of how the parameters have been chosen in this process. Maybe we can resolve that, when comparing our past results, so let me start.</p>\n<p>Method: </p>\n<pre><code>create_chopped_dataset on validation.zarr\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\n-&gt; 94694 agents to predict\n</code></pre>\n<p>Validation:</p>\n<pre><code>neg_multi_log_likelihood 15.90226142426006\ntime_displace [0.03825686 0.05979496 0.07908053 0.09720777 0.11376195 0.13041205\n 0.14533137 0.16254424 0.17893563 0.19681822 0.21209035 0.22800589\n 0.24817092 0.2667516  0.28417274 0.29907279 0.3146695  0.32590776\n 0.33657942 0.34686728 0.35946204 0.37245056 0.38405101 0.39074047\n 0.40032023 0.4101571  0.42140985 0.43188392 0.43951849 0.44799483\n 0.45873317 0.46933649 0.47870115 0.48539507 0.49432693 0.50711085\n 0.51884592 0.52831929 0.53900899 0.55048483 0.56373485 0.57716685\n 0.59198183 0.60644297 0.62096873 0.63373832 0.65085737 0.66873676\n 0.68705682 0.70606301]\n</code></pre>\n<p>Public LB: </p>\n<pre><code>15.768\n</code></pre>\n<p>I think this validation strategy agrees well with the public LB score and probably will translate to the private LB. To further validate that hypothesis, please post your results ;)</p>",
  "messages": [
    {
      "id": 1037244,
      "postDate": "2020-10-04T18:11:28.717Z",
      "content": "<p>I would like to get a little more light onto the validation strategy that everyone is using. </p>\n<p>The hosts are suggesting to use a chopped dataset for validation, as this was also used to create the test dataset.</p>\n<p>There have been some discussions of how the parameters have been chosen in this process. Maybe we can resolve that, when comparing our past results, so let me start.</p>\n<p>Method: </p>\n<pre><code>create_chopped_dataset on validation.zarr\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\n-&gt; 94694 agents to predict\n</code></pre>\n<p>Validation:</p>\n<pre><code>neg_multi_log_likelihood 15.90226142426006\ntime_displace [0.03825686 0.05979496 0.07908053 0.09720777 0.11376195 0.13041205\n 0.14533137 0.16254424 0.17893563 0.19681822 0.21209035 0.22800589\n 0.24817092 0.2667516  0.28417274 0.29907279 0.3146695  0.32590776\n 0.33657942 0.34686728 0.35946204 0.37245056 0.38405101 0.39074047\n 0.40032023 0.4101571  0.42140985 0.43188392 0.43951849 0.44799483\n 0.45873317 0.46933649 0.47870115 0.48539507 0.49432693 0.50711085\n 0.51884592 0.52831929 0.53900899 0.55048483 0.56373485 0.57716685\n 0.59198183 0.60644297 0.62096873 0.63373832 0.65085737 0.66873676\n 0.68705682 0.70606301]\n</code></pre>\n<p>Public LB: </p>\n<pre><code>15.768\n</code></pre>\n<p>I think this validation strategy agrees well with the public LB score and probably will translate to the private LB. To further validate that hypothesis, please post your results ;)</p>",
      "rawMarkdown": "I would like to get a little more light onto the validation strategy that everyone is using. \n\nThe hosts are suggesting to use a chopped dataset for validation, as this was also used to create the test dataset.\n\nThere have been some discussions of how the parameters have been chosen in this process. Maybe we can resolve that, when comparing our past results, so let me start.\n\nMethod: \n```\ncreate_chopped_dataset on validation.zarr\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\n-> 94694 agents to predict\n```\n\nValidation:\n```\nneg_multi_log_likelihood 15.90226142426006\ntime_displace [0.03825686 0.05979496 0.07908053 0.09720777 0.11376195 0.13041205\n 0.14533137 0.16254424 0.17893563 0.19681822 0.21209035 0.22800589\n 0.24817092 0.2667516  0.28417274 0.29907279 0.3146695  0.32590776\n 0.33657942 0.34686728 0.35946204 0.37245056 0.38405101 0.39074047\n 0.40032023 0.4101571  0.42140985 0.43188392 0.43951849 0.44799483\n 0.45873317 0.46933649 0.47870115 0.48539507 0.49432693 0.50711085\n 0.51884592 0.52831929 0.53900899 0.55048483 0.56373485 0.57716685\n 0.59198183 0.60644297 0.62096873 0.63373832 0.65085737 0.66873676\n 0.68705682 0.70606301]\n```\n\nPublic LB: \n```\n15.768\n```\n\nI think this validation strategy agrees well with the public LB score and probably will translate to the private LB. To further validate that hypothesis, please post your results ;)",
      "votes": 45
    },
    {
      "id": 1038000,
      "postDate": "2020-10-05T13:23:32.010Z",
      "content": "<p>Provided validation set nll and lb score is very close for me as well. But full validation takes a good amount of time. Does anybody know a subset of the validation set that is reasonably good as well?</p>",
      "rawMarkdown": "Provided validation set nll and lb score is very close for me as well. But full validation takes a good amount of time. Does anybody know a subset of the validation set that is reasonably good as well?",
      "votes": 3,
      "replies": [
        {
          "id": 1038236,
          "postDate": "2020-10-05T16:58:55.673Z",
          "content": "<p>I've found that the data is very stable. Taking a 10K subset of the chopped validation set gives me a slightly lower LB than CV but the preferences are always consistent: better CV =&gt; better LB in every test I've run so far.</p>",
          "rawMarkdown": "I've found that the data is very stable. Taking a 10K subset of the chopped validation set gives me a slightly lower LB than CV but the preferences are always consistent: better CV => better LB in every test I've run so far.",
          "votes": 7
        },
        {
          "id": 1043404,
          "postDate": "2020-10-09T00:24:28.400Z",
          "content": "<p>May be a good idea to cache rendered images at least for the validation set, should make it quite a bit faster.</p>",
          "rawMarkdown": "May be a good idea to cache rendered images at least for the validation set, should make it quite a bit faster.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1037359,
      "postDate": "2020-10-04T22:39:06.810Z",
      "content": "<p>Thanks for sharing. Same conclusion on my side. <br>\nPublic LB <code>17.449</code> and local validation (same setting as above) </p>\n<pre><code>neg_multi_log_likelihood 17.639541738821357\ntime_displace [0.043765   0.06239079 0.08461178 0.1045251  0.12230116 0.13607619\n 0.15364686 0.17026302 0.18461843 0.20065285 0.21922876 0.23798571\n 0.25381047 0.26798241 0.28671686 0.30309549 0.32077815 0.33500339\n 0.34464988 0.35790864 0.37138817 0.38228462 0.39201591 0.40421648\n 0.41869074 0.4339096  0.44838353 0.45712691 0.46366225 0.47801506\n 0.49306092 0.4955435  0.50572086 0.51511343 0.52933362 0.53734122\n 0.55293909 0.56172705 0.57153331 0.58700482 0.59916493 0.61534454\n 0.62890407 0.64430313 0.66137071 0.67384805 0.68635556 0.70236678\n 0.71957936 0.73788754]\n</code></pre>",
      "rawMarkdown": "Thanks for sharing. Same conclusion on my side. \nPublic LB `17.449` and local validation (same setting as above) \n```\nneg_multi_log_likelihood 17.639541738821357\ntime_displace [0.043765   0.06239079 0.08461178 0.1045251  0.12230116 0.13607619\n 0.15364686 0.17026302 0.18461843 0.20065285 0.21922876 0.23798571\n 0.25381047 0.26798241 0.28671686 0.30309549 0.32077815 0.33500339\n 0.34464988 0.35790864 0.37138817 0.38228462 0.39201591 0.40421648\n 0.41869074 0.4339096  0.44838353 0.45712691 0.46366225 0.47801506\n 0.49306092 0.4955435  0.50572086 0.51511343 0.52933362 0.53734122\n 0.55293909 0.56172705 0.57153331 0.58700482 0.59916493 0.61534454\n 0.62890407 0.64430313 0.66137071 0.67384805 0.68635556 0.70236678\n 0.71957936 0.73788754]\n```",
      "votes": 3
    },
    {
      "id": 1041857,
      "postDate": "2020-10-08T00:11:38.053Z",
      "content": "<p>If There is someone looking for reference to predict using Validate.Zarr and chopped data. You can refer this Notebook. The Notebook also saves gt.csv from temp file. Big thanks to <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and <a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> . The Pipeline Notebook is from based on their work</p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/deepakrajpurushothaman/predict-using-chopped-data-and-know-your-score</a></p>",
      "rawMarkdown": "If There is someone looking for reference to predict using Validate.Zarr and chopped data. You can refer this Notebook. The Notebook also saves gt.csv from temp file. Big thanks to @corochann and @thomasbrandon . The Pipeline Notebook is from based on their work\n\n[https://www.kaggle.com/deepakrajpurushothaman/predict-using-chopped-data-and-know-your-score](url)",
      "votes": 4,
      "replies": [
        {
          "id": 1044241,
          "postDate": "2020-10-09T16:09:57.523Z",
          "content": "<p>It seems that your notebook is private.please make it public.</p>",
          "rawMarkdown": "It seems that your notebook is private.please make it public.",
          "votes": 4
        },
        {
          "id": 1053447,
          "postDate": "2020-10-19T02:37:22.850Z",
          "content": "<p>Thanks for mention <a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a> :)</p>",
          "rawMarkdown": "Thanks for mention @deepakrajpurushothaman :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1079312,
      "postDate": "2020-11-15T22:24:03.230Z",
      "content": "<p>Check my notebook, I wrote a short script, which you can verify your model LB score. It takes only 5000k sample from the chopped dataset, and it has around +-0.3 accuracy.<br>\n<a href=\"https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min\" target=\"_blank\">https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min</a></p>",
      "rawMarkdown": "Check my notebook, I wrote a short script, which you can verify your model LB score. It takes only 5000k sample from the chopped dataset, and it has around +-0.3 accuracy.\nhttps://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min",
      "votes": 2,
      "replies": [
        {
          "id": 1079545,
          "postDate": "2020-11-16T07:21:21.303Z",
          "content": "<p>It's looks fine! I don't know why somebody vote down</p>",
          "rawMarkdown": "It's looks fine! I don't know why somebody vote down",
          "votes": 1
        }
      ]
    },
    {
      "id": 1074981,
      "postDate": "2020-11-11T09:36:51.170Z",
      "content": "<p>Hi folks, <br>\n  Thank for you sharing out the val &amp; PB scores. I'm wondering if your training loss aligns with validation loss? Mine shows a pretty large discrepancy. For example, <code>training loss ~ 20+, validation loss ~ 60+</code>. Currently, I'm iterating through <code>train.zarr</code> for training <code>1000 steps x batch_size=128</code> and use <code>chopp_dataset of full validate.zarr</code> for validation. Any experience in such issue here? </p>\n<p>Thank you! </p>",
      "rawMarkdown": "Hi folks, \n  Thank for you sharing out the val & PB scores. I'm wondering if your training loss aligns with validation loss? Mine shows a pretty large discrepancy. For example, `training loss ~ 20+, validation loss ~ 60+`. Currently, I'm iterating through `train.zarr` for training `1000 steps x batch_size=128` and use `chopp_dataset of full validate.zarr` for validation. Any experience in such issue here? \n\n Thank you! ",
      "votes": 1,
      "replies": [
        {
          "id": 1075078,
          "postDate": "2020-11-11T11:25:07.273Z",
          "content": "<p>I think 1000 iterations is a very small dataset to achieve better validation score. If its okay to share, How many epochs have you trained on the 1000 batches? </p>",
          "rawMarkdown": "I think 1000 iterations is a very small dataset to achieve better validation score. If its okay to share, How many epochs have you trained on the 1000 batches? "
        },
        {
          "id": 1075094,
          "postDate": "2020-11-11T11:35:49.810Z",
          "content": "<p>Hi there,  1000 iterations are for my training, and my training batch size is 128. On the contrary, many folks may use batch_size=32 for 10K iterations…. Maybe that's the difference?</p>\n<p>Sure thing, my latest experiment is still running and it has bee training for 9 epochs… Maybe that's not enough either?</p>\n<p>Thanks!</p>",
          "rawMarkdown": "Hi there,  1000 iterations are for my training, and my training batch size is 128. On the contrary, many folks may use batch_size=32 for 10K iterations.... Maybe that's the difference?\n\nSure thing, my latest experiment is still running and it has bee training for 9 epochs... Maybe that's not enough either?\n\nThanks!",
          "votes": 1,
          "replies": [
            {
              "id": 1075110,
              "postDate": "2020-11-11T11:55:07.030Z",
              "content": "<p>I think you have overfit the training set. 1000 iterations on batch size of 128 for 9 epochs means your model has kinda memorized the training set and hence the lower training score. It's also the reason your model gives larger error on validation set as it doesn't generalise well. </p>",
              "rawMarkdown": "I think you have overfit the training set. 1000 iterations on batch size of 128 for 9 epochs means your model has kinda memorized the training set and hence the lower training score. It's also the reason your model gives larger error on validation set as it doesn't generalise well. "
            },
            {
              "id": 1075136,
              "postDate": "2020-11-11T12:24:57.427Z",
              "content": "<p>I'm kind of lost in my experiment results. </p>\n<p>Previously,  my training loss &amp; validate loss moved downwards together and matched each other because I trained &amp; validated on the original train.zarr &amp; validate.zarr dataset.  But when I used <strong>chopped validate</strong> dataset, the discrepancy started to appear… </p>\n<p>And I think maybe your overfitting theory may be correct, but I probably need to restart a new training session and see what happens. </p>\n<p>Thanks for your help!</p>",
              "rawMarkdown": "I'm kind of lost in my experiment results. \n\nPreviously,  my training loss & validate loss moved downwards together and matched each other because I trained & validated on the original train.zarr & validate.zarr dataset.  But when I used **chopped validate** dataset, the discrepancy started to appear... \n\nAnd I think maybe your overfitting theory may be correct, but I probably need to restart a new training session and see what happens. \n\nThanks for your help!"
            },
            {
              "id": 1075140,
              "postDate": "2020-11-11T12:36:41.893Z",
              "content": "<p>Sure the chopped dataset clearly correlates with LB and not Full data</p>",
              "rawMarkdown": "Sure the chopped dataset clearly correlates with LB and not Full data"
            },
            {
              "id": 1075143,
              "postDate": "2020-11-11T12:40:46.467Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>, sorry I'm kinda lost in you reply. Could you elaborate a litter bit?</p>",
              "rawMarkdown": "Hi @morizin, sorry I'm kinda lost in you reply. Could you elaborate a litter bit?"
            },
            {
              "id": 1075178,
              "postDate": "2020-11-11T13:17:33.307Z",
              "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>, sorry I'm kinda lost in you reply. Could you elaborate a litter bit?<br>\n  Sorry to hear that</p>\n</blockquote>\n<p>I even tried to use validate without chopped version. while we do eval after certain training in non chopped datset , its sure the validation correlates with training score and but not LB<br>\nthe case is different in chopped validate dataset. I might not be correlating verwell on training score but the score on LB is better than what you see in Validation score</p>",
              "rawMarkdown": "> Hi @morizin, sorry I'm kinda lost in you reply. Could you elaborate a litter bit?\nSorry to hear that\n\nI even tried to use validate without chopped version. while we do eval after certain training in non chopped datset , its sure the validation correlates with training score and but not LB\nthe case is different in chopped validate dataset. I might not be correlating verwell on training score but the score on LB is better than what you see in Validation score"
            }
          ]
        }
      ]
    },
    {
      "id": 1037576,
      "postDate": "2020-10-05T06:19:23.607Z",
      "content": "<pre><code>neg_multi_log_likelihood 14.785969823499773\ntime_displace [0.0381546  0.05812458 0.07756729 0.09666313 0.11373225 0.13055148\n 0.14453314 0.16217867 0.17796376 0.19276053 0.21084736 0.22936855\n 0.24707832 0.26275199 0.27871745 0.29462635 0.30701878 0.31865779\n 0.32920492 0.33891281 0.34976245 0.36036229 0.36911364 0.37830991\n 0.38650796 0.39458455 0.403425   0.41191151 0.42092284 0.42943664\n 0.43824895 0.44682792 0.45746868 0.46702098 0.47547127 0.48541446\n 0.49275159 0.50045443 0.51071008 0.52505713 0.53472998 0.5485291\n 0.56223968 0.57483904 0.58859108 0.60047564 0.61616895 0.63170392\n 0.64937508 0.66797937]\n</code></pre>\n<p>Public LB:</p>\n<pre><code>14.803\n</code></pre>",
      "rawMarkdown": "```\nneg_multi_log_likelihood 14.785969823499773\ntime_displace [0.0381546  0.05812458 0.07756729 0.09666313 0.11373225 0.13055148\n 0.14453314 0.16217867 0.17796376 0.19276053 0.21084736 0.22936855\n 0.24707832 0.26275199 0.27871745 0.29462635 0.30701878 0.31865779\n 0.32920492 0.33891281 0.34976245 0.36036229 0.36911364 0.37830991\n 0.38650796 0.39458455 0.403425   0.41191151 0.42092284 0.42943664\n 0.43824895 0.44682792 0.45746868 0.46702098 0.47547127 0.48541446\n 0.49275159 0.50045443 0.51071008 0.52505713 0.53472998 0.5485291\n 0.56223968 0.57483904 0.58859108 0.60047564 0.61616895 0.63170392\n 0.64937508 0.66797937]\n``` \n\nPublic LB:\n```\n14.803\n```",
      "votes": 2,
      "replies": [
        {
          "id": 1044427,
          "postDate": "2020-10-09T18:47:54.983Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>,</p>\n<p>Thanks for sharing. I am using the same validation scheme with the same results as to agreement with LB. What about your training score/loss? Does this agree with valid/LB?</p>",
          "rawMarkdown": "Hi @ilu000,\n\nThanks for sharing. I am using the same validation scheme with the same results as to agreement with LB. What about your training score/loss? Does this agree with valid/LB?"
        },
        {
          "id": 1044492,
          "postDate": "2020-10-09T20:16:42.597Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a> depending on the backbone and the NN structure, the training loss agrees quite well with the validation score and LB score. But that is only true if you are not using too much training data. I would think maybe around 10% and as mentioned above not for all NN structures. </p>\n<p>So, validation is important to get a sense of how good a model is actually doing. </p>",
          "rawMarkdown": "Hi @yousof9 depending on the backbone and the NN structure, the training loss agrees quite well with the validation score and LB score. But that is only true if you are not using too much training data. I would think maybe around 10% and as mentioned above not for all NN structures. \n\nSo, validation is important to get a sense of how good a model is actually doing. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1037450,
      "postDate": "2020-10-05T03:24:06.620Z",
      "content": "<p>I can confirm this validation strategy matched LB for me as well, but only on the single sample/submit so far</p>",
      "rawMarkdown": "I can confirm this validation strategy matched LB for me as well, but only on the single sample/submit so far",
      "votes": 2
    },
    {
      "id": 1078553,
      "postDate": "2020-11-15T01:24:55.700Z",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> Thanks for sharing a lot of useful informations in this competition!<br>\nMay I ask one more maybe silly question?<br>\nWhen I created val dataset by following <a href=\"https://github.com/lyft/l5kit/blob/03eb9e037d23940a134e27c9f124021e18982020/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">this</a> or <a href=\"https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset/comments\" target=\"_blank\">this</a>, it outputs data with all target_availabilities = 0. Then how can I calculate a correct loss without target_availabilities ? Or do you have any idea where I made mistake??</p>\n<p>Thanks in advance, good luck for the competition end!</p>",
      "rawMarkdown": "@ilu000 Thanks for sharing a lot of useful informations in this competition!\nMay I ask one more maybe silly question?\nWhen I created val dataset by following [this](https://github.com/lyft/l5kit/blob/03eb9e037d23940a134e27c9f124021e18982020/examples/agent_motion_prediction/agent_motion_prediction.ipynb) or [this](https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset/comments), it outputs data with all target_availabilities = 0. Then how can I calculate a correct loss without target_availabilities ? Or do you have any idea where I made mistake??\n\nThanks in advance, good luck for the competition end!\n",
      "replies": [
        {
          "id": 1078558,
          "postDate": "2020-11-15T01:49:34.777Z",
          "content": "<p>I think creating the validation set using create_chopped_dataset generates a mask.npz file which you can use to get agents for prediction</p>",
          "rawMarkdown": "I think creating the validation set using create_chopped_dataset generates a mask.npz file which you can use to get agents for prediction"
        },
        {
          "id": 1078562,
          "postDate": "2020-11-15T02:02:10.147Z",
          "content": "<p><a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a> <br>\nThanks for help!<br>\nI created the dataset by following official notebook like that.<br>\nI believe it uses the mask created by create_chopped_dataset.</p>\n<pre><code>eval_zarr_path = str(Path(eval_base_path) / Path(dm.require(eval_cfg[\"key\"])).name)\neval_mask_path = str(Path(eval_base_path) / \"mask.npz\")\neval_gt_path = str(Path(eval_base_path) / \"gt.csv\")\n\neval_zarr = ChunkedDataset(eval_zarr_path).open()\neval_mask = np.load(eval_mask_path)[\"arr_0\"]\n# ===== INIT DATASET AND LOAD MASK\neval_dataset = AgentDataset(cfg, eval_zarr, rasterizer, agents_mask=eval_mask)\neval_dataloader = DataLoader(eval_dataset, shuffle=eval_cfg[\"shuffle\"], batch_size=eval_cfg[\"batch_size\"], \n                             num_workers=eval_cfg[\"num_workers\"])\nprint(eval_dataset)\n</code></pre>\n<p>But when I check the data in the dataloader, all of the target_availabilities is zero.</p>\n<pre><code>data = next(iter(eval_dataloader))\nprint(data['target_availabilities'].sum())\n# tensor(0.)\n</code></pre>\n<blockquote>\n  <p>I think creating the validation set using create_chopped_dataset generates a mask.npz file which you can use to get agents for prediction</p>\n</blockquote>\n<p>That means this dataset cannot be directly used for validation? Do we need to map target availability manually??</p>\n<p>Thanks!</p>",
          "rawMarkdown": "@yousof9 \nThanks for help!\nI created the dataset by following official notebook like that.\nI believe it uses the mask created by create_chopped_dataset.\n```\neval_zarr_path = str(Path(eval_base_path) / Path(dm.require(eval_cfg[\"key\"])).name)\neval_mask_path = str(Path(eval_base_path) / \"mask.npz\")\neval_gt_path = str(Path(eval_base_path) / \"gt.csv\")\n\neval_zarr = ChunkedDataset(eval_zarr_path).open()\neval_mask = np.load(eval_mask_path)[\"arr_0\"]\n# ===== INIT DATASET AND LOAD MASK\neval_dataset = AgentDataset(cfg, eval_zarr, rasterizer, agents_mask=eval_mask)\neval_dataloader = DataLoader(eval_dataset, shuffle=eval_cfg[\"shuffle\"], batch_size=eval_cfg[\"batch_size\"], \n                             num_workers=eval_cfg[\"num_workers\"])\nprint(eval_dataset)\n```\n\nBut when I check the data in the dataloader, all of the target_availabilities is zero.\n```\ndata = next(iter(eval_dataloader))\nprint(data['target_availabilities'].sum())\n# tensor(0.)\n```\n\n> I think creating the validation set using create_chopped_dataset generates a mask.npz file which you can use to get agents for prediction\n\nThat means this dataset cannot be directly used for validation? Do we need to map target availability manually??\n\nThanks!"
        },
        {
          "id": 1078720,
          "postDate": "2020-11-15T08:25:21.140Z",
          "content": "<p>True, the created mask.npz must be used. Target availabilities are given through the gt.csv which is also created during the chopping process. This is the same during with the test set. We don't know anything about the target avails there. But it is used for the metric calculation on the kaggle scoreboard. </p>",
          "rawMarkdown": "True, the created mask.npz must be used. Target availabilities are given through the gt.csv which is also created during the chopping process. This is the same during with the test set. We don't know anything about the target avails there. But it is used for the metric calculation on the kaggle scoreboard. "
        },
        {
          "id": 1079309,
          "postDate": "2020-11-15T22:18:06.657Z",
          "content": "<p>Thanks, now I decided to validate separately from my training loop.</p>",
          "rawMarkdown": "Thanks, now I decided to validate separately from my training loop."
        }
      ]
    },
    {
      "id": 1062729,
      "postDate": "2020-10-28T06:24:16.183Z",
      "content": "<p>I tried implementing the validation from the config file   </p>\n<pre><code>'val_data_loader': {\n  'key': 'scenes/validate.zarr',\n  'batch_size': 16,\n  'shuffle': False,\n  'num_workers': 4\n}\n</code></pre>\n<p>num_frames_to_chop = 100<br>\neval_cfg = cfg[\"val_data_loader\"]<br>\neval_base_path = create_chopped_dataset(dm.require(eval_cfg[\"key\"]), cfg[\"raster_params\"][\"filter_agents_threshold\"], <br>\n                              num_frames_to_chop, cfg[\"model_params\"][\"future_num_frames\"], MIN_FUTURE_STEPS)</p>\n<p>but it gave me an OSError: [Errno 30] Read-only file system: '/kaggle/input/lyft-motion-prediction-autonomous-vehicles/scenes/sample_chopped_100'<br>\nas Kaggle doesn't provide access to write files in the input directory but how to change the /kaggle/input to /kaggle/working ?</p>\n<p>What am I doing wrong ?</p>",
      "rawMarkdown": "I tried implementing the validation from the config file   \n\n    'val_data_loader': {\n      'key': 'scenes/validate.zarr',\n      'batch_size': 16,\n      'shuffle': False,\n      'num_workers': 4\n    }\nnum_frames_to_chop = 100\neval_cfg = cfg[\"val_data_loader\"]\neval_base_path = create_chopped_dataset(dm.require(eval_cfg[\"key\"]), cfg[\"raster_params\"][\"filter_agents_threshold\"], \n                              num_frames_to_chop, cfg[\"model_params\"][\"future_num_frames\"], MIN_FUTURE_STEPS)\n\nbut it gave me an OSError: [Errno 30] Read-only file system: '/kaggle/input/lyft-motion-prediction-autonomous-vehicles/scenes/sample_chopped_100'\nas Kaggle doesn't provide access to write files in the input directory but how to change the /kaggle/input to /kaggle/working ?\n\nWhat am I doing wrong ?",
      "replies": [
        {
          "id": 1062741,
          "postDate": "2020-10-28T06:34:50.323Z",
          "content": "<p>You can try using this book. This will definitely help. <a href=\"https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset\" target=\"_blank\">https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset</a></p>",
          "rawMarkdown": "You can try using this book. This will definitely help. https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset",
          "votes": 3
        },
        {
          "id": 1062745,
          "postDate": "2020-10-28T06:37:11.553Z",
          "content": "<p>Thank You, this is what I was looking for.<br>\nOne more thing, do we have to use 'sample.zarr' or 'validate.zarr' for the validation ?</p>",
          "rawMarkdown": "Thank You, this is what I was looking for.\nOne more thing, do we have to use 'sample.zarr' or 'validate.zarr' for the validation ?",
          "votes": 1,
          "replies": [
            {
              "id": 1062754,
              "postDate": "2020-10-28T06:46:57.193Z",
              "content": "<p>We have to used Validate.zarr, sample.zarr is just for reference.</p>",
              "rawMarkdown": "We have to used Validate.zarr, sample.zarr is just for reference.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 1047626,
      "postDate": "2020-10-12T19:12:48.297Z",
      "content": "<p>I'm interested how you all are preparing training data now. With the targets in meters or pixels and how do you get validation to align with training numbers now? Right now I have configured my training loop to operate on meters by simply dividing the output from the dataloader by the pixel size and then multiplying by the pixel size at inference time. Training and validation no longer seem to perfectly line up but it performs the same on validation as when I was using the old dataloader and transforming the data myself. </p>",
      "rawMarkdown": "I'm interested how you all are preparing training data now. With the targets in meters or pixels and how do you get validation to align with training numbers now? Right now I have configured my training loop to operate on meters by simply dividing the output from the dataloader by the pixel size and then multiplying by the pixel size at inference time. Training and validation no longer seem to perfectly line up but it performs the same on validation as when I was using the old dataloader and transforming the data myself. ",
      "replies": [
        {
          "id": 1047659,
          "postDate": "2020-10-12T20:20:09.093Z",
          "content": "<p>I stayed at meters. No further transformation done, yet. </p>",
          "rawMarkdown": "I stayed at meters. No further transformation done, yet. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 1044086,
      "postDate": "2020-10-09T14:08:30.517Z",
      "content": "<p>Can you give me the notebook of validation if it is available in public .i am just going through the discussion but cannot understand where you implemented</p>",
      "rawMarkdown": "Can you give me the notebook of validation if it is available in public .i am just going through the discussion but cannot understand where you implemented",
      "replies": [
        {
          "id": 1044517,
          "postDate": "2020-10-09T20:43:16.270Z",
          "content": "<p>Do you mean <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb</a> (Under the point Evaluation)</p>",
          "rawMarkdown": "Do you mean https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb (Under the point Evaluation)",
          "votes": 6
        },
        {
          "id": 1045298,
          "postDate": "2020-10-10T13:56:24.263Z",
          "content": "<p>Thank you for this comment</p>",
          "rawMarkdown": "Thank you for this comment"
        }
      ]
    },
    {
      "id": 1041807,
      "postDate": "2020-10-07T23:20:50.913Z",
      "content": "<p>With new L5kit, I also confirm its true. Only difference is method. I guess you and some people were putting other values for Min_Future_Frame like 10. Mine remains at default 1.   The problem with train loss and eval loss or test loss still remians. Thats Huge difference. </p>\n<p>Method:<br>\ncreate_chopped_dataset on validation.zarr<br>\nnum_frames_to_chop = 100<br>\nMIN_FUTURE_STEPS = 10<br>\nMIN_FUTURE_FRAME=1<br>\n-&gt; 94694 agents to predict</p>\n<p>Train: Trin_Full.zarr for 1.5M samples<br>\nl5kit version: 1.1.0<br>\nLB Score: 279.xx<br>\nEval Score :<br>\nneg_multi_log_likelihood 267.17160372981397<br>\ntime_displace [0.11229513 0.16690064 0.22574588 0.28469734 0.34264851 0.40680175<br>\n 0.46570777 0.52782559 0.58903438 0.64639612 0.70205706 0.75062746<br>\n 0.80532094 0.85498418 0.90345951 0.94995913 0.9971724  1.04381954<br>\n 1.08859232 1.1342783  1.17871579 1.21726209 1.26382939 1.31105303<br>\n 1.3562235  1.39718102 1.44048208 1.48720213 1.52771844 1.56625857<br>\n 1.61256805 1.65881987 1.70607001 1.74212776 1.77483601 1.81985783<br>\n 1.86126298 1.90600642 1.94457169 1.99220685 2.04489769 2.07543546<br>\n 2.11432232 2.15322523 2.19668563 2.23679864 2.26960409 2.30307122<br>\n 2.36240464 2.39407803]</p>",
      "rawMarkdown": "With new L5kit, I also confirm its true. Only difference is method. I guess you and some people were putting other values for Min_Future_Frame like 10. Mine remains at default 1.   The problem with train loss and eval loss or test loss still remians. Thats Huge difference. \n\nMethod:\ncreate_chopped_dataset on validation.zarr\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\nMIN_FUTURE_FRAME=1\n-> 94694 agents to predict\n\nTrain: Trin_Full.zarr for 1.5M samples\nl5kit version: 1.1.0\nLB Score: 279.xx\nEval Score :\nneg_multi_log_likelihood 267.17160372981397\ntime_displace [0.11229513 0.16690064 0.22574588 0.28469734 0.34264851 0.40680175\n 0.46570777 0.52782559 0.58903438 0.64639612 0.70205706 0.75062746\n 0.80532094 0.85498418 0.90345951 0.94995913 0.9971724  1.04381954\n 1.08859232 1.1342783  1.17871579 1.21726209 1.26382939 1.31105303\n 1.3562235  1.39718102 1.44048208 1.48720213 1.52771844 1.56625857\n 1.61256805 1.65881987 1.70607001 1.74212776 1.77483601 1.81985783\n 1.86126298 1.90600642 1.94457169 1.99220685 2.04489769 2.07543546\n 2.11432232 2.15322523 2.19668563 2.23679864 2.26960409 2.30307122\n 2.36240464 2.39407803]"
    },
    {
      "id": 1037824,
      "postDate": "2020-10-05T11:11:25.973Z",
      "content": "<p>I just iter over a valid_dataloader for 1500 Frames every 5000 Training Frames:</p>\n<p>Validation (NLLLoss): 39.49<br>\nPublic-LB: 44.66</p>",
      "rawMarkdown": "I just iter over a valid_dataloader for 1500 Frames every 5000 Training Frames:\n\nValidation (NLLLoss): 39.49\nPublic-LB: 44.66",
      "replies": [
        {
          "id": 1038214,
          "postDate": "2020-10-05T16:45:45.973Z",
          "content": "<p>You need to use <code>create_chopped_dataset</code> as the dataset is serially correlated in time</p>",
          "rawMarkdown": "You need to use `create_chopped_dataset` as the dataset is serially correlated in time",
          "votes": 1
        },
        {
          "id": 1038317,
          "postDate": "2020-10-05T18:14:20.723Z",
          "content": "<p>and it will save you a lot of time ;)</p>",
          "rawMarkdown": "and it will save you a lot of time ;)",
          "votes": 1
        },
        {
          "id": 1038493,
          "postDate": "2020-10-05T20:33:55.077Z",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> saving time is a good argument, but I wonder how it worked so well.</p>",
          "rawMarkdown": "@ilu000 saving time is a good argument, but I wonder how it worked so well."
        }
      ]
    },
    {
      "id": 1037670,
      "postDate": "2020-10-05T08:19:48.690Z",
      "content": "<p>Same for me. Had a similar experience with correlation between submits and validating as they show in the example notebook with the chopped datasets</p>",
      "rawMarkdown": "Same for me. Had a similar experience with correlation between submits and validating as they show in the example notebook with the chopped datasets"
    },
    {
      "id": 1037367,
      "postDate": "2020-10-04T23:15:24.447Z",
      "content": "<p>Wow, solid cross validation :) CV ~ Public LB. Thanks for sharing <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/ggrizzly\" target=\"_blank\">@ggrizzly</a> </p>",
      "rawMarkdown": "Wow, solid cross validation :) CV ~ Public LB. Thanks for sharing @ilu000 @ggrizzly "
    },
    {
      "id": 1068502,
      "postDate": "2020-11-03T13:28:48.583Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1038000,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2020-10-05T13:23:32.010000",
      "content": "<p>Provided validation set nll and lb score is very close for me as well. But full validation takes a good amount of time. Does anybody know a subset of the validation set that is reasonably good as well?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1038236,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-10-05T16:58:55.673000",
          "content": "<p>I've found that the data is very stable. Taking a 10K subset of the chopped validation set gives me a slightly lower LB than CV but the preferences are always consistent: better CV =&gt; better LB in every test I've run so far.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1043404,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-10-09T00:24:28.400000",
          "content": "<p>May be a good idea to cache rendered images at least for the validation set, should make it quite a bit faster.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1037359,
      "author_name": "A_Elsheikh",
      "author_url": "",
      "post_date": "2020-10-04T22:39:06.810000",
      "content": "<p>Thanks for sharing. Same conclusion on my side. <br>\nPublic LB <code>17.449</code> and local validation (same setting as above) </p>\n<pre><code>neg_multi_log_likelihood 17.639541738821357\ntime_displace [0.043765   0.06239079 0.08461178 0.1045251  0.12230116 0.13607619\n 0.15364686 0.17026302 0.18461843 0.20065285 0.21922876 0.23798571\n 0.25381047 0.26798241 0.28671686 0.30309549 0.32077815 0.33500339\n 0.34464988 0.35790864 0.37138817 0.38228462 0.39201591 0.40421648\n 0.41869074 0.4339096  0.44838353 0.45712691 0.46366225 0.47801506\n 0.49306092 0.4955435  0.50572086 0.51511343 0.52933362 0.53734122\n 0.55293909 0.56172705 0.57153331 0.58700482 0.59916493 0.61534454\n 0.62890407 0.64430313 0.66137071 0.67384805 0.68635556 0.70236678\n 0.71957936 0.73788754]\n</code></pre>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1041857,
      "author_name": "The Brown Iceman",
      "author_url": "",
      "post_date": "2020-10-08T00:11:38.053000",
      "content": "<p>If There is someone looking for reference to predict using Validate.Zarr and chopped data. You can refer this Notebook. The Notebook also saves gt.csv from temp file. Big thanks to <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and <a href=\"https://www.kaggle.com/thomasbrandon\" target=\"_blank\">@thomasbrandon</a> . The Pipeline Notebook is from based on their work</p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/deepakrajpurushothaman/predict-using-chopped-data-and-know-your-score</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1044241,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-09T16:09:57.523000",
          "content": "<p>It seems that your notebook is private.please make it public.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1053447,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-10-19T02:37:22.850000",
          "content": "<p>Thanks for mention <a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a> :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1079312,
      "author_name": "Bessenyei Szilárd",
      "author_url": "",
      "post_date": "2020-11-15T22:24:03.230000",
      "content": "<p>Check my notebook, I wrote a short script, which you can verify your model LB score. It takes only 5000k sample from the chopped dataset, and it has around +-0.3 accuracy.<br>\n<a href=\"https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min\" target=\"_blank\">https://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1079545,
          "author_name": "Gzl0506",
          "author_url": "",
          "post_date": "2020-11-16T07:21:21.303000",
          "content": "<p>It's looks fine! I don't know why somebody vote down</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1074981,
      "author_name": "豆柴金鯱",
      "author_url": "",
      "post_date": "2020-11-11T09:36:51.170000",
      "content": "<p>Hi folks, <br>\n  Thank for you sharing out the val &amp; PB scores. I'm wondering if your training loss aligns with validation loss? Mine shows a pretty large discrepancy. For example, <code>training loss ~ 20+, validation loss ~ 60+</code>. Currently, I'm iterating through <code>train.zarr</code> for training <code>1000 steps x batch_size=128</code> and use <code>chopp_dataset of full validate.zarr</code> for validation. Any experience in such issue here? </p>\n<p>Thank you! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1075078,
          "author_name": "SuryaJR_Rafl",
          "author_url": "",
          "post_date": "2020-11-11T11:25:07.273000",
          "content": "<p>I think 1000 iterations is a very small dataset to achieve better validation score. If its okay to share, How many epochs have you trained on the 1000 batches? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1075094,
          "author_name": "豆柴金鯱",
          "author_url": "",
          "post_date": "2020-11-11T11:35:49.810000",
          "content": "<p>Hi there,  1000 iterations are for my training, and my training batch size is 128. On the contrary, many folks may use batch_size=32 for 10K iterations…. Maybe that's the difference?</p>\n<p>Sure thing, my latest experiment is still running and it has bee training for 9 epochs… Maybe that's not enough either?</p>\n<p>Thanks!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1075110,
              "author_name": "SuryaJR_Rafl",
              "author_url": "",
              "post_date": "2020-11-11T11:55:07.030000",
              "content": "<p>I think you have overfit the training set. 1000 iterations on batch size of 128 for 9 epochs means your model has kinda memorized the training set and hence the lower training score. It's also the reason your model gives larger error on validation set as it doesn't generalise well. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1075136,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2020-11-11T12:24:57.427000",
              "content": "<p>I'm kind of lost in my experiment results. </p>\n<p>Previously,  my training loss &amp; validate loss moved downwards together and matched each other because I trained &amp; validated on the original train.zarr &amp; validate.zarr dataset.  But when I used <strong>chopped validate</strong> dataset, the discrepancy started to appear… </p>\n<p>And I think maybe your overfitting theory may be correct, but I probably need to restart a new training session and see what happens. </p>\n<p>Thanks for your help!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1075140,
              "author_name": "Mohammed Rizin V K",
              "author_url": "",
              "post_date": "2020-11-11T12:36:41.893000",
              "content": "<p>Sure the chopped dataset clearly correlates with LB and not Full data</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1075143,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2020-11-11T12:40:46.467000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>, sorry I'm kinda lost in you reply. Could you elaborate a litter bit?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1075178,
              "author_name": "Mohammed Rizin V K",
              "author_url": "",
              "post_date": "2020-11-11T13:17:33.307000",
              "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>, sorry I'm kinda lost in you reply. Could you elaborate a litter bit?<br>\n  Sorry to hear that</p>\n</blockquote>\n<p>I even tried to use validate without chopped version. while we do eval after certain training in non chopped datset , its sure the validation correlates with training score and but not LB<br>\nthe case is different in chopped validate dataset. I might not be correlating verwell on training score but the score on LB is better than what you see in Validation score</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1037576,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2020-10-05T06:19:23.607000",
      "content": "<pre><code>neg_multi_log_likelihood 14.785969823499773\ntime_displace [0.0381546  0.05812458 0.07756729 0.09666313 0.11373225 0.13055148\n 0.14453314 0.16217867 0.17796376 0.19276053 0.21084736 0.22936855\n 0.24707832 0.26275199 0.27871745 0.29462635 0.30701878 0.31865779\n 0.32920492 0.33891281 0.34976245 0.36036229 0.36911364 0.37830991\n 0.38650796 0.39458455 0.403425   0.41191151 0.42092284 0.42943664\n 0.43824895 0.44682792 0.45746868 0.46702098 0.47547127 0.48541446\n 0.49275159 0.50045443 0.51071008 0.52505713 0.53472998 0.5485291\n 0.56223968 0.57483904 0.58859108 0.60047564 0.61616895 0.63170392\n 0.64937508 0.66797937]\n</code></pre>\n<p>Public LB:</p>\n<pre><code>14.803\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 1044427,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-10-09T18:47:54.983000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>,</p>\n<p>Thanks for sharing. I am using the same validation scheme with the same results as to agreement with LB. What about your training score/loss? Does this agree with valid/LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1044492,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-09T20:16:42.597000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a> depending on the backbone and the NN structure, the training loss agrees quite well with the validation score and LB score. But that is only true if you are not using too much training data. I would think maybe around 10% and as mentioned above not for all NN structures. </p>\n<p>So, validation is important to get a sense of how good a model is actually doing. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1037450,
      "author_name": "Dmytro Poplavskiy",
      "author_url": "",
      "post_date": "2020-10-05T03:24:06.620000",
      "content": "<p>I can confirm this validation strategy matched LB for me as well, but only on the single sample/submit so far</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1078553,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-11-15T01:24:55.700000",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> Thanks for sharing a lot of useful informations in this competition!<br>\nMay I ask one more maybe silly question?<br>\nWhen I created val dataset by following <a href=\"https://github.com/lyft/l5kit/blob/03eb9e037d23940a134e27c9f124021e18982020/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">this</a> or <a href=\"https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset/comments\" target=\"_blank\">this</a>, it outputs data with all target_availabilities = 0. Then how can I calculate a correct loss without target_availabilities ? Or do you have any idea where I made mistake??</p>\n<p>Thanks in advance, good luck for the competition end!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1078558,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-11-15T01:49:34.777000",
          "content": "<p>I think creating the validation set using create_chopped_dataset generates a mask.npz file which you can use to get agents for prediction</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078562,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-11-15T02:02:10.147000",
          "content": "<p><a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a> <br>\nThanks for help!<br>\nI created the dataset by following official notebook like that.<br>\nI believe it uses the mask created by create_chopped_dataset.</p>\n<pre><code>eval_zarr_path = str(Path(eval_base_path) / Path(dm.require(eval_cfg[\"key\"])).name)\neval_mask_path = str(Path(eval_base_path) / \"mask.npz\")\neval_gt_path = str(Path(eval_base_path) / \"gt.csv\")\n\neval_zarr = ChunkedDataset(eval_zarr_path).open()\neval_mask = np.load(eval_mask_path)[\"arr_0\"]\n# ===== INIT DATASET AND LOAD MASK\neval_dataset = AgentDataset(cfg, eval_zarr, rasterizer, agents_mask=eval_mask)\neval_dataloader = DataLoader(eval_dataset, shuffle=eval_cfg[\"shuffle\"], batch_size=eval_cfg[\"batch_size\"], \n                             num_workers=eval_cfg[\"num_workers\"])\nprint(eval_dataset)\n</code></pre>\n<p>But when I check the data in the dataloader, all of the target_availabilities is zero.</p>\n<pre><code>data = next(iter(eval_dataloader))\nprint(data['target_availabilities'].sum())\n# tensor(0.)\n</code></pre>\n<blockquote>\n  <p>I think creating the validation set using create_chopped_dataset generates a mask.npz file which you can use to get agents for prediction</p>\n</blockquote>\n<p>That means this dataset cannot be directly used for validation? Do we need to map target availability manually??</p>\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078720,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-11-15T08:25:21.140000",
          "content": "<p>True, the created mask.npz must be used. Target availabilities are given through the gt.csv which is also created during the chopping process. This is the same during with the test set. We don't know anything about the target avails there. But it is used for the metric calculation on the kaggle scoreboard. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1079309,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-11-15T22:18:06.657000",
          "content": "<p>Thanks, now I decided to validate separately from my training loop.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1062729,
      "author_name": "ilovepotatoes",
      "author_url": "",
      "post_date": "2020-10-28T06:24:16.183000",
      "content": "<p>I tried implementing the validation from the config file   </p>\n<pre><code>'val_data_loader': {\n  'key': 'scenes/validate.zarr',\n  'batch_size': 16,\n  'shuffle': False,\n  'num_workers': 4\n}\n</code></pre>\n<p>num_frames_to_chop = 100<br>\neval_cfg = cfg[\"val_data_loader\"]<br>\neval_base_path = create_chopped_dataset(dm.require(eval_cfg[\"key\"]), cfg[\"raster_params\"][\"filter_agents_threshold\"], <br>\n                              num_frames_to_chop, cfg[\"model_params\"][\"future_num_frames\"], MIN_FUTURE_STEPS)</p>\n<p>but it gave me an OSError: [Errno 30] Read-only file system: '/kaggle/input/lyft-motion-prediction-autonomous-vehicles/scenes/sample_chopped_100'<br>\nas Kaggle doesn't provide access to write files in the input directory but how to change the /kaggle/input to /kaggle/working ?</p>\n<p>What am I doing wrong ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1062741,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-10-28T06:34:50.323000",
          "content": "<p>You can try using this book. This will definitely help. <a href=\"https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset\" target=\"_blank\">https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1062745,
          "author_name": "ilovepotatoes",
          "author_url": "",
          "post_date": "2020-10-28T06:37:11.553000",
          "content": "<p>Thank You, this is what I was looking for.<br>\nOne more thing, do we have to use 'sample.zarr' or 'validate.zarr' for the validation ?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1062754,
              "author_name": "SuryaJR_Rafl",
              "author_url": "",
              "post_date": "2020-10-28T06:46:57.193000",
              "content": "<p>We have to used Validate.zarr, sample.zarr is just for reference.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1047626,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-10-12T19:12:48.297000",
      "content": "<p>I'm interested how you all are preparing training data now. With the targets in meters or pixels and how do you get validation to align with training numbers now? Right now I have configured my training loop to operate on meters by simply dividing the output from the dataloader by the pixel size and then multiplying by the pixel size at inference time. Training and validation no longer seem to perfectly line up but it performs the same on validation as when I was using the old dataloader and transforming the data myself. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1047659,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-12T20:20:09.093000",
          "content": "<p>I stayed at meters. No further transformation done, yet. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1044086,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-09T14:08:30.517000",
      "content": "<p>Can you give me the notebook of validation if it is available in public .i am just going through the discussion but cannot understand where you implemented</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1044517,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-09T20:43:16.270000",
          "content": "<p>Do you mean <a href=\"https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb\" target=\"_blank\">https://github.com/lyft/l5kit/blob/master/examples/agent_motion_prediction/agent_motion_prediction.ipynb</a> (Under the point Evaluation)</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1045298,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-10T13:56:24.263000",
          "content": "<p>Thank you for this comment</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1041807,
      "author_name": "The Brown Iceman",
      "author_url": "",
      "post_date": "2020-10-07T23:20:50.913000",
      "content": "<p>With new L5kit, I also confirm its true. Only difference is method. I guess you and some people were putting other values for Min_Future_Frame like 10. Mine remains at default 1.   The problem with train loss and eval loss or test loss still remians. Thats Huge difference. </p>\n<p>Method:<br>\ncreate_chopped_dataset on validation.zarr<br>\nnum_frames_to_chop = 100<br>\nMIN_FUTURE_STEPS = 10<br>\nMIN_FUTURE_FRAME=1<br>\n-&gt; 94694 agents to predict</p>\n<p>Train: Trin_Full.zarr for 1.5M samples<br>\nl5kit version: 1.1.0<br>\nLB Score: 279.xx<br>\nEval Score :<br>\nneg_multi_log_likelihood 267.17160372981397<br>\ntime_displace [0.11229513 0.16690064 0.22574588 0.28469734 0.34264851 0.40680175<br>\n 0.46570777 0.52782559 0.58903438 0.64639612 0.70205706 0.75062746<br>\n 0.80532094 0.85498418 0.90345951 0.94995913 0.9971724  1.04381954<br>\n 1.08859232 1.1342783  1.17871579 1.21726209 1.26382939 1.31105303<br>\n 1.3562235  1.39718102 1.44048208 1.48720213 1.52771844 1.56625857<br>\n 1.61256805 1.65881987 1.70607001 1.74212776 1.77483601 1.81985783<br>\n 1.86126298 1.90600642 1.94457169 1.99220685 2.04489769 2.07543546<br>\n 2.11432232 2.15322523 2.19668563 2.23679864 2.26960409 2.30307122<br>\n 2.36240464 2.39407803]</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1037824,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-10-05T11:11:25.973000",
      "content": "<p>I just iter over a valid_dataloader for 1500 Frames every 5000 Training Frames:</p>\n<p>Validation (NLLLoss): 39.49<br>\nPublic-LB: 44.66</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1038214,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-10-05T16:45:45.973000",
          "content": "<p>You need to use <code>create_chopped_dataset</code> as the dataset is serially correlated in time</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1038317,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-05T18:14:20.723000",
          "content": "<p>and it will save you a lot of time ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1038493,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-05T20:33:55.077000",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> saving time is a good argument, but I wonder how it worked so well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1037670,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-10-05T08:19:48.690000",
      "content": "<p>Same for me. Had a similar experience with correlation between submits and validating as they show in the example notebook with the chopped datasets</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1037367,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2020-10-04T23:15:24.447000",
      "content": "<p>Wow, solid cross validation :) CV ~ Public LB. Thanks for sharing <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/ggrizzly\" target=\"_blank\">@ggrizzly</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1068502,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-03T13:28:48.583000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1037244": "I would like to get a little more light onto the validation strategy that everyone is using. \n\nThe hosts are suggesting to use a chopped dataset for validation, as this was also used to create the test dataset.\n\nThere have been some discussions of how the parameters have been chosen in this process. Maybe we can resolve that, when comparing our past results, so let me start.\n\nMethod: \n```\ncreate_chopped_dataset on validation.zarr\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\n-> 94694 agents to predict\n```\n\nValidation:\n```\nneg_multi_log_likelihood 15.90226142426006\ntime_displace [0.03825686 0.05979496 0.07908053 0.09720777 0.11376195 0.13041205\n 0.14533137 0.16254424 0.17893563 0.19681822 0.21209035 0.22800589\n 0.24817092 0.2667516  0.28417274 0.29907279 0.3146695  0.32590776\n 0.33657942 0.34686728 0.35946204 0.37245056 0.38405101 0.39074047\n 0.40032023 0.4101571  0.42140985 0.43188392 0.43951849 0.44799483\n 0.45873317 0.46933649 0.47870115 0.48539507 0.49432693 0.50711085\n 0.51884592 0.52831929 0.53900899 0.55048483 0.56373485 0.57716685\n 0.59198183 0.60644297 0.62096873 0.63373832 0.65085737 0.66873676\n 0.68705682 0.70606301]\n```\n\nPublic LB: \n```\n15.768\n```\n\nI think this validation strategy agrees well with the public LB score and probably will translate to the private LB. To further validate that hypothesis, please post your results ;)",
    "1038000": "Provided validation set nll and lb score is very close for me as well. But full validation takes a good amount of time. Does anybody know a subset of the validation set that is reasonably good as well?",
    "1037359": "Thanks for sharing. Same conclusion on my side. \nPublic LB `17.449` and local validation (same setting as above) \n```\nneg_multi_log_likelihood 17.639541738821357\ntime_displace [0.043765   0.06239079 0.08461178 0.1045251  0.12230116 0.13607619\n 0.15364686 0.17026302 0.18461843 0.20065285 0.21922876 0.23798571\n 0.25381047 0.26798241 0.28671686 0.30309549 0.32077815 0.33500339\n 0.34464988 0.35790864 0.37138817 0.38228462 0.39201591 0.40421648\n 0.41869074 0.4339096  0.44838353 0.45712691 0.46366225 0.47801506\n 0.49306092 0.4955435  0.50572086 0.51511343 0.52933362 0.53734122\n 0.55293909 0.56172705 0.57153331 0.58700482 0.59916493 0.61534454\n 0.62890407 0.64430313 0.66137071 0.67384805 0.68635556 0.70236678\n 0.71957936 0.73788754]\n```",
    "1041857": "If There is someone looking for reference to predict using Validate.Zarr and chopped data. You can refer this Notebook. The Notebook also saves gt.csv from temp file. Big thanks to @corochann and @thomasbrandon . The Pipeline Notebook is from based on their work\n\n[https://www.kaggle.com/deepakrajpurushothaman/predict-using-chopped-data-and-know-your-score](url)",
    "1079312": "Check my notebook, I wrote a short script, which you can verify your model LB score. It takes only 5000k sample from the chopped dataset, and it has around +-0.3 accuracy.\nhttps://www.kaggle.com/bessenyeiszilrd/get-lb-score-under-10-min",
    "1074981": "Hi folks, \n  Thank for you sharing out the val & PB scores. I'm wondering if your training loss aligns with validation loss? Mine shows a pretty large discrepancy. For example, `training loss ~ 20+, validation loss ~ 60+`. Currently, I'm iterating through `train.zarr` for training `1000 steps x batch_size=128` and use `chopp_dataset of full validate.zarr` for validation. Any experience in such issue here? \n\n Thank you! ",
    "1037576": "```\nneg_multi_log_likelihood 14.785969823499773\ntime_displace [0.0381546  0.05812458 0.07756729 0.09666313 0.11373225 0.13055148\n 0.14453314 0.16217867 0.17796376 0.19276053 0.21084736 0.22936855\n 0.24707832 0.26275199 0.27871745 0.29462635 0.30701878 0.31865779\n 0.32920492 0.33891281 0.34976245 0.36036229 0.36911364 0.37830991\n 0.38650796 0.39458455 0.403425   0.41191151 0.42092284 0.42943664\n 0.43824895 0.44682792 0.45746868 0.46702098 0.47547127 0.48541446\n 0.49275159 0.50045443 0.51071008 0.52505713 0.53472998 0.5485291\n 0.56223968 0.57483904 0.58859108 0.60047564 0.61616895 0.63170392\n 0.64937508 0.66797937]\n``` \n\nPublic LB:\n```\n14.803\n```",
    "1037450": "I can confirm this validation strategy matched LB for me as well, but only on the single sample/submit so far",
    "1078553": "@ilu000 Thanks for sharing a lot of useful informations in this competition!\nMay I ask one more maybe silly question?\nWhen I created val dataset by following [this](https://github.com/lyft/l5kit/blob/03eb9e037d23940a134e27c9f124021e18982020/examples/agent_motion_prediction/agent_motion_prediction.ipynb) or [this](https://www.kaggle.com/thomasbrandon/l5kit-chopped-dataset/comments), it outputs data with all target_availabilities = 0. Then how can I calculate a correct loss without target_availabilities ? Or do you have any idea where I made mistake??\n\nThanks in advance, good luck for the competition end!\n",
    "1062729": "I tried implementing the validation from the config file   \n\n    'val_data_loader': {\n      'key': 'scenes/validate.zarr',\n      'batch_size': 16,\n      'shuffle': False,\n      'num_workers': 4\n    }\nnum_frames_to_chop = 100\neval_cfg = cfg[\"val_data_loader\"]\neval_base_path = create_chopped_dataset(dm.require(eval_cfg[\"key\"]), cfg[\"raster_params\"][\"filter_agents_threshold\"], \n                              num_frames_to_chop, cfg[\"model_params\"][\"future_num_frames\"], MIN_FUTURE_STEPS)\n\nbut it gave me an OSError: [Errno 30] Read-only file system: '/kaggle/input/lyft-motion-prediction-autonomous-vehicles/scenes/sample_chopped_100'\nas Kaggle doesn't provide access to write files in the input directory but how to change the /kaggle/input to /kaggle/working ?\n\nWhat am I doing wrong ?",
    "1047626": "I'm interested how you all are preparing training data now. With the targets in meters or pixels and how do you get validation to align with training numbers now? Right now I have configured my training loop to operate on meters by simply dividing the output from the dataloader by the pixel size and then multiplying by the pixel size at inference time. Training and validation no longer seem to perfectly line up but it performs the same on validation as when I was using the old dataloader and transforming the data myself. ",
    "1044086": "Can you give me the notebook of validation if it is available in public .i am just going through the discussion but cannot understand where you implemented",
    "1041807": "With new L5kit, I also confirm its true. Only difference is method. I guess you and some people were putting other values for Min_Future_Frame like 10. Mine remains at default 1.   The problem with train loss and eval loss or test loss still remians. Thats Huge difference. \n\nMethod:\ncreate_chopped_dataset on validation.zarr\nnum_frames_to_chop = 100\nMIN_FUTURE_STEPS = 10\nMIN_FUTURE_FRAME=1\n-> 94694 agents to predict\n\nTrain: Trin_Full.zarr for 1.5M samples\nl5kit version: 1.1.0\nLB Score: 279.xx\nEval Score :\nneg_multi_log_likelihood 267.17160372981397\ntime_displace [0.11229513 0.16690064 0.22574588 0.28469734 0.34264851 0.40680175\n 0.46570777 0.52782559 0.58903438 0.64639612 0.70205706 0.75062746\n 0.80532094 0.85498418 0.90345951 0.94995913 0.9971724  1.04381954\n 1.08859232 1.1342783  1.17871579 1.21726209 1.26382939 1.31105303\n 1.3562235  1.39718102 1.44048208 1.48720213 1.52771844 1.56625857\n 1.61256805 1.65881987 1.70607001 1.74212776 1.77483601 1.81985783\n 1.86126298 1.90600642 1.94457169 1.99220685 2.04489769 2.07543546\n 2.11432232 2.15322523 2.19668563 2.23679864 2.26960409 2.30307122\n 2.36240464 2.39407803]",
    "1037824": "I just iter over a valid_dataloader for 1500 Frames every 5000 Training Frames:\n\nValidation (NLLLoss): 39.49\nPublic-LB: 44.66",
    "1037670": "Same for me. Had a similar experience with correlation between submits and validating as they show in the example notebook with the chopped datasets",
    "1037367": "Wow, solid cross validation :) CV ~ Public LB. Thanks for sharing @ilu000 @ggrizzly ",
    "1068502": ""
  }
}