{
  "id": 177662,
  "title": "Reduce submission time to ~10-15 minutes",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/177662",
  "author_name": "Peter",
  "post_date": "2020-08-26T20:21:22.975000",
  "votes": 54,
  "comment_count": 15,
  "views": 0,
  "content": "<p>In this competition, we don't have a separate private test set. In the test.zarr we have all the samples, so we can predict locally. For me, it is much faster: instead of ~2 hours, I submitted my last result in less than 15 minutes.</p>\n<ul>\n<li>Predict locally (using the test.zarr and the mask.npz)</li>\n<li>Store the result csv (I used the official method from l5kit)</li>\n<li>Create a private dataset (you only have to do it once)</li>\n<li>Upload the locally saved csv</li>\n<li>Create a submission notebook (see the script below)</li>\n<li>Commit, submit.</li>\n</ul>\n<pre><code>import pandas as pd\n\n# Change these to your dataset/submission.csv\nSUBMISSION_FOLDER = '/kaggle/input/lyft-submissions-private'\nSUBMISSION_FILE = 'best_valid__submission.csv'\n\n\nsubmissions = pd.read_csv(f\"{SUBMISSION_FOLDER}/{SUBMISSION_FILE}\")\nsubmissions.to_csv(\"submission.csv\", index=False)\n</code></pre>",
  "messages": [
    {
      "id": 986830,
      "postDate": "2020-08-26T20:21:22.977Z",
      "content": "<p>In this competition, we don't have a separate private test set. In the test.zarr we have all the samples, so we can predict locally. For me, it is much faster: instead of ~2 hours, I submitted my last result in less than 15 minutes.</p>\n<ul>\n<li>Predict locally (using the test.zarr and the mask.npz)</li>\n<li>Store the result csv (I used the official method from l5kit)</li>\n<li>Create a private dataset (you only have to do it once)</li>\n<li>Upload the locally saved csv</li>\n<li>Create a submission notebook (see the script below)</li>\n<li>Commit, submit.</li>\n</ul>\n<pre><code>import pandas as pd\n\n# Change these to your dataset/submission.csv\nSUBMISSION_FOLDER = '/kaggle/input/lyft-submissions-private'\nSUBMISSION_FILE = 'best_valid__submission.csv'\n\n\nsubmissions = pd.read_csv(f\"{SUBMISSION_FOLDER}/{SUBMISSION_FILE}\")\nsubmissions.to_csv(\"submission.csv\", index=False)\n</code></pre>",
      "rawMarkdown": "In this competition, we don't have a separate private test set. In the test.zarr we have all the samples, so we can predict locally. For me, it is much faster: instead of ~2 hours, I submitted my last result in less than 15 minutes.\n\n- Predict locally (using the test.zarr and the mask.npz)\n- Store the result csv (I used the official method from l5kit)\n- Create a private dataset (you only have to do it once)\n- Upload the locally saved csv\n- Create a submission notebook (see the script below)\n- Commit, submit.\n\n```\nimport pandas as pd\n\n# Change these to your dataset/submission.csv\nSUBMISSION_FOLDER = '/kaggle/input/lyft-submissions-private'\nSUBMISSION_FILE = 'best_valid__submission.csv'\n\n\nsubmissions = pd.read_csv(f\"{SUBMISSION_FOLDER}/{SUBMISSION_FILE}\")\nsubmissions.to_csv(\"submission.csv\", index=False)\n```\n",
      "votes": 53
    },
    {
      "id": 988276,
      "postDate": "2020-08-28T00:04:51.027Z",
      "content": "<p>Thank you for the post <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> , I demonstrated it in the kernel <a href=\"https://www.kaggle.com/corochann/save-your-time-submit-without-kernel-inference\" target=\"_blank\">Save your time, submit without kernel inference</a> using your kernel output, it worked very well. Thanks again!</p>",
      "rawMarkdown": "Thank you for the post @pestipeti , I demonstrated it in the kernel [Save your time, submit without kernel inference](https://www.kaggle.com/corochann/save-your-time-submit-without-kernel-inference) using your kernel output, it worked very well. Thanks again!",
      "votes": 5
    },
    {
      "id": 1046757,
      "postDate": "2020-10-12T01:29:31.073Z",
      "content": "<p>Hmm… If you can create a private dataset which contains the predictions offline, then what is the point of being a Code Competition? Wouldn't it be easier for Kaggle just provide the regular submit csv as prediction?</p>",
      "rawMarkdown": "Hmm... If you can create a private dataset which contains the predictions offline, then what is the point of being a Code Competition? Wouldn't it be easier for Kaggle just provide the regular submit csv as prediction?",
      "votes": 1
    },
    {
      "id": 1074840,
      "postDate": "2020-11-11T06:00:45.007Z",
      "content": "<p>Or simply <br>\n<code>!cp ../input/lyft-submissions-private/best_valid__submission.csv submission.csv</code><br>\ninstead of reading in by pandas since read then write csv don't will introduce precision error.</p>",
      "rawMarkdown": "Or simply \n`!cp ../input/lyft-submissions-private/best_valid__submission.csv submission.csv`\ninstead of reading in by pandas since read then write csv don't will introduce precision error.",
      "votes": 2
    },
    {
      "id": 1025396,
      "postDate": "2020-09-24T14:27:24.753Z",
      "content": "<p>it can be ever faster ~4 min <a href=\"https://www.kaggle.com/lao777/fast-submission-valid-for-public-private-lb\" target=\"_blank\">https://www.kaggle.com/lao777/fast-submission-valid-for-public-private-lb</a></p>",
      "rawMarkdown": "it can be ever faster ~4 min https://www.kaggle.com/lao777/fast-submission-valid-for-public-private-lb",
      "votes": 2
    },
    {
      "id": 1090877,
      "postDate": "2020-11-25T16:30:02.793Z",
      "content": "<p>I am new to Kaggle and I have already stored submission.csv as a private dataset and it's in the input folder. I get an error saying \"your notebook tried to allocate more memory than is available\". Is there any work around it?</p>",
      "rawMarkdown": "I am new to Kaggle and I have already stored submission.csv as a private dataset and it's in the input folder. I get an error saying \"your notebook tried to allocate more memory than is available\". Is there any work around it?"
    },
    {
      "id": 1074971,
      "postDate": "2020-11-11T09:23:47.257Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>, thanks for pointing out this technique. I'm wondering if we need to generate a local submission file from all test.zarr or the <code>chopped_dataset of test.zarr</code>? I think we should do the first option because Kaggle kernel will select the \"100th frames\" and show a metric score? Am I right?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hi @pestipeti, thanks for pointing out this technique. I'm wondering if we need to generate a local submission file from all test.zarr or the `chopped_dataset of test.zarr`? I think we should do the first option because Kaggle kernel will select the \"100th frames\" and show a metric score? Am I right?\n\nThank you!",
      "replies": [
        {
          "id": 1075026,
          "postDate": "2020-11-11T10:30:38.597Z",
          "content": "<p>You should predict the samples from test.zarr, it is already chopped.</p>",
          "rawMarkdown": "You should predict the samples from test.zarr, it is already chopped."
        },
        {
          "id": 1075101,
          "postDate": "2020-11-11T11:40:41.287Z",
          "content": "<p>Hi Peter, thank you for your confirmation. Now I get it! Thank you!</p>",
          "rawMarkdown": "Hi Peter, thank you for your confirmation. Now I get it! Thank you!"
        }
      ]
    },
    {
      "id": 990750,
      "postDate": "2020-08-29T20:19:29.740Z",
      "content": "<p>Good share</p>",
      "rawMarkdown": "Good share"
    },
    {
      "id": 990158,
      "postDate": "2020-08-29T11:33:55.297Z",
      "content": "<p>Thanks you it's very useful <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  ,It's saves time …</p>",
      "rawMarkdown": "Thanks you it's very useful @pestipeti  ,It's saves time ..."
    },
    {
      "id": 988091,
      "postDate": "2020-08-27T19:11:05.890Z",
      "content": "<p>Very useful indeed. I have noticed that that was the case for few other competitions as well, i.e. you can predict offline by building a private dataset. Thanks for sharing!</p>",
      "rawMarkdown": "Very useful indeed. I have noticed that that was the case for few other competitions as well, i.e. you can predict offline by building a private dataset. Thanks for sharing!"
    },
    {
      "id": 987391,
      "postDate": "2020-08-27T08:12:12.360Z",
      "content": "<p>This is a great idea! The test dataset is SO big it takes a ridiculous amount of time to create a submission from a kernel that can only run on CPU (since there's a bug in the l5kit script you have to include that means you can't run in GPU). </p>",
      "rawMarkdown": "This is a great idea! The test dataset is SO big it takes a ridiculous amount of time to create a submission from a kernel that can only run on CPU (since there's a bug in the l5kit script you have to include that means you can't run in GPU). ",
      "replies": [
        {
          "id": 987399,
          "postDate": "2020-08-27T08:20:40.813Z",
          "content": "<p>Even if you are able to use the GPU the process is slow, because of the rasterization. The speedup on my local machine is because I have 12 CPU cores (instead of the 4 on the kernels). The rasterization is faster, but still, it is the bottleneck. </p>",
          "rawMarkdown": "Even if you are able to use the GPU the process is slow, because of the rasterization. The speedup on my local machine is because I have 12 CPU cores (instead of the 4 on the kernels). The rasterization is faster, but still, it is the bottleneck. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 1052824,
      "postDate": "2020-10-18T10:18:38.723Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1047071,
      "postDate": "2020-10-12T08:29:54.567Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 988276,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-08-28T00:04:51.027000",
      "content": "<p>Thank you for the post <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> , I demonstrated it in the kernel <a href=\"https://www.kaggle.com/corochann/save-your-time-submit-without-kernel-inference\" target=\"_blank\">Save your time, submit without kernel inference</a> using your kernel output, it worked very well. Thanks again!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1046757,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2020-10-12T01:29:31.073000",
      "content": "<p>Hmm… If you can create a private dataset which contains the predictions offline, then what is the point of being a Code Competition? Wouldn't it be easier for Kaggle just provide the regular submit csv as prediction?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1074840,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2020-11-11T06:00:45.007000",
      "content": "<p>Or simply <br>\n<code>!cp ../input/lyft-submissions-private/best_valid__submission.csv submission.csv</code><br>\ninstead of reading in by pandas since read then write csv don't will introduce precision error.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1025396,
      "author_name": "laol",
      "author_url": "",
      "post_date": "2020-09-24T14:27:24.753000",
      "content": "<p>it can be ever faster ~4 min <a href=\"https://www.kaggle.com/lao777/fast-submission-valid-for-public-private-lb\" target=\"_blank\">https://www.kaggle.com/lao777/fast-submission-valid-for-public-private-lb</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1090877,
      "author_name": "Praveen Phatate",
      "author_url": "",
      "post_date": "2020-11-25T16:30:02.793000",
      "content": "<p>I am new to Kaggle and I have already stored submission.csv as a private dataset and it's in the input folder. I get an error saying \"your notebook tried to allocate more memory than is available\". Is there any work around it?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1074971,
      "author_name": "豆柴金鯱",
      "author_url": "",
      "post_date": "2020-11-11T09:23:47.257000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>, thanks for pointing out this technique. I'm wondering if we need to generate a local submission file from all test.zarr or the <code>chopped_dataset of test.zarr</code>? I think we should do the first option because Kaggle kernel will select the \"100th frames\" and show a metric score? Am I right?</p>\n<p>Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1075026,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-11-11T10:30:38.597000",
          "content": "<p>You should predict the samples from test.zarr, it is already chopped.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1075101,
          "author_name": "豆柴金鯱",
          "author_url": "",
          "post_date": "2020-11-11T11:40:41.287000",
          "content": "<p>Hi Peter, thank you for your confirmation. Now I get it! Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 990750,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-08-29T20:19:29.740000",
      "content": "<p>Good share</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 990158,
      "author_name": "Surekha Ramireddy",
      "author_url": "",
      "post_date": "2020-08-29T11:33:55.297000",
      "content": "<p>Thanks you it's very useful <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  ,It's saves time …</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 988091,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-08-27T19:11:05.890000",
      "content": "<p>Very useful indeed. I have noticed that that was the case for few other competitions as well, i.e. you can predict offline by building a private dataset. Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 987391,
      "author_name": "Ian Ormesher",
      "author_url": "",
      "post_date": "2020-08-27T08:12:12.360000",
      "content": "<p>This is a great idea! The test dataset is SO big it takes a ridiculous amount of time to create a submission from a kernel that can only run on CPU (since there's a bug in the l5kit script you have to include that means you can't run in GPU). </p>",
      "votes": 0,
      "replies": [
        {
          "id": 987399,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-08-27T08:20:40.813000",
          "content": "<p>Even if you are able to use the GPU the process is slow, because of the rasterization. The speedup on my local machine is because I have 12 CPU cores (instead of the 4 on the kernels). The rasterization is faster, but still, it is the bottleneck. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1052824,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-18T10:18:38.723000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1047071,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-12T08:29:54.567000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "986830": "In this competition, we don't have a separate private test set. In the test.zarr we have all the samples, so we can predict locally. For me, it is much faster: instead of ~2 hours, I submitted my last result in less than 15 minutes.\n\n- Predict locally (using the test.zarr and the mask.npz)\n- Store the result csv (I used the official method from l5kit)\n- Create a private dataset (you only have to do it once)\n- Upload the locally saved csv\n- Create a submission notebook (see the script below)\n- Commit, submit.\n\n```\nimport pandas as pd\n\n# Change these to your dataset/submission.csv\nSUBMISSION_FOLDER = '/kaggle/input/lyft-submissions-private'\nSUBMISSION_FILE = 'best_valid__submission.csv'\n\n\nsubmissions = pd.read_csv(f\"{SUBMISSION_FOLDER}/{SUBMISSION_FILE}\")\nsubmissions.to_csv(\"submission.csv\", index=False)\n```\n",
    "988276": "Thank you for the post @pestipeti , I demonstrated it in the kernel [Save your time, submit without kernel inference](https://www.kaggle.com/corochann/save-your-time-submit-without-kernel-inference) using your kernel output, it worked very well. Thanks again!",
    "1046757": "Hmm... If you can create a private dataset which contains the predictions offline, then what is the point of being a Code Competition? Wouldn't it be easier for Kaggle just provide the regular submit csv as prediction?",
    "1074840": "Or simply \n`!cp ../input/lyft-submissions-private/best_valid__submission.csv submission.csv`\ninstead of reading in by pandas since read then write csv don't will introduce precision error.",
    "1025396": "it can be ever faster ~4 min https://www.kaggle.com/lao777/fast-submission-valid-for-public-private-lb",
    "1090877": "I am new to Kaggle and I have already stored submission.csv as a private dataset and it's in the input folder. I get an error saying \"your notebook tried to allocate more memory than is available\". Is there any work around it?",
    "1074971": "Hi @pestipeti, thanks for pointing out this technique. I'm wondering if we need to generate a local submission file from all test.zarr or the `chopped_dataset of test.zarr`? I think we should do the first option because Kaggle kernel will select the \"100th frames\" and show a metric score? Am I right?\n\nThank you!",
    "990750": "Good share",
    "990158": "Thanks you it's very useful @pestipeti  ,It's saves time ...",
    "988091": "Very useful indeed. I have noticed that that was the case for few other competitions as well, i.e. you can predict offline by building a private dataset. Thanks for sharing!",
    "987391": "This is a great idea! The test dataset is SO big it takes a ridiculous amount of time to create a submission from a kernel that can only run on CPU (since there's a bug in the l5kit script you have to include that means you can't run in GPU). ",
    "1052824": "",
    "1047071": ""
  }
}