{
  "id": 198700,
  "title": "Why it takes more than 2 hours for simply reading all test timestamps and agent ids?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/198700",
  "author_name": "",
  "post_date": "2020-11-22T15:06:14.814713600Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm new in Kaggle and for this competition as well. Last night, I tried to make my first submission, but it will take about 12 hours to finish all calculations, then it stopped at 9 hours. So, I'm still trying to make my first submission. </p>\n<p>I made a simple test to check why it takes so long to complete. I'm not using CNN but read image information as model inputs. When I make a very simple test to just read all test dataset's timestamps and track_ids, it will still take 2 hours without doing other computations. This doesn't make much sense to me but I don't know why. Could you guys help and check it out? Any suggestions to improve the efficiency so that I can make my first submission? </p>\n<p>I also tried to use multi processing (just learned last night how to do it). But it doesn't help to reduce the time. </p>\n<p><a href=\"https://www.kaggle.com/noahxi/lyft-question\" target=\"_blank\">https://www.kaggle.com/noahxi/lyft-question</a></p>",
  "messages": [
    {
      "id": "1087274",
      "postDate": "11/22/2020 15:06:14",
      "content": "<p>I'm new in Kaggle and for this competition as well. Last night, I tried to make my first submission, but it will take about 12 hours to finish all calculations, then it stopped at 9 hours. So, I'm still trying to make my first submission. </p>\n<p>I made a simple test to check why it takes so long to complete. I'm not using CNN but read image information as model inputs. When I make a very simple test to just read all test dataset's timestamps and track_ids, it will still take 2 hours without doing other computations. This doesn't make much sense to me but I don't know why. Could you guys help and check it out? Any suggestions to improve the efficiency so that I can make my first submission? </p>\n<p>I also tried to use multi processing (just learned last night how to do it). But it doesn't help to reduce the time. </p>\n<p><a href=\"https://www.kaggle.com/noahxi/lyft-question\" target=\"_blank\">https://www.kaggle.com/noahxi/lyft-question</a></p>",
      "rawMarkdown": "I'm new in Kaggle and for this competition as well. Last night, I tried to make my first submission, but it will take about 12 hours to finish all calculations, then it stopped at 9 hours. So, I'm still trying to make my first submission. \n\nI made a simple test to check why it takes so long to complete. I'm not using CNN but read image information as model inputs. When I make a very simple test to just read all test dataset's timestamps and track_ids, it will still take 2 hours without doing other computations. This doesn't make much sense to me but I don't know why. Could you guys help and check it out? Any suggestions to improve the efficiency so that I can make my first submission? \n\nI also tried to use multi processing (just learned last night how to do it). But it doesn't help to reduce the time. \n\nhttps://www.kaggle.com/noahxi/lyft-question",
      "votes": null
    },
    {
      "id": "1087281",
      "postDate": "11/22/2020 15:12:09",
      "content": "<p>Any comments are welcome</p>",
      "rawMarkdown": "Any comments are welcome",
      "votes": null
    },
    {
      "id": "1087498",
      "postDate": "11/22/2020 19:26:01",
      "content": "<p>Looping over all 75k samples of test with the l5kit can take around 2h for reasonable raster sizes in kaggle kernels iirc (doing it offline for a while now). The speed highly depends on your rasterizer. Maybe you are using very high resolution rasters?</p>",
      "rawMarkdown": "Looping over all 75k samples of test with the l5kit can take around 2h for reasonable raster sizes in kaggle kernels iirc (doing it offline for a while now). The speed highly depends on your rasterizer. Maybe you are using very high resolution rasters?",
      "votes": null
    },
    {
      "id": "1087518",
      "postDate": "11/22/2020 19:57:38",
      "content": "<p>This is because the l5kit (this is from Lyft) is very slow on rendering the dataset. Note the lyft dataset don't have the dataset ready in the final form on the disk. It always try to compute the images on the fly for each sample when you call to get any element of the AgentDataset. (even if you are not using the image)<br>\nYou should do the scoring offline (which can reduce the inference time to lower than 10mins) and only upload the submission csv as private dataset to kaggle, then read it from the dataset in your submission notebook. This will reduce the submission notebook runtime to &lt;10s.<br>\nThe submission notebook only need to have one line like</p>\n<pre><code>!cp ../input/your_kaggle_private_dataset/your_submision.csv submission.csv\n</code></pre>",
      "rawMarkdown": "This is because the l5kit (this is from Lyft) is very slow on rendering the dataset. Note the lyft dataset don't have the dataset ready in the final form on the disk. It always try to compute the images on the fly for each sample when you call to get any element of the AgentDataset. (even if you are not using the image)\nYou should do the scoring offline (which can reduce the inference time to lower than 10mins) and only upload the submission csv as private dataset to kaggle, then read it from the dataset in your submission notebook. This will reduce the submission notebook runtime to <10s.\nThe submission notebook only need to have one line like\n```\n!cp ../input/your_kaggle_private_dataset/your_submision.csv submission.csv\n```",
      "votes": null
    },
    {
      "id": "1088856",
      "postDate": "11/24/2020 01:58:05",
      "content": "<p>200*200. When I reduce the size, it's better but still slow. Louis's comments make sense. Thank you guys!</p>",
      "rawMarkdown": "200*200. When I reduce the size, it's better but still slow. Louis's comments make sense. Thank you guys!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1087281,
      "author_name": "noahxi",
      "author_url": "",
      "post_date": "11/22/2020 15:12:09",
      "content": "<p>Any comments are welcome</p>",
      "votes": null,
      "replies": [
        {
          "id": 1087498,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "11/22/2020 19:26:01",
          "content": "<p>Looping over all 75k samples of test with the l5kit can take around 2h for reasonable raster sizes in kaggle kernels iirc (doing it offline for a while now). The speed highly depends on your rasterizer. Maybe you are using very high resolution rasters?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088856,
          "author_name": "noahxi",
          "author_url": "",
          "post_date": "11/24/2020 01:58:05",
          "content": "<p>200*200. When I reduce the size, it's better but still slow. Louis's comments make sense. Thank you guys!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1087518,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/22/2020 19:57:38",
      "content": "<p>This is because the l5kit (this is from Lyft) is very slow on rendering the dataset. Note the lyft dataset don't have the dataset ready in the final form on the disk. It always try to compute the images on the fly for each sample when you call to get any element of the AgentDataset. (even if you are not using the image)<br>\nYou should do the scoring offline (which can reduce the inference time to lower than 10mins) and only upload the submission csv as private dataset to kaggle, then read it from the dataset in your submission notebook. This will reduce the submission notebook runtime to &lt;10s.<br>\nThe submission notebook only need to have one line like</p>\n<pre><code>!cp ../input/your_kaggle_private_dataset/your_submision.csv submission.csv\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1087274": "I'm new in Kaggle and for this competition as well. Last night, I tried to make my first submission, but it will take about 12 hours to finish all calculations, then it stopped at 9 hours. So, I'm still trying to make my first submission. \n\nI made a simple test to check why it takes so long to complete. I'm not using CNN but read image information as model inputs. When I make a very simple test to just read all test dataset's timestamps and track_ids, it will still take 2 hours without doing other computations. This doesn't make much sense to me but I don't know why. Could you guys help and check it out? Any suggestions to improve the efficiency so that I can make my first submission? \n\nI also tried to use multi processing (just learned last night how to do it). But it doesn't help to reduce the time. \n\nhttps://www.kaggle.com/noahxi/lyft-question",
    "1087281": "Any comments are welcome",
    "1087498": "Looping over all 75k samples of test with the l5kit can take around 2h for reasonable raster sizes in kaggle kernels iirc (doing it offline for a while now). The speed highly depends on your rasterizer. Maybe you are using very high resolution rasters?",
    "1087518": "This is because the l5kit (this is from Lyft) is very slow on rendering the dataset. Note the lyft dataset don't have the dataset ready in the final form on the disk. It always try to compute the images on the fly for each sample when you call to get any element of the AgentDataset. (even if you are not using the image)\nYou should do the scoring offline (which can reduce the inference time to lower than 10mins) and only upload the submission csv as private dataset to kaggle, then read it from the dataset in your submission notebook. This will reduce the submission notebook runtime to <10s.\nThe submission notebook only need to have one line like\n```\n!cp ../input/your_kaggle_private_dataset/your_submision.csv submission.csv\n```",
    "1088856": "200*200. When I reduce the size, it's better but still slow. Louis's comments make sense. Thank you guys!"
  },
  "source": "meta"
}