{
  "id": 177637,
  "title": "How to improve training time",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/177637",
  "author_name": "Peter",
  "post_date": "2020-08-26T18:12:03.127000",
  "votes": 60,
  "comment_count": 29,
  "views": 0,
  "content": "<p>In this competition, we have to generate the image for our models. Unfortunately, the rasterization is a very slow process (if we do it on the CPU).</p>\n<p>My idea is to create a preprocessed image dataset. Generate a million samples (images and additional data). Train the model without rasterizing new images. We can fine-tune later with new samples (with rasterization)</p>\n<h2>Experiments:</h2>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Size</th>\n<th>History</th>\n<th># of samples</th>\n<th>Storage size</th>\n<th>Training time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Rasterizer</td>\n<td>300</td>\n<td>10</td>\n<td>32000</td>\n<td>0</td>\n<td>4:48</td>\n<td>3.47</td>\n</tr>\n<tr>\n<td>Rasterizer</td>\n<td>300</td>\n<td>5</td>\n<td>32000</td>\n<td>0</td>\n<td>3:46</td>\n<td>4.42</td>\n</tr>\n<tr>\n<td>Pre-generated data</td>\n<td>300</td>\n<td>10</td>\n<td>32000</td>\n<td>1.9 Gb</td>\n<td>3:26</td>\n<td>4.85</td>\n</tr>\n<tr>\n<td>Pre-generated data</td>\n<td>300</td>\n<td>5</td>\n<td>32000</td>\n<td>1.8 Gb</td>\n<td>2:56</td>\n<td>5.68</td>\n</tr>\n</tbody>\n</table>\n<h2>Advantages</h2>\n<ul>\n<li>We can speed up the training process</li>\n<li>We can save the validation and the test data as well</li>\n</ul>\n<h2>Disadvantages</h2>\n<ul>\n<li>Without compression, it takes too much space; with compression, the data loading is slower.</li>\n<li>We have to re-generate the data if we change the config. (image size, raster size or the number of historical frames)</li>\n</ul>\n<h2>Code: Generate and save</h2>\n<pre><code>... imports ...\n... lyft configuration ...\n\ndef save_sample(i):\n    idx = random.randint(0, len(dataset))\n\n    # 300px, 0.5 raster size, 5 historical frames\n    obj_save(dataset[idx], f'sample_{i}', './cache/pre_300px__0_5__5')\n\ndef obj_save(obj, name, dir_cache):\n    with bz2.BZ2File(f'{dir_cache}/{name}.pbz', 'wb') as f:\n        pickle.dump(obj, f)\n\ndm = LocalDataManager()\ndataset_path = dm.require(cfg[\"train_data_loader\"][\"key\"])\nzarr_dataset = ChunkedDataset(dataset_path)\nzarr_dataset.open()\n\nrast = build_rasterizer(cfg, dm)\ndataset = AgentDataset(cfg, zarr_dataset, rast)\n\nwith Pool(processes=4) as p:\n    max_ = 32000\n    with tqdm(total=max_) as pbar:\n        for i, _ in enumerate(p.imap_unordered(save_sample, range(0, max_))):\n            pbar.update()\n</code></pre>\n<h2>Code: Data loader</h2>\n<pre><code>class LyftImageDataset(Dataset):\n\n    def __init__(self, data_folder):\n        super().__init__()\n        self.data_folder = data_folder\n        self.files = []\n\n        for filename in os.listdir(self.data_folder):\n            if filename.endswith(\".pbz\"):\n                self.files.append(filename)\n\n        print(len(self.files))\n        print(self.files[0])\n\n    def __getitem__(self, index: int):\n        return self.obj_load(self.files[index])\n\n    def obj_load(self, name):\n        with bz2.BZ2File(f'{self.data_folder}/{name}', 'rb') as f:\n            return pickle.load(f)\n\n    def __len__(self):\n        return len(self.files)\n\n...\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = LyftImageDataset('./cache/pre_300px__0_5__10')\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=train_cfg[\"shuffle\"],\n                              batch_size=train_cfg[\"batch_size\"],\n                              num_workers=train_cfg[\"num_workers\"])\n</code></pre>",
  "messages": [
    {
      "id": 986716,
      "postDate": "2020-08-26T18:12:03.127Z",
      "content": "<p>In this competition, we have to generate the image for our models. Unfortunately, the rasterization is a very slow process (if we do it on the CPU).</p>\n<p>My idea is to create a preprocessed image dataset. Generate a million samples (images and additional data). Train the model without rasterizing new images. We can fine-tune later with new samples (with rasterization)</p>\n<h2>Experiments:</h2>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Size</th>\n<th>History</th>\n<th># of samples</th>\n<th>Storage size</th>\n<th>Training time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Rasterizer</td>\n<td>300</td>\n<td>10</td>\n<td>32000</td>\n<td>0</td>\n<td>4:48</td>\n<td>3.47</td>\n</tr>\n<tr>\n<td>Rasterizer</td>\n<td>300</td>\n<td>5</td>\n<td>32000</td>\n<td>0</td>\n<td>3:46</td>\n<td>4.42</td>\n</tr>\n<tr>\n<td>Pre-generated data</td>\n<td>300</td>\n<td>10</td>\n<td>32000</td>\n<td>1.9 Gb</td>\n<td>3:26</td>\n<td>4.85</td>\n</tr>\n<tr>\n<td>Pre-generated data</td>\n<td>300</td>\n<td>5</td>\n<td>32000</td>\n<td>1.8 Gb</td>\n<td>2:56</td>\n<td>5.68</td>\n</tr>\n</tbody>\n</table>\n<h2>Advantages</h2>\n<ul>\n<li>We can speed up the training process</li>\n<li>We can save the validation and the test data as well</li>\n</ul>\n<h2>Disadvantages</h2>\n<ul>\n<li>Without compression, it takes too much space; with compression, the data loading is slower.</li>\n<li>We have to re-generate the data if we change the config. (image size, raster size or the number of historical frames)</li>\n</ul>\n<h2>Code: Generate and save</h2>\n<pre><code>... imports ...\n... lyft configuration ...\n\ndef save_sample(i):\n    idx = random.randint(0, len(dataset))\n\n    # 300px, 0.5 raster size, 5 historical frames\n    obj_save(dataset[idx], f'sample_{i}', './cache/pre_300px__0_5__5')\n\ndef obj_save(obj, name, dir_cache):\n    with bz2.BZ2File(f'{dir_cache}/{name}.pbz', 'wb') as f:\n        pickle.dump(obj, f)\n\ndm = LocalDataManager()\ndataset_path = dm.require(cfg[\"train_data_loader\"][\"key\"])\nzarr_dataset = ChunkedDataset(dataset_path)\nzarr_dataset.open()\n\nrast = build_rasterizer(cfg, dm)\ndataset = AgentDataset(cfg, zarr_dataset, rast)\n\nwith Pool(processes=4) as p:\n    max_ = 32000\n    with tqdm(total=max_) as pbar:\n        for i, _ in enumerate(p.imap_unordered(save_sample, range(0, max_))):\n            pbar.update()\n</code></pre>\n<h2>Code: Data loader</h2>\n<pre><code>class LyftImageDataset(Dataset):\n\n    def __init__(self, data_folder):\n        super().__init__()\n        self.data_folder = data_folder\n        self.files = []\n\n        for filename in os.listdir(self.data_folder):\n            if filename.endswith(\".pbz\"):\n                self.files.append(filename)\n\n        print(len(self.files))\n        print(self.files[0])\n\n    def __getitem__(self, index: int):\n        return self.obj_load(self.files[index])\n\n    def obj_load(self, name):\n        with bz2.BZ2File(f'{self.data_folder}/{name}', 'rb') as f:\n            return pickle.load(f)\n\n    def __len__(self):\n        return len(self.files)\n\n...\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = LyftImageDataset('./cache/pre_300px__0_5__10')\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=train_cfg[\"shuffle\"],\n                              batch_size=train_cfg[\"batch_size\"],\n                              num_workers=train_cfg[\"num_workers\"])\n</code></pre>",
      "rawMarkdown": "In this competition, we have to generate the image for our models. Unfortunately, the rasterization is a very slow process (if we do it on the CPU).\n\nMy idea is to create a preprocessed image dataset. Generate a million samples (images and additional data). Train the model without rasterizing new images. We can fine-tune later with new samples (with rasterization)\n\n## Experiments:\n| Method             | Size | History | \\# of samples | Storage size | Training time | it/sec |\n| ------------------ | ---- | ------- | ------------- | ------------ | ------------- | ------ |\n| Rasterizer         | 300  | 10      | 32000         | 0            | 4:48          | 3.47   |\n| Rasterizer         | 300  | 5       | 32000         | 0            | 3:46          | 4.42   |\n| Pre-generated data | 300  | 10      | 32000         | 1.9 Gb       | 3:26          | 4.85   |\n| Pre-generated data | 300  | 5       | 32000         | 1.8 Gb       | 2:56          | 5.68   |\n\n\n## Advantages\n- We can speed up the training process\n- We can save the validation and the test data as well\n\n\n## Disadvantages\n- Without compression, it takes too much space; with compression, the data loading is slower.\n- We have to re-generate the data if we change the config. (image size, raster size or the number of historical frames)\n\n\n## Code: Generate and save\n```\n... imports ...\n... lyft configuration ...\n\ndef save_sample(i):\n    idx = random.randint(0, len(dataset))\n\n    # 300px, 0.5 raster size, 5 historical frames\n    obj_save(dataset[idx], f'sample_{i}', './cache/pre_300px__0_5__5')\n\ndef obj_save(obj, name, dir_cache):\n    with bz2.BZ2File(f'{dir_cache}/{name}.pbz', 'wb') as f:\n        pickle.dump(obj, f)\n\ndm = LocalDataManager()\ndataset_path = dm.require(cfg[\"train_data_loader\"][\"key\"])\nzarr_dataset = ChunkedDataset(dataset_path)\nzarr_dataset.open()\n\nrast = build_rasterizer(cfg, dm)\ndataset = AgentDataset(cfg, zarr_dataset, rast)\n\nwith Pool(processes=4) as p:\n    max_ = 32000\n    with tqdm(total=max_) as pbar:\n        for i, _ in enumerate(p.imap_unordered(save_sample, range(0, max_))):\n            pbar.update()\n\n```\n\n## Code: Data loader\n```\nclass LyftImageDataset(Dataset):\n\n    def __init__(self, data_folder):\n        super().__init__()\n        self.data_folder = data_folder\n        self.files = []\n\n        for filename in os.listdir(self.data_folder):\n            if filename.endswith(\".pbz\"):\n                self.files.append(filename)\n\n        print(len(self.files))\n        print(self.files[0])\n\n    def __getitem__(self, index: int):\n        return self.obj_load(self.files[index])\n\n    def obj_load(self, name):\n        with bz2.BZ2File(f'{self.data_folder}/{name}', 'rb') as f:\n            return pickle.load(f)\n\n    def __len__(self):\n        return len(self.files)\n\n...\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = LyftImageDataset('./cache/pre_300px__0_5__10')\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=train_cfg[\"shuffle\"],\n                              batch_size=train_cfg[\"batch_size\"],\n                              num_workers=train_cfg[\"num_workers\"])\n```\n\n",
      "votes": 59
    },
    {
      "id": 997174,
      "postDate": "2020-09-03T19:52:46.113Z",
      "content": "<p>You could also try:</p>\n<p>Create one notebook/script to extract and save the raster images from the <code>.zarr</code> files into a <code>local folder</code> using Peter's above code (on the CPU). Then create another notebook/script that reads and loads the processed raster images from the <code>local folder</code> for GPU training. After the first notebook processes a couple thousand images, you run the training notebook. This way, you create a pre-processed dataset to quickly train and experiment with in the future, but you can also train on GPU while waiting for the <code>.zarr</code> files to be rasterized by the CPU</p>",
      "rawMarkdown": "You could also try:\n\nCreate one notebook/script to extract and save the raster images from the `.zarr` files into a `local folder` using Peter's above code (on the CPU). Then create another notebook/script that reads and loads the processed raster images from the `local folder` for GPU training. After the first notebook processes a couple thousand images, you run the training notebook. This way, you create a pre-processed dataset to quickly train and experiment with in the future, but you can also train on GPU while waiting for the `.zarr` files to be rasterized by the CPU",
      "votes": 8
    },
    {
      "id": 1052152,
      "postDate": "2020-10-17T11:50:37.013Z",
      "content": "<p>I'm getting 17 days training time for my model on a decent system (rtx2080ti, 64gb of ram and i9 36 core cpu). I'm wondering how long has it taken for top leaderboard scores to train their models.</p>",
      "rawMarkdown": "I'm getting 17 days training time for my model on a decent system (rtx2080ti, 64gb of ram and i9 36 core cpu). I'm wondering how long has it taken for top leaderboard scores to train their models.",
      "votes": 1,
      "replies": [
        {
          "id": 1053424,
          "postDate": "2020-10-19T01:52:12.757Z",
          "content": "<p>For me usually a day (sometimes a few days for larger morels/more samples) on 2080ti, I train on the subset of the training dataset, usually 10-20 hours is sufficient to test if idea works or not.</p>",
          "rawMarkdown": "For me usually a day (sometimes a few days for larger morels/more samples) on 2080ti, I train on the subset of the training dataset, usually 10-20 hours is sufficient to test if idea works or not.",
          "votes": 9
        },
        {
          "id": 1053431,
          "postDate": "2020-10-19T02:06:15.623Z",
          "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> did you train on triain_full too or only train ? </p>",
          "rawMarkdown": "@dmytropoplavskiy did you train on triain_full too or only train ? ",
          "votes": 1
        },
        {
          "id": 1053477,
          "postDate": "2020-10-19T03:24:39.273Z",
          "content": "<p>only train, usually not even using all samples from train</p>",
          "rawMarkdown": "only train, usually not even using all samples from train",
          "votes": 2
        },
        {
          "id": 1053651,
          "postDate": "2020-10-19T07:42:14.060Z",
          "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a>  the default number of samples in the train is about 22 million, could I ask the number of samples in ur subset? and ur selecting method?</p>",
          "rawMarkdown": "@dmytropoplavskiy  the default number of samples in the train is about 22 million, could I ask the number of samples in ur subset? and ur selecting method?",
          "votes": 1
        },
        {
          "id": 1053824,
          "postDate": "2020-10-19T11:58:25.057Z",
          "content": "<p>For me it's faster to load pre-rendered images from ssd drive, so I preprocessed agents from each 4th frame from the dataset, so I'm running experiments on the 1/4 of the training dataset. During training, simple random sampling from pre-rendered samples.</p>",
          "rawMarkdown": "For me it's faster to load pre-rendered images from ssd drive, so I preprocessed agents from each 4th frame from the dataset, so I'm running experiments on the 1/4 of the training dataset. During training, simple random sampling from pre-rendered samples.",
          "votes": 13
        },
        {
          "id": 1056531,
          "postDate": "2020-10-21T19:43:37.767Z",
          "content": "<p>If you don't mind me asking, how long did preprocessing the data take <a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> ?</p>",
          "rawMarkdown": "If you don't mind me asking, how long did preprocessing the data take @dmytropoplavskiy ?",
          "votes": 1
        },
        {
          "id": 1056946,
          "postDate": "2020-10-22T08:06:30.880Z",
          "content": "<p>It depends on the resolution and number of channels, but usually around half of the day for 1/4th of the training dataset, on 16 cores CPU.</p>",
          "rawMarkdown": "It depends on the resolution and number of channels, but usually around half of the day for 1/4th of the training dataset, on 16 cores CPU.",
          "votes": 2
        },
        {
          "id": 1058633,
          "postDate": "2020-10-24T02:19:32.400Z",
          "content": "<p>I try to do this (preprocessing 1/4 of the dataset) with 267 resolution and default number of channels (10 history) using multiprocessing on 12 cores, and I am getting estimated times of 100ish hours. I am fetching one agent at a time from AgentDataset per process like the above code. Anything jump out to you?</p>",
          "rawMarkdown": "I try to do this (preprocessing 1/4 of the dataset) with 267 resolution and default number of channels (10 history) using multiprocessing on 12 cores, and I am getting estimated times of 100ish hours. I am fetching one agent at a time from AgentDataset per process like the above code. Anything jump out to you?",
          "votes": 1
        },
        {
          "id": 1059809,
          "postDate": "2020-10-25T13:44:03.997Z",
          "content": "<p>I was running it in parallel to utilize all cores, should make it faster. Maybe to have the separate AgentDataset instances per process? May be useful to uncompress the dataset files, the size difference is small but the random access is faster.</p>",
          "rawMarkdown": "I was running it in parallel to utilize all cores, should make it faster. Maybe to have the separate AgentDataset instances per process? May be useful to uncompress the dataset files, the size difference is small but the random access is faster.",
          "votes": 1
        },
        {
          "id": 1067401,
          "postDate": "2020-11-02T14:10:14.417Z",
          "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> is it possible to do the preprocessing of the AgentDataset to obtain the images on a Kaggle kernel? And if it is, one would not need GPU for it right as I believe raasterization only requires CPU?</p>",
          "rawMarkdown": "@dmytropoplavskiy is it possible to do the preprocessing of the AgentDataset to obtain the images on a Kaggle kernel? And if it is, one would not need GPU for it right as I believe raasterization only requires CPU?"
        }
      ]
    },
    {
      "id": 1018701,
      "postDate": "2020-09-19T21:25:55.590Z",
      "content": "<p>Hi friends,<br>\nmy training phase in the Lyft dataset is very time-consuming with high CPU usage ( **and very low GPU usage **)<br>\nIs it just my problem or u experienced it too? ( I'm just running the baseline they have provided )<br>\nand how to deal with it?<br>\nfor example, I have to about one week to iterate just about half of the train.zarr for the baseline.</p>\n<p>thanks in advance</p>\n<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a></p>",
      "rawMarkdown": "Hi friends,\nmy training phase in the Lyft dataset is very time-consuming with high CPU usage ( **and very low GPU usage **)\nIs it just my problem or u experienced it too? ( I'm just running the baseline they have provided )\nand how to deal with it?\nfor example, I have to about one week to iterate just about half of the train.zarr for the baseline.\n\nthanks in advance\n\n@pestipeti",
      "votes": 1
    },
    {
      "id": 1071579,
      "postDate": "2020-11-07T04:25:06.917Z",
      "content": "<p>Is the parameter max_ the number of samples to be cached? Thanks!</p>",
      "rawMarkdown": "Is the parameter max_ the number of samples to be cached? Thanks!",
      "replies": [
        {
          "id": 1071675,
          "postDate": "2020-11-07T08:31:55.143Z",
          "content": "<p>Yes, it is.</p>",
          "rawMarkdown": "Yes, it is."
        }
      ]
    },
    {
      "id": 999557,
      "postDate": "2020-09-05T18:50:48.333Z",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> \"./cache/pre_300px__0_5__5\" --&gt; this is the location where our rasterized images are stored, right? and one more thing, can you tell me what's the significance of 'historical frames'?  </p>",
      "rawMarkdown": "@pestipeti \"./cache/pre_300px__0_5__5\" --> this is the location where our rasterized images are stored, right? and one more thing, can you tell me what's the significance of 'historical frames'?  ",
      "replies": [
        {
          "id": 999794,
          "postDate": "2020-09-06T02:57:21.423Z",
          "content": "<p><a href=\"https://www.kaggle.com/pawankumarsahu\" target=\"_blank\">@pawankumarsahu</a> <br>\nYes, that is the dest folder, you can change it if you want. The historical frame is an l5kit feature, see their doc/sample code for more details.</p>",
          "rawMarkdown": "@pawankumarsahu \nYes, that is the dest folder, you can change it if you want. The historical frame is an l5kit feature, see their doc/sample code for more details.",
          "votes": 3
        }
      ]
    },
    {
      "id": 993896,
      "postDate": "2020-09-01T08:03:48.517Z",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> isn't the training time for 32000 (1000 iterations for batch size 32) samples approximately 40 minutes? </p>",
      "rawMarkdown": "@pestipeti isn't the training time for 32000 (1000 iterations for batch size 32) samples approximately 40 minutes? ",
      "replies": [
        {
          "id": 997198,
          "postDate": "2020-09-03T20:08:16.527Z",
          "content": "<p><a href=\"https://www.kaggle.com/axel81\" target=\"_blank\">@axel81</a> I did the experiments on my local computer. I have 12 cores (24threads) and I optimized the l5kit rasterizer a lot. ~20% faster with 4 cores. </p>",
          "rawMarkdown": "@axel81 I did the experiments on my local computer. I have 12 cores (24threads) and I optimized the l5kit rasterizer a lot. ~20% faster with 4 cores. ",
          "votes": 2
        },
        {
          "id": 997604,
          "postDate": "2020-09-04T06:07:17.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> the improvement is quite impressive!</p>",
          "rawMarkdown": "@pestipeti the improvement is quite impressive!"
        },
        {
          "id": 1051743,
          "postDate": "2020-10-16T19:34:55.020Z",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> Do you mind sharing how you optimized the l5kit  rasterization?</p>",
          "rawMarkdown": "@pestipeti Do you mind sharing how you optimized the l5kit  rasterization?",
          "votes": 2
        }
      ]
    },
    {
      "id": 987443,
      "postDate": "2020-08-27T09:07:58.720Z",
      "content": "<p>Thanks for sharing your experiments <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a><br>\n~2gb for 32k samples; that means for the approximately 1.4m train+val samples in the competition dataset, it would take up around 90gb. And if you were to use the full 70gb lyft dataset that'll be even more!<br>\nAnd all this for an improvement from ~1 it/s to ~5 it/s, still not being able to fully utilise the gpu. Is this really worth it? Does it make sense to generate and save an image dataset?</p>\n<p>It'd be nice if we can get the GPU to do some of the rasterization as well.</p>",
      "rawMarkdown": "Thanks for sharing your experiments @pestipeti\n~2gb for 32k samples; that means for the approximately 1.4m train+val samples in the competition dataset, it would take up around 90gb. And if you were to use the full 70gb lyft dataset that'll be even more!\nAnd all this for an improvement from ~1 it/s to ~5 it/s, still not being able to fully utilise the gpu. Is this really worth it? Does it make sense to generate and save an image dataset?\n\nIt'd be nice if we can get the GPU to do some of the rasterization as well.",
      "replies": [
        {
          "id": 987451,
          "postDate": "2020-08-27T09:20:09.743Z",
          "content": "<p>I think it is much worse. 1.4m is not the size of your dataloader? I am guessing you use batch=16, so the actual number of samples ~22.5M, it would be ~1.3Tb</p>\n<p>You are right, we should use the GPU for rasterization. OpenCV supports GPU (I think; I never used) and OpenGL worth a try as well. I am looking into options, I'll post if I find a better solution.</p>",
          "rawMarkdown": "I think it is much worse. 1.4m is not the size of your dataloader? I am guessing you use batch=16, so the actual number of samples ~22.5M, it would be ~1.3Tb\n\nYou are right, we should use the GPU for rasterization. OpenCV supports GPU (I think; I never used) and OpenGL worth a try as well. I am looking into options, I'll post if I find a better solution.",
          "votes": 2
        },
        {
          "id": 987900,
          "postDate": "2020-08-27T15:49:25.897Z",
          "content": "<p>Oh you're right, 1.4m is the train+val dataloaders combined with batch size 32, so even worse!</p>\n<p>I'll create a feature request for GPU rasterization on <code>l5kit</code>. If they can add something like <code>device = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")</code> in the <code>l5kit</code>, that'll be really convenient. I've not worked with rasterization before so I'm not sure how much effort it would take to change all the rasterization code and other dependencies for GPU utilization.</p>",
          "rawMarkdown": "Oh you're right, 1.4m is the train+val dataloaders combined with batch size 32, so even worse!\n\nI'll create a feature request for GPU rasterization on `l5kit`. If they can add something like `device = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")` in the `l5kit`, that'll be really convenient. I've not worked with rasterization before so I'm not sure how much effort it would take to change all the rasterization code and other dependencies for GPU utilization."
        }
      ]
    },
    {
      "id": 1058624,
      "postDate": "2020-10-24T01:39:47.530Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1043094,
      "postDate": "2020-10-08T17:18:31.680Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1043233,
          "postDate": "2020-10-08T19:35:32.913Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 989953,
      "postDate": "2020-08-29T08:32:11.987Z",
      "content": "<p>thanks for sharing ! </p>",
      "rawMarkdown": "thanks for sharing ! "
    },
    {
      "id": 988863,
      "postDate": "2020-08-28T10:49:53.030Z",
      "content": "<p>thanks a lot :)</p>",
      "rawMarkdown": "thanks a lot :)"
    }
  ],
  "comments": [
    {
      "id": 997174,
      "author_name": "Tucker Arrants",
      "author_url": "",
      "post_date": "2020-09-03T19:52:46.113000",
      "content": "<p>You could also try:</p>\n<p>Create one notebook/script to extract and save the raster images from the <code>.zarr</code> files into a <code>local folder</code> using Peter's above code (on the CPU). Then create another notebook/script that reads and loads the processed raster images from the <code>local folder</code> for GPU training. After the first notebook processes a couple thousand images, you run the training notebook. This way, you create a pre-processed dataset to quickly train and experiment with in the future, but you can also train on GPU while waiting for the <code>.zarr</code> files to be rasterized by the CPU</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1052152,
      "author_name": "NTL",
      "author_url": "",
      "post_date": "2020-10-17T11:50:37.013000",
      "content": "<p>I'm getting 17 days training time for my model on a decent system (rtx2080ti, 64gb of ram and i9 36 core cpu). I'm wondering how long has it taken for top leaderboard scores to train their models.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1053424,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-10-19T01:52:12.757000",
          "content": "<p>For me usually a day (sometimes a few days for larger morels/more samples) on 2080ti, I train on the subset of the training dataset, usually 10-20 hours is sufficient to test if idea works or not.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1053431,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2020-10-19T02:06:15.623000",
          "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> did you train on triain_full too or only train ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1053477,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-10-19T03:24:39.273000",
          "content": "<p>only train, usually not even using all samples from train</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1053651,
          "author_name": "sakht977",
          "author_url": "",
          "post_date": "2020-10-19T07:42:14.060000",
          "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a>  the default number of samples in the train is about 22 million, could I ask the number of samples in ur subset? and ur selecting method?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1053824,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-10-19T11:58:25.057000",
          "content": "<p>For me it's faster to load pre-rendered images from ssd drive, so I preprocessed agents from each 4th frame from the dataset, so I'm running experiments on the 1/4 of the training dataset. During training, simple random sampling from pre-rendered samples.</p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 1056531,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-10-21T19:43:37.767000",
          "content": "<p>If you don't mind me asking, how long did preprocessing the data take <a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1056946,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-10-22T08:06:30.880000",
          "content": "<p>It depends on the resolution and number of channels, but usually around half of the day for 1/4th of the training dataset, on 16 cores CPU.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1058633,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-10-24T02:19:32.400000",
          "content": "<p>I try to do this (preprocessing 1/4 of the dataset) with 267 resolution and default number of channels (10 history) using multiprocessing on 12 cores, and I am getting estimated times of 100ish hours. I am fetching one agent at a time from AgentDataset per process like the above code. Anything jump out to you?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1059809,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-10-25T13:44:03.997000",
          "content": "<p>I was running it in parallel to utilize all cores, should make it faster. Maybe to have the separate AgentDataset instances per process? May be useful to uncompress the dataset files, the size difference is small but the random access is faster.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1067401,
          "author_name": "Arpit",
          "author_url": "",
          "post_date": "2020-11-02T14:10:14.417000",
          "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> is it possible to do the preprocessing of the AgentDataset to obtain the images on a Kaggle kernel? And if it is, one would not need GPU for it right as I believe raasterization only requires CPU?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1018701,
      "author_name": "sakht977",
      "author_url": "",
      "post_date": "2020-09-19T21:25:55.590000",
      "content": "<p>Hi friends,<br>\nmy training phase in the Lyft dataset is very time-consuming with high CPU usage ( **and very low GPU usage **)<br>\nIs it just my problem or u experienced it too? ( I'm just running the baseline they have provided )<br>\nand how to deal with it?<br>\nfor example, I have to about one week to iterate just about half of the train.zarr for the baseline.</p>\n<p>thanks in advance</p>\n<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1071579,
      "author_name": "Gold Retriever",
      "author_url": "",
      "post_date": "2020-11-07T04:25:06.917000",
      "content": "<p>Is the parameter max_ the number of samples to be cached? Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1071675,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-11-07T08:31:55.143000",
          "content": "<p>Yes, it is.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 999557,
      "author_name": "Pawan KS",
      "author_url": "",
      "post_date": "2020-09-05T18:50:48.333000",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> \"./cache/pre_300px__0_5__5\" --&gt; this is the location where our rasterized images are stored, right? and one more thing, can you tell me what's the significance of 'historical frames'?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 999794,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-06T02:57:21.423000",
          "content": "<p><a href=\"https://www.kaggle.com/pawankumarsahu\" target=\"_blank\">@pawankumarsahu</a> <br>\nYes, that is the dest folder, you can change it if you want. The historical frame is an l5kit feature, see their doc/sample code for more details.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 993896,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2020-09-01T08:03:48.517000",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> isn't the training time for 32000 (1000 iterations for batch size 32) samples approximately 40 minutes? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 997198,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-03T20:08:16.527000",
          "content": "<p><a href=\"https://www.kaggle.com/axel81\" target=\"_blank\">@axel81</a> I did the experiments on my local computer. I have 12 cores (24threads) and I optimized the l5kit rasterizer a lot. ~20% faster with 4 cores. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 997604,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-09-04T06:07:17.807000",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> the improvement is quite impressive!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1051743,
          "author_name": "Artem.Sanakoev",
          "author_url": "",
          "post_date": "2020-10-16T19:34:55.020000",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> Do you mind sharing how you optimized the l5kit  rasterization?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 987443,
      "author_name": "Khushal B",
      "author_url": "",
      "post_date": "2020-08-27T09:07:58.720000",
      "content": "<p>Thanks for sharing your experiments <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a><br>\n~2gb for 32k samples; that means for the approximately 1.4m train+val samples in the competition dataset, it would take up around 90gb. And if you were to use the full 70gb lyft dataset that'll be even more!<br>\nAnd all this for an improvement from ~1 it/s to ~5 it/s, still not being able to fully utilise the gpu. Is this really worth it? Does it make sense to generate and save an image dataset?</p>\n<p>It'd be nice if we can get the GPU to do some of the rasterization as well.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 987451,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-08-27T09:20:09.743000",
          "content": "<p>I think it is much worse. 1.4m is not the size of your dataloader? I am guessing you use batch=16, so the actual number of samples ~22.5M, it would be ~1.3Tb</p>\n<p>You are right, we should use the GPU for rasterization. OpenCV supports GPU (I think; I never used) and OpenGL worth a try as well. I am looking into options, I'll post if I find a better solution.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 987900,
          "author_name": "Khushal B",
          "author_url": "",
          "post_date": "2020-08-27T15:49:25.897000",
          "content": "<p>Oh you're right, 1.4m is the train+val dataloaders combined with batch size 32, so even worse!</p>\n<p>I'll create a feature request for GPU rasterization on <code>l5kit</code>. If they can add something like <code>device = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")</code> in the <code>l5kit</code>, that'll be really convenient. I've not worked with rasterization before so I'm not sure how much effort it would take to change all the rasterization code and other dependencies for GPU utilization.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1058624,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-24T01:39:47.530000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1043094,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-08T17:18:31.680000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1043233,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-08T19:35:32.913000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 989953,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2020-08-29T08:32:11.987000",
      "content": "<p>thanks for sharing ! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 988863,
      "author_name": "efe.eroglu",
      "author_url": "",
      "post_date": "2020-08-28T10:49:53.030000",
      "content": "<p>thanks a lot :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "986716": "In this competition, we have to generate the image for our models. Unfortunately, the rasterization is a very slow process (if we do it on the CPU).\n\nMy idea is to create a preprocessed image dataset. Generate a million samples (images and additional data). Train the model without rasterizing new images. We can fine-tune later with new samples (with rasterization)\n\n## Experiments:\n| Method             | Size | History | \\# of samples | Storage size | Training time | it/sec |\n| ------------------ | ---- | ------- | ------------- | ------------ | ------------- | ------ |\n| Rasterizer         | 300  | 10      | 32000         | 0            | 4:48          | 3.47   |\n| Rasterizer         | 300  | 5       | 32000         | 0            | 3:46          | 4.42   |\n| Pre-generated data | 300  | 10      | 32000         | 1.9 Gb       | 3:26          | 4.85   |\n| Pre-generated data | 300  | 5       | 32000         | 1.8 Gb       | 2:56          | 5.68   |\n\n\n## Advantages\n- We can speed up the training process\n- We can save the validation and the test data as well\n\n\n## Disadvantages\n- Without compression, it takes too much space; with compression, the data loading is slower.\n- We have to re-generate the data if we change the config. (image size, raster size or the number of historical frames)\n\n\n## Code: Generate and save\n```\n... imports ...\n... lyft configuration ...\n\ndef save_sample(i):\n    idx = random.randint(0, len(dataset))\n\n    # 300px, 0.5 raster size, 5 historical frames\n    obj_save(dataset[idx], f'sample_{i}', './cache/pre_300px__0_5__5')\n\ndef obj_save(obj, name, dir_cache):\n    with bz2.BZ2File(f'{dir_cache}/{name}.pbz', 'wb') as f:\n        pickle.dump(obj, f)\n\ndm = LocalDataManager()\ndataset_path = dm.require(cfg[\"train_data_loader\"][\"key\"])\nzarr_dataset = ChunkedDataset(dataset_path)\nzarr_dataset.open()\n\nrast = build_rasterizer(cfg, dm)\ndataset = AgentDataset(cfg, zarr_dataset, rast)\n\nwith Pool(processes=4) as p:\n    max_ = 32000\n    with tqdm(total=max_) as pbar:\n        for i, _ in enumerate(p.imap_unordered(save_sample, range(0, max_))):\n            pbar.update()\n\n```\n\n## Code: Data loader\n```\nclass LyftImageDataset(Dataset):\n\n    def __init__(self, data_folder):\n        super().__init__()\n        self.data_folder = data_folder\n        self.files = []\n\n        for filename in os.listdir(self.data_folder):\n            if filename.endswith(\".pbz\"):\n                self.files.append(filename)\n\n        print(len(self.files))\n        print(self.files[0])\n\n    def __getitem__(self, index: int):\n        return self.obj_load(self.files[index])\n\n    def obj_load(self, name):\n        with bz2.BZ2File(f'{self.data_folder}/{name}', 'rb') as f:\n            return pickle.load(f)\n\n    def __len__(self):\n        return len(self.files)\n\n...\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = LyftImageDataset('./cache/pre_300px__0_5__10')\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=train_cfg[\"shuffle\"],\n                              batch_size=train_cfg[\"batch_size\"],\n                              num_workers=train_cfg[\"num_workers\"])\n```\n\n",
    "997174": "You could also try:\n\nCreate one notebook/script to extract and save the raster images from the `.zarr` files into a `local folder` using Peter's above code (on the CPU). Then create another notebook/script that reads and loads the processed raster images from the `local folder` for GPU training. After the first notebook processes a couple thousand images, you run the training notebook. This way, you create a pre-processed dataset to quickly train and experiment with in the future, but you can also train on GPU while waiting for the `.zarr` files to be rasterized by the CPU",
    "1052152": "I'm getting 17 days training time for my model on a decent system (rtx2080ti, 64gb of ram and i9 36 core cpu). I'm wondering how long has it taken for top leaderboard scores to train their models.",
    "1018701": "Hi friends,\nmy training phase in the Lyft dataset is very time-consuming with high CPU usage ( **and very low GPU usage **)\nIs it just my problem or u experienced it too? ( I'm just running the baseline they have provided )\nand how to deal with it?\nfor example, I have to about one week to iterate just about half of the train.zarr for the baseline.\n\nthanks in advance\n\n@pestipeti",
    "1071579": "Is the parameter max_ the number of samples to be cached? Thanks!",
    "999557": "@pestipeti \"./cache/pre_300px__0_5__5\" --> this is the location where our rasterized images are stored, right? and one more thing, can you tell me what's the significance of 'historical frames'?  ",
    "993896": "@pestipeti isn't the training time for 32000 (1000 iterations for batch size 32) samples approximately 40 minutes? ",
    "987443": "Thanks for sharing your experiments @pestipeti\n~2gb for 32k samples; that means for the approximately 1.4m train+val samples in the competition dataset, it would take up around 90gb. And if you were to use the full 70gb lyft dataset that'll be even more!\nAnd all this for an improvement from ~1 it/s to ~5 it/s, still not being able to fully utilise the gpu. Is this really worth it? Does it make sense to generate and save an image dataset?\n\nIt'd be nice if we can get the GPU to do some of the rasterization as well.",
    "1058624": "",
    "1043094": "",
    "989953": "thanks for sharing ! ",
    "988863": "thanks a lot :)"
  }
}