{
  "id": 196602,
  "title": "Bag of optimization tricks",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/196602",
  "author_name": "",
  "post_date": "2020-11-11T21:51:14.805938200Z",
  "votes": 40,
  "comment_count": 29,
  "views": 0,
  "content": "<p>As you probably know, one of the difficulties in this competition is speed. You can improve it if you optimize the data-loading and the rasterization process. In the past weeks/months, I've tried to improve the training speed as much as possible. In this post, I'd like to share some details. </p>\n<p>Leave me a comment if you have any other idea/technique.</p>\n<h2>Optimizations</h2>\n<h3>Basic</h3>\n<p>For these tricks, you don't have to modify the l5kit.</p>\n<ul>\n<li><strong>Pre-generate training dataset (partial)</strong> For more details, take a look at <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this post</a></li>\n<li><strong>Pre-generate validation/test dataset (full)</strong></li>\n<li><strong>Implement a faster metric calculation method</strong>. The official metric saves and re-loads your predictions. It is faster if you keep everything in memory.</li>\n<li><strong>Validate less frequently</strong>. Reduce your validation rounds. I validate after 5000 iterations for new experiments, but during a long training process, I reduce it to every 25000.</li>\n<li><strong>Validate using part of the official validation set.</strong> Make sure that the subset's distribution does not change!</li>\n<li><strong>Validate on public LB.</strong> Fortunately, our datasets are stable. You can skip the local validation.  Instead of predicting the validation set, predict and submit the test set during training. You only need to \"validate\" on the test set 5 times per day!</li>\n<li><strong>Reduce the height of the image</strong> The agents are always moving towards positive-x direction. You can reduce the height of the input images.</li>\n</ul>\n<h3>Intermediate</h3>\n<p>For a bit more speed, you'll have to modify the l5kit.</p>\n<ul>\n<li><strong>Optimize the box rasterizer.</strong> Try to simplify the <code>draw_boxes</code> method - especially the matrix multiplications, point transformations. Try to eliminate the costly Numpy calls like <code>vstack</code>, <code>transpose</code>, <code>concatenate</code>, etc. Take a look at the <code>cv2</code> drawing calls as well. Get rid of everything that does not add information to the training.</li>\n<li><strong>Optimize the semantic rasterizer.</strong> You can gain 6-8% if you optimize this. Same as above. Hint: Check out my <a href=\"https://github.com/lyft/l5kit/pull/140\" target=\"_blank\">pull-request</a> (pending).</li>\n<li><strong>Data caching.</strong> Besides the rasterization, the other bottleneck is the data loading. L5Kit has to do a lot of work to load and prepare the data for the rasterizers. You can speed things up if you save the prepared data inside l5kit (<code>agent_sampling</code>; without the rasterized images).</li>\n</ul>\n<h3>Advanced</h3>\n<p>If you want significant speed improvement, you'll have to rewrite the rasterization in C++ and CUDA. And you'll need a few other tricks as well. We'll share the details, my C++/CUDA implementation, and our training method after the competition.</p>\n<h2>Training time</h2>\n<p>There are lots of things that could affect the training time. If you optimize your training process, you can speed up the training process by 2.0-2.5-3.0x</p>\n<p>Stats of my latest experiment:</p>\n<ul>\n<li>16,000,000 samples (250,000 iterations; 64 batch)</li>\n<li>Validating: 100% of the validation set after every 5000 iterations</li>\n<li>Running time: 25 hours (~23 hours with 1/5 of validation)</li>\n<li>Image size: Bigger than I used in the experiments below</li>\n<li>Model: Bigger than in the experiments below</li>\n<li>Public LB: 16.180</li>\n</ul>\n<p><em>I used 20 CPU cores, I have 128Gb RAM, 1Tb SSD, and an RTX-2080ti</em></p>\n<h2>Results</h2>\n<p>For all of the experiments below I used:</p>\n<ul>\n<li>300x300px input size</li>\n<li>0.5 pixel size</li>\n<li>10 historical frames</li>\n<li>32 batch</li>\n<li>1000 iterations</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Net</th>\n<th># of workers</th>\n<th># of samples</th>\n<th>Storage size</th>\n<th>Training time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>l5kit</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>-</td>\n<td>4:48</td>\n<td>3.47</td>\n</tr>\n<tr>\n<td>optimized l5kit</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>-</td>\n<td>3:35*</td>\n<td>4.65*</td>\n</tr>\n<tr>\n<td>Pre-generated (w images)</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>1.9 Gb</td>\n<td>3:26</td>\n<td>4.85</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>0.36 Gb</td>\n<td>1:42</td>\n<td>9.81</td>\n</tr>\n</tbody>\n</table>\n<p>* Estimated values (I used different settings for this experiment).</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Net</th>\n<th># of workers</th>\n<th>Storage</th>\n<th>Training time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>l5kit</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>-</td>\n<td>4:48</td>\n<td>3.47</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>0.36 Gb</td>\n<td>1:42</td>\n<td>9.81</td>\n</tr>\n<tr>\n<td>l5kit</td>\n<td>Resnet-18</td>\n<td>4</td>\n<td>-</td>\n<td>6:45</td>\n<td>2.46</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Resnet-18</td>\n<td>4</td>\n<td>0.36 Gb</td>\n<td>1:43</td>\n<td>9.79</td>\n</tr>\n<tr>\n<td>l5kit</td>\n<td>Eff-B0</td>\n<td>20</td>\n<td>-</td>\n<td>4:45</td>\n<td>3.51</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Eff-B0</td>\n<td>20</td>\n<td>0.36 Gb</td>\n<td>2:43</td>\n<td>6.13</td>\n</tr>\n<tr>\n<td>l5kit</td>\n<td>Eff-B0</td>\n<td>4</td>\n<td>-</td>\n<td>6:48</td>\n<td>2.44</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Eff-B0</td>\n<td>4</td>\n<td>0.36 Gb</td>\n<td>2:43</td>\n<td>6.13</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "1075680",
      "postDate": "11/11/2020 21:51:14",
      "content": "<p>As you probably know, one of the difficulties in this competition is speed. You can improve it if you optimize the data-loading and the rasterization process. In the past weeks/months, I've tried to improve the training speed as much as possible. In this post, I'd like to share some details. </p>\n<p>Leave me a comment if you have any other idea/technique.</p>\n<h2>Optimizations</h2>\n<h3>Basic</h3>\n<p>For these tricks, you don't have to modify the l5kit.</p>\n<ul>\n<li><strong>Pre-generate training dataset (partial)</strong> For more details, take a look at <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this post</a></li>\n<li><strong>Pre-generate validation/test dataset (full)</strong></li>\n<li><strong>Implement a faster metric calculation method</strong>. The official metric saves and re-loads your predictions. It is faster if you keep everything in memory.</li>\n<li><strong>Validate less frequently</strong>. Reduce your validation rounds. I validate after 5000 iterations for new experiments, but during a long training process, I reduce it to every 25000.</li>\n<li><strong>Validate using part of the official validation set.</strong> Make sure that the subset's distribution does not change!</li>\n<li><strong>Validate on public LB.</strong> Fortunately, our datasets are stable. You can skip the local validation.  Instead of predicting the validation set, predict and submit the test set during training. You only need to \"validate\" on the test set 5 times per day!</li>\n<li><strong>Reduce the height of the image</strong> The agents are always moving towards positive-x direction. You can reduce the height of the input images.</li>\n</ul>\n<h3>Intermediate</h3>\n<p>For a bit more speed, you'll have to modify the l5kit.</p>\n<ul>\n<li><strong>Optimize the box rasterizer.</strong> Try to simplify the <code>draw_boxes</code> method - especially the matrix multiplications, point transformations. Try to eliminate the costly Numpy calls like <code>vstack</code>, <code>transpose</code>, <code>concatenate</code>, etc. Take a look at the <code>cv2</code> drawing calls as well. Get rid of everything that does not add information to the training.</li>\n<li><strong>Optimize the semantic rasterizer.</strong> You can gain 6-8% if you optimize this. Same as above. Hint: Check out my <a href=\"https://github.com/lyft/l5kit/pull/140\" target=\"_blank\">pull-request</a> (pending).</li>\n<li><strong>Data caching.</strong> Besides the rasterization, the other bottleneck is the data loading. L5Kit has to do a lot of work to load and prepare the data for the rasterizers. You can speed things up if you save the prepared data inside l5kit (<code>agent_sampling</code>; without the rasterized images).</li>\n</ul>\n<h3>Advanced</h3>\n<p>If you want significant speed improvement, you'll have to rewrite the rasterization in C++ and CUDA. And you'll need a few other tricks as well. We'll share the details, my C++/CUDA implementation, and our training method after the competition.</p>\n<h2>Training time</h2>\n<p>There are lots of things that could affect the training time. If you optimize your training process, you can speed up the training process by 2.0-2.5-3.0x</p>\n<p>Stats of my latest experiment:</p>\n<ul>\n<li>16,000,000 samples (250,000 iterations; 64 batch)</li>\n<li>Validating: 100% of the validation set after every 5000 iterations</li>\n<li>Running time: 25 hours (~23 hours with 1/5 of validation)</li>\n<li>Image size: Bigger than I used in the experiments below</li>\n<li>Model: Bigger than in the experiments below</li>\n<li>Public LB: 16.180</li>\n</ul>\n<p><em>I used 20 CPU cores, I have 128Gb RAM, 1Tb SSD, and an RTX-2080ti</em></p>\n<h2>Results</h2>\n<p>For all of the experiments below I used:</p>\n<ul>\n<li>300x300px input size</li>\n<li>0.5 pixel size</li>\n<li>10 historical frames</li>\n<li>32 batch</li>\n<li>1000 iterations</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Net</th>\n<th># of workers</th>\n<th># of samples</th>\n<th>Storage size</th>\n<th>Training time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>l5kit</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>-</td>\n<td>4:48</td>\n<td>3.47</td>\n</tr>\n<tr>\n<td>optimized l5kit</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>-</td>\n<td>3:35*</td>\n<td>4.65*</td>\n</tr>\n<tr>\n<td>Pre-generated (w images)</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>1.9 Gb</td>\n<td>3:26</td>\n<td>4.85</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>32000</td>\n<td>0.36 Gb</td>\n<td>1:42</td>\n<td>9.81</td>\n</tr>\n</tbody>\n</table>\n<p>* Estimated values (I used different settings for this experiment).</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Net</th>\n<th># of workers</th>\n<th>Storage</th>\n<th>Training time</th>\n<th>it/sec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>l5kit</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>-</td>\n<td>4:48</td>\n<td>3.47</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Resnet-18</td>\n<td>20</td>\n<td>0.36 Gb</td>\n<td>1:42</td>\n<td>9.81</td>\n</tr>\n<tr>\n<td>l5kit</td>\n<td>Resnet-18</td>\n<td>4</td>\n<td>-</td>\n<td>6:45</td>\n<td>2.46</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Resnet-18</td>\n<td>4</td>\n<td>0.36 Gb</td>\n<td>1:43</td>\n<td>9.79</td>\n</tr>\n<tr>\n<td>l5kit</td>\n<td>Eff-B0</td>\n<td>20</td>\n<td>-</td>\n<td>4:45</td>\n<td>3.51</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Eff-B0</td>\n<td>20</td>\n<td>0.36 Gb</td>\n<td>2:43</td>\n<td>6.13</td>\n</tr>\n<tr>\n<td>l5kit</td>\n<td>Eff-B0</td>\n<td>4</td>\n<td>-</td>\n<td>6:48</td>\n<td>2.44</td>\n</tr>\n<tr>\n<td>advanced</td>\n<td>Eff-B0</td>\n<td>4</td>\n<td>0.36 Gb</td>\n<td>2:43</td>\n<td>6.13</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "As you probably know, one of the difficulties in this competition is speed. You can improve it if you optimize the data-loading and the rasterization process. In the past weeks/months, I've tried to improve the training speed as much as possible. In this post, I'd like to share some details. \n\nLeave me a comment if you have any other idea/technique.\n\n## Optimizations\n\n### Basic\n\nFor these tricks, you don't have to modify the l5kit.\n- **Pre-generate training dataset (partial)** For more details, take a look at [this post](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637)\n- **Pre-generate validation/test dataset (full)**\n- **Implement a faster metric calculation method**. The official metric saves and re-loads your predictions. It is faster if you keep everything in memory.\n- **Validate less frequently**. Reduce your validation rounds. I validate after 5000 iterations for new experiments, but during a long training process, I reduce it to every 25000.\n- **Validate using part of the official validation set.** Make sure that the subset's distribution does not change!\n- **Validate on public LB.** Fortunately, our datasets are stable. You can skip the local validation.  Instead of predicting the validation set, predict and submit the test set during training. You only need to \"validate\" on the test set 5 times per day!\n- **Reduce the height of the image** The agents are always moving towards positive-x direction. You can reduce the height of the input images.\n\n### Intermediate\n\nFor a bit more speed, you'll have to modify the l5kit.\n\n- **Optimize the box rasterizer.** Try to simplify the `draw_boxes` method - especially the matrix multiplications, point transformations. Try to eliminate the costly Numpy calls like `vstack`, `transpose`, `concatenate`, etc. Take a look at the `cv2` drawing calls as well. Get rid of everything that does not add information to the training.\n- **Optimize the semantic rasterizer.** You can gain 6-8% if you optimize this. Same as above. Hint: Check out my [pull-request](https://github.com/lyft/l5kit/pull/140) (pending).\n- **Data caching.** Besides the rasterization, the other bottleneck is the data loading. L5Kit has to do a lot of work to load and prepare the data for the rasterizers. You can speed things up if you save the prepared data inside l5kit (`agent_sampling`; without the rasterized images).\n\n\n\n### Advanced\n\nIf you want significant speed improvement, you'll have to rewrite the rasterization in C++ and CUDA. And you'll need a few other tricks as well. We'll share the details, my C++/CUDA implementation, and our training method after the competition.\n\n## Training time\n\nThere are lots of things that could affect the training time. If you optimize your training process, you can speed up the training process by 2.0-2.5-3.0x\n\nStats of my latest experiment:\n\n- 16,000,000 samples (250,000 iterations; 64 batch)\n- Validating: 100% of the validation set after every 5000 iterations\n- Running time: 25 hours (~23 hours with 1/5 of validation)\n- Image size: Bigger than I used in the experiments below\n- Model: Bigger than in the experiments below\n- Public LB: 16.180\n\n\n*I used 20 CPU cores, I have 128Gb RAM, 1Tb SSD, and an RTX-2080ti*\n\n## Results\n\nFor all of the experiments below I used:\n\n- 300x300px input size\n- 0.5 pixel size\n- 10 historical frames\n- 32 batch\n- 1000 iterations\n\n| Method                           | Net       | \\# of workers | \\# of samples | Storage size | Training time | it/sec |\n| -------------------------------- | --------- | ------------- | ------------- | ------------ | ------------- | ------ |\n| l5kit                            | Resnet-18 | 20            | 32000         | -            | 4:48          | 3.47   |\n| optimized l5kit                  | Resnet-18 | 20            | 32000         | -            | 3:35\\*        | 4.65\\* |\n| Pre-generated (w images) | Resnet-18 | 20            | 32000         | 1.9 Gb       | 3:26          | 4.85   |\n| advanced                         | Resnet-18 | 20            | 32000         | 0.36 Gb      | 1:42          | 9.81   |\n\n\\* Estimated values (I used different settings for this experiment).\n\n\n\n| Method   | Net       | \\# of workers | Storage | Training time | it/sec |\n| -------- | --------- | ------------- | ------- | ------------- | ------ |\n| l5kit    | Resnet-18 | 20            | -       | 4:48          | 3.47   |\n| advanced | Resnet-18 | 20            | 0.36 Gb | 1:42          | 9.81   |\n| l5kit    | Resnet-18 | 4             | -       | 6:45          | 2.46   |\n| advanced | Resnet-18 | 4             | 0.36 Gb | 1:43          | 9.79   |\n| l5kit    | Eff-B0    | 20            | -       | 4:45          | 3.51   |\n| advanced | Eff-B0    | 20            | 0.36 Gb | 2:43          | 6.13   |\n| l5kit    | Eff-B0    | 4             | -       | 6:48          | 2.44   |\n| advanced | Eff-B0    | 4             | 0.36 Gb | 2:43          | 6.13   |",
      "votes": null
    },
    {
      "id": "1075712",
      "postDate": "11/11/2020 23:05:11",
      "content": "<p>Thank you! This is amazing! I was wondering if your current train stats are based on pre-generate cached data or not?</p>",
      "rawMarkdown": "Thank you! This is amazing! I was wondering if your current train stats are based on pre-generate cached data or not?",
      "votes": null
    },
    {
      "id": "1075717",
      "postDate": "11/11/2020 23:09:37",
      "content": "<p>We cache data, but without the rasterized image.</p>",
      "rawMarkdown": "We cache data, but without the rasterized image.",
      "votes": null
    },
    {
      "id": "1075739",
      "postDate": "11/11/2020 23:28:14",
      "content": "<p>Could you explain a bit what's the difference between caching data without and caching data with rasterized image? (Or a link to another discussion post that contains this info). Thank you very much!</p>",
      "rawMarkdown": "Could you explain a bit what's the difference between caching data without and caching data with rasterized image? (Or a link to another discussion post that contains this info). Thank you very much!",
      "votes": null
    },
    {
      "id": "1075751",
      "postDate": "11/11/2020 23:40:23",
      "content": "<p>Caching with images is simple (see <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this post</a>). Its drawback is the size. You have to compress, but in that case, the loading is slow.</p>\n<p>You can cache without the image, but it is harder to implement (you need to modify the l5kit). With this solution, you can save the l5kit's data-loading/zarr-reading time, but you still have to rasterize. <br>\nFor us, this one is optimal, but I implement the rasterization in C++/CUDA. It could work with CPU rasterization, but I haven't tested it.</p>",
      "rawMarkdown": "Caching with images is simple (see [this post](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637)). Its drawback is the size. You have to compress, but in that case, the loading is slow.\n\nYou can cache without the image, but it is harder to implement (you need to modify the l5kit). With this solution, you can save the l5kit's data-loading/zarr-reading time, but you still have to rasterize. \nFor us, this one is optimal, but I implement the rasterization in C++/CUDA. It could work with CPU rasterization, but I haven't tested it.",
      "votes": null
    },
    {
      "id": "1075769",
      "postDate": "11/12/2020 00:17:47",
      "content": "<p>Do you have performance numbers from eff-b0? I haven't explore other backbones at all. Not sure if it is worth my time or just wasted compute. </p>",
      "rawMarkdown": "Do you have performance numbers from eff-b0? I haven't explore other backbones at all. Not sure if it is worth my time or just wasted compute.",
      "votes": null
    },
    {
      "id": "1075887",
      "postDate": "11/12/2020 04:02:50",
      "content": "<p>Sounds like this is more a competition to help Lyft rewrite and optimize their l5kit package?</p>",
      "rawMarkdown": "Sounds like this is more a competition to help Lyft rewrite and optimize their l5kit package?",
      "votes": null
    },
    {
      "id": "1075889",
      "postDate": "11/12/2020 04:10:02",
      "content": "<p>That does seem to be the primary value lyft is getting out of this. A big audit of their package by dangling a competition in front of us to motivate us to interact with it. </p>",
      "rawMarkdown": "That does seem to be the primary value lyft is getting out of this. A big audit of their package by dangling a competition in front of us to motivate us to interact with it.",
      "votes": null
    },
    {
      "id": "1076217",
      "postDate": "11/12/2020 10:42:09",
      "content": "<p>I tried eff-b0 but it underperformed resnet considerably</p>",
      "rawMarkdown": "I tried eff-b0 but it underperformed resnet considerably",
      "votes": null
    },
    {
      "id": "1076320",
      "postDate": "11/12/2020 12:49:51",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1076321",
      "postDate": "11/12/2020 12:50:47",
      "content": "<p>May I ask if your current score is based on resnet or not? <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> </p>",
      "rawMarkdown": "May I ask if your current score is based on resnet or not? @fergusoci",
      "votes": null
    },
    {
      "id": "1076576",
      "postDate": "11/12/2020 16:58:27",
      "content": "<p>Yes, it's based on resnet</p>",
      "rawMarkdown": "Yes, it's based on resnet",
      "votes": null
    },
    {
      "id": "1076663",
      "postDate": "11/12/2020 18:15:34",
      "content": "<p>From my experiments, I understand that :</p>\n<ul>\n<li>Larger network will improve in a small amount but improves forever</li>\n<li>But in Smaller networks like resnet18 it stops at point and lags behind it</li>\n</ul>\n<p>are my experiments true or not?</p>",
      "rawMarkdown": "From my experiments, I understand that :\n- Larger network will improve in a small amount but improves forever\n- But in Smaller networks like resnet18 it stops at point and lags behind it\n\nare my experiments true or not?",
      "votes": null
    },
    {
      "id": "1076664",
      "postDate": "11/12/2020 18:15:57",
      "content": "<p>what l5kit version are you using</p>",
      "rawMarkdown": "what l5kit version are you using",
      "votes": null
    },
    {
      "id": "1076674",
      "postDate": "11/12/2020 18:23:06",
      "content": "<p>The latest one (think it's 1.1.0)</p>",
      "rawMarkdown": "The latest one (think it's 1.1.0)",
      "votes": null
    },
    {
      "id": "1076943",
      "postDate": "11/13/2020 04:09:04",
      "content": "<p>I am not sure larger network is better. Looks like you can easily overfit in the large network. I have tried resnet 18, 34, and 50. The resnet50 has the best train score but the worse LB and validation score. On the other hand, resnet 18 has almost no gap between train, val, and LB score while give me the best score on LB so far.</p>",
      "rawMarkdown": "I am not sure larger network is better. Looks like you can easily overfit in the large network. I have tried resnet 18, 34, and 50. The resnet50 has the best train score but the worse LB and validation score. On the other hand, resnet 18 has almost no gap between train, val, and LB score while give me the best score on LB so far.",
      "votes": null
    },
    {
      "id": "1076946",
      "postDate": "11/13/2020 04:11:35",
      "content": "<p>It's only my experiments.<br>\nThanks, i felt that large network are easy to fit in train set but overfit in LB or Validation</p>",
      "rawMarkdown": "It's only my experiments.\nThanks, i felt that large network are easy to fit in train set but overfit in LB or Validation",
      "votes": null
    },
    {
      "id": "1076961",
      "postDate": "11/13/2020 04:41:39",
      "content": "<p>That is amazing <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>. I think it was so hard to recode all of them to C++ from python.<br>\nI just was working in converting slower functions to C++ like Rasterization</p>",
      "rawMarkdown": "That is amazing @pestipeti. I think it was so hard to recode all of them to C++ from python.\nI just was working in converting slower functions to C++ like Rasterization",
      "votes": null
    },
    {
      "id": "1076975",
      "postDate": "11/13/2020 05:26:00",
      "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> I implemented the rasterization only in C++ </p>",
      "rawMarkdown": "morizin I implemented the rasterization only in C++",
      "votes": null
    },
    {
      "id": "1076979",
      "postDate": "11/13/2020 05:37:16",
      "content": "<p>you mean the Semantic_rasterizer, Box_rasterizer, MapAPI<br>\nif i have 96 CPUs and 4 GPU and running in C++ optimized l5kit. will there be any speed changes</p>",
      "rawMarkdown": "you mean the Semantic_rasterizer, Box_rasterizer, MapAPI\nif i have 96 CPUs and 4 GPU and running in C++ optimized l5kit. will there be any speed changes",
      "votes": null
    },
    {
      "id": "1077789",
      "postDate": "11/13/2020 23:29:48",
      "content": "<p>If you have lots of memory you can do some online caching and sample reuse </p>\n<pre><code>class ReplayMemory(object):\n    ''' storage class for sample reuse '''\n    def __init__(self, capacity):\n        self.capacity = capacity\n        self.memory = []\n\n    def push(self, item):\n        \"\"\"Saves a transition.\"\"\"\n        self.memory.append(item)\n        if len(self.memory) &gt; self.capacity:\n            old_item = self.memory.pop(0)\n            del old_item\n\n    def sample(self, last_item=False):\n        if last_item:\n            item = self.memory[-1]\n        else:\n            item = random.sample(self.memory, 1)[0]\n        return item\n\n    def __len__(self):\n        return len(self.memory)\n</code></pre>\n<p>Then in your code </p>\n<pre><code>replay_db = ReplayMemory(capacity=100)  # depending on your system memory size\n\ndata_iter = iter(train_dataloader)\nprogress_bar = tqdm(range(len(train_dataloader)))\nfor _ in progress_bar:\n    try:\n        data = next(data_iter)\n    except StopIteration:\n        break\n    replay_db.push(data)\n\n    max_sub_itr = 2  # if you increase it then you risk overfitting\n\n    for sub_itr_idx in range(max_sub_itr):\n        data = replay_db.sample(sub_itr_idx == 0)\n        # train here\n</code></pre>",
      "rawMarkdown": "If you have lots of memory you can do some online caching and sample reuse \n```\nclass ReplayMemory(object):\n    ''' storage class for sample reuse '''\n    def __init__(self, capacity):\n        self.capacity = capacity\n        self.memory = []\n\n    def push(self, item):\n        \"\"\"Saves a transition.\"\"\"\n        self.memory.append(item)\n        if len(self.memory) > self.capacity:\n            old_item = self.memory.pop(0)\n            del old_item\n\n    def sample(self, last_item=False):\n        if last_item:\n            item = self.memory[-1]\n        else:\n            item = random.sample(self.memory, 1)[0]\n        return item\n\n    def __len__(self):\n        return len(self.memory)\n\n```\nThen in your code \n```\nreplay_db = ReplayMemory(capacity=100)  # depending on your system memory size\n\ndata_iter = iter(train_dataloader)\nprogress_bar = tqdm(range(len(train_dataloader)))\nfor _ in progress_bar:\n    try:\n        data = next(data_iter)\n    except StopIteration:\n        break\n    replay_db.push(data)\n\n    max_sub_itr = 2  # if you increase it then you risk overfitting\n\n    for sub_itr_idx in range(max_sub_itr):\n        data = replay_db.sample(sub_itr_idx == 0)\n        # train here\n```",
      "votes": null
    },
    {
      "id": "1078633",
      "postDate": "11/15/2020 05:34:51",
      "content": "<p>Interesting idea! But this seems to give me worse result even with <code>max_sub_itr = 2</code>.</p>",
      "rawMarkdown": "Interesting idea! But this seems to give me worse result even with `max_sub_itr = 2`.",
      "votes": null
    },
    {
      "id": "1079191",
      "postDate": "11/15/2020 18:20:21",
      "content": "<p>Do you mean it is slower?  Is the capacity of the ReplayMemory with 100 elements too much for your system?</p>",
      "rawMarkdown": "Do you mean it is slower?  Is the capacity of the ReplayMemory with 100 elements too much for your system?",
      "votes": null
    },
    {
      "id": "1079352",
      "postDate": "11/16/2020 00:41:20",
      "content": "<p>No, for the same amount of data, it gave me worse LB and validation score than not doing the replay somehow…<br>\nSay I load 300k batches with <code>max_sub_itr = 2</code> (so the model went through 600k batches). the LB score was worse than doing simple 300k batches without replay. Not sure if it already overfit, or it is because 300k is still too little data?<br>\nWhat is the total number of samples or batches that you train?</p>",
      "rawMarkdown": "No, for the same amount of data, it gave me worse LB and validation score than not doing the replay somehow...\nSay I load 300k batches with `max_sub_itr = 2` (so the model went through 600k batches). the LB score was worse than doing simple 300k batches without replay. Not sure if it already overfit, or it is because 300k is still too little data?\nWhat is the total number of samples or batches that you train?",
      "votes": null
    },
    {
      "id": "1079390",
      "postDate": "11/16/2020 02:26:01",
      "content": "<p>Yes, that is understandable. Sample reuse too soon could lead to overfitting but using a batch twice shouldn't be a problem. You have to mitigate the serial correlation problem in the data first.</p>",
      "rawMarkdown": "Yes, that is understandable. Sample reuse too soon could lead to overfitting but using a batch twice shouldn't be a problem. You have to mitigate the serial correlation problem in the data first.",
      "votes": null
    },
    {
      "id": "1079427",
      "postDate": "11/16/2020 03:52:05",
      "content": "<p>I see! Do you filter out the training data that are too similar to each other?</p>",
      "rawMarkdown": "I see! Do you filter out the training data that are too similar to each other?",
      "votes": null
    },
    {
      "id": "1079577",
      "postDate": "11/16/2020 08:39:52",
      "content": "<p>there is a simple trick one can use to move the bottleneck from CPU to GPU (while rasterizing on the fly, no caching, storing to drive or C++). In my experiments, I used 16GB RAM and 1x1080Ti and the GPU was utilized 100%. My feeling is that revealing this trick now would not be wise, so i will wait until the competition ends.</p>\n<p>Anyone else out there who found the magic?</p>\n<p>PS: I hoped this trick could get me a chance to team up with someone who has more resources (1080Ti ain't cut it). It turns out, some think this could be done and a few others think my email is not worth replying. [I didn't share it with anyone, i just wrote i know how to do it - to obey with private sharing rule]</p>",
      "rawMarkdown": "there is a simple trick one can use to move the bottleneck from CPU to GPU (while rasterizing on the fly, no caching, storing to drive or C++). In my experiments, I used 16GB RAM and 1x1080Ti and the GPU was utilized 100%. My feeling is that revealing this trick now would not be wise, so i will wait until the competition ends.\n\nAnyone else out there who found the magic?\n\nPS: I hoped this trick could get me a chance to team up with someone who has more resources (1080Ti ain't cut it). It turns out, some think this could be done and a few others think my email is not worth replying. [I didn't share it with anyone, i just wrote i know how to do it - to obey with private sharing rule]",
      "votes": null
    },
    {
      "id": "1079696",
      "postDate": "11/16/2020 11:54:54",
      "content": "<p>Perhaps, I'll never understand what the downvoters tried to say here. </p>\n<p>I haven't said what the trick is so the trick itself shouldn't be the reason for a downvote. Unless they think it's impossible to do it.</p>\n<p>Could it be the fact that i tried to leverage what i believe to be a good idea for teaming up (many reported rasterization as the main bottleneck)? </p>",
      "rawMarkdown": "Perhaps, I'll never understand what the downvoters tried to say here. \n\nI haven't said what the trick is so the trick itself shouldn't be the reason for a downvote. Unless they think it's impossible to do it.\n\nCould it be the fact that i tried to leverage what i believe to be a good idea for teaming up (many reported rasterization as the main bottleneck)?",
      "votes": null
    },
    {
      "id": "1081043",
      "postDate": "11/16/2020 18:46:02",
      "content": "<p>Looking forward to see what is the trick when the competition ends.</p>",
      "rawMarkdown": "Looking forward to see what is the trick when the competition ends.",
      "votes": null
    },
    {
      "id": "1082478",
      "postDate": "11/17/2020 23:55:31",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  are you facing the issue of overfitting </p>",
      "rawMarkdown": "pestipeti  are you facing the issue of overfitting",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1075712,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "11/11/2020 23:05:11",
      "content": "<p>Thank you! This is amazing! I was wondering if your current train stats are based on pre-generate cached data or not?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075717,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/11/2020 23:09:37",
          "content": "<p>We cache data, but without the rasterized image.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075739,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "11/11/2020 23:28:14",
          "content": "<p>Could you explain a bit what's the difference between caching data without and caching data with rasterized image? (Or a link to another discussion post that contains this info). Thank you very much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075751,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/11/2020 23:40:23",
          "content": "<p>Caching with images is simple (see <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this post</a>). Its drawback is the size. You have to compress, but in that case, the loading is slow.</p>\n<p>You can cache without the image, but it is harder to implement (you need to modify the l5kit). With this solution, you can save the l5kit's data-loading/zarr-reading time, but you still have to rasterize. <br>\nFor us, this one is optimal, but I implement the rasterization in C++/CUDA. It could work with CPU rasterization, but I haven't tested it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076320,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "11/12/2020 12:49:51",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1075769,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/12/2020 00:17:47",
      "content": "<p>Do you have performance numbers from eff-b0? I haven't explore other backbones at all. Not sure if it is worth my time or just wasted compute. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1076217,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/12/2020 10:42:09",
          "content": "<p>I tried eff-b0 but it underperformed resnet considerably</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076321,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "11/12/2020 12:50:47",
          "content": "<p>May I ask if your current score is based on resnet or not? <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076576,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/12/2020 16:58:27",
          "content": "<p>Yes, it's based on resnet</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076664,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "11/12/2020 18:15:57",
          "content": "<p>what l5kit version are you using</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076674,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "11/12/2020 18:23:06",
          "content": "<p>The latest one (think it's 1.1.0)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1075887,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/12/2020 04:02:50",
      "content": "<p>Sounds like this is more a competition to help Lyft rewrite and optimize their l5kit package?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075889,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/12/2020 04:10:02",
          "content": "<p>That does seem to be the primary value lyft is getting out of this. A big audit of their package by dangling a competition in front of us to motivate us to interact with it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1076663,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "11/12/2020 18:15:34",
      "content": "<p>From my experiments, I understand that :</p>\n<ul>\n<li>Larger network will improve in a small amount but improves forever</li>\n<li>But in Smaller networks like resnet18 it stops at point and lags behind it</li>\n</ul>\n<p>are my experiments true or not?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1076943,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/13/2020 04:09:04",
          "content": "<p>I am not sure larger network is better. Looks like you can easily overfit in the large network. I have tried resnet 18, 34, and 50. The resnet50 has the best train score but the worse LB and validation score. On the other hand, resnet 18 has almost no gap between train, val, and LB score while give me the best score on LB so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076946,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "11/13/2020 04:11:35",
          "content": "<p>It's only my experiments.<br>\nThanks, i felt that large network are easy to fit in train set but overfit in LB or Validation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1076961,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "11/13/2020 04:41:39",
      "content": "<p>That is amazing <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>. I think it was so hard to recode all of them to C++ from python.<br>\nI just was working in converting slower functions to C++ like Rasterization</p>",
      "votes": null,
      "replies": [
        {
          "id": 1076975,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/13/2020 05:26:00",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> I implemented the rasterization only in C++ </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076979,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "11/13/2020 05:37:16",
          "content": "<p>you mean the Semantic_rasterizer, Box_rasterizer, MapAPI<br>\nif i have 96 CPUs and 4 GPU and running in C++ optimized l5kit. will there be any speed changes</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1077789,
      "author_name": "ggrizzly",
      "author_url": "",
      "post_date": "11/13/2020 23:29:48",
      "content": "<p>If you have lots of memory you can do some online caching and sample reuse </p>\n<pre><code>class ReplayMemory(object):\n    ''' storage class for sample reuse '''\n    def __init__(self, capacity):\n        self.capacity = capacity\n        self.memory = []\n\n    def push(self, item):\n        \"\"\"Saves a transition.\"\"\"\n        self.memory.append(item)\n        if len(self.memory) &gt; self.capacity:\n            old_item = self.memory.pop(0)\n            del old_item\n\n    def sample(self, last_item=False):\n        if last_item:\n            item = self.memory[-1]\n        else:\n            item = random.sample(self.memory, 1)[0]\n        return item\n\n    def __len__(self):\n        return len(self.memory)\n</code></pre>\n<p>Then in your code </p>\n<pre><code>replay_db = ReplayMemory(capacity=100)  # depending on your system memory size\n\ndata_iter = iter(train_dataloader)\nprogress_bar = tqdm(range(len(train_dataloader)))\nfor _ in progress_bar:\n    try:\n        data = next(data_iter)\n    except StopIteration:\n        break\n    replay_db.push(data)\n\n    max_sub_itr = 2  # if you increase it then you risk overfitting\n\n    for sub_itr_idx in range(max_sub_itr):\n        data = replay_db.sample(sub_itr_idx == 0)\n        # train here\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1078633,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/15/2020 05:34:51",
          "content": "<p>Interesting idea! But this seems to give me worse result even with <code>max_sub_itr = 2</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079191,
          "author_name": "ggrizzly",
          "author_url": "",
          "post_date": "11/15/2020 18:20:21",
          "content": "<p>Do you mean it is slower?  Is the capacity of the ReplayMemory with 100 elements too much for your system?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079352,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/16/2020 00:41:20",
          "content": "<p>No, for the same amount of data, it gave me worse LB and validation score than not doing the replay somehow…<br>\nSay I load 300k batches with <code>max_sub_itr = 2</code> (so the model went through 600k batches). the LB score was worse than doing simple 300k batches without replay. Not sure if it already overfit, or it is because 300k is still too little data?<br>\nWhat is the total number of samples or batches that you train?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079390,
          "author_name": "ggrizzly",
          "author_url": "",
          "post_date": "11/16/2020 02:26:01",
          "content": "<p>Yes, that is understandable. Sample reuse too soon could lead to overfitting but using a batch twice shouldn't be a problem. You have to mitigate the serial correlation problem in the data first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079427,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/16/2020 03:52:05",
          "content": "<p>I see! Do you filter out the training data that are too similar to each other?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1079577,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "11/16/2020 08:39:52",
      "content": "<p>there is a simple trick one can use to move the bottleneck from CPU to GPU (while rasterizing on the fly, no caching, storing to drive or C++). In my experiments, I used 16GB RAM and 1x1080Ti and the GPU was utilized 100%. My feeling is that revealing this trick now would not be wise, so i will wait until the competition ends.</p>\n<p>Anyone else out there who found the magic?</p>\n<p>PS: I hoped this trick could get me a chance to team up with someone who has more resources (1080Ti ain't cut it). It turns out, some think this could be done and a few others think my email is not worth replying. [I didn't share it with anyone, i just wrote i know how to do it - to obey with private sharing rule]</p>",
      "votes": null,
      "replies": [
        {
          "id": 1079696,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "11/16/2020 11:54:54",
          "content": "<p>Perhaps, I'll never understand what the downvoters tried to say here. </p>\n<p>I haven't said what the trick is so the trick itself shouldn't be the reason for a downvote. Unless they think it's impossible to do it.</p>\n<p>Could it be the fact that i tried to leverage what i believe to be a good idea for teaming up (many reported rasterization as the main bottleneck)? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1081043,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/16/2020 18:46:02",
          "content": "<p>Looking forward to see what is the trick when the competition ends.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1082478,
      "author_name": "",
      "author_url": "",
      "post_date": "11/17/2020 23:55:31",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  are you facing the issue of overfitting </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1075680": "As you probably know, one of the difficulties in this competition is speed. You can improve it if you optimize the data-loading and the rasterization process. In the past weeks/months, I've tried to improve the training speed as much as possible. In this post, I'd like to share some details. \n\nLeave me a comment if you have any other idea/technique.\n\n## Optimizations\n\n### Basic\n\nFor these tricks, you don't have to modify the l5kit.\n- **Pre-generate training dataset (partial)** For more details, take a look at [this post](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637)\n- **Pre-generate validation/test dataset (full)**\n- **Implement a faster metric calculation method**. The official metric saves and re-loads your predictions. It is faster if you keep everything in memory.\n- **Validate less frequently**. Reduce your validation rounds. I validate after 5000 iterations for new experiments, but during a long training process, I reduce it to every 25000.\n- **Validate using part of the official validation set.** Make sure that the subset's distribution does not change!\n- **Validate on public LB.** Fortunately, our datasets are stable. You can skip the local validation.  Instead of predicting the validation set, predict and submit the test set during training. You only need to \"validate\" on the test set 5 times per day!\n- **Reduce the height of the image** The agents are always moving towards positive-x direction. You can reduce the height of the input images.\n\n### Intermediate\n\nFor a bit more speed, you'll have to modify the l5kit.\n\n- **Optimize the box rasterizer.** Try to simplify the `draw_boxes` method - especially the matrix multiplications, point transformations. Try to eliminate the costly Numpy calls like `vstack`, `transpose`, `concatenate`, etc. Take a look at the `cv2` drawing calls as well. Get rid of everything that does not add information to the training.\n- **Optimize the semantic rasterizer.** You can gain 6-8% if you optimize this. Same as above. Hint: Check out my [pull-request](https://github.com/lyft/l5kit/pull/140) (pending).\n- **Data caching.** Besides the rasterization, the other bottleneck is the data loading. L5Kit has to do a lot of work to load and prepare the data for the rasterizers. You can speed things up if you save the prepared data inside l5kit (`agent_sampling`; without the rasterized images).\n\n\n\n### Advanced\n\nIf you want significant speed improvement, you'll have to rewrite the rasterization in C++ and CUDA. And you'll need a few other tricks as well. We'll share the details, my C++/CUDA implementation, and our training method after the competition.\n\n## Training time\n\nThere are lots of things that could affect the training time. If you optimize your training process, you can speed up the training process by 2.0-2.5-3.0x\n\nStats of my latest experiment:\n\n- 16,000,000 samples (250,000 iterations; 64 batch)\n- Validating: 100% of the validation set after every 5000 iterations\n- Running time: 25 hours (~23 hours with 1/5 of validation)\n- Image size: Bigger than I used in the experiments below\n- Model: Bigger than in the experiments below\n- Public LB: 16.180\n\n\n*I used 20 CPU cores, I have 128Gb RAM, 1Tb SSD, and an RTX-2080ti*\n\n## Results\n\nFor all of the experiments below I used:\n\n- 300x300px input size\n- 0.5 pixel size\n- 10 historical frames\n- 32 batch\n- 1000 iterations\n\n| Method                           | Net       | \\# of workers | \\# of samples | Storage size | Training time | it/sec |\n| -------------------------------- | --------- | ------------- | ------------- | ------------ | ------------- | ------ |\n| l5kit                            | Resnet-18 | 20            | 32000         | -            | 4:48          | 3.47   |\n| optimized l5kit                  | Resnet-18 | 20            | 32000         | -            | 3:35\\*        | 4.65\\* |\n| Pre-generated (w images) | Resnet-18 | 20            | 32000         | 1.9 Gb       | 3:26          | 4.85   |\n| advanced                         | Resnet-18 | 20            | 32000         | 0.36 Gb      | 1:42          | 9.81   |\n\n\\* Estimated values (I used different settings for this experiment).\n\n\n\n| Method   | Net       | \\# of workers | Storage | Training time | it/sec |\n| -------- | --------- | ------------- | ------- | ------------- | ------ |\n| l5kit    | Resnet-18 | 20            | -       | 4:48          | 3.47   |\n| advanced | Resnet-18 | 20            | 0.36 Gb | 1:42          | 9.81   |\n| l5kit    | Resnet-18 | 4             | -       | 6:45          | 2.46   |\n| advanced | Resnet-18 | 4             | 0.36 Gb | 1:43          | 9.79   |\n| l5kit    | Eff-B0    | 20            | -       | 4:45          | 3.51   |\n| advanced | Eff-B0    | 20            | 0.36 Gb | 2:43          | 6.13   |\n| l5kit    | Eff-B0    | 4             | -       | 6:48          | 2.44   |\n| advanced | Eff-B0    | 4             | 0.36 Gb | 2:43          | 6.13   |",
    "1075712": "Thank you! This is amazing! I was wondering if your current train stats are based on pre-generate cached data or not?",
    "1075717": "We cache data, but without the rasterized image.",
    "1075739": "Could you explain a bit what's the difference between caching data without and caching data with rasterized image? (Or a link to another discussion post that contains this info). Thank you very much!",
    "1075751": "Caching with images is simple (see [this post](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637)). Its drawback is the size. You have to compress, but in that case, the loading is slow.\n\nYou can cache without the image, but it is harder to implement (you need to modify the l5kit). With this solution, you can save the l5kit's data-loading/zarr-reading time, but you still have to rasterize. \nFor us, this one is optimal, but I implement the rasterization in C++/CUDA. It could work with CPU rasterization, but I haven't tested it.",
    "1075769": "Do you have performance numbers from eff-b0? I haven't explore other backbones at all. Not sure if it is worth my time or just wasted compute.",
    "1075887": "Sounds like this is more a competition to help Lyft rewrite and optimize their l5kit package?",
    "1075889": "That does seem to be the primary value lyft is getting out of this. A big audit of their package by dangling a competition in front of us to motivate us to interact with it.",
    "1076217": "I tried eff-b0 but it underperformed resnet considerably",
    "1076320": "Thank you!",
    "1076321": "May I ask if your current score is based on resnet or not? @fergusoci",
    "1076576": "Yes, it's based on resnet",
    "1076663": "From my experiments, I understand that :\n- Larger network will improve in a small amount but improves forever\n- But in Smaller networks like resnet18 it stops at point and lags behind it\n\nare my experiments true or not?",
    "1076664": "what l5kit version are you using",
    "1076674": "The latest one (think it's 1.1.0)",
    "1076943": "I am not sure larger network is better. Looks like you can easily overfit in the large network. I have tried resnet 18, 34, and 50. The resnet50 has the best train score but the worse LB and validation score. On the other hand, resnet 18 has almost no gap between train, val, and LB score while give me the best score on LB so far.",
    "1076946": "It's only my experiments.\nThanks, i felt that large network are easy to fit in train set but overfit in LB or Validation",
    "1076961": "That is amazing @pestipeti. I think it was so hard to recode all of them to C++ from python.\nI just was working in converting slower functions to C++ like Rasterization",
    "1076975": "morizin I implemented the rasterization only in C++",
    "1076979": "you mean the Semantic_rasterizer, Box_rasterizer, MapAPI\nif i have 96 CPUs and 4 GPU and running in C++ optimized l5kit. will there be any speed changes",
    "1077789": "If you have lots of memory you can do some online caching and sample reuse \n```\nclass ReplayMemory(object):\n    ''' storage class for sample reuse '''\n    def __init__(self, capacity):\n        self.capacity = capacity\n        self.memory = []\n\n    def push(self, item):\n        \"\"\"Saves a transition.\"\"\"\n        self.memory.append(item)\n        if len(self.memory) > self.capacity:\n            old_item = self.memory.pop(0)\n            del old_item\n\n    def sample(self, last_item=False):\n        if last_item:\n            item = self.memory[-1]\n        else:\n            item = random.sample(self.memory, 1)[0]\n        return item\n\n    def __len__(self):\n        return len(self.memory)\n\n```\nThen in your code \n```\nreplay_db = ReplayMemory(capacity=100)  # depending on your system memory size\n\ndata_iter = iter(train_dataloader)\nprogress_bar = tqdm(range(len(train_dataloader)))\nfor _ in progress_bar:\n    try:\n        data = next(data_iter)\n    except StopIteration:\n        break\n    replay_db.push(data)\n\n    max_sub_itr = 2  # if you increase it then you risk overfitting\n\n    for sub_itr_idx in range(max_sub_itr):\n        data = replay_db.sample(sub_itr_idx == 0)\n        # train here\n```",
    "1078633": "Interesting idea! But this seems to give me worse result even with `max_sub_itr = 2`.",
    "1079191": "Do you mean it is slower?  Is the capacity of the ReplayMemory with 100 elements too much for your system?",
    "1079352": "No, for the same amount of data, it gave me worse LB and validation score than not doing the replay somehow...\nSay I load 300k batches with `max_sub_itr = 2` (so the model went through 600k batches). the LB score was worse than doing simple 300k batches without replay. Not sure if it already overfit, or it is because 300k is still too little data?\nWhat is the total number of samples or batches that you train?",
    "1079390": "Yes, that is understandable. Sample reuse too soon could lead to overfitting but using a batch twice shouldn't be a problem. You have to mitigate the serial correlation problem in the data first.",
    "1079427": "I see! Do you filter out the training data that are too similar to each other?",
    "1079577": "there is a simple trick one can use to move the bottleneck from CPU to GPU (while rasterizing on the fly, no caching, storing to drive or C++). In my experiments, I used 16GB RAM and 1x1080Ti and the GPU was utilized 100%. My feeling is that revealing this trick now would not be wise, so i will wait until the competition ends.\n\nAnyone else out there who found the magic?\n\nPS: I hoped this trick could get me a chance to team up with someone who has more resources (1080Ti ain't cut it). It turns out, some think this could be done and a few others think my email is not worth replying. [I didn't share it with anyone, i just wrote i know how to do it - to obey with private sharing rule]",
    "1079696": "Perhaps, I'll never understand what the downvoters tried to say here. \n\nI haven't said what the trick is so the trick itself shouldn't be the reason for a downvote. Unless they think it's impossible to do it.\n\nCould it be the fact that i tried to leverage what i believe to be a good idea for teaming up (many reported rasterization as the main bottleneck)?",
    "1081043": "Looking forward to see what is the trick when the competition ends.",
    "1082478": "pestipeti  are you facing the issue of overfitting"
  },
  "source": "meta"
}