{
  "id": 182787,
  "title": "Dataset Size / Model Performance Tradeoff",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/182787",
  "author_name": "fergusoci",
  "post_date": "2020-09-14T10:54:19.362000",
  "votes": 33,
  "comment_count": 35,
  "views": 0,
  "content": "<p>From playing around it seems like one of the largest issues we’ll contend with is training time. The possibility of using train_full.zarr seems fanciful to me right now!  That said, using the full dataset seems like it will be important for performance. </p>\n<p>Here are the results of some experiments that I have run so far:</p>\n<ul>\n<li><p>resnet18 backbone, 6M samples (approx 26% of train.zarr), train time ~20 hrs: LB 35.9</p></li>\n<li><p>resnet18 backbone, 18M samples (approx 80% of train.zarr), train time ~60 hrs: LB 29.8</p></li>\n</ul>\n<p>This is using input image sizes of 320, training locally. I think going with a larger input size will help, but this will push training times to a point where it’s taking a week to cycle once through the small training dataset.</p>\n<p>Doing well in this competition is clearing going to require some strategising around input size/history_num_frames/training time trade offs. </p>\n<p>My current strategy is to test ideas with a small image size and approx 13% of train.zarr, then run the ideas that work on a more extended dataset. It seems that this smaller dataset is stable enough to determine relative performance.</p>\n<p>Does anybody have an idea of the impact of using more of the training data - either more than one cycle through train.zarr, or of using train_full.zarr?</p>",
  "messages": [
    {
      "id": 1009894,
      "postDate": "2020-09-14T10:54:19.363Z",
      "content": "<p>From playing around it seems like one of the largest issues we’ll contend with is training time. The possibility of using train_full.zarr seems fanciful to me right now!  That said, using the full dataset seems like it will be important for performance. </p>\n<p>Here are the results of some experiments that I have run so far:</p>\n<ul>\n<li><p>resnet18 backbone, 6M samples (approx 26% of train.zarr), train time ~20 hrs: LB 35.9</p></li>\n<li><p>resnet18 backbone, 18M samples (approx 80% of train.zarr), train time ~60 hrs: LB 29.8</p></li>\n</ul>\n<p>This is using input image sizes of 320, training locally. I think going with a larger input size will help, but this will push training times to a point where it’s taking a week to cycle once through the small training dataset.</p>\n<p>Doing well in this competition is clearing going to require some strategising around input size/history_num_frames/training time trade offs. </p>\n<p>My current strategy is to test ideas with a small image size and approx 13% of train.zarr, then run the ideas that work on a more extended dataset. It seems that this smaller dataset is stable enough to determine relative performance.</p>\n<p>Does anybody have an idea of the impact of using more of the training data - either more than one cycle through train.zarr, or of using train_full.zarr?</p>",
      "rawMarkdown": "From playing around it seems like one of the largest issues we’ll contend with is training time. The possibility of using train_full.zarr seems fanciful to me right now!  That said, using the full dataset seems like it will be important for performance. \n\nHere are the results of some experiments that I have run so far:\n\n- resnet18 backbone, 6M samples (approx 26% of train.zarr), train time ~20 hrs: LB 35.9\n\n- resnet18 backbone, 18M samples (approx 80% of train.zarr), train time ~60 hrs: LB 29.8\n\nThis is using input image sizes of 320, training locally. I think going with a larger input size will help, but this will push training times to a point where it’s taking a week to cycle once through the small training dataset.\n\nDoing well in this competition is clearing going to require some strategising around input size/history_num_frames/training time trade offs. \n\nMy current strategy is to test ideas with a small image size and approx 13% of train.zarr, then run the ideas that work on a more extended dataset. It seems that this smaller dataset is stable enough to determine relative performance.\n\nDoes anybody have an idea of the impact of using more of the training data - either more than one cycle through train.zarr, or of using train_full.zarr?",
      "votes": 33
    },
    {
      "id": 1069257,
      "postDate": "2020-11-04T09:02:05.727Z",
      "content": "<p>we have some much data (e.g. train_full). we could capitalize for stacking (i.e. build second or third layer simple blending/ensembling model)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F364aacbfc23535361367c549b3567df5%2FSelection_091.png?generation=1604480959249507&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "we have some much data (e.g. train\\_full). we could capitalize for stacking (i.e. build second or third layer simple blending/ensembling model)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F364aacbfc23535361367c549b3567df5%2FSelection_091.png?generation=1604480959249507&alt=media)",
      "votes": 3
    },
    {
      "id": 1043371,
      "postDate": "2020-10-08T22:59:54.107Z",
      "content": "<p>I'd expect it would depend quite a bit on the optimiser and the LR schedule as well. I'm currently training on 25% of train.zarr to have faster experiments feedback, seems to be good enough to compare models/approaches</p>",
      "rawMarkdown": "I'd expect it would depend quite a bit on the optimiser and the LR schedule as well. I'm currently training on 25% of train.zarr to have faster experiments feedback, seems to be good enough to compare models/approaches",
      "votes": 3,
      "replies": [
        {
          "id": 1043715,
          "postDate": "2020-10-09T07:29:13.210Z",
          "content": "<p>Yep, if there's one thing to be said about the data it seems to be very stable. You can run on small training sets, small raster sizes, small history_num_frames and the result preferences seem to consistently hold for larger versions of each.</p>",
          "rawMarkdown": "Yep, if there's one thing to be said about the data it seems to be very stable. You can run on small training sets, small raster sizes, small history_num_frames and the result preferences seem to consistently hold for larger versions of each.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1043946,
      "postDate": "2020-10-09T10:58:39.693Z",
      "content": "<p>I already experiment with some models and also I'm expecting bigger models to get more accuracy.<br>\nbut it's failed.(ResNet18 &gt; ResNet34 &gt; ResNet50). I don't use efficientnet until now.(will experiment)</p>",
      "rawMarkdown": "I already experiment with some models and also I'm expecting bigger models to get more accuracy.\nbut it's failed.(ResNet18 > ResNet34 > ResNet50). I don't use efficientnet until now.(will experiment)",
      "votes": 4,
      "replies": [
        {
          "id": 1043969,
          "postDate": "2020-10-09T11:26:54.773Z",
          "content": "<p>It behaves against the fundamentals of Deep-Learning, but it is very interesting to spectate such relations and it would be way more interesting to understand why it behaves like that. </p>",
          "rawMarkdown": "It behaves against the fundamentals of Deep-Learning, but it is very interesting to spectate such relations and it would be way more interesting to understand why it behaves like that. "
        },
        {
          "id": 1066009,
          "postDate": "2020-11-01T07:25:02.017Z",
          "content": "<p>I think it is because bigger models learn well, but since a lot of data is in some way duplicated, they are learning to the point of \"memorisation\" - Hence it looks like a case of overfitting. </p>",
          "rawMarkdown": "I think it is because bigger models learn well, but since a lot of data is in some way duplicated, they are learning to the point of \"memorisation\" - Hence it looks like a case of overfitting. "
        }
      ]
    },
    {
      "id": 1010130,
      "postDate": "2020-09-14T14:22:38.393Z",
      "content": "<p>Since we are talking. how far left and down you think we should go to cover this dataset? 😄 Any dense layer will make mathematically more difficult. or do you think others will have much better perf or behaviour?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F344232%2F051fecd3130949270d56134c6db78c54%2FScreenshot%202020-09-11%20110341.png?generation=1600093059987424&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Since we are talking. how far left and down you think we should go to cover this dataset? 😄 Any dense layer will make mathematically more difficult. or do you think others will have much better perf or behaviour?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F344232%2F051fecd3130949270d56134c6db78c54%2FScreenshot%202020-09-11%20110341.png?generation=1600093059987424&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1010164,
          "postDate": "2020-09-14T14:51:53.243Z",
          "content": "<p>I haven't experimented with it, but you can probably implement some complex models without affecting your runtime too much, given that rasterization is the bottleneck. i.e. if you want to cover the dataset, choosing a slightly more complex model isn't going to stop you doing that. </p>\n<p>The only experimenting around backbone that I've done was with Efficientnet - which others have noted <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178245\" target=\"_blank\">underperformed</a>, and I found the same.</p>",
          "rawMarkdown": "I haven't experimented with it, but you can probably implement some complex models without affecting your runtime too much, given that rasterization is the bottleneck. i.e. if you want to cover the dataset, choosing a slightly more complex model isn't going to stop you doing that. \n\nThe only experimenting around backbone that I've done was with Efficientnet - which others have noted [underperformed](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178245), and I found the same.",
          "votes": 3
        },
        {
          "id": 1010177,
          "postDate": "2020-09-14T15:02:13.670Z",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> Ja! I think because of lot of parameters. Its difficult to say which model is better. I have run into making lot of premature conclusion. Probably my inference about dense layer is one such.    </p>",
          "rawMarkdown": "@fergusoci Ja! I think because of lot of parameters. Its difficult to say which model is better. I have run into making lot of premature conclusion. Probably my inference about dense layer is one such.    "
        },
        {
          "id": 1043351,
          "postDate": "2020-10-08T22:31:45.990Z",
          "content": "<p>I had similar results with small Efficientnets comparing to resnet18, much worse with a few larger models I tried</p>",
          "rawMarkdown": "I had similar results with small Efficientnets comparing to resnet18, much worse with a few larger models I tried",
          "votes": 1
        }
      ]
    },
    {
      "id": 1013113,
      "postDate": "2020-09-16T14:07:07.903Z",
      "content": "<blockquote>\n  <p>it was more important to try to cover the full dataset than have bigger image resolution</p>\n</blockquote>\n<p>I got only 31 after training 128*128 for 1 epoch (one cycle through train.zarr). I think image size matters.</p>",
      "rawMarkdown": "> it was more important to try to cover the full dataset than have bigger image resolution\n\nI got only 31 after training 128*128 for 1 epoch (one cycle through train.zarr). I think image size matters.",
      "votes": 2,
      "replies": [
        {
          "id": 1013270,
          "postDate": "2020-09-16T15:42:32.150Z",
          "content": "<p>31 seems like a great score for 128*128! For most people getting one cycle through train.zarr is tricky as it takes so long. How long did it take you to run? Have you experimented with &gt; 1 epoch, and if so, did it add much value?</p>",
          "rawMarkdown": "31 seems like a great score for 128*128! For most people getting one cycle through train.zarr is tricky as it takes so long. How long did it take you to run? Have you experimented with > 1 epoch, and if so, did it add much value?",
          "votes": 2
        },
        {
          "id": 1013283,
          "postDate": "2020-09-16T15:53:31.717Z",
          "content": "<p>I am also curious to know. I had been experimenting with 180 * 180 and yesterday I chose to exp with 300 * 300 and my iteration/s went down by a factor of 4. and thats the only parameter i changed. Is it only me or anyone else observing the same?</p>",
          "rawMarkdown": "I am also curious to know. I had been experimenting with 180 * 180 and yesterday I chose to exp with 300 * 300 and my iteration/s went down by a factor of 4. and thats the only parameter i changed. Is it only me or anyone else observing the same?"
        },
        {
          "id": 1013293,
          "postDate": "2020-09-16T16:07:05.727Z",
          "content": "<ul>\n<li>It took 2-3 days. </li>\n<li>Overfitting after ~ 1.2 epochs.</li>\n</ul>",
          "rawMarkdown": "- It took 2-3 days. \n- Overfitting after ~ 1.2 epochs.",
          "votes": 4
        },
        {
          "id": 1013433,
          "postDate": "2020-09-16T17:32:34.997Z",
          "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> Good to know, thank you!</p>",
          "rawMarkdown": "@zzy990106 Good to know, thank you!"
        }
      ]
    },
    {
      "id": 1010561,
      "postDate": "2020-09-14T21:29:20.720Z",
      "content": "<p>Are you able to run that many samples continuously? I have been running into a memory leak on the dataloader. Curious what version of l5kit and pytorch you might be running. </p>",
      "rawMarkdown": "Are you able to run that many samples continuously? I have been running into a memory leak on the dataloader. Curious what version of l5kit and pytorch you might be running. ",
      "votes": 2,
      "replies": [
        {
          "id": 1010950,
          "postDate": "2020-09-15T06:48:05.063Z",
          "content": "<p>Yes, there doesn't seem to be any problems there. I saw your thread but was scratching my head a bit, as it doesn't appear to affect me.</p>",
          "rawMarkdown": "Yes, there doesn't seem to be any problems there. I saw your thread but was scratching my head a bit, as it doesn't appear to affect me.",
          "votes": 3
        },
        {
          "id": 1011038,
          "postDate": "2020-09-15T08:03:00.340Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> I had a look at your code. I dont see issue over there. Just to share some input. I run L5kit 1.0.6 and pytorch 1.5.1. Even i dont see any problem while loading or iterating. I have also tried to load train_full. Which by the way has 191M(10x more than train) samples and the usage shot up only for few sec. I did encounter the problem you mentioned a week or two back not now. But that was only during creation of chopped dataset. so my recommendation would be obvious ones. Use  <code>import gc</code> <code>gc.collect()</code> once on top. 34 batch and 4 worker should work on kaggle. I have even tried with 64 and 4. Maybe play around and see reaction. just for debug purpose set <code>shuffle</code>to <code>false</code>and see reaction. And if the problem still persist try this thread once. <code>https://discuss.pytorch.org/t/guidelines-for-assigning-num-workers-to-dataloader/813/3</code> </p>",
          "rawMarkdown": "@ryches I had a look at your code. I dont see issue over there. Just to share some input. I run L5kit 1.0.6 and pytorch 1.5.1. Even i dont see any problem while loading or iterating. I have also tried to load train_full. Which by the way has 191M(10x more than train) samples and the usage shot up only for few sec. I did encounter the problem you mentioned a week or two back not now. But that was only during creation of chopped dataset. so my recommendation would be obvious ones. Use  `import gc` `gc.collect()` once on top. 34 batch and 4 worker should work on kaggle. I have even tried with 64 and 4. Maybe play around and see reaction. just for debug purpose set `shuffle `to `false `and see reaction. And if the problem still persist try this thread once. `https://discuss.pytorch.org/t/guidelines-for-assigning-num-workers-to-dataloader/813/3` ",
          "votes": 1
        },
        {
          "id": 1011044,
          "postDate": "2020-09-15T08:05:38.940Z",
          "content": "<p>odd because I am running the same l5kit and pytorch versions. Unsure of what else might be causing that. I've looked pretty deeply at the various dataloader parameters and run gc.collect and cleared torch cache and it seems to be something accumulating inside of the dataloader itself so gc and other methods dont clear it. </p>",
          "rawMarkdown": "odd because I am running the same l5kit and pytorch versions. Unsure of what else might be causing that. I've looked pretty deeply at the various dataloader parameters and run gc.collect and cleared torch cache and it seems to be something accumulating inside of the dataloader itself so gc and other methods dont clear it. "
        },
        {
          "id": 1011773,
          "postDate": "2020-09-15T17:14:57.450Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Not sure if it helps, but one thing that I'm doing slightly differently to what you have is wrapping the rasterizer dataset. Maybe this leads to a different interaction with the dataloader.</p>\n<p>Sample shell wrapper:</p>\n<p>class MotionPredictDataset(Dataset):</p>\n<pre><code>def __init__(self,\n             cfg,\n             str_loader='train_data_loader',\n             fn_rasterizer=build_rasterizer,\n             test_mask_path=os.path.join(BASE_DIR, 'scenes/mask.npz')):\n\n    self.cfg = cfg\n    self.str_loader = str_loader\n    self.fn_rasterizer = fn_rasterizer\n    self.test_mask_path = test_mask_path\n\n    self.setup()\n\ndef setup(self):\n\n    self.dm = LocalDataManager(None)\n    self.rasterizer = self.fn_rasterizer(self.cfg, self.dm)\n    self.data_zarr = ChunkedDataset(self.dm.require(self.cfg[self.str_loader][\"key\"])).open()\n\n    if self.str_loader == 'test_data_loader':\n        test_mask = np.load(self.test_mask_path)[\"arr_0\"]\n        self.ds = AgentDataset(self.cfg, self.data_zarr, self.rasterizer, agents_mask=test_mask)\n    else:\n        self.ds = AgentDataset(self.cfg, self.data_zarr, self.rasterizer)\n\ndef __getitem__(self, index):\n    return self.ds[index]\n\ndef __len__(self):\n    return self.ds\n</code></pre>",
          "rawMarkdown": "@ryches Not sure if it helps, but one thing that I'm doing slightly differently to what you have is wrapping the rasterizer dataset. Maybe this leads to a different interaction with the dataloader.\n\nSample shell wrapper:\n\nclass MotionPredictDataset(Dataset):\n\n    def __init__(self,\n                 cfg,\n                 str_loader='train_data_loader',\n                 fn_rasterizer=build_rasterizer,\n                 test_mask_path=os.path.join(BASE_DIR, 'scenes/mask.npz')):\n\n        self.cfg = cfg\n        self.str_loader = str_loader\n        self.fn_rasterizer = fn_rasterizer\n        self.test_mask_path = test_mask_path\n        \n        self.setup()\n\n    def setup(self):\n\n        self.dm = LocalDataManager(None)\n        self.rasterizer = self.fn_rasterizer(self.cfg, self.dm)\n        self.data_zarr = ChunkedDataset(self.dm.require(self.cfg[self.str_loader][\"key\"])).open()\n        \n        if self.str_loader == 'test_data_loader':\n            test_mask = np.load(self.test_mask_path)[\"arr_0\"]\n            self.ds = AgentDataset(self.cfg, self.data_zarr, self.rasterizer, agents_mask=test_mask)\n        else:\n            self.ds = AgentDataset(self.cfg, self.data_zarr, self.rasterizer)\n\n    def __getitem__(self, index):\n        return self.ds[index]\n\n    def __len__(self):\n        return self.ds\n"
        }
      ]
    },
    {
      "id": 1010322,
      "postDate": "2020-09-14T17:05:59.197Z",
      "content": "<p>Have you considered pre-rasterizing your images? </p>\n<p>I mean, if rasterization is your bottleneck, why not try and run through the rasterization process once and then saving just the images. <br>\nThen you can just read in the images. </p>\n<p>Admittedly, this might take up quite a lot of space. But you are now swapping cpu cycles for diskspace. </p>\n<p>I'll have to check how much space this would take up, but it might be doable. Might be worth it to store it in a compressed format like png, especially if you are using the semantic rasterizer. </p>\n<p>I haven't actually tried this, has anyone? </p>",
      "rawMarkdown": "Have you considered pre-rasterizing your images? \n\nI mean, if rasterization is your bottleneck, why not try and run through the rasterization process once and then saving just the images. \nThen you can just read in the images. \n\nAdmittedly, this might take up quite a lot of space. But you are now swapping cpu cycles for diskspace. \n\nI'll have to check how much space this would take up, but it might be doable. Might be worth it to store it in a compressed format like png, especially if you are using the semantic rasterizer. \n\nI haven't actually tried this, has anyone? ",
      "votes": 2,
      "replies": [
        {
          "id": 1010332,
          "postDate": "2020-09-14T17:15:43.060Z",
          "content": "<p><a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> Take a look at <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this thread</a></p>",
          "rawMarkdown": "@fnands Take a look at [this thread](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637)",
          "votes": 3
        },
        {
          "id": 1010388,
          "postDate": "2020-09-14T17:54:29.833Z",
          "content": "<p>Ah thanks, makes sense that someone already did it. That's quite a lot of space, and a decent increase in speed.</p>\n<p>Should be better if you store compressed, seeing as the semantic images are mostly large blocks of white, but then you lose a bit again in the decoding. </p>",
          "rawMarkdown": "Ah thanks, makes sense that someone already did it. That's quite a lot of space, and a decent increase in speed.\n\nShould be better if you store compressed, seeing as the semantic images are mostly large blocks of white, but then you lose a bit again in the decoding. \n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1009945,
      "postDate": "2020-09-14T11:46:52.430Z",
      "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> For me it took 70hrs for 3M. Think it depends on lot of factors. For strategy with image size or resolution. You can refer to <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> experiments. Here is the link for it. <a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/tuckerarrants/lyftpretrainedmodels\" target=\"_blank\">https://www.kaggle.com/tuckerarrants/lyftpretrainedmodels</a>    . I believe he has done these experiment for unimodel output. You can transfer the idea maybe.  </p>",
      "rawMarkdown": "@fergusoci For me it took 70hrs for 3M. Think it depends on lot of factors. For strategy with image size or resolution. You can refer to @tuckerarrants experiments. Here is the link for it. [https://www.kaggle.com/tuckerarrants/lyftpretrainedmodels   ](url) . I believe he has done these experiment for unimodel output. You can transfer the idea maybe.  ",
      "votes": 2,
      "replies": [
        {
          "id": 1009977,
          "postDate": "2020-09-14T12:18:37.893Z",
          "content": "<p>Yes, in this case I think it's more dependent on what you have in terms of CPU resources. I'm running locally on 18 cores, 128GB ram, SSD. Thanks for the <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> link</p>",
          "rawMarkdown": "Yes, in this case I think it's more dependent on what you have in terms of CPU resources. I'm running locally on 18 cores, 128GB ram, SSD. Thanks for the @tuckerarrants link",
          "votes": 1
        },
        {
          "id": 1010066,
          "postDate": "2020-09-14T13:36:46.563Z",
          "content": "<p>My current hardware setup is far from ideal and it takes me significantly longer than the numbers you report above to train. Increasing raster size and lowering pixel size will increase training time drastically. I tried to train on <code>raster_size = (600, 600)</code> and <code>pixel_size = (.2, .2)</code> and it took me days (like 4) to train around 2 million samples, which is just too long. </p>\n<p>Do you mind me asking, are you running Windows as OS or Linux?</p>",
          "rawMarkdown": "My current hardware setup is far from ideal and it takes me significantly longer than the numbers you report above to train. Increasing raster size and lowering pixel size will increase training time drastically. I tried to train on `raster_size = (600, 600)` and `pixel_size = (.2, .2)` and it took me days (like 4) to train around 2 million samples, which is just too long. \n\nDo you mind me asking, are you running Windows as OS or Linux?",
          "votes": 2
        },
        {
          "id": 1010102,
          "postDate": "2020-09-14T14:03:25.860Z",
          "content": "<p>I'm on Linux. </p>\n<p>I figured any image above 450x450 (pixel size 0.5) was just going to take to long to play around with, so I haven't explored that at all. Thought it was more important to try to cover the full dataset than have bigger image resolution. Who knows, though…</p>",
          "rawMarkdown": "I'm on Linux. \n\nI figured any image above 450x450 (pixel size 0.5) was just going to take to long to play around with, so I haven't explored that at all. Thought it was more important to try to cover the full dataset than have bigger image resolution. Who knows, though...",
          "votes": 3
        },
        {
          "id": 1010106,
          "postDate": "2020-09-14T14:07:52.113Z",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> that's an exciting update! In your case, I believe rasterization is not an issue. So do you just increase the batch size to as large as possible to fasten training convergence?</p>",
          "rawMarkdown": "@fergusoci that's an exciting update! In your case, I believe rasterization is not an issue. So do you just increase the batch size to as large as possible to fasten training convergence?",
          "votes": 3
        },
        {
          "id": 1010155,
          "postDate": "2020-09-14T14:41:49.893Z",
          "content": "<p>Well, rasterization is the issue - it's definitely the bottleneck. I think that it's going to be the bottleneck for everyone unless we can get it running on GPU (I know that <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> had done a lot of great work on this <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359\" target=\"_blank\">here</a>, but to no avail). What I'm doing so far is to try to weigh up the trade offs: raster size / history vs dataset size, etc. If you can get a feel for what's worth more you can prioritise your experiments, then make a final decision as to what you're going to invest a long runtime in for your final set of models.</p>",
          "rawMarkdown": "Well, rasterization is the issue - it's definitely the bottleneck. I think that it's going to be the bottleneck for everyone unless we can get it running on GPU (I know that @pestipeti had done a lot of great work on this [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359), but to no avail). What I'm doing so far is to try to weigh up the trade offs: raster size / history vs dataset size, etc. If you can get a feel for what's worth more you can prioritise your experiments, then make a final decision as to what you're going to invest a long runtime in for your final set of models.",
          "votes": 3
        },
        {
          "id": 1043816,
          "postDate": "2020-10-09T09:09:24.423Z",
          "content": "<p>I think someone was also working on a <code>numba</code> porting of some critical functions used during rasterisation. I'm happy to review any PRs that improve rasterisation's performance but sadly I'm not of much help there as I've already implemented all the tricks I was aware of :( </p>",
          "rawMarkdown": "I think someone was also working on a `numba` porting of some critical functions used during rasterisation. I'm happy to review any PRs that improve rasterisation's performance but sadly I'm not of much help there as I've already implemented all the tricks I was aware of :( "
        }
      ]
    },
    {
      "id": 1041658,
      "postDate": "2020-10-07T21:30:53.677Z",
      "content": "<p>What are your results after the trajectory fix?</p>",
      "rawMarkdown": "What are your results after the trajectory fix?",
      "replies": [
        {
          "id": 1042337,
          "postDate": "2020-10-08T07:21:28.360Z",
          "content": "<p>I didn't repeat the full experiment. My current LB (17.9) is with the trajectory fix.</p>",
          "rawMarkdown": "I didn't repeat the full experiment. My current LB (17.9) is with the trajectory fix."
        }
      ]
    },
    {
      "id": 1011909,
      "postDate": "2020-09-15T18:45:36.010Z",
      "content": "<p>Wow! 6m in 20 hours. How many GPU's are you using for training? and are you using mixed precision?</p>",
      "rawMarkdown": "Wow! 6m in 20 hours. How many GPU's are you using for training? and are you using mixed precision?",
      "replies": [
        {
          "id": 1011929,
          "postDate": "2020-09-15T19:03:37.947Z",
          "content": "<p>Yes, mixed precision. 3 GTX 1080s, but actually, I don't think that's necessary. Should work on a single GPU. Speed is from 16 cores, ssd.</p>",
          "rawMarkdown": "Yes, mixed precision. 3 GTX 1080s, but actually, I don't think that's necessary. Should work on a single GPU. Speed is from 16 cores, ssd.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1010666,
      "postDate": "2020-09-15T01:34:47.413Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1069257,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-04T09:02:05.727000",
      "content": "<p>we have some much data (e.g. train_full). we could capitalize for stacking (i.e. build second or third layer simple blending/ensembling model)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F364aacbfc23535361367c549b3567df5%2FSelection_091.png?generation=1604480959249507&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1043371,
      "author_name": "Dmytro Poplavskiy",
      "author_url": "",
      "post_date": "2020-10-08T22:59:54.107000",
      "content": "<p>I'd expect it would depend quite a bit on the optimiser and the LR schedule as well. I'm currently training on 25% of train.zarr to have faster experiments feedback, seems to be good enough to compare models/approaches</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1043715,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-10-09T07:29:13.210000",
          "content": "<p>Yep, if there's one thing to be said about the data it seems to be very stable. You can run on small training sets, small raster sizes, small history_num_frames and the result preferences seem to consistently hold for larger versions of each.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1043946,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "2020-10-09T10:58:39.693000",
      "content": "<p>I already experiment with some models and also I'm expecting bigger models to get more accuracy.<br>\nbut it's failed.(ResNet18 &gt; ResNet34 &gt; ResNet50). I don't use efficientnet until now.(will experiment)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1043969,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-09T11:26:54.773000",
          "content": "<p>It behaves against the fundamentals of Deep-Learning, but it is very interesting to spectate such relations and it would be way more interesting to understand why it behaves like that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1066009,
          "author_name": "Vee",
          "author_url": "",
          "post_date": "2020-11-01T07:25:02.017000",
          "content": "<p>I think it is because bigger models learn well, but since a lot of data is in some way duplicated, they are learning to the point of \"memorisation\" - Hence it looks like a case of overfitting. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1010130,
      "author_name": "The Brown Iceman",
      "author_url": "",
      "post_date": "2020-09-14T14:22:38.393000",
      "content": "<p>Since we are talking. how far left and down you think we should go to cover this dataset? 😄 Any dense layer will make mathematically more difficult. or do you think others will have much better perf or behaviour?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F344232%2F051fecd3130949270d56134c6db78c54%2FScreenshot%202020-09-11%20110341.png?generation=1600093059987424&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1010164,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-14T14:51:53.243000",
          "content": "<p>I haven't experimented with it, but you can probably implement some complex models without affecting your runtime too much, given that rasterization is the bottleneck. i.e. if you want to cover the dataset, choosing a slightly more complex model isn't going to stop you doing that. </p>\n<p>The only experimenting around backbone that I've done was with Efficientnet - which others have noted <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178245\" target=\"_blank\">underperformed</a>, and I found the same.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1010177,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-09-14T15:02:13.670000",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> Ja! I think because of lot of parameters. Its difficult to say which model is better. I have run into making lot of premature conclusion. Probably my inference about dense layer is one such.    </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1043351,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-10-08T22:31:45.990000",
          "content": "<p>I had similar results with small Efficientnets comparing to resnet18, much worse with a few larger models I tried</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1013113,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2020-09-16T14:07:07.903000",
      "content": "<blockquote>\n  <p>it was more important to try to cover the full dataset than have bigger image resolution</p>\n</blockquote>\n<p>I got only 31 after training 128*128 for 1 epoch (one cycle through train.zarr). I think image size matters.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1013270,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-16T15:42:32.150000",
          "content": "<p>31 seems like a great score for 128*128! For most people getting one cycle through train.zarr is tricky as it takes so long. How long did it take you to run? Have you experimented with &gt; 1 epoch, and if so, did it add much value?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1013283,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-09-16T15:53:31.717000",
          "content": "<p>I am also curious to know. I had been experimenting with 180 * 180 and yesterday I chose to exp with 300 * 300 and my iteration/s went down by a factor of 4. and thats the only parameter i changed. Is it only me or anyone else observing the same?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1013293,
          "author_name": "Leon",
          "author_url": "",
          "post_date": "2020-09-16T16:07:05.727000",
          "content": "<ul>\n<li>It took 2-3 days. </li>\n<li>Overfitting after ~ 1.2 epochs.</li>\n</ul>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1013433,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-16T17:32:34.997000",
          "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> Good to know, thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1010561,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-09-14T21:29:20.720000",
      "content": "<p>Are you able to run that many samples continuously? I have been running into a memory leak on the dataloader. Curious what version of l5kit and pytorch you might be running. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1010950,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-15T06:48:05.063000",
          "content": "<p>Yes, there doesn't seem to be any problems there. I saw your thread but was scratching my head a bit, as it doesn't appear to affect me.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1011038,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-09-15T08:03:00.340000",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> I had a look at your code. I dont see issue over there. Just to share some input. I run L5kit 1.0.6 and pytorch 1.5.1. Even i dont see any problem while loading or iterating. I have also tried to load train_full. Which by the way has 191M(10x more than train) samples and the usage shot up only for few sec. I did encounter the problem you mentioned a week or two back not now. But that was only during creation of chopped dataset. so my recommendation would be obvious ones. Use  <code>import gc</code> <code>gc.collect()</code> once on top. 34 batch and 4 worker should work on kaggle. I have even tried with 64 and 4. Maybe play around and see reaction. just for debug purpose set <code>shuffle</code>to <code>false</code>and see reaction. And if the problem still persist try this thread once. <code>https://discuss.pytorch.org/t/guidelines-for-assigning-num-workers-to-dataloader/813/3</code> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1011044,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-15T08:05:38.940000",
          "content": "<p>odd because I am running the same l5kit and pytorch versions. Unsure of what else might be causing that. I've looked pretty deeply at the various dataloader parameters and run gc.collect and cleared torch cache and it seems to be something accumulating inside of the dataloader itself so gc and other methods dont clear it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1011773,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-15T17:14:57.450000",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Not sure if it helps, but one thing that I'm doing slightly differently to what you have is wrapping the rasterizer dataset. Maybe this leads to a different interaction with the dataloader.</p>\n<p>Sample shell wrapper:</p>\n<p>class MotionPredictDataset(Dataset):</p>\n<pre><code>def __init__(self,\n             cfg,\n             str_loader='train_data_loader',\n             fn_rasterizer=build_rasterizer,\n             test_mask_path=os.path.join(BASE_DIR, 'scenes/mask.npz')):\n\n    self.cfg = cfg\n    self.str_loader = str_loader\n    self.fn_rasterizer = fn_rasterizer\n    self.test_mask_path = test_mask_path\n\n    self.setup()\n\ndef setup(self):\n\n    self.dm = LocalDataManager(None)\n    self.rasterizer = self.fn_rasterizer(self.cfg, self.dm)\n    self.data_zarr = ChunkedDataset(self.dm.require(self.cfg[self.str_loader][\"key\"])).open()\n\n    if self.str_loader == 'test_data_loader':\n        test_mask = np.load(self.test_mask_path)[\"arr_0\"]\n        self.ds = AgentDataset(self.cfg, self.data_zarr, self.rasterizer, agents_mask=test_mask)\n    else:\n        self.ds = AgentDataset(self.cfg, self.data_zarr, self.rasterizer)\n\ndef __getitem__(self, index):\n    return self.ds[index]\n\ndef __len__(self):\n    return self.ds\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1010322,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "2020-09-14T17:05:59.197000",
      "content": "<p>Have you considered pre-rasterizing your images? </p>\n<p>I mean, if rasterization is your bottleneck, why not try and run through the rasterization process once and then saving just the images. <br>\nThen you can just read in the images. </p>\n<p>Admittedly, this might take up quite a lot of space. But you are now swapping cpu cycles for diskspace. </p>\n<p>I'll have to check how much space this would take up, but it might be doable. Might be worth it to store it in a compressed format like png, especially if you are using the semantic rasterizer. </p>\n<p>I haven't actually tried this, has anyone? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1010332,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-14T17:15:43.060000",
          "content": "<p><a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> Take a look at <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">this thread</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1010388,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "2020-09-14T17:54:29.833000",
          "content": "<p>Ah thanks, makes sense that someone already did it. That's quite a lot of space, and a decent increase in speed.</p>\n<p>Should be better if you store compressed, seeing as the semantic images are mostly large blocks of white, but then you lose a bit again in the decoding. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1009945,
      "author_name": "The Brown Iceman",
      "author_url": "",
      "post_date": "2020-09-14T11:46:52.430000",
      "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> For me it took 70hrs for 3M. Think it depends on lot of factors. For strategy with image size or resolution. You can refer to <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> experiments. Here is the link for it. <a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/tuckerarrants/lyftpretrainedmodels\" target=\"_blank\">https://www.kaggle.com/tuckerarrants/lyftpretrainedmodels</a>    . I believe he has done these experiment for unimodel output. You can transfer the idea maybe.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1009977,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-14T12:18:37.893000",
          "content": "<p>Yes, in this case I think it's more dependent on what you have in terms of CPU resources. I'm running locally on 18 cores, 128GB ram, SSD. Thanks for the <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> link</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1010066,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2020-09-14T13:36:46.563000",
          "content": "<p>My current hardware setup is far from ideal and it takes me significantly longer than the numbers you report above to train. Increasing raster size and lowering pixel size will increase training time drastically. I tried to train on <code>raster_size = (600, 600)</code> and <code>pixel_size = (.2, .2)</code> and it took me days (like 4) to train around 2 million samples, which is just too long. </p>\n<p>Do you mind me asking, are you running Windows as OS or Linux?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1010102,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-14T14:03:25.860000",
          "content": "<p>I'm on Linux. </p>\n<p>I figured any image above 450x450 (pixel size 0.5) was just going to take to long to play around with, so I haven't explored that at all. Thought it was more important to try to cover the full dataset than have bigger image resolution. Who knows, though…</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1010106,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2020-09-14T14:07:52.113000",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> that's an exciting update! In your case, I believe rasterization is not an issue. So do you just increase the batch size to as large as possible to fasten training convergence?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1010155,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-14T14:41:49.893000",
          "content": "<p>Well, rasterization is the issue - it's definitely the bottleneck. I think that it's going to be the bottleneck for everyone unless we can get it running on GPU (I know that <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> had done a lot of great work on this <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359\" target=\"_blank\">here</a>, but to no avail). What I'm doing so far is to try to weigh up the trade offs: raster size / history vs dataset size, etc. If you can get a feel for what's worth more you can prioritise your experiments, then make a final decision as to what you're going to invest a long runtime in for your final set of models.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1043816,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-10-09T09:09:24.423000",
          "content": "<p>I think someone was also working on a <code>numba</code> porting of some critical functions used during rasterisation. I'm happy to review any PRs that improve rasterisation's performance but sadly I'm not of much help there as I've already implemented all the tricks I was aware of :( </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1041658,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-10-07T21:30:53.677000",
      "content": "<p>What are your results after the trajectory fix?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1042337,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-10-08T07:21:28.360000",
          "content": "<p>I didn't repeat the full experiment. My current LB (17.9) is with the trajectory fix.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1011909,
      "author_name": "SujaydKhandekar",
      "author_url": "",
      "post_date": "2020-09-15T18:45:36.010000",
      "content": "<p>Wow! 6m in 20 hours. How many GPU's are you using for training? and are you using mixed precision?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1011929,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-15T19:03:37.947000",
          "content": "<p>Yes, mixed precision. 3 GTX 1080s, but actually, I don't think that's necessary. Should work on a single GPU. Speed is from 16 cores, ssd.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1010666,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-15T01:34:47.413000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1009894": "From playing around it seems like one of the largest issues we’ll contend with is training time. The possibility of using train_full.zarr seems fanciful to me right now!  That said, using the full dataset seems like it will be important for performance. \n\nHere are the results of some experiments that I have run so far:\n\n- resnet18 backbone, 6M samples (approx 26% of train.zarr), train time ~20 hrs: LB 35.9\n\n- resnet18 backbone, 18M samples (approx 80% of train.zarr), train time ~60 hrs: LB 29.8\n\nThis is using input image sizes of 320, training locally. I think going with a larger input size will help, but this will push training times to a point where it’s taking a week to cycle once through the small training dataset.\n\nDoing well in this competition is clearing going to require some strategising around input size/history_num_frames/training time trade offs. \n\nMy current strategy is to test ideas with a small image size and approx 13% of train.zarr, then run the ideas that work on a more extended dataset. It seems that this smaller dataset is stable enough to determine relative performance.\n\nDoes anybody have an idea of the impact of using more of the training data - either more than one cycle through train.zarr, or of using train_full.zarr?",
    "1069257": "we have some much data (e.g. train\\_full). we could capitalize for stacking (i.e. build second or third layer simple blending/ensembling model)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F364aacbfc23535361367c549b3567df5%2FSelection_091.png?generation=1604480959249507&alt=media)",
    "1043371": "I'd expect it would depend quite a bit on the optimiser and the LR schedule as well. I'm currently training on 25% of train.zarr to have faster experiments feedback, seems to be good enough to compare models/approaches",
    "1043946": "I already experiment with some models and also I'm expecting bigger models to get more accuracy.\nbut it's failed.(ResNet18 > ResNet34 > ResNet50). I don't use efficientnet until now.(will experiment)",
    "1010130": "Since we are talking. how far left and down you think we should go to cover this dataset? 😄 Any dense layer will make mathematically more difficult. or do you think others will have much better perf or behaviour?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F344232%2F051fecd3130949270d56134c6db78c54%2FScreenshot%202020-09-11%20110341.png?generation=1600093059987424&alt=media)",
    "1013113": "> it was more important to try to cover the full dataset than have bigger image resolution\n\nI got only 31 after training 128*128 for 1 epoch (one cycle through train.zarr). I think image size matters.",
    "1010561": "Are you able to run that many samples continuously? I have been running into a memory leak on the dataloader. Curious what version of l5kit and pytorch you might be running. ",
    "1010322": "Have you considered pre-rasterizing your images? \n\nI mean, if rasterization is your bottleneck, why not try and run through the rasterization process once and then saving just the images. \nThen you can just read in the images. \n\nAdmittedly, this might take up quite a lot of space. But you are now swapping cpu cycles for diskspace. \n\nI'll have to check how much space this would take up, but it might be doable. Might be worth it to store it in a compressed format like png, especially if you are using the semantic rasterizer. \n\nI haven't actually tried this, has anyone? ",
    "1009945": "@fergusoci For me it took 70hrs for 3M. Think it depends on lot of factors. For strategy with image size or resolution. You can refer to @tuckerarrants experiments. Here is the link for it. [https://www.kaggle.com/tuckerarrants/lyftpretrainedmodels   ](url) . I believe he has done these experiment for unimodel output. You can transfer the idea maybe.  ",
    "1041658": "What are your results after the trajectory fix?",
    "1011909": "Wow! 6m in 20 hours. How many GPU's are you using for training? and are you using mixed precision?",
    "1010666": ""
  }
}