{
  "id": 189996,
  "title": "How can we train for long time?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/189996",
  "author_name": "Mohammed Rizin V K",
  "post_date": "2020-10-09T16:21:39.438000",
  "votes": 4,
  "comment_count": 63,
  "views": 0,
  "content": "<p>Can anyone suggest a way to run 10k samples on TPU or GPU faster</p>",
  "messages": [
    {
      "id": 1044250,
      "postDate": "2020-10-09T16:21:39.437Z",
      "content": "<p>Can anyone suggest a way to run 10k samples on TPU or GPU faster</p>",
      "rawMarkdown": "Can anyone suggest a way to run 10k samples on TPU or GPU faster",
      "votes": 4
    },
    {
      "id": 1045189,
      "postDate": "2020-10-10T12:23:27.673Z",
      "content": "<p>Can anyone have a way to split the batches that is <br>\nif we cant train for 1million batches at one go<br>\ncan we train one session with 10000 and refresh the kernel and again 10000<br>\nCan anyone know how to do it<br>\nQuestion to everyone also <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> <a href=\"https://www.kaggle.com/pondruska\" target=\"_blank\">@pondruska</a> <br>\nSorry if the question is already answered</p>",
      "rawMarkdown": "Can anyone have a way to split the batches that is \nif we cant train for 1million batches at one go\ncan we train one session with 10000 and refresh the kernel and again 10000\nCan anyone know how to do it\nQuestion to everyone also @iglovikov @lucabergamini @pondruska \nSorry if the question is already answered",
      "votes": 1,
      "replies": [
        {
          "id": 1045274,
          "postDate": "2020-10-10T13:31:15.783Z",
          "content": "<p>If you dont use shuffeling you could do it with something like:</p>\n<pre><code>for itr in progress_bar:\n        if itr &lt;= pretrain_iteration:\n            continue\n...\n</code></pre>\n<p>with <code>pretrain_iteration</code> being 10000 in your case. </p>\n<p>With shuffeling off you make sure, that you dont let the model learn the same sample again.</p>",
          "rawMarkdown": "If you dont use shuffeling you could do it with something like:\n\n```\nfor itr in progress_bar:\n        if itr <= pretrain_iteration:\n            continue\n...\n```\nwith `pretrain_iteration` being 10000 in your case. \n\nWith shuffeling off you make sure, that you dont let the model learn the same sample again.",
          "votes": 2
        },
        {
          "id": 1045302,
          "postDate": "2020-10-10T13:58:14.707Z",
          "content": "<p>Can we do seed and shuffle it will not any problem doesnt it</p>",
          "rawMarkdown": "Can we do seed and shuffle it will not any problem doesnt it"
        },
        {
          "id": 1045371,
          "postDate": "2020-10-10T15:03:09.230Z",
          "content": "<p>No it wont cause any problems, you might depending on the size of <code>pretrain_iteration</code>, teach your model the same input multiple times but thats equal to training it over a set amount of epochs.</p>",
          "rawMarkdown": "No it wont cause any problems, you might depending on the size of `pretrain_iteration`, teach your model the same input multiple times but thats equal to training it over a set amount of epochs.",
          "votes": 1
        },
        {
          "id": 1046105,
          "postDate": "2020-10-11T10:53:15.573Z",
          "content": "<p>I have similar problem. I am working from home so I can only use my PC for training at night. I hope to resume my training at the same place of dataloader as I stopped last time.</p>\n<p>Here is how I do it. Hope it makes sense (if it doesn't please let me know):</p>\n<pre><code>random.seed(2020)\n\nclass MySampler(Sampler):\n    def __init__(self, data, i=0):\n        self.seq = list(range(len(data)))\n        random.shuffle(self.seq)\n        self.seq = self.seq[i * train_cfg[\"batch_size\"]:]\n\n    def __iter__(self):\n        return iter(self.seq)\n\n    def __len__(self):\n        return len(self.seq)\n\nsampler = MySampler(train_dataset, i = starting)\ntrain_dataloader = DataLoader(train_dataset, batch_size=train_cfg[\"batch_size\"],        \n                         sampler=sampler, num_workers=train_cfg[\"num_workers\"])\n</code></pre>\n<p>if you stopped at 100k iterations last time, let starting = 100000</p>",
          "rawMarkdown": "I have similar problem. I am working from home so I can only use my PC for training at night. I hope to resume my training at the same place of dataloader as I stopped last time.\n\nHere is how I do it. Hope it makes sense (if it doesn't please let me know):\n```\nrandom.seed(2020)\n\nclass MySampler(Sampler):\n    def __init__(self, data, i=0):\n        self.seq = list(range(len(data)))\n        random.shuffle(self.seq)\n        self.seq = self.seq[i * train_cfg[\"batch_size\"]:]\n\n    def __iter__(self):\n        return iter(self.seq)\n\n    def __len__(self):\n        return len(self.seq)\n\nsampler = MySampler(train_dataset, i = starting)\ntrain_dataloader = DataLoader(train_dataset, batch_size=train_cfg[\"batch_size\"],        \n                         sampler=sampler, num_workers=train_cfg[\"num_workers\"])\n```\nif you stopped at 100k iterations last time, let starting = 100000",
          "votes": 1
        },
        {
          "id": 1046260,
          "postDate": "2020-10-11T14:05:46.210Z",
          "content": "<p>Here is some solution i could find out from kernels hope it works</p>\n<pre><code>tr_it = iter(train_dataloader)\n\nprogress_bar = tqdm(range(iterno,cfg[\"train_params\"][\"max_num_steps\"]))\nlosses_train = []\nfor itr in progress_bar:\n    try:\n        data = next(tr_it)\n\n    except StopIteration:\n\n        tr_it = iter(train_dataloader)\n        data = next(tr_it)\n\n    model.train()\n    torch.set_grad_enabled(True)\n</code></pre>\n<p>where we save iter no with the checkpoint to not lost our iterations</p>",
          "rawMarkdown": "Here is some solution i could find out from kernels hope it works\n\n```\n\ntr_it = iter(train_dataloader)\n\nprogress_bar = tqdm(range(iterno,cfg[\"train_params\"][\"max_num_steps\"]))\nlosses_train = []\nfor itr in progress_bar:\n    try:\n        data = next(tr_it)\n\n    except StopIteration:\n\n        tr_it = iter(train_dataloader)\n        data = next(tr_it)\n\n    model.train()\n    torch.set_grad_enabled(True)\n```\n\nwhere we save iter no with the checkpoint to not lost our iterations"
        },
        {
          "id": 1046549,
          "postDate": "2020-10-11T19:03:49.533Z",
          "content": "<p>don't forget to also save the optimizer state and the current step for the scheduler.<br>\nsee <a href=\"https://pytorch.org/tutorials/beginner/saving_loading_models.html\" target=\"_blank\">here</a> for details how this is done in pytorch. </p>",
          "rawMarkdown": "don't forget to also save the optimizer state and the current step for the scheduler.\nsee [here](https://pytorch.org/tutorials/beginner/saving_loading_models.html) for details how this is done in pytorch. ",
          "votes": 2
        },
        {
          "id": 1046792,
          "postDate": "2020-10-12T02:40:48.763Z",
          "content": "<p>Ya  I forgot about it</p>",
          "rawMarkdown": "Ya  I forgot about it"
        },
        {
          "id": 1062067,
          "postDate": "2020-10-27T14:44:58.317Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>  I'm having problem with loading optimizer's state. It throws up errors like </p>\n<pre><code>runtimeerror: expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!\n</code></pre>\n<p>After some browsing I found a way to move all the tensors to cuda by</p>\n<pre><code>    checkpoint = torch.load(weight_path)\n    model.load_state_dict(checkpoint['model_state_dict'])\n    model.cuda()\n    optimizer = optim.Adadelta(model.parameters())\n    optimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n    for state in optimizer.state.values():\n        for k, v in state.items():\n            if isinstance(v, torch.Tensor):\n                state[k] = v.cuda()\n</code></pre>\n<p>But still I'm unable to resume my training and loss shoots up to a large value (~4000) but drops after few iterations (~400)</p>\n<p>Did you face any such issues </p>\n<pre><code>This happened when I used the weights trained over Kaggle's GPU on my PC \nMy Setup\nOS : Win 10, torch 1.6.0\nGPU : RTX 2060\n</code></pre>\n<p>Any ideas??</p>\n<p>In simple I just wanna resume training offline……..</p>\n<p>Thanks </p>",
          "rawMarkdown": "Hi @ilu000  I'm having problem with loading optimizer's state. It throws up errors like \n```\nruntimeerror: expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!\n```\nAfter some browsing I found a way to move all the tensors to cuda by\n```   \n    checkpoint = torch.load(weight_path)\n    model.load_state_dict(checkpoint['model_state_dict'])\n    model.cuda()\n    optimizer = optim.Adadelta(model.parameters())\n    optimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n    for state in optimizer.state.values():\n        for k, v in state.items():\n            if isinstance(v, torch.Tensor):\n                state[k] = v.cuda()\n```\n\n\nBut still I'm unable to resume my training and loss shoots up to a large value (~4000) but drops after few iterations (~400)\n\nDid you face any such issues \n\n```\nThis happened when I used the weights trained over Kaggle's GPU on my PC \nMy Setup\nOS : Win 10, torch 1.6.0\nGPU : RTX 2060\n```\nAny ideas??\n\nIn simple I just wanna resume training offline........\n\nThanks "
        },
        {
          "id": 1062076,
          "postDate": "2020-10-27T14:59:21.653Z",
          "content": "<p>Right now i added a sampler</p>",
          "rawMarkdown": "Right now i added a sampler"
        },
        {
          "id": 1062097,
          "postDate": "2020-10-27T15:09:45.993Z",
          "content": "<p>Were you able to continue from the checkpoint?</p>",
          "rawMarkdown": "Were you able to continue from the checkpoint?"
        },
        {
          "id": 1062137,
          "postDate": "2020-10-27T15:46:27.407Z",
          "content": "<p>Yes, I had a successful in them after a long time</p>",
          "rawMarkdown": "Yes, I had a successful in them after a long time"
        },
        {
          "id": 1062141,
          "postDate": "2020-10-27T15:48:18.900Z",
          "content": "<p>Nice ..Did it require saving optimizer and scheduler state ?  </p>",
          "rawMarkdown": "Nice ..Did it require saving optimizer and scheduler state ?  "
        },
        {
          "id": 1062151,
          "postDate": "2020-10-27T15:55:14.360Z",
          "content": "<p>i save optimizer , scheduler , sampler, model weights all the things</p>",
          "rawMarkdown": "i save optimizer , scheduler , sampler, model weights all the things"
        },
        {
          "id": 1065800,
          "postDate": "2020-10-31T19:04:42.377Z",
          "content": "<p><a href=\"https://www.kaggle.com/ak0210\" target=\"_blank\">@ak0210</a> you should try this</p>\n<p>device = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")<br>\ncheckpoint = torch.load(weight_path)<br>\nmodel.load_state_dict(checkpoint['model_state_dict'])<br>\nmodel.to(device)<br>\noptimizer = optim.Adadelta(model.parameters())<br>\noptimizer.to(device) <strong><em>if you want your optimizer in cuda else no need to this</em></strong><br>\noptimizer.load_state_dict(checkpoint['optimizer_state_dict'])</p>\n<p><strong>This is the procedure for loading model and optimizer in pytorch</strong></p>",
          "rawMarkdown": "@ak0210 you should try this\n\ndevice = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\ncheckpoint = torch.load(weight_path)\nmodel.load_state_dict(checkpoint['model_state_dict'])\nmodel.to(device)\noptimizer = optim.Adadelta(model.parameters())\noptimizer.to(device) ***if you want your optimizer in cuda else no need to this***\noptimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n\n**This is the procedure for loading model and optimizer in pytorch**"
        },
        {
          "id": 1067694,
          "postDate": "2020-11-02T16:26:09.437Z",
          "content": "<p>Thanks for the suggestion <a href=\"https://www.kaggle.com/tanishgupta34\" target=\"_blank\">@tanishgupta34</a> but still ain't working, getting the error..<br>\nThe error pops up when I try to load to GPU only</p>\n<pre><code>&lt;ipython-input-17-4a4c8090f612&gt; in &lt;module&gt;\n      6 model.to(device)\n      7 optimizer = optim.Adadelta(model.parameters())\n----&gt; 8 optimizer.to(device)\n      9 optimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n\nAttributeError: 'Adadelta' object has no attribute 'to'\n</code></pre>\n<p>My solution on previous reply does the trick </p>",
          "rawMarkdown": "Thanks for the suggestion @tanishgupta34 but still ain't working, getting the error..\nThe error pops up when I try to load to GPU only\n\n```\n<ipython-input-17-4a4c8090f612> in <module>\n      6 model.to(device)\n      7 optimizer = optim.Adadelta(model.parameters())\n----> 8 optimizer.to(device)\n      9 optimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n\nAttributeError: 'Adadelta' object has no attribute 'to'\n```\n\nMy solution on previous reply does the trick "
        },
        {
          "id": 1070085,
          "postDate": "2020-11-05T10:24:57.767Z",
          "content": "<p><a href=\"https://www.kaggle.com/ak0210\" target=\"_blank\">@ak0210</a> yeah this was working fine on 1.5.1 but now in version 1.6.0 it is showing error.You can try once by removing optimizer.to(device) and it will work fine</p>",
          "rawMarkdown": "@ak0210 yeah this was working fine on 1.5.1 but now in version 1.6.0 it is showing error.You can try once by removing optimizer.to(device) and it will work fine"
        },
        {
          "id": 1070406,
          "postDate": "2020-11-05T18:14:10.920Z",
          "content": "<p><a href=\"https://www.kaggle.com/tanishgupta34\" target=\"_blank\">@tanishgupta34</a>  yes, it works fine without <code>optimizer.to(device)</code> but I find the loss values exploding.<br>\nWhich is partly resolved in by the solution in the above comment. </p>",
          "rawMarkdown": "@tanishgupta34  yes, it works fine without `optimizer.to(device)` but I find the loss values exploding.\nWhich is partly resolved in by the solution in the above comment. "
        }
      ]
    },
    {
      "id": 1044325,
      "postDate": "2020-10-09T17:11:11.260Z",
      "content": "<p>With more dataloader workers your rasterization step which is the main bottleneck can be executed faster. I am running 40 workers on a 6 core machine.</p>",
      "rawMarkdown": "With more dataloader workers your rasterization step which is the main bottleneck can be executed faster. I am running 40 workers on a 6 core machine.",
      "votes": -1,
      "replies": [
        {
          "id": 1044519,
          "postDate": "2020-10-09T20:44:33.030Z",
          "content": "<p>Is the difference noticeable?</p>\n<p>When I put <code>num_workers</code> higher than 4 I cant see a difference anymore.</p>",
          "rawMarkdown": "Is the difference noticeable?\n\nWhen I put `num_workers` higher than 4 I cant see a difference anymore.",
          "votes": 1
        },
        {
          "id": 1044531,
          "postDate": "2020-10-09T21:08:43.067Z",
          "content": "<p>From my (limited) experience more than 2 workers per core are usually not beneficial. <br>\nHere, I didn't even notice a difference between 1 worker per core and 2 workers per core. </p>\n<p>Kaggle kernels have 2 cores, if you were wondering. </p>",
          "rawMarkdown": "From my (limited) experience more than 2 workers per core are usually not beneficial. \nHere, I didn't even notice a difference between 1 worker per core and 2 workers per core. \n\nKaggle kernels have 2 cores, if you were wondering. ",
          "votes": 1
        },
        {
          "id": 1045135,
          "postDate": "2020-10-10T11:23:23.857Z",
          "content": "<p>That makes sense, so tweaking the <code>num_worker</code> count seems to be useless change. </p>",
          "rawMarkdown": "That makes sense, so tweaking the `num_worker` count seems to be useless change. ",
          "votes": -1
        }
      ]
    },
    {
      "id": 1059374,
      "postDate": "2020-10-25T02:44:41.320Z",
      "content": "<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> you said you training loss was 24.xxx in this Public notebook. But it seems to be 324.xxx when we start the training with that loss</p>",
      "rawMarkdown": "@huanvo you said you training loss was 24.xxx in this Public notebook. But it seems to be 324.xxx when we start the training with that loss"
    },
    {
      "id": 1059364,
      "postDate": "2020-10-25T02:18:26.950Z",
      "content": "<p>If we training train_full.zarr by using l5kit do we need 64GB of Ram to actually store them</p>",
      "rawMarkdown": "If we training train_full.zarr by using l5kit do we need 64GB of Ram to actually store them",
      "replies": [
        {
          "id": 1059624,
          "postDate": "2020-10-25T09:41:03.267Z",
          "content": "<p>32 GB seems to work fine. Load the train set before the eval set. After finishing the loading process, the blocked memory will drop significantly. </p>",
          "rawMarkdown": "32 GB seems to work fine. Load the train set before the eval set. After finishing the loading process, the blocked memory will drop significantly. "
        },
        {
          "id": 1059653,
          "postDate": "2020-10-25T10:16:32.957Z",
          "content": "<p>i know that it will get down after finishing. I just looked up in the kaggle kernel but it was allocating more memory</p>",
          "rawMarkdown": "i know that it will get down after finishing. I just looked up in the kaggle kernel but it was allocating more memory",
          "votes": 1
        },
        {
          "id": 1059720,
          "postDate": "2020-10-25T11:39:22.440Z",
          "content": "<p>Yes, for kaggle kernels it's too large (probably the reason why they have the smaller train.zarr included here by default).<br>\nI was just saying that the minimum requirement is 32 GB and not 64 GB. </p>",
          "rawMarkdown": "Yes, for kaggle kernels it's too large (probably the reason why they have the smaller train.zarr included here by default).\nI was just saying that the minimum requirement is 32 GB and not 64 GB. "
        },
        {
          "id": 1059725,
          "postDate": "2020-10-25T11:46:06.203Z",
          "content": "<p>Thank you, But how long it takes to run 1 epoch?</p>",
          "rawMarkdown": "Thank you, But how long it takes to run 1 epoch?"
        },
        {
          "id": 1059804,
          "postDate": "2020-10-25T13:33:59.613Z",
          "content": "<blockquote>\n  <p>1 epoch of train_full.zarr would take me about 22days</p>\n</blockquote>\n<p>you could just scroll down and look at my comment from 15 days ago</p>",
          "rawMarkdown": "> 1 epoch of train_full.zarr would take me about 22days\n\nyou could just scroll down and look at my comment from 15 days ago\n",
          "votes": 2
        },
        {
          "id": 1059821,
          "postDate": "2020-10-25T13:54:14.360Z",
          "content": "<p>Oh sorry, I thought it was for train.zarr. Okay how long it takes for train.zarr.<br>\nAlso, I have a doubt on my solution itself. I have given a <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/189996#1046260\" target=\"_blank\">code solution</a> in this discussion for resuming the iteration<br>\nbut I have a doubt, when we resume is it starts from point1 or from the iter no <br>\nsince I am doing <br>\n<code>data = next(tr_it)</code></p>",
          "rawMarkdown": "Oh sorry, I thought it was for train.zarr. Okay how long it takes for train.zarr.\nAlso, I have a doubt on my solution itself. I have given a [code solution](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/189996#1046260) in this discussion for resuming the iteration\nbut I have a doubt, when we resume is it starts from point1 or from the iter no \nsince I am doing \n`data = next(tr_it)`"
        },
        {
          "id": 1059834,
          "postDate": "2020-10-25T14:09:58.003Z",
          "content": "<p>train_full has 191_177_863 items. Standard train has what, ~23 million? <br>\nSize ~1/9, so time would also be ~1/9</p>\n<p>22 days / 9 = 2.5 days</p>\n<p>edit: seeing your edit above, you need to save the current point of your dataloader as well (and save your shuffle seed) if you want to continue training. </p>",
          "rawMarkdown": "train_full has 191_177_863 items. Standard train has what, ~23 million? \nSize ~1/9, so time would also be ~1/9\n\n22 days / 9 = 2.5 days\n\nedit: seeing your edit above, you need to save the current point of your dataloader as well (and save your shuffle seed) if you want to continue training. ",
          "votes": 3
        },
        {
          "id": 1059839,
          "postDate": "2020-10-25T14:15:37.533Z",
          "content": "<p>I have unique seed. But what about saving DataLoader<br>\nWhen I try to train train.zarr it shows it take 624hours on kaggle kernel<br>\nbut on colab its 3x faster<br>\nand 4x faster in my Tesla T4 (4GPU)<br>\nAnd are you saying you are running on GPU<br>\nWhats your GPU</p>",
          "rawMarkdown": "I have unique seed. But what about saving DataLoader\nWhen I try to train train.zarr it shows it take 624hours on kaggle kernel\nbut on colab its 3x faster\nand 4x faster in my Tesla T4 (4GPU)\nAnd are you saying you are running on GPU\nWhats your GPU"
        },
        {
          "id": 1059951,
          "postDate": "2020-10-25T16:05:40.597Z",
          "content": "<p>Why is my code takes 20 days to finish on train.zarr? Can you fix them please</p>",
          "rawMarkdown": "Why is my code takes 20 days to finish on train.zarr? Can you fix them please"
        },
        {
          "id": 1059964,
          "postDate": "2020-10-25T16:18:31.843Z",
          "content": "<p>mixed precision, raster_size, CPU speed, lots of things why your runtime may be longer. <br>\nI suggest you go ahead and learn how to profile your code, find the bottleneck and then ask more specific questions. </p>",
          "rawMarkdown": "mixed precision, raster_size, CPU speed, lots of things why your runtime may be longer. \nI suggest you go ahead and learn how to profile your code, find the bottleneck and then ask more specific questions. ",
          "votes": 2
        },
        {
          "id": 1059974,
          "postDate": "2020-10-25T16:30:54.253Z",
          "rawMarkdown": "",
          "votes": -4,
          "isDeleted": true
        },
        {
          "id": 1060541,
          "postDate": "2020-10-26T09:40:18.217Z",
          "rawMarkdown": "",
          "votes": -4,
          "isDeleted": true
        },
        {
          "id": 1060645,
          "postDate": "2020-10-26T12:39:42.077Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 1061823,
          "postDate": "2020-10-27T11:01:04.687Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 1062159,
          "postDate": "2020-10-27T16:00:19.213Z",
          "content": "<p>The CPU Speed of kaggle kernels is very low (and this is sadly the bottleneck here) and you only have two cores available. Even on nowadays multicore consumer CPUs (combined with an OK GPU) you can get to much faster training times.</p>",
          "rawMarkdown": "The CPU Speed of kaggle kernels is very low (and this is sadly the bottleneck here) and you only have two cores available. Even on nowadays multicore consumer CPUs (combined with an OK GPU) you can get to much faster training times."
        },
        {
          "id": 1062165,
          "postDate": "2020-10-27T16:04:42.693Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        },
        {
          "id": 1062222,
          "postDate": "2020-10-27T17:05:08.397Z",
          "content": "<p>As stated <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/192458#1057112\" target=\"_blank\">here</a> 2 cores (kaggle kernels) give about 1 batch (16 frames) per second, which gets you to 16.6 days of training (23M frames, 16 frames/batch). Colab Pro high-memory VM for GPU: CPU cores = 4 can already make that 2.7 times faster. 32 cores should give you a 16 times speed-up: 1.1 days (maybe a little less because of the necessary overhead). I don't know why your training is not faster than on kaggle kernels. </p>\n<p>mixed precision makes use of less accurate floats (lesser memory footprint, faster I/O). You can read about it <a href=\"https://pytorch.org/blog/accelerating-training-on-nvidia-gpus-with-pytorch-automatic-mixed-precision/\" target=\"_blank\">here</a>. It can also make your training slightly faster. Decreasing the raster_size has an influence on your training time as your arrays can be smaller but it's still limited by the transformations as stated <a href=\"https://github.com/lyft/l5kit/issues/136\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "As stated [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/192458#1057112) 2 cores (kaggle kernels) give about 1 batch (16 frames) per second, which gets you to 16.6 days of training (23M frames, 16 frames/batch). Colab Pro high-memory VM for GPU: CPU cores = 4 can already make that 2.7 times faster. 32 cores should give you a 16 times speed-up: 1.1 days (maybe a little less because of the necessary overhead). I don't know why your training is not faster than on kaggle kernels. \n\nmixed precision makes use of less accurate floats (lesser memory footprint, faster I/O). You can read about it [here](https://pytorch.org/blog/accelerating-training-on-nvidia-gpus-with-pytorch-automatic-mixed-precision/). It can also make your training slightly faster. Decreasing the raster_size has an influence on your training time as your arrays can be smaller but it's still limited by the transformations as stated [here](https://github.com/lyft/l5kit/issues/136).",
          "votes": 2
        },
        {
          "id": 1062226,
          "postDate": "2020-10-27T17:08:53.733Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        },
        {
          "id": 1066238,
          "postDate": "2020-11-01T14:17:04.750Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        },
        {
          "id": 1071115,
          "postDate": "2020-11-06T14:14:54.610Z",
          "content": "<p>You mean if i have 32 vCPU then I should select num_workers = 32 will use all my CPU or 64</p>",
          "rawMarkdown": "You mean if i have 32 vCPU then I should select num_workers = 32 will use all my CPU or 64"
        },
        {
          "id": 1071502,
          "postDate": "2020-11-07T00:12:26.093Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> he meant exactly the same you told and you can check the numworkers by using <strong>os.cpu_count()</strong> if you have any confusion </p>",
          "rawMarkdown": "@morizin he meant exactly the same you told and you can check the numworkers by using **os.cpu_count()** if you have any confusion "
        }
      ]
    },
    {
      "id": 1057203,
      "postDate": "2020-10-22T13:19:40.063Z",
      "content": "<p>Do anyone know that is this <a href=\"https://www.kaggle.com/ryanholbrook/lyft-motion-prediction-tfrecords\" target=\"_blank\">data</a> is train_full.zarr or train.zarr file?</p>",
      "rawMarkdown": "Do anyone know that is this [data](https://www.kaggle.com/ryanholbrook/lyft-motion-prediction-tfrecords) is train_full.zarr or train.zarr file?\n",
      "replies": [
        {
          "id": 1057205,
          "postDate": "2020-10-22T13:23:13.643Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> you will find full data here <a href=\"url\" target=\"_blank\">https://www.kaggle.com/philculliton/lyft-full-training-set</a>. And train.zarr you can download from data section of commpetition.</p>",
          "rawMarkdown": "@morizin you will find full data here [https://www.kaggle.com/philculliton/lyft-full-training-set](url). And train.zarr you can download from data section of commpetition."
        },
        {
          "id": 1057210,
          "postDate": "2020-10-22T13:25:59.490Z",
          "content": "<p><a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a> is that data is Full data, So that I could decide to train on zarr file or TfRec Files</p>",
          "rawMarkdown": "@deepakrajpurushothaman is that data is Full data, So that I could decide to train on zarr file or TfRec Files"
        }
      ]
    },
    {
      "id": 1047356,
      "postDate": "2020-10-12T14:08:30.050Z",
      "content": "<p>Is anyone is using Tensorflow if so how did you manage to do Dataset</p>",
      "rawMarkdown": "Is anyone is using Tensorflow if so how did you manage to do Dataset"
    },
    {
      "id": 1046031,
      "postDate": "2020-10-11T09:30:24.197Z",
      "content": "<p>Does the score on eval set while training will project out to leaderboard</p>",
      "rawMarkdown": "Does the score on eval set while training will project out to leaderboard",
      "replies": [
        {
          "id": 1046100,
          "postDate": "2020-10-11T10:48:23.680Z",
          "content": "<p>See the discussion <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695\" target=\"_blank\">here</a>.</p>\n<p>Using the correct validation method the LB score correlates very well with the validation score</p>",
          "rawMarkdown": "See the discussion [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695).\n\nUsing the correct validation method the LB score correlates very well with the validation score",
          "votes": 2
        }
      ]
    },
    {
      "id": 1045166,
      "postDate": "2020-10-10T11:56:59.667Z",
      "content": "<p>If you aren't training all the lower layers you might be able to forward prop once and store/run from there</p>",
      "rawMarkdown": "If you aren't training all the lower layers you might be able to forward prop once and store/run from there"
    },
    {
      "id": 1044498,
      "postDate": "2020-10-09T20:23:24.607Z",
      "content": "<p>Write the transformation for the rasterizer in C++ or port the routine to GPU as this is the bottleneck. <br>\nsee discussion here <a href=\"https://github.com/lyft/l5kit/issues/136\" target=\"_blank\">https://github.com/lyft/l5kit/issues/136</a></p>",
      "rawMarkdown": "Write the transformation for the rasterizer in C++ or port the routine to GPU as this is the bottleneck. \nsee discussion here https://github.com/lyft/l5kit/issues/136",
      "replies": [
        {
          "id": 1044509,
          "postDate": "2020-10-09T20:34:03.740Z",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> Have you done it? (I mean the c++ transform/rasterizing)</p>",
          "rawMarkdown": "@ilu000 Have you done it? (I mean the c++ transform/rasterizing)"
        },
        {
          "id": 1044527,
          "postDate": "2020-10-09T21:01:05.403Z",
          "content": "<p>Not yet, I was thinking about doing so, but I think it will need a bigger rewrite as the I/O from python to C++ and back would become the new bottleneck. Numpy is already pretty fast, but here we are using non-contiguous arrays when calling the function, which makes it a little tricky. </p>",
          "rawMarkdown": "Not yet, I was thinking about doing so, but I think it will need a bigger rewrite as the I/O from python to C++ and back would become the new bottleneck. Numpy is already pretty fast, but here we are using non-contiguous arrays when calling the function, which makes it a little tricky. "
        },
        {
          "id": 1045315,
          "postDate": "2020-10-10T14:07:02.947Z",
          "content": "<p>C++ is a great on its speed butwe need to recode all of the python code to it is pretty well harder<br>\nWhat are you you people use GPU or TPU can you share what time it take for some specific batches and time</p>",
          "rawMarkdown": "C++ is a great on its speed butwe need to recode all of the python code to it is pretty well harder\nWhat are you you people use GPU or TPU can you share what time it take for some specific batches and time"
        },
        {
          "id": 1045515,
          "postDate": "2020-10-10T17:58:47.533Z",
          "content": "<p>1 epoch of train_full.zarr would take me about 22days. As the bottleneck is the rasterization it wouldn't matter if you use GPU or TPU</p>",
          "rawMarkdown": "1 epoch of train_full.zarr would take me about 22days. As the bottleneck is the rasterization it wouldn't matter if you use GPU or TPU"
        },
        {
          "id": 1045734,
          "postDate": "2020-10-11T01:28:49.633Z",
          "content": "<p>But when i trained on TPU 10000k batches was pretty well faster than GPU which takes 40 days</p>",
          "rawMarkdown": "But when i trained on TPU 10000k batches was pretty well faster than GPU which takes 40 days"
        },
        {
          "id": 1046555,
          "postDate": "2020-10-11T19:11:10.083Z",
          "content": "<p>my training stopped if i use pytorch xla on tpu.. do u face similar issue <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> ?</p>",
          "rawMarkdown": "my training stopped if i use pytorch xla on tpu.. do u face similar issue @morizin ?"
        },
        {
          "id": 1046795,
          "postDate": "2020-10-12T02:42:02.860Z",
          "content": "<p>No <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> </p>",
          "rawMarkdown": "No @seshurajup "
        },
        {
          "id": 1056993,
          "postDate": "2020-10-22T08:45:51.700Z",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> how did you load the full dataset.its alerting me \"Allocating more memory in Kaggle\" Can you say how you do it? either if online learning </p>",
          "rawMarkdown": "@ilu000 how did you load the full dataset.its alerting me \"Allocating more memory in Kaggle\" Can you say how you do it? either if online learning "
        },
        {
          "id": 1057019,
          "postDate": "2020-10-22T09:32:44.083Z",
          "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> sounds like a PyTorch XLA related problem, it gets reported a lot. What notebook did you use to get into PyTorch XLA?</p>",
          "rawMarkdown": "@seshurajup sounds like a PyTorch XLA related problem, it gets reported a lot. What notebook did you use to get into PyTorch XLA?"
        },
        {
          "id": 1057031,
          "postDate": "2020-10-22T09:40:58.427Z",
          "content": "<p>it was just looked , the inference is bit unstable</p>",
          "rawMarkdown": "it was just looked , the inference is bit unstable"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1045189,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-10T12:23:27.673000",
      "content": "<p>Can anyone have a way to split the batches that is <br>\nif we cant train for 1million batches at one go<br>\ncan we train one session with 10000 and refresh the kernel and again 10000<br>\nCan anyone know how to do it<br>\nQuestion to everyone also <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> <a href=\"https://www.kaggle.com/pondruska\" target=\"_blank\">@pondruska</a> <br>\nSorry if the question is already answered</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1045274,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-10T13:31:15.783000",
          "content": "<p>If you dont use shuffeling you could do it with something like:</p>\n<pre><code>for itr in progress_bar:\n        if itr &lt;= pretrain_iteration:\n            continue\n...\n</code></pre>\n<p>with <code>pretrain_iteration</code> being 10000 in your case. </p>\n<p>With shuffeling off you make sure, that you dont let the model learn the same sample again.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1045302,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-10T13:58:14.707000",
          "content": "<p>Can we do seed and shuffle it will not any problem doesnt it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1045371,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-10T15:03:09.230000",
          "content": "<p>No it wont cause any problems, you might depending on the size of <code>pretrain_iteration</code>, teach your model the same input multiple times but thats equal to training it over a set amount of epochs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1046105,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-10-11T10:53:15.573000",
          "content": "<p>I have similar problem. I am working from home so I can only use my PC for training at night. I hope to resume my training at the same place of dataloader as I stopped last time.</p>\n<p>Here is how I do it. Hope it makes sense (if it doesn't please let me know):</p>\n<pre><code>random.seed(2020)\n\nclass MySampler(Sampler):\n    def __init__(self, data, i=0):\n        self.seq = list(range(len(data)))\n        random.shuffle(self.seq)\n        self.seq = self.seq[i * train_cfg[\"batch_size\"]:]\n\n    def __iter__(self):\n        return iter(self.seq)\n\n    def __len__(self):\n        return len(self.seq)\n\nsampler = MySampler(train_dataset, i = starting)\ntrain_dataloader = DataLoader(train_dataset, batch_size=train_cfg[\"batch_size\"],        \n                         sampler=sampler, num_workers=train_cfg[\"num_workers\"])\n</code></pre>\n<p>if you stopped at 100k iterations last time, let starting = 100000</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1046260,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-11T14:05:46.210000",
          "content": "<p>Here is some solution i could find out from kernels hope it works</p>\n<pre><code>tr_it = iter(train_dataloader)\n\nprogress_bar = tqdm(range(iterno,cfg[\"train_params\"][\"max_num_steps\"]))\nlosses_train = []\nfor itr in progress_bar:\n    try:\n        data = next(tr_it)\n\n    except StopIteration:\n\n        tr_it = iter(train_dataloader)\n        data = next(tr_it)\n\n    model.train()\n    torch.set_grad_enabled(True)\n</code></pre>\n<p>where we save iter no with the checkpoint to not lost our iterations</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046549,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-11T19:03:49.533000",
          "content": "<p>don't forget to also save the optimizer state and the current step for the scheduler.<br>\nsee <a href=\"https://pytorch.org/tutorials/beginner/saving_loading_models.html\" target=\"_blank\">here</a> for details how this is done in pytorch. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1046792,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-12T02:40:48.763000",
          "content": "<p>Ya  I forgot about it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062067,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-10-27T14:44:58.317000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>  I'm having problem with loading optimizer's state. It throws up errors like </p>\n<pre><code>runtimeerror: expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!\n</code></pre>\n<p>After some browsing I found a way to move all the tensors to cuda by</p>\n<pre><code>    checkpoint = torch.load(weight_path)\n    model.load_state_dict(checkpoint['model_state_dict'])\n    model.cuda()\n    optimizer = optim.Adadelta(model.parameters())\n    optimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n    for state in optimizer.state.values():\n        for k, v in state.items():\n            if isinstance(v, torch.Tensor):\n                state[k] = v.cuda()\n</code></pre>\n<p>But still I'm unable to resume my training and loss shoots up to a large value (~4000) but drops after few iterations (~400)</p>\n<p>Did you face any such issues </p>\n<pre><code>This happened when I used the weights trained over Kaggle's GPU on my PC \nMy Setup\nOS : Win 10, torch 1.6.0\nGPU : RTX 2060\n</code></pre>\n<p>Any ideas??</p>\n<p>In simple I just wanna resume training offline……..</p>\n<p>Thanks </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062076,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-27T14:59:21.653000",
          "content": "<p>Right now i added a sampler</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062097,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-10-27T15:09:45.993000",
          "content": "<p>Were you able to continue from the checkpoint?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062137,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-27T15:46:27.407000",
          "content": "<p>Yes, I had a successful in them after a long time</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062141,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-10-27T15:48:18.900000",
          "content": "<p>Nice ..Did it require saving optimizer and scheduler state ?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062151,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-27T15:55:14.360000",
          "content": "<p>i save optimizer , scheduler , sampler, model weights all the things</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1065800,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-31T19:04:42.377000",
          "content": "<p><a href=\"https://www.kaggle.com/ak0210\" target=\"_blank\">@ak0210</a> you should try this</p>\n<p>device = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")<br>\ncheckpoint = torch.load(weight_path)<br>\nmodel.load_state_dict(checkpoint['model_state_dict'])<br>\nmodel.to(device)<br>\noptimizer = optim.Adadelta(model.parameters())<br>\noptimizer.to(device) <strong><em>if you want your optimizer in cuda else no need to this</em></strong><br>\noptimizer.load_state_dict(checkpoint['optimizer_state_dict'])</p>\n<p><strong>This is the procedure for loading model and optimizer in pytorch</strong></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1067694,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-11-02T16:26:09.437000",
          "content": "<p>Thanks for the suggestion <a href=\"https://www.kaggle.com/tanishgupta34\" target=\"_blank\">@tanishgupta34</a> but still ain't working, getting the error..<br>\nThe error pops up when I try to load to GPU only</p>\n<pre><code>&lt;ipython-input-17-4a4c8090f612&gt; in &lt;module&gt;\n      6 model.to(device)\n      7 optimizer = optim.Adadelta(model.parameters())\n----&gt; 8 optimizer.to(device)\n      9 optimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n\nAttributeError: 'Adadelta' object has no attribute 'to'\n</code></pre>\n<p>My solution on previous reply does the trick </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1070085,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-05T10:24:57.767000",
          "content": "<p><a href=\"https://www.kaggle.com/ak0210\" target=\"_blank\">@ak0210</a> yeah this was working fine on 1.5.1 but now in version 1.6.0 it is showing error.You can try once by removing optimizer.to(device) and it will work fine</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1070406,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-11-05T18:14:10.920000",
          "content": "<p><a href=\"https://www.kaggle.com/tanishgupta34\" target=\"_blank\">@tanishgupta34</a>  yes, it works fine without <code>optimizer.to(device)</code> but I find the loss values exploding.<br>\nWhich is partly resolved in by the solution in the above comment. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1044325,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2020-10-09T17:11:11.260000",
      "content": "<p>With more dataloader workers your rasterization step which is the main bottleneck can be executed faster. I am running 40 workers on a 6 core machine.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 1044519,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-09T20:44:33.030000",
          "content": "<p>Is the difference noticeable?</p>\n<p>When I put <code>num_workers</code> higher than 4 I cant see a difference anymore.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1044531,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-09T21:08:43.067000",
          "content": "<p>From my (limited) experience more than 2 workers per core are usually not beneficial. <br>\nHere, I didn't even notice a difference between 1 worker per core and 2 workers per core. </p>\n<p>Kaggle kernels have 2 cores, if you were wondering. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1045135,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-10T11:23:23.857000",
          "content": "<p>That makes sense, so tweaking the <code>num_worker</code> count seems to be useless change. </p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1059374,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-25T02:44:41.320000",
      "content": "<p><a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> you said you training loss was 24.xxx in this Public notebook. But it seems to be 324.xxx when we start the training with that loss</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1059364,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-25T02:18:26.950000",
      "content": "<p>If we training train_full.zarr by using l5kit do we need 64GB of Ram to actually store them</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1059624,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-25T09:41:03.267000",
          "content": "<p>32 GB seems to work fine. Load the train set before the eval set. After finishing the loading process, the blocked memory will drop significantly. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059653,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-25T10:16:32.957000",
          "content": "<p>i know that it will get down after finishing. I just looked up in the kaggle kernel but it was allocating more memory</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1059720,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-25T11:39:22.440000",
          "content": "<p>Yes, for kaggle kernels it's too large (probably the reason why they have the smaller train.zarr included here by default).<br>\nI was just saying that the minimum requirement is 32 GB and not 64 GB. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059725,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-25T11:46:06.203000",
          "content": "<p>Thank you, But how long it takes to run 1 epoch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059804,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-25T13:33:59.613000",
          "content": "<blockquote>\n  <p>1 epoch of train_full.zarr would take me about 22days</p>\n</blockquote>\n<p>you could just scroll down and look at my comment from 15 days ago</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1059821,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-25T13:54:14.360000",
          "content": "<p>Oh sorry, I thought it was for train.zarr. Okay how long it takes for train.zarr.<br>\nAlso, I have a doubt on my solution itself. I have given a <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/189996#1046260\" target=\"_blank\">code solution</a> in this discussion for resuming the iteration<br>\nbut I have a doubt, when we resume is it starts from point1 or from the iter no <br>\nsince I am doing <br>\n<code>data = next(tr_it)</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059834,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-25T14:09:58.003000",
          "content": "<p>train_full has 191_177_863 items. Standard train has what, ~23 million? <br>\nSize ~1/9, so time would also be ~1/9</p>\n<p>22 days / 9 = 2.5 days</p>\n<p>edit: seeing your edit above, you need to save the current point of your dataloader as well (and save your shuffle seed) if you want to continue training. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1059839,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-25T14:15:37.533000",
          "content": "<p>I have unique seed. But what about saving DataLoader<br>\nWhen I try to train train.zarr it shows it take 624hours on kaggle kernel<br>\nbut on colab its 3x faster<br>\nand 4x faster in my Tesla T4 (4GPU)<br>\nAnd are you saying you are running on GPU<br>\nWhats your GPU</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059951,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-25T16:05:40.597000",
          "content": "<p>Why is my code takes 20 days to finish on train.zarr? Can you fix them please</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059964,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-25T16:18:31.843000",
          "content": "<p>mixed precision, raster_size, CPU speed, lots of things why your runtime may be longer. <br>\nI suggest you go ahead and learn how to profile your code, find the bottleneck and then ask more specific questions. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1059974,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-25T16:30:54.253000",
          "content": "",
          "votes": -4,
          "replies": []
        },
        {
          "id": 1060541,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-26T09:40:18.217000",
          "content": "",
          "votes": -4,
          "replies": []
        },
        {
          "id": 1060645,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-26T12:39:42.077000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1061823,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-27T11:01:04.687000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1062159,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-27T16:00:19.213000",
          "content": "<p>The CPU Speed of kaggle kernels is very low (and this is sadly the bottleneck here) and you only have two cores available. Even on nowadays multicore consumer CPUs (combined with an OK GPU) you can get to much faster training times.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1062165,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-27T16:04:42.693000",
          "content": "",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1062222,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-27T17:05:08.397000",
          "content": "<p>As stated <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/192458#1057112\" target=\"_blank\">here</a> 2 cores (kaggle kernels) give about 1 batch (16 frames) per second, which gets you to 16.6 days of training (23M frames, 16 frames/batch). Colab Pro high-memory VM for GPU: CPU cores = 4 can already make that 2.7 times faster. 32 cores should give you a 16 times speed-up: 1.1 days (maybe a little less because of the necessary overhead). I don't know why your training is not faster than on kaggle kernels. </p>\n<p>mixed precision makes use of less accurate floats (lesser memory footprint, faster I/O). You can read about it <a href=\"https://pytorch.org/blog/accelerating-training-on-nvidia-gpus-with-pytorch-automatic-mixed-precision/\" target=\"_blank\">here</a>. It can also make your training slightly faster. Decreasing the raster_size has an influence on your training time as your arrays can be smaller but it's still limited by the transformations as stated <a href=\"https://github.com/lyft/l5kit/issues/136\" target=\"_blank\">here</a>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1062226,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-27T17:08:53.733000",
          "content": "",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1066238,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-01T14:17:04.750000",
          "content": "",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1071115,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-11-06T14:14:54.610000",
          "content": "<p>You mean if i have 32 vCPU then I should select num_workers = 32 will use all my CPU or 64</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1071502,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-07T00:12:26.093000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> he meant exactly the same you told and you can check the numworkers by using <strong>os.cpu_count()</strong> if you have any confusion </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1057203,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-22T13:19:40.063000",
      "content": "<p>Do anyone know that is this <a href=\"https://www.kaggle.com/ryanholbrook/lyft-motion-prediction-tfrecords\" target=\"_blank\">data</a> is train_full.zarr or train.zarr file?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1057205,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-10-22T13:23:13.643000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> you will find full data here <a href=\"url\" target=\"_blank\">https://www.kaggle.com/philculliton/lyft-full-training-set</a>. And train.zarr you can download from data section of commpetition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057210,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-22T13:25:59.490000",
          "content": "<p><a href=\"https://www.kaggle.com/deepakrajpurushothaman\" target=\"_blank\">@deepakrajpurushothaman</a> is that data is Full data, So that I could decide to train on zarr file or TfRec Files</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1047356,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-12T14:08:30.050000",
      "content": "<p>Is anyone is using Tensorflow if so how did you manage to do Dataset</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1046031,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-10-11T09:30:24.197000",
      "content": "<p>Does the score on eval set while training will project out to leaderboard</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1046100,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-11T10:48:23.680000",
          "content": "<p>See the discussion <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695\" target=\"_blank\">here</a>.</p>\n<p>Using the correct validation method the LB score correlates very well with the validation score</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1045166,
      "author_name": "n3n7i",
      "author_url": "",
      "post_date": "2020-10-10T11:56:59.667000",
      "content": "<p>If you aren't training all the lower layers you might be able to forward prop once and store/run from there</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1044498,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2020-10-09T20:23:24.607000",
      "content": "<p>Write the transformation for the rasterizer in C++ or port the routine to GPU as this is the bottleneck. <br>\nsee discussion here <a href=\"https://github.com/lyft/l5kit/issues/136\" target=\"_blank\">https://github.com/lyft/l5kit/issues/136</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1044509,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-10-09T20:34:03.740000",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> Have you done it? (I mean the c++ transform/rasterizing)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1044527,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-09T21:01:05.403000",
          "content": "<p>Not yet, I was thinking about doing so, but I think it will need a bigger rewrite as the I/O from python to C++ and back would become the new bottleneck. Numpy is already pretty fast, but here we are using non-contiguous arrays when calling the function, which makes it a little tricky. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1045315,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-10T14:07:02.947000",
          "content": "<p>C++ is a great on its speed butwe need to recode all of the python code to it is pretty well harder<br>\nWhat are you you people use GPU or TPU can you share what time it take for some specific batches and time</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1045515,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-10-10T17:58:47.533000",
          "content": "<p>1 epoch of train_full.zarr would take me about 22days. As the bottleneck is the rasterization it wouldn't matter if you use GPU or TPU</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1045734,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-11T01:28:49.633000",
          "content": "<p>But when i trained on TPU 10000k batches was pretty well faster than GPU which takes 40 days</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046555,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2020-10-11T19:11:10.083000",
          "content": "<p>my training stopped if i use pytorch xla on tpu.. do u face similar issue <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046795,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-12T02:42:02.860000",
          "content": "<p>No <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1056993,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-22T08:45:51.700000",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> how did you load the full dataset.its alerting me \"Allocating more memory in Kaggle\" Can you say how you do it? either if online learning </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057019,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-10-22T09:32:44.083000",
          "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> sounds like a PyTorch XLA related problem, it gets reported a lot. What notebook did you use to get into PyTorch XLA?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057031,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-22T09:40:58.427000",
          "content": "<p>it was just looked , the inference is bit unstable</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1044250": "Can anyone suggest a way to run 10k samples on TPU or GPU faster",
    "1045189": "Can anyone have a way to split the batches that is \nif we cant train for 1million batches at one go\ncan we train one session with 10000 and refresh the kernel and again 10000\nCan anyone know how to do it\nQuestion to everyone also @iglovikov @lucabergamini @pondruska \nSorry if the question is already answered",
    "1044325": "With more dataloader workers your rasterization step which is the main bottleneck can be executed faster. I am running 40 workers on a 6 core machine.",
    "1059374": "@huanvo you said you training loss was 24.xxx in this Public notebook. But it seems to be 324.xxx when we start the training with that loss",
    "1059364": "If we training train_full.zarr by using l5kit do we need 64GB of Ram to actually store them",
    "1057203": "Do anyone know that is this [data](https://www.kaggle.com/ryanholbrook/lyft-motion-prediction-tfrecords) is train_full.zarr or train.zarr file?\n",
    "1047356": "Is anyone is using Tensorflow if so how did you manage to do Dataset",
    "1046031": "Does the score on eval set while training will project out to leaderboard",
    "1045166": "If you aren't training all the lower layers you might be able to forward prop once and store/run from there",
    "1044498": "Write the transformation for the rasterizer in C++ or port the routine to GPU as this is the bottleneck. \nsee discussion here https://github.com/lyft/l5kit/issues/136"
  }
}