{
  "id": 401580,
  "title": "cannot utilize GPU resources in pytorch (using datasets and dataloaders)",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/401580",
  "author_name": "",
  "post_date": "2023-04-13T23:43:13.619596500Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>have been running my notebook for almost an hour unaccelerated, how long should the submission normally take? how can i view log traces on the submission itself to see what is taking so long?</p>\n<p>EDIT:<br>\nit seems i am unable to utilize GPU resources even though I am setting torch.device(\"cuda:0\") when available and .to(DEVICE) on both model and tensors.</p>\n<p>this is my starter notebook <a href=\"https://www.kaggle.com/alelat/start-cuda\" target=\"_blank\">https://www.kaggle.com/alelat/start-cuda</a> </p>\n<p>in the code what takes so long is looping through my dataloader. i know you can play with num_workers on the dataloader but when i do i get this error</p>\n<p>RuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method</p>\n<p>given this open issue on pytorch github</p>\n<p><a href=\"https://github.com/pytorch/pytorch/issues/40403\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/40403</a></p>\n<p>i tried forcing spawn on multiprocessing </p>\n<p>multiprocessing.set_start_method('spawn', force=True)</p>\n<p>to which i get this error</p>\n<p>File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 105, in spawn_main<br>\n    exitcode = _main(fd)<br>\n  File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 115, in _main<br>\n    self = reduction.pickle.load(from_parent)<br>\nAttributeError: Can't get attribute 'FOGDataset' on </p>\n<p>Out of the 400+ subs I can imagine many have successfully used GPU's in pytorch esp given that on CPU only my submission on a simple LSTM gives runtime errors.</p>\n<p>does anyone have a better way of using GPU's with Dataloaders and Datasets?</p>",
  "messages": [
    {
      "id": "2221051",
      "postDate": "04/13/2023 23:43:13",
      "content": "<p>have been running my notebook for almost an hour unaccelerated, how long should the submission normally take? how can i view log traces on the submission itself to see what is taking so long?</p>\n<p>EDIT:<br>\nit seems i am unable to utilize GPU resources even though I am setting torch.device(\"cuda:0\") when available and .to(DEVICE) on both model and tensors.</p>\n<p>this is my starter notebook <a href=\"https://www.kaggle.com/alelat/start-cuda\" target=\"_blank\">https://www.kaggle.com/alelat/start-cuda</a> </p>\n<p>in the code what takes so long is looping through my dataloader. i know you can play with num_workers on the dataloader but when i do i get this error</p>\n<p>RuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method</p>\n<p>given this open issue on pytorch github</p>\n<p><a href=\"https://github.com/pytorch/pytorch/issues/40403\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/40403</a></p>\n<p>i tried forcing spawn on multiprocessing </p>\n<p>multiprocessing.set_start_method('spawn', force=True)</p>\n<p>to which i get this error</p>\n<p>File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 105, in spawn_main<br>\n    exitcode = _main(fd)<br>\n  File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 115, in _main<br>\n    self = reduction.pickle.load(from_parent)<br>\nAttributeError: Can't get attribute 'FOGDataset' on </p>\n<p>Out of the 400+ subs I can imagine many have successfully used GPU's in pytorch esp given that on CPU only my submission on a simple LSTM gives runtime errors.</p>\n<p>does anyone have a better way of using GPU's with Dataloaders and Datasets?</p>",
      "rawMarkdown": "have been running my notebook for almost an hour unaccelerated, how long should the submission normally take? how can i view log traces on the submission itself to see what is taking so long?\n\nEDIT:\nit seems i am unable to utilize GPU resources even though I am setting torch.device(\"cuda:0\") when available and .to(DEVICE) on both model and tensors.\n\nthis is my starter notebook https://www.kaggle.com/alelat/start-cuda \n\nin the code what takes so long is looping through my dataloader. i know you can play with num_workers on the dataloader but when i do i get this error\n\nRuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method\n\ngiven this open issue on pytorch github\n\nhttps://github.com/pytorch/pytorch/issues/40403\n\ni tried forcing spawn on multiprocessing \n\nmultiprocessing.set_start_method('spawn', force=True)\n\nto which i get this error\n\n File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 105, in spawn_main\n    exitcode = _main(fd)\n  File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 115, in _main\n    self = reduction.pickle.load(from_parent)\nAttributeError: Can't get attribute 'FOGDataset' on <module '__main__' (built-in)>\n\nOut of the 400+ subs I can imagine many have successfully used GPU's in pytorch esp given that on CPU only my submission on a simple LSTM gives runtime errors.\n\ndoes anyone have a better way of using GPU's with Dataloaders and Datasets?",
      "votes": null
    },
    {
      "id": "2221696",
      "postDate": "04/14/2023 13:30:07",
      "content": "<p>It depends on your code. Your code will be run, and the submission file will be evaluated. It is better if you can run your code on your computer first to make sure it works. Then you submit it in a notebook for evaluation. If you run your notebook on a CPU, whatever time you got from your personal computer, double or triple it for the submission notebook.</p>",
      "rawMarkdown": "It depends on your code. Your code will be run, and the submission file will be evaluated. It is better if you can run your code on your computer first to make sure it works. Then you submit it in a notebook for evaluation. If you run your notebook on a CPU, whatever time you got from your personal computer, double or triple it for the submission notebook.",
      "votes": null
    },
    {
      "id": "2221767",
      "postDate": "04/14/2023 14:49:41",
      "content": "<p>So my notebook takes less than 10 mins to run from start to finish where it outputs a submission.csv file correctly formatted. </p>\n<p>With CPU only it took 16 hrs, notebook timed-out without score.</p>\n<p>I am now 7 hrs into a GPU run submission.</p>\n<p>Quite confused by the difference as it a very basic LSTM model, reducing mem on datasets, 1000 test batch size. its very basic uncomplicated code that as I said runs all cells in less that 10 mins normally.</p>\n<p>The frustration is not being able to see the output. the log trace goes up to 755s or 12m and nothing more. There are no while loops where it can go into infinite recursion. Quite perplexed by this experience as it my first competition sub</p>",
      "rawMarkdown": "So my notebook takes less than 10 mins to run from start to finish where it outputs a submission.csv file correctly formatted. \n\nWith CPU only it took 16 hrs, notebook timed-out without score.\n\nI am now 7 hrs into a GPU run submission.\n\nQuite confused by the difference as it a very basic LSTM model, reducing mem on datasets, 1000 test batch size. its very basic uncomplicated code that as I said runs all cells in less that 10 mins normally.\n\nThe frustration is not being able to see the output. the log trace goes up to 755s or 12m and nothing more. There are no while loops where it can go into infinite recursion. Quite perplexed by this experience as it my first competition sub",
      "votes": null
    },
    {
      "id": "2221773",
      "postDate": "04/14/2023 14:55:58",
      "content": "<p>what is ever more strange is that 755 seconds is actually where the test ends. as in it has run all the code required and saved the submission dataframe like so sub.to_csv('submission.csv', index=False)</p>\n<p>perhaps the file path is incorrect?</p>",
      "rawMarkdown": "what is ever more strange is that 755 seconds is actually where the test ends. as in it has run all the code required and saved the submission dataframe like so sub.to_csv('submission.csv', index=False)\n\nperhaps the file path is incorrect?",
      "votes": null
    },
    {
      "id": "2222561",
      "postDate": "04/15/2023 11:06:47",
      "content": "<p>One good metric to use is to see how much time it takes to do a prediction on the <code>defog</code> and <code>tdcsfog</code> folders of the train dataset. It could weed out some inefficiencies in you code. For me a single model inference GPU notebook takes about 23-25 minutes to complete with the hidden data set while the visible test data takes 2 min and 31 seconds.</p>",
      "rawMarkdown": "One good metric to use is to see how much time it takes to do a prediction on the `defog` and `tdcsfog` folders of the train dataset. It could weed out some inefficiencies in you code. For me a single model inference GPU notebook takes about 23-25 minutes to complete with the hidden data set while the visible test data takes 2 min and 31 seconds.",
      "votes": null
    },
    {
      "id": "2222811",
      "postDate": "04/15/2023 15:23:17",
      "content": "<p>ok i have completely missed the mark here with regards to utilizing GPU resources in pytorch, would you mind helping me out pls.</p>\n<p>it seems my notebook <a href=\"https://www.kaggle.com/alelat/start-cuda\" target=\"_blank\">https://www.kaggle.com/alelat/start-cuda</a> does not utilize GPU resources</p>\n<p>when i am specifying </p>\n<p>DEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")</p>\n<p>and then after either model.train() or model.eval() i specify</p>\n<p>model = model.to(DEVICE)</p>\n<p>annoyingly everytime i make a new tensor I also have to cast it tensor.to(DEVICE) and this is repeated in my custom Dataset class as well as every time i need to convert something into a tensor.</p>\n<p>what is the best way utilize the GPU correctly in pytorch given how i structured my code in the above notebook? </p>",
      "rawMarkdown": "ok i have completely missed the mark here with regards to utilizing GPU resources in pytorch, would you mind helping me out pls.\n\nit seems my notebook https://www.kaggle.com/alelat/start-cuda does not utilize GPU resources\n\nwhen i am specifying \n\nDEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\nand then after either model.train() or model.eval() i specify\n\nmodel = model.to(DEVICE)\n\nannoyingly everytime i make a new tensor I also have to cast it tensor.to(DEVICE) and this is repeated in my custom Dataset class as well as every time i need to convert something into a tensor.\n\nwhat is the best way utilize the GPU correctly in pytorch given how i structured my code in the above notebook?",
      "votes": null
    },
    {
      "id": "2222813",
      "postDate": "04/15/2023 15:26:11",
      "content": "<p>has anyone been able to train this on CPU only?</p>",
      "rawMarkdown": "has anyone been able to train this on CPU only?",
      "votes": null
    },
    {
      "id": "2222817",
      "postDate": "04/15/2023 15:28:58",
      "content": "<p>currently on CPU only this is training on average 45s per iteration and needs to run through 7000 ish iterations so it would mean 118 hours of training time. is this dataset impossible to train on CPU alone? Or am i being silly inefficient in my code somewhere?</p>",
      "rawMarkdown": "currently on CPU only this is training on average 45s per iteration and needs to run through 7000 ish iterations so it would mean 118 hours of training time. is this dataset impossible to train on CPU alone? Or am i being silly inefficient in my code somewhere?",
      "votes": null
    },
    {
      "id": "2222827",
      "postDate": "04/15/2023 15:40:28",
      "content": "<p>it seems my biggest time sink is the dataloader itself and it takes ages to iterate through each batch. i have tried specifying num_workers=MAX_GPU_WORKERS </p>\n<p>MAX_GPU_WORKERS = torch.cuda.device_count()</p>\n<p>but it is still crazy slow and also doesnt change the fact my GPU usage is 0</p>",
      "rawMarkdown": "it seems my biggest time sink is the dataloader itself and it takes ages to iterate through each batch. i have tried specifying num_workers=MAX_GPU_WORKERS \n\nMAX_GPU_WORKERS = torch.cuda.device_count()\n\nbut it is still crazy slow and also doesnt change the fact my GPU usage is 0",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2221696,
      "author_name": "jdeneva",
      "author_url": "",
      "post_date": "04/14/2023 13:30:07",
      "content": "<p>It depends on your code. Your code will be run, and the submission file will be evaluated. It is better if you can run your code on your computer first to make sure it works. Then you submit it in a notebook for evaluation. If you run your notebook on a CPU, whatever time you got from your personal computer, double or triple it for the submission notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2221767,
          "author_name": "alelat",
          "author_url": "",
          "post_date": "04/14/2023 14:49:41",
          "content": "<p>So my notebook takes less than 10 mins to run from start to finish where it outputs a submission.csv file correctly formatted. </p>\n<p>With CPU only it took 16 hrs, notebook timed-out without score.</p>\n<p>I am now 7 hrs into a GPU run submission.</p>\n<p>Quite confused by the difference as it a very basic LSTM model, reducing mem on datasets, 1000 test batch size. its very basic uncomplicated code that as I said runs all cells in less that 10 mins normally.</p>\n<p>The frustration is not being able to see the output. the log trace goes up to 755s or 12m and nothing more. There are no while loops where it can go into infinite recursion. Quite perplexed by this experience as it my first competition sub</p>",
          "votes": null,
          "replies": [
            {
              "id": 2221773,
              "author_name": "alelat",
              "author_url": "",
              "post_date": "04/14/2023 14:55:58",
              "content": "<p>what is ever more strange is that 755 seconds is actually where the test ends. as in it has run all the code required and saved the submission dataframe like so sub.to_csv('submission.csv', index=False)</p>\n<p>perhaps the file path is incorrect?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2222561,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "04/15/2023 11:06:47",
      "content": "<p>One good metric to use is to see how much time it takes to do a prediction on the <code>defog</code> and <code>tdcsfog</code> folders of the train dataset. It could weed out some inefficiencies in you code. For me a single model inference GPU notebook takes about 23-25 minutes to complete with the hidden data set while the visible test data takes 2 min and 31 seconds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2222811,
          "author_name": "alelat",
          "author_url": "",
          "post_date": "04/15/2023 15:23:17",
          "content": "<p>ok i have completely missed the mark here with regards to utilizing GPU resources in pytorch, would you mind helping me out pls.</p>\n<p>it seems my notebook <a href=\"https://www.kaggle.com/alelat/start-cuda\" target=\"_blank\">https://www.kaggle.com/alelat/start-cuda</a> does not utilize GPU resources</p>\n<p>when i am specifying </p>\n<p>DEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")</p>\n<p>and then after either model.train() or model.eval() i specify</p>\n<p>model = model.to(DEVICE)</p>\n<p>annoyingly everytime i make a new tensor I also have to cast it tensor.to(DEVICE) and this is repeated in my custom Dataset class as well as every time i need to convert something into a tensor.</p>\n<p>what is the best way utilize the GPU correctly in pytorch given how i structured my code in the above notebook? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2222817,
              "author_name": "alelat",
              "author_url": "",
              "post_date": "04/15/2023 15:28:58",
              "content": "<p>currently on CPU only this is training on average 45s per iteration and needs to run through 7000 ish iterations so it would mean 118 hours of training time. is this dataset impossible to train on CPU alone? Or am i being silly inefficient in my code somewhere?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2222827,
              "author_name": "alelat",
              "author_url": "",
              "post_date": "04/15/2023 15:40:28",
              "content": "<p>it seems my biggest time sink is the dataloader itself and it takes ages to iterate through each batch. i have tried specifying num_workers=MAX_GPU_WORKERS </p>\n<p>MAX_GPU_WORKERS = torch.cuda.device_count()</p>\n<p>but it is still crazy slow and also doesnt change the fact my GPU usage is 0</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2222813,
      "author_name": "alelat",
      "author_url": "",
      "post_date": "04/15/2023 15:26:11",
      "content": "<p>has anyone been able to train this on CPU only?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2221051": "have been running my notebook for almost an hour unaccelerated, how long should the submission normally take? how can i view log traces on the submission itself to see what is taking so long?\n\nEDIT:\nit seems i am unable to utilize GPU resources even though I am setting torch.device(\"cuda:0\") when available and .to(DEVICE) on both model and tensors.\n\nthis is my starter notebook https://www.kaggle.com/alelat/start-cuda \n\nin the code what takes so long is looping through my dataloader. i know you can play with num_workers on the dataloader but when i do i get this error\n\nRuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method\n\ngiven this open issue on pytorch github\n\nhttps://github.com/pytorch/pytorch/issues/40403\n\ni tried forcing spawn on multiprocessing \n\nmultiprocessing.set_start_method('spawn', force=True)\n\nto which i get this error\n\n File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 105, in spawn_main\n    exitcode = _main(fd)\n  File \"/opt/conda/lib/python3.7/multiprocessing/spawn.py\", line 115, in _main\n    self = reduction.pickle.load(from_parent)\nAttributeError: Can't get attribute 'FOGDataset' on <module '__main__' (built-in)>\n\nOut of the 400+ subs I can imagine many have successfully used GPU's in pytorch esp given that on CPU only my submission on a simple LSTM gives runtime errors.\n\ndoes anyone have a better way of using GPU's with Dataloaders and Datasets?",
    "2221696": "It depends on your code. Your code will be run, and the submission file will be evaluated. It is better if you can run your code on your computer first to make sure it works. Then you submit it in a notebook for evaluation. If you run your notebook on a CPU, whatever time you got from your personal computer, double or triple it for the submission notebook.",
    "2221767": "So my notebook takes less than 10 mins to run from start to finish where it outputs a submission.csv file correctly formatted. \n\nWith CPU only it took 16 hrs, notebook timed-out without score.\n\nI am now 7 hrs into a GPU run submission.\n\nQuite confused by the difference as it a very basic LSTM model, reducing mem on datasets, 1000 test batch size. its very basic uncomplicated code that as I said runs all cells in less that 10 mins normally.\n\nThe frustration is not being able to see the output. the log trace goes up to 755s or 12m and nothing more. There are no while loops where it can go into infinite recursion. Quite perplexed by this experience as it my first competition sub",
    "2221773": "what is ever more strange is that 755 seconds is actually where the test ends. as in it has run all the code required and saved the submission dataframe like so sub.to_csv('submission.csv', index=False)\n\nperhaps the file path is incorrect?",
    "2222561": "One good metric to use is to see how much time it takes to do a prediction on the `defog` and `tdcsfog` folders of the train dataset. It could weed out some inefficiencies in you code. For me a single model inference GPU notebook takes about 23-25 minutes to complete with the hidden data set while the visible test data takes 2 min and 31 seconds.",
    "2222811": "ok i have completely missed the mark here with regards to utilizing GPU resources in pytorch, would you mind helping me out pls.\n\nit seems my notebook https://www.kaggle.com/alelat/start-cuda does not utilize GPU resources\n\nwhen i am specifying \n\nDEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\nand then after either model.train() or model.eval() i specify\n\nmodel = model.to(DEVICE)\n\nannoyingly everytime i make a new tensor I also have to cast it tensor.to(DEVICE) and this is repeated in my custom Dataset class as well as every time i need to convert something into a tensor.\n\nwhat is the best way utilize the GPU correctly in pytorch given how i structured my code in the above notebook?",
    "2222813": "has anyone been able to train this on CPU only?",
    "2222817": "currently on CPU only this is training on average 45s per iteration and needs to run through 7000 ish iterations so it would mean 118 hours of training time. is this dataset impossible to train on CPU alone? Or am i being silly inefficient in my code somewhere?",
    "2222827": "it seems my biggest time sink is the dataloader itself and it takes ages to iterate through each batch. i have tried specifying num_workers=MAX_GPU_WORKERS \n\nMAX_GPU_WORKERS = torch.cuda.device_count()\n\nbut it is still crazy slow and also doesnt change the fact my GPU usage is 0"
  },
  "source": "meta"
}