{
  "id": 307135,
  "title": "RAM is overloaded pytorch-lightning TPU training ",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307135",
  "author_name": "",
  "post_date": "2022-02-12T18:50:21.326839900Z",
  "votes": 8,
  "comment_count": 16,
  "views": 0,
  "content": "<p>hi can you please take a look at my notebook Idk why it's crashing (the ram is getting overloaded even the batch size is 8 ) <br>\nI am trying the same as you did but with Pytorch lightning <br>\nit'll be a great help if you did <br>\n<a href=\"https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface\" target=\"_blank\">https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface</a></p>\n<p>thanks in advance</p>\n<p>I am trying to create the same architecture as <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> <br>\nhis notebook(tf) <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061</a>'<br>\nhis notebook is having a batch size of 32 * 8 and mine has 8 can anyone help? </p>",
  "messages": [
    {
      "id": "1687345",
      "postDate": "02/12/2022 18:50:21",
      "content": "<p>hi can you please take a look at my notebook Idk why it's crashing (the ram is getting overloaded even the batch size is 8 ) <br>\nI am trying the same as you did but with Pytorch lightning <br>\nit'll be a great help if you did <br>\n<a href=\"https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface\" target=\"_blank\">https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface</a></p>\n<p>thanks in advance</p>\n<p>I am trying to create the same architecture as <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> <br>\nhis notebook(tf) <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061</a>'<br>\nhis notebook is having a batch size of 32 * 8 and mine has 8 can anyone help? </p>",
      "rawMarkdown": "hi can you please take a look at my notebook Idk why it's crashing (the ram is getting overloaded even the batch size is 8 ) \nI am trying the same as you did but with Pytorch lightning \nit'll be a great help if you did \n[https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface](https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface)\n\nthanks in advance\n\nI am trying to create the same architecture as @ks2019 \nhis notebook(tf) [https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061](https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061)'\nhis notebook is having a batch size of 32 * 8 and mine has 8 can anyone help?",
      "votes": null
    },
    {
      "id": "1688188",
      "postDate": "02/13/2022 13:00:37",
      "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> have you tried using PyTorch profiler to check? </p>\n<p>I didn't get a chance to check your code yet-it might be due to image sizes or architecture or possibly not cleaning your memory before starting training. </p>\n<p><a href=\"https://wandb.ai/wandb/trace/reports/Using-the-PyTorch-Profiler-with-W-B--Vmlldzo5MDE3NjU\" target=\"_blank\">Here's an article on how to use PyTorch profiler with Weights and Biases</a></p>",
      "rawMarkdown": "somesh88 have you tried using PyTorch profiler to check? \n\nI didn't get a chance to check your code yet-it might be due to image sizes or architecture or possibly not cleaning your memory before starting training. \n\n[Here's an article on how to use PyTorch profiler with Weights and Biases](https://wandb.ai/wandb/trace/reports/Using-the-PyTorch-Profiler-with-W-B--Vmlldzo5MDE3NjU)",
      "votes": null
    },
    {
      "id": "1688223",
      "postDate": "02/13/2022 13:52:31",
      "content": "<p>hey thanks for recommendation I'll surely check it out before I used wandb just for plotting various metrics and creating some reports but this is new thing for me ( It might take some more time to implement and learn how to use it haha ) </p>",
      "rawMarkdown": "hey thanks for recommendation I'll surely check it out before I used wandb just for plotting various metrics and creating some reports but this is new thing for me ( It might take some more time to implement and learn how to use it haha )",
      "votes": null
    },
    {
      "id": "1688234",
      "postDate": "02/13/2022 14:11:03",
      "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> I just saw the tutorial for the profiler <img src=\"https://i.ibb.co/KNyxRCm/image.png\" alt=\"\"></p>\n<p>it says it's because of num workers but if we used num workers as 0 then won't it be really slow ? </p>",
      "rawMarkdown": "init27 I just saw the tutorial for the profiler ![](https://i.ibb.co/KNyxRCm/image.png)\n\nit says it's because of num workers but if we used num workers as 0 then won't it be really slow ?",
      "votes": null
    },
    {
      "id": "1688235",
      "postDate": "02/13/2022 14:11:31",
      "content": "<p>If anyone wants tutorial link <a href=\"https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e\" target=\"_blank\">https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e</a></p>",
      "rawMarkdown": "If anyone wants tutorial link [https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e](https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e)",
      "votes": null
    },
    {
      "id": "1688240",
      "postDate": "02/13/2022 14:19:58",
      "content": "<p>Oh wow, this wasn't written by me so let me check with my colleagues, I usually run with <code>-1</code> argument or <code>NUM_CPU_THREADS-2</code> to allow 2 threads for System IO. </p>\n<p>Either its a bug or we learned something new today 😄</p>",
      "rawMarkdown": "Oh wow, this wasn't written by me so let me check with my colleagues, I usually run with `-1` argument or `NUM_CPU_THREADS-2` to allow 2 threads for System IO. \n\nEither its a bug or we learned something new today 😄",
      "votes": null
    },
    {
      "id": "1688298",
      "postDate": "02/13/2022 15:06:47",
      "content": "<p>I even tried with num_workers = 0 but not working I raised issue on pytorch lightning's github discussion let's see haha</p>",
      "rawMarkdown": "I even tried with num_workers = 0 but not working I raised issue on pytorch lightning's github discussion let's see haha",
      "votes": null
    },
    {
      "id": "1688322",
      "postDate": "02/13/2022 15:22:23",
      "content": "<p>Keep us posted :)</p>",
      "rawMarkdown": "Keep us posted :)",
      "votes": null
    },
    {
      "id": "1688503",
      "postDate": "02/13/2022 17:02:02",
      "content": "<p>yes sure I'll try without TPU and let's see </p>",
      "rawMarkdown": "yes sure I'll try without TPU and let's see",
      "votes": null
    },
    {
      "id": "1688942",
      "postDate": "02/13/2022 22:47:18",
      "content": "<p>Hello! somesh88.</p>\n<p>Thank you for answer.</p>",
      "rawMarkdown": "Hello! somesh88.\n\nThank you for answer.",
      "votes": null
    },
    {
      "id": "1695433",
      "postDate": "02/18/2022 06:31:37",
      "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> This problem can arise if you are accumulating values somewhere (ex. appending any values in a list) and not being collected by garbage collector.  Try to free cuda memory also.</p>",
      "rawMarkdown": "somesh88 This problem can arise if you are accumulating values somewhere (ex. appending any values in a list) and not being collected by garbage collector.  Try to free cuda memory also.",
      "votes": null
    },
    {
      "id": "1695447",
      "postDate": "02/18/2022 06:42:11",
      "content": "<p>welcome :) </p>",
      "rawMarkdown": "welcome :)",
      "votes": null
    },
    {
      "id": "1695453",
      "postDate": "02/18/2022 06:43:08",
      "content": "<p>ohh I see but I am not collecting any gradients <br>\neverything is logged on by wandb lightning callback not sure why it's happening</p>",
      "rawMarkdown": "ohh I see but I am not collecting any gradients \neverything is logged on by wandb lightning callback not sure why it's happening",
      "votes": null
    },
    {
      "id": "1696149",
      "postDate": "02/18/2022 16:12:27",
      "content": "<blockquote>\n  <p>it says it's because of num workers but if we used num workers as 0 then won't it be really slow ?</p>\n</blockquote>\n<p>Regarding num_workers, once I faced a similar issue with RAM. The issue was related to python multiprocessing philosophy - multiprocessing just copies the whole python process N-times to achieve the power of the N-subprocess. So, if you loaded large object into the RAM (like dataset) and use multiprocessing - you are likely to run of memory</p>",
      "rawMarkdown": "> it says it's because of num workers but if we used num workers as 0 then won't it be really slow ?\n\nRegarding num_workers, once I faced a similar issue with RAM. The issue was related to python multiprocessing philosophy - multiprocessing just copies the whole python process N-times to achieve the power of the N-subprocess. So, if you loaded large object into the RAM (like dataset) and use multiprocessing - you are likely to run of memory",
      "votes": null
    },
    {
      "id": "1696238",
      "postDate": "02/18/2022 17:22:45",
      "content": "<p>hey, thanks for the clarification. I got it. but still, I think it's not using its whole potential while TensorFlow is working perfectly fine without any of ram full or other errors ;( </p>",
      "rawMarkdown": "hey, thanks for the clarification. I got it. but still, I think it's not using its whole potential while TensorFlow is working perfectly fine without any of ram full or other errors ;(",
      "votes": null
    },
    {
      "id": "1698168",
      "postDate": "02/20/2022 07:44:29",
      "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> Can you share the notebook with the outputs and where exactly the code is breaking?</p>",
      "rawMarkdown": "somesh88 Can you share the notebook with the outputs and where exactly the code is breaking?",
      "votes": null
    },
    {
      "id": "1698489",
      "postDate": "02/20/2022 12:17:46",
      "content": "<p>the notebook is not getting saved (even if I saved it it'll encounter error and you can't see whole notebook unless you copy and edit) </p>\n<p>I have shared quick save version in the post above you might take a look at that. <br>\n<a href=\"https://www.kaggle.com/jainishsavalia\" target=\"_blank\">@jainishsavalia</a> </p>",
      "rawMarkdown": "the notebook is not getting saved (even if I saved it it'll encounter error and you can't see whole notebook unless you copy and edit) \n\nI have shared quick save version in the post above you might take a look at that. \n@jainishsavalia",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1688188,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/13/2022 13:00:37",
      "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> have you tried using PyTorch profiler to check? </p>\n<p>I didn't get a chance to check your code yet-it might be due to image sizes or architecture or possibly not cleaning your memory before starting training. </p>\n<p><a href=\"https://wandb.ai/wandb/trace/reports/Using-the-PyTorch-Profiler-with-W-B--Vmlldzo5MDE3NjU\" target=\"_blank\">Here's an article on how to use PyTorch profiler with Weights and Biases</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1688223,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/13/2022 13:52:31",
          "content": "<p>hey thanks for recommendation I'll surely check it out before I used wandb just for plotting various metrics and creating some reports but this is new thing for me ( It might take some more time to implement and learn how to use it haha ) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688234,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/13/2022 14:11:03",
          "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> I just saw the tutorial for the profiler <img src=\"https://i.ibb.co/KNyxRCm/image.png\" alt=\"\"></p>\n<p>it says it's because of num workers but if we used num workers as 0 then won't it be really slow ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688235,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/13/2022 14:11:31",
          "content": "<p>If anyone wants tutorial link <a href=\"https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e\" target=\"_blank\">https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688240,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/13/2022 14:19:58",
          "content": "<p>Oh wow, this wasn't written by me so let me check with my colleagues, I usually run with <code>-1</code> argument or <code>NUM_CPU_THREADS-2</code> to allow 2 threads for System IO. </p>\n<p>Either its a bug or we learned something new today 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688298,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/13/2022 15:06:47",
          "content": "<p>I even tried with num_workers = 0 but not working I raised issue on pytorch lightning's github discussion let's see haha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688322,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/13/2022 15:22:23",
          "content": "<p>Keep us posted :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688503,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/13/2022 17:02:02",
          "content": "<p>yes sure I'll try without TPU and let's see </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696149,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "02/18/2022 16:12:27",
          "content": "<blockquote>\n  <p>it says it's because of num workers but if we used num workers as 0 then won't it be really slow ?</p>\n</blockquote>\n<p>Regarding num_workers, once I faced a similar issue with RAM. The issue was related to python multiprocessing philosophy - multiprocessing just copies the whole python process N-times to achieve the power of the N-subprocess. So, if you loaded large object into the RAM (like dataset) and use multiprocessing - you are likely to run of memory</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696238,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/18/2022 17:22:45",
          "content": "<p>hey, thanks for the clarification. I got it. but still, I think it's not using its whole potential while TensorFlow is working perfectly fine without any of ram full or other errors ;( </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1688942,
      "author_name": "ishizakireo",
      "author_url": "",
      "post_date": "02/13/2022 22:47:18",
      "content": "<p>Hello! somesh88.</p>\n<p>Thank you for answer.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1695447,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/18/2022 06:42:11",
          "content": "<p>welcome :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1695433,
      "author_name": "jainishsavalia",
      "author_url": "",
      "post_date": "02/18/2022 06:31:37",
      "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> This problem can arise if you are accumulating values somewhere (ex. appending any values in a list) and not being collected by garbage collector.  Try to free cuda memory also.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1695453,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/18/2022 06:43:08",
          "content": "<p>ohh I see but I am not collecting any gradients <br>\neverything is logged on by wandb lightning callback not sure why it's happening</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698168,
          "author_name": "jainishsavalia",
          "author_url": "",
          "post_date": "02/20/2022 07:44:29",
          "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> Can you share the notebook with the outputs and where exactly the code is breaking?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698489,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/20/2022 12:17:46",
          "content": "<p>the notebook is not getting saved (even if I saved it it'll encounter error and you can't see whole notebook unless you copy and edit) </p>\n<p>I have shared quick save version in the post above you might take a look at that. <br>\n<a href=\"https://www.kaggle.com/jainishsavalia\" target=\"_blank\">@jainishsavalia</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1687345": "hi can you please take a look at my notebook Idk why it's crashing (the ram is getting overloaded even the batch size is 8 ) \nI am trying the same as you did but with Pytorch lightning \nit'll be a great help if you did \n[https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface](https://www.kaggle.com/somesh88/happywhale-pytorch-lightning-yolo-arcface)\n\nthanks in advance\n\nI am trying to create the same architecture as @ks2019 \nhis notebook(tf) [https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061](https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu/notebook?scriptVersionId=86802061)'\nhis notebook is having a batch size of 32 * 8 and mine has 8 can anyone help?",
    "1688188": "somesh88 have you tried using PyTorch profiler to check? \n\nI didn't get a chance to check your code yet-it might be due to image sizes or architecture or possibly not cleaning your memory before starting training. \n\n[Here's an article on how to use PyTorch profiler with Weights and Biases](https://wandb.ai/wandb/trace/reports/Using-the-PyTorch-Profiler-with-W-B--Vmlldzo5MDE3NjU)",
    "1688223": "hey thanks for recommendation I'll surely check it out before I used wandb just for plotting various metrics and creating some reports but this is new thing for me ( It might take some more time to implement and learn how to use it haha )",
    "1688234": "init27 I just saw the tutorial for the profiler ![](https://i.ibb.co/KNyxRCm/image.png)\n\nit says it's because of num workers but if we used num workers as 0 then won't it be really slow ?",
    "1688235": "If anyone wants tutorial link [https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e](https://colab.research.google.com/github/wandb/examples/blob/master/colabs/pytorch-lightning/Profile_PyTorch_Code.ipynb#scrollTo=81_Q15GTYy3e)",
    "1688240": "Oh wow, this wasn't written by me so let me check with my colleagues, I usually run with `-1` argument or `NUM_CPU_THREADS-2` to allow 2 threads for System IO. \n\nEither its a bug or we learned something new today 😄",
    "1688298": "I even tried with num_workers = 0 but not working I raised issue on pytorch lightning's github discussion let's see haha",
    "1688322": "Keep us posted :)",
    "1688503": "yes sure I'll try without TPU and let's see",
    "1688942": "Hello! somesh88.\n\nThank you for answer.",
    "1695433": "somesh88 This problem can arise if you are accumulating values somewhere (ex. appending any values in a list) and not being collected by garbage collector.  Try to free cuda memory also.",
    "1695447": "welcome :)",
    "1695453": "ohh I see but I am not collecting any gradients \neverything is logged on by wandb lightning callback not sure why it's happening",
    "1696149": "> it says it's because of num workers but if we used num workers as 0 then won't it be really slow ?\n\nRegarding num_workers, once I faced a similar issue with RAM. The issue was related to python multiprocessing philosophy - multiprocessing just copies the whole python process N-times to achieve the power of the N-subprocess. So, if you loaded large object into the RAM (like dataset) and use multiprocessing - you are likely to run of memory",
    "1696238": "hey, thanks for the clarification. I got it. but still, I think it's not using its whole potential while TensorFlow is working perfectly fine without any of ram full or other errors ;(",
    "1698168": "somesh88 Can you share the notebook with the outputs and where exactly the code is breaking?",
    "1698489": "the notebook is not getting saved (even if I saved it it'll encounter error and you can't see whole notebook unless you copy and edit) \n\nI have shared quick save version in the post above you might take a look at that. \n@jainishsavalia"
  },
  "source": "meta"
}