{
  "id": 270612,
  "title": "How long to train one epoch?",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/270612",
  "author_name": "",
  "post_date": "2021-09-06T10:12:41.991087800Z",
  "votes": 15,
  "comment_count": 73,
  "views": 0,
  "content": "<p>So far, for one epoch and one single model on a 5 folds split, it takes around 1 hour to run on an<br>\n<strong>RTX 3090</strong>. </p>\n<p>So around <strong>5 hours</strong> for one full (all splits) <strong>epoch</strong>.</p>\n<p>I was wondering if these times are closer to what you have or if I can further optimize the code. Thanks for your help!</p>",
  "messages": [
    {
      "id": "1504365",
      "postDate": "09/06/2021 10:12:41",
      "content": "<p>So far, for one epoch and one single model on a 5 folds split, it takes around 1 hour to run on an<br>\n<strong>RTX 3090</strong>. </p>\n<p>So around <strong>5 hours</strong> for one full (all splits) <strong>epoch</strong>.</p>\n<p>I was wondering if these times are closer to what you have or if I can further optimize the code. Thanks for your help!</p>",
      "rawMarkdown": "So far, for one epoch and one single model on a 5 folds split, it takes around 1 hour to run on an\n**RTX 3090**. \n\nSo around **5 hours** for one full (all splits) **epoch**.\n\nI was wondering if these times are closer to what you have or if I can further optimize the code. Thanks for your help!",
      "votes": null
    },
    {
      "id": "1504370",
      "postDate": "09/06/2021 10:18:15",
      "content": "<p>that's a long time - I also have a 3090 and can train an epoch of b0 in less than 3mins. My best model is a b2 and it's current setup takes around 12mins an epoch</p>\n<p>my guess is you're either using a huge model or huge image or you have some bottleneck?</p>",
      "rawMarkdown": "that's a long time - I also have a 3090 and can train an epoch of b0 in less than 3mins. My best model is a b2 and it's current setup takes around 12mins an epoch\n\nmy guess is you're either using a huge model or huge image or you have some bottleneck?",
      "votes": null
    },
    {
      "id": "1504372",
      "postDate": "09/06/2021 10:19:02",
      "content": "<p>I assume you're doing all the usual mixed precision/pin_memory/etc optimisations?</p>",
      "rawMarkdown": "I assume you're doing all the usual mixed precision/pin_memory/etc optimisations?",
      "votes": null
    },
    {
      "id": "1504400",
      "postDate": "09/06/2021 10:45:22",
      "content": "<p>Not that big only a B3. Also, I am generating image features on the fly so probably not optimal for now. Thanks for providing these numbers, I will check what can be improved. 👌</p>",
      "rawMarkdown": "Not that big only a B3. Also, I am generating image features on the fly so probably not optimal for now. Thanks for providing these numbers, I will check what can be improved. 👌",
      "votes": null
    },
    {
      "id": "1504405",
      "postDate": "09/06/2021 10:51:49",
      "content": "<p><a href=\"https://www.kaggle.com/hamishdickson\" target=\"_blank\">@hamishdickson</a> <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> i am also experimenting using rtx3090 but when i converted wave signal into image and resized it down to 512x512 it takes around 1 hour to train a b0 model for 1 epoch and then my pc shuts down,i can't figure out why pc getting turned off during training,,,when i was using wave data only,,at that time training was fast and pc never got shut down,,but after converting it to 512x512 image size,pc shutting down after few hours of training(sometimes less than 1 hour)<br>\ni was working using this kernel : <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch</a><br>\nduring training i checked nvidia-smi and i was using less than 20gb vram out of 24</p>",
      "rawMarkdown": "hamishdickson @yassinealouini i am also experimenting using rtx3090 but when i converted wave signal into image and resized it down to 512x512 it takes around 1 hour to train a b0 model for 1 epoch and then my pc shuts down,i can't figure out why pc getting turned off during training,,,when i was using wave data only,,at that time training was fast and pc never got shut down,,but after converting it to 512x512 image size,pc shutting down after few hours of training(sometimes less than 1 hour)\ni was working using this kernel : https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\nduring training i checked nvidia-smi and i was using less than 20gb vram out of 24",
      "votes": null
    },
    {
      "id": "1504415",
      "postDate": "09/06/2021 11:02:37",
      "content": "<blockquote>\n  <p>and then my pc shuts down</p>\n</blockquote>\n<p>Undervolt your GPU and check if the problem goes away. I had the same problem when my old PSU gradually died - big power spikes during initialization forced PSU to shut down.</p>\n<p>3090, depending on model, can generate power spikes up to ~600W. Undervolting with 5-10% performance loss might keep them to ~350-400W max.</p>",
      "rawMarkdown": "> and then my pc shuts down\n\nUndervolt your GPU and check if the problem goes away. I had the same problem when my old PSU gradually died - big power spikes during initialization forced PSU to shut down.\n\n3090, depending on model, can generate power spikes up to ~600W. Undervolting with 5-10% performance loss might keep them to ~350-400W max.",
      "votes": null
    },
    {
      "id": "1504429",
      "postDate": "09/06/2021 11:14:08",
      "content": "<p>That's very interesting. I will check my wandb dashboard but I don't have those spikes I think. Thanks for sharing <a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a>. 👌 </p>",
      "rawMarkdown": "That's very interesting. I will check my wandb dashboard but I don't have those spikes I think. Thanks for sharing @fffrrt. 👌",
      "votes": null
    },
    {
      "id": "1504434",
      "postDate": "09/06/2021 11:25:50",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> thank you for sharing<br>\ni saw that my vram usage was low(less than 20gb from 24gb),so i thought cpu or ram or something else is causing this issue but i can't figure out,i checked kern log and found no problem,i thought it's num_workers that is generating more heat and causing this issue so set num_workers = 0 but no luck,i have gold certified good power supply,good cpu cooler but still for some reason my pc getting shut down,i hope your idea of Undervolt works for me,,otherwise i don't know how to solve this issue,thank you</p>",
      "rawMarkdown": "fffrrt thank you for sharing\ni saw that my vram usage was low(less than 20gb from 24gb),so i thought cpu or ram or something else is causing this issue but i can't figure out,i checked kern log and found no problem,i thought it's num_workers that is generating more heat and causing this issue so set num_workers = 0 but no luck,i have gold certified good power supply,good cpu cooler but still for some reason my pc getting shut down,i hope your idea of Undervolt works for me,,otherwise i don't know how to solve this issue,thank you",
      "votes": null
    },
    {
      "id": "1504438",
      "postDate": "09/06/2021 11:30:22",
      "content": "<blockquote>\n  <p>I will check my wandb dashboard but I don't have those spikes I think.</p>\n</blockquote>\n<p>They are not visible on monitoring software, because they are too short. Long enough for some PSUs to trip overcurrent protection, though.</p>\n<p>I could not find original igorslab article about them (when 3080/3090 on release crashed many PCs), only his testing after the drivers fix - <a href=\"https://www.igorslab.de/en/wonder-how-invidia-the-crashes-of-the-force-rtx-3080-andrtx-3090-will-be-removed-and-still-will-be-removed-even-from-the-power-supplies-analysis/\" target=\"_blank\">https://www.igorslab.de/en/wonder-how-invidia-the-crashes-of-the-force-rtx-3080-andrtx-3090-will-be-removed-and-still-will-be-removed-even-from-the-power-supplies-analysis/</a></p>\n<p>Illustration from there (3080, should be more or less actual behavior): <a href=\"https://www.igorslab.de/wp-content/uploads/2020/09/12a-Gaming-Zoom-Power-New-Drivers.png\" target=\"_blank\">https://www.igorslab.de/wp-content/uploads/2020/09/12a-Gaming-Zoom-Power-New-Drivers.png</a></p>\n<p>Edit: Found this topics that might prove helpful, i read them during my own issues - </p>\n<p><a href=\"https://discuss.pytorch.org/t/multi-gpu-2080-ti-training-crashes-pc/52032/9\" target=\"_blank\">https://discuss.pytorch.org/t/multi-gpu-2080-ti-training-crashes-pc/52032/9</a></p>\n<p><a href=\"https://github.com/pytorch/pytorch/issues/3022\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/3022</a></p>",
      "rawMarkdown": "> I will check my wandb dashboard but I don't have those spikes I think.\n\nThey are not visible on monitoring software, because they are too short. Long enough for some PSUs to trip overcurrent protection, though.\n\nI could not find original igorslab article about them (when 3080/3090 on release crashed many PCs), only his testing after the drivers fix - https://www.igorslab.de/en/wonder-how-invidia-the-crashes-of-the-force-rtx-3080-andrtx-3090-will-be-removed-and-still-will-be-removed-even-from-the-power-supplies-analysis/\n\nIllustration from there (3080, should be more or less actual behavior): https://www.igorslab.de/wp-content/uploads/2020/09/12a-Gaming-Zoom-Power-New-Drivers.png\n\nEdit: Found this topics that might prove helpful, i read them during my own issues - \n\nhttps://discuss.pytorch.org/t/multi-gpu-2080-ti-training-crashes-pc/52032/9\n\nhttps://github.com/pytorch/pytorch/issues/3022",
      "votes": null
    },
    {
      "id": "1504480",
      "postDate": "09/06/2021 12:18:15",
      "content": "<p>So I'm not alone. I had been having 3090 issues as well for the past few months. In my case, the computer does not freeze but training with segfault and I cannot reclaim the vram until I do a restart. Also it's usually GPU specific, that is, I can still train on the other GPU. Though occasionally when one blows, it takes both down. I have 1600W PSU though so I even if both pull spike their load, I wouldn't expect that to take down the ship. It is so annoying with these longer model train times, you let the models sail all night long only to wake up and discover stuff froze after the first fold.</p>",
      "rawMarkdown": "So I'm not alone. I had been having 3090 issues as well for the past few months. In my case, the computer does not freeze but training with segfault and I cannot reclaim the vram until I do a restart. Also it's usually GPU specific, that is, I can still train on the other GPU. Though occasionally when one blows, it takes both down. I have 1600W PSU though so I even if both pull spike their load, I wouldn't expect that to take down the ship. It is so annoying with these longer model train times, you let the models sail all night long only to wake up and discover stuff froze after the first fold.",
      "votes": null
    },
    {
      "id": "1504520",
      "postDate": "09/06/2021 12:52:03",
      "content": "<p><strong>you let the models sail all night long only to wake up and discover stuff froze after the first fold.</strong> 😪</p>",
      "rawMarkdown": "**you let the models sail all night long only to wake up and discover stuff froze after the first fold.** 😪",
      "votes": null
    },
    {
      "id": "1504580",
      "postDate": "09/06/2021 13:35:56",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> image size is also a factor.  Depending on the image size your timing may just be good, or bad.  One way to check is to look at the memory and compute use of your GPU using nvidia-smi.  </p>\n<p>For instance, if use is close to 100% and memory close to GPU memory then you are probably fine.  </p>\n<p>Another example: if use is 100% but memory is only half of GPU memory then you can probably double batch size and divide time by 2.  Of course, changing batch size may impact other hyper parameter values.</p>",
      "rawMarkdown": "yassinealouini image size is also a factor.  Depending on the image size your timing may just be good, or bad.  One way to check is to look at the memory and compute use of your GPU using nvidia-smi.  \n\nFor instance, if use is close to 100% and memory close to GPU memory then you are probably fine.  \n\nAnother example: if use is 100% but memory is only half of GPU memory then you can probably double batch size and divide time by 2.  Of course, changing batch size may impact other hyper parameter values.",
      "votes": null
    },
    {
      "id": "1504583",
      "postDate": "09/06/2021 13:41:24",
      "content": "<p>Not to hijack the thread and perhaps too close to competition deadline, but <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> if I might inquire, what bs are you currently utilizing in your model? It was something I'd actually been wondering. I'm currently using a modest 32 bs resulting in 98% utilization, however +85% vram is free… so I wonder.</p>",
      "rawMarkdown": "Not to hijack the thread and perhaps too close to competition deadline, but @cpmpml if I might inquire, what bs are you currently utilizing in your model? It was something I'd actually been wondering. I'm currently using a modest 32 bs resulting in 98% utilization, however +85% vram is free... so I wonder.",
      "votes": null
    },
    {
      "id": "1504588",
      "postDate": "09/06/2021 13:45:52",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> It looks like you should upgrade to the latest driver, and also check pour power unit.</p>",
      "rawMarkdown": "mobassir It looks like you should upgrade to the latest driver, and also check pour power unit.",
      "votes": null
    },
    {
      "id": "1504595",
      "postDate": "09/06/2021 13:51:44",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for the details. I will check these and report back later. 👌 </p>",
      "rawMarkdown": "Thanks @cpmpml for the details. I will check these and report back later. 👌",
      "votes": null
    },
    {
      "id": "1504597",
      "postDate": "09/06/2021 13:54:01",
      "content": "<p><a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a> previously i had 370watt limit for gpu and now after undervolting 70watt i have this : <br>\nMon Sep  6 19:12:16 2021       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 460.91.03    Driver Version: 460.91.03    CUDA Version: 11.2     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  GeForce RTX 3090    Off  | 00000000:01:00.0 Off |                  N/A |<br>\n|  0%   28C    P8    12W / 300W |     19MiB / 24268MiB |      0%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n|    0   N/A  N/A      1140      G   /usr/lib/xorg/Xorg                  9MiB |<br>\n|    0   N/A  N/A      1319      G   /usr/bin/gnome-shell                8MiB |<br>\n+-----------------------------------------------------------------------------+</p>\n<p><strong>nvcc --version</strong></p>\n<p>nvcc: NVIDIA (R) Cuda compiler driver<br>\nCopyright (c) 2005-2021 NVIDIA Corporation<br>\nBuilt on Sun_Feb_14_21:12:58_PST_2021<br>\nCuda compilation tools, release 11.2, V11.2.152<br>\nBuild cuda_11.2.r11.2/compiler.29618528_0</p>\n<p>can you suggest me which driver version i should try now? thank you</p>",
      "rawMarkdown": "cpmp previously i had 370watt limit for gpu and now after undervolting 70watt i have this : \nMon Sep  6 19:12:16 2021       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 460.91.03    Driver Version: 460.91.03    CUDA Version: 11.2     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce RTX 3090    Off  | 00000000:01:00.0 Off |                  N/A |\n|  0%   28C    P8    12W / 300W |     19MiB / 24268MiB |      0%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n|    0   N/A  N/A      1140      G   /usr/lib/xorg/Xorg                  9MiB |\n|    0   N/A  N/A      1319      G   /usr/bin/gnome-shell                8MiB |\n+-----------------------------------------------------------------------------+\n\n**nvcc --version**\n\nnvcc: NVIDIA (R) Cuda compiler driver\nCopyright (c) 2005-2021 NVIDIA Corporation\nBuilt on Sun_Feb_14_21:12:58_PST_2021\nCuda compilation tools, release 11.2, V11.2.152\nBuild cuda_11.2.r11.2/compiler.29618528_0\n\ncan you suggest me which driver version i should try now? thank you",
      "votes": null
    },
    {
      "id": "1504670",
      "postDate": "09/06/2021 14:58:03",
      "content": "<blockquote>\n  <p>however +85% vram is free… so I wonder.</p>\n</blockquote>\n<p>You should use  a larger batch size IMHO.  It is not what cause your problem of course, but you are under using your GPU it seems</p>",
      "rawMarkdown": "> however +85% vram is free… so I wonder.\n\nYou should use  a larger batch size IMHO.  It is not what cause your problem of course, but you are under using your GPU it seems",
      "votes": null
    },
    {
      "id": "1504693",
      "postDate": "09/06/2021 15:25:05",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> 😪<br>\npreviously i had 370watt limit for gpu and then i undervolted 70watt and limited gpu to use only 300watt.<br>\ni reduced image size from 512 to 256<br>\ni increased batch size (but vram usage was still less than 19gb out of 24)<br>\ni used mixed precision this time<br>\n<strong>PC SHUTS DOWN AGAIN AFTER 2nd EPOCH</strong>  😭😪</p>",
      "rawMarkdown": "fffrrt 😪\npreviously i had 370watt limit for gpu and then i undervolted 70watt and limited gpu to use only 300watt.\ni reduced image size from 512 to 256\ni increased batch size (but vram usage was still less than 19gb out of 24)\ni used mixed precision this time\n**PC SHUTS DOWN AGAIN AFTER 2nd EPOCH**  😭😪",
      "votes": null
    },
    {
      "id": "1504721",
      "postDate": "09/06/2021 15:54:25",
      "content": "<p>Are you caching tf valid dataset? If so try disabling it. Also try empty cache after every epoch. I haven't encountered this issue so I am just throwing ideas. </p>",
      "rawMarkdown": "Are you caching tf valid dataset? If so try disabling it. Also try empty cache after every epoch. I haven't encountered this issue so I am just throwing ideas.",
      "votes": null
    },
    {
      "id": "1504725",
      "postDate": "09/06/2021 15:56:52",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> could it be a hardware defect in the GPU? Have you tried contacting NVIDIA support? </p>",
      "rawMarkdown": "mobassir could it be a hardware defect in the GPU? Have you tried contacting NVIDIA support?",
      "votes": null
    },
    {
      "id": "1504766",
      "postDate": "09/06/2021 16:30:21",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I am using an image of size <strong>256 x 256</strong> and here are some utilization metrics over time. So not optimal for now, I will try to debug a bit and optimize tonight and update with new findings. </p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/16PVPs6/gpu-vram-and-utilization.png\" alt=\"gpu-vram-and-utilization\"></a></p>",
      "rawMarkdown": "cpmpml I am using an image of size **256 x 256** and here are some utilization metrics over time. So not optimal for now, I will try to debug a bit and optimize tonight and update with new findings. \n\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/16PVPs6/gpu-vram-and-utilization.png\" alt=\"gpu-vram-and-utilization\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "1504774",
      "postDate": "09/06/2021 16:33:50",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> i am not caching tf valid dataset,,i tried another baseline and encountered same issue,,for example if i experiment using this notebook : <a href=\"https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training</a><br>\nreplace b7 with b0<br>\nin custom model class use F.interpolate to convert wave data into 512x512 image size<br>\nadjust batch size that takes just less than 20gb vram and train <strong>(KEEP EVERYTHING AS IT IS IN THAT PUBLIC KERNEL)</strong></p>\n<p>still my pc turns off <strong>(this time after 5 epoch)</strong></p>\n<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> i Have not tried contacting NVIDIA support yet,when i do not convert wave data into 512x512 image data i don't face this issue,,i tried to train few nlp models and advance scene text detection/recognition models and didn't face this issue<br>\ni face this issue occasionally and i can't understand how to solve this issue</p>\n<p>instead of shutdown if i could use a command/setting that will restart the pc instead of freezing/turning it off then it could be good for me but i don't know if it's possible to do so </p>",
      "rawMarkdown": "pheadrus i am not caching tf valid dataset,,i tried another baseline and encountered same issue,,for example if i experiment using this notebook : https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\nreplace b7 with b0\nin custom model class use F.interpolate to convert wave data into 512x512 image size\nadjust batch size that takes just less than 20gb vram and train **(KEEP EVERYTHING AS IT IS IN THAT PUBLIC KERNEL)**\n\nstill my pc turns off **(this time after 5 epoch)**\n\n@yassinealouini i Have not tried contacting NVIDIA support yet,when i do not convert wave data into 512x512 image data i don't face this issue,,i tried to train few nlp models and advance scene text detection/recognition models and didn't face this issue\ni face this issue occasionally and i can't understand how to solve this issue\n\ninstead of shutdown if i could use a command/setting that will restart the pc instead of freezing/turning it off then it could be good for me but i don't know if it's possible to do so",
      "votes": null
    },
    {
      "id": "1504810",
      "postDate": "09/06/2021 17:03:10",
      "content": "<p>obviously you monitoring gpu temperatures? nvidia gpus have shutdown thresholds at ~90c.   it should not lead to pc shutdown thou</p>",
      "rawMarkdown": "obviously you monitoring gpu temperatures? nvidia gpus have shutdown thresholds at ~90c.   it should not lead to pc shutdown thou",
      "votes": null
    },
    {
      "id": "1504834",
      "postDate": "09/06/2021 17:24:27",
      "content": "<blockquote>\n  <p>try empty cache after every epoch</p>\n</blockquote>\n<p>I used to do this when I had dual 2080's and it would balloon train times like crazzzyyyyyyyyyyy. Now, I just let the thing take care of the thing by itself.</p>",
      "rawMarkdown": "> try empty cache after every epoch\n\nI used to do this when I had dual 2080's and it would balloon train times like crazzzyyyyyyyyyyy. Now, I just let the thing take care of the thing by itself.",
      "votes": null
    },
    {
      "id": "1504858",
      "postDate": "09/06/2021 17:48:07",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I see, thanks for the clarification. So it seems to be a specific issue to this specific model setup. Have you tried very small image size (let's say 64 x 64) and/or removing the EfficientNet and replacing it by the identity or another model? </p>\n<p>Not sure how to help but if I were in your situation, I would try various combinations until I find something. Best of luck!</p>",
      "rawMarkdown": "mobassir I see, thanks for the clarification. So it seems to be a specific issue to this specific model setup. Have you tried very small image size (let's say 64 x 64) and/or removing the EfficientNet and replacing it by the identity or another model? \n\nNot sure how to help but if I were in your situation, I would try various combinations until I find something. Best of luck!",
      "votes": null
    },
    {
      "id": "1504890",
      "postDate": "09/06/2021 18:39:57",
      "content": "<p>Here is my use (V100 in NVIDIA cluster, first generation at 163W):</p>\n<p><img src=\"https://i.imgur.com/eljWSbK.png\" alt=\"gpu use\"></p>\n<p>I actually use almost all memory, for some reason memory use is divided by 2 in the report.</p>",
      "rawMarkdown": "Here is my use (V100 in NVIDIA cluster, first generation at 163W):\n\n![gpu use](https://i.imgur.com/eljWSbK.png)\n\nI actually use almost all memory, for some reason memory use is divided by 2 in the report.",
      "votes": null
    },
    {
      "id": "1504928",
      "postDate": "09/06/2021 19:03:17",
      "content": "<p>Thanks for the plot. 👌</p>",
      "rawMarkdown": "Thanks for the plot. 👌",
      "votes": null
    },
    {
      "id": "1505145",
      "postDate": "09/07/2021 02:07:06",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> In case of 3090 it also might be caused by high memory junction temperatures. I think they are only accessible to monitoring software under Windows, for example in HWinfo64. GDDR6X runs very hot, and sometimes reaches thermal throttle at 110 degrees.</p>\n<p>If this is the case (you can quickly run any Eth mining software on Windows as a way to stress test VRAM), then changing thermal pads usually improves the situation, but that involves disassembly of the cooling system and might complicate warranty. Another way could be just propping up your 3090 in the case - <a href=\"https://www.reddit.com/r/nvidia/comments/pihwzo/3090_vram_temperature_fix_without_replacing_pads/\" target=\"_blank\">https://www.reddit.com/r/nvidia/comments/pihwzo/3090_vram_temperature_fix_without_replacing_pads/</a></p>\n<p>Another relatively common mistake is connecting GPU with one cable from PSU &gt; two 8-pin connectors on GPU. Always use separate cables for separate 8-pin connectors, i think it is even in installation manual.</p>\n<p>However, if VRAM junction temperature is not above 100C, undervolting does not help with your shutdowns, and card is connected properly, and you are still having shutdowns - then it might be a general hardware instability, which is a nightmare to debug. I would still check with a ridiculously overpowered PSU just in case, but at that point i would probably resort to changing components one by one to find the root cause.</p>",
      "rawMarkdown": "mobassir In case of 3090 it also might be caused by high memory junction temperatures. I think they are only accessible to monitoring software under Windows, for example in HWinfo64. GDDR6X runs very hot, and sometimes reaches thermal throttle at 110 degrees.\n\nIf this is the case (you can quickly run any Eth mining software on Windows as a way to stress test VRAM), then changing thermal pads usually improves the situation, but that involves disassembly of the cooling system and might complicate warranty. Another way could be just propping up your 3090 in the case - https://www.reddit.com/r/nvidia/comments/pihwzo/3090_vram_temperature_fix_without_replacing_pads/\n\nAnother relatively common mistake is connecting GPU with one cable from PSU > two 8-pin connectors on GPU. Always use separate cables for separate 8-pin connectors, i think it is even in installation manual.\n\nHowever, if VRAM junction temperature is not above 100C, undervolting does not help with your shutdowns, and card is connected properly, and you are still having shutdowns - then it might be a general hardware instability, which is a nightmare to debug. I would still check with a ridiculously overpowered PSU just in case, but at that point i would probably resort to changing components one by one to find the root cause.",
      "votes": null
    },
    {
      "id": "1505296",
      "postDate": "09/07/2021 06:10:34",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> <a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> </p>\n<p>apology for a little mistake from my end,</p>\n<p><strong>The pc is not getting turned off,the problem is it is getting frozen every time and forcing me to restart everytime</strong></p>\n<p>for reproducing the result i just took the public kernel of <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> san and modified a little bit to train and face the freeze issue so that i can share the code and log with you guys for better investigation..</p>\n<p>here is the code and log file : <a href=\"https://pastebin.com/XEusHppZ\" target=\"_blank\">https://pastebin.com/XEusHppZ</a><br>\nyou can see from my log that during 7th epoch  i faced 'freeze' issue and i had to restart the pc.<br>\ni was using around 19gb vram out of 24 iirc while training this model (full code is provided).</p>\n<p>could this be related to pytorch version or something else?</p>\n<p>it will be highly appreciated if any kaggler help me solve this issue (i've shared the full code),thank you a lot in advance!</p>",
      "rawMarkdown": "fffrrt @cpmpml @yassinealouini @authman @bakeryproducts @pheadrus \n\napology for a little mistake from my end,\n\n**The pc is not getting turned off,the problem is it is getting frozen every time and forcing me to restart everytime**\n\nfor reproducing the result i just took the public kernel of @yasufuminakama san and modified a little bit to train and face the freeze issue so that i can share the code and log with you guys for better investigation..\n\nhere is the code and log file : https://pastebin.com/XEusHppZ\nyou can see from my log that during 7th epoch  i faced 'freeze' issue and i had to restart the pc.\ni was using around 19gb vram out of 24 iirc while training this model (full code is provided).\n\ncould this be related to pytorch version or something else?\n\nit will be highly appreciated if any kaggler help me solve this issue (i've shared the full code),thank you a lot in advance!",
      "votes": null
    },
    {
      "id": "1505311",
      "postDate": "09/07/2021 06:33:51",
      "content": "<p>One \"cheap\" solution is to stop after 3 or 4 epochs maybe? </p>\n<p>Otherwise, maybe try to reinstall everything? I've taken few notes here: <a href=\"https://gist.github.com/yassineAlouini/e843531219e1b9d978aa4fe2a5e71940\" target=\"_blank\">https://gist.github.com/yassineAlouini/e843531219e1b9d978aa4fe2a5e71940</a>. There is also this good blog post: <a href=\"https://www.itzgeek.com/post/how-to-install-nvidia-drivers-on-ubuntu-20-04-ubuntu-18-04.html\" target=\"_blank\">https://www.itzgeek.com/post/how-to-install-nvidia-drivers-on-ubuntu-20-04-ubuntu-18-04.html</a></p>",
      "rawMarkdown": "One \"cheap\" solution is to stop after 3 or 4 epochs maybe? \n\nOtherwise, maybe try to reinstall everything? I've taken few notes here: https://gist.github.com/yassineAlouini/e843531219e1b9d978aa4fe2a5e71940. There is also this good blog post: https://www.itzgeek.com/post/how-to-install-nvidia-drivers-on-ubuntu-20-04-ubuntu-18-04.html",
      "votes": null
    },
    {
      "id": "1505379",
      "postDate": "09/07/2021 07:42:24",
      "content": "<p>I had same issues as yours. I changed the PSU but the problem was still persistent. Later, I changed the motherboard boom it started working again, No freeze or random shutdown. Might be problem due to current spike causing damage to weaker motherboard capacitors/components.   </p>",
      "rawMarkdown": "I had same issues as yours. I changed the PSU but the problem was still persistent. Later, I changed the motherboard boom it started working again, No freeze or random shutdown. Might be problem due to current spike causing damage to weaker motherboard capacitors/components.",
      "votes": null
    },
    {
      "id": "1505606",
      "postDate": "09/07/2021 12:03:05",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> what was the GPU temperature when it froze?  </p>\n<p>Also, as <a href=\"https://www.kaggle.com/karmaz\" target=\"_blank\">@karmaz</a> indicated, it maybe a bad motherboard.  However, before changing it, have you tried to remove the GPU and disconnect everything, then connect and plug it back?  Sometimes a defective connection is the cause of such issues.</p>",
      "rawMarkdown": "mobassir what was the GPU temperature when it froze?  \n\nAlso, as @karmaz indicated, it maybe a bad motherboard.  However, before changing it, have you tried to remove the GPU and disconnect everything, then connect and plug it back?  Sometimes a defective connection is the cause of such issues.",
      "votes": null
    },
    {
      "id": "1505714",
      "postDate": "09/07/2021 13:52:42",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> i am  not sure if it is really a hardware related issue,,i thought this issue is related to pytorch version maybe,i am using the following packages : <br>\nName Version Build Channel<br>\n_libgcc_mutex 0.1 main<br>\n_openmp_mutex 4.5 1_gnu<br>\naddict 2.4.0 pypi_0 pypi<br>\nalabaster 0.7.12 py_0 conda-forge<br>\nalbumentations 1.0.3 pypi_0 pypi<br>\nalsa-lib 1.2.3 h516909a_0 conda-forge<br>\nappdirs 1.4.4 pyh9f0ad1d_0 conda-forge<br>\nargh 0.26.2 pyh9f0ad1d_1002 conda-forge<br>\nastroid 2.6.0 py38h578d9bd_0 conda-forge<br>\nasync_generator 1.10 py_0 conda-forge<br>\nasynctest 0.13.0 pypi_0 pypi<br>\natomicwrites 1.4.0 pyh9f0ad1d_0 conda-forge<br>\nattrs 21.2.0 pyhd8ed1ab_0 conda-forge<br>\nautopep8 1.5.6 pyhd8ed1ab_0 conda-forge<br>\nbabel 2.9.1 pyh44b312d_0 conda-forge<br>\nbackcall 0.2.0 pyhd3eb1b0_0<br>\nblack 21.5b2 pyhd8ed1ab_0 conda-forge<br>\nbleach 3.3.0 pyh44b312d_0 conda-forge<br>\nbrotlipy 0.7.0 py38h497a2fe_1001 conda-forge<br>\nca-certificates 2021.5.30 ha878542_0 conda-forge<br>\ncertifi 2021.5.30 py38h578d9bd_0 conda-forge<br>\ncffi 1.14.5 py38ha65f79e_0 conda-forge<br>\ncfgv 3.3.0 pypi_0 pypi<br>\nchardet 4.0.0 py38h578d9bd_1 conda-forge<br>\nclick 8.0.1 py38h578d9bd_0 conda-forge<br>\ncloudpickle 1.6.0 py_0 conda-forge<br>\ncodecov 2.1.11 pypi_0 pypi<br>\ncolorama 0.4.4 pyh9f0ad1d_0 conda-forge<br>\nconfigparser 5.0.2 pypi_0 pypi<br>\ncoverage 5.5 pypi_0 pypi<br>\ncryptography 3.4.7 py38ha5dfef3_0 conda-forge<br>\ncudatoolkit 11.1.74 h6bb024c_0 nvidia<br>\ncycler 0.10.0 pypi_0 pypi<br>\ncython 0.29.23 pypi_0 pypi<br>\ndataclasses 0.8 pyhc8e2a94_1 conda-forge<br>\ndbus 1.13.6 h48d8840_2 conda-forge<br>\ndecorator 4.4.2 pypi_0 pypi<br>\ndefusedxml 0.7.1 pyhd8ed1ab_0 conda-forge<br>\ndiff-match-patch 20200713 pyh9f0ad1d_0 conda-forge<br>\ndill 0.3.4 pypi_0 pypi<br>\ndistlib 0.3.2 pypi_0 pypi<br>\ndocker-pycreds 0.4.0 pypi_0 pypi<br>\ndocutils 0.17.1 py38h578d9bd_0 conda-forge<br>\nentrypoints 0.3 pyhd8ed1ab_1003 conda-forge<br>\nexpat 2.4.1 h9c3ff4c_0 conda-forge<br>\nfilelock 3.0.12 pypi_0 pypi<br>\nflake8 3.9.2 pypi_0 pypi<br>\nfontconfig 2.13.1 hba837de_1005 conda-forge<br>\nfreetype 2.10.4 h0708190_1 conda-forge<br>\nfuture 0.18.2 py38h578d9bd_3 conda-forge<br>\ngast 0.3.3 pypi_0 pypi<br>\ngettext 0.19.8.1 h0b5b191_1005 conda-forge<br>\ngitdb 4.0.7 pypi_0 pypi<br>\ngitpython 3.1.18 pypi_0 pypi<br>\nglib 2.68.3 h9c3ff4c_0 conda-forge<br>\nglib-tools 2.68.3 h9c3ff4c_0 conda-forge<br>\ngoogleapis-common-protos 1.53.0 pypi_0 pypi<br>\ngrad-cam 1.3.1 pypi_0 pypi<br>\ngrpcio 1.32.0 pypi_0 pypi<br>\ngst-plugins-base 1.18.4 hf529b03_2 conda-forge<br>\ngstreamer 1.18.4 h76c114f_2 conda-forge<br>\nh5py 2.10.0 pypi_0 pypi<br>\nhelpdev 0.7.1 pyhd8ed1ab_0 conda-forge<br>\nicu 68.1 h58526e2_0 conda-forge<br>\nidentify 2.2.10 pypi_0 pypi<br>\nidna 2.10 pyh9f0ad1d_0 conda-forge<br>\nimageio 2.9.0 pypi_0 pypi<br>\nimagesize 1.2.0 py_0 conda-forge<br>\nimgaug 0.4.0 pypi_0 pypi<br>\nimportlib-metadata 4.5.0 py38h578d9bd_0 conda-forge<br>\nimportlib-resources 5.2.2 pypi_0 pypi<br>\nimportlib_metadata 4.5.0 hd8ed1ab_0 conda-forge<br>\niniconfig 1.1.1 pypi_0 pypi<br>\nintervaltree 3.0.2 py_0 conda-forge<br>\nipykernel 5.5.5 py38hd0cf306_0 conda-forge<br>\nipython 7.18.1 py38h5ca1d4c_0 anaconda<br>\nipython_genutils 0.2.0 pyhd3eb1b0_1<br>\nipywidgets 7.6.3 pypi_0 pypi<br>\nisort 5.8.0 pypi_0 pypi<br>\njedi 0.17.2 py38h578d9bd_1 conda-forge<br>\njeepney 0.6.0 pyhd8ed1ab_0 conda-forge<br>\njinja2 3.0.1 pyhd8ed1ab_0 conda-forge<br>\njpeg 9d h36c2ea0_0 conda-forge<br>\njsonschema 3.2.0 pyhd8ed1ab_3 conda-forge<br>\njupyter_client 6.1.12 pyhd8ed1ab_0 conda-forge<br>\njupyter_core 4.7.1 py38h578d9bd_0 conda-forge<br>\njupyterlab-widgets 1.0.0 pypi_0 pypi<br>\njupyterlab_pygments 0.1.2 pyh9f0ad1d_0 conda-forge<br>\nkaggle 1.5.12 pypi_0 pypi<br>\nkaggledatasets 0.0.1 pypi_0 pypi<br>\nkeyring 23.0.1 py38h578d9bd_0 conda-forge<br>\nkiwisolver 1.3.1 pypi_0 pypi<br>\nkrb5 1.19.1 hcc1bbae_0 conda-forge<br>\nkwarray 0.5.19 pypi_0 pypi<br>\nlanms-proper 1.0.1 pypi_0 pypi<br>\nlazy-object-proxy 1.6.0 py38h497a2fe_0 conda-forge<br>\nld_impl_linux-64 2.35.1 h7274673_9<br>\nlibclang 11.1.0 default_ha53f305_1 conda-forge<br>\nlibedit 3.1.20191231 he28a2e2_2 conda-forge<br>\nlibevent 2.1.10 hcdb4288_3 conda-forge<br>\nlibffi 3.3 he6710b0_2<br>\nlibgcc-ng 9.3.0 h5101ec6_17<br>\nlibglib 2.68.3 h3e27bee_0 conda-forge<br>\nlibgomp 9.3.0 h5101ec6_17<br>\nlibiconv 1.16 h516909a_0 conda-forge<br>\nlibllvm11 11.1.0 hf817b99_2 conda-forge<br>\nlibogg 1.3.4 h7f98852_1 conda-forge<br>\nlibopus 1.3.1 h7f98852_1 conda-forge<br>\nlibpng 1.6.37 h21135ba_2 conda-forge<br>\nlibpq 13.3 hd57d9b9_0 conda-forge<br>\nlibsodium 1.0.18 h36c2ea0_1 conda-forge<br>\nlibspatialindex 1.9.3 h9c3ff4c_3 conda-forge<br>\nlibstdcxx-ng 9.3.0 hd4cf53a_17<br>\nlibuuid 2.32.1 h7f98852_1000 conda-forge<br>\nlibvorbis 1.3.7 h9c3ff4c_0 conda-forge<br>\nlibxcb 1.13 h7f98852_1003 conda-forge<br>\nlibxkbcommon 1.0.3 he3ba5ed_0 conda-forge<br>\nlibxml2 2.9.12 h72842e0_0 conda-forge<br>\nllvmlite 0.36.0 pypi_0 pypi<br>\nlmdb 1.2.1 pypi_0 pypi<br>\nlz4-c 1.9.3 h9c3ff4c_0 conda-forge<br>\nmarkupsafe 2.0.1 py38h497a2fe_0 conda-forge<br>\nmatplotlib 3.4.2 pypi_0 pypi<br>\nmccabe 0.6.1 pypi_0 pypi<br>\nmistune 0.8.4 py38h497a2fe_1003 conda-forge<br>\nmmcv-full 1.3.7 dev_0<br>\nmmdet 2.11.0 pypi_0 pypi<br>\nmmocr 0.2.0 dev_0<br>\nmmpycocotools 12.0.3 pypi_0 pypi<br>\nmypy_extensions 0.4.3 py38h578d9bd_3 conda-forge<br>\nmysql-common 8.0.25 ha770c72_2 conda-forge<br>\nmysql-libs 8.0.25 hfa10184_2 conda-forge<br>\nnbclient 0.5.3 pyhd8ed1ab_0 conda-forge<br>\nnbconvert 6.0.7 pypi_0 pypi<br>\nnbformat 5.1.3 pyhd8ed1ab_0 conda-forge<br>\nncurses 6.2 he6710b0_1<br>\nnest-asyncio 1.5.1 pyhd8ed1ab_0 conda-forge<br>\nnetworkx 2.5.1 pypi_0 pypi<br>\nnnaudio 0.2.5 pypi_0 pypi<br>\nnodeenv 1.6.0 pypi_0 pypi<br>\nnspr 4.30 h9c3ff4c_0 conda-forge<br>\nnss 3.64 hb5efdd6_0 conda-forge<br>\nnumba 0.53.1 pypi_0 pypi<br>\nnumpy 1.19.5 pypi_0 pypi<br>\nnumpydoc 1.1.0 py_1 conda-forge<br>\noauthlib 3.1.1 pypi_0 pypi<br>\nopencv-python 4.5.2.54 pypi_0 pypi<br>\nopencv-python-headless 4.5.3.56 pypi_0 pypi<br>\nopenssl 1.1.1k h7f98852_0 conda-forge<br>\nordered-set 4.0.2 pypi_0 pypi<br>\npackaging 20.9 pyh44b312d_0 conda-forge<br>\npandas 1.3.0 pypi_0 pypi<br>\npandoc 2.14.0.3 h7f98852_0 conda-forge<br>\npandocfilters 1.4.3 pypi_0 pypi<br>\nparso 0.7.0 pyh9f0ad1d_0 conda-forge<br>\npathspec 0.8.1 pyhd3deb0d_0 conda-forge<br>\npathtools 0.1.2 pypi_0 pypi<br>\npcre 8.45 h9c3ff4c_0 conda-forge<br>\npexpect 4.8.0 pyhd3eb1b0_3<br>\npickleshare 0.7.5 pyhd3eb1b0_1003<br>\npillow 8.2.0 pypi_0 pypi<br>\npip 21.1.2 py38h06a4308_0<br>\npluggy 0.13.1 py38h578d9bd_4 conda-forge<br>\npolygon3 3.0.9.1 pypi_0 pypi<br>\npre-commit 2.13.0 pypi_0 pypi<br>\npromise 2.3 pypi_0 pypi<br>\nprompt-toolkit 3.0.18 pypi_0 pypi<br>\nprotobuf 3.17.3 pypi_0 pypi<br>\npsutil 5.8.0 py38h497a2fe_1 conda-forge<br>\npthread-stubs 0.4 h36c2ea0_1001 conda-forge<br>\nptyprocess 0.7.0 pyhd3eb1b0_2<br>\npy 1.10.0 pypi_0 pypi<br>\npyclipper 1.2.1 pypi_0 pypi<br>\npycocotools 2.0.2 pypi_0 pypi<br>\npycodestyle 2.7.0 pypi_0 pypi<br>\npycparser 2.20 pyh9f0ad1d_2 conda-forge<br>\npydocstyle 6.1.1 pyhd8ed1ab_0 conda-forge<br>\npyflakes 2.3.1 pypi_0 pypi<br>\npygments 2.9.0 pyhd3eb1b0_0<br>\npylint 2.8.2 pyhd8ed1ab_0 conda-forge<br>\npyls-black 0.4.6 pyh9f0ad1d_0 conda-forge<br>\npyls-spyder 0.3.2 pyhd8ed1ab_0 conda-forge<br>\npyopenssl 20.0.1 pyhd8ed1ab_0 conda-forge<br>\npyparsing 2.4.7 pyh9f0ad1d_0 conda-forge<br>\npyqt 5.12.3 py38h578d9bd_7 conda-forge<br>\npyqt-impl 5.12.3 py38h7400c14_7 conda-forge<br>\npyqt5-sip 4.19.18 py38h709712a_7 conda-forge<br>\npyqtchart 5.12 py38h7400c14_7 conda-forge<br>\npyqtwebengine 5.12.1 py38h7400c14_7 conda-forge<br>\npyrsistent 0.17.3 py38h497a2fe_2 conda-forge<br>\npysocks 1.7.1 py38h578d9bd_3 conda-forge<br>\npytest 6.2.4 pypi_0 pypi<br>\npytest-cov 2.12.1 pypi_0 pypi<br>\npytest-runner 5.3.1 pypi_0 pypi<br>\npython 3.8.5 h7579374_1<br>\npython-dateutil 2.8.1 py_0 conda-forge<br>\npython-jsonrpc-server 0.4.0 pyh9f0ad1d_0 conda-forge<br>\npython-language-server 0.36.2 pyhd8ed1ab_0 conda-forge<br>\npython-slugify 5.0.2 pypi_0 pypi<br>\npython_abi 3.8 2_cp38 conda-forge<br>\npytz 2021.1 pyhd8ed1ab_0 conda-forge<br>\npywavelets 1.1.1 pypi_0 pypi<br>\npyxdg 0.27 pyhd8ed1ab_0 conda-forge<br>\npyyaml 5.4.1 pypi_0 pypi<br>\npyzmq 22.1.0 py38h2035c66_0 conda-forge<br>\nqdarkstyle 2.8.1 pyhd8ed1ab_2 conda-forge<br>\nqt 5.12.9 hda022c4_4 conda-forge<br>\nqtawesome 1.0.3 pyhd8ed1ab_0 conda-forge<br>\nqtconsole 5.1.0 pyhd8ed1ab_0 conda-forge<br>\nqtpy 1.9.0 py_0 conda-forge<br>\nqudida 0.0.4 pypi_0 pypi<br>\nrapidfuzz 1.4.1 pypi_0 pypi<br>\nreadline 8.1 h27cfd23_0<br>\nregex 2021.4.4 py38h497a2fe_0 conda-forge<br>\nrequests 2.25.1 pyhd3deb0d_0 conda-forge<br>\nrope 0.19.0 pyhd8ed1ab_0 conda-forge<br>\nrtree 0.9.7 py38h02d302b_1 conda-forge<br>\nscikit-image 0.18.1 pypi_0 pypi<br>\nscipy 1.6.3 pypi_0 pypi<br>\nseaborn 0.11.2 pypi_0 pypi<br>\nsecretstorage 3.3.1 py38h578d9bd_0 conda-forge<br>\nsend2trash 1.5.0 pypi_0 pypi<br>\nsentry-sdk 1.3.1 pypi_0 pypi<br>\nsetuptools 52.0.0 py38h06a4308_0<br>\nshapely 1.7.1 pypi_0 pypi<br>\nshortuuid 1.0.1 pypi_0 pypi<br>\nsix 1.16.0 pyh6c4a22f_0 conda-forge<br>\nsmmap 4.0.0 pypi_0 pypi<br>\nsnowballstemmer 2.1.0 pyhd8ed1ab_0 conda-forge<br>\nsortedcontainers 2.4.0 pyhd8ed1ab_0 conda-forge<br>\nsphinx 4.0.2 pyh6c4a22f_1 conda-forge<br>\nsphinxcontrib-applehelp 1.0.2 py_0 conda-forge<br>\nsphinxcontrib-devhelp 1.0.2 py_0 conda-forge<br>\nsphinxcontrib-htmlhelp 2.0.0 pyhd8ed1ab_0 conda-forge<br>\nsphinxcontrib-jsmath 1.0.1 py_0 conda-forge<br>\nsphinxcontrib-qthelp 1.0.3 py_0 conda-forge<br>\nsphinxcontrib-serializinghtml 1.1.5 pyhd8ed1ab_0 conda-forge<br>\nspyder 4.2.5 py38h578d9bd_0 conda-forge<br>\nspyder-kernels 1.10.2 py38h578d9bd_0 conda-forge<br>\nsqlite 3.35.4 hdfb4753_0<br>\nsubprocess32 3.5.4 pypi_0 pypi<br>\ntensorboard 2.6.0 pypi_0 pypi<br>\ntensorflow-datasets 4.4.0 pypi_0 pypi<br>\ntensorflow-gpu 2.4.0 pypi_0 pypi<br>\ntensorflow-metadata 1.2.0 pypi_0 pypi<br>\nterminado 0.10.0 pypi_0 pypi<br>\nterminaltables 3.1.0 pypi_0 pypi<br>\ntestpath 0.5.0 pyhd8ed1ab_0 conda-forge<br>\ntext-unidecode 1.3 pypi_0 pypi<br>\ntextdistance 4.2.1 pyhd8ed1ab_0 conda-forge<br>\nthree-merge 0.1.1 pyh9f0ad1d_0 conda-forge<br>\ntifffile 2021.6.6 pypi_0 pypi<br>\ntimm 0.4.13 pypi_0 pypi<br>\ntk 8.6.10 hbc83047_0<br>\ntoml 0.10.2 pyhd8ed1ab_0 conda-forge<br>\ntorch 1.10.0.dev20210623+cu111 pypi_0 pypi<br>\ntorchaudio 0.8.1 pypi_0 pypi<br>\ntorchvision 0.11.0.dev20210623+cu111 pypi_0 pypi<br>\ntornado 6.1 py38h497a2fe_1 conda-forge<br>\ntraitlets 5.0.5 pyhd3eb1b0_0<br>\nttach 0.0.3 pypi_0 pypi<br>\ntyped-ast 1.4.3 py38h497a2fe_0 conda-forge<br>\ntyping_extensions 3.10.0.0 pyha770c72_0 conda-forge<br>\nubelt 0.9.5 pypi_0 pypi<br>\nujson 4.0.2 py38h709712a_0 conda-forge<br>\nurllib3 1.26.5 pyhd8ed1ab_0 conda-forge<br>\nvirtualenv 20.4.7 pypi_0 pypi<br>\nwandb 0.12.1 pypi_0 pypi<br>\nwatchdog 1.0.2 py38h578d9bd_1 conda-forge<br>\nwcwidth 0.2.5 py_0<br>\nwebencodings 0.5.1 pypi_0 pypi<br>\nwheel 0.36.2 pyhd3eb1b0_0<br>\nwidgetsnbextension 3.5.1 pypi_0 pypi<br>\nwrapt 1.12.1 py38h497a2fe_3 conda-forge<br>\nwurlitzer 2.1.0 py38h578d9bd_0 conda-forge<br>\nxdoctest 0.15.4 pypi_0 pypi<br>\nxorg-libxau 1.0.9 h7f98852_0 conda-forge<br>\nxorg-libxdmcp 1.1.3 h7f98852_0 conda-forge<br>\nxz 5.2.5 h7b6447c_0<br>\nyaml 0.2.5 h516909a_0 conda-forge<br>\nyapf 0.31.0 pyhd8ed1ab_0 conda-forge<br>\nzeromq 4.3.4 h9c3ff4c_0 conda-forge<br>\nzipp 3.4.1 pyhd8ed1ab_0 conda-forge<br>\nzlib 1.2.11 h7b6447c_3<br>\nzstd 1.5.0 ha95c52a_0 conda-forge</p>\n<p>and using this i trained deeper models like robustscanner,nrtr,fcenet etc<br>\nin another pytorch environment where i have pytorch 1.8 installed i trained many powerful models like xlm roberta,deberta large etc by utilizing 24gb vram and i didn't face this freezing issue<br>\nso if it is hardware problem then i should experience this freezing issue in every model training(at least in most model training),no?<br>\nfor example,,if i do not convert the wave signal into 512x512 image size for training,,then i do not face this freezing issue!</p>\n<p>most of the times i face this freezing issue while trying nfnets,i think i am probably missing a line of code which could fix this freezing issue,i wonder if it is necessary to use torch.cuda.empty_cache() and gc.collect() always for getting rid of this freezing issue or something else that is causing the problem! 😪</p>",
      "rawMarkdown": "cpmpml i am  not sure if it is really a hardware related issue,,i thought this issue is related to pytorch version maybe,i am using the following packages : \nName Version Build Channel\n_libgcc_mutex 0.1 main\n_openmp_mutex 4.5 1_gnu\naddict 2.4.0 pypi_0 pypi\nalabaster 0.7.12 py_0 conda-forge\nalbumentations 1.0.3 pypi_0 pypi\nalsa-lib 1.2.3 h516909a_0 conda-forge\nappdirs 1.4.4 pyh9f0ad1d_0 conda-forge\nargh 0.26.2 pyh9f0ad1d_1002 conda-forge\nastroid 2.6.0 py38h578d9bd_0 conda-forge\nasync_generator 1.10 py_0 conda-forge\nasynctest 0.13.0 pypi_0 pypi\natomicwrites 1.4.0 pyh9f0ad1d_0 conda-forge\nattrs 21.2.0 pyhd8ed1ab_0 conda-forge\nautopep8 1.5.6 pyhd8ed1ab_0 conda-forge\nbabel 2.9.1 pyh44b312d_0 conda-forge\nbackcall 0.2.0 pyhd3eb1b0_0\nblack 21.5b2 pyhd8ed1ab_0 conda-forge\nbleach 3.3.0 pyh44b312d_0 conda-forge\nbrotlipy 0.7.0 py38h497a2fe_1001 conda-forge\nca-certificates 2021.5.30 ha878542_0 conda-forge\ncertifi 2021.5.30 py38h578d9bd_0 conda-forge\ncffi 1.14.5 py38ha65f79e_0 conda-forge\ncfgv 3.3.0 pypi_0 pypi\nchardet 4.0.0 py38h578d9bd_1 conda-forge\nclick 8.0.1 py38h578d9bd_0 conda-forge\ncloudpickle 1.6.0 py_0 conda-forge\ncodecov 2.1.11 pypi_0 pypi\ncolorama 0.4.4 pyh9f0ad1d_0 conda-forge\nconfigparser 5.0.2 pypi_0 pypi\ncoverage 5.5 pypi_0 pypi\ncryptography 3.4.7 py38ha5dfef3_0 conda-forge\ncudatoolkit 11.1.74 h6bb024c_0 nvidia\ncycler 0.10.0 pypi_0 pypi\ncython 0.29.23 pypi_0 pypi\ndataclasses 0.8 pyhc8e2a94_1 conda-forge\ndbus 1.13.6 h48d8840_2 conda-forge\ndecorator 4.4.2 pypi_0 pypi\ndefusedxml 0.7.1 pyhd8ed1ab_0 conda-forge\ndiff-match-patch 20200713 pyh9f0ad1d_0 conda-forge\ndill 0.3.4 pypi_0 pypi\ndistlib 0.3.2 pypi_0 pypi\ndocker-pycreds 0.4.0 pypi_0 pypi\ndocutils 0.17.1 py38h578d9bd_0 conda-forge\nentrypoints 0.3 pyhd8ed1ab_1003 conda-forge\nexpat 2.4.1 h9c3ff4c_0 conda-forge\nfilelock 3.0.12 pypi_0 pypi\nflake8 3.9.2 pypi_0 pypi\nfontconfig 2.13.1 hba837de_1005 conda-forge\nfreetype 2.10.4 h0708190_1 conda-forge\nfuture 0.18.2 py38h578d9bd_3 conda-forge\ngast 0.3.3 pypi_0 pypi\ngettext 0.19.8.1 h0b5b191_1005 conda-forge\ngitdb 4.0.7 pypi_0 pypi\ngitpython 3.1.18 pypi_0 pypi\nglib 2.68.3 h9c3ff4c_0 conda-forge\nglib-tools 2.68.3 h9c3ff4c_0 conda-forge\ngoogleapis-common-protos 1.53.0 pypi_0 pypi\ngrad-cam 1.3.1 pypi_0 pypi\ngrpcio 1.32.0 pypi_0 pypi\ngst-plugins-base 1.18.4 hf529b03_2 conda-forge\ngstreamer 1.18.4 h76c114f_2 conda-forge\nh5py 2.10.0 pypi_0 pypi\nhelpdev 0.7.1 pyhd8ed1ab_0 conda-forge\nicu 68.1 h58526e2_0 conda-forge\nidentify 2.2.10 pypi_0 pypi\nidna 2.10 pyh9f0ad1d_0 conda-forge\nimageio 2.9.0 pypi_0 pypi\nimagesize 1.2.0 py_0 conda-forge\nimgaug 0.4.0 pypi_0 pypi\nimportlib-metadata 4.5.0 py38h578d9bd_0 conda-forge\nimportlib-resources 5.2.2 pypi_0 pypi\nimportlib_metadata 4.5.0 hd8ed1ab_0 conda-forge\niniconfig 1.1.1 pypi_0 pypi\nintervaltree 3.0.2 py_0 conda-forge\nipykernel 5.5.5 py38hd0cf306_0 conda-forge\nipython 7.18.1 py38h5ca1d4c_0 anaconda\nipython_genutils 0.2.0 pyhd3eb1b0_1\nipywidgets 7.6.3 pypi_0 pypi\nisort 5.8.0 pypi_0 pypi\njedi 0.17.2 py38h578d9bd_1 conda-forge\njeepney 0.6.0 pyhd8ed1ab_0 conda-forge\njinja2 3.0.1 pyhd8ed1ab_0 conda-forge\njpeg 9d h36c2ea0_0 conda-forge\njsonschema 3.2.0 pyhd8ed1ab_3 conda-forge\njupyter_client 6.1.12 pyhd8ed1ab_0 conda-forge\njupyter_core 4.7.1 py38h578d9bd_0 conda-forge\njupyterlab-widgets 1.0.0 pypi_0 pypi\njupyterlab_pygments 0.1.2 pyh9f0ad1d_0 conda-forge\nkaggle 1.5.12 pypi_0 pypi\nkaggledatasets 0.0.1 pypi_0 pypi\nkeyring 23.0.1 py38h578d9bd_0 conda-forge\nkiwisolver 1.3.1 pypi_0 pypi\nkrb5 1.19.1 hcc1bbae_0 conda-forge\nkwarray 0.5.19 pypi_0 pypi\nlanms-proper 1.0.1 pypi_0 pypi\nlazy-object-proxy 1.6.0 py38h497a2fe_0 conda-forge\nld_impl_linux-64 2.35.1 h7274673_9\nlibclang 11.1.0 default_ha53f305_1 conda-forge\nlibedit 3.1.20191231 he28a2e2_2 conda-forge\nlibevent 2.1.10 hcdb4288_3 conda-forge\nlibffi 3.3 he6710b0_2\nlibgcc-ng 9.3.0 h5101ec6_17\nlibglib 2.68.3 h3e27bee_0 conda-forge\nlibgomp 9.3.0 h5101ec6_17\nlibiconv 1.16 h516909a_0 conda-forge\nlibllvm11 11.1.0 hf817b99_2 conda-forge\nlibogg 1.3.4 h7f98852_1 conda-forge\nlibopus 1.3.1 h7f98852_1 conda-forge\nlibpng 1.6.37 h21135ba_2 conda-forge\nlibpq 13.3 hd57d9b9_0 conda-forge\nlibsodium 1.0.18 h36c2ea0_1 conda-forge\nlibspatialindex 1.9.3 h9c3ff4c_3 conda-forge\nlibstdcxx-ng 9.3.0 hd4cf53a_17\nlibuuid 2.32.1 h7f98852_1000 conda-forge\nlibvorbis 1.3.7 h9c3ff4c_0 conda-forge\nlibxcb 1.13 h7f98852_1003 conda-forge\nlibxkbcommon 1.0.3 he3ba5ed_0 conda-forge\nlibxml2 2.9.12 h72842e0_0 conda-forge\nllvmlite 0.36.0 pypi_0 pypi\nlmdb 1.2.1 pypi_0 pypi\nlz4-c 1.9.3 h9c3ff4c_0 conda-forge\nmarkupsafe 2.0.1 py38h497a2fe_0 conda-forge\nmatplotlib 3.4.2 pypi_0 pypi\nmccabe 0.6.1 pypi_0 pypi\nmistune 0.8.4 py38h497a2fe_1003 conda-forge\nmmcv-full 1.3.7 dev_0\nmmdet 2.11.0 pypi_0 pypi\nmmocr 0.2.0 dev_0\nmmpycocotools 12.0.3 pypi_0 pypi\nmypy_extensions 0.4.3 py38h578d9bd_3 conda-forge\nmysql-common 8.0.25 ha770c72_2 conda-forge\nmysql-libs 8.0.25 hfa10184_2 conda-forge\nnbclient 0.5.3 pyhd8ed1ab_0 conda-forge\nnbconvert 6.0.7 pypi_0 pypi\nnbformat 5.1.3 pyhd8ed1ab_0 conda-forge\nncurses 6.2 he6710b0_1\nnest-asyncio 1.5.1 pyhd8ed1ab_0 conda-forge\nnetworkx 2.5.1 pypi_0 pypi\nnnaudio 0.2.5 pypi_0 pypi\nnodeenv 1.6.0 pypi_0 pypi\nnspr 4.30 h9c3ff4c_0 conda-forge\nnss 3.64 hb5efdd6_0 conda-forge\nnumba 0.53.1 pypi_0 pypi\nnumpy 1.19.5 pypi_0 pypi\nnumpydoc 1.1.0 py_1 conda-forge\noauthlib 3.1.1 pypi_0 pypi\nopencv-python 4.5.2.54 pypi_0 pypi\nopencv-python-headless 4.5.3.56 pypi_0 pypi\nopenssl 1.1.1k h7f98852_0 conda-forge\nordered-set 4.0.2 pypi_0 pypi\npackaging 20.9 pyh44b312d_0 conda-forge\npandas 1.3.0 pypi_0 pypi\npandoc 2.14.0.3 h7f98852_0 conda-forge\npandocfilters 1.4.3 pypi_0 pypi\nparso 0.7.0 pyh9f0ad1d_0 conda-forge\npathspec 0.8.1 pyhd3deb0d_0 conda-forge\npathtools 0.1.2 pypi_0 pypi\npcre 8.45 h9c3ff4c_0 conda-forge\npexpect 4.8.0 pyhd3eb1b0_3\npickleshare 0.7.5 pyhd3eb1b0_1003\npillow 8.2.0 pypi_0 pypi\npip 21.1.2 py38h06a4308_0\npluggy 0.13.1 py38h578d9bd_4 conda-forge\npolygon3 3.0.9.1 pypi_0 pypi\npre-commit 2.13.0 pypi_0 pypi\npromise 2.3 pypi_0 pypi\nprompt-toolkit 3.0.18 pypi_0 pypi\nprotobuf 3.17.3 pypi_0 pypi\npsutil 5.8.0 py38h497a2fe_1 conda-forge\npthread-stubs 0.4 h36c2ea0_1001 conda-forge\nptyprocess 0.7.0 pyhd3eb1b0_2\npy 1.10.0 pypi_0 pypi\npyclipper 1.2.1 pypi_0 pypi\npycocotools 2.0.2 pypi_0 pypi\npycodestyle 2.7.0 pypi_0 pypi\npycparser 2.20 pyh9f0ad1d_2 conda-forge\npydocstyle 6.1.1 pyhd8ed1ab_0 conda-forge\npyflakes 2.3.1 pypi_0 pypi\npygments 2.9.0 pyhd3eb1b0_0\npylint 2.8.2 pyhd8ed1ab_0 conda-forge\npyls-black 0.4.6 pyh9f0ad1d_0 conda-forge\npyls-spyder 0.3.2 pyhd8ed1ab_0 conda-forge\npyopenssl 20.0.1 pyhd8ed1ab_0 conda-forge\npyparsing 2.4.7 pyh9f0ad1d_0 conda-forge\npyqt 5.12.3 py38h578d9bd_7 conda-forge\npyqt-impl 5.12.3 py38h7400c14_7 conda-forge\npyqt5-sip 4.19.18 py38h709712a_7 conda-forge\npyqtchart 5.12 py38h7400c14_7 conda-forge\npyqtwebengine 5.12.1 py38h7400c14_7 conda-forge\npyrsistent 0.17.3 py38h497a2fe_2 conda-forge\npysocks 1.7.1 py38h578d9bd_3 conda-forge\npytest 6.2.4 pypi_0 pypi\npytest-cov 2.12.1 pypi_0 pypi\npytest-runner 5.3.1 pypi_0 pypi\npython 3.8.5 h7579374_1\npython-dateutil 2.8.1 py_0 conda-forge\npython-jsonrpc-server 0.4.0 pyh9f0ad1d_0 conda-forge\npython-language-server 0.36.2 pyhd8ed1ab_0 conda-forge\npython-slugify 5.0.2 pypi_0 pypi\npython_abi 3.8 2_cp38 conda-forge\npytz 2021.1 pyhd8ed1ab_0 conda-forge\npywavelets 1.1.1 pypi_0 pypi\npyxdg 0.27 pyhd8ed1ab_0 conda-forge\npyyaml 5.4.1 pypi_0 pypi\npyzmq 22.1.0 py38h2035c66_0 conda-forge\nqdarkstyle 2.8.1 pyhd8ed1ab_2 conda-forge\nqt 5.12.9 hda022c4_4 conda-forge\nqtawesome 1.0.3 pyhd8ed1ab_0 conda-forge\nqtconsole 5.1.0 pyhd8ed1ab_0 conda-forge\nqtpy 1.9.0 py_0 conda-forge\nqudida 0.0.4 pypi_0 pypi\nrapidfuzz 1.4.1 pypi_0 pypi\nreadline 8.1 h27cfd23_0\nregex 2021.4.4 py38h497a2fe_0 conda-forge\nrequests 2.25.1 pyhd3deb0d_0 conda-forge\nrope 0.19.0 pyhd8ed1ab_0 conda-forge\nrtree 0.9.7 py38h02d302b_1 conda-forge\nscikit-image 0.18.1 pypi_0 pypi\nscipy 1.6.3 pypi_0 pypi\nseaborn 0.11.2 pypi_0 pypi\nsecretstorage 3.3.1 py38h578d9bd_0 conda-forge\nsend2trash 1.5.0 pypi_0 pypi\nsentry-sdk 1.3.1 pypi_0 pypi\nsetuptools 52.0.0 py38h06a4308_0\nshapely 1.7.1 pypi_0 pypi\nshortuuid 1.0.1 pypi_0 pypi\nsix 1.16.0 pyh6c4a22f_0 conda-forge\nsmmap 4.0.0 pypi_0 pypi\nsnowballstemmer 2.1.0 pyhd8ed1ab_0 conda-forge\nsortedcontainers 2.4.0 pyhd8ed1ab_0 conda-forge\nsphinx 4.0.2 pyh6c4a22f_1 conda-forge\nsphinxcontrib-applehelp 1.0.2 py_0 conda-forge\nsphinxcontrib-devhelp 1.0.2 py_0 conda-forge\nsphinxcontrib-htmlhelp 2.0.0 pyhd8ed1ab_0 conda-forge\nsphinxcontrib-jsmath 1.0.1 py_0 conda-forge\nsphinxcontrib-qthelp 1.0.3 py_0 conda-forge\nsphinxcontrib-serializinghtml 1.1.5 pyhd8ed1ab_0 conda-forge\nspyder 4.2.5 py38h578d9bd_0 conda-forge\nspyder-kernels 1.10.2 py38h578d9bd_0 conda-forge\nsqlite 3.35.4 hdfb4753_0\nsubprocess32 3.5.4 pypi_0 pypi\ntensorboard 2.6.0 pypi_0 pypi\ntensorflow-datasets 4.4.0 pypi_0 pypi\ntensorflow-gpu 2.4.0 pypi_0 pypi\ntensorflow-metadata 1.2.0 pypi_0 pypi\nterminado 0.10.0 pypi_0 pypi\nterminaltables 3.1.0 pypi_0 pypi\ntestpath 0.5.0 pyhd8ed1ab_0 conda-forge\ntext-unidecode 1.3 pypi_0 pypi\ntextdistance 4.2.1 pyhd8ed1ab_0 conda-forge\nthree-merge 0.1.1 pyh9f0ad1d_0 conda-forge\ntifffile 2021.6.6 pypi_0 pypi\ntimm 0.4.13 pypi_0 pypi\ntk 8.6.10 hbc83047_0\ntoml 0.10.2 pyhd8ed1ab_0 conda-forge\ntorch 1.10.0.dev20210623+cu111 pypi_0 pypi\ntorchaudio 0.8.1 pypi_0 pypi\ntorchvision 0.11.0.dev20210623+cu111 pypi_0 pypi\ntornado 6.1 py38h497a2fe_1 conda-forge\ntraitlets 5.0.5 pyhd3eb1b0_0\nttach 0.0.3 pypi_0 pypi\ntyped-ast 1.4.3 py38h497a2fe_0 conda-forge\ntyping_extensions 3.10.0.0 pyha770c72_0 conda-forge\nubelt 0.9.5 pypi_0 pypi\nujson 4.0.2 py38h709712a_0 conda-forge\nurllib3 1.26.5 pyhd8ed1ab_0 conda-forge\nvirtualenv 20.4.7 pypi_0 pypi\nwandb 0.12.1 pypi_0 pypi\nwatchdog 1.0.2 py38h578d9bd_1 conda-forge\nwcwidth 0.2.5 py_0\nwebencodings 0.5.1 pypi_0 pypi\nwheel 0.36.2 pyhd3eb1b0_0\nwidgetsnbextension 3.5.1 pypi_0 pypi\nwrapt 1.12.1 py38h497a2fe_3 conda-forge\nwurlitzer 2.1.0 py38h578d9bd_0 conda-forge\nxdoctest 0.15.4 pypi_0 pypi\nxorg-libxau 1.0.9 h7f98852_0 conda-forge\nxorg-libxdmcp 1.1.3 h7f98852_0 conda-forge\nxz 5.2.5 h7b6447c_0\nyaml 0.2.5 h516909a_0 conda-forge\nyapf 0.31.0 pyhd8ed1ab_0 conda-forge\nzeromq 4.3.4 h9c3ff4c_0 conda-forge\nzipp 3.4.1 pyhd8ed1ab_0 conda-forge\nzlib 1.2.11 h7b6447c_3\nzstd 1.5.0 ha95c52a_0 conda-forge\n\n\nand using this i trained deeper models like robustscanner,nrtr,fcenet etc\nin another pytorch environment where i have pytorch 1.8 installed i trained many powerful models like xlm roberta,deberta large etc by utilizing 24gb vram and i didn't face this freezing issue\nso if it is hardware problem then i should experience this freezing issue in every model training(at least in most model training),no?\nfor example,,if i do not convert the wave signal into 512x512 image size for training,,then i do not face this freezing issue!\n\nmost of the times i face this freezing issue while trying nfnets,i think i am probably missing a line of code which could fix this freezing issue,i wonder if it is necessary to use torch.cuda.empty_cache() and gc.collect() always for getting rid of this freezing issue or something else that is causing the problem! 😪",
      "votes": null
    },
    {
      "id": "1505790",
      "postDate": "09/07/2021 14:50:41",
      "content": "<p>Are you using AMD cpus and ubuntu?<br>\nIs the freeze happening during an epoch or at start/end?<br>\nDid you track the temperature?</p>",
      "rawMarkdown": "Are you using AMD cpus and ubuntu?\nIs the freeze happening during an epoch or at start/end?\nDid you track the temperature?",
      "votes": null
    },
    {
      "id": "1505903",
      "postDate": "09/07/2021 15:55:11",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> </p>\n<ol>\n<li>intel cpu core i9 with ubuntu</li>\n<li>during an epoch as you can see from the log at the bottom here : <a href=\"https://pastebin.com/XEusHppZ\" target=\"_blank\">https://pastebin.com/XEusHppZ</a></li>\n<li>when the model was training i was not in front of the computer and i don't know if i can monitor and save temperate overtime when i am not in-front of the computer,sorry for this</li>\n</ol>",
      "rawMarkdown": "philippsinger \n1. intel cpu core i9 with ubuntu\n2. during an epoch as you can see from the log at the bottom here : https://pastebin.com/XEusHppZ\n3. when the model was training i was not in front of the computer and i don't know if i can monitor and save temperate overtime when i am not in-front of the computer,sorry for this",
      "votes": null
    },
    {
      "id": "1505939",
      "postDate": "09/07/2021 16:41:46",
      "content": "<p>If you use wandb, you can track some of your hardware's indicators. Not sure how reliable this is but it is better than nothing I guess?</p>",
      "rawMarkdown": "If you use wandb, you can track some of your hardware's indicators. Not sure how reliable this is but it is better than nothing I guess?",
      "votes": null
    },
    {
      "id": "1506099",
      "postDate": "09/07/2021 20:55:00",
      "content": "<p>and How Much it takes  on non RTX ones like  GTX 1660ti </p>",
      "rawMarkdown": "and How Much it takes  on non RTX ones like  GTX 1660ti",
      "votes": null
    },
    {
      "id": "1506106",
      "postDate": "09/07/2021 21:18:16",
      "content": "<p>I work on waveform only, with tfrecords and 8 TPU, I do one epoch into 27 seconds.</p>",
      "rawMarkdown": "I work on waveform only, with tfrecords and 8 TPU, I do one epoch into 27 seconds.",
      "votes": null
    },
    {
      "id": "1506250",
      "postDate": "09/08/2021 04:20:04",
      "content": "<p>I also have a rtx  3090, I average 20 min per epoch with preprocessed data and ~120-180 min when processing data on the fly. Pytorch</p>",
      "rawMarkdown": "I also have a rtx  3090, I average 20 min per epoch with preprocessed data and ~120-180 min when processing data on the fly. Pytorch",
      "votes": null
    },
    {
      "id": "1506318",
      "postDate": "09/08/2021 05:58:54",
      "content": "<p>Interesting, so you get a 6 or 8 factor from saving the images. I am planning to do it as well, will let you know how it goes.</p>",
      "rawMarkdown": "Interesting, so you get a 6 or 8 factor from saving the images. I am planning to do it as well, will let you know how it goes.",
      "votes": null
    },
    {
      "id": "1506321",
      "postDate": "09/08/2021 06:02:23",
      "content": "<p>Probably forever. 😄</p>\n<p>Joke aside, try to run on small images (128 x 128) and a very small model (EfficientNet B0 maybe or even something else). Also try to add gradient accumulation to avoid memory issues.</p>",
      "rawMarkdown": "Probably forever. 😄\n\nJoke aside, try to run on small images (128 x 128) and a very small model (EfficientNet B0 maybe or even something else). Also try to add gradient accumulation to avoid memory issues.",
      "votes": null
    },
    {
      "id": "1506333",
      "postDate": "09/08/2021 06:19:03",
      "content": "<p>You can also do it on the fly on the GPU with nnAudio</p>",
      "rawMarkdown": "You can also do it on the fly on the GPU with nnAudio",
      "votes": null
    },
    {
      "id": "1506400",
      "postDate": "09/08/2021 07:20:20",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> how many Watts is your PSU ?</p>",
      "rawMarkdown": "mobassir how many Watts is your PSU ?",
      "votes": null
    },
    {
      "id": "1506470",
      "postDate": "09/08/2021 08:55:21",
      "content": "<p>between this two it is not much about RT part, but tensor cores. 1660 has none. But again 1080ti for example is still a solid performer even without them.</p>",
      "rawMarkdown": "between this two it is not much about RT part, but tensor cores. 1660 has none. But again 1080ti for example is still a solid performer even without them.",
      "votes": null
    },
    {
      "id": "1506495",
      "postDate": "09/08/2021 09:25:43",
      "content": "<p>Is wandb not tracking the temperature? Whats your temperature after you trained it for lets say 5-10 minutes? To me it sounds very much like overheating, this is typical for freezes.</p>",
      "rawMarkdown": "Is wandb not tracking the temperature? Whats your temperature after you trained it for lets say 5-10 minutes? To me it sounds very much like overheating, this is typical for freezes.",
      "votes": null
    },
    {
      "id": "1506573",
      "postDate": "09/08/2021 11:08:16",
      "content": "<p>Maybe its a issue with the memory junction temps. GDDR6X gets really hot. That might be a problem</p>",
      "rawMarkdown": "Maybe its a issue with the memory junction temps. GDDR6X gets really hot. That might be a problem",
      "votes": null
    },
    {
      "id": "1507130",
      "postDate": "09/08/2021 21:16:58",
      "content": "<p>I was not using wandb.soon i will use it and will let you know.thank you</p>",
      "rawMarkdown": "I was not using wandb.soon i will use it and will let you know.thank you",
      "votes": null
    },
    {
      "id": "1507131",
      "postDate": "09/08/2021 21:18:33",
      "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> using this one : <a href=\"https://www.coolermaster.com/catalog/power-supplies/mwe-series/mwe-gold-850-v2-full-modular/\" target=\"_blank\">https://www.coolermaster.com/catalog/power-supplies/mwe-series/mwe-gold-850-v2-full-modular/</a></p>\n<p>And this cooler : <a href=\"https://www.coolermaster.com/catalog/coolers/cpu-air-coolers/hyper-212-led-turbo/\" target=\"_blank\">https://www.coolermaster.com/catalog/coolers/cpu-air-coolers/hyper-212-led-turbo/</a></p>",
      "rawMarkdown": "mithilsalunkhe using this one : https://www.coolermaster.com/catalog/power-supplies/mwe-series/mwe-gold-850-v2-full-modular/\n\nAnd this cooler : https://www.coolermaster.com/catalog/coolers/cpu-air-coolers/hyper-212-led-turbo/",
      "votes": null
    },
    {
      "id": "1507255",
      "postDate": "09/09/2021 03:00:27",
      "content": "<p>Does your case have enough airflow ?</p>",
      "rawMarkdown": "Does your case have enough airflow ?",
      "votes": null
    },
    {
      "id": "1509883",
      "postDate": "09/11/2021 18:44:40",
      "content": "<p>I have preprocessed and saved the data but without any improvements so far. I guess either the bottleneck is in the IO operations (i.e. opening files) since the processed images are on a \"classical\" HDD (and not SSD) or I am doing something wrong? </p>\n<p>Notice that I am using the EfficienNet B0 model and 512 x 512 images. </p>",
      "rawMarkdown": "I have preprocessed and saved the data but without any improvements so far. I guess either the bottleneck is in the IO operations (i.e. opening files) since the processed images are on a \"classical\" HDD (and not SSD) or I am doing something wrong? \n\nNotice that I am using the EfficienNet B0 model and 512 x 512 images.",
      "votes": null
    },
    {
      "id": "1509884",
      "postDate": "09/11/2021 18:44:59",
      "content": "<p>That's quite impressive!</p>",
      "rawMarkdown": "That's quite impressive!",
      "votes": null
    },
    {
      "id": "1509885",
      "postDate": "09/11/2021 18:46:19",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> any update on your situation? 👀</p>",
      "rawMarkdown": "mobassir any update on your situation? 👀",
      "votes": null
    },
    {
      "id": "1509902",
      "postDate": "09/11/2021 19:08:48",
      "content": "<p>how much RAM do you have?</p>\n<p>It maybe better to upload all waves in ram and compute images on the fly using packages that run on GPU like torch.fft, nnaudio, etc.</p>",
      "rawMarkdown": "how much RAM do you have?\n\nIt maybe better to upload all waves in ram and compute images on the fly using packages that run on GPU like torch.fft, nnaudio, etc.",
      "votes": null
    },
    {
      "id": "1509917",
      "postDate": "09/11/2021 19:26:08",
      "content": "<p>RAM of 64GB. </p>\n<p>Yes, if my SSD test doesn't give better results, I will for sure move to this option. I am also checking on a cloud GPU to see if the problem comes from my current setup. </p>\n<p>Will let you know.</p>",
      "rawMarkdown": "RAM of 64GB. \n\nYes, if my SSD test doesn't give better results, I will for sure move to this option. I am also checking on a cloud GPU to see if the problem comes from my current setup. \n\nWill let you know.",
      "votes": null
    },
    {
      "id": "1509946",
      "postDate": "09/11/2021 20:00:42",
      "content": "<p>SSD should definitely help. Move all your training data to it.</p>",
      "rawMarkdown": "SSD should definitely help. Move all your training data to it.",
      "votes": null
    },
    {
      "id": "1509976",
      "postDate": "09/11/2021 21:12:15",
      "content": "<p>Update: I have resized the images to 256 x 256 and still using EfficientNet B0 and now can train one epoch for one fold in 1<strong>2 minutes</strong> (GPU usage is almost at max). </p>\n<p>The secret was in getting a much bigger batch size: around 300.</p>\n<p>I can probably optimize even further but at least now I can run more experiments. 👌</p>",
      "rawMarkdown": "Update: I have resized the images to 256 x 256 and still using EfficientNet B0 and now can train one epoch for one fold in 1**2 minutes** (GPU usage is almost at max). \n\nThe secret was in getting a much bigger batch size: around 300.\n\nI can probably optimize even further but at least now I can run more experiments. 👌",
      "votes": null
    },
    {
      "id": "1510154",
      "postDate": "09/12/2021 05:56:25",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealtuned\" target=\"_blank\">@yassinealtuned</a><br>\nCurrently i am using kaggle tpu for this competition… </p>\n<p>sorry i am yet to train anything again with rtx3090.<br>\nProblem is, i am using a server pc through remote desktop connection which has rtx3090.that server pc is hosted in my office which is far away from me and i am working from home!<br>\nI use that pc via remote desktop connection <br>\nSo It's difficult for me to restart the pc over and over again for this painful freezing issue &gt;_&lt;<br>\nWhen it freezes,i lose connection and can't Connect to pc via remote desktop connection and then i disturb office colleagues/staff for restarting the pc(who is there in the office at that Moment)<br>\nLast Thursday the pc got disconnected (maybe because of either shutdown/freeze issue)<br>\nThen here in my territory Friday and Saturday’s are off day,so i found no office staff to restart the pc<br>\nToday is sunday,so one office staff just turned on the pc again and now during daytime my office colleagues are using that pc for generating synthetic data and train models related to optical character recognition (ocr) for daytime office Project. <br>\nThey(my office team mates) do not compete in kaggle<br>\nI plan to try wandb again within next 2-3 days when the pc is not busy for training model for official works! <br>\nI Just hate this freeze issue,restarting the pc for me is very difficult task as i work for home :(<br>\nStay tuned,thanks</p>",
      "rawMarkdown": "yassinealtuned\nCurrently i am using kaggle tpu for this competition... \n\n sorry i am yet to train anything again with rtx3090.\nProblem is, i am using a server pc through remote desktop connection which has rtx3090.that server pc is hosted in my office which is far away from me and i am working from home!\nI use that pc via remote desktop connection \nSo It's difficult for me to restart the pc over and over again for this painful freezing issue >_<\nWhen it freezes,i lose connection and can't Connect to pc via remote desktop connection and then i disturb office colleagues/staff for restarting the pc(who is there in the office at that Moment)\nLast Thursday the pc got disconnected (maybe because of either shutdown/freeze issue)\nThen here in my territory Friday and Saturday’s are off day,so i found no office staff to restart the pc\nToday is sunday,so one office staff just turned on the pc again and now during daytime my office colleagues are using that pc for generating synthetic data and train models related to optical character recognition (ocr) for daytime office Project. \nThey(my office team mates) do not compete in kaggle\nI plan to try wandb again within next 2-3 days when the pc is not busy for training model for official works! \nI Just hate this freeze issue,restarting the pc for me is very difficult task as i work for home :(\nStay tuned,thanks",
      "votes": null
    },
    {
      "id": "1510947",
      "postDate": "09/12/2021 23:58:41",
      "content": "<p>I switched to nn Audio and I also get similar results but with a smaller batchsize</p>",
      "rawMarkdown": "I switched to nn Audio and I also get similar results but with a smaller batchsize",
      "votes": null
    },
    {
      "id": "1513397",
      "postDate": "09/15/2021 06:30:47",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> this is the first time i used wandb,sorry if i am missing something.<br>\nhere is the wandb folder where i have log : <a href=\"https://www.kaggle.com/mobassir/wandb-log\" target=\"_blank\">https://www.kaggle.com/mobassir/wandb-log</a><br>\ni can't see any temperature related log,don't know how to interpret,let me know if you need the code that i tried to generate these logs,thanks</p>",
      "rawMarkdown": "yassinealouini this is the first time i used wandb,sorry if i am missing something.\nhere is the wandb folder where i have log : https://www.kaggle.com/mobassir/wandb-log\ni can't see any temperature related log,don't know how to interpret,let me know if you need the code that i tried to generate these logs,thanks",
      "votes": null
    },
    {
      "id": "1513451",
      "postDate": "09/15/2021 07:23:39",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> do you have access to the online dashboard? This is the easiest way to get some system information. I will have a look at your logs and let you know later. 👌</p>",
      "rawMarkdown": "mobassir do you have access to the online dashboard? This is the easiest way to get some system information. I will have a look at your logs and let you know later. 👌",
      "votes": null
    },
    {
      "id": "1513498",
      "postDate": "09/15/2021 07:56:54",
      "content": "<p>sorry <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> <br>\ni don't have any online dashboard,i am simply using wandb like this kernel : <a href=\"https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training</a></p>\n<p>i just modified his code a little bit for locally training,,i use like this : </p>\n<p>import wandb<br>\nwandb.login()</p>\n<p>and i do not use these : </p>\n<h1>from kaggle_secrets import UserSecretsClient</h1>\n<h1>user_secrets = UserSecretsClient()</h1>\n<h1>wandb_api = user_secrets.get_secret(\"wandb_api\")</h1>\n<h1>wandb.login(key=wandb_api)</h1>\n<p>and then he was doing this : </p>\n<p>run = wandb.init(project=\"G2Net-Public-experiments\", <br>\n                 name=\"exp1\",<br>\n                 config=class2dict(CFG),<br>\n                 group=CFG.model_name,<br>\n                 job_type=\"train\")</p>\n<p>where i did this : </p>\n<p>run = wandb.init(project=\"G2Net-Public-experiments\", <br>\n                 name=\"exp1\",<br>\n                 config=class2dict(CFG),<br>\n                 group=CFG.model_name,<br>\n                 job_type=\"train\",<br>\n                 settings=wandb.Settings(_disable_stats=False))</p>\n<p>everything else is same for wandb logging like him</p>",
      "rawMarkdown": "sorry @yassinealouini \ni don't have any online dashboard,i am simply using wandb like this kernel : https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\n\ni just modified his code a little bit for locally training,,i use like this : \n\n\nimport wandb\nwandb.login()\n\nand i do not use these : \n\n#from kaggle_secrets import UserSecretsClient\n# user_secrets = UserSecretsClient()\n# wandb_api = user_secrets.get_secret(\"wandb_api\")\n# wandb.login(key=wandb_api)\n\nand then he was doing this : \n\nrun = wandb.init(project=\"G2Net-Public-experiments\", \n                 name=\"exp1\",\n                 config=class2dict(CFG),\n                 group=CFG.model_name,\n                 job_type=\"train\")\n\nwhere i did this : \n\nrun = wandb.init(project=\"G2Net-Public-experiments\", \n                 name=\"exp1\",\n                 config=class2dict(CFG),\n                 group=CFG.model_name,\n                 job_type=\"train\",\n                 settings=wandb.Settings(_disable_stats=False))\n\neverything else is same for wandb logging like him",
      "votes": null
    },
    {
      "id": "1513499",
      "postDate": "09/15/2021 07:58:32",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> Are there any other GPUs in that machine? The reason I ask is that when the 2080Ti came out, I wanted to use it alongside my old GTX 1080. But during training it would randomly freeze even if the 1080 was idle and nothing was showing on any logs (therefore a catastrophic kernel issue). The only way I could fix it was by removing the 1080 from the system, and I think the crashing was probably caused by some driver issue caused by mixing different generation cards.</p>",
      "rawMarkdown": "mobassir Are there any other GPUs in that machine? The reason I ask is that when the 2080Ti came out, I wanted to use it alongside my old GTX 1080. But during training it would randomly freeze even if the 1080 was idle and nothing was showing on any logs (therefore a catastrophic kernel issue). The only way I could fix it was by removing the 1080 from the system, and I think the crashing was probably caused by some driver issue caused by mixing different generation cards.",
      "votes": null
    },
    {
      "id": "1513515",
      "postDate": "09/15/2021 08:10:26",
      "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> i am using only rtx3090(single gpu)</p>\n<p>Processor    : i9 10th Gen 10Core<br>\nRAM          : 32GB (3200Mhz)<br>\nHDD          : 1TB SSD Nvme<br>\nGPU          : MSI Nvidia RTX 3090 24GB</p>",
      "rawMarkdown": "anjum48 i am using only rtx3090(single gpu)\n\nProcessor    : i9 10th Gen 10Core\nRAM          : 32GB (3200Mhz)\nHDD          : 1TB SSD Nvme\nGPU          : MSI Nvidia RTX 3090 24GB",
      "votes": null
    },
    {
      "id": "1513726",
      "postDate": "09/15/2021 11:07:29",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I only used wandb once so far, and outside Kaggle, but it seems you need to create a wandb account and use your credentials.  I have not user kaggle secret either but it looks like a way to publicly share notebooks without revealing your credentials.</p>",
      "rawMarkdown": "mobassir I only used wandb once so far, and outside Kaggle, but it seems you need to create a wandb account and use your credentials.  I have not user kaggle secret either but it looks like a way to publicly share notebooks without revealing your credentials.",
      "votes": null
    },
    {
      "id": "1514121",
      "postDate": "09/15/2021 17:55:19",
      "content": "<p>dear <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> here you can see my system metric(screenshot) : <a href=\"https://www.kaggle.com/mobassir/wandb-log?select=1.PNG\" target=\"_blank\">https://www.kaggle.com/mobassir/wandb-log?select=1.PNG</a><br>\ncheck 1.png,2.png,3.png and 4.png<br>\nthank you</p>",
      "rawMarkdown": "dear @cpmpml @yassinealouini @philippsinger here you can see my system metric(screenshot) : https://www.kaggle.com/mobassir/wandb-log?select=1.PNG\ncheck 1.png,2.png,3.png and 4.png\nthank you",
      "votes": null
    },
    {
      "id": "1517966",
      "postDate": "09/20/2021 08:50:19",
      "content": "<p>For those still following this thread, the fastest way I have found so far is using <strong>TFRecords</strong> and <strong>TPUs</strong>. I will share how long it takes once the competition is over. 👌</p>",
      "rawMarkdown": "For those still following this thread, the fastest way I have found so far is using **TFRecords** and **TPUs**. I will share how long it takes once the competition is over. 👌",
      "votes": null
    },
    {
      "id": "1517967",
      "postDate": "09/20/2021 08:51:50",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> it looks like the graphs are fine except the first dip we see. I don't have other options to offer for now. 😬</p>",
      "rawMarkdown": "mobassir it looks like the graphs are fine except the first dip we see. I don't have other options to offer for now. 😬",
      "votes": null
    },
    {
      "id": "1520899",
      "postDate": "09/22/2021 17:22:20",
      "content": "<p>Something that I've now noticed. It's always the model being trained on gpu:0 that crashes. I'm training on a desktop, so I wonder if it has something to do with X / the display system interfering. Once this competition is done, I plan on swapping the physical positions of the two cards to see if the data loader thread crashing persists on the same GPU device or on the same slot to better debug.</p>\n<p>Also, does anyone see:</p>\n<p><code>[W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool)</code> being spammed to the console where jupyter notebook is invoked?</p>",
      "rawMarkdown": "Something that I've now noticed. It's always the model being trained on gpu:0 that crashes. I'm training on a desktop, so I wonder if it has something to do with X / the display system interfering. Once this competition is done, I plan on swapping the physical positions of the two cards to see if the data loader thread crashing persists on the same GPU device or on the same slot to better debug.\n\nAlso, does anyone see:\n\n`[W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool)` being spammed to the console where jupyter notebook is invoked?",
      "votes": null
    },
    {
      "id": "1520942",
      "postDate": "09/22/2021 17:59:41",
      "content": "<blockquote>\n  <p>Also, does anyone see:</p>\n  <p>[W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool) being spammed to the console where jupyter notebook is invoked?</p>\n</blockquote>\n<p>Yes, that's a very annoying bug in pytorch (which looks like it might be fixed in 1.10) <a href=\"https://github.com/pytorch/pytorch/issues/57273\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/57273</a></p>",
      "rawMarkdown": "> Also, does anyone see:\n\n> [W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool) being spammed to the console where jupyter notebook is invoked?\n\nYes, that's a very annoying bug in pytorch (which looks like it might be fixed in 1.10) https://github.com/pytorch/pytorch/issues/57273",
      "votes": null
    },
    {
      "id": "1521032",
      "postDate": "09/22/2021 20:04:11",
      "content": "<p>Wow, I took out <code>pin_memory=True</code> and behold, all my crashing issues have been resolved.</p>",
      "rawMarkdown": "Wow, I took out `pin_memory=True` and behold, all my crashing issues have been resolved.",
      "votes": null
    },
    {
      "id": "1521047",
      "postDate": "09/22/2021 20:20:05",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Yes, I had a similar warning.</p>",
      "rawMarkdown": "authman Yes, I had a similar warning.",
      "votes": null
    },
    {
      "id": "1524715",
      "postDate": "09/26/2021 19:01:33",
      "content": "<p>Great! :-) Thanks to your discussion post I started to learn much more about how to optimise my GPU and TPU performances and it still goes on…  </p>",
      "rawMarkdown": "Great! :-) Thanks to your discussion post I started to learn much more about how to optimise my GPU and TPU performances and it still goes on...",
      "votes": null
    },
    {
      "id": "1524777",
      "postDate": "09/26/2021 20:38:00",
      "content": "<p>That's awesome! I have learned many things as well, might share them once the competition ends. 👌</p>",
      "rawMarkdown": "That's awesome! I have learned many things as well, might share them once the competition ends. 👌",
      "votes": null
    },
    {
      "id": "1559736",
      "postDate": "10/27/2021 07:10:28",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1504370,
      "author_name": "hamishdickson",
      "author_url": "",
      "post_date": "09/06/2021 10:18:15",
      "content": "<p>that's a long time - I also have a 3090 and can train an epoch of b0 in less than 3mins. My best model is a b2 and it's current setup takes around 12mins an epoch</p>\n<p>my guess is you're either using a huge model or huge image or you have some bottleneck?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1504372,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "09/06/2021 10:19:02",
          "content": "<p>I assume you're doing all the usual mixed precision/pin_memory/etc optimisations?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504400,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/06/2021 10:45:22",
          "content": "<p>Not that big only a B3. Also, I am generating image features on the fly so probably not optimal for now. Thanks for providing these numbers, I will check what can be improved. 👌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504405,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/06/2021 10:51:49",
          "content": "<p><a href=\"https://www.kaggle.com/hamishdickson\" target=\"_blank\">@hamishdickson</a> <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> i am also experimenting using rtx3090 but when i converted wave signal into image and resized it down to 512x512 it takes around 1 hour to train a b0 model for 1 epoch and then my pc shuts down,i can't figure out why pc getting turned off during training,,,when i was using wave data only,,at that time training was fast and pc never got shut down,,but after converting it to 512x512 image size,pc shutting down after few hours of training(sometimes less than 1 hour)<br>\ni was working using this kernel : <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch</a><br>\nduring training i checked nvidia-smi and i was using less than 20gb vram out of 24</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504415,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "09/06/2021 11:02:37",
          "content": "<blockquote>\n  <p>and then my pc shuts down</p>\n</blockquote>\n<p>Undervolt your GPU and check if the problem goes away. I had the same problem when my old PSU gradually died - big power spikes during initialization forced PSU to shut down.</p>\n<p>3090, depending on model, can generate power spikes up to ~600W. Undervolting with 5-10% performance loss might keep them to ~350-400W max.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504429,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/06/2021 11:14:08",
          "content": "<p>That's very interesting. I will check my wandb dashboard but I don't have those spikes I think. Thanks for sharing <a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a>. 👌 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504434,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/06/2021 11:25:50",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> thank you for sharing<br>\ni saw that my vram usage was low(less than 20gb from 24gb),so i thought cpu or ram or something else is causing this issue but i can't figure out,i checked kern log and found no problem,i thought it's num_workers that is generating more heat and causing this issue so set num_workers = 0 but no luck,i have gold certified good power supply,good cpu cooler but still for some reason my pc getting shut down,i hope your idea of Undervolt works for me,,otherwise i don't know how to solve this issue,thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504438,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "09/06/2021 11:30:22",
          "content": "<blockquote>\n  <p>I will check my wandb dashboard but I don't have those spikes I think.</p>\n</blockquote>\n<p>They are not visible on monitoring software, because they are too short. Long enough for some PSUs to trip overcurrent protection, though.</p>\n<p>I could not find original igorslab article about them (when 3080/3090 on release crashed many PCs), only his testing after the drivers fix - <a href=\"https://www.igorslab.de/en/wonder-how-invidia-the-crashes-of-the-force-rtx-3080-andrtx-3090-will-be-removed-and-still-will-be-removed-even-from-the-power-supplies-analysis/\" target=\"_blank\">https://www.igorslab.de/en/wonder-how-invidia-the-crashes-of-the-force-rtx-3080-andrtx-3090-will-be-removed-and-still-will-be-removed-even-from-the-power-supplies-analysis/</a></p>\n<p>Illustration from there (3080, should be more or less actual behavior): <a href=\"https://www.igorslab.de/wp-content/uploads/2020/09/12a-Gaming-Zoom-Power-New-Drivers.png\" target=\"_blank\">https://www.igorslab.de/wp-content/uploads/2020/09/12a-Gaming-Zoom-Power-New-Drivers.png</a></p>\n<p>Edit: Found this topics that might prove helpful, i read them during my own issues - </p>\n<p><a href=\"https://discuss.pytorch.org/t/multi-gpu-2080-ti-training-crashes-pc/52032/9\" target=\"_blank\">https://discuss.pytorch.org/t/multi-gpu-2080-ti-training-crashes-pc/52032/9</a></p>\n<p><a href=\"https://github.com/pytorch/pytorch/issues/3022\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/3022</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504480,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/06/2021 12:18:15",
          "content": "<p>So I'm not alone. I had been having 3090 issues as well for the past few months. In my case, the computer does not freeze but training with segfault and I cannot reclaim the vram until I do a restart. Also it's usually GPU specific, that is, I can still train on the other GPU. Though occasionally when one blows, it takes both down. I have 1600W PSU though so I even if both pull spike their load, I wouldn't expect that to take down the ship. It is so annoying with these longer model train times, you let the models sail all night long only to wake up and discover stuff froze after the first fold.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504520,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/06/2021 12:52:03",
          "content": "<p><strong>you let the models sail all night long only to wake up and discover stuff froze after the first fold.</strong> 😪</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504580,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/06/2021 13:35:56",
          "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> image size is also a factor.  Depending on the image size your timing may just be good, or bad.  One way to check is to look at the memory and compute use of your GPU using nvidia-smi.  </p>\n<p>For instance, if use is close to 100% and memory close to GPU memory then you are probably fine.  </p>\n<p>Another example: if use is 100% but memory is only half of GPU memory then you can probably double batch size and divide time by 2.  Of course, changing batch size may impact other hyper parameter values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504583,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/06/2021 13:41:24",
          "content": "<p>Not to hijack the thread and perhaps too close to competition deadline, but <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> if I might inquire, what bs are you currently utilizing in your model? It was something I'd actually been wondering. I'm currently using a modest 32 bs resulting in 98% utilization, however +85% vram is free… so I wonder.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504588,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/06/2021 13:45:52",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> It looks like you should upgrade to the latest driver, and also check pour power unit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504595,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/06/2021 13:51:44",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for the details. I will check these and report back later. 👌 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504597,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/06/2021 13:54:01",
          "content": "<p><a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a> previously i had 370watt limit for gpu and now after undervolting 70watt i have this : <br>\nMon Sep  6 19:12:16 2021       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 460.91.03    Driver Version: 460.91.03    CUDA Version: 11.2     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  GeForce RTX 3090    Off  | 00000000:01:00.0 Off |                  N/A |<br>\n|  0%   28C    P8    12W / 300W |     19MiB / 24268MiB |      0%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n|    0   N/A  N/A      1140      G   /usr/lib/xorg/Xorg                  9MiB |<br>\n|    0   N/A  N/A      1319      G   /usr/bin/gnome-shell                8MiB |<br>\n+-----------------------------------------------------------------------------+</p>\n<p><strong>nvcc --version</strong></p>\n<p>nvcc: NVIDIA (R) Cuda compiler driver<br>\nCopyright (c) 2005-2021 NVIDIA Corporation<br>\nBuilt on Sun_Feb_14_21:12:58_PST_2021<br>\nCuda compilation tools, release 11.2, V11.2.152<br>\nBuild cuda_11.2.r11.2/compiler.29618528_0</p>\n<p>can you suggest me which driver version i should try now? thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504670,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/06/2021 14:58:03",
          "content": "<blockquote>\n  <p>however +85% vram is free… so I wonder.</p>\n</blockquote>\n<p>You should use  a larger batch size IMHO.  It is not what cause your problem of course, but you are under using your GPU it seems</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504693,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/06/2021 15:25:05",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> 😪<br>\npreviously i had 370watt limit for gpu and then i undervolted 70watt and limited gpu to use only 300watt.<br>\ni reduced image size from 512 to 256<br>\ni increased batch size (but vram usage was still less than 19gb out of 24)<br>\ni used mixed precision this time<br>\n<strong>PC SHUTS DOWN AGAIN AFTER 2nd EPOCH</strong>  😭😪</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504721,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "09/06/2021 15:54:25",
          "content": "<p>Are you caching tf valid dataset? If so try disabling it. Also try empty cache after every epoch. I haven't encountered this issue so I am just throwing ideas. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504725,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/06/2021 15:56:52",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> could it be a hardware defect in the GPU? Have you tried contacting NVIDIA support? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504774,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/06/2021 16:33:50",
          "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> i am not caching tf valid dataset,,i tried another baseline and encountered same issue,,for example if i experiment using this notebook : <a href=\"https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training</a><br>\nreplace b7 with b0<br>\nin custom model class use F.interpolate to convert wave data into 512x512 image size<br>\nadjust batch size that takes just less than 20gb vram and train <strong>(KEEP EVERYTHING AS IT IS IN THAT PUBLIC KERNEL)</strong></p>\n<p>still my pc turns off <strong>(this time after 5 epoch)</strong></p>\n<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> i Have not tried contacting NVIDIA support yet,when i do not convert wave data into 512x512 image data i don't face this issue,,i tried to train few nlp models and advance scene text detection/recognition models and didn't face this issue<br>\ni face this issue occasionally and i can't understand how to solve this issue</p>\n<p>instead of shutdown if i could use a command/setting that will restart the pc instead of freezing/turning it off then it could be good for me but i don't know if it's possible to do so </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504810,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "09/06/2021 17:03:10",
          "content": "<p>obviously you monitoring gpu temperatures? nvidia gpus have shutdown thresholds at ~90c.   it should not lead to pc shutdown thou</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504834,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/06/2021 17:24:27",
          "content": "<blockquote>\n  <p>try empty cache after every epoch</p>\n</blockquote>\n<p>I used to do this when I had dual 2080's and it would balloon train times like crazzzyyyyyyyyyyy. Now, I just let the thing take care of the thing by itself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504858,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/06/2021 17:48:07",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I see, thanks for the clarification. So it seems to be a specific issue to this specific model setup. Have you tried very small image size (let's say 64 x 64) and/or removing the EfficientNet and replacing it by the identity or another model? </p>\n<p>Not sure how to help but if I were in your situation, I would try various combinations until I find something. Best of luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505145,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "09/07/2021 02:07:06",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> In case of 3090 it also might be caused by high memory junction temperatures. I think they are only accessible to monitoring software under Windows, for example in HWinfo64. GDDR6X runs very hot, and sometimes reaches thermal throttle at 110 degrees.</p>\n<p>If this is the case (you can quickly run any Eth mining software on Windows as a way to stress test VRAM), then changing thermal pads usually improves the situation, but that involves disassembly of the cooling system and might complicate warranty. Another way could be just propping up your 3090 in the case - <a href=\"https://www.reddit.com/r/nvidia/comments/pihwzo/3090_vram_temperature_fix_without_replacing_pads/\" target=\"_blank\">https://www.reddit.com/r/nvidia/comments/pihwzo/3090_vram_temperature_fix_without_replacing_pads/</a></p>\n<p>Another relatively common mistake is connecting GPU with one cable from PSU &gt; two 8-pin connectors on GPU. Always use separate cables for separate 8-pin connectors, i think it is even in installation manual.</p>\n<p>However, if VRAM junction temperature is not above 100C, undervolting does not help with your shutdowns, and card is connected properly, and you are still having shutdowns - then it might be a general hardware instability, which is a nightmare to debug. I would still check with a ridiculously overpowered PSU just in case, but at that point i would probably resort to changing components one by one to find the root cause.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505296,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/07/2021 06:10:34",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> <a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> </p>\n<p>apology for a little mistake from my end,</p>\n<p><strong>The pc is not getting turned off,the problem is it is getting frozen every time and forcing me to restart everytime</strong></p>\n<p>for reproducing the result i just took the public kernel of <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> san and modified a little bit to train and face the freeze issue so that i can share the code and log with you guys for better investigation..</p>\n<p>here is the code and log file : <a href=\"https://pastebin.com/XEusHppZ\" target=\"_blank\">https://pastebin.com/XEusHppZ</a><br>\nyou can see from my log that during 7th epoch  i faced 'freeze' issue and i had to restart the pc.<br>\ni was using around 19gb vram out of 24 iirc while training this model (full code is provided).</p>\n<p>could this be related to pytorch version or something else?</p>\n<p>it will be highly appreciated if any kaggler help me solve this issue (i've shared the full code),thank you a lot in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505311,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/07/2021 06:33:51",
          "content": "<p>One \"cheap\" solution is to stop after 3 or 4 epochs maybe? </p>\n<p>Otherwise, maybe try to reinstall everything? I've taken few notes here: <a href=\"https://gist.github.com/yassineAlouini/e843531219e1b9d978aa4fe2a5e71940\" target=\"_blank\">https://gist.github.com/yassineAlouini/e843531219e1b9d978aa4fe2a5e71940</a>. There is also this good blog post: <a href=\"https://www.itzgeek.com/post/how-to-install-nvidia-drivers-on-ubuntu-20-04-ubuntu-18-04.html\" target=\"_blank\">https://www.itzgeek.com/post/how-to-install-nvidia-drivers-on-ubuntu-20-04-ubuntu-18-04.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505379,
          "author_name": "karmaz",
          "author_url": "",
          "post_date": "09/07/2021 07:42:24",
          "content": "<p>I had same issues as yours. I changed the PSU but the problem was still persistent. Later, I changed the motherboard boom it started working again, No freeze or random shutdown. Might be problem due to current spike causing damage to weaker motherboard capacitors/components.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505606,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/07/2021 12:03:05",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> what was the GPU temperature when it froze?  </p>\n<p>Also, as <a href=\"https://www.kaggle.com/karmaz\" target=\"_blank\">@karmaz</a> indicated, it maybe a bad motherboard.  However, before changing it, have you tried to remove the GPU and disconnect everything, then connect and plug it back?  Sometimes a defective connection is the cause of such issues.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505714,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/07/2021 13:52:42",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> i am  not sure if it is really a hardware related issue,,i thought this issue is related to pytorch version maybe,i am using the following packages : <br>\nName Version Build Channel<br>\n_libgcc_mutex 0.1 main<br>\n_openmp_mutex 4.5 1_gnu<br>\naddict 2.4.0 pypi_0 pypi<br>\nalabaster 0.7.12 py_0 conda-forge<br>\nalbumentations 1.0.3 pypi_0 pypi<br>\nalsa-lib 1.2.3 h516909a_0 conda-forge<br>\nappdirs 1.4.4 pyh9f0ad1d_0 conda-forge<br>\nargh 0.26.2 pyh9f0ad1d_1002 conda-forge<br>\nastroid 2.6.0 py38h578d9bd_0 conda-forge<br>\nasync_generator 1.10 py_0 conda-forge<br>\nasynctest 0.13.0 pypi_0 pypi<br>\natomicwrites 1.4.0 pyh9f0ad1d_0 conda-forge<br>\nattrs 21.2.0 pyhd8ed1ab_0 conda-forge<br>\nautopep8 1.5.6 pyhd8ed1ab_0 conda-forge<br>\nbabel 2.9.1 pyh44b312d_0 conda-forge<br>\nbackcall 0.2.0 pyhd3eb1b0_0<br>\nblack 21.5b2 pyhd8ed1ab_0 conda-forge<br>\nbleach 3.3.0 pyh44b312d_0 conda-forge<br>\nbrotlipy 0.7.0 py38h497a2fe_1001 conda-forge<br>\nca-certificates 2021.5.30 ha878542_0 conda-forge<br>\ncertifi 2021.5.30 py38h578d9bd_0 conda-forge<br>\ncffi 1.14.5 py38ha65f79e_0 conda-forge<br>\ncfgv 3.3.0 pypi_0 pypi<br>\nchardet 4.0.0 py38h578d9bd_1 conda-forge<br>\nclick 8.0.1 py38h578d9bd_0 conda-forge<br>\ncloudpickle 1.6.0 py_0 conda-forge<br>\ncodecov 2.1.11 pypi_0 pypi<br>\ncolorama 0.4.4 pyh9f0ad1d_0 conda-forge<br>\nconfigparser 5.0.2 pypi_0 pypi<br>\ncoverage 5.5 pypi_0 pypi<br>\ncryptography 3.4.7 py38ha5dfef3_0 conda-forge<br>\ncudatoolkit 11.1.74 h6bb024c_0 nvidia<br>\ncycler 0.10.0 pypi_0 pypi<br>\ncython 0.29.23 pypi_0 pypi<br>\ndataclasses 0.8 pyhc8e2a94_1 conda-forge<br>\ndbus 1.13.6 h48d8840_2 conda-forge<br>\ndecorator 4.4.2 pypi_0 pypi<br>\ndefusedxml 0.7.1 pyhd8ed1ab_0 conda-forge<br>\ndiff-match-patch 20200713 pyh9f0ad1d_0 conda-forge<br>\ndill 0.3.4 pypi_0 pypi<br>\ndistlib 0.3.2 pypi_0 pypi<br>\ndocker-pycreds 0.4.0 pypi_0 pypi<br>\ndocutils 0.17.1 py38h578d9bd_0 conda-forge<br>\nentrypoints 0.3 pyhd8ed1ab_1003 conda-forge<br>\nexpat 2.4.1 h9c3ff4c_0 conda-forge<br>\nfilelock 3.0.12 pypi_0 pypi<br>\nflake8 3.9.2 pypi_0 pypi<br>\nfontconfig 2.13.1 hba837de_1005 conda-forge<br>\nfreetype 2.10.4 h0708190_1 conda-forge<br>\nfuture 0.18.2 py38h578d9bd_3 conda-forge<br>\ngast 0.3.3 pypi_0 pypi<br>\ngettext 0.19.8.1 h0b5b191_1005 conda-forge<br>\ngitdb 4.0.7 pypi_0 pypi<br>\ngitpython 3.1.18 pypi_0 pypi<br>\nglib 2.68.3 h9c3ff4c_0 conda-forge<br>\nglib-tools 2.68.3 h9c3ff4c_0 conda-forge<br>\ngoogleapis-common-protos 1.53.0 pypi_0 pypi<br>\ngrad-cam 1.3.1 pypi_0 pypi<br>\ngrpcio 1.32.0 pypi_0 pypi<br>\ngst-plugins-base 1.18.4 hf529b03_2 conda-forge<br>\ngstreamer 1.18.4 h76c114f_2 conda-forge<br>\nh5py 2.10.0 pypi_0 pypi<br>\nhelpdev 0.7.1 pyhd8ed1ab_0 conda-forge<br>\nicu 68.1 h58526e2_0 conda-forge<br>\nidentify 2.2.10 pypi_0 pypi<br>\nidna 2.10 pyh9f0ad1d_0 conda-forge<br>\nimageio 2.9.0 pypi_0 pypi<br>\nimagesize 1.2.0 py_0 conda-forge<br>\nimgaug 0.4.0 pypi_0 pypi<br>\nimportlib-metadata 4.5.0 py38h578d9bd_0 conda-forge<br>\nimportlib-resources 5.2.2 pypi_0 pypi<br>\nimportlib_metadata 4.5.0 hd8ed1ab_0 conda-forge<br>\niniconfig 1.1.1 pypi_0 pypi<br>\nintervaltree 3.0.2 py_0 conda-forge<br>\nipykernel 5.5.5 py38hd0cf306_0 conda-forge<br>\nipython 7.18.1 py38h5ca1d4c_0 anaconda<br>\nipython_genutils 0.2.0 pyhd3eb1b0_1<br>\nipywidgets 7.6.3 pypi_0 pypi<br>\nisort 5.8.0 pypi_0 pypi<br>\njedi 0.17.2 py38h578d9bd_1 conda-forge<br>\njeepney 0.6.0 pyhd8ed1ab_0 conda-forge<br>\njinja2 3.0.1 pyhd8ed1ab_0 conda-forge<br>\njpeg 9d h36c2ea0_0 conda-forge<br>\njsonschema 3.2.0 pyhd8ed1ab_3 conda-forge<br>\njupyter_client 6.1.12 pyhd8ed1ab_0 conda-forge<br>\njupyter_core 4.7.1 py38h578d9bd_0 conda-forge<br>\njupyterlab-widgets 1.0.0 pypi_0 pypi<br>\njupyterlab_pygments 0.1.2 pyh9f0ad1d_0 conda-forge<br>\nkaggle 1.5.12 pypi_0 pypi<br>\nkaggledatasets 0.0.1 pypi_0 pypi<br>\nkeyring 23.0.1 py38h578d9bd_0 conda-forge<br>\nkiwisolver 1.3.1 pypi_0 pypi<br>\nkrb5 1.19.1 hcc1bbae_0 conda-forge<br>\nkwarray 0.5.19 pypi_0 pypi<br>\nlanms-proper 1.0.1 pypi_0 pypi<br>\nlazy-object-proxy 1.6.0 py38h497a2fe_0 conda-forge<br>\nld_impl_linux-64 2.35.1 h7274673_9<br>\nlibclang 11.1.0 default_ha53f305_1 conda-forge<br>\nlibedit 3.1.20191231 he28a2e2_2 conda-forge<br>\nlibevent 2.1.10 hcdb4288_3 conda-forge<br>\nlibffi 3.3 he6710b0_2<br>\nlibgcc-ng 9.3.0 h5101ec6_17<br>\nlibglib 2.68.3 h3e27bee_0 conda-forge<br>\nlibgomp 9.3.0 h5101ec6_17<br>\nlibiconv 1.16 h516909a_0 conda-forge<br>\nlibllvm11 11.1.0 hf817b99_2 conda-forge<br>\nlibogg 1.3.4 h7f98852_1 conda-forge<br>\nlibopus 1.3.1 h7f98852_1 conda-forge<br>\nlibpng 1.6.37 h21135ba_2 conda-forge<br>\nlibpq 13.3 hd57d9b9_0 conda-forge<br>\nlibsodium 1.0.18 h36c2ea0_1 conda-forge<br>\nlibspatialindex 1.9.3 h9c3ff4c_3 conda-forge<br>\nlibstdcxx-ng 9.3.0 hd4cf53a_17<br>\nlibuuid 2.32.1 h7f98852_1000 conda-forge<br>\nlibvorbis 1.3.7 h9c3ff4c_0 conda-forge<br>\nlibxcb 1.13 h7f98852_1003 conda-forge<br>\nlibxkbcommon 1.0.3 he3ba5ed_0 conda-forge<br>\nlibxml2 2.9.12 h72842e0_0 conda-forge<br>\nllvmlite 0.36.0 pypi_0 pypi<br>\nlmdb 1.2.1 pypi_0 pypi<br>\nlz4-c 1.9.3 h9c3ff4c_0 conda-forge<br>\nmarkupsafe 2.0.1 py38h497a2fe_0 conda-forge<br>\nmatplotlib 3.4.2 pypi_0 pypi<br>\nmccabe 0.6.1 pypi_0 pypi<br>\nmistune 0.8.4 py38h497a2fe_1003 conda-forge<br>\nmmcv-full 1.3.7 dev_0<br>\nmmdet 2.11.0 pypi_0 pypi<br>\nmmocr 0.2.0 dev_0<br>\nmmpycocotools 12.0.3 pypi_0 pypi<br>\nmypy_extensions 0.4.3 py38h578d9bd_3 conda-forge<br>\nmysql-common 8.0.25 ha770c72_2 conda-forge<br>\nmysql-libs 8.0.25 hfa10184_2 conda-forge<br>\nnbclient 0.5.3 pyhd8ed1ab_0 conda-forge<br>\nnbconvert 6.0.7 pypi_0 pypi<br>\nnbformat 5.1.3 pyhd8ed1ab_0 conda-forge<br>\nncurses 6.2 he6710b0_1<br>\nnest-asyncio 1.5.1 pyhd8ed1ab_0 conda-forge<br>\nnetworkx 2.5.1 pypi_0 pypi<br>\nnnaudio 0.2.5 pypi_0 pypi<br>\nnodeenv 1.6.0 pypi_0 pypi<br>\nnspr 4.30 h9c3ff4c_0 conda-forge<br>\nnss 3.64 hb5efdd6_0 conda-forge<br>\nnumba 0.53.1 pypi_0 pypi<br>\nnumpy 1.19.5 pypi_0 pypi<br>\nnumpydoc 1.1.0 py_1 conda-forge<br>\noauthlib 3.1.1 pypi_0 pypi<br>\nopencv-python 4.5.2.54 pypi_0 pypi<br>\nopencv-python-headless 4.5.3.56 pypi_0 pypi<br>\nopenssl 1.1.1k h7f98852_0 conda-forge<br>\nordered-set 4.0.2 pypi_0 pypi<br>\npackaging 20.9 pyh44b312d_0 conda-forge<br>\npandas 1.3.0 pypi_0 pypi<br>\npandoc 2.14.0.3 h7f98852_0 conda-forge<br>\npandocfilters 1.4.3 pypi_0 pypi<br>\nparso 0.7.0 pyh9f0ad1d_0 conda-forge<br>\npathspec 0.8.1 pyhd3deb0d_0 conda-forge<br>\npathtools 0.1.2 pypi_0 pypi<br>\npcre 8.45 h9c3ff4c_0 conda-forge<br>\npexpect 4.8.0 pyhd3eb1b0_3<br>\npickleshare 0.7.5 pyhd3eb1b0_1003<br>\npillow 8.2.0 pypi_0 pypi<br>\npip 21.1.2 py38h06a4308_0<br>\npluggy 0.13.1 py38h578d9bd_4 conda-forge<br>\npolygon3 3.0.9.1 pypi_0 pypi<br>\npre-commit 2.13.0 pypi_0 pypi<br>\npromise 2.3 pypi_0 pypi<br>\nprompt-toolkit 3.0.18 pypi_0 pypi<br>\nprotobuf 3.17.3 pypi_0 pypi<br>\npsutil 5.8.0 py38h497a2fe_1 conda-forge<br>\npthread-stubs 0.4 h36c2ea0_1001 conda-forge<br>\nptyprocess 0.7.0 pyhd3eb1b0_2<br>\npy 1.10.0 pypi_0 pypi<br>\npyclipper 1.2.1 pypi_0 pypi<br>\npycocotools 2.0.2 pypi_0 pypi<br>\npycodestyle 2.7.0 pypi_0 pypi<br>\npycparser 2.20 pyh9f0ad1d_2 conda-forge<br>\npydocstyle 6.1.1 pyhd8ed1ab_0 conda-forge<br>\npyflakes 2.3.1 pypi_0 pypi<br>\npygments 2.9.0 pyhd3eb1b0_0<br>\npylint 2.8.2 pyhd8ed1ab_0 conda-forge<br>\npyls-black 0.4.6 pyh9f0ad1d_0 conda-forge<br>\npyls-spyder 0.3.2 pyhd8ed1ab_0 conda-forge<br>\npyopenssl 20.0.1 pyhd8ed1ab_0 conda-forge<br>\npyparsing 2.4.7 pyh9f0ad1d_0 conda-forge<br>\npyqt 5.12.3 py38h578d9bd_7 conda-forge<br>\npyqt-impl 5.12.3 py38h7400c14_7 conda-forge<br>\npyqt5-sip 4.19.18 py38h709712a_7 conda-forge<br>\npyqtchart 5.12 py38h7400c14_7 conda-forge<br>\npyqtwebengine 5.12.1 py38h7400c14_7 conda-forge<br>\npyrsistent 0.17.3 py38h497a2fe_2 conda-forge<br>\npysocks 1.7.1 py38h578d9bd_3 conda-forge<br>\npytest 6.2.4 pypi_0 pypi<br>\npytest-cov 2.12.1 pypi_0 pypi<br>\npytest-runner 5.3.1 pypi_0 pypi<br>\npython 3.8.5 h7579374_1<br>\npython-dateutil 2.8.1 py_0 conda-forge<br>\npython-jsonrpc-server 0.4.0 pyh9f0ad1d_0 conda-forge<br>\npython-language-server 0.36.2 pyhd8ed1ab_0 conda-forge<br>\npython-slugify 5.0.2 pypi_0 pypi<br>\npython_abi 3.8 2_cp38 conda-forge<br>\npytz 2021.1 pyhd8ed1ab_0 conda-forge<br>\npywavelets 1.1.1 pypi_0 pypi<br>\npyxdg 0.27 pyhd8ed1ab_0 conda-forge<br>\npyyaml 5.4.1 pypi_0 pypi<br>\npyzmq 22.1.0 py38h2035c66_0 conda-forge<br>\nqdarkstyle 2.8.1 pyhd8ed1ab_2 conda-forge<br>\nqt 5.12.9 hda022c4_4 conda-forge<br>\nqtawesome 1.0.3 pyhd8ed1ab_0 conda-forge<br>\nqtconsole 5.1.0 pyhd8ed1ab_0 conda-forge<br>\nqtpy 1.9.0 py_0 conda-forge<br>\nqudida 0.0.4 pypi_0 pypi<br>\nrapidfuzz 1.4.1 pypi_0 pypi<br>\nreadline 8.1 h27cfd23_0<br>\nregex 2021.4.4 py38h497a2fe_0 conda-forge<br>\nrequests 2.25.1 pyhd3deb0d_0 conda-forge<br>\nrope 0.19.0 pyhd8ed1ab_0 conda-forge<br>\nrtree 0.9.7 py38h02d302b_1 conda-forge<br>\nscikit-image 0.18.1 pypi_0 pypi<br>\nscipy 1.6.3 pypi_0 pypi<br>\nseaborn 0.11.2 pypi_0 pypi<br>\nsecretstorage 3.3.1 py38h578d9bd_0 conda-forge<br>\nsend2trash 1.5.0 pypi_0 pypi<br>\nsentry-sdk 1.3.1 pypi_0 pypi<br>\nsetuptools 52.0.0 py38h06a4308_0<br>\nshapely 1.7.1 pypi_0 pypi<br>\nshortuuid 1.0.1 pypi_0 pypi<br>\nsix 1.16.0 pyh6c4a22f_0 conda-forge<br>\nsmmap 4.0.0 pypi_0 pypi<br>\nsnowballstemmer 2.1.0 pyhd8ed1ab_0 conda-forge<br>\nsortedcontainers 2.4.0 pyhd8ed1ab_0 conda-forge<br>\nsphinx 4.0.2 pyh6c4a22f_1 conda-forge<br>\nsphinxcontrib-applehelp 1.0.2 py_0 conda-forge<br>\nsphinxcontrib-devhelp 1.0.2 py_0 conda-forge<br>\nsphinxcontrib-htmlhelp 2.0.0 pyhd8ed1ab_0 conda-forge<br>\nsphinxcontrib-jsmath 1.0.1 py_0 conda-forge<br>\nsphinxcontrib-qthelp 1.0.3 py_0 conda-forge<br>\nsphinxcontrib-serializinghtml 1.1.5 pyhd8ed1ab_0 conda-forge<br>\nspyder 4.2.5 py38h578d9bd_0 conda-forge<br>\nspyder-kernels 1.10.2 py38h578d9bd_0 conda-forge<br>\nsqlite 3.35.4 hdfb4753_0<br>\nsubprocess32 3.5.4 pypi_0 pypi<br>\ntensorboard 2.6.0 pypi_0 pypi<br>\ntensorflow-datasets 4.4.0 pypi_0 pypi<br>\ntensorflow-gpu 2.4.0 pypi_0 pypi<br>\ntensorflow-metadata 1.2.0 pypi_0 pypi<br>\nterminado 0.10.0 pypi_0 pypi<br>\nterminaltables 3.1.0 pypi_0 pypi<br>\ntestpath 0.5.0 pyhd8ed1ab_0 conda-forge<br>\ntext-unidecode 1.3 pypi_0 pypi<br>\ntextdistance 4.2.1 pyhd8ed1ab_0 conda-forge<br>\nthree-merge 0.1.1 pyh9f0ad1d_0 conda-forge<br>\ntifffile 2021.6.6 pypi_0 pypi<br>\ntimm 0.4.13 pypi_0 pypi<br>\ntk 8.6.10 hbc83047_0<br>\ntoml 0.10.2 pyhd8ed1ab_0 conda-forge<br>\ntorch 1.10.0.dev20210623+cu111 pypi_0 pypi<br>\ntorchaudio 0.8.1 pypi_0 pypi<br>\ntorchvision 0.11.0.dev20210623+cu111 pypi_0 pypi<br>\ntornado 6.1 py38h497a2fe_1 conda-forge<br>\ntraitlets 5.0.5 pyhd3eb1b0_0<br>\nttach 0.0.3 pypi_0 pypi<br>\ntyped-ast 1.4.3 py38h497a2fe_0 conda-forge<br>\ntyping_extensions 3.10.0.0 pyha770c72_0 conda-forge<br>\nubelt 0.9.5 pypi_0 pypi<br>\nujson 4.0.2 py38h709712a_0 conda-forge<br>\nurllib3 1.26.5 pyhd8ed1ab_0 conda-forge<br>\nvirtualenv 20.4.7 pypi_0 pypi<br>\nwandb 0.12.1 pypi_0 pypi<br>\nwatchdog 1.0.2 py38h578d9bd_1 conda-forge<br>\nwcwidth 0.2.5 py_0<br>\nwebencodings 0.5.1 pypi_0 pypi<br>\nwheel 0.36.2 pyhd3eb1b0_0<br>\nwidgetsnbextension 3.5.1 pypi_0 pypi<br>\nwrapt 1.12.1 py38h497a2fe_3 conda-forge<br>\nwurlitzer 2.1.0 py38h578d9bd_0 conda-forge<br>\nxdoctest 0.15.4 pypi_0 pypi<br>\nxorg-libxau 1.0.9 h7f98852_0 conda-forge<br>\nxorg-libxdmcp 1.1.3 h7f98852_0 conda-forge<br>\nxz 5.2.5 h7b6447c_0<br>\nyaml 0.2.5 h516909a_0 conda-forge<br>\nyapf 0.31.0 pyhd8ed1ab_0 conda-forge<br>\nzeromq 4.3.4 h9c3ff4c_0 conda-forge<br>\nzipp 3.4.1 pyhd8ed1ab_0 conda-forge<br>\nzlib 1.2.11 h7b6447c_3<br>\nzstd 1.5.0 ha95c52a_0 conda-forge</p>\n<p>and using this i trained deeper models like robustscanner,nrtr,fcenet etc<br>\nin another pytorch environment where i have pytorch 1.8 installed i trained many powerful models like xlm roberta,deberta large etc by utilizing 24gb vram and i didn't face this freezing issue<br>\nso if it is hardware problem then i should experience this freezing issue in every model training(at least in most model training),no?<br>\nfor example,,if i do not convert the wave signal into 512x512 image size for training,,then i do not face this freezing issue!</p>\n<p>most of the times i face this freezing issue while trying nfnets,i think i am probably missing a line of code which could fix this freezing issue,i wonder if it is necessary to use torch.cuda.empty_cache() and gc.collect() always for getting rid of this freezing issue or something else that is causing the problem! 😪</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505790,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/07/2021 14:50:41",
          "content": "<p>Are you using AMD cpus and ubuntu?<br>\nIs the freeze happening during an epoch or at start/end?<br>\nDid you track the temperature?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505903,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/07/2021 15:55:11",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> </p>\n<ol>\n<li>intel cpu core i9 with ubuntu</li>\n<li>during an epoch as you can see from the log at the bottom here : <a href=\"https://pastebin.com/XEusHppZ\" target=\"_blank\">https://pastebin.com/XEusHppZ</a></li>\n<li>when the model was training i was not in front of the computer and i don't know if i can monitor and save temperate overtime when i am not in-front of the computer,sorry for this</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505939,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/07/2021 16:41:46",
          "content": "<p>If you use wandb, you can track some of your hardware's indicators. Not sure how reliable this is but it is better than nothing I guess?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1506400,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "09/08/2021 07:20:20",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> how many Watts is your PSU ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1506495,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/08/2021 09:25:43",
          "content": "<p>Is wandb not tracking the temperature? Whats your temperature after you trained it for lets say 5-10 minutes? To me it sounds very much like overheating, this is typical for freezes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1506573,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "09/08/2021 11:08:16",
          "content": "<p>Maybe its a issue with the memory junction temps. GDDR6X gets really hot. That might be a problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507130,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/08/2021 21:16:58",
          "content": "<p>I was not using wandb.soon i will use it and will let you know.thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507131,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/08/2021 21:18:33",
          "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> using this one : <a href=\"https://www.coolermaster.com/catalog/power-supplies/mwe-series/mwe-gold-850-v2-full-modular/\" target=\"_blank\">https://www.coolermaster.com/catalog/power-supplies/mwe-series/mwe-gold-850-v2-full-modular/</a></p>\n<p>And this cooler : <a href=\"https://www.coolermaster.com/catalog/coolers/cpu-air-coolers/hyper-212-led-turbo/\" target=\"_blank\">https://www.coolermaster.com/catalog/coolers/cpu-air-coolers/hyper-212-led-turbo/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507255,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "09/09/2021 03:00:27",
          "content": "<p>Does your case have enough airflow ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1509885,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/11/2021 18:46:19",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> any update on your situation? 👀</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1510154,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/12/2021 05:56:25",
          "content": "<p><a href=\"https://www.kaggle.com/yassinealtuned\" target=\"_blank\">@yassinealtuned</a><br>\nCurrently i am using kaggle tpu for this competition… </p>\n<p>sorry i am yet to train anything again with rtx3090.<br>\nProblem is, i am using a server pc through remote desktop connection which has rtx3090.that server pc is hosted in my office which is far away from me and i am working from home!<br>\nI use that pc via remote desktop connection <br>\nSo It's difficult for me to restart the pc over and over again for this painful freezing issue &gt;_&lt;<br>\nWhen it freezes,i lose connection and can't Connect to pc via remote desktop connection and then i disturb office colleagues/staff for restarting the pc(who is there in the office at that Moment)<br>\nLast Thursday the pc got disconnected (maybe because of either shutdown/freeze issue)<br>\nThen here in my territory Friday and Saturday’s are off day,so i found no office staff to restart the pc<br>\nToday is sunday,so one office staff just turned on the pc again and now during daytime my office colleagues are using that pc for generating synthetic data and train models related to optical character recognition (ocr) for daytime office Project. <br>\nThey(my office team mates) do not compete in kaggle<br>\nI plan to try wandb again within next 2-3 days when the pc is not busy for training model for official works! <br>\nI Just hate this freeze issue,restarting the pc for me is very difficult task as i work for home :(<br>\nStay tuned,thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513397,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/15/2021 06:30:47",
          "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> this is the first time i used wandb,sorry if i am missing something.<br>\nhere is the wandb folder where i have log : <a href=\"https://www.kaggle.com/mobassir/wandb-log\" target=\"_blank\">https://www.kaggle.com/mobassir/wandb-log</a><br>\ni can't see any temperature related log,don't know how to interpret,let me know if you need the code that i tried to generate these logs,thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513451,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/15/2021 07:23:39",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> do you have access to the online dashboard? This is the easiest way to get some system information. I will have a look at your logs and let you know later. 👌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513498,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/15/2021 07:56:54",
          "content": "<p>sorry <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> <br>\ni don't have any online dashboard,i am simply using wandb like this kernel : <a href=\"https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training</a></p>\n<p>i just modified his code a little bit for locally training,,i use like this : </p>\n<p>import wandb<br>\nwandb.login()</p>\n<p>and i do not use these : </p>\n<h1>from kaggle_secrets import UserSecretsClient</h1>\n<h1>user_secrets = UserSecretsClient()</h1>\n<h1>wandb_api = user_secrets.get_secret(\"wandb_api\")</h1>\n<h1>wandb.login(key=wandb_api)</h1>\n<p>and then he was doing this : </p>\n<p>run = wandb.init(project=\"G2Net-Public-experiments\", <br>\n                 name=\"exp1\",<br>\n                 config=class2dict(CFG),<br>\n                 group=CFG.model_name,<br>\n                 job_type=\"train\")</p>\n<p>where i did this : </p>\n<p>run = wandb.init(project=\"G2Net-Public-experiments\", <br>\n                 name=\"exp1\",<br>\n                 config=class2dict(CFG),<br>\n                 group=CFG.model_name,<br>\n                 job_type=\"train\",<br>\n                 settings=wandb.Settings(_disable_stats=False))</p>\n<p>everything else is same for wandb logging like him</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513499,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "09/15/2021 07:58:32",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> Are there any other GPUs in that machine? The reason I ask is that when the 2080Ti came out, I wanted to use it alongside my old GTX 1080. But during training it would randomly freeze even if the 1080 was idle and nothing was showing on any logs (therefore a catastrophic kernel issue). The only way I could fix it was by removing the 1080 from the system, and I think the crashing was probably caused by some driver issue caused by mixing different generation cards.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513515,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/15/2021 08:10:26",
          "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> i am using only rtx3090(single gpu)</p>\n<p>Processor    : i9 10th Gen 10Core<br>\nRAM          : 32GB (3200Mhz)<br>\nHDD          : 1TB SSD Nvme<br>\nGPU          : MSI Nvidia RTX 3090 24GB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513726,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/15/2021 11:07:29",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I only used wandb once so far, and outside Kaggle, but it seems you need to create a wandb account and use your credentials.  I have not user kaggle secret either but it looks like a way to publicly share notebooks without revealing your credentials.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1514121,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/15/2021 17:55:19",
          "content": "<p>dear <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> here you can see my system metric(screenshot) : <a href=\"https://www.kaggle.com/mobassir/wandb-log?select=1.PNG\" target=\"_blank\">https://www.kaggle.com/mobassir/wandb-log?select=1.PNG</a><br>\ncheck 1.png,2.png,3.png and 4.png<br>\nthank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1517967,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/20/2021 08:51:50",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> it looks like the graphs are fine except the first dip we see. I don't have other options to offer for now. 😬</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1520899,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/22/2021 17:22:20",
          "content": "<p>Something that I've now noticed. It's always the model being trained on gpu:0 that crashes. I'm training on a desktop, so I wonder if it has something to do with X / the display system interfering. Once this competition is done, I plan on swapping the physical positions of the two cards to see if the data loader thread crashing persists on the same GPU device or on the same slot to better debug.</p>\n<p>Also, does anyone see:</p>\n<p><code>[W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool)</code> being spammed to the console where jupyter notebook is invoked?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1520942,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "09/22/2021 17:59:41",
          "content": "<blockquote>\n  <p>Also, does anyone see:</p>\n  <p>[W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool) being spammed to the console where jupyter notebook is invoked?</p>\n</blockquote>\n<p>Yes, that's a very annoying bug in pytorch (which looks like it might be fixed in 1.10) <a href=\"https://github.com/pytorch/pytorch/issues/57273\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/57273</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1521032,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/22/2021 20:04:11",
          "content": "<p>Wow, I took out <code>pin_memory=True</code> and behold, all my crashing issues have been resolved.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1521047,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/22/2021 20:20:05",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Yes, I had a similar warning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1504766,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "09/06/2021 16:30:21",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I am using an image of size <strong>256 x 256</strong> and here are some utilization metrics over time. So not optimal for now, I will try to debug a bit and optimize tonight and update with new findings. </p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/16PVPs6/gpu-vram-and-utilization.png\" alt=\"gpu-vram-and-utilization\"></a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1504890,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/06/2021 18:39:57",
          "content": "<p>Here is my use (V100 in NVIDIA cluster, first generation at 163W):</p>\n<p><img src=\"https://i.imgur.com/eljWSbK.png\" alt=\"gpu use\"></p>\n<p>I actually use almost all memory, for some reason memory use is divided by 2 in the report.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504928,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/06/2021 19:03:17",
          "content": "<p>Thanks for the plot. 👌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1506099,
      "author_name": "codereaper",
      "author_url": "",
      "post_date": "09/07/2021 20:55:00",
      "content": "<p>and How Much it takes  on non RTX ones like  GTX 1660ti </p>",
      "votes": null,
      "replies": [
        {
          "id": 1506321,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/08/2021 06:02:23",
          "content": "<p>Probably forever. 😄</p>\n<p>Joke aside, try to run on small images (128 x 128) and a very small model (EfficientNet B0 maybe or even something else). Also try to add gradient accumulation to avoid memory issues.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1506470,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "09/08/2021 08:55:21",
          "content": "<p>between this two it is not much about RT part, but tensor cores. 1660 has none. But again 1080ti for example is still a solid performer even without them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1506106,
      "author_name": "benjamin35",
      "author_url": "",
      "post_date": "09/07/2021 21:18:16",
      "content": "<p>I work on waveform only, with tfrecords and 8 TPU, I do one epoch into 27 seconds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1509884,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/11/2021 18:44:59",
          "content": "<p>That's quite impressive!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1506250,
      "author_name": "haplophyrne",
      "author_url": "",
      "post_date": "09/08/2021 04:20:04",
      "content": "<p>I also have a rtx  3090, I average 20 min per epoch with preprocessed data and ~120-180 min when processing data on the fly. Pytorch</p>",
      "votes": null,
      "replies": [
        {
          "id": 1506318,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/08/2021 05:58:54",
          "content": "<p>Interesting, so you get a 6 or 8 factor from saving the images. I am planning to do it as well, will let you know how it goes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1506333,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "09/08/2021 06:19:03",
          "content": "<p>You can also do it on the fly on the GPU with nnAudio</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1509883,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/11/2021 18:44:40",
          "content": "<p>I have preprocessed and saved the data but without any improvements so far. I guess either the bottleneck is in the IO operations (i.e. opening files) since the processed images are on a \"classical\" HDD (and not SSD) or I am doing something wrong? </p>\n<p>Notice that I am using the EfficienNet B0 model and 512 x 512 images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1509902,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/11/2021 19:08:48",
          "content": "<p>how much RAM do you have?</p>\n<p>It maybe better to upload all waves in ram and compute images on the fly using packages that run on GPU like torch.fft, nnaudio, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1509917,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/11/2021 19:26:08",
          "content": "<p>RAM of 64GB. </p>\n<p>Yes, if my SSD test doesn't give better results, I will for sure move to this option. I am also checking on a cloud GPU to see if the problem comes from my current setup. </p>\n<p>Will let you know.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1509946,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/11/2021 20:00:42",
          "content": "<p>SSD should definitely help. Move all your training data to it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1509976,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "09/11/2021 21:12:15",
      "content": "<p>Update: I have resized the images to 256 x 256 and still using EfficientNet B0 and now can train one epoch for one fold in 1<strong>2 minutes</strong> (GPU usage is almost at max). </p>\n<p>The secret was in getting a much bigger batch size: around 300.</p>\n<p>I can probably optimize even further but at least now I can run more experiments. 👌</p>",
      "votes": null,
      "replies": [
        {
          "id": 1510947,
          "author_name": "haplophyrne",
          "author_url": "",
          "post_date": "09/12/2021 23:58:41",
          "content": "<p>I switched to nn Audio and I also get similar results but with a smaller batchsize</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1517966,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "09/20/2021 08:50:19",
      "content": "<p>For those still following this thread, the fastest way I have found so far is using <strong>TFRecords</strong> and <strong>TPUs</strong>. I will share how long it takes once the competition is over. 👌</p>",
      "votes": null,
      "replies": [
        {
          "id": 1524715,
          "author_name": "allunia",
          "author_url": "",
          "post_date": "09/26/2021 19:01:33",
          "content": "<p>Great! :-) Thanks to your discussion post I started to learn much more about how to optimise my GPU and TPU performances and it still goes on…  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1524777,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/26/2021 20:38:00",
          "content": "<p>That's awesome! I have learned many things as well, might share them once the competition ends. 👌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559736,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:10:28",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1504365": "So far, for one epoch and one single model on a 5 folds split, it takes around 1 hour to run on an\n**RTX 3090**. \n\nSo around **5 hours** for one full (all splits) **epoch**.\n\nI was wondering if these times are closer to what you have or if I can further optimize the code. Thanks for your help!",
    "1504370": "that's a long time - I also have a 3090 and can train an epoch of b0 in less than 3mins. My best model is a b2 and it's current setup takes around 12mins an epoch\n\nmy guess is you're either using a huge model or huge image or you have some bottleneck?",
    "1504372": "I assume you're doing all the usual mixed precision/pin_memory/etc optimisations?",
    "1504400": "Not that big only a B3. Also, I am generating image features on the fly so probably not optimal for now. Thanks for providing these numbers, I will check what can be improved. 👌",
    "1504405": "hamishdickson @yassinealouini i am also experimenting using rtx3090 but when i converted wave signal into image and resized it down to 512x512 it takes around 1 hour to train a b0 model for 1 epoch and then my pc shuts down,i can't figure out why pc getting turned off during training,,,when i was using wave data only,,at that time training was fast and pc never got shut down,,but after converting it to 512x512 image size,pc shutting down after few hours of training(sometimes less than 1 hour)\ni was working using this kernel : https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\nduring training i checked nvidia-smi and i was using less than 20gb vram out of 24",
    "1504415": "> and then my pc shuts down\n\nUndervolt your GPU and check if the problem goes away. I had the same problem when my old PSU gradually died - big power spikes during initialization forced PSU to shut down.\n\n3090, depending on model, can generate power spikes up to ~600W. Undervolting with 5-10% performance loss might keep them to ~350-400W max.",
    "1504429": "That's very interesting. I will check my wandb dashboard but I don't have those spikes I think. Thanks for sharing @fffrrt. 👌",
    "1504434": "fffrrt thank you for sharing\ni saw that my vram usage was low(less than 20gb from 24gb),so i thought cpu or ram or something else is causing this issue but i can't figure out,i checked kern log and found no problem,i thought it's num_workers that is generating more heat and causing this issue so set num_workers = 0 but no luck,i have gold certified good power supply,good cpu cooler but still for some reason my pc getting shut down,i hope your idea of Undervolt works for me,,otherwise i don't know how to solve this issue,thank you",
    "1504438": "> I will check my wandb dashboard but I don't have those spikes I think.\n\nThey are not visible on monitoring software, because they are too short. Long enough for some PSUs to trip overcurrent protection, though.\n\nI could not find original igorslab article about them (when 3080/3090 on release crashed many PCs), only his testing after the drivers fix - https://www.igorslab.de/en/wonder-how-invidia-the-crashes-of-the-force-rtx-3080-andrtx-3090-will-be-removed-and-still-will-be-removed-even-from-the-power-supplies-analysis/\n\nIllustration from there (3080, should be more or less actual behavior): https://www.igorslab.de/wp-content/uploads/2020/09/12a-Gaming-Zoom-Power-New-Drivers.png\n\nEdit: Found this topics that might prove helpful, i read them during my own issues - \n\nhttps://discuss.pytorch.org/t/multi-gpu-2080-ti-training-crashes-pc/52032/9\n\nhttps://github.com/pytorch/pytorch/issues/3022",
    "1504480": "So I'm not alone. I had been having 3090 issues as well for the past few months. In my case, the computer does not freeze but training with segfault and I cannot reclaim the vram until I do a restart. Also it's usually GPU specific, that is, I can still train on the other GPU. Though occasionally when one blows, it takes both down. I have 1600W PSU though so I even if both pull spike their load, I wouldn't expect that to take down the ship. It is so annoying with these longer model train times, you let the models sail all night long only to wake up and discover stuff froze after the first fold.",
    "1504520": "**you let the models sail all night long only to wake up and discover stuff froze after the first fold.** 😪",
    "1504580": "yassinealouini image size is also a factor.  Depending on the image size your timing may just be good, or bad.  One way to check is to look at the memory and compute use of your GPU using nvidia-smi.  \n\nFor instance, if use is close to 100% and memory close to GPU memory then you are probably fine.  \n\nAnother example: if use is 100% but memory is only half of GPU memory then you can probably double batch size and divide time by 2.  Of course, changing batch size may impact other hyper parameter values.",
    "1504583": "Not to hijack the thread and perhaps too close to competition deadline, but @cpmpml if I might inquire, what bs are you currently utilizing in your model? It was something I'd actually been wondering. I'm currently using a modest 32 bs resulting in 98% utilization, however +85% vram is free... so I wonder.",
    "1504588": "mobassir It looks like you should upgrade to the latest driver, and also check pour power unit.",
    "1504595": "Thanks @cpmpml for the details. I will check these and report back later. 👌",
    "1504597": "cpmp previously i had 370watt limit for gpu and now after undervolting 70watt i have this : \nMon Sep  6 19:12:16 2021       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 460.91.03    Driver Version: 460.91.03    CUDA Version: 11.2     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce RTX 3090    Off  | 00000000:01:00.0 Off |                  N/A |\n|  0%   28C    P8    12W / 300W |     19MiB / 24268MiB |      0%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n|    0   N/A  N/A      1140      G   /usr/lib/xorg/Xorg                  9MiB |\n|    0   N/A  N/A      1319      G   /usr/bin/gnome-shell                8MiB |\n+-----------------------------------------------------------------------------+\n\n**nvcc --version**\n\nnvcc: NVIDIA (R) Cuda compiler driver\nCopyright (c) 2005-2021 NVIDIA Corporation\nBuilt on Sun_Feb_14_21:12:58_PST_2021\nCuda compilation tools, release 11.2, V11.2.152\nBuild cuda_11.2.r11.2/compiler.29618528_0\n\ncan you suggest me which driver version i should try now? thank you",
    "1504670": "> however +85% vram is free… so I wonder.\n\nYou should use  a larger batch size IMHO.  It is not what cause your problem of course, but you are under using your GPU it seems",
    "1504693": "fffrrt 😪\npreviously i had 370watt limit for gpu and then i undervolted 70watt and limited gpu to use only 300watt.\ni reduced image size from 512 to 256\ni increased batch size (but vram usage was still less than 19gb out of 24)\ni used mixed precision this time\n**PC SHUTS DOWN AGAIN AFTER 2nd EPOCH**  😭😪",
    "1504721": "Are you caching tf valid dataset? If so try disabling it. Also try empty cache after every epoch. I haven't encountered this issue so I am just throwing ideas.",
    "1504725": "mobassir could it be a hardware defect in the GPU? Have you tried contacting NVIDIA support?",
    "1504766": "cpmpml I am using an image of size **256 x 256** and here are some utilization metrics over time. So not optimal for now, I will try to debug a bit and optimize tonight and update with new findings. \n\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/16PVPs6/gpu-vram-and-utilization.png\" alt=\"gpu-vram-and-utilization\" border=\"0\"></a>",
    "1504774": "pheadrus i am not caching tf valid dataset,,i tried another baseline and encountered same issue,,for example if i experiment using this notebook : https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\nreplace b7 with b0\nin custom model class use F.interpolate to convert wave data into 512x512 image size\nadjust batch size that takes just less than 20gb vram and train **(KEEP EVERYTHING AS IT IS IN THAT PUBLIC KERNEL)**\n\nstill my pc turns off **(this time after 5 epoch)**\n\n@yassinealouini i Have not tried contacting NVIDIA support yet,when i do not convert wave data into 512x512 image data i don't face this issue,,i tried to train few nlp models and advance scene text detection/recognition models and didn't face this issue\ni face this issue occasionally and i can't understand how to solve this issue\n\ninstead of shutdown if i could use a command/setting that will restart the pc instead of freezing/turning it off then it could be good for me but i don't know if it's possible to do so",
    "1504810": "obviously you monitoring gpu temperatures? nvidia gpus have shutdown thresholds at ~90c.   it should not lead to pc shutdown thou",
    "1504834": "> try empty cache after every epoch\n\nI used to do this when I had dual 2080's and it would balloon train times like crazzzyyyyyyyyyyy. Now, I just let the thing take care of the thing by itself.",
    "1504858": "mobassir I see, thanks for the clarification. So it seems to be a specific issue to this specific model setup. Have you tried very small image size (let's say 64 x 64) and/or removing the EfficientNet and replacing it by the identity or another model? \n\nNot sure how to help but if I were in your situation, I would try various combinations until I find something. Best of luck!",
    "1504890": "Here is my use (V100 in NVIDIA cluster, first generation at 163W):\n\n![gpu use](https://i.imgur.com/eljWSbK.png)\n\nI actually use almost all memory, for some reason memory use is divided by 2 in the report.",
    "1504928": "Thanks for the plot. 👌",
    "1505145": "mobassir In case of 3090 it also might be caused by high memory junction temperatures. I think they are only accessible to monitoring software under Windows, for example in HWinfo64. GDDR6X runs very hot, and sometimes reaches thermal throttle at 110 degrees.\n\nIf this is the case (you can quickly run any Eth mining software on Windows as a way to stress test VRAM), then changing thermal pads usually improves the situation, but that involves disassembly of the cooling system and might complicate warranty. Another way could be just propping up your 3090 in the case - https://www.reddit.com/r/nvidia/comments/pihwzo/3090_vram_temperature_fix_without_replacing_pads/\n\nAnother relatively common mistake is connecting GPU with one cable from PSU > two 8-pin connectors on GPU. Always use separate cables for separate 8-pin connectors, i think it is even in installation manual.\n\nHowever, if VRAM junction temperature is not above 100C, undervolting does not help with your shutdowns, and card is connected properly, and you are still having shutdowns - then it might be a general hardware instability, which is a nightmare to debug. I would still check with a ridiculously overpowered PSU just in case, but at that point i would probably resort to changing components one by one to find the root cause.",
    "1505296": "fffrrt @cpmpml @yassinealouini @authman @bakeryproducts @pheadrus \n\napology for a little mistake from my end,\n\n**The pc is not getting turned off,the problem is it is getting frozen every time and forcing me to restart everytime**\n\nfor reproducing the result i just took the public kernel of @yasufuminakama san and modified a little bit to train and face the freeze issue so that i can share the code and log with you guys for better investigation..\n\nhere is the code and log file : https://pastebin.com/XEusHppZ\nyou can see from my log that during 7th epoch  i faced 'freeze' issue and i had to restart the pc.\ni was using around 19gb vram out of 24 iirc while training this model (full code is provided).\n\ncould this be related to pytorch version or something else?\n\nit will be highly appreciated if any kaggler help me solve this issue (i've shared the full code),thank you a lot in advance!",
    "1505311": "One \"cheap\" solution is to stop after 3 or 4 epochs maybe? \n\nOtherwise, maybe try to reinstall everything? I've taken few notes here: https://gist.github.com/yassineAlouini/e843531219e1b9d978aa4fe2a5e71940. There is also this good blog post: https://www.itzgeek.com/post/how-to-install-nvidia-drivers-on-ubuntu-20-04-ubuntu-18-04.html",
    "1505379": "I had same issues as yours. I changed the PSU but the problem was still persistent. Later, I changed the motherboard boom it started working again, No freeze or random shutdown. Might be problem due to current spike causing damage to weaker motherboard capacitors/components.",
    "1505606": "mobassir what was the GPU temperature when it froze?  \n\nAlso, as @karmaz indicated, it maybe a bad motherboard.  However, before changing it, have you tried to remove the GPU and disconnect everything, then connect and plug it back?  Sometimes a defective connection is the cause of such issues.",
    "1505714": "cpmpml i am  not sure if it is really a hardware related issue,,i thought this issue is related to pytorch version maybe,i am using the following packages : \nName Version Build Channel\n_libgcc_mutex 0.1 main\n_openmp_mutex 4.5 1_gnu\naddict 2.4.0 pypi_0 pypi\nalabaster 0.7.12 py_0 conda-forge\nalbumentations 1.0.3 pypi_0 pypi\nalsa-lib 1.2.3 h516909a_0 conda-forge\nappdirs 1.4.4 pyh9f0ad1d_0 conda-forge\nargh 0.26.2 pyh9f0ad1d_1002 conda-forge\nastroid 2.6.0 py38h578d9bd_0 conda-forge\nasync_generator 1.10 py_0 conda-forge\nasynctest 0.13.0 pypi_0 pypi\natomicwrites 1.4.0 pyh9f0ad1d_0 conda-forge\nattrs 21.2.0 pyhd8ed1ab_0 conda-forge\nautopep8 1.5.6 pyhd8ed1ab_0 conda-forge\nbabel 2.9.1 pyh44b312d_0 conda-forge\nbackcall 0.2.0 pyhd3eb1b0_0\nblack 21.5b2 pyhd8ed1ab_0 conda-forge\nbleach 3.3.0 pyh44b312d_0 conda-forge\nbrotlipy 0.7.0 py38h497a2fe_1001 conda-forge\nca-certificates 2021.5.30 ha878542_0 conda-forge\ncertifi 2021.5.30 py38h578d9bd_0 conda-forge\ncffi 1.14.5 py38ha65f79e_0 conda-forge\ncfgv 3.3.0 pypi_0 pypi\nchardet 4.0.0 py38h578d9bd_1 conda-forge\nclick 8.0.1 py38h578d9bd_0 conda-forge\ncloudpickle 1.6.0 py_0 conda-forge\ncodecov 2.1.11 pypi_0 pypi\ncolorama 0.4.4 pyh9f0ad1d_0 conda-forge\nconfigparser 5.0.2 pypi_0 pypi\ncoverage 5.5 pypi_0 pypi\ncryptography 3.4.7 py38ha5dfef3_0 conda-forge\ncudatoolkit 11.1.74 h6bb024c_0 nvidia\ncycler 0.10.0 pypi_0 pypi\ncython 0.29.23 pypi_0 pypi\ndataclasses 0.8 pyhc8e2a94_1 conda-forge\ndbus 1.13.6 h48d8840_2 conda-forge\ndecorator 4.4.2 pypi_0 pypi\ndefusedxml 0.7.1 pyhd8ed1ab_0 conda-forge\ndiff-match-patch 20200713 pyh9f0ad1d_0 conda-forge\ndill 0.3.4 pypi_0 pypi\ndistlib 0.3.2 pypi_0 pypi\ndocker-pycreds 0.4.0 pypi_0 pypi\ndocutils 0.17.1 py38h578d9bd_0 conda-forge\nentrypoints 0.3 pyhd8ed1ab_1003 conda-forge\nexpat 2.4.1 h9c3ff4c_0 conda-forge\nfilelock 3.0.12 pypi_0 pypi\nflake8 3.9.2 pypi_0 pypi\nfontconfig 2.13.1 hba837de_1005 conda-forge\nfreetype 2.10.4 h0708190_1 conda-forge\nfuture 0.18.2 py38h578d9bd_3 conda-forge\ngast 0.3.3 pypi_0 pypi\ngettext 0.19.8.1 h0b5b191_1005 conda-forge\ngitdb 4.0.7 pypi_0 pypi\ngitpython 3.1.18 pypi_0 pypi\nglib 2.68.3 h9c3ff4c_0 conda-forge\nglib-tools 2.68.3 h9c3ff4c_0 conda-forge\ngoogleapis-common-protos 1.53.0 pypi_0 pypi\ngrad-cam 1.3.1 pypi_0 pypi\ngrpcio 1.32.0 pypi_0 pypi\ngst-plugins-base 1.18.4 hf529b03_2 conda-forge\ngstreamer 1.18.4 h76c114f_2 conda-forge\nh5py 2.10.0 pypi_0 pypi\nhelpdev 0.7.1 pyhd8ed1ab_0 conda-forge\nicu 68.1 h58526e2_0 conda-forge\nidentify 2.2.10 pypi_0 pypi\nidna 2.10 pyh9f0ad1d_0 conda-forge\nimageio 2.9.0 pypi_0 pypi\nimagesize 1.2.0 py_0 conda-forge\nimgaug 0.4.0 pypi_0 pypi\nimportlib-metadata 4.5.0 py38h578d9bd_0 conda-forge\nimportlib-resources 5.2.2 pypi_0 pypi\nimportlib_metadata 4.5.0 hd8ed1ab_0 conda-forge\niniconfig 1.1.1 pypi_0 pypi\nintervaltree 3.0.2 py_0 conda-forge\nipykernel 5.5.5 py38hd0cf306_0 conda-forge\nipython 7.18.1 py38h5ca1d4c_0 anaconda\nipython_genutils 0.2.0 pyhd3eb1b0_1\nipywidgets 7.6.3 pypi_0 pypi\nisort 5.8.0 pypi_0 pypi\njedi 0.17.2 py38h578d9bd_1 conda-forge\njeepney 0.6.0 pyhd8ed1ab_0 conda-forge\njinja2 3.0.1 pyhd8ed1ab_0 conda-forge\njpeg 9d h36c2ea0_0 conda-forge\njsonschema 3.2.0 pyhd8ed1ab_3 conda-forge\njupyter_client 6.1.12 pyhd8ed1ab_0 conda-forge\njupyter_core 4.7.1 py38h578d9bd_0 conda-forge\njupyterlab-widgets 1.0.0 pypi_0 pypi\njupyterlab_pygments 0.1.2 pyh9f0ad1d_0 conda-forge\nkaggle 1.5.12 pypi_0 pypi\nkaggledatasets 0.0.1 pypi_0 pypi\nkeyring 23.0.1 py38h578d9bd_0 conda-forge\nkiwisolver 1.3.1 pypi_0 pypi\nkrb5 1.19.1 hcc1bbae_0 conda-forge\nkwarray 0.5.19 pypi_0 pypi\nlanms-proper 1.0.1 pypi_0 pypi\nlazy-object-proxy 1.6.0 py38h497a2fe_0 conda-forge\nld_impl_linux-64 2.35.1 h7274673_9\nlibclang 11.1.0 default_ha53f305_1 conda-forge\nlibedit 3.1.20191231 he28a2e2_2 conda-forge\nlibevent 2.1.10 hcdb4288_3 conda-forge\nlibffi 3.3 he6710b0_2\nlibgcc-ng 9.3.0 h5101ec6_17\nlibglib 2.68.3 h3e27bee_0 conda-forge\nlibgomp 9.3.0 h5101ec6_17\nlibiconv 1.16 h516909a_0 conda-forge\nlibllvm11 11.1.0 hf817b99_2 conda-forge\nlibogg 1.3.4 h7f98852_1 conda-forge\nlibopus 1.3.1 h7f98852_1 conda-forge\nlibpng 1.6.37 h21135ba_2 conda-forge\nlibpq 13.3 hd57d9b9_0 conda-forge\nlibsodium 1.0.18 h36c2ea0_1 conda-forge\nlibspatialindex 1.9.3 h9c3ff4c_3 conda-forge\nlibstdcxx-ng 9.3.0 hd4cf53a_17\nlibuuid 2.32.1 h7f98852_1000 conda-forge\nlibvorbis 1.3.7 h9c3ff4c_0 conda-forge\nlibxcb 1.13 h7f98852_1003 conda-forge\nlibxkbcommon 1.0.3 he3ba5ed_0 conda-forge\nlibxml2 2.9.12 h72842e0_0 conda-forge\nllvmlite 0.36.0 pypi_0 pypi\nlmdb 1.2.1 pypi_0 pypi\nlz4-c 1.9.3 h9c3ff4c_0 conda-forge\nmarkupsafe 2.0.1 py38h497a2fe_0 conda-forge\nmatplotlib 3.4.2 pypi_0 pypi\nmccabe 0.6.1 pypi_0 pypi\nmistune 0.8.4 py38h497a2fe_1003 conda-forge\nmmcv-full 1.3.7 dev_0\nmmdet 2.11.0 pypi_0 pypi\nmmocr 0.2.0 dev_0\nmmpycocotools 12.0.3 pypi_0 pypi\nmypy_extensions 0.4.3 py38h578d9bd_3 conda-forge\nmysql-common 8.0.25 ha770c72_2 conda-forge\nmysql-libs 8.0.25 hfa10184_2 conda-forge\nnbclient 0.5.3 pyhd8ed1ab_0 conda-forge\nnbconvert 6.0.7 pypi_0 pypi\nnbformat 5.1.3 pyhd8ed1ab_0 conda-forge\nncurses 6.2 he6710b0_1\nnest-asyncio 1.5.1 pyhd8ed1ab_0 conda-forge\nnetworkx 2.5.1 pypi_0 pypi\nnnaudio 0.2.5 pypi_0 pypi\nnodeenv 1.6.0 pypi_0 pypi\nnspr 4.30 h9c3ff4c_0 conda-forge\nnss 3.64 hb5efdd6_0 conda-forge\nnumba 0.53.1 pypi_0 pypi\nnumpy 1.19.5 pypi_0 pypi\nnumpydoc 1.1.0 py_1 conda-forge\noauthlib 3.1.1 pypi_0 pypi\nopencv-python 4.5.2.54 pypi_0 pypi\nopencv-python-headless 4.5.3.56 pypi_0 pypi\nopenssl 1.1.1k h7f98852_0 conda-forge\nordered-set 4.0.2 pypi_0 pypi\npackaging 20.9 pyh44b312d_0 conda-forge\npandas 1.3.0 pypi_0 pypi\npandoc 2.14.0.3 h7f98852_0 conda-forge\npandocfilters 1.4.3 pypi_0 pypi\nparso 0.7.0 pyh9f0ad1d_0 conda-forge\npathspec 0.8.1 pyhd3deb0d_0 conda-forge\npathtools 0.1.2 pypi_0 pypi\npcre 8.45 h9c3ff4c_0 conda-forge\npexpect 4.8.0 pyhd3eb1b0_3\npickleshare 0.7.5 pyhd3eb1b0_1003\npillow 8.2.0 pypi_0 pypi\npip 21.1.2 py38h06a4308_0\npluggy 0.13.1 py38h578d9bd_4 conda-forge\npolygon3 3.0.9.1 pypi_0 pypi\npre-commit 2.13.0 pypi_0 pypi\npromise 2.3 pypi_0 pypi\nprompt-toolkit 3.0.18 pypi_0 pypi\nprotobuf 3.17.3 pypi_0 pypi\npsutil 5.8.0 py38h497a2fe_1 conda-forge\npthread-stubs 0.4 h36c2ea0_1001 conda-forge\nptyprocess 0.7.0 pyhd3eb1b0_2\npy 1.10.0 pypi_0 pypi\npyclipper 1.2.1 pypi_0 pypi\npycocotools 2.0.2 pypi_0 pypi\npycodestyle 2.7.0 pypi_0 pypi\npycparser 2.20 pyh9f0ad1d_2 conda-forge\npydocstyle 6.1.1 pyhd8ed1ab_0 conda-forge\npyflakes 2.3.1 pypi_0 pypi\npygments 2.9.0 pyhd3eb1b0_0\npylint 2.8.2 pyhd8ed1ab_0 conda-forge\npyls-black 0.4.6 pyh9f0ad1d_0 conda-forge\npyls-spyder 0.3.2 pyhd8ed1ab_0 conda-forge\npyopenssl 20.0.1 pyhd8ed1ab_0 conda-forge\npyparsing 2.4.7 pyh9f0ad1d_0 conda-forge\npyqt 5.12.3 py38h578d9bd_7 conda-forge\npyqt-impl 5.12.3 py38h7400c14_7 conda-forge\npyqt5-sip 4.19.18 py38h709712a_7 conda-forge\npyqtchart 5.12 py38h7400c14_7 conda-forge\npyqtwebengine 5.12.1 py38h7400c14_7 conda-forge\npyrsistent 0.17.3 py38h497a2fe_2 conda-forge\npysocks 1.7.1 py38h578d9bd_3 conda-forge\npytest 6.2.4 pypi_0 pypi\npytest-cov 2.12.1 pypi_0 pypi\npytest-runner 5.3.1 pypi_0 pypi\npython 3.8.5 h7579374_1\npython-dateutil 2.8.1 py_0 conda-forge\npython-jsonrpc-server 0.4.0 pyh9f0ad1d_0 conda-forge\npython-language-server 0.36.2 pyhd8ed1ab_0 conda-forge\npython-slugify 5.0.2 pypi_0 pypi\npython_abi 3.8 2_cp38 conda-forge\npytz 2021.1 pyhd8ed1ab_0 conda-forge\npywavelets 1.1.1 pypi_0 pypi\npyxdg 0.27 pyhd8ed1ab_0 conda-forge\npyyaml 5.4.1 pypi_0 pypi\npyzmq 22.1.0 py38h2035c66_0 conda-forge\nqdarkstyle 2.8.1 pyhd8ed1ab_2 conda-forge\nqt 5.12.9 hda022c4_4 conda-forge\nqtawesome 1.0.3 pyhd8ed1ab_0 conda-forge\nqtconsole 5.1.0 pyhd8ed1ab_0 conda-forge\nqtpy 1.9.0 py_0 conda-forge\nqudida 0.0.4 pypi_0 pypi\nrapidfuzz 1.4.1 pypi_0 pypi\nreadline 8.1 h27cfd23_0\nregex 2021.4.4 py38h497a2fe_0 conda-forge\nrequests 2.25.1 pyhd3deb0d_0 conda-forge\nrope 0.19.0 pyhd8ed1ab_0 conda-forge\nrtree 0.9.7 py38h02d302b_1 conda-forge\nscikit-image 0.18.1 pypi_0 pypi\nscipy 1.6.3 pypi_0 pypi\nseaborn 0.11.2 pypi_0 pypi\nsecretstorage 3.3.1 py38h578d9bd_0 conda-forge\nsend2trash 1.5.0 pypi_0 pypi\nsentry-sdk 1.3.1 pypi_0 pypi\nsetuptools 52.0.0 py38h06a4308_0\nshapely 1.7.1 pypi_0 pypi\nshortuuid 1.0.1 pypi_0 pypi\nsix 1.16.0 pyh6c4a22f_0 conda-forge\nsmmap 4.0.0 pypi_0 pypi\nsnowballstemmer 2.1.0 pyhd8ed1ab_0 conda-forge\nsortedcontainers 2.4.0 pyhd8ed1ab_0 conda-forge\nsphinx 4.0.2 pyh6c4a22f_1 conda-forge\nsphinxcontrib-applehelp 1.0.2 py_0 conda-forge\nsphinxcontrib-devhelp 1.0.2 py_0 conda-forge\nsphinxcontrib-htmlhelp 2.0.0 pyhd8ed1ab_0 conda-forge\nsphinxcontrib-jsmath 1.0.1 py_0 conda-forge\nsphinxcontrib-qthelp 1.0.3 py_0 conda-forge\nsphinxcontrib-serializinghtml 1.1.5 pyhd8ed1ab_0 conda-forge\nspyder 4.2.5 py38h578d9bd_0 conda-forge\nspyder-kernels 1.10.2 py38h578d9bd_0 conda-forge\nsqlite 3.35.4 hdfb4753_0\nsubprocess32 3.5.4 pypi_0 pypi\ntensorboard 2.6.0 pypi_0 pypi\ntensorflow-datasets 4.4.0 pypi_0 pypi\ntensorflow-gpu 2.4.0 pypi_0 pypi\ntensorflow-metadata 1.2.0 pypi_0 pypi\nterminado 0.10.0 pypi_0 pypi\nterminaltables 3.1.0 pypi_0 pypi\ntestpath 0.5.0 pyhd8ed1ab_0 conda-forge\ntext-unidecode 1.3 pypi_0 pypi\ntextdistance 4.2.1 pyhd8ed1ab_0 conda-forge\nthree-merge 0.1.1 pyh9f0ad1d_0 conda-forge\ntifffile 2021.6.6 pypi_0 pypi\ntimm 0.4.13 pypi_0 pypi\ntk 8.6.10 hbc83047_0\ntoml 0.10.2 pyhd8ed1ab_0 conda-forge\ntorch 1.10.0.dev20210623+cu111 pypi_0 pypi\ntorchaudio 0.8.1 pypi_0 pypi\ntorchvision 0.11.0.dev20210623+cu111 pypi_0 pypi\ntornado 6.1 py38h497a2fe_1 conda-forge\ntraitlets 5.0.5 pyhd3eb1b0_0\nttach 0.0.3 pypi_0 pypi\ntyped-ast 1.4.3 py38h497a2fe_0 conda-forge\ntyping_extensions 3.10.0.0 pyha770c72_0 conda-forge\nubelt 0.9.5 pypi_0 pypi\nujson 4.0.2 py38h709712a_0 conda-forge\nurllib3 1.26.5 pyhd8ed1ab_0 conda-forge\nvirtualenv 20.4.7 pypi_0 pypi\nwandb 0.12.1 pypi_0 pypi\nwatchdog 1.0.2 py38h578d9bd_1 conda-forge\nwcwidth 0.2.5 py_0\nwebencodings 0.5.1 pypi_0 pypi\nwheel 0.36.2 pyhd3eb1b0_0\nwidgetsnbextension 3.5.1 pypi_0 pypi\nwrapt 1.12.1 py38h497a2fe_3 conda-forge\nwurlitzer 2.1.0 py38h578d9bd_0 conda-forge\nxdoctest 0.15.4 pypi_0 pypi\nxorg-libxau 1.0.9 h7f98852_0 conda-forge\nxorg-libxdmcp 1.1.3 h7f98852_0 conda-forge\nxz 5.2.5 h7b6447c_0\nyaml 0.2.5 h516909a_0 conda-forge\nyapf 0.31.0 pyhd8ed1ab_0 conda-forge\nzeromq 4.3.4 h9c3ff4c_0 conda-forge\nzipp 3.4.1 pyhd8ed1ab_0 conda-forge\nzlib 1.2.11 h7b6447c_3\nzstd 1.5.0 ha95c52a_0 conda-forge\n\n\nand using this i trained deeper models like robustscanner,nrtr,fcenet etc\nin another pytorch environment where i have pytorch 1.8 installed i trained many powerful models like xlm roberta,deberta large etc by utilizing 24gb vram and i didn't face this freezing issue\nso if it is hardware problem then i should experience this freezing issue in every model training(at least in most model training),no?\nfor example,,if i do not convert the wave signal into 512x512 image size for training,,then i do not face this freezing issue!\n\nmost of the times i face this freezing issue while trying nfnets,i think i am probably missing a line of code which could fix this freezing issue,i wonder if it is necessary to use torch.cuda.empty_cache() and gc.collect() always for getting rid of this freezing issue or something else that is causing the problem! 😪",
    "1505790": "Are you using AMD cpus and ubuntu?\nIs the freeze happening during an epoch or at start/end?\nDid you track the temperature?",
    "1505903": "philippsinger \n1. intel cpu core i9 with ubuntu\n2. during an epoch as you can see from the log at the bottom here : https://pastebin.com/XEusHppZ\n3. when the model was training i was not in front of the computer and i don't know if i can monitor and save temperate overtime when i am not in-front of the computer,sorry for this",
    "1505939": "If you use wandb, you can track some of your hardware's indicators. Not sure how reliable this is but it is better than nothing I guess?",
    "1506099": "and How Much it takes  on non RTX ones like  GTX 1660ti",
    "1506106": "I work on waveform only, with tfrecords and 8 TPU, I do one epoch into 27 seconds.",
    "1506250": "I also have a rtx  3090, I average 20 min per epoch with preprocessed data and ~120-180 min when processing data on the fly. Pytorch",
    "1506318": "Interesting, so you get a 6 or 8 factor from saving the images. I am planning to do it as well, will let you know how it goes.",
    "1506321": "Probably forever. 😄\n\nJoke aside, try to run on small images (128 x 128) and a very small model (EfficientNet B0 maybe or even something else). Also try to add gradient accumulation to avoid memory issues.",
    "1506333": "You can also do it on the fly on the GPU with nnAudio",
    "1506400": "mobassir how many Watts is your PSU ?",
    "1506470": "between this two it is not much about RT part, but tensor cores. 1660 has none. But again 1080ti for example is still a solid performer even without them.",
    "1506495": "Is wandb not tracking the temperature? Whats your temperature after you trained it for lets say 5-10 minutes? To me it sounds very much like overheating, this is typical for freezes.",
    "1506573": "Maybe its a issue with the memory junction temps. GDDR6X gets really hot. That might be a problem",
    "1507130": "I was not using wandb.soon i will use it and will let you know.thank you",
    "1507131": "mithilsalunkhe using this one : https://www.coolermaster.com/catalog/power-supplies/mwe-series/mwe-gold-850-v2-full-modular/\n\nAnd this cooler : https://www.coolermaster.com/catalog/coolers/cpu-air-coolers/hyper-212-led-turbo/",
    "1507255": "Does your case have enough airflow ?",
    "1509883": "I have preprocessed and saved the data but without any improvements so far. I guess either the bottleneck is in the IO operations (i.e. opening files) since the processed images are on a \"classical\" HDD (and not SSD) or I am doing something wrong? \n\nNotice that I am using the EfficienNet B0 model and 512 x 512 images.",
    "1509884": "That's quite impressive!",
    "1509885": "mobassir any update on your situation? 👀",
    "1509902": "how much RAM do you have?\n\nIt maybe better to upload all waves in ram and compute images on the fly using packages that run on GPU like torch.fft, nnaudio, etc.",
    "1509917": "RAM of 64GB. \n\nYes, if my SSD test doesn't give better results, I will for sure move to this option. I am also checking on a cloud GPU to see if the problem comes from my current setup. \n\nWill let you know.",
    "1509946": "SSD should definitely help. Move all your training data to it.",
    "1509976": "Update: I have resized the images to 256 x 256 and still using EfficientNet B0 and now can train one epoch for one fold in 1**2 minutes** (GPU usage is almost at max). \n\nThe secret was in getting a much bigger batch size: around 300.\n\nI can probably optimize even further but at least now I can run more experiments. 👌",
    "1510154": "yassinealtuned\nCurrently i am using kaggle tpu for this competition... \n\n sorry i am yet to train anything again with rtx3090.\nProblem is, i am using a server pc through remote desktop connection which has rtx3090.that server pc is hosted in my office which is far away from me and i am working from home!\nI use that pc via remote desktop connection \nSo It's difficult for me to restart the pc over and over again for this painful freezing issue >_<\nWhen it freezes,i lose connection and can't Connect to pc via remote desktop connection and then i disturb office colleagues/staff for restarting the pc(who is there in the office at that Moment)\nLast Thursday the pc got disconnected (maybe because of either shutdown/freeze issue)\nThen here in my territory Friday and Saturday’s are off day,so i found no office staff to restart the pc\nToday is sunday,so one office staff just turned on the pc again and now during daytime my office colleagues are using that pc for generating synthetic data and train models related to optical character recognition (ocr) for daytime office Project. \nThey(my office team mates) do not compete in kaggle\nI plan to try wandb again within next 2-3 days when the pc is not busy for training model for official works! \nI Just hate this freeze issue,restarting the pc for me is very difficult task as i work for home :(\nStay tuned,thanks",
    "1510947": "I switched to nn Audio and I also get similar results but with a smaller batchsize",
    "1513397": "yassinealouini this is the first time i used wandb,sorry if i am missing something.\nhere is the wandb folder where i have log : https://www.kaggle.com/mobassir/wandb-log\ni can't see any temperature related log,don't know how to interpret,let me know if you need the code that i tried to generate these logs,thanks",
    "1513451": "mobassir do you have access to the online dashboard? This is the easiest way to get some system information. I will have a look at your logs and let you know later. 👌",
    "1513498": "sorry @yassinealouini \ni don't have any online dashboard,i am simply using wandb like this kernel : https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\n\ni just modified his code a little bit for locally training,,i use like this : \n\n\nimport wandb\nwandb.login()\n\nand i do not use these : \n\n#from kaggle_secrets import UserSecretsClient\n# user_secrets = UserSecretsClient()\n# wandb_api = user_secrets.get_secret(\"wandb_api\")\n# wandb.login(key=wandb_api)\n\nand then he was doing this : \n\nrun = wandb.init(project=\"G2Net-Public-experiments\", \n                 name=\"exp1\",\n                 config=class2dict(CFG),\n                 group=CFG.model_name,\n                 job_type=\"train\")\n\nwhere i did this : \n\nrun = wandb.init(project=\"G2Net-Public-experiments\", \n                 name=\"exp1\",\n                 config=class2dict(CFG),\n                 group=CFG.model_name,\n                 job_type=\"train\",\n                 settings=wandb.Settings(_disable_stats=False))\n\neverything else is same for wandb logging like him",
    "1513499": "mobassir Are there any other GPUs in that machine? The reason I ask is that when the 2080Ti came out, I wanted to use it alongside my old GTX 1080. But during training it would randomly freeze even if the 1080 was idle and nothing was showing on any logs (therefore a catastrophic kernel issue). The only way I could fix it was by removing the 1080 from the system, and I think the crashing was probably caused by some driver issue caused by mixing different generation cards.",
    "1513515": "anjum48 i am using only rtx3090(single gpu)\n\nProcessor    : i9 10th Gen 10Core\nRAM          : 32GB (3200Mhz)\nHDD          : 1TB SSD Nvme\nGPU          : MSI Nvidia RTX 3090 24GB",
    "1513726": "mobassir I only used wandb once so far, and outside Kaggle, but it seems you need to create a wandb account and use your credentials.  I have not user kaggle secret either but it looks like a way to publicly share notebooks without revealing your credentials.",
    "1514121": "dear @cpmpml @yassinealouini @philippsinger here you can see my system metric(screenshot) : https://www.kaggle.com/mobassir/wandb-log?select=1.PNG\ncheck 1.png,2.png,3.png and 4.png\nthank you",
    "1517966": "For those still following this thread, the fastest way I have found so far is using **TFRecords** and **TPUs**. I will share how long it takes once the competition is over. 👌",
    "1517967": "mobassir it looks like the graphs are fine except the first dip we see. I don't have other options to offer for now. 😬",
    "1520899": "Something that I've now noticed. It's always the model being trained on gpu:0 that crashes. I'm training on a desktop, so I wonder if it has something to do with X / the display system interfering. Once this competition is done, I plan on swapping the physical positions of the two cards to see if the data loader thread crashing persists on the same GPU device or on the same slot to better debug.\n\nAlso, does anyone see:\n\n`[W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool)` being spammed to the console where jupyter notebook is invoked?",
    "1520942": "> Also, does anyone see:\n\n> [W pthreadpool-cpp.cc:90] Warning: Leaking Caffe2 thread-pool after fork. (function pthreadpool) being spammed to the console where jupyter notebook is invoked?\n\nYes, that's a very annoying bug in pytorch (which looks like it might be fixed in 1.10) https://github.com/pytorch/pytorch/issues/57273",
    "1521032": "Wow, I took out `pin_memory=True` and behold, all my crashing issues have been resolved.",
    "1521047": "authman Yes, I had a similar warning.",
    "1524715": "Great! :-) Thanks to your discussion post I started to learn much more about how to optimise my GPU and TPU performances and it still goes on...",
    "1524777": "That's awesome! I have learned many things as well, might share them once the competition ends. 👌",
    "1559736": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}