{
  "id": 276146,
  "title": "Few things I have learned",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/276146",
  "author_name": "",
  "post_date": "2021-10-03T11:28:16.226109100Z",
  "votes": 14,
  "comment_count": 18,
  "views": 0,
  "content": "<p>First, thanks to all the participants and organizers for this great and tough competition. </p>\n<p>I am to some extent new to signal processing in general and to <a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/e3ad18d4c345806aa42732ffc7b8f60b3ab0e071/Installation/nnAudio/Spectrogram.py#L750\" target=\"_blank\"><strong>CQT</strong></a> and <a href=\"https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration\" target=\"_blank\"><strong>CWT</strong></a> methods in particular. </p>\n<p>Here are some of the things in no particular order I've learned:</p>\n<ul>\n<li>make your baseline <strong>run as fast as possible</strong>: I've written about this <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612\" target=\"_blank\"><strong>here</strong></a> and lots of discussions happened so thanks to all those that have shared ideas. At the end of the day, the best option was using <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\" target=\"_blank\"><strong>TFRecord</strong></a> and <a href=\"https://cloud.google.com/tpu\" target=\"_blank\">TPU</a> via colab pro + and I am glad I was able to learn more about these.</li>\n<li>don't underestimate <strong>possible solutions</strong>: I've focused mainly on 2D models and put aside 1D models. This, in hindsight, was a mistake, I should have split my time between the two approaches.</li>\n<li>public notebooks <strong>are useful to some extent</strong>: use those that you can reproduce and improve upon, avoid as the plague the \"ensemble of ensemble of ensemble\" ones. Most importantly, try to understand why one notebook works and how you can improve it.</li>\n<li>sometimes, you have to work with <strong>things you don't master</strong>: I prefer working with <a href=\"https://pytorch.org/\" target=\"_blank\">PyTorch</a> as much as possible but due to lack of time, I had to work with TensorFlow and <a href=\"https://keras.io/\" target=\"_blank\">Keras</a> since I had a baseline that worked well with TFRecord and TPU. </li>\n<li><strong>it takes time,</strong> especially for deep learning competitions: 1 month for a deep learning competition isn't enough (well at least if the competition is longer), the sweet spot is probably 6 weeks. Indeed, I've started to see some promising results after 3 weeks and at the end of 4 weeks, I had few other ideas I wanted to try.</li>\n</ul>\n<p>Finally, thanks to <a href=\"https://www.kaggle.com/headsortails\" target=\"_blank\">@headsortails</a> for teaming up with me, it was a privilege working with him and sharing ideas. Hopefully our next collaboration will lead to a medal. 😊</p>",
  "messages": [
    {
      "id": "1532772",
      "postDate": "10/03/2021 11:28:16",
      "content": "<p>First, thanks to all the participants and organizers for this great and tough competition. </p>\n<p>I am to some extent new to signal processing in general and to <a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/e3ad18d4c345806aa42732ffc7b8f60b3ab0e071/Installation/nnAudio/Spectrogram.py#L750\" target=\"_blank\"><strong>CQT</strong></a> and <a href=\"https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration\" target=\"_blank\"><strong>CWT</strong></a> methods in particular. </p>\n<p>Here are some of the things in no particular order I've learned:</p>\n<ul>\n<li>make your baseline <strong>run as fast as possible</strong>: I've written about this <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612\" target=\"_blank\"><strong>here</strong></a> and lots of discussions happened so thanks to all those that have shared ideas. At the end of the day, the best option was using <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\" target=\"_blank\"><strong>TFRecord</strong></a> and <a href=\"https://cloud.google.com/tpu\" target=\"_blank\">TPU</a> via colab pro + and I am glad I was able to learn more about these.</li>\n<li>don't underestimate <strong>possible solutions</strong>: I've focused mainly on 2D models and put aside 1D models. This, in hindsight, was a mistake, I should have split my time between the two approaches.</li>\n<li>public notebooks <strong>are useful to some extent</strong>: use those that you can reproduce and improve upon, avoid as the plague the \"ensemble of ensemble of ensemble\" ones. Most importantly, try to understand why one notebook works and how you can improve it.</li>\n<li>sometimes, you have to work with <strong>things you don't master</strong>: I prefer working with <a href=\"https://pytorch.org/\" target=\"_blank\">PyTorch</a> as much as possible but due to lack of time, I had to work with TensorFlow and <a href=\"https://keras.io/\" target=\"_blank\">Keras</a> since I had a baseline that worked well with TFRecord and TPU. </li>\n<li><strong>it takes time,</strong> especially for deep learning competitions: 1 month for a deep learning competition isn't enough (well at least if the competition is longer), the sweet spot is probably 6 weeks. Indeed, I've started to see some promising results after 3 weeks and at the end of 4 weeks, I had few other ideas I wanted to try.</li>\n</ul>\n<p>Finally, thanks to <a href=\"https://www.kaggle.com/headsortails\" target=\"_blank\">@headsortails</a> for teaming up with me, it was a privilege working with him and sharing ideas. Hopefully our next collaboration will lead to a medal. 😊</p>",
      "rawMarkdown": "First, thanks to all the participants and organizers for this great and tough competition. \n\nI am to some extent new to signal processing in general and to [**CQT**](https://github.com/KinWaiCheuk/nnAudio/blob/e3ad18d4c345806aa42732ffc7b8f60b3ab0e071/Installation/nnAudio/Spectrogram.py#L750) and [**CWT**](https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration) methods in particular. \n\nHere are some of the things in no particular order I've learned:\n\n- make your baseline **run as fast as possible**: I've written about this [**here**](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612) and lots of discussions happened so thanks to all those that have shared ideas. At the end of the day, the best option was using [**TFRecord**](https://www.tensorflow.org/tutorials/load_data/tfrecord) and [TPU](https://cloud.google.com/tpu) via colab pro + and I am glad I was able to learn more about these.\n- don't underestimate **possible solutions**: I've focused mainly on 2D models and put aside 1D models. This, in hindsight, was a mistake, I should have split my time between the two approaches.\n- public notebooks **are useful to some extent**: use those that you can reproduce and improve upon, avoid as the plague the \"ensemble of ensemble of ensemble\" ones. Most importantly, try to understand why one notebook works and how you can improve it.\n- sometimes, you have to work with **things you don't master**: I prefer working with [PyTorch](https://pytorch.org/) as much as possible but due to lack of time, I had to work with TensorFlow and [Keras](https://keras.io/) since I had a baseline that worked well with TFRecord and TPU. \n- **it takes time,** especially for deep learning competitions: 1 month for a deep learning competition isn't enough (well at least if the competition is longer), the sweet spot is probably 6 weeks. Indeed, I've started to see some promising results after 3 weeks and at the end of 4 weeks, I had few other ideas I wanted to try.\n\n\nFinally, thanks to @headsortails for teaming up with me, it was a privilege working with him and sharing ideas. Hopefully our next collaboration will lead to a medal. 😊",
      "votes": null
    },
    {
      "id": "1533497",
      "postDate": "10/04/2021 04:55:11",
      "content": "<p><strong>things that i couldn't learn</strong> : </p>\n<ol>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n</ol>",
      "rawMarkdown": "**things that i couldn't learn** : \n\n1. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n2. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n3. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n4. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n5. how to solve my rtx3090 freezing issue during training pytorch models 😭😪",
      "votes": null
    },
    {
      "id": "1533563",
      "postDate": "10/04/2021 06:31:31",
      "content": "<p>That's sad indeed. I hope you find a solution very soon. 💪</p>",
      "rawMarkdown": "That's sad indeed. I hope you find a solution very soon. 💪",
      "votes": null
    },
    {
      "id": "1533682",
      "postDate": "10/04/2021 08:46:28",
      "content": "<p>Thank you for your sharing.👍</p>",
      "rawMarkdown": "Thank you for your sharing.👍",
      "votes": null
    },
    {
      "id": "1534162",
      "postDate": "10/04/2021 17:01:24",
      "content": "<p>TdrDelay ?</p>",
      "rawMarkdown": "TdrDelay ?",
      "votes": null
    },
    {
      "id": "1534199",
      "postDate": "10/04/2021 17:45:40",
      "content": "<p><a href=\"https://www.kaggle.com/glimmung\" target=\"_blank\">@glimmung</a>  full discussion regarding the freezing issue that i faced and can't solve can be found here : <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612#1504405\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612#1504405</a></p>",
      "rawMarkdown": "glimmung  full discussion regarding the freezing issue that i faced and can't solve can be found here : https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612#1504405",
      "votes": null
    },
    {
      "id": "1534210",
      "postDate": "10/04/2021 17:57:58",
      "content": "<p>Glad it helps. 👌 </p>",
      "rawMarkdown": "Glad it helps. 👌",
      "votes": null
    },
    {
      "id": "1534376",
      "postDate": "10/04/2021 20:40:20",
      "content": "<p>Did the <code>pin_memory=False</code> trick not work for you? It has completely solved my 3090 freezing issue. Now that I'm between competitions I'm very tempted to clean install cuda/cudnn/torch/etc. but I'm too afraid since I finally got stable…</p>\n<p><a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> if possible, could you share here also some of those cpu cache clearing scripts you use?</p>",
      "rawMarkdown": "Did the `pin_memory=False` trick not work for you? It has completely solved my 3090 freezing issue. Now that I'm between competitions I'm very tempted to clean install cuda/cudnn/torch/etc. but I'm too afraid since I finally got stable...\n\n@jpison if possible, could you share here also some of those cpu cache clearing scripts you use?",
      "votes": null
    },
    {
      "id": "1534587",
      "postDate": "10/05/2021 04:33:21",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> here is the code that i tried last time locally that failed : <a href=\"https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing</a><br>\ninstead of this : <br>\ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(196,512)).squeeze().numpy()</p>\n<p>i did this : <br>\ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(512,512)).squeeze().numpy()</p>\n<p>increased the batch size,tuned cqt and changed the backbone (i realized that higher the batch size,faster the pc will freeze)</p>\n<p>please check the code,i am not using pin_memory = True and by default pytorch uses pin_memory = False</p>\n<p>this code works everywhere(colab pro and in kaggle) but not in my PC 😭<br>\ni tried 3 different pytorch baseline in this competition and all got frozen</p>\n<ol>\n<li>y.nakama's pipeline</li>\n<li>preprocess tfrecord and train with pytorch</li>\n<li>this one : <a href=\"https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\" target=\"_blank\">https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E</a></li>\n</ol>\n<p>i just used  F.interpolate() in those 3 notebooks and they all crashed,,,last time i tried resnet34d in this baseline :  <a href=\"https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\" target=\"_blank\">https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E</a><br>\ni tuned cqt and used batch size 128 which took around 8 gb vram iirc and it crashed within 5 minutes,,,<br>\njust to make sure i retried and this time with batch size 256 and it crashed(pc got frozen) within 3-4 minutes!</p>",
      "rawMarkdown": "authman here is the code that i tried last time locally that failed : https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing\ninstead of this : \ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(196,512)).squeeze().numpy()\n\ni did this : \ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(512,512)).squeeze().numpy()\n\nincreased the batch size,tuned cqt and changed the backbone (i realized that higher the batch size,faster the pc will freeze)\n\nplease check the code,i am not using pin_memory = True and by default pytorch uses pin_memory = False\n\nthis code works everywhere(colab pro and in kaggle) but not in my PC 😭\ni tried 3 different pytorch baseline in this competition and all got frozen\n1. y.nakama's pipeline\n2. preprocess tfrecord and train with pytorch\n3. this one : https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\n\ni just used  F.interpolate() in those 3 notebooks and they all crashed,,,last time i tried resnet34d in this baseline :  https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\ni tuned cqt and used batch size 128 which took around 8 gb vram iirc and it crashed within 5 minutes,,,\njust to make sure i retried and this time with batch size 256 and it crashed(pc got frozen) within 3-4 minutes!",
      "votes": null
    },
    {
      "id": "1534743",
      "postDate": "10/05/2021 08:04:21",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> never reinstall CUDA if it is already working for you. 😄<br>\nJoke aside, yes, understanding how to get a stable install is a valuable skill to have for sure.</p>",
      "rawMarkdown": "authman never reinstall CUDA if it is already working for you. 😄\nJoke aside, yes, understanding how to get a stable install is a valuable skill to have for sure.",
      "votes": null
    },
    {
      "id": "1534745",
      "postDate": "10/05/2021 08:05:10",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> it seems the saga is continuing. :(<br>\nHave you tried contacting the NVIDIA support? Maybe you need to return the GPU?</p>",
      "rawMarkdown": "mobassir it seems the saga is continuing. :(\nHave you tried contacting the NVIDIA support? Maybe you need to return the GPU?",
      "votes": null
    },
    {
      "id": "1534777",
      "postDate": "10/05/2021 08:43:54",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> i can't conclude that this problem is occurring because of faulty gpu hence can't claim replacement warranty,i tried to train many other networks for different kaggle competitions using this gpu they are xlm roberta large,deberta large,robustscanner etc etc and not always i face this freezing issue like this competition and while using models like nfnets,,and in this competition even the resnet34d was using only 8gb vram out of 24 and it made the pc frozen,,maybe gpu is okay(we also saw wandb system matrices where gpu was working fine)</p>",
      "rawMarkdown": "yassinealouini i can't conclude that this problem is occurring because of faulty gpu hence can't claim replacement warranty,i tried to train many other networks for different kaggle competitions using this gpu they are xlm roberta large,deberta large,robustscanner etc etc and not always i face this freezing issue like this competition and while using models like nfnets,,and in this competition even the resnet34d was using only 8gb vram out of 24 and it made the pc frozen,,maybe gpu is okay(we also saw wandb system matrices where gpu was working fine)",
      "votes": null
    },
    {
      "id": "1538531",
      "postDate": "10/08/2021 14:12:57",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> what Cuda version are you using ?</p>",
      "rawMarkdown": "mobassir what Cuda version are you using ?",
      "votes": null
    },
    {
      "id": "1538594",
      "postDate": "10/08/2021 15:08:20",
      "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> </p>\n<p><strong>(mobassir) apsisdev@ML:~$ nvcc --version</strong><br>\nnvcc: NVIDIA (R) Cuda compiler driver<br>\nCopyright (c) 2005-2021 NVIDIA Corporation<br>\nBuilt on Sun_Feb_14_21:12:58_PST_2021<br>\nCuda compilation tools, release 11.2, V11.2.152<br>\nBuild cuda_11.2.r11.2/compiler.29618528_0</p>\n<pre><code>import torch\nprint(torch.version.cuda)\n</code></pre>\n<p><strong>output : 11.1</strong></p>",
      "rawMarkdown": "mithilsalunkhe \n\n**(mobassir) apsisdev@ML:~$ nvcc --version**\nnvcc: NVIDIA (R) Cuda compiler driver\nCopyright (c) 2005-2021 NVIDIA Corporation\nBuilt on Sun_Feb_14_21:12:58_PST_2021\nCuda compilation tools, release 11.2, V11.2.152\nBuild cuda_11.2.r11.2/compiler.29618528_0\n\n```\nimport torch\nprint(torch.version.cuda)\n```\n**output : 11.1**",
      "votes": null
    },
    {
      "id": "1538620",
      "postDate": "10/08/2021 15:30:15",
      "content": "<p>Pytorch version <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> </p>",
      "rawMarkdown": "Pytorch version @mobassir",
      "votes": null
    },
    {
      "id": "1538623",
      "postDate": "10/08/2021 15:32:00",
      "content": "<p>Also try out tensorflow. If tensorflow also does not work that means it is a GPU problem. If it works then it is a Pytorch problem</p>",
      "rawMarkdown": "Also try out tensorflow. If tensorflow also does not work that means it is a GPU problem. If it works then it is a Pytorch problem",
      "votes": null
    },
    {
      "id": "1538718",
      "postDate": "10/08/2021 16:49:21",
      "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>  tried both pytorch 1.8 and 1.10<br>\nno luck. i never faced this issued with tf model training but unfortunately i don't use tensorflow much,even if i use,i mostly use kaggle tpu for tf model</p>",
      "rawMarkdown": "mithilsalunkhe  tried both pytorch 1.8 and 1.10\nno luck. i never faced this issued with tf model training but unfortunately i don't use tensorflow much,even if i use,i mostly use kaggle tpu for tf model",
      "votes": null
    },
    {
      "id": "1539009",
      "postDate": "10/09/2021 02:52:07",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> try out this command <br>\n<code>pip3 install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html</code></p>",
      "rawMarkdown": "mobassir try out this command \n`pip3 install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html`",
      "votes": null
    },
    {
      "id": "1559911",
      "postDate": "10/27/2021 08:08:23",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1533497,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "10/04/2021 04:55:11",
      "content": "<p><strong>things that i couldn't learn</strong> : </p>\n<ol>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n<li>how to solve my rtx3090 freezing issue during training pytorch models 😭😪</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1533563,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "10/04/2021 06:31:31",
          "content": "<p>That's sad indeed. I hope you find a solution very soon. 💪</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1534162,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "10/04/2021 17:01:24",
          "content": "<p>TdrDelay ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1534199,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "10/04/2021 17:45:40",
          "content": "<p><a href=\"https://www.kaggle.com/glimmung\" target=\"_blank\">@glimmung</a>  full discussion regarding the freezing issue that i faced and can't solve can be found here : <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612#1504405\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612#1504405</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1534376,
          "author_name": "authman",
          "author_url": "",
          "post_date": "10/04/2021 20:40:20",
          "content": "<p>Did the <code>pin_memory=False</code> trick not work for you? It has completely solved my 3090 freezing issue. Now that I'm between competitions I'm very tempted to clean install cuda/cudnn/torch/etc. but I'm too afraid since I finally got stable…</p>\n<p><a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> if possible, could you share here also some of those cpu cache clearing scripts you use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1534587,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "10/05/2021 04:33:21",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> here is the code that i tried last time locally that failed : <a href=\"https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing</a><br>\ninstead of this : <br>\ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(196,512)).squeeze().numpy()</p>\n<p>i did this : <br>\ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(512,512)).squeeze().numpy()</p>\n<p>increased the batch size,tuned cqt and changed the backbone (i realized that higher the batch size,faster the pc will freeze)</p>\n<p>please check the code,i am not using pin_memory = True and by default pytorch uses pin_memory = False</p>\n<p>this code works everywhere(colab pro and in kaggle) but not in my PC 😭<br>\ni tried 3 different pytorch baseline in this competition and all got frozen</p>\n<ol>\n<li>y.nakama's pipeline</li>\n<li>preprocess tfrecord and train with pytorch</li>\n<li>this one : <a href=\"https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\" target=\"_blank\">https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E</a></li>\n</ol>\n<p>i just used  F.interpolate() in those 3 notebooks and they all crashed,,,last time i tried resnet34d in this baseline :  <a href=\"https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\" target=\"_blank\">https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E</a><br>\ni tuned cqt and used batch size 128 which took around 8 gb vram iirc and it crashed within 5 minutes,,,<br>\njust to make sure i retried and this time with batch size 256 and it crashed(pc got frozen) within 3-4 minutes!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1534743,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "10/05/2021 08:04:21",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> never reinstall CUDA if it is already working for you. 😄<br>\nJoke aside, yes, understanding how to get a stable install is a valuable skill to have for sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1534745,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "10/05/2021 08:05:10",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> it seems the saga is continuing. :(<br>\nHave you tried contacting the NVIDIA support? Maybe you need to return the GPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1534777,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "10/05/2021 08:43:54",
          "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> i can't conclude that this problem is occurring because of faulty gpu hence can't claim replacement warranty,i tried to train many other networks for different kaggle competitions using this gpu they are xlm roberta large,deberta large,robustscanner etc etc and not always i face this freezing issue like this competition and while using models like nfnets,,and in this competition even the resnet34d was using only 8gb vram out of 24 and it made the pc frozen,,maybe gpu is okay(we also saw wandb system matrices where gpu was working fine)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1538531,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "10/08/2021 14:12:57",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> what Cuda version are you using ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1538594,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "10/08/2021 15:08:20",
          "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> </p>\n<p><strong>(mobassir) apsisdev@ML:~$ nvcc --version</strong><br>\nnvcc: NVIDIA (R) Cuda compiler driver<br>\nCopyright (c) 2005-2021 NVIDIA Corporation<br>\nBuilt on Sun_Feb_14_21:12:58_PST_2021<br>\nCuda compilation tools, release 11.2, V11.2.152<br>\nBuild cuda_11.2.r11.2/compiler.29618528_0</p>\n<pre><code>import torch\nprint(torch.version.cuda)\n</code></pre>\n<p><strong>output : 11.1</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1538620,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "10/08/2021 15:30:15",
          "content": "<p>Pytorch version <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1538623,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "10/08/2021 15:32:00",
          "content": "<p>Also try out tensorflow. If tensorflow also does not work that means it is a GPU problem. If it works then it is a Pytorch problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1538718,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "10/08/2021 16:49:21",
          "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>  tried both pytorch 1.8 and 1.10<br>\nno luck. i never faced this issued with tf model training but unfortunately i don't use tensorflow much,even if i use,i mostly use kaggle tpu for tf model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1539009,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "10/09/2021 02:52:07",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> try out this command <br>\n<code>pip3 install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1533682,
      "author_name": "ohseokkim",
      "author_url": "",
      "post_date": "10/04/2021 08:46:28",
      "content": "<p>Thank you for your sharing.👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1534210,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "10/04/2021 17:57:58",
          "content": "<p>Glad it helps. 👌 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559911,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:08:23",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1532772": "First, thanks to all the participants and organizers for this great and tough competition. \n\nI am to some extent new to signal processing in general and to [**CQT**](https://github.com/KinWaiCheuk/nnAudio/blob/e3ad18d4c345806aa42732ffc7b8f60b3ab0e071/Installation/nnAudio/Spectrogram.py#L750) and [**CWT**](https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration) methods in particular. \n\nHere are some of the things in no particular order I've learned:\n\n- make your baseline **run as fast as possible**: I've written about this [**here**](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612) and lots of discussions happened so thanks to all those that have shared ideas. At the end of the day, the best option was using [**TFRecord**](https://www.tensorflow.org/tutorials/load_data/tfrecord) and [TPU](https://cloud.google.com/tpu) via colab pro + and I am glad I was able to learn more about these.\n- don't underestimate **possible solutions**: I've focused mainly on 2D models and put aside 1D models. This, in hindsight, was a mistake, I should have split my time between the two approaches.\n- public notebooks **are useful to some extent**: use those that you can reproduce and improve upon, avoid as the plague the \"ensemble of ensemble of ensemble\" ones. Most importantly, try to understand why one notebook works and how you can improve it.\n- sometimes, you have to work with **things you don't master**: I prefer working with [PyTorch](https://pytorch.org/) as much as possible but due to lack of time, I had to work with TensorFlow and [Keras](https://keras.io/) since I had a baseline that worked well with TFRecord and TPU. \n- **it takes time,** especially for deep learning competitions: 1 month for a deep learning competition isn't enough (well at least if the competition is longer), the sweet spot is probably 6 weeks. Indeed, I've started to see some promising results after 3 weeks and at the end of 4 weeks, I had few other ideas I wanted to try.\n\n\nFinally, thanks to @headsortails for teaming up with me, it was a privilege working with him and sharing ideas. Hopefully our next collaboration will lead to a medal. 😊",
    "1533497": "**things that i couldn't learn** : \n\n1. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n2. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n3. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n4. how to solve my rtx3090 freezing issue during training pytorch models 😭😪\n5. how to solve my rtx3090 freezing issue during training pytorch models 😭😪",
    "1533563": "That's sad indeed. I hope you find a solution very soon. 💪",
    "1533682": "Thank you for your sharing.👍",
    "1534162": "TdrDelay ?",
    "1534199": "glimmung  full discussion regarding the freezing issue that i faced and can't solve can be found here : https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270612#1504405",
    "1534210": "Glad it helps. 👌",
    "1534376": "Did the `pin_memory=False` trick not work for you? It has completely solved my 3090 freezing issue. Now that I'm between competitions I'm very tempted to clean install cuda/cudnn/torch/etc. but I'm too afraid since I finally got stable...\n\n@jpison if possible, could you share here also some of those cpu cache clearing scripts you use?",
    "1534587": "authman here is the code that i tried last time locally that failed : https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing\ninstead of this : \ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(196,512)).squeeze().numpy()\n\ni did this : \ndata = F.interpolate(torch.tensor(data).expand(1,3, 46, 513), size=(512,512)).squeeze().numpy()\n\nincreased the batch size,tuned cqt and changed the backbone (i realized that higher the batch size,faster the pc will freeze)\n\nplease check the code,i am not using pin_memory = True and by default pytorch uses pin_memory = False\n\nthis code works everywhere(colab pro and in kaggle) but not in my PC 😭\ni tried 3 different pytorch baseline in this competition and all got frozen\n1. y.nakama's pipeline\n2. preprocess tfrecord and train with pytorch\n3. this one : https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\n\ni just used  F.interpolate() in those 3 notebooks and they all crashed,,,last time i tried resnet34d in this baseline :  https://colab.research.google.com/drive/1BA1_LGYjRSxhgaPJNdbg3WxiHMT6Ctxn?usp=sharing#scrollTo=Fsb74FTCGs9E\ni tuned cqt and used batch size 128 which took around 8 gb vram iirc and it crashed within 5 minutes,,,\njust to make sure i retried and this time with batch size 256 and it crashed(pc got frozen) within 3-4 minutes!",
    "1534743": "authman never reinstall CUDA if it is already working for you. 😄\nJoke aside, yes, understanding how to get a stable install is a valuable skill to have for sure.",
    "1534745": "mobassir it seems the saga is continuing. :(\nHave you tried contacting the NVIDIA support? Maybe you need to return the GPU?",
    "1534777": "yassinealouini i can't conclude that this problem is occurring because of faulty gpu hence can't claim replacement warranty,i tried to train many other networks for different kaggle competitions using this gpu they are xlm roberta large,deberta large,robustscanner etc etc and not always i face this freezing issue like this competition and while using models like nfnets,,and in this competition even the resnet34d was using only 8gb vram out of 24 and it made the pc frozen,,maybe gpu is okay(we also saw wandb system matrices where gpu was working fine)",
    "1538531": "mobassir what Cuda version are you using ?",
    "1538594": "mithilsalunkhe \n\n**(mobassir) apsisdev@ML:~$ nvcc --version**\nnvcc: NVIDIA (R) Cuda compiler driver\nCopyright (c) 2005-2021 NVIDIA Corporation\nBuilt on Sun_Feb_14_21:12:58_PST_2021\nCuda compilation tools, release 11.2, V11.2.152\nBuild cuda_11.2.r11.2/compiler.29618528_0\n\n```\nimport torch\nprint(torch.version.cuda)\n```\n**output : 11.1**",
    "1538620": "Pytorch version @mobassir",
    "1538623": "Also try out tensorflow. If tensorflow also does not work that means it is a GPU problem. If it works then it is a Pytorch problem",
    "1538718": "mithilsalunkhe  tried both pytorch 1.8 and 1.10\nno luck. i never faced this issued with tf model training but unfortunately i don't use tensorflow much,even if i use,i mostly use kaggle tpu for tf model",
    "1539009": "mobassir try out this command \n`pip3 install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html`",
    "1559911": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}