{
  "id": 239044,
  "title": "Help needed in solving the below error. It is realted to Cudnn and Cuda.",
  "url": "/competitions/birdclef-2021/discussion/239044",
  "author_name": "",
  "post_date": "2021-05-14T12:20:23.432836Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi guys,</p>\n<p>I am running the model training on Kaggle. I am using Pytorch as my framework. The training process works well for say 30 epochs and sometimes more than that but after that suddenly it throws an error. Below is the trackback of it.</p>\n<pre><code>---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\n&lt;ipython-input-26-7a1e5b630b5c&gt; in &lt;module&gt;\n     51                 device=device,\n     52                 input_key=\"image\",\n---&gt; 53                 input_target_key=\"targets\"\n     54             )\n     55             scheduler.step()\n\n&lt;ipython-input-25-7bd46f9c870e&gt; in train_one_epoch(model, dataloader, optimizer, criterion, device, input_key, input_target_key)\n     18         loss = criterion(outputs, y)\n     19         optimizer.zero_grad()\n---&gt; 20         loss.backward()\n     21         optimizer.step()\n     22 \n\n/opt/conda/lib/python3.7/site-packages/torch/tensor.py in backward(self, gradient, retain_graph, create_graph)\n    219                 retain_graph=retain_graph,\n    220                 create_graph=create_graph)\n--&gt; 221         torch.autograd.backward(self, gradient, retain_graph, create_graph)\n    222 \n    223     def register_hook(self, hook):\n\n/opt/conda/lib/python3.7/site-packages/torch/autograd/__init__.py in backward(tensors, grad_tensors, retain_graph, create_graph, grad_variables)\n    130     Variable._execution_engine.run_backward(\n    131         tensors, grad_tensors_, retain_graph, create_graph,\n--&gt; 132         allow_unreachable=True)  # allow_unreachable flag\n    133 \n    134 \n\nRuntimeError: cuDNN error: CUDNN_STATUS_NOT_SUPPORTED. This error may appear if you passed in a non-contiguous input.\nYou can try to repro this exception using the following code snippet. If that doesn't trigger the error, please include your original repro script when reporting this issue.\n</code></pre>\n<p>So far what I understood is this </p>\n<ol>\n<li>It is due to <code>loss.backward()</code> call.</li>\n<li>Something has to be off either in the outputs of the models or targets to get this error. </li>\n</ol>\n<p>I am not sure what could go wrong at 30+ epoch and trigger the error. It is very unusual to see such an error.</p>",
  "messages": [
    {
      "id": "1307387",
      "postDate": "05/14/2021 12:20:23",
      "content": "<p>Hi guys,</p>\n<p>I am running the model training on Kaggle. I am using Pytorch as my framework. The training process works well for say 30 epochs and sometimes more than that but after that suddenly it throws an error. Below is the trackback of it.</p>\n<pre><code>---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\n&lt;ipython-input-26-7a1e5b630b5c&gt; in &lt;module&gt;\n     51                 device=device,\n     52                 input_key=\"image\",\n---&gt; 53                 input_target_key=\"targets\"\n     54             )\n     55             scheduler.step()\n\n&lt;ipython-input-25-7bd46f9c870e&gt; in train_one_epoch(model, dataloader, optimizer, criterion, device, input_key, input_target_key)\n     18         loss = criterion(outputs, y)\n     19         optimizer.zero_grad()\n---&gt; 20         loss.backward()\n     21         optimizer.step()\n     22 \n\n/opt/conda/lib/python3.7/site-packages/torch/tensor.py in backward(self, gradient, retain_graph, create_graph)\n    219                 retain_graph=retain_graph,\n    220                 create_graph=create_graph)\n--&gt; 221         torch.autograd.backward(self, gradient, retain_graph, create_graph)\n    222 \n    223     def register_hook(self, hook):\n\n/opt/conda/lib/python3.7/site-packages/torch/autograd/__init__.py in backward(tensors, grad_tensors, retain_graph, create_graph, grad_variables)\n    130     Variable._execution_engine.run_backward(\n    131         tensors, grad_tensors_, retain_graph, create_graph,\n--&gt; 132         allow_unreachable=True)  # allow_unreachable flag\n    133 \n    134 \n\nRuntimeError: cuDNN error: CUDNN_STATUS_NOT_SUPPORTED. This error may appear if you passed in a non-contiguous input.\nYou can try to repro this exception using the following code snippet. If that doesn't trigger the error, please include your original repro script when reporting this issue.\n</code></pre>\n<p>So far what I understood is this </p>\n<ol>\n<li>It is due to <code>loss.backward()</code> call.</li>\n<li>Something has to be off either in the outputs of the models or targets to get this error. </li>\n</ol>\n<p>I am not sure what could go wrong at 30+ epoch and trigger the error. It is very unusual to see such an error.</p>",
      "rawMarkdown": "Hi guys,\n\nI am running the model training on Kaggle. I am using Pytorch as my framework. The training process works well for say 30 epochs and sometimes more than that but after that suddenly it throws an error. Below is the trackback of it.\n\n```\n---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\n<ipython-input-26-7a1e5b630b5c> in <module>\n     51                 device=device,\n     52                 input_key=\"image\",\n---> 53                 input_target_key=\"targets\"\n     54             )\n     55             scheduler.step()\n\n<ipython-input-25-7bd46f9c870e> in train_one_epoch(model, dataloader, optimizer, criterion, device, input_key, input_target_key)\n     18         loss = criterion(outputs, y)\n     19         optimizer.zero_grad()\n---> 20         loss.backward()\n     21         optimizer.step()\n     22 \n\n/opt/conda/lib/python3.7/site-packages/torch/tensor.py in backward(self, gradient, retain_graph, create_graph)\n    219                 retain_graph=retain_graph,\n    220                 create_graph=create_graph)\n--> 221         torch.autograd.backward(self, gradient, retain_graph, create_graph)\n    222 \n    223     def register_hook(self, hook):\n\n/opt/conda/lib/python3.7/site-packages/torch/autograd/__init__.py in backward(tensors, grad_tensors, retain_graph, create_graph, grad_variables)\n    130     Variable._execution_engine.run_backward(\n    131         tensors, grad_tensors_, retain_graph, create_graph,\n--> 132         allow_unreachable=True)  # allow_unreachable flag\n    133 \n    134 \n\nRuntimeError: cuDNN error: CUDNN_STATUS_NOT_SUPPORTED. This error may appear if you passed in a non-contiguous input.\nYou can try to repro this exception using the following code snippet. If that doesn't trigger the error, please include your original repro script when reporting this issue.\n```\nSo far what I understood is this \n1. It is due to `loss.backward()` call.\n2. Something has to be off either in the outputs of the models or targets to get this error. \n\nI am not sure what could go wrong at 30+ epoch and trigger the error. It is very unusual to see such an error.",
      "votes": null
    },
    {
      "id": "1307800",
      "postDate": "05/14/2021 17:11:22",
      "content": "<p>I get something similar on local machine and root cause is always memory issue.  Reduce your batch size and see if it will than run more than 30.</p>",
      "rawMarkdown": "I get something similar on local machine and root cause is always memory issue.  Reduce your batch size and see if it will than run more than 30.",
      "votes": null
    },
    {
      "id": "1308133",
      "postDate": "05/15/2021 03:19:13",
      "content": "<p>Hi, yeah I think it gets resolved by using small batch size. </p>",
      "rawMarkdown": "Hi, yeah I think it gets resolved by using small batch size.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1307800,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/14/2021 17:11:22",
      "content": "<p>I get something similar on local machine and root cause is always memory issue.  Reduce your batch size and see if it will than run more than 30.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1308133,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "05/15/2021 03:19:13",
          "content": "<p>Hi, yeah I think it gets resolved by using small batch size. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1307387": "Hi guys,\n\nI am running the model training on Kaggle. I am using Pytorch as my framework. The training process works well for say 30 epochs and sometimes more than that but after that suddenly it throws an error. Below is the trackback of it.\n\n```\n---------------------------------------------------------------------------\nRuntimeError                              Traceback (most recent call last)\n<ipython-input-26-7a1e5b630b5c> in <module>\n     51                 device=device,\n     52                 input_key=\"image\",\n---> 53                 input_target_key=\"targets\"\n     54             )\n     55             scheduler.step()\n\n<ipython-input-25-7bd46f9c870e> in train_one_epoch(model, dataloader, optimizer, criterion, device, input_key, input_target_key)\n     18         loss = criterion(outputs, y)\n     19         optimizer.zero_grad()\n---> 20         loss.backward()\n     21         optimizer.step()\n     22 \n\n/opt/conda/lib/python3.7/site-packages/torch/tensor.py in backward(self, gradient, retain_graph, create_graph)\n    219                 retain_graph=retain_graph,\n    220                 create_graph=create_graph)\n--> 221         torch.autograd.backward(self, gradient, retain_graph, create_graph)\n    222 \n    223     def register_hook(self, hook):\n\n/opt/conda/lib/python3.7/site-packages/torch/autograd/__init__.py in backward(tensors, grad_tensors, retain_graph, create_graph, grad_variables)\n    130     Variable._execution_engine.run_backward(\n    131         tensors, grad_tensors_, retain_graph, create_graph,\n--> 132         allow_unreachable=True)  # allow_unreachable flag\n    133 \n    134 \n\nRuntimeError: cuDNN error: CUDNN_STATUS_NOT_SUPPORTED. This error may appear if you passed in a non-contiguous input.\nYou can try to repro this exception using the following code snippet. If that doesn't trigger the error, please include your original repro script when reporting this issue.\n```\nSo far what I understood is this \n1. It is due to `loss.backward()` call.\n2. Something has to be off either in the outputs of the models or targets to get this error. \n\nI am not sure what could go wrong at 30+ epoch and trigger the error. It is very unusual to see such an error.",
    "1307800": "I get something similar on local machine and root cause is always memory issue.  Reduce your batch size and see if it will than run more than 30.",
    "1308133": "Hi, yeah I think it gets resolved by using small batch size."
  },
  "source": "meta"
}