{
  "id": 398769,
  "title": "CUDA error's SED model [solved]",
  "url": "/competitions/birdclef-2023/discussion/398769",
  "author_name": "Maximiliano Diaz Battan",
  "post_date": "2023-03-31T16:55:03.691000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In the last couple of days, I experience a lot of CUDA errors that I don't remember ever experiencing before on Kaggle notebooks, like the following 2: </p>\n<ul>\n<li><p>RuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling <code>cublasSgemm( handle, opa, opb, m, n, k, &amp;alpha, a, lda, b, ldb, &amp;beta, c, ldc)</code> </p></li>\n<li><p>RuntimeError: cuDNN error: CUDNN_STATUS_EXECUTION_FAILED. </p></li>\n</ul>\n<p>The thing is errors occur once the training it's started, in the beginning around 10 epochs, and now that it's reduced to 5 or less. I think could be a gradient problem, but I don't know, because I don't remember making any changes to the code, and previously I could train the SED model without any issues for 20 epochs or more. It's just me, or is the same thing happening to someone else?</p>\n<p>Solution: After a lot of reading in Pytorch forums, I was able to find the problem, apparently, is due to output inconsistencies, I solved it by clamping the values between 0. and 1., before the loss function.</p>\n<pre><code>class PANNsLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        self.bce = nn.BCELoss()\n\n    def forward(self, inputs, target):\n\n        input_ = inputs[\"clipwise_output\"]\n        input_ = torch.where(torch.isnan(input_), torch.zeros_like(input_), input_)\n        input_ = torch.where(torch.isinf(input_), torch.zeros_like(input_), input_)\n\n        input_ = torch.clamp(input_, min=0., max=1.)\n\n        return self.bce(input_, target)\n</code></pre>",
  "messages": [
    {
      "id": 2204547,
      "postDate": "2023-03-31T16:55:03.690Z",
      "content": "<p>In the last couple of days, I experience a lot of CUDA errors that I don't remember ever experiencing before on Kaggle notebooks, like the following 2: </p>\n<ul>\n<li><p>RuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling <code>cublasSgemm( handle, opa, opb, m, n, k, &amp;alpha, a, lda, b, ldb, &amp;beta, c, ldc)</code> </p></li>\n<li><p>RuntimeError: cuDNN error: CUDNN_STATUS_EXECUTION_FAILED. </p></li>\n</ul>\n<p>The thing is errors occur once the training it's started, in the beginning around 10 epochs, and now that it's reduced to 5 or less. I think could be a gradient problem, but I don't know, because I don't remember making any changes to the code, and previously I could train the SED model without any issues for 20 epochs or more. It's just me, or is the same thing happening to someone else?</p>\n<p>Solution: After a lot of reading in Pytorch forums, I was able to find the problem, apparently, is due to output inconsistencies, I solved it by clamping the values between 0. and 1., before the loss function.</p>\n<pre><code>class PANNsLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        self.bce = nn.BCELoss()\n\n    def forward(self, inputs, target):\n\n        input_ = inputs[\"clipwise_output\"]\n        input_ = torch.where(torch.isnan(input_), torch.zeros_like(input_), input_)\n        input_ = torch.where(torch.isinf(input_), torch.zeros_like(input_), input_)\n\n        input_ = torch.clamp(input_, min=0., max=1.)\n\n        return self.bce(input_, target)\n</code></pre>",
      "rawMarkdown": "In the last couple of days, I experience a lot of CUDA errors that I don't remember ever experiencing before on Kaggle notebooks, like the following 2: \n\n- RuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)` \n\n- RuntimeError: cuDNN error: CUDNN_STATUS_EXECUTION_FAILED. \n\nThe thing is errors occur once the training it's started, in the beginning around 10 epochs, and now that it's reduced to 5 or less. I think could be a gradient problem, but I don't know, because I don't remember making any changes to the code, and previously I could train the SED model without any issues for 20 epochs or more. It's just me, or is the same thing happening to someone else?\n\nSolution: After a lot of reading in Pytorch forums, I was able to find the problem, apparently, is due to output inconsistencies, I solved it by clamping the values between 0. and 1., before the loss function.\n\n~~~\nclass PANNsLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        self.bce = nn.BCELoss()\n\n    def forward(self, inputs, target):\n        \n        input_ = inputs[\"clipwise_output\"]\n        input_ = torch.where(torch.isnan(input_), torch.zeros_like(input_), input_)\n        input_ = torch.where(torch.isinf(input_), torch.zeros_like(input_), input_)\n        \n        input_ = torch.clamp(input_, min=0., max=1.)\n\n        return self.bce(input_, target)\n~~~\n\n",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2204547": "In the last couple of days, I experience a lot of CUDA errors that I don't remember ever experiencing before on Kaggle notebooks, like the following 2: \n\n- RuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)` \n\n- RuntimeError: cuDNN error: CUDNN_STATUS_EXECUTION_FAILED. \n\nThe thing is errors occur once the training it's started, in the beginning around 10 epochs, and now that it's reduced to 5 or less. I think could be a gradient problem, but I don't know, because I don't remember making any changes to the code, and previously I could train the SED model without any issues for 20 epochs or more. It's just me, or is the same thing happening to someone else?\n\nSolution: After a lot of reading in Pytorch forums, I was able to find the problem, apparently, is due to output inconsistencies, I solved it by clamping the values between 0. and 1., before the loss function.\n\n~~~\nclass PANNsLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        self.bce = nn.BCELoss()\n\n    def forward(self, inputs, target):\n        \n        input_ = inputs[\"clipwise_output\"]\n        input_ = torch.where(torch.isnan(input_), torch.zeros_like(input_), input_)\n        input_ = torch.where(torch.isinf(input_), torch.zeros_like(input_), input_)\n        \n        input_ = torch.clamp(input_, min=0., max=1.)\n\n        return self.bce(input_, target)\n~~~\n\n"
  }
}