{
  "id": 229475,
  "title": "Gradient overflow.  Skipping step, and result validation loss Nan",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/229475",
  "author_name": "",
  "post_date": "2021-03-30T10:28:58.352469300Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am trying use bce + dice loss, but there's an issue for me. I use bce+ dice in train stage, and bce loss in validation stage, then after i train my model 30 epochs, it suddenly appear train gradient overfolw, and valid loss Nan, i know it because of i use mix-precision, and gradient overfolw, but i don't know how to fix it, it's normal when i just use bce loss in train. Anyone know how to fix it?</p>\n<p>Epoch: [39][78/131]     Time 0.738 (0.654)      Speed 173.551 (195.654) Loss 0.1971025020 (0.2156)     <br>\nEpoch: [39][91/131]     Time 0.747 (0.666)      Speed 171.303 (192.238) Loss 0.2891663313 (0.2248)    <br>\nEpoch: [39][104/131]    Time 0.738 (0.674)      Speed 173.395 (189.944) Loss 0.2159770578 (0.2238)      <br>\nEpoch: [39][117/131]    Time 0.723 (0.679)      Speed 177.123 (188.579) Loss 0.2206496298 (0.2235)      <br>\nGradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625</p>\n<p>Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625<br>\nEpoch: [39][130/131]    Time 0.694 (0.680)      Speed 184.342 (188.186) Loss 0.2696546316 (0.2250)     </p>\n<p>Test: [0/52]    Time 0.700 (0.700)      Speed 182.733 (182.733) Loss nan (nan)<br>\nTest: [13/52]   Time 0.183 (0.230)      Speed 701.031 (555.920) Loss nan (nan)<br>\nTest: [26/52]   Time 0.178 (0.209)      Speed 718.416 (612.396) Loss nan (nan)<br>\nTest: [39/52]   Time 0.197 (0.203)      Speed 650.795 (630.060) Loss nan (nan)<br>\nTest: [51/52]   Time 0.116 (0.198)      Speed 1098.909 (647.312)        Loss nan (nan)</p>\n<p>this is my bce + dice loss code:</p>\n<pre><code>class dice_bce(nn.Module):\n    def __init__(self, weights=None, reduction='mean'):\n        super().__init__()\n        self.dice_loss = SoftDiceLossV1(reduction=reduction)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction=reduction, weight=weights)\n\n    def forward(self, logits, target):\n        loss = self.dice_loss(logits.squeeze(1), target.squeeze(1).long()) + self.bce_loss(logits, target)\n        return loss\n\nclass SoftDiceLossV1(nn.Module):\n    '''\n    soft-dice loss, useful in binary segmentation\n    '''\n    def __init__(self,\n                 p=1,\n                 smooth=1,\n                 reduction='mean'):\n        super(SoftDiceLossV1, self).__init__()\n        self.p = p\n        self.smooth = smooth\n        self.reduction = reduction\n\n    def forward(self, logits, labels):\n        '''\n        args: logits: tensor of shape (N, H, W)\n        args: label: tensor of shape(N, H, W)\n        '''\n        probs = torch.sigmoid(logits)\n        numer = (probs * labels).sum(dim=(1, 2))\n        denor = (probs.pow(self.p) + labels).sum(dim=(1, 2))\n        loss = 1. - (2 * numer + self.smooth) / (denor + self.smooth)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n        return loss\n</code></pre>",
  "messages": [
    {
      "id": "1256881",
      "postDate": "03/30/2021 10:28:58",
      "content": "<p>I am trying use bce + dice loss, but there's an issue for me. I use bce+ dice in train stage, and bce loss in validation stage, then after i train my model 30 epochs, it suddenly appear train gradient overfolw, and valid loss Nan, i know it because of i use mix-precision, and gradient overfolw, but i don't know how to fix it, it's normal when i just use bce loss in train. Anyone know how to fix it?</p>\n<p>Epoch: [39][78/131]     Time 0.738 (0.654)      Speed 173.551 (195.654) Loss 0.1971025020 (0.2156)     <br>\nEpoch: [39][91/131]     Time 0.747 (0.666)      Speed 171.303 (192.238) Loss 0.2891663313 (0.2248)    <br>\nEpoch: [39][104/131]    Time 0.738 (0.674)      Speed 173.395 (189.944) Loss 0.2159770578 (0.2238)      <br>\nEpoch: [39][117/131]    Time 0.723 (0.679)      Speed 177.123 (188.579) Loss 0.2206496298 (0.2235)      <br>\nGradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625</p>\n<p>Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625<br>\nEpoch: [39][130/131]    Time 0.694 (0.680)      Speed 184.342 (188.186) Loss 0.2696546316 (0.2250)     </p>\n<p>Test: [0/52]    Time 0.700 (0.700)      Speed 182.733 (182.733) Loss nan (nan)<br>\nTest: [13/52]   Time 0.183 (0.230)      Speed 701.031 (555.920) Loss nan (nan)<br>\nTest: [26/52]   Time 0.178 (0.209)      Speed 718.416 (612.396) Loss nan (nan)<br>\nTest: [39/52]   Time 0.197 (0.203)      Speed 650.795 (630.060) Loss nan (nan)<br>\nTest: [51/52]   Time 0.116 (0.198)      Speed 1098.909 (647.312)        Loss nan (nan)</p>\n<p>this is my bce + dice loss code:</p>\n<pre><code>class dice_bce(nn.Module):\n    def __init__(self, weights=None, reduction='mean'):\n        super().__init__()\n        self.dice_loss = SoftDiceLossV1(reduction=reduction)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction=reduction, weight=weights)\n\n    def forward(self, logits, target):\n        loss = self.dice_loss(logits.squeeze(1), target.squeeze(1).long()) + self.bce_loss(logits, target)\n        return loss\n\nclass SoftDiceLossV1(nn.Module):\n    '''\n    soft-dice loss, useful in binary segmentation\n    '''\n    def __init__(self,\n                 p=1,\n                 smooth=1,\n                 reduction='mean'):\n        super(SoftDiceLossV1, self).__init__()\n        self.p = p\n        self.smooth = smooth\n        self.reduction = reduction\n\n    def forward(self, logits, labels):\n        '''\n        args: logits: tensor of shape (N, H, W)\n        args: label: tensor of shape(N, H, W)\n        '''\n        probs = torch.sigmoid(logits)\n        numer = (probs * labels).sum(dim=(1, 2))\n        denor = (probs.pow(self.p) + labels).sum(dim=(1, 2))\n        loss = 1. - (2 * numer + self.smooth) / (denor + self.smooth)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n        return loss\n</code></pre>",
      "rawMarkdown": "I am trying use bce + dice loss, but there's an issue for me. I use bce+ dice in train stage, and bce loss in validation stage, then after i train my model 30 epochs, it suddenly appear train gradient overfolw, and valid loss Nan, i know it because of i use mix-precision, and gradient overfolw, but i don't know how to fix it, it's normal when i just use bce loss in train. Anyone know how to fix it?\n\nEpoch: [39][78/131]     Time 0.738 (0.654)      Speed 173.551 (195.654) Loss 0.1971025020 (0.2156)     \nEpoch: [39][91/131]     Time 0.747 (0.666)      Speed 171.303 (192.238) Loss 0.2891663313 (0.2248)    \nEpoch: [39][104/131]    Time 0.738 (0.674)      Speed 173.395 (189.944) Loss 0.2159770578 (0.2238)      \nEpoch: [39][117/131]    Time 0.723 (0.679)      Speed 177.123 (188.579) Loss 0.2206496298 (0.2235)      \nGradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625\n\nGradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625\nEpoch: [39][130/131]    Time 0.694 (0.680)      Speed 184.342 (188.186) Loss 0.2696546316 (0.2250)     \n\nTest: [0/52]    Time 0.700 (0.700)      Speed 182.733 (182.733) Loss nan (nan)\nTest: [13/52]   Time 0.183 (0.230)      Speed 701.031 (555.920) Loss nan (nan)\nTest: [26/52]   Time 0.178 (0.209)      Speed 718.416 (612.396) Loss nan (nan)\nTest: [39/52]   Time 0.197 (0.203)      Speed 650.795 (630.060) Loss nan (nan)\nTest: [51/52]   Time 0.116 (0.198)      Speed 1098.909 (647.312)        Loss nan (nan)\n\nthis is my bce + dice loss code:\n\n```\nclass dice_bce(nn.Module):\n    def __init__(self, weights=None, reduction='mean'):\n        super().__init__()\n        self.dice_loss = SoftDiceLossV1(reduction=reduction)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction=reduction, weight=weights)\n\n    def forward(self, logits, target):\n        loss = self.dice_loss(logits.squeeze(1), target.squeeze(1).long()) + self.bce_loss(logits, target)\n        return loss\n\nclass SoftDiceLossV1(nn.Module):\n    '''\n    soft-dice loss, useful in binary segmentation\n    '''\n    def __init__(self,\n                 p=1,\n                 smooth=1,\n                 reduction='mean'):\n        super(SoftDiceLossV1, self).__init__()\n        self.p = p\n        self.smooth = smooth\n        self.reduction = reduction\n\n    def forward(self, logits, labels):\n        '''\n        args: logits: tensor of shape (N, H, W)\n        args: label: tensor of shape(N, H, W)\n        '''\n        probs = torch.sigmoid(logits)\n        numer = (probs * labels).sum(dim=(1, 2))\n        denor = (probs.pow(self.p) + labels).sum(dim=(1, 2))\n        loss = 1. - (2 * numer + self.smooth) / (denor + self.smooth)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n        return loss\n\n\n```",
      "votes": null
    },
    {
      "id": "1256912",
      "postDate": "03/30/2021 11:06:02",
      "content": "<p><a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> Try using gradient clipping with <code>torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm)</code>. </p>\n<p>You can refer to the original documentation here <a href=\"url\" target=\"_blank\">https://pytorch.org/docs/stable/notes/amp_examples.html</a></p>",
      "rawMarkdown": "cswwp347724 Try using gradient clipping with `torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm)`. \n\nYou can refer to the original documentation here [https://pytorch.org/docs/stable/notes/amp_examples.html](url)",
      "votes": null
    },
    {
      "id": "1256967",
      "postDate": "03/30/2021 12:20:08",
      "content": "<p><a href=\"https://www.kaggle.com/yovinyahathugoda\" target=\"_blank\">@yovinyahathugoda</a> Thank you, i am checking clip_grad_norm_</p>",
      "rawMarkdown": "yovinyahathugoda Thank you, i am checking clip_grad_norm_",
      "votes": null
    },
    {
      "id": "1257038",
      "postDate": "03/30/2021 13:26:54",
      "content": "<p>Another question: if i use torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm). how can i set max_norm value? just depend on experience or some rules? <a href=\"https://www.kaggle.com/yovinyahathugoda\" target=\"_blank\">@yovinyahathugoda</a> </p>",
      "rawMarkdown": "Another question: if i use torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm). how can i set max_norm value? just depend on experience or some rules? @yovinyahathugoda",
      "votes": null
    },
    {
      "id": "1257113",
      "postDate": "03/30/2021 14:34:10",
      "content": "<p><a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> I usually use values such as 1.0 or 2.0 for <code>max_norm</code> with mixed precision. I'm not sure if there is a specific rule as to how this value should be calculated. Depending on the training parameters such as learning rate and batch size, I think you can go up to values such as 5.0.</p>",
      "rawMarkdown": "cswwp347724 I usually use values such as 1.0 or 2.0 for `max_norm` with mixed precision. I'm not sure if there is a specific rule as to how this value should be calculated. Depending on the training parameters such as learning rate and batch size, I think you can go up to values such as 5.0.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1256912,
      "author_name": "yovinyahathugoda",
      "author_url": "",
      "post_date": "03/30/2021 11:06:02",
      "content": "<p><a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> Try using gradient clipping with <code>torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm)</code>. </p>\n<p>You can refer to the original documentation here <a href=\"url\" target=\"_blank\">https://pytorch.org/docs/stable/notes/amp_examples.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1256967,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "03/30/2021 12:20:08",
          "content": "<p><a href=\"https://www.kaggle.com/yovinyahathugoda\" target=\"_blank\">@yovinyahathugoda</a> Thank you, i am checking clip_grad_norm_</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1257038,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "03/30/2021 13:26:54",
          "content": "<p>Another question: if i use torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm). how can i set max_norm value? just depend on experience or some rules? <a href=\"https://www.kaggle.com/yovinyahathugoda\" target=\"_blank\">@yovinyahathugoda</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1257113,
          "author_name": "yovinyahathugoda",
          "author_url": "",
          "post_date": "03/30/2021 14:34:10",
          "content": "<p><a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> I usually use values such as 1.0 or 2.0 for <code>max_norm</code> with mixed precision. I'm not sure if there is a specific rule as to how this value should be calculated. Depending on the training parameters such as learning rate and batch size, I think you can go up to values such as 5.0.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1256881": "I am trying use bce + dice loss, but there's an issue for me. I use bce+ dice in train stage, and bce loss in validation stage, then after i train my model 30 epochs, it suddenly appear train gradient overfolw, and valid loss Nan, i know it because of i use mix-precision, and gradient overfolw, but i don't know how to fix it, it's normal when i just use bce loss in train. Anyone know how to fix it?\n\nEpoch: [39][78/131]     Time 0.738 (0.654)      Speed 173.551 (195.654) Loss 0.1971025020 (0.2156)     \nEpoch: [39][91/131]     Time 0.747 (0.666)      Speed 171.303 (192.238) Loss 0.2891663313 (0.2248)    \nEpoch: [39][104/131]    Time 0.738 (0.674)      Speed 173.395 (189.944) Loss 0.2159770578 (0.2238)      \nEpoch: [39][117/131]    Time 0.723 (0.679)      Speed 177.123 (188.579) Loss 0.2206496298 (0.2235)      \nGradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625\n\nGradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625Gradient overflow.  Skipping step, loss scaler 0 reducing loss scale to 0.0625\nEpoch: [39][130/131]    Time 0.694 (0.680)      Speed 184.342 (188.186) Loss 0.2696546316 (0.2250)     \n\nTest: [0/52]    Time 0.700 (0.700)      Speed 182.733 (182.733) Loss nan (nan)\nTest: [13/52]   Time 0.183 (0.230)      Speed 701.031 (555.920) Loss nan (nan)\nTest: [26/52]   Time 0.178 (0.209)      Speed 718.416 (612.396) Loss nan (nan)\nTest: [39/52]   Time 0.197 (0.203)      Speed 650.795 (630.060) Loss nan (nan)\nTest: [51/52]   Time 0.116 (0.198)      Speed 1098.909 (647.312)        Loss nan (nan)\n\nthis is my bce + dice loss code:\n\n```\nclass dice_bce(nn.Module):\n    def __init__(self, weights=None, reduction='mean'):\n        super().__init__()\n        self.dice_loss = SoftDiceLossV1(reduction=reduction)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction=reduction, weight=weights)\n\n    def forward(self, logits, target):\n        loss = self.dice_loss(logits.squeeze(1), target.squeeze(1).long()) + self.bce_loss(logits, target)\n        return loss\n\nclass SoftDiceLossV1(nn.Module):\n    '''\n    soft-dice loss, useful in binary segmentation\n    '''\n    def __init__(self,\n                 p=1,\n                 smooth=1,\n                 reduction='mean'):\n        super(SoftDiceLossV1, self).__init__()\n        self.p = p\n        self.smooth = smooth\n        self.reduction = reduction\n\n    def forward(self, logits, labels):\n        '''\n        args: logits: tensor of shape (N, H, W)\n        args: label: tensor of shape(N, H, W)\n        '''\n        probs = torch.sigmoid(logits)\n        numer = (probs * labels).sum(dim=(1, 2))\n        denor = (probs.pow(self.p) + labels).sum(dim=(1, 2))\n        loss = 1. - (2 * numer + self.smooth) / (denor + self.smooth)\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n        return loss\n\n\n```",
    "1256912": "cswwp347724 Try using gradient clipping with `torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm)`. \n\nYou can refer to the original documentation here [https://pytorch.org/docs/stable/notes/amp_examples.html](url)",
    "1256967": "yovinyahathugoda Thank you, i am checking clip_grad_norm_",
    "1257038": "Another question: if i use torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm). how can i set max_norm value? just depend on experience or some rules? @yovinyahathugoda",
    "1257113": "cswwp347724 I usually use values such as 1.0 or 2.0 for `max_norm` with mixed precision. I'm not sure if there is a specific rule as to how this value should be calculated. Depending on the training parameters such as learning rate and batch size, I think you can go up to values such as 5.0."
  },
  "source": "meta"
}