{
  "id": 122569,
  "title": "Why this loss gives NAN",
  "url": "/competitions/pku-autonomous-driving/discussion/122569",
  "author_name": "",
  "post_date": "2019-12-21T03:17:19.358289200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I took it below loss from Phoenix notebook .\n<code>mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)</code>\nThis gives a high value in thousands although comes down fast but after few iters  it ends in nan.\nCompared pytorch standard f.bce with logits which dsnt gives nan and also gives less than 1\nDo you see any trouble ? I use mixed fp</p>",
  "messages": [
    {
      "id": "699861",
      "postDate": "12/21/2019 03:17:19",
      "content": "<p>I took it below loss from Phoenix notebook .\n<code>mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)</code>\nThis gives a high value in thousands although comes down fast but after few iters  it ends in nan.\nCompared pytorch standard f.bce with logits which dsnt gives nan and also gives less than 1\nDo you see any trouble ? I use mixed fp</p>",
      "rawMarkdown": "I took it below loss from Phoenix notebook .\n`mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)`\nThis gives a high value in thousands although comes down fast but after few iters  it ends in nan.\nCompared pytorch standard f.bce with logits which dsnt gives nan and also gives less than 1\nDo you see any trouble ? I use mixed fp",
      "votes": null
    },
    {
      "id": "699977",
      "postDate": "12/21/2019 07:59:10",
      "content": "<p>Mixed fp can suffer from numerical stability problems. <code>1e-12</code> can not be represented in half precision so that mixed fp might lead <code>pred_mask + 1e-12</code> to 0. Finally, the <code>torch.log(0)</code> would give the nan error.</p>",
      "rawMarkdown": "Mixed fp can suffer from numerical stability problems. `1e-12` can not be represented in half precision so that mixed fp might lead `pred_mask + 1e-12` to 0. Finally, the `torch.log(0)` would give the nan error.",
      "votes": null
    },
    {
      "id": "699999",
      "postDate": "12/21/2019 08:38:51",
      "content": "<p>@syoya thanks\n1)i checked tensor(1e-7).half() returns a non zero value so same should be the case every time\n2)why this loss result dsnt matches to that f.bce where i get  0.5 to 1 but this give more than 1000,2000</p>",
      "rawMarkdown": "syoya thanks\n1)i checked tensor(1e-7).half() returns a non zero value so same should be the case every time\n2)why this loss result dsnt matches to that f.bce where i get  0.5 to 1 but this give more than 1000,2000",
      "votes": null
    },
    {
      "id": "716109",
      "postDate": "01/11/2020 09:09:06",
      "content": "<p>I also experienced instability when using the hand-written bce, seems like torch's BCEWithLogits work better. The magnitude is still similar to the hand-written loss in the kernel. I do still run into issues while I try to use mixed precision, kinda giving up on it...</p>",
      "rawMarkdown": "I also experienced instability when using the hand-written bce, seems like torch's BCEWithLogits work better. The magnitude is still similar to the hand-written loss in the kernel. I do still run into issues while I try to use mixed precision, kinda giving up on it...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 699977,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "12/21/2019 07:59:10",
      "content": "<p>Mixed fp can suffer from numerical stability problems. <code>1e-12</code> can not be represented in half precision so that mixed fp might lead <code>pred_mask + 1e-12</code> to 0. Finally, the <code>torch.log(0)</code> would give the nan error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 699999,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/21/2019 08:38:51",
          "content": "<p>@syoya thanks\n1)i checked tensor(1e-7).half() returns a non zero value so same should be the case every time\n2)why this loss result dsnt matches to that f.bce where i get  0.5 to 1 but this give more than 1000,2000</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 716109,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "01/11/2020 09:09:06",
      "content": "<p>I also experienced instability when using the hand-written bce, seems like torch's BCEWithLogits work better. The magnitude is still similar to the hand-written loss in the kernel. I do still run into issues while I try to use mixed precision, kinda giving up on it...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "699861": "I took it below loss from Phoenix notebook .\n`mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)`\nThis gives a high value in thousands although comes down fast but after few iters  it ends in nan.\nCompared pytorch standard f.bce with logits which dsnt gives nan and also gives less than 1\nDo you see any trouble ? I use mixed fp",
    "699977": "Mixed fp can suffer from numerical stability problems. `1e-12` can not be represented in half precision so that mixed fp might lead `pred_mask + 1e-12` to 0. Finally, the `torch.log(0)` would give the nan error.",
    "699999": "syoya thanks\n1)i checked tensor(1e-7).half() returns a non zero value so same should be the case every time\n2)why this loss result dsnt matches to that f.bce where i get  0.5 to 1 but this give more than 1000,2000",
    "716109": "I also experienced instability when using the hand-written bce, seems like torch's BCEWithLogits work better. The magnitude is still similar to the hand-written loss in the kernel. I do still run into issues while I try to use mixed precision, kinda giving up on it..."
  },
  "source": "meta"
}