{
  "id": 165488,
  "title": "ZeroDivisionError: float division by zero",
  "url": "/competitions/alaska2-image-steganalysis/discussion/165488",
  "author_name": "",
  "post_date": "2020-07-09T22:33:32.444525100Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Did anyone meet this error while training?\nI think this is caused by gradient explosion or vanishing.\nI setup larger eps and gradient clip but not helps</p>\n\n<p><code>\nTraceback (most recent call last):\n  File \"/mnt/ufs18/home-114/dunan/Learn/Kaggle/alaska2/kaggle_alaska2/train.py\", line 428, in &lt;module&gt;\n  File \"/mnt/ufs18/home-114/dunan/Learn/Kaggle/alaska2/kaggle_alaska2/train.py\", line 313, in train_epoch\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/contextlib.py\", line 119, in __exit__\n    next(self.gen)\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/handle.py\", line 123, in scale_loss\n    optimizer._post_amp_backward(loss_scaler)\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/_process_optimizer.py\", line 249, in post_backward_no_master_weights\n    post_backward_models_are_masters(scaler, params, stashed_grads)\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/_process_optimizer.py\", line 135, in post_backward_models_are_masters\n    scale_override=(grads_have_scale, stashed_have_scale, out_scale))\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/scaler.py\", line 176, in unscale_with_stashed\n    out_scale/grads_have_scale,   # 1./scale,\nZeroDivisionError: float division by zero\n</code></p>",
  "messages": [
    {
      "id": "922203",
      "postDate": "07/09/2020 22:33:32",
      "content": "<p>Did anyone meet this error while training?\nI think this is caused by gradient explosion or vanishing.\nI setup larger eps and gradient clip but not helps</p>\n\n<p><code>\nTraceback (most recent call last):\n  File \"/mnt/ufs18/home-114/dunan/Learn/Kaggle/alaska2/kaggle_alaska2/train.py\", line 428, in &lt;module&gt;\n  File \"/mnt/ufs18/home-114/dunan/Learn/Kaggle/alaska2/kaggle_alaska2/train.py\", line 313, in train_epoch\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/contextlib.py\", line 119, in __exit__\n    next(self.gen)\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/handle.py\", line 123, in scale_loss\n    optimizer._post_amp_backward(loss_scaler)\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/_process_optimizer.py\", line 249, in post_backward_no_master_weights\n    post_backward_models_are_masters(scaler, params, stashed_grads)\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/_process_optimizer.py\", line 135, in post_backward_models_are_masters\n    scale_override=(grads_have_scale, stashed_have_scale, out_scale))\n  File \"/mnt/home/dunan/anaconda3/envs/kaggle/lib/python3.7/site-packages/apex/amp/scaler.py\", line 176, in unscale_with_stashed\n    out_scale/grads_have_scale,   # 1./scale,\nZeroDivisionError: float division by zero\n</code></p>",
      "rawMarkdown": "Did anyone meet this error while training?\nI think this is caused by gradient explosion or vanishing.\nI setup larger eps and gradient clip but not helps\n\n```\nTraceback (most recent call last):\n  File \"/mnt/ufs18/home-114/dunan/Learn/Kaggle/alaska2/kaggle_alaska2/train.py\", line 428, in",
      "votes": null
    },
    {
      "id": "922391",
      "postDate": "07/10/2020 04:41:19",
      "content": "<p>Any suggestion?</p>",
      "rawMarkdown": "Any suggestion?",
      "votes": null
    },
    {
      "id": "924130",
      "postDate": "07/11/2020 09:13:18",
      "content": "<p><a href=\"/strideradu\">@strideradu</a> I encountered this error in metric calculation. I casted the values to float32 to solve the problem.</p>",
      "rawMarkdown": "strideradu I encountered this error in metric calculation. I casted the values to float32 to solve the problem.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 922391,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "07/10/2020 04:41:19",
      "content": "<p>Any suggestion?</p>",
      "votes": null,
      "replies": [
        {
          "id": 924130,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "07/11/2020 09:13:18",
          "content": "<p><a href=\"/strideradu\">@strideradu</a> I encountered this error in metric calculation. I casted the values to float32 to solve the problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "922203": "Did anyone meet this error while training?\nI think this is caused by gradient explosion or vanishing.\nI setup larger eps and gradient clip but not helps\n\n```\nTraceback (most recent call last):\n  File \"/mnt/ufs18/home-114/dunan/Learn/Kaggle/alaska2/kaggle_alaska2/train.py\", line 428, in",
    "922391": "Any suggestion?",
    "924130": "strideradu I encountered this error in metric calculation. I casted the values to float32 to solve the problem."
  },
  "source": "meta"
}