{
  "id": 316260,
  "title": "loos is nan...",
  "url": "/competitions/happy-whale-and-dolphin/discussion/316260",
  "author_name": "",
  "post_date": "2022-04-01T02:58:31.014310Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I ran into a problem, where the loss would be nan…</p>\n<p>Has anyone else encountered the same problem?<br>\nAlso, how did you solve the problem?</p>\n<p>・I am running Tensorflow on colab.<br>\n・batch size is 32<br>\n・learning rate 0.001</p>\n<p><br>\nI modified class Arc Margin as follows.<br>\n1 -&gt; 1.00000001<br>\n<code>sine = tf.math.sqrt(1.00000001 - tf.math.pow(cosine, 2))</code></p>",
  "messages": [
    {
      "id": "1741637",
      "postDate": "04/01/2022 02:58:31",
      "content": "<p>I ran into a problem, where the loss would be nan…</p>\n<p>Has anyone else encountered the same problem?<br>\nAlso, how did you solve the problem?</p>\n<p>・I am running Tensorflow on colab.<br>\n・batch size is 32<br>\n・learning rate 0.001</p>\n<p><br>\nI modified class Arc Margin as follows.<br>\n1 -&gt; 1.00000001<br>\n<code>sine = tf.math.sqrt(1.00000001 - tf.math.pow(cosine, 2))</code></p>",
      "rawMarkdown": "I ran into a problem, where the loss would be nan...\n\nHas anyone else encountered the same problem?\nAlso, how did you solve the problem?\n\n・I am running Tensorflow on colab.\n・batch size is 32\n・learning rate 0.001\n\n<What we have tried>\nI modified class Arc Margin as follows.\n1 -> 1.00000001\n`sine = tf.math.sqrt(1.00000001 - tf.math.pow(cosine, 2))`",
      "votes": null
    },
    {
      "id": "1741939",
      "postDate": "04/01/2022 10:04:22",
      "content": "<p>If loss starts nan from start - you have bug, if it is nan after several iterations - try to reduce lr (especially if you use Adam optimizer)</p>",
      "rawMarkdown": "If loss starts nan from start - you have bug, if it is nan after several iterations - try to reduce lr (especially if you use Adam optimizer)",
      "votes": null
    },
    {
      "id": "1742053",
      "postDate": "04/01/2022 12:20:34",
      "content": "<p>Getting a NaN loss is a well known problem when training neural networks, which is mostly caused by wrong optimizer configurations (exploding/vanishing gradients). I would recommend using the SGD optimizer with a small learning rate and gradient clipping to exclude the optimizer as problem source:</p>\n<pre><code>tf.keras.optimizers.SGD(\n    learning_rate=1e-4,\n    momentum=0.9,\n    clipnorm=5.0,\n)\n</code></pre>\n<p>Out of curiosity, what did you try to achieve with your modification?</p>",
      "rawMarkdown": "Getting a NaN loss is a well known problem when training neural networks, which is mostly caused by wrong optimizer configurations (exploding/vanishing gradients). I would recommend using the SGD optimizer with a small learning rate and gradient clipping to exclude the optimizer as problem source:\n\n```\ntf.keras.optimizers.SGD(\n    learning_rate=1e-4,\n    momentum=0.9,\n    clipnorm=5.0,\n)\n```\n\nOut of curiosity, what did you try to achieve with your modification?",
      "votes": null
    },
    {
      "id": "1742232",
      "postDate": "04/01/2022 15:43:42",
      "content": "<p>Check the number of classes.  It should be 15,587.  If the number of classes is not correct you can get a loss of Nan.</p>",
      "rawMarkdown": "Check the number of classes.  It should be 15,587.  If the number of classes is not correct you can get a loss of Nan.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1741939,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "04/01/2022 10:04:22",
      "content": "<p>If loss starts nan from start - you have bug, if it is nan after several iterations - try to reduce lr (especially if you use Adam optimizer)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1742053,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "04/01/2022 12:20:34",
      "content": "<p>Getting a NaN loss is a well known problem when training neural networks, which is mostly caused by wrong optimizer configurations (exploding/vanishing gradients). I would recommend using the SGD optimizer with a small learning rate and gradient clipping to exclude the optimizer as problem source:</p>\n<pre><code>tf.keras.optimizers.SGD(\n    learning_rate=1e-4,\n    momentum=0.9,\n    clipnorm=5.0,\n)\n</code></pre>\n<p>Out of curiosity, what did you try to achieve with your modification?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1742232,
      "author_name": "ufosoftwarellc",
      "author_url": "",
      "post_date": "04/01/2022 15:43:42",
      "content": "<p>Check the number of classes.  It should be 15,587.  If the number of classes is not correct you can get a loss of Nan.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1741637": "I ran into a problem, where the loss would be nan...\n\nHas anyone else encountered the same problem?\nAlso, how did you solve the problem?\n\n・I am running Tensorflow on colab.\n・batch size is 32\n・learning rate 0.001\n\n<What we have tried>\nI modified class Arc Margin as follows.\n1 -> 1.00000001\n`sine = tf.math.sqrt(1.00000001 - tf.math.pow(cosine, 2))`",
    "1741939": "If loss starts nan from start - you have bug, if it is nan after several iterations - try to reduce lr (especially if you use Adam optimizer)",
    "1742053": "Getting a NaN loss is a well known problem when training neural networks, which is mostly caused by wrong optimizer configurations (exploding/vanishing gradients). I would recommend using the SGD optimizer with a small learning rate and gradient clipping to exclude the optimizer as problem source:\n\n```\ntf.keras.optimizers.SGD(\n    learning_rate=1e-4,\n    momentum=0.9,\n    clipnorm=5.0,\n)\n```\n\nOut of curiosity, what did you try to achieve with your modification?",
    "1742232": "Check the number of classes.  It should be 15,587.  If the number of classes is not correct you can get a loss of Nan."
  },
  "source": "meta"
}