{
  "id": 272279,
  "title": "NaN output with using torch.amp",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/272279",
  "author_name": "",
  "post_date": "2021-09-14T19:31:33.721473500Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Does anyone have a problem with NaN outputs with timm's efficientnet_bx at .eval() mode while using native torch AMP after several epoch?</p>",
  "messages": [
    {
      "id": "1513084",
      "postDate": "09/14/2021 19:31:33",
      "content": "<p>Does anyone have a problem with NaN outputs with timm's efficientnet_bx at .eval() mode while using native torch AMP after several epoch?</p>",
      "rawMarkdown": "Does anyone have a problem with NaN outputs with timm's efficientnet_bx at .eval() mode while using native torch AMP after several epoch?",
      "votes": null
    },
    {
      "id": "1513109",
      "postDate": "09/14/2021 20:22:24",
      "content": "<p>Your BatchNorm layers are being corrupted. Overwrite them with any working version and you will be fine.</p>\n<p><a href=\"https://discuss.pytorch.org/t/batch-norm-instability/32159\" target=\"_blank\">https://discuss.pytorch.org/t/batch-norm-instability/32159</a> - probably the same bug, but in his case without amp it \"just\" tanks the training completely.</p>\n<p>(I spent three days banging my head against this cryptic bug, and after seeing this post and realizing that it exists for others - i pin it in fifteen minutes… oh irony.)</p>",
      "rawMarkdown": "Your BatchNorm layers are being corrupted. Overwrite them with any working version and you will be fine.\n\nhttps://discuss.pytorch.org/t/batch-norm-instability/32159 - probably the same bug, but in his case without amp it \"just\" tanks the training completely.\n\n(I spent three days banging my head against this cryptic bug, and after seeing this post and realizing that it exists for others - i pin it in fifteen minutes... oh irony.)",
      "votes": null
    },
    {
      "id": "1513141",
      "postDate": "09/14/2021 21:28:14",
      "content": "<p>Thank you for your answer! What do you mean about overwriting BatchNorm? I don't know other realizations of batchnorm layers for pytorch. Did you mean changing parameters like track_running_stats=False after a while?</p>",
      "rawMarkdown": "Thank you for your answer! What do you mean about overwriting BatchNorm? I don't know other realizations of batchnorm layers for pytorch. Did you mean changing parameters like track_running_stats=False after a while?",
      "votes": null
    },
    {
      "id": "1513143",
      "postDate": "09/14/2021 21:37:51",
      "content": "<p>Literal overwriting. Create a new pretrained model (or load your own earlier checkpoint) and copy all BatchNorm layers from there to your corrupted model.</p>\n<p>I tested this on half the epoch worth of training, and after ~100 batches everything looked normal once more. No guarantees that accuracy would not suffer slightly in a long run, though.</p>",
      "rawMarkdown": "Literal overwriting. Create a new pretrained model (or load your own earlier checkpoint) and copy all BatchNorm layers from there to your corrupted model.\n\nI tested this on half the epoch worth of training, and after ~100 batches everything looked normal once more. No guarantees that accuracy would not suffer slightly in a long run, though.",
      "votes": null
    },
    {
      "id": "1513153",
      "postDate": "09/14/2021 21:58:44",
      "content": "<p>Thank you very much for sharing!</p>",
      "rawMarkdown": "Thank you very much for sharing!",
      "votes": null
    },
    {
      "id": "1559715",
      "postDate": "10/27/2021 07:08:55",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1513109,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "09/14/2021 20:22:24",
      "content": "<p>Your BatchNorm layers are being corrupted. Overwrite them with any working version and you will be fine.</p>\n<p><a href=\"https://discuss.pytorch.org/t/batch-norm-instability/32159\" target=\"_blank\">https://discuss.pytorch.org/t/batch-norm-instability/32159</a> - probably the same bug, but in his case without amp it \"just\" tanks the training completely.</p>\n<p>(I spent three days banging my head against this cryptic bug, and after seeing this post and realizing that it exists for others - i pin it in fifteen minutes… oh irony.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1513141,
          "author_name": "trytolose",
          "author_url": "",
          "post_date": "09/14/2021 21:28:14",
          "content": "<p>Thank you for your answer! What do you mean about overwriting BatchNorm? I don't know other realizations of batchnorm layers for pytorch. Did you mean changing parameters like track_running_stats=False after a while?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513143,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "09/14/2021 21:37:51",
          "content": "<p>Literal overwriting. Create a new pretrained model (or load your own earlier checkpoint) and copy all BatchNorm layers from there to your corrupted model.</p>\n<p>I tested this on half the epoch worth of training, and after ~100 batches everything looked normal once more. No guarantees that accuracy would not suffer slightly in a long run, though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513153,
          "author_name": "trytolose",
          "author_url": "",
          "post_date": "09/14/2021 21:58:44",
          "content": "<p>Thank you very much for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559715,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:08:55",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1513084": "Does anyone have a problem with NaN outputs with timm's efficientnet_bx at .eval() mode while using native torch AMP after several epoch?",
    "1513109": "Your BatchNorm layers are being corrupted. Overwrite them with any working version and you will be fine.\n\nhttps://discuss.pytorch.org/t/batch-norm-instability/32159 - probably the same bug, but in his case without amp it \"just\" tanks the training completely.\n\n(I spent three days banging my head against this cryptic bug, and after seeing this post and realizing that it exists for others - i pin it in fifteen minutes... oh irony.)",
    "1513141": "Thank you for your answer! What do you mean about overwriting BatchNorm? I don't know other realizations of batchnorm layers for pytorch. Did you mean changing parameters like track_running_stats=False after a while?",
    "1513143": "Literal overwriting. Create a new pretrained model (or load your own earlier checkpoint) and copy all BatchNorm layers from there to your corrupted model.\n\nI tested this on half the epoch worth of training, and after ~100 batches everything looked normal once more. No guarantees that accuracy would not suffer slightly in a long run, though.",
    "1513153": "Thank you very much for sharing!",
    "1559715": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}