{
  "id": 165338,
  "title": "ValueError: Input contains NaN, infinity or a value too large for dtype('float32')",
  "url": "/competitions/alaska2-image-steganalysis/discussion/165338",
  "author_name": "",
  "post_date": "2020-07-09T10:30:39.910684900Z",
  "votes": 8,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Did anyone meet this error while training?\nWe believe it's caused by NaN value in logits and CNN weights but didn't know why.\nWe've tried clip gradiant but didn't help.</p>\n\n<p><code>\n  File \"train.py\", line 247, in alaska_weighted_auc\n    fpr, tpr, thresholds = metrics.roc_curve(y_true, y_valid, pos_label=1)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 771, in roc_curve\n    y_true, y_score, pos_label=pos_label, sample_weight=sample_weight)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 542, in _binary_clf_curve\n    assert_all_finite(y_score)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 77, in assert_all_finite\n    _assert_all_finite(X.data if sp.issparse(X) else X, allow_nan)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 60, in _assert_all_finite\n    msg_dtype if msg_dtype is not None else X.dtype)\nValueError: Input contains NaN, infinity or a value too large for dtype('float32').\n</code></p>",
  "messages": [
    {
      "id": "921492",
      "postDate": "07/09/2020 10:30:39",
      "content": "<p>Did anyone meet this error while training?\nWe believe it's caused by NaN value in logits and CNN weights but didn't know why.\nWe've tried clip gradiant but didn't help.</p>\n\n<p><code>\n  File \"train.py\", line 247, in alaska_weighted_auc\n    fpr, tpr, thresholds = metrics.roc_curve(y_true, y_valid, pos_label=1)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 771, in roc_curve\n    y_true, y_score, pos_label=pos_label, sample_weight=sample_weight)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 542, in _binary_clf_curve\n    assert_all_finite(y_score)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 77, in assert_all_finite\n    _assert_all_finite(X.data if sp.issparse(X) else X, allow_nan)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 60, in _assert_all_finite\n    msg_dtype if msg_dtype is not None else X.dtype)\nValueError: Input contains NaN, infinity or a value too large for dtype('float32').\n</code></p>",
      "rawMarkdown": "Did anyone meet this error while training?\nWe believe it's caused by NaN value in logits and CNN weights but didn't know why.\nWe've tried clip gradiant but didn't help.\n\n```\n  File \"train.py\", line 247, in alaska_weighted_auc\n    fpr, tpr, thresholds = metrics.roc_curve(y_true, y_valid, pos_label=1)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 771, in roc_curve\n    y_true, y_score, pos_label=pos_label, sample_weight=sample_weight)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 542, in _binary_clf_curve\n    assert_all_finite(y_score)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 77, in assert_all_finite\n    _assert_all_finite(X.data if sp.issparse(X) else X, allow_nan)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 60, in _assert_all_finite\n    msg_dtype if msg_dtype is not None else X.dtype)\nValueError: Input contains NaN, infinity or a value too large for dtype('float32').\n```",
      "votes": null
    },
    {
      "id": "921527",
      "postDate": "07/09/2020 11:01:44",
      "content": "<p>yes,   when i trained with fp16, some models crash like that. \nHowever, some models are fine. For example, efficientnetb2 is ok,  effcientnetb0 will comes to NAN. </p>",
      "rawMarkdown": "yes,   when i trained with fp16, some models crash like that. \nHowever, some models are fine. For example, efficientnetb2 is ok,  effcientnetb0 will comes to NAN.",
      "votes": null
    },
    {
      "id": "921556",
      "postDate": "07/09/2020 11:28:05",
      "content": "<p>I occasionally get NANs as predictions and replace those accordingly.</p>",
      "rawMarkdown": "I occasionally get NANs as predictions and replace those accordingly.",
      "votes": null
    },
    {
      "id": "921568",
      "postDate": "07/09/2020 11:43:38",
      "content": "<p>When training EfficientNet-B7 with mixed precision, I would start getting NaN loss. Couldn't figure it out, so just gave up.</p>",
      "rawMarkdown": "When training EfficientNet-B7 with mixed precision, I would start getting NaN loss. Couldn't figure it out, so just gave up.",
      "votes": null
    },
    {
      "id": "921619",
      "postDate": "07/09/2020 12:30:19",
      "content": "<p>I also got an error when using fp16.\nI solved it by increasing the eps value.\n<code>\ntorch.optim.AdamW(self.model.parameters(), lr=0.001, eps=1e-4)\n</code></p>",
      "rawMarkdown": "I also got an error when using fp16.\nI solved it by increasing the eps value.\n```\ntorch.optim.AdamW(self.model.parameters(), lr=0.001, eps=1e-4)\n```",
      "votes": null
    },
    {
      "id": "921757",
      "postDate": "07/09/2020 14:24:26",
      "content": "<p>Thanks! I'll try it.</p>",
      "rawMarkdown": "Thanks! I'll try it.",
      "votes": null
    },
    {
      "id": "922475",
      "postDate": "07/10/2020 05:59:03",
      "content": "<p>It still crashes occasionally... 🤕 \nThat's harsh...</p>",
      "rawMarkdown": "It still crashes occasionally... 🤕 \nThat's harsh...",
      "votes": null
    },
    {
      "id": "922504",
      "postDate": "07/10/2020 06:40:06",
      "content": "<p>There is one full black image in the data that gave me troubles before, just another random idea.</p>",
      "rawMarkdown": "There is one full black image in the data that gave me troubles before, just another random idea.",
      "votes": null
    },
    {
      "id": "922625",
      "postDate": "07/10/2020 08:37:21",
      "content": "<p>Thanks! We'll have a try.</p>",
      "rawMarkdown": "Thanks! We'll have a try.",
      "votes": null
    },
    {
      "id": "922733",
      "postDate": "07/10/2020 09:55:23",
      "content": "<p>EfficientNet larger than b5 gives me the same error in validating (interestingly I haven't got this in training),  but it happens randomly and only occasionally, so I just ignored these batches.</p>",
      "rawMarkdown": "EfficientNet larger than b5 gives me the same error in validating (interestingly I haven't got this in training),  but it happens randomly and only occasionally, so I just ignored these batches.",
      "votes": null
    },
    {
      "id": "926228",
      "postDate": "07/12/2020 15:32:32",
      "content": "<p>When we adding eps=1e-4, clip grad = 0.5, and avoiding all-zero input, it still crashed occasionally but the probability is much lower than before.\nThanks for you guys help!</p>",
      "rawMarkdown": "When we adding eps=1e-4, clip grad = 0.5, and avoiding all-zero input, it still crashed occasionally but the probability is much lower than before.\nThanks for you guys help!",
      "votes": null
    },
    {
      "id": "926745",
      "postDate": "07/13/2020 00:44:46",
      "content": "<p>this don't the problem, but suppress the error so that the training can go on:</p>\n\n<p><code>\nprobability[np.isnan(probability)]=0\n</code></p>",
      "rawMarkdown": "this don't the problem, but suppress the error so that the training can go on:\n\n```\nprobability[np.isnan(probability)]=0\n```",
      "votes": null
    },
    {
      "id": "931752",
      "postDate": "07/16/2020 12:26:02",
      "content": "<p>i find one image that causes NAN</p>\n\n<p>0670.jpg (test)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2dfbb94c4987745980533fe9c10e6260%2F0670.jpg?generation=1594902348609780&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i find one image that causes NAN\n\n0670.jpg (test)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2dfbb94c4987745980533fe9c10e6260%2F0670.jpg?generation=1594902348609780&amp;alt=media)",
      "votes": null
    },
    {
      "id": "932104",
      "postDate": "07/16/2020 17:41:00",
      "content": "<p>Interesting. But why it cause Nan?</p>",
      "rawMarkdown": "Interesting. But why it cause Nan?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 921527,
      "author_name": "cooolz",
      "author_url": "",
      "post_date": "07/09/2020 11:01:44",
      "content": "<p>yes,   when i trained with fp16, some models crash like that. \nHowever, some models are fine. For example, efficientnetb2 is ok,  effcientnetb0 will comes to NAN. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 921556,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "07/09/2020 11:28:05",
      "content": "<p>I occasionally get NANs as predictions and replace those accordingly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 921568,
      "author_name": "vaillant",
      "author_url": "",
      "post_date": "07/09/2020 11:43:38",
      "content": "<p>When training EfficientNet-B7 with mixed precision, I would start getting NaN loss. Couldn't figure it out, so just gave up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 921619,
      "author_name": "kani23",
      "author_url": "",
      "post_date": "07/09/2020 12:30:19",
      "content": "<p>I also got an error when using fp16.\nI solved it by increasing the eps value.\n<code>\ntorch.optim.AdamW(self.model.parameters(), lr=0.001, eps=1e-4)\n</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 921757,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "07/09/2020 14:24:26",
          "content": "<p>Thanks! I'll try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922475,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "07/10/2020 05:59:03",
          "content": "<p>It still crashes occasionally... 🤕 \nThat's harsh...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922504,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/10/2020 06:40:06",
          "content": "<p>There is one full black image in the data that gave me troubles before, just another random idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922625,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "07/10/2020 08:37:21",
          "content": "<p>Thanks! We'll have a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926228,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "07/12/2020 15:32:32",
          "content": "<p>When we adding eps=1e-4, clip grad = 0.5, and avoiding all-zero input, it still crashed occasionally but the probability is much lower than before.\nThanks for you guys help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 922733,
      "author_name": "hawkey",
      "author_url": "",
      "post_date": "07/10/2020 09:55:23",
      "content": "<p>EfficientNet larger than b5 gives me the same error in validating (interestingly I haven't got this in training),  but it happens randomly and only occasionally, so I just ignored these batches.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 926745,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/13/2020 00:44:46",
      "content": "<p>this don't the problem, but suppress the error so that the training can go on:</p>\n\n<p><code>\nprobability[np.isnan(probability)]=0\n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 931752,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/16/2020 12:26:02",
      "content": "<p>i find one image that causes NAN</p>\n\n<p>0670.jpg (test)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2dfbb94c4987745980533fe9c10e6260%2F0670.jpg?generation=1594902348609780&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 932104,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "07/16/2020 17:41:00",
          "content": "<p>Interesting. But why it cause Nan?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "921492": "Did anyone meet this error while training?\nWe believe it's caused by NaN value in logits and CNN weights but didn't know why.\nWe've tried clip gradiant but didn't help.\n\n```\n  File \"train.py\", line 247, in alaska_weighted_auc\n    fpr, tpr, thresholds = metrics.roc_curve(y_true, y_valid, pos_label=1)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 771, in roc_curve\n    y_true, y_score, pos_label=pos_label, sample_weight=sample_weight)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/metrics/_ranking.py\", line 542, in _binary_clf_curve\n    assert_all_finite(y_score)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 77, in assert_all_finite\n    _assert_all_finite(X.data if sp.issparse(X) else X, allow_nan)\n  File \"/home/yuan/miniconda3/envs/ML/lib/python3.7/site-packages/sklearn/utils/validation.py\", line 60, in _assert_all_finite\n    msg_dtype if msg_dtype is not None else X.dtype)\nValueError: Input contains NaN, infinity or a value too large for dtype('float32').\n```",
    "921527": "yes,   when i trained with fp16, some models crash like that. \nHowever, some models are fine. For example, efficientnetb2 is ok,  effcientnetb0 will comes to NAN.",
    "921556": "I occasionally get NANs as predictions and replace those accordingly.",
    "921568": "When training EfficientNet-B7 with mixed precision, I would start getting NaN loss. Couldn't figure it out, so just gave up.",
    "921619": "I also got an error when using fp16.\nI solved it by increasing the eps value.\n```\ntorch.optim.AdamW(self.model.parameters(), lr=0.001, eps=1e-4)\n```",
    "921757": "Thanks! I'll try it.",
    "922475": "It still crashes occasionally... 🤕 \nThat's harsh...",
    "922504": "There is one full black image in the data that gave me troubles before, just another random idea.",
    "922625": "Thanks! We'll have a try.",
    "922733": "EfficientNet larger than b5 gives me the same error in validating (interestingly I haven't got this in training),  but it happens randomly and only occasionally, so I just ignored these batches.",
    "926228": "When we adding eps=1e-4, clip grad = 0.5, and avoiding all-zero input, it still crashed occasionally but the probability is much lower than before.\nThanks for you guys help!",
    "926745": "this don't the problem, but suppress the error so that the training can go on:\n\n```\nprobability[np.isnan(probability)]=0\n```",
    "931752": "i find one image that causes NAN\n\n0670.jpg (test)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2dfbb94c4987745980533fe9c10e6260%2F0670.jpg?generation=1594902348609780&amp;alt=media)",
    "932104": "Interesting. But why it cause Nan?"
  },
  "source": "meta"
}