{
  "id": 306916,
  "title": "Unexpected error while training with Batch_SIZE 8",
  "url": "/competitions/happy-whale-and-dolphin/discussion/306916",
  "author_name": "",
  "post_date": "2022-02-11T13:56:31.694103500Z",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>When I train my model with Batch_SIZE 8<br>\nI get this error:<br>\n<code>ValueError: Input contains NaN, infinity or a value too large for dtype('float32')</code><br>\nIn this code cell:<br>\n<img src=\"https://i.imgur.com/udxd1nn.png\" alt=\"\"></p>\n<h5>Can anyone help me to solve this error?</h5>",
  "messages": [
    {
      "id": "1685721",
      "postDate": "02/11/2022 13:56:31",
      "content": "<p>When I train my model with Batch_SIZE 8<br>\nI get this error:<br>\n<code>ValueError: Input contains NaN, infinity or a value too large for dtype('float32')</code><br>\nIn this code cell:<br>\n<img src=\"https://i.imgur.com/udxd1nn.png\" alt=\"\"></p>\n<h5>Can anyone help me to solve this error?</h5>",
      "rawMarkdown": "When I train my model with Batch_SIZE 8\nI get this error:\n`ValueError: Input contains NaN, infinity or a value too large for dtype('float32')`\nIn this code cell:\n![](https://i.imgur.com/udxd1nn.png)\n##### Can anyone help me to solve this error?",
      "votes": null
    },
    {
      "id": "1686188",
      "postDate": "02/11/2022 20:47:55",
      "content": "<p>embeddings = np.nan_to_num(embeddings).Using this the error will be resolved</p>",
      "rawMarkdown": "embeddings = np.nan_to_num(embeddings).Using this the error will be resolved",
      "votes": null
    },
    {
      "id": "1686406",
      "postDate": "02/12/2022 04:04:35",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/darkravager\" target=\"_blank\">@darkravager</a> for your great sharing.<br>\nIt resolves my error, but can you please tell me what is reason behind this error?</p>",
      "rawMarkdown": "Thank you @darkravager for your great sharing.\nIt resolves my error, but can you please tell me what is reason behind this error?",
      "votes": null
    },
    {
      "id": "1687972",
      "postDate": "02/13/2022 09:22:41",
      "content": "<p>You can find the answer in the <a href=\"https://numpy.org/doc/stable/reference/generated/numpy.nan_to_num.html\" target=\"_blank\">docs</a></p>\n<p>From the page:</p>\n<blockquote>\n  <p>Replace NaN with zero and infinity with large finite numbers (default behaviour) or with the numbers defined by the user using the nan, posinf and/or neginf keywords.</p>\n</blockquote>\n<p>We have NaNs in the dataset </p>",
      "rawMarkdown": "You can find the answer in the [docs](https://numpy.org/doc/stable/reference/generated/numpy.nan_to_num.html)\n\nFrom the page:\n\n> Replace NaN with zero and infinity with large finite numbers (default behaviour) or with the numbers defined by the user using the nan, posinf and/or neginf keywords.\n\nWe have NaNs in the dataset",
      "votes": null
    },
    {
      "id": "1688016",
      "postDate": "02/13/2022 10:18:55",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> for your great sharing and response.<br>\nAs you previously stated, the dataset contains NaN values.<br>\nBut when I train my model with Batch_SIZE 8, it finds NaN values.<br>\nBut if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.<br>\nCan you please tell me why this happens?</p>",
      "rawMarkdown": "Thank you @init27 for your great sharing and response.\nAs you previously stated, the dataset contains NaN values.\nBut when I train my model with Batch_SIZE 8, it finds NaN values.\nBut if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.\nCan you please tell me why this happens?",
      "votes": null
    },
    {
      "id": "1688023",
      "postDate": "02/13/2022 10:27:45",
      "content": "<blockquote>\n  <p>But if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.</p>\n</blockquote>\n<p>This is really strange though 😱, let me try to replicate it</p>",
      "rawMarkdown": "> But if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.\n\nThis is really strange though 😱, let me try to replicate it",
      "votes": null
    },
    {
      "id": "1688364",
      "postDate": "02/13/2022 15:39:31",
      "content": "<p>Try to decrease learning rate when you use a smaller batch size.</p>",
      "rawMarkdown": "Try to decrease learning rate when you use a smaller batch size.",
      "votes": null
    },
    {
      "id": "1688399",
      "postDate": "02/13/2022 15:49:15",
      "content": "<p><a href=\"https://www.kaggle.com/mknzfr\" target=\"_blank\">@mknzfr</a> Sorry if I don't understand-how might that help with the current issue?</p>",
      "rawMarkdown": "mknzfr Sorry if I don't understand-how might that help with the current issue?",
      "votes": null
    },
    {
      "id": "1696088",
      "postDate": "02/18/2022 15:28:46",
      "content": "<p>I am not sure if it is relevant for now, but… Do you use triplet loss? If you use a small batch size there is a big chance to make a batch with all different individuals. So, you cannot form Anchor-Positive pairs. That may cause you unexpected computations. I am not sure, it is just a hypothesis </p>",
      "rawMarkdown": "I am not sure if it is relevant for now, but... Do you use triplet loss? If you use a small batch size there is a big chance to make a batch with all different individuals. So, you cannot form Anchor-Positive pairs. That may cause you unexpected computations. I am not sure, it is just a hypothesis",
      "votes": null
    },
    {
      "id": "1750998",
      "postDate": "04/10/2022 09:26:02",
      "content": "<p>Can it be because batchnorm becomes unstable with too small a batch size?</p>",
      "rawMarkdown": "Can it be because batchnorm becomes unstable with too small a batch size?",
      "votes": null
    },
    {
      "id": "1752093",
      "postDate": "04/11/2022 11:52:02",
      "content": "<p>Another solution is to upgrade tf like in this <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315363\" target=\"_blank\">discussion post by Andrjj</a>. As to why increasing batch size removes those NaNs, it may be due to arcface. ArcFace trains well with higher batch size so having too small of a batch size may result in unstable arcface and weird embeddings (idk for sure if this is true however, just an idea).</p>",
      "rawMarkdown": "Another solution is to upgrade tf like in this [discussion post by Andrjj](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315363). As to why increasing batch size removes those NaNs, it may be due to arcface. ArcFace trains well with higher batch size so having too small of a batch size may result in unstable arcface and weird embeddings (idk for sure if this is true however, just an idea).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1686188,
      "author_name": "darkravager",
      "author_url": "",
      "post_date": "02/11/2022 20:47:55",
      "content": "<p>embeddings = np.nan_to_num(embeddings).Using this the error will be resolved</p>",
      "votes": null,
      "replies": [
        {
          "id": 1686406,
          "author_name": "iftiben10",
          "author_url": "",
          "post_date": "02/12/2022 04:04:35",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/darkravager\" target=\"_blank\">@darkravager</a> for your great sharing.<br>\nIt resolves my error, but can you please tell me what is reason behind this error?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1687972,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/13/2022 09:22:41",
          "content": "<p>You can find the answer in the <a href=\"https://numpy.org/doc/stable/reference/generated/numpy.nan_to_num.html\" target=\"_blank\">docs</a></p>\n<p>From the page:</p>\n<blockquote>\n  <p>Replace NaN with zero and infinity with large finite numbers (default behaviour) or with the numbers defined by the user using the nan, posinf and/or neginf keywords.</p>\n</blockquote>\n<p>We have NaNs in the dataset </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688016,
          "author_name": "iftiben10",
          "author_url": "",
          "post_date": "02/13/2022 10:18:55",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> for your great sharing and response.<br>\nAs you previously stated, the dataset contains NaN values.<br>\nBut when I train my model with Batch_SIZE 8, it finds NaN values.<br>\nBut if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.<br>\nCan you please tell me why this happens?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688023,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/13/2022 10:27:45",
          "content": "<blockquote>\n  <p>But if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.</p>\n</blockquote>\n<p>This is really strange though 😱, let me try to replicate it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688364,
          "author_name": "mknzfr",
          "author_url": "",
          "post_date": "02/13/2022 15:39:31",
          "content": "<p>Try to decrease learning rate when you use a smaller batch size.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688399,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/13/2022 15:49:15",
          "content": "<p><a href=\"https://www.kaggle.com/mknzfr\" target=\"_blank\">@mknzfr</a> Sorry if I don't understand-how might that help with the current issue?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696088,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "02/18/2022 15:28:46",
          "content": "<p>I am not sure if it is relevant for now, but… Do you use triplet loss? If you use a small batch size there is a big chance to make a batch with all different individuals. So, you cannot form Anchor-Positive pairs. That may cause you unexpected computations. I am not sure, it is just a hypothesis </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1752093,
          "author_name": "vexxingbanana",
          "author_url": "",
          "post_date": "04/11/2022 11:52:02",
          "content": "<p>Another solution is to upgrade tf like in this <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315363\" target=\"_blank\">discussion post by Andrjj</a>. As to why increasing batch size removes those NaNs, it may be due to arcface. ArcFace trains well with higher batch size so having too small of a batch size may result in unstable arcface and weird embeddings (idk for sure if this is true however, just an idea).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1750998,
      "author_name": "jackchungchiehyu",
      "author_url": "",
      "post_date": "04/10/2022 09:26:02",
      "content": "<p>Can it be because batchnorm becomes unstable with too small a batch size?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1685721": "When I train my model with Batch_SIZE 8\nI get this error:\n`ValueError: Input contains NaN, infinity or a value too large for dtype('float32')`\nIn this code cell:\n![](https://i.imgur.com/udxd1nn.png)\n##### Can anyone help me to solve this error?",
    "1686188": "embeddings = np.nan_to_num(embeddings).Using this the error will be resolved",
    "1686406": "Thank you @darkravager for your great sharing.\nIt resolves my error, but can you please tell me what is reason behind this error?",
    "1687972": "You can find the answer in the [docs](https://numpy.org/doc/stable/reference/generated/numpy.nan_to_num.html)\n\nFrom the page:\n\n> Replace NaN with zero and infinity with large finite numbers (default behaviour) or with the numbers defined by the user using the nan, posinf and/or neginf keywords.\n\nWe have NaNs in the dataset",
    "1688016": "Thank you @init27 for your great sharing and response.\nAs you previously stated, the dataset contains NaN values.\nBut when I train my model with Batch_SIZE 8, it finds NaN values.\nBut if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.\nCan you please tell me why this happens?",
    "1688023": "> But if I train my model with Batch_SIZE 16 instead of 8, it gets trained perfectly.\n\nThis is really strange though 😱, let me try to replicate it",
    "1688364": "Try to decrease learning rate when you use a smaller batch size.",
    "1688399": "mknzfr Sorry if I don't understand-how might that help with the current issue?",
    "1696088": "I am not sure if it is relevant for now, but... Do you use triplet loss? If you use a small batch size there is a big chance to make a batch with all different individuals. So, you cannot form Anchor-Positive pairs. That may cause you unexpected computations. I am not sure, it is just a hypothesis",
    "1750998": "Can it be because batchnorm becomes unstable with too small a batch size?",
    "1752093": "Another solution is to upgrade tf like in this [discussion post by Andrjj](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315363). As to why increasing batch size removes those NaNs, it may be due to arcface. ArcFace trains well with higher batch size so having too small of a batch size may result in unstable arcface and weird embeddings (idk for sure if this is true however, just an idea)."
  },
  "source": "meta"
}