{
  "id": 309478,
  "title": "why  training loss becomes NAN when reducing batch size?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/309478",
  "author_name": "dragon zhang",
  "post_date": "2022-02-23T16:26:58.168000",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\" target=\"_blank\">https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop</a></p>\n<p>The notebook in above link  runs perfectly on Kaggle, however  I ran out of TPU hours quickly.</p>\n<p>when I reduce the batch size and run on Colab pro TPU,  the first 5/6 epochs run fine, however then the loss becomes NAN suddenly.</p>\n<p>I  search and find that It says that adding a small number to avoid it.</p>\n<p>My question is that why It happens when reducing batch size for TPU training?</p>",
  "messages": [
    {
      "id": 1702473,
      "postDate": "2022-02-23T16:26:58.167Z",
      "content": "<p><a href=\"https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\" target=\"_blank\">https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop</a></p>\n<p>The notebook in above link  runs perfectly on Kaggle, however  I ran out of TPU hours quickly.</p>\n<p>when I reduce the batch size and run on Colab pro TPU,  the first 5/6 epochs run fine, however then the loss becomes NAN suddenly.</p>\n<p>I  search and find that It says that adding a small number to avoid it.</p>\n<p>My question is that why It happens when reducing batch size for TPU training?</p>",
      "rawMarkdown": "https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\n\nThe notebook in above link  runs perfectly on Kaggle, however  I ran out of TPU hours quickly.\n\nwhen I reduce the batch size and run on Colab pro TPU,  the first 5/6 epochs run fine, however then the loss becomes NAN suddenly.\n\nI  search and find that It says that adding a small number to avoid it.\n\nMy question is that why It happens when reducing batch size for TPU training?\n",
      "votes": 4
    },
    {
      "id": 1702953,
      "postDate": "2022-02-24T05:14:00.467Z",
      "content": "<p>I'm not running that exact notebook on Colab Pro, but I am running something <em>very</em> similar, which is my mod of the initial version  <a href=\"https://www.kaggle.com/manojprabhaakr/effnet-b6-whale-comp\" target=\"_blank\">EFFNET B6 WHALE COMP</a> by <a href=\"https://www.kaggle.com/manojprabhaakr\" target=\"_blank\">@manojprabhaakr</a> </p>\n<p>I have also added in the changes to be able to use the dectic cropped images.</p>\n<p>I'm able to run it at <code>BATCH_SIZE = 4 * strategy.num_replicas_in_sync</code> (32) and I'm pushing it harder than the NB you are using. I don't get NaN errors. </p>",
      "rawMarkdown": "I'm not running that exact notebook on Colab Pro, but I am running something *very* similar, which is my mod of the initial version  [EFFNET B6 WHALE COMP](https://www.kaggle.com/manojprabhaakr/effnet-b6-whale-comp) by @manojprabhaakr \n\nI have also added in the changes to be able to use the dectic cropped images.\n\nI'm able to run it at `BATCH_SIZE = 4 * strategy.num_replicas_in_sync` (32) and I'm pushing it harder than the NB you are using. I don't get NaN errors. ",
      "votes": 1,
      "replies": [
        {
          "id": 1702955,
          "postDate": "2022-02-24T05:19:25.173Z",
          "content": "<p>thank you for reply.</p>\n<p>That is why I don't get it.</p>",
          "rawMarkdown": "thank you for reply.\n\nThat is why I don't get it."
        }
      ]
    },
    {
      "id": 1706517,
      "postDate": "2022-02-27T14:31:22.237Z",
      "content": "<p>I have the same issue when I reduce batch size. Moreover I noticed that train embedings after go through the models also had some nan value. </p>\n<p>Have you solved this yet?</p>",
      "rawMarkdown": "I have the same issue when I reduce batch size. Moreover I noticed that train embedings after go through the models also had some nan value. \n\nHave you solved this yet?"
    },
    {
      "id": 1702927,
      "postDate": "2022-02-24T04:27:05.153Z",
      "content": "<p>Hi, I use batch size=32 on colab pro, everything works fine.</p>",
      "rawMarkdown": "Hi, I use batch size=32 on colab pro, everything works fine.",
      "replies": [
        {
          "id": 1702934,
          "postDate": "2022-02-24T04:43:28.023Z",
          "content": "<p>pro or pro+?  your env</p>",
          "rawMarkdown": "pro or pro+?  your env"
        },
        {
          "id": 1702936,
          "postDate": "2022-02-24T04:44:30.667Z",
          "content": "<p>colab pro.</p>",
          "rawMarkdown": "colab pro."
        }
      ]
    },
    {
      "id": 1702717,
      "postDate": "2022-02-23T21:03:41.777Z",
      "content": "<p>What batch size does that causes the issue?</p>",
      "rawMarkdown": "What batch size does that causes the issue?",
      "replies": [
        {
          "id": 1702933,
          "postDate": "2022-02-24T04:42:47.377Z",
          "content": "<p>2xreplication=16</p>",
          "rawMarkdown": "2xreplication=16"
        }
      ]
    },
    {
      "id": 1705437,
      "postDate": "2022-02-26T13:44:13.750Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1702953,
      "author_name": "Bruce Young",
      "author_url": "",
      "post_date": "2022-02-24T05:14:00.467000",
      "content": "<p>I'm not running that exact notebook on Colab Pro, but I am running something <em>very</em> similar, which is my mod of the initial version  <a href=\"https://www.kaggle.com/manojprabhaakr/effnet-b6-whale-comp\" target=\"_blank\">EFFNET B6 WHALE COMP</a> by <a href=\"https://www.kaggle.com/manojprabhaakr\" target=\"_blank\">@manojprabhaakr</a> </p>\n<p>I have also added in the changes to be able to use the dectic cropped images.</p>\n<p>I'm able to run it at <code>BATCH_SIZE = 4 * strategy.num_replicas_in_sync</code> (32) and I'm pushing it harder than the NB you are using. I don't get NaN errors. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1702955,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-02-24T05:19:25.173000",
          "content": "<p>thank you for reply.</p>\n<p>That is why I don't get it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1706517,
      "author_name": "Mạnh Đỗ",
      "author_url": "",
      "post_date": "2022-02-27T14:31:22.237000",
      "content": "<p>I have the same issue when I reduce batch size. Moreover I noticed that train embedings after go through the models also had some nan value. </p>\n<p>Have you solved this yet?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1702927,
      "author_name": "Time Master",
      "author_url": "",
      "post_date": "2022-02-24T04:27:05.153000",
      "content": "<p>Hi, I use batch size=32 on colab pro, everything works fine.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1702934,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-02-24T04:43:28.023000",
          "content": "<p>pro or pro+?  your env</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1702936,
          "author_name": "Time Master",
          "author_url": "",
          "post_date": "2022-02-24T04:44:30.667000",
          "content": "<p>colab pro.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1702717,
      "author_name": "Mykola",
      "author_url": "",
      "post_date": "2022-02-23T21:03:41.777000",
      "content": "<p>What batch size does that causes the issue?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1702933,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-02-24T04:42:47.377000",
          "content": "<p>2xreplication=16</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1705437,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-26T13:44:13.750000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1702473": "https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\n\nThe notebook in above link  runs perfectly on Kaggle, however  I ran out of TPU hours quickly.\n\nwhen I reduce the batch size and run on Colab pro TPU,  the first 5/6 epochs run fine, however then the loss becomes NAN suddenly.\n\nI  search and find that It says that adding a small number to avoid it.\n\nMy question is that why It happens when reducing batch size for TPU training?\n",
    "1702953": "I'm not running that exact notebook on Colab Pro, but I am running something *very* similar, which is my mod of the initial version  [EFFNET B6 WHALE COMP](https://www.kaggle.com/manojprabhaakr/effnet-b6-whale-comp) by @manojprabhaakr \n\nI have also added in the changes to be able to use the dectic cropped images.\n\nI'm able to run it at `BATCH_SIZE = 4 * strategy.num_replicas_in_sync` (32) and I'm pushing it harder than the NB you are using. I don't get NaN errors. ",
    "1706517": "I have the same issue when I reduce batch size. Moreover I noticed that train embedings after go through the models also had some nan value. \n\nHave you solved this yet?",
    "1702927": "Hi, I use batch size=32 on colab pro, everything works fine.",
    "1702717": "What batch size does that causes the issue?",
    "1705437": ""
  }
}