{
  "id": 180278,
  "title": "has anybody encountered nan loss during training? any solution to this??? ",
  "url": "/competitions/landmark-recognition-2020/discussion/180278",
  "author_name": "",
  "post_date": "2020-09-04T13:19:46.780219500Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>i am training using simple effnet b6  and frequently ending up with nan loss during training? any solution to this? i am not able to trace down where this is coming from???</p>",
  "messages": [
    {
      "id": "998064",
      "postDate": "09/04/2020 13:19:46",
      "content": "<p>i am training using simple effnet b6  and frequently ending up with nan loss during training? any solution to this? i am not able to trace down where this is coming from???</p>",
      "rawMarkdown": "i am training using simple effnet b6  and frequently ending up with nan loss during training? any solution to this? i am not able to trace down where this is coming from???",
      "votes": null
    },
    {
      "id": "998095",
      "postDate": "09/04/2020 13:50:02",
      "content": "<p>I believe this write-up might be useful for you: <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306</a></p>\n<blockquote>\n  <p>At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that the dataset is a subset so the labels is also a subset. I overcome it with tf.lookup.statichashtable.</p>\n</blockquote>",
      "rawMarkdown": "I believe this write-up might be useful for you: [https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306)\n\n> At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that the dataset is a subset so the labels is also a subset. I overcome it with tf.lookup.statichashtable.",
      "votes": null
    },
    {
      "id": "998143",
      "postDate": "09/04/2020 14:20:50",
      "content": "<p>this is how i do:<br>\n1)take the subset of the data from train.csv<br>\n2)map the ids to labels using dict around 25000 classes<br>\n3)feed the effnet. </p>\n<p>i do not understand your statement that labels is also a subset. yes it is a subset but how that would create the problem can you please explain in more detail? what about the above procedure? still do u think it will create the nan probelm because of the subset of labels??</p>",
      "rawMarkdown": "this is how i do:\n1)take the subset of the data from train.csv\n2)map the ids to labels using dict around 25000 classes\n3)feed the effnet. \n\ni do not understand your statement that labels is also a subset. yes it is a subset but how that would create the problem can you please explain in more detail? what about the above procedure? still do u think it will create the nan probelm because of the subset of labels??",
      "votes": null
    },
    {
      "id": "998190",
      "postDate": "09/04/2020 15:02:39",
      "content": "<p>may be you are training with fp16 ?</p>",
      "rawMarkdown": "may be you are training with fp16 ?",
      "votes": null
    },
    {
      "id": "998330",
      "postDate": "09/04/2020 16:59:09",
      "content": "<p>yes off-course!! wow how did you guess it?? is this the reason for nan error????</p>",
      "rawMarkdown": "yes off-course!! wow how did you guess it?? is this the reason for nan error????",
      "votes": null
    },
    {
      "id": "998705",
      "postDate": "09/05/2020 01:34:12",
      "content": "<p>Have you tried turning it off and monitor the loss?<br>\nHere is the <a href=\"https://forums.fast.ai/t/mixed-precision-training/29601/12\" target=\"_blank\">discussion</a> about nan loss and mixed precision training that they explained it pretty well </p>",
      "rawMarkdown": "Have you tried turning it off and monitor the loss?\nHere is the [discussion](https://forums.fast.ai/t/mixed-precision-training/29601/12 ) about nan loss and mixed precision training that they explained it pretty well",
      "votes": null
    },
    {
      "id": "1002451",
      "postDate": "09/08/2020 06:21:10",
      "content": "<p>not yet..I will try and update </p>",
      "rawMarkdown": "not yet..I will try and update",
      "votes": null
    },
    {
      "id": "1002453",
      "postDate": "09/08/2020 06:23:13",
      "content": "<p>you saved my day! after this comment i started to look at the labels side and found that sparse categorical cross entropy was creating problem. i changed to just categorical cross entropy and it works well. thanks </p>",
      "rawMarkdown": "you saved my day! after this comment i started to look at the labels side and found that sparse categorical cross entropy was creating problem. i changed to just categorical cross entropy and it works well. thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 998095,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "09/04/2020 13:50:02",
      "content": "<p>I believe this write-up might be useful for you: <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306</a></p>\n<blockquote>\n  <p>At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that the dataset is a subset so the labels is also a subset. I overcome it with tf.lookup.statichashtable.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 998143,
          "author_name": "udaygurugubelli",
          "author_url": "",
          "post_date": "09/04/2020 14:20:50",
          "content": "<p>this is how i do:<br>\n1)take the subset of the data from train.csv<br>\n2)map the ids to labels using dict around 25000 classes<br>\n3)feed the effnet. </p>\n<p>i do not understand your statement that labels is also a subset. yes it is a subset but how that would create the problem can you please explain in more detail? what about the above procedure? still do u think it will create the nan probelm because of the subset of labels??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1002453,
          "author_name": "udaygurugubelli",
          "author_url": "",
          "post_date": "09/08/2020 06:23:13",
          "content": "<p>you saved my day! after this comment i started to look at the labels side and found that sparse categorical cross entropy was creating problem. i changed to just categorical cross entropy and it works well. thanks </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 998190,
      "author_name": "dttung2905",
      "author_url": "",
      "post_date": "09/04/2020 15:02:39",
      "content": "<p>may be you are training with fp16 ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 998330,
          "author_name": "udaygurugubelli",
          "author_url": "",
          "post_date": "09/04/2020 16:59:09",
          "content": "<p>yes off-course!! wow how did you guess it?? is this the reason for nan error????</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998705,
          "author_name": "dttung2905",
          "author_url": "",
          "post_date": "09/05/2020 01:34:12",
          "content": "<p>Have you tried turning it off and monitor the loss?<br>\nHere is the <a href=\"https://forums.fast.ai/t/mixed-precision-training/29601/12\" target=\"_blank\">discussion</a> about nan loss and mixed precision training that they explained it pretty well </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1002451,
          "author_name": "udaygurugubelli",
          "author_url": "",
          "post_date": "09/08/2020 06:21:10",
          "content": "<p>not yet..I will try and update </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "998064": "i am training using simple effnet b6  and frequently ending up with nan loss during training? any solution to this? i am not able to trace down where this is coming from???",
    "998095": "I believe this write-up might be useful for you: [https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306)\n\n> At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that the dataset is a subset so the labels is also a subset. I overcome it with tf.lookup.statichashtable.",
    "998143": "this is how i do:\n1)take the subset of the data from train.csv\n2)map the ids to labels using dict around 25000 classes\n3)feed the effnet. \n\ni do not understand your statement that labels is also a subset. yes it is a subset but how that would create the problem can you please explain in more detail? what about the above procedure? still do u think it will create the nan probelm because of the subset of labels??",
    "998190": "may be you are training with fp16 ?",
    "998330": "yes off-course!! wow how did you guess it?? is this the reason for nan error????",
    "998705": "Have you tried turning it off and monitor the loss?\nHere is the [discussion](https://forums.fast.ai/t/mixed-precision-training/29601/12 ) about nan loss and mixed precision training that they explained it pretty well",
    "1002451": "not yet..I will try and update",
    "1002453": "you saved my day! after this comment i started to look at the labels side and found that sparse categorical cross entropy was creating problem. i changed to just categorical cross entropy and it works well. thanks"
  },
  "source": "meta"
}