{
  "id": 190096,
  "title": "Error: Confidences should sum to 1",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/190096",
  "author_name": "",
  "post_date": "2020-10-10T06:34:29.405683800Z",
  "votes": 8,
  "comment_count": 14,
  "views": 0,
  "content": "<p>From time to time I encounter the following error while training with a part of the trainings data:</p>\n<pre><code>assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\n\nAssertionError: confidences should sum to 1\n</code></pre>\n<p>Did anyone receive the same error yet? Mostly, I just run the same script with the same config again and it works with a second try but I guess there must be a reason for this.</p>",
  "messages": [
    {
      "id": "1044823",
      "postDate": "10/10/2020 06:34:29",
      "content": "<p>From time to time I encounter the following error while training with a part of the trainings data:</p>\n<pre><code>assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\n\nAssertionError: confidences should sum to 1\n</code></pre>\n<p>Did anyone receive the same error yet? Mostly, I just run the same script with the same config again and it works with a second try but I guess there must be a reason for this.</p>",
      "rawMarkdown": "From time to time I encounter the following error while training with a part of the trainings data:\n\n```\nassert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\n\nAssertionError: confidences should sum to 1\n```\n\nDid anyone receive the same error yet? Mostly, I just run the same script with the same config again and it works with a second try but I guess there must be a reason for this.",
      "votes": null
    },
    {
      "id": "1045110",
      "postDate": "10/10/2020 10:57:51",
      "content": "<p>I got the same error sometimes, and i am not sure what causes it. It's mostly preceeded by a sudden jump in the training jump which may cause the weights of the net to diverge.</p>\n<p>Made me wonder if there are any broken records in the training data. I've only seen it in train_full.zarr yet. Where did it happen for you?</p>",
      "rawMarkdown": "I got the same error sometimes, and i am not sure what causes it. It's mostly preceeded by a sudden jump in the training jump which may cause the weights of the net to diverge.\n\nMade me wonder if there are any broken records in the training data. I've only seen it in train_full.zarr yet. Where did it happen for you?",
      "votes": null
    },
    {
      "id": "1045153",
      "postDate": "10/10/2020 11:41:32",
      "content": "<p>It only occurs when I take a part of the total data set (train_zarr not train_full.zarr!). I have added a param to the cfg which defines the percentage I want to use for train -&gt; for example to use only 5% of the training data. This is where the error often occurs. So far the error did not occur when I used all data. But basically it seems to have been just luck in that case…</p>\n<p>To be honest, I simply started training without any form of quality checks. Purely from the theory it is already possible that sensors etc. sometimes deliver wrong values that have found there way into the data set. For example, we could check if there were unusually high changes in the coordinates between two frames and drop them from the dataset. What do you think?</p>",
      "rawMarkdown": "It only occurs when I take a part of the total data set (train_zarr not train_full.zarr!). I have added a param to the cfg which defines the percentage I want to use for train -> for example to use only 5% of the training data. This is where the error often occurs. So far the error did not occur when I used all data. But basically it seems to have been just luck in that case...\n\nTo be honest, I simply started training without any form of quality checks. Purely from the theory it is already possible that sensors etc. sometimes deliver wrong values that have found there way into the data set. For example, we could check if there were unusually high changes in the coordinates between two frames and drop them from the dataset. What do you think?",
      "votes": null
    },
    {
      "id": "1045175",
      "postDate": "10/10/2020 12:10:01",
      "content": "<p>yes doing a quality check over the training set (and also the test set) was on my list of TODOs ;)</p>",
      "rawMarkdown": "yes doing a quality check over the training set (and also the test set) was on my list of TODOs ;)",
      "votes": null
    },
    {
      "id": "1045280",
      "postDate": "10/10/2020 13:40:59",
      "content": "<p>I had the same error. The error appeared because the model puts in somecases <code>NaN</code> values out, which resulted to <code>AssertionError: confidences should sum to 1</code> <br>\nThis issue is also known under the term <em>Exploding Gradient</em> </p>\n<p>To fix it you can adapt different hyperparameters, I have opened a topic before to report about the issue:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773</a></p>",
      "rawMarkdown": "I had the same error. The error appeared because the model puts in somecases `NaN` values out, which resulted to `AssertionError: confidences should sum to 1` \nThis issue is also known under the term *Exploding Gradient* \n\n\nTo fix it you can adapt different hyperparameters, I have opened a topic before to report about the issue:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773",
      "votes": null
    },
    {
      "id": "1045289",
      "postDate": "10/10/2020 13:49:26",
      "content": "<p>Wow, thanks. That is a great finding. Thanks for sharing. I even saw your thread but didn´t pay enough attention..</p>",
      "rawMarkdown": "Wow, thanks. That is a great finding. Thanks for sharing. I even saw your thread but didn´t pay enough attention..",
      "votes": null
    },
    {
      "id": "1046618",
      "postDate": "10/11/2020 21:05:04",
      "content": "<p>maybe we can think about putting a Nan check in L5Kit evaluation functions….that error is a little bit misleading</p>",
      "rawMarkdown": "maybe we can think about putting a Nan check in L5Kit evaluation functions....that error is a little bit misleading",
      "votes": null
    },
    {
      "id": "1046929",
      "postDate": "10/12/2020 05:39:25",
      "content": "<p>This happens when train loader is not set properly</p>",
      "rawMarkdown": "This happens when train loader is not set properly",
      "votes": null
    },
    {
      "id": "1046959",
      "postDate": "10/12/2020 06:18:24",
      "content": "<p>would you mind explaining in more detail?</p>",
      "rawMarkdown": "would you mind explaining in more detail?",
      "votes": null
    },
    {
      "id": "1047310",
      "postDate": "10/12/2020 13:14:48",
      "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> . That will be useful. It seems that error occurs often</p>",
      "rawMarkdown": "lucabergamini . That will be useful. It seems that error occurs often",
      "votes": null
    },
    {
      "id": "1047598",
      "postDate": "10/12/2020 18:18:46",
      "content": "<p>I tried to avoid it by putting a try catch for the assertion errors and in the catch the dataloader will go to the next iteration, without making the weights update. What is interesting is that after this errors appears once, it appears on all future dataloader iteration, so the only way that I see how to go pass this errors is to start the training again from the last saved checkpoint. It is pretty annoying. What approach are you using ?<br>\nAlso, changing hyperparams, resolution and other variables did not help, sooner or later this assertion appeared </p>",
      "rawMarkdown": "I tried to avoid it by putting a try catch for the assertion errors and in the catch the dataloader will go to the next iteration, without making the weights update. What is interesting is that after this errors appears once, it appears on all future dataloader iteration, so the only way that I see how to go pass this errors is to start the training again from the last saved checkpoint. It is pretty annoying. What approach are you using ?\nAlso, changing hyperparams, resolution and other variables did not help, sooner or later this assertion appeared",
      "votes": null
    },
    {
      "id": "1048153",
      "postDate": "10/13/2020 08:37:20",
      "content": "<p>I am using a try / except block as well. Not a beautiful solutions but it works. Unfortunately the error appears more often in my latest tests when playing around with the pixel sizes..</p>",
      "rawMarkdown": "I am using a try / except block as well. Not a beautiful solutions but it works. Unfortunately the error appears more often in my latest tests when playing around with the pixel sizes..",
      "votes": null
    },
    {
      "id": "1048205",
      "postDate": "10/13/2020 09:26:28",
      "content": "<p>For me after I catch the assertion error I use a continue statement so it will go on the next iteration on the for loop of the dataloader. But also the next batches after the catch keep giving the same assertion error. It is the same for you too <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> ?</p>",
      "rawMarkdown": "For me after I catch the assertion error I use a continue statement so it will go on the next iteration on the for loop of the dataloader. But also the next batches after the catch keep giving the same assertion error. It is the same for you too @benbla ?",
      "votes": null
    },
    {
      "id": "1048227",
      "postDate": "10/13/2020 09:55:43",
      "content": "<p>Actually I was very lucky with training for larger episodes. I encounter the error mostly in the first ~2 hours of training. For testing different parameters with smaller parts of the trainings data, I use something like that:</p>\n<pre><code>for test in tests:\n   while True:\n      try:\n         #do stuff here and use test parameters\n         break\n      except:         \n         continue\n</code></pre>\n<p>So the dataloader is also integrated in the \"try\"-block, so basically yes, I have to load the whole dataset again.</p>",
      "rawMarkdown": "Actually I was very lucky with training for larger episodes. I encounter the error mostly in the first ~2 hours of training. For testing different parameters with smaller parts of the trainings data, I use something like that:\n\n```\nfor test in tests:\n   while True:\n      try:\n         #do stuff here and use test parameters\n         break\n      except:         \n         continue\n```\n\nSo the dataloader is also integrated in the \"try\"-block, so basically yes, I have to load the whole dataset again.",
      "votes": null
    },
    {
      "id": "1048297",
      "postDate": "10/13/2020 11:17:18",
      "content": "<p>For me it appears random, sometimes at beginning and sometimes even after over 80 hours of training. I have tried adjusting params but sooner or later it happens. It is annoying not be able to train on the complete dataset. <br>\n<a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a>  Do you know more details about this type of behavior or how can we avoid ? </p>",
      "rawMarkdown": "For me it appears random, sometimes at beginning and sometimes even after over 80 hours of training. I have tried adjusting params but sooner or later it happens. It is annoying not be able to train on the complete dataset. \n@lucabergamini  Do you know more details about this type of behavior or how can we avoid ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1045110,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "10/10/2020 10:57:51",
      "content": "<p>I got the same error sometimes, and i am not sure what causes it. It's mostly preceeded by a sudden jump in the training jump which may cause the weights of the net to diverge.</p>\n<p>Made me wonder if there are any broken records in the training data. I've only seen it in train_full.zarr yet. Where did it happen for you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1045153,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "10/10/2020 11:41:32",
          "content": "<p>It only occurs when I take a part of the total data set (train_zarr not train_full.zarr!). I have added a param to the cfg which defines the percentage I want to use for train -&gt; for example to use only 5% of the training data. This is where the error often occurs. So far the error did not occur when I used all data. But basically it seems to have been just luck in that case…</p>\n<p>To be honest, I simply started training without any form of quality checks. Purely from the theory it is already possible that sensors etc. sometimes deliver wrong values that have found there way into the data set. For example, we could check if there were unusually high changes in the coordinates between two frames and drop them from the dataset. What do you think?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045175,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/10/2020 12:10:01",
          "content": "<p>yes doing a quality check over the training set (and also the test set) was on my list of TODOs ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045280,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "10/10/2020 13:40:59",
          "content": "<p>I had the same error. The error appeared because the model puts in somecases <code>NaN</code> values out, which resulted to <code>AssertionError: confidences should sum to 1</code> <br>\nThis issue is also known under the term <em>Exploding Gradient</em> </p>\n<p>To fix it you can adapt different hyperparameters, I have opened a topic before to report about the issue:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045289,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "10/10/2020 13:49:26",
          "content": "<p>Wow, thanks. That is a great finding. Thanks for sharing. I even saw your thread but didn´t pay enough attention..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046618,
          "author_name": "lucabergamini",
          "author_url": "",
          "post_date": "10/11/2020 21:05:04",
          "content": "<p>maybe we can think about putting a Nan check in L5Kit evaluation functions….that error is a little bit misleading</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047310,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "10/12/2020 13:14:48",
          "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> . That will be useful. It seems that error occurs often</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046929,
      "author_name": "sanjaydsb",
      "author_url": "",
      "post_date": "10/12/2020 05:39:25",
      "content": "<p>This happens when train loader is not set properly</p>",
      "votes": null,
      "replies": [
        {
          "id": 1046959,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/12/2020 06:18:24",
          "content": "<p>would you mind explaining in more detail?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1047598,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "10/12/2020 18:18:46",
      "content": "<p>I tried to avoid it by putting a try catch for the assertion errors and in the catch the dataloader will go to the next iteration, without making the weights update. What is interesting is that after this errors appears once, it appears on all future dataloader iteration, so the only way that I see how to go pass this errors is to start the training again from the last saved checkpoint. It is pretty annoying. What approach are you using ?<br>\nAlso, changing hyperparams, resolution and other variables did not help, sooner or later this assertion appeared </p>",
      "votes": null,
      "replies": [
        {
          "id": 1048153,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "10/13/2020 08:37:20",
          "content": "<p>I am using a try / except block as well. Not a beautiful solutions but it works. Unfortunately the error appears more often in my latest tests when playing around with the pixel sizes..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048205,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "10/13/2020 09:26:28",
          "content": "<p>For me after I catch the assertion error I use a continue statement so it will go on the next iteration on the for loop of the dataloader. But also the next batches after the catch keep giving the same assertion error. It is the same for you too <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048227,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "10/13/2020 09:55:43",
          "content": "<p>Actually I was very lucky with training for larger episodes. I encounter the error mostly in the first ~2 hours of training. For testing different parameters with smaller parts of the trainings data, I use something like that:</p>\n<pre><code>for test in tests:\n   while True:\n      try:\n         #do stuff here and use test parameters\n         break\n      except:         \n         continue\n</code></pre>\n<p>So the dataloader is also integrated in the \"try\"-block, so basically yes, I have to load the whole dataset again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048297,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "10/13/2020 11:17:18",
          "content": "<p>For me it appears random, sometimes at beginning and sometimes even after over 80 hours of training. I have tried adjusting params but sooner or later it happens. It is annoying not be able to train on the complete dataset. <br>\n<a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a>  Do you know more details about this type of behavior or how can we avoid ? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1044823": "From time to time I encounter the following error while training with a part of the trainings data:\n\n```\nassert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\n\nAssertionError: confidences should sum to 1\n```\n\nDid anyone receive the same error yet? Mostly, I just run the same script with the same config again and it works with a second try but I guess there must be a reason for this.",
    "1045110": "I got the same error sometimes, and i am not sure what causes it. It's mostly preceeded by a sudden jump in the training jump which may cause the weights of the net to diverge.\n\nMade me wonder if there are any broken records in the training data. I've only seen it in train_full.zarr yet. Where did it happen for you?",
    "1045153": "It only occurs when I take a part of the total data set (train_zarr not train_full.zarr!). I have added a param to the cfg which defines the percentage I want to use for train -> for example to use only 5% of the training data. This is where the error often occurs. So far the error did not occur when I used all data. But basically it seems to have been just luck in that case...\n\nTo be honest, I simply started training without any form of quality checks. Purely from the theory it is already possible that sensors etc. sometimes deliver wrong values that have found there way into the data set. For example, we could check if there were unusually high changes in the coordinates between two frames and drop them from the dataset. What do you think?",
    "1045175": "yes doing a quality check over the training set (and also the test set) was on my list of TODOs ;)",
    "1045280": "I had the same error. The error appeared because the model puts in somecases `NaN` values out, which resulted to `AssertionError: confidences should sum to 1` \nThis issue is also known under the term *Exploding Gradient* \n\n\nTo fix it you can adapt different hyperparameters, I have opened a topic before to report about the issue:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773",
    "1045289": "Wow, thanks. That is a great finding. Thanks for sharing. I even saw your thread but didn´t pay enough attention..",
    "1046618": "maybe we can think about putting a Nan check in L5Kit evaluation functions....that error is a little bit misleading",
    "1046929": "This happens when train loader is not set properly",
    "1046959": "would you mind explaining in more detail?",
    "1047310": "lucabergamini . That will be useful. It seems that error occurs often",
    "1047598": "I tried to avoid it by putting a try catch for the assertion errors and in the catch the dataloader will go to the next iteration, without making the weights update. What is interesting is that after this errors appears once, it appears on all future dataloader iteration, so the only way that I see how to go pass this errors is to start the training again from the last saved checkpoint. It is pretty annoying. What approach are you using ?\nAlso, changing hyperparams, resolution and other variables did not help, sooner or later this assertion appeared",
    "1048153": "I am using a try / except block as well. Not a beautiful solutions but it works. Unfortunately the error appears more often in my latest tests when playing around with the pixel sizes..",
    "1048205": "For me after I catch the assertion error I use a continue statement so it will go on the next iteration on the for loop of the dataloader. But also the next batches after the catch keep giving the same assertion error. It is the same for you too @benbla ?",
    "1048227": "Actually I was very lucky with training for larger episodes. I encounter the error mostly in the first ~2 hours of training. For testing different parameters with smaller parts of the trainings data, I use something like that:\n\n```\nfor test in tests:\n   while True:\n      try:\n         #do stuff here and use test parameters\n         break\n      except:         \n         continue\n```\n\nSo the dataloader is also integrated in the \"try\"-block, so basically yes, I have to load the whole dataset again.",
    "1048297": "For me it appears random, sometimes at beginning and sometimes even after over 80 hours of training. I have tried adjusting params but sooner or later it happens. It is annoying not be able to train on the complete dataset. \n@lucabergamini  Do you know more details about this type of behavior or how can we avoid ?"
  },
  "source": "meta"
}