{
  "id": 135474,
  "title": "Val loss and Val accuracy both increasing with training, how is this possible?",
  "url": "/competitions/bengaliai-cv19/discussion/135474",
  "author_name": "",
  "post_date": "2020-03-14T03:08:25.313151900Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The below are snips of my training session.\nCan anyone explain me how validation loss and validation accuracy are increasing at the same time while training. How is this possible? Ideally, the val loss should decrease with more number of epochs and thus correspondingly val acc would increase. </p>\n\n<p>The model was trained for 50-60 epochs with mixup, now i am trying to use gridmask on the pretrained mixup model, over in this training session.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2Ff195b1bb3983402f9bbaa42adbbff618%2Fval_loss.JPG?generation=1584154373076461&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2F773a327d7f656b700e5aa27484d89758%2Fval_acc.JPG?generation=1584154477514714&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "771329",
      "postDate": "03/14/2020 03:08:25",
      "content": "<p>The below are snips of my training session.\nCan anyone explain me how validation loss and validation accuracy are increasing at the same time while training. How is this possible? Ideally, the val loss should decrease with more number of epochs and thus correspondingly val acc would increase. </p>\n\n<p>The model was trained for 50-60 epochs with mixup, now i am trying to use gridmask on the pretrained mixup model, over in this training session.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2Ff195b1bb3983402f9bbaa42adbbff618%2Fval_loss.JPG?generation=1584154373076461&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2F773a327d7f656b700e5aa27484d89758%2Fval_acc.JPG?generation=1584154477514714&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The below are snips of my training session.\nCan anyone explain me how validation loss and validation accuracy are increasing at the same time while training. How is this possible? Ideally, the val loss should decrease with more number of epochs and thus correspondingly val acc would increase. \n\nThe model was trained for 50-60 epochs with mixup, now i am trying to use gridmask on the pretrained mixup model, over in this training session.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2Ff195b1bb3983402f9bbaa42adbbff618%2Fval_loss.JPG?generation=1584154373076461&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2F773a327d7f656b700e5aa27484d89758%2Fval_acc.JPG?generation=1584154477514714&amp;alt=media)",
      "votes": null
    },
    {
      "id": "771332",
      "postDate": "03/14/2020 03:19:12",
      "content": "<p>Building on Ankur's answer and the comment underneath it, I think the following scenario is possible, while I have no proof of it. Two phenomenons might be happening at the same time :</p>\n\n<pre><code>Some examples with borderline predictions get predicted better and so their output class changes (eg a cat image predicted at 0.4 to be a cat and 0.6 to be a horse becomes predicted 0.4 to be a horse and 0.6 to be a cat). Thanks to this, accuracy increases while loss decreases.\n\nSome examples with very bad predictions keep getting worse (eg a cat image predicted at 0.8 to be a horse becomes predicted at 0.9 to be a horse) AND/OR (more probable, in particular for multi-class ?) some examples with very good predictions get a little worse (eg a cat image predicted at 0.9 to be a cat becomes predicted at 0.8 to be a cat). With this phenomenon, loss increases while accuracy stays the same.\n</code></pre>\n\n<p>So if phenomenon 2 kicks in at some point, on lots of examples (eg for a specific class which is not well understood for some reason) and/or with a loss increase stonger than the loss decrease you gain from 1., then you might find yourself in your scenario.\nOnce again, maybe this is not what's happening, but I think that being able to come up with such scenarios must remind us of the sometimes slippery relationship between (cross-entropy)loss and accuracy.</p>",
      "rawMarkdown": "Building on Ankur's answer and the comment underneath it, I think the following scenario is possible, while I have no proof of it. Two phenomenons might be happening at the same time :\n\n    Some examples with borderline predictions get predicted better and so their output class changes (eg a cat image predicted at 0.4 to be a cat and 0.6 to be a horse becomes predicted 0.4 to be a horse and 0.6 to be a cat). Thanks to this, accuracy increases while loss decreases.\n\n    Some examples with very bad predictions keep getting worse (eg a cat image predicted at 0.8 to be a horse becomes predicted at 0.9 to be a horse) AND/OR (more probable, in particular for multi-class ?) some examples with very good predictions get a little worse (eg a cat image predicted at 0.9 to be a cat becomes predicted at 0.8 to be a cat). With this phenomenon, loss increases while accuracy stays the same.\n\nSo if phenomenon 2 kicks in at some point, on lots of examples (eg for a specific class which is not well understood for some reason) and/or with a loss increase stonger than the loss decrease you gain from 1., then you might find yourself in your scenario.\nOnce again, maybe this is not what's happening, but I think that being able to come up with such scenarios must remind us of the sometimes slippery relationship between (cross-entropy)loss and accuracy.",
      "votes": null
    },
    {
      "id": "771333",
      "postDate": "03/14/2020 03:20:29",
      "content": "<p><strong>A model can overfit to cross entropy loss without over overfitting to accuracy.</strong></p>\n\n<p>```\nThere is a key difference between the two types of loss:</p>\n\n<pre><code>Accuracy measures whether you get the prediction right\nCross entropy measures how confident you are about a prediction\n</code></pre>\n\n<p>For example, if an image of a cat is passed into two models. Model A predicts {cat: 0.9, dog: 0.1} and model B predicts {cat: 0.6, dog: 0.4}. Both model will score the same accuracy, but model A will have a lower loss.</p>\n\n<p>Because of this the model will try to be more and more confident to minimize loss. It works fine in training stage, but in validation stage it will perform poorly in term of loss. For example, for some borderline images, being confident e.g. {cat: 0.9, dog: 0.1} will give higher loss than being uncertain e.g. {cat: 0.6, dog: 0.4}</p>\n\n<p>In short, cross entropy loss measures the calibration of a model. Mis-calibration is a common issue to modern neuronal networks. They tend to be over-confident. On Calibration of Modern Neural Networks talks about it in great details.\n```</p>",
      "rawMarkdown": "**A model can overfit to cross entropy loss without over overfitting to accuracy.**\n\n```\nThere is a key difference between the two types of loss:\n\n    Accuracy measures whether you get the prediction right\n    Cross entropy measures how confident you are about a prediction\n\nFor example, if an image of a cat is passed into two models. Model A predicts {cat: 0.9, dog: 0.1} and model B predicts {cat: 0.6, dog: 0.4}. Both model will score the same accuracy, but model A will have a lower loss.\n\nBecause of this the model will try to be more and more confident to minimize loss. It works fine in training stage, but in validation stage it will perform poorly in term of loss. For example, for some borderline images, being confident e.g. {cat: 0.9, dog: 0.1} will give higher loss than being uncertain e.g. {cat: 0.6, dog: 0.4}\n\nIn short, cross entropy loss measures the calibration of a model. Mis-calibration is a common issue to modern neuronal networks. They tend to be over-confident. On Calibration of Modern Neural Networks talks about it in great details.\n```",
      "votes": null
    },
    {
      "id": "771334",
      "postDate": "03/14/2020 03:21:45",
      "content": "<p>```\nMany answers focus on the mathematical calculation explaining how is this possible. But they don't explain why it becomes so. And they cannot suggest how to digger further to be more clear.</p>\n\n<p>I have 3 hypothesis. And suggest some experiments to verify them. Hopefully it can help explain this problem.</p>\n\n<ol>\n<li>Label is noisy. Compare the false predictions when val_loss is minimum and val_acc is maximum. Check whether these sample are correctly labelled.</li>\n<li>[Less likely] The model doesn't have enough aspect of information to be certain. Experiment with more and larger hidden layers.\n3.[A very wild guess] This is a case where the model is less certain about certain things as being trained longer. Such situation happens to human as well. When someone started to learn a technique, he is told exactly what is good or bad, what is certain things for (high certainty). When he goes through more cases and examples, he realizes sometimes certain border can be blur (less certain, higher loss), even though he can make better decisions (more accuracy). And he may eventually gets more certain when he becomes a master after going through a huge list of samples and lots of trial and errors (more training data). So in this case, I suggest experiment with adding more noise to the training data (not label) may be helpful.</li>\n</ol>\n\n<p>Don't argue about this by just saying if you disagree with these hypothesis. It will be more meaningful to discuss with experiments to verify them, no matter the results prove them right, or prove them wrong.\n```</p>",
      "rawMarkdown": "```\nMany answers focus on the mathematical calculation explaining how is this possible. But they don't explain why it becomes so. And they cannot suggest how to digger further to be more clear.\n\nI have 3 hypothesis. And suggest some experiments to verify them. Hopefully it can help explain this problem.\n\n   1. Label is noisy. Compare the false predictions when val_loss is minimum and val_acc is maximum. Check whether these sample are correctly labelled.\n   2.  [Less likely] The model doesn't have enough aspect of information to be certain. Experiment with more and larger hidden layers.\n    3.[A very wild guess] This is a case where the model is less certain about certain things as being trained longer. Such situation happens to human as well. When someone started to learn a technique, he is told exactly what is good or bad, what is certain things for (high certainty). When he goes through more cases and examples, he realizes sometimes certain border can be blur (less certain, higher loss), even though he can make better decisions (more accuracy). And he may eventually gets more certain when he becomes a master after going through a huge list of samples and lots of trial and errors (more training data). So in this case, I suggest experiment with adding more noise to the training data (not label) may be helpful.\n\nDon't argue about this by just saying if you disagree with these hypothesis. It will be more meaningful to discuss with experiments to verify them, no matter the results prove them right, or prove them wrong.\n```",
      "votes": null
    },
    {
      "id": "771335",
      "postDate": "03/14/2020 03:22:55",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4244482%2F5fd6210a0f2e2b576fdab25083327fee%2Ftttt.jpg?generation=1584156173865554&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4244482%2F5fd6210a0f2e2b576fdab25083327fee%2Ftttt.jpg?generation=1584156173865554&amp;alt=media)",
      "votes": null
    },
    {
      "id": "771336",
      "postDate": "03/14/2020 03:23:25",
      "content": "<p>Accuracy of a set is evaluated by just cross-checking the highest softmax output and the correct labeled class.It is not depended on how high is the softmax output. To make it clearer, here are some numbers.</p>\n\n<p>Suppose there are 3 classes- dog, cat and horse. For our case, the correct class is horse . Now, the output of the softmax is [0.9, 0.1]. For this loss ~0.37. The classifier will predict that it is a horse. Take another case where softmax output is [0.6, 0.4]. Loss ~0.6. The classifier will still predict that it is a horse. But surely, the loss has increased. So, it is all about the output distribution.</p>",
      "rawMarkdown": "Accuracy of a set is evaluated by just cross-checking the highest softmax output and the correct labeled class.It is not depended on how high is the softmax output. To make it clearer, here are some numbers.\n\nSuppose there are 3 classes- dog, cat and horse. For our case, the correct class is horse . Now, the output of the softmax is [0.9, 0.1]. For this loss ~0.37. The classifier will predict that it is a horse. Take another case where softmax output is [0.6, 0.4]. Loss ~0.6. The classifier will still predict that it is a horse. But surely, the loss has increased. So, it is all about the output distribution.",
      "votes": null
    },
    {
      "id": "771359",
      "postDate": "03/14/2020 04:39:33",
      "content": "<p><a href=\"/tikoboss\">@tikoboss</a> Thanks a lot. That was insightful.</p>",
      "rawMarkdown": "tikoboss Thanks a lot. That was insightful.",
      "votes": null
    },
    {
      "id": "772250",
      "postDate": "03/15/2020 08:28:09",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 771332,
      "author_name": "tikoboss",
      "author_url": "",
      "post_date": "03/14/2020 03:19:12",
      "content": "<p>Building on Ankur's answer and the comment underneath it, I think the following scenario is possible, while I have no proof of it. Two phenomenons might be happening at the same time :</p>\n\n<pre><code>Some examples with borderline predictions get predicted better and so their output class changes (eg a cat image predicted at 0.4 to be a cat and 0.6 to be a horse becomes predicted 0.4 to be a horse and 0.6 to be a cat). Thanks to this, accuracy increases while loss decreases.\n\nSome examples with very bad predictions keep getting worse (eg a cat image predicted at 0.8 to be a horse becomes predicted at 0.9 to be a horse) AND/OR (more probable, in particular for multi-class ?) some examples with very good predictions get a little worse (eg a cat image predicted at 0.9 to be a cat becomes predicted at 0.8 to be a cat). With this phenomenon, loss increases while accuracy stays the same.\n</code></pre>\n\n<p>So if phenomenon 2 kicks in at some point, on lots of examples (eg for a specific class which is not well understood for some reason) and/or with a loss increase stonger than the loss decrease you gain from 1., then you might find yourself in your scenario.\nOnce again, maybe this is not what's happening, but I think that being able to come up with such scenarios must remind us of the sometimes slippery relationship between (cross-entropy)loss and accuracy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771333,
      "author_name": "tikoboss",
      "author_url": "",
      "post_date": "03/14/2020 03:20:29",
      "content": "<p><strong>A model can overfit to cross entropy loss without over overfitting to accuracy.</strong></p>\n\n<p>```\nThere is a key difference between the two types of loss:</p>\n\n<pre><code>Accuracy measures whether you get the prediction right\nCross entropy measures how confident you are about a prediction\n</code></pre>\n\n<p>For example, if an image of a cat is passed into two models. Model A predicts {cat: 0.9, dog: 0.1} and model B predicts {cat: 0.6, dog: 0.4}. Both model will score the same accuracy, but model A will have a lower loss.</p>\n\n<p>Because of this the model will try to be more and more confident to minimize loss. It works fine in training stage, but in validation stage it will perform poorly in term of loss. For example, for some borderline images, being confident e.g. {cat: 0.9, dog: 0.1} will give higher loss than being uncertain e.g. {cat: 0.6, dog: 0.4}</p>\n\n<p>In short, cross entropy loss measures the calibration of a model. Mis-calibration is a common issue to modern neuronal networks. They tend to be over-confident. On Calibration of Modern Neural Networks talks about it in great details.\n```</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771334,
      "author_name": "tikoboss",
      "author_url": "",
      "post_date": "03/14/2020 03:21:45",
      "content": "<p>```\nMany answers focus on the mathematical calculation explaining how is this possible. But they don't explain why it becomes so. And they cannot suggest how to digger further to be more clear.</p>\n\n<p>I have 3 hypothesis. And suggest some experiments to verify them. Hopefully it can help explain this problem.</p>\n\n<ol>\n<li>Label is noisy. Compare the false predictions when val_loss is minimum and val_acc is maximum. Check whether these sample are correctly labelled.</li>\n<li>[Less likely] The model doesn't have enough aspect of information to be certain. Experiment with more and larger hidden layers.\n3.[A very wild guess] This is a case where the model is less certain about certain things as being trained longer. Such situation happens to human as well. When someone started to learn a technique, he is told exactly what is good or bad, what is certain things for (high certainty). When he goes through more cases and examples, he realizes sometimes certain border can be blur (less certain, higher loss), even though he can make better decisions (more accuracy). And he may eventually gets more certain when he becomes a master after going through a huge list of samples and lots of trial and errors (more training data). So in this case, I suggest experiment with adding more noise to the training data (not label) may be helpful.</li>\n</ol>\n\n<p>Don't argue about this by just saying if you disagree with these hypothesis. It will be more meaningful to discuss with experiments to verify them, no matter the results prove them right, or prove them wrong.\n```</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771335,
      "author_name": "tikoboss",
      "author_url": "",
      "post_date": "03/14/2020 03:22:55",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4244482%2F5fd6210a0f2e2b576fdab25083327fee%2Ftttt.jpg?generation=1584156173865554&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771336,
      "author_name": "tikoboss",
      "author_url": "",
      "post_date": "03/14/2020 03:23:25",
      "content": "<p>Accuracy of a set is evaluated by just cross-checking the highest softmax output and the correct labeled class.It is not depended on how high is the softmax output. To make it clearer, here are some numbers.</p>\n\n<p>Suppose there are 3 classes- dog, cat and horse. For our case, the correct class is horse . Now, the output of the softmax is [0.9, 0.1]. For this loss ~0.37. The classifier will predict that it is a horse. Take another case where softmax output is [0.6, 0.4]. Loss ~0.6. The classifier will still predict that it is a horse. But surely, the loss has increased. So, it is all about the output distribution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771359,
      "author_name": "ankitsajwan",
      "author_url": "",
      "post_date": "03/14/2020 04:39:33",
      "content": "<p><a href=\"/tikoboss\">@tikoboss</a> Thanks a lot. That was insightful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 772250,
      "author_name": "mielek",
      "author_url": "",
      "post_date": "03/15/2020 08:28:09",
      "content": "<p>Nice</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "771329": "The below are snips of my training session.\nCan anyone explain me how validation loss and validation accuracy are increasing at the same time while training. How is this possible? Ideally, the val loss should decrease with more number of epochs and thus correspondingly val acc would increase. \n\nThe model was trained for 50-60 epochs with mixup, now i am trying to use gridmask on the pretrained mixup model, over in this training session.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2Ff195b1bb3983402f9bbaa42adbbff618%2Fval_loss.JPG?generation=1584154373076461&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4283293%2F773a327d7f656b700e5aa27484d89758%2Fval_acc.JPG?generation=1584154477514714&amp;alt=media)",
    "771332": "Building on Ankur's answer and the comment underneath it, I think the following scenario is possible, while I have no proof of it. Two phenomenons might be happening at the same time :\n\n    Some examples with borderline predictions get predicted better and so their output class changes (eg a cat image predicted at 0.4 to be a cat and 0.6 to be a horse becomes predicted 0.4 to be a horse and 0.6 to be a cat). Thanks to this, accuracy increases while loss decreases.\n\n    Some examples with very bad predictions keep getting worse (eg a cat image predicted at 0.8 to be a horse becomes predicted at 0.9 to be a horse) AND/OR (more probable, in particular for multi-class ?) some examples with very good predictions get a little worse (eg a cat image predicted at 0.9 to be a cat becomes predicted at 0.8 to be a cat). With this phenomenon, loss increases while accuracy stays the same.\n\nSo if phenomenon 2 kicks in at some point, on lots of examples (eg for a specific class which is not well understood for some reason) and/or with a loss increase stonger than the loss decrease you gain from 1., then you might find yourself in your scenario.\nOnce again, maybe this is not what's happening, but I think that being able to come up with such scenarios must remind us of the sometimes slippery relationship between (cross-entropy)loss and accuracy.",
    "771333": "**A model can overfit to cross entropy loss without over overfitting to accuracy.**\n\n```\nThere is a key difference between the two types of loss:\n\n    Accuracy measures whether you get the prediction right\n    Cross entropy measures how confident you are about a prediction\n\nFor example, if an image of a cat is passed into two models. Model A predicts {cat: 0.9, dog: 0.1} and model B predicts {cat: 0.6, dog: 0.4}. Both model will score the same accuracy, but model A will have a lower loss.\n\nBecause of this the model will try to be more and more confident to minimize loss. It works fine in training stage, but in validation stage it will perform poorly in term of loss. For example, for some borderline images, being confident e.g. {cat: 0.9, dog: 0.1} will give higher loss than being uncertain e.g. {cat: 0.6, dog: 0.4}\n\nIn short, cross entropy loss measures the calibration of a model. Mis-calibration is a common issue to modern neuronal networks. They tend to be over-confident. On Calibration of Modern Neural Networks talks about it in great details.\n```",
    "771334": "```\nMany answers focus on the mathematical calculation explaining how is this possible. But they don't explain why it becomes so. And they cannot suggest how to digger further to be more clear.\n\nI have 3 hypothesis. And suggest some experiments to verify them. Hopefully it can help explain this problem.\n\n   1. Label is noisy. Compare the false predictions when val_loss is minimum and val_acc is maximum. Check whether these sample are correctly labelled.\n   2.  [Less likely] The model doesn't have enough aspect of information to be certain. Experiment with more and larger hidden layers.\n    3.[A very wild guess] This is a case where the model is less certain about certain things as being trained longer. Such situation happens to human as well. When someone started to learn a technique, he is told exactly what is good or bad, what is certain things for (high certainty). When he goes through more cases and examples, he realizes sometimes certain border can be blur (less certain, higher loss), even though he can make better decisions (more accuracy). And he may eventually gets more certain when he becomes a master after going through a huge list of samples and lots of trial and errors (more training data). So in this case, I suggest experiment with adding more noise to the training data (not label) may be helpful.\n\nDon't argue about this by just saying if you disagree with these hypothesis. It will be more meaningful to discuss with experiments to verify them, no matter the results prove them right, or prove them wrong.\n```",
    "771335": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4244482%2F5fd6210a0f2e2b576fdab25083327fee%2Ftttt.jpg?generation=1584156173865554&amp;alt=media)",
    "771336": "Accuracy of a set is evaluated by just cross-checking the highest softmax output and the correct labeled class.It is not depended on how high is the softmax output. To make it clearer, here are some numbers.\n\nSuppose there are 3 classes- dog, cat and horse. For our case, the correct class is horse . Now, the output of the softmax is [0.9, 0.1]. For this loss ~0.37. The classifier will predict that it is a horse. Take another case where softmax output is [0.6, 0.4]. Loss ~0.6. The classifier will still predict that it is a horse. But surely, the loss has increased. So, it is all about the output distribution.",
    "771359": "tikoboss Thanks a lot. That was insightful.",
    "772250": "Nice"
  },
  "source": "meta"
}