{
  "id": 133197,
  "title": "bad validation and \"jumpy\" predictions ...why?",
  "url": "/competitions/deepfake-detection-challenge/discussion/133197",
  "author_name": "",
  "post_date": "2020-03-01T09:21:55.863017400Z",
  "votes": 1,
  "comment_count": 14,
  "views": 0,
  "content": "<p>In many of my models, I noticed a similar behavior: the training loss continuously decreases, while the validation loss jumps around like crazy.</p>\n\n<p><em>Overfitting</em>!  ...would be the obvious response. However, upon inspecting the validation predictions a little closer, I noticed something very peculiar: it's completely unbalanced! Here is the prediction distribution of two consecutive training epochs (on a balanced shuffled faces images):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2F81e270347b4af1fc1eafbfc59573e845%2Fepoch_1.png?generation=1583053969500091&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2Faf71bb5cf8b4a0c39c1eceec89f82a98%2Fepoch_2.png?generation=1583053987577979&amp;alt=media\" alt=\"\"></p>\n\n<p>So, it does not look like overfitting per-se, but rather like \"unstable\" learning ...the question is of course: why is the training/prediction so unstable? (despite the training loss looks great)</p>\n\n<p>What I'm using is one of the many pre-trained models at <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>, followed by a single sigmoid dense neuron. I also tried with an additional in-between dense layer and dropouts, a few other optimizers, but results always seem to have similar behavior. Predictions mostly wanders from one extreme to the other.</p>\n\n<p>Can you help me understand why it is so?</p>",
  "messages": [
    {
      "id": "760428",
      "postDate": "03/01/2020 09:21:55",
      "content": "<p>In many of my models, I noticed a similar behavior: the training loss continuously decreases, while the validation loss jumps around like crazy.</p>\n\n<p><em>Overfitting</em>!  ...would be the obvious response. However, upon inspecting the validation predictions a little closer, I noticed something very peculiar: it's completely unbalanced! Here is the prediction distribution of two consecutive training epochs (on a balanced shuffled faces images):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2F81e270347b4af1fc1eafbfc59573e845%2Fepoch_1.png?generation=1583053969500091&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2Faf71bb5cf8b4a0c39c1eceec89f82a98%2Fepoch_2.png?generation=1583053987577979&amp;alt=media\" alt=\"\"></p>\n\n<p>So, it does not look like overfitting per-se, but rather like \"unstable\" learning ...the question is of course: why is the training/prediction so unstable? (despite the training loss looks great)</p>\n\n<p>What I'm using is one of the many pre-trained models at <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>, followed by a single sigmoid dense neuron. I also tried with an additional in-between dense layer and dropouts, a few other optimizers, but results always seem to have similar behavior. Predictions mostly wanders from one extreme to the other.</p>\n\n<p>Can you help me understand why it is so?</p>",
      "rawMarkdown": "In many of my models, I noticed a similar behavior: the training loss continuously decreases, while the validation loss jumps around like crazy.\n\n*Overfitting*!  ...would be the obvious response. However, upon inspecting the validation predictions a little closer, I noticed something very peculiar: it's completely unbalanced! Here is the prediction distribution of two consecutive training epochs (on a balanced shuffled faces images):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2F81e270347b4af1fc1eafbfc59573e845%2Fepoch_1.png?generation=1583053969500091&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2Faf71bb5cf8b4a0c39c1eceec89f82a98%2Fepoch_2.png?generation=1583053987577979&amp;alt=media)\n\nSo, it does not look like overfitting per-se, but rather like \"unstable\" learning ...the question is of course: why is the training/prediction so unstable? (despite the training loss looks great)\n\nWhat I'm using is one of the many pre-trained models at https://keras.io/applications/, followed by a single sigmoid dense neuron. I also tried with an additional in-between dense layer and dropouts, a few other optimizers, but results always seem to have similar behavior. Predictions mostly wanders from one extreme to the other.\n\nCan you help me understand why it is so?",
      "votes": null
    },
    {
      "id": "760599",
      "postDate": "03/01/2020 14:11:23",
      "content": "<p>Equal the number of FAKE and REAL photos</p>",
      "rawMarkdown": "Equal the number of FAKE and REAL photos",
      "votes": null
    },
    {
      "id": "760617",
      "postDate": "03/01/2020 14:29:29",
      "content": "<p>It's already the case ;)</p>",
      "rawMarkdown": "It's already the case ;)",
      "votes": null
    },
    {
      "id": "760793",
      "postDate": "03/01/2020 18:16:59",
      "content": "<p>This is pretty normal. labels gather around 0.0 and 1.0 means the model is pretty confident. what is NOT normal(or proves the model not good) is a even distribution from 0-1.</p>",
      "rawMarkdown": "This is pretty normal. labels gather around 0.0 and 1.0 means the model is pretty confident. what is NOT normal(or proves the model not good) is a even distribution from 0-1.",
      "votes": null
    },
    {
      "id": "760849",
      "postDate": "03/01/2020 19:52:09",
      "content": "<p>Well, I would find it \"normal\" if <em>both</em> 0 and 1 would be present in the prediction. ...but in one epoch, everything is predicted as 0, and after the next epoch everything is suddenly predicted as 1 ...I wouldn't call it confident, I'd rather call it nonsensical. ;) :P</p>",
      "rawMarkdown": "Well, I would find it \"normal\" if *both* 0 and 1 would be present in the prediction. ...but in one epoch, everything is predicted as 0, and after the next epoch everything is suddenly predicted as 1 ...I wouldn't call it confident, I'd rather call it nonsensical. ;) :P",
      "votes": null
    },
    {
      "id": "760862",
      "postDate": "03/01/2020 20:06:28",
      "content": "<p>oh, I thought that's one picture lol. xD. my apologies. Well that is pretty abnormal. One possibility is that your training/validation generator is not right. </p>",
      "rawMarkdown": "oh, I thought that's one picture lol. xD. my apologies. Well that is pretty abnormal. One possibility is that your training/validation generator is not right.",
      "votes": null
    },
    {
      "id": "760945",
      "postDate": "03/02/2020 00:13:00",
      "content": "<p>Is there anything off in your image augmentation that might cause an overflow (e.g.if doing with uint8, going over 255) or unusable images? Just a thought, that perhaps the input is meaningless and so it settles on a meaningless back and forth since loss will be 50% at 0 and 1.</p>",
      "rawMarkdown": "Is there anything off in your image augmentation that might cause an overflow (e.g.if doing with uint8, going over 255) or unusable images? Just a thought, that perhaps the input is meaningless and so it settles on a meaningless back and forth since loss will be 50% at 0 and 1.",
      "votes": null
    },
    {
      "id": "761156",
      "postDate": "03/02/2020 07:40:23",
      "content": "<p>Most likely, your network is unstable. Lower your learning rates. Put in some BatchNorm layers and check what happens. It could also be that your y_train or y_val or both is faulty.</p>",
      "rawMarkdown": "Most likely, your network is unstable. Lower your learning rates. Put in some BatchNorm layers and check what happens. It could also be that your y_train or y_val or both is faulty.",
      "votes": null
    },
    {
      "id": "761288",
      "postDate": "03/02/2020 11:07:13",
      "content": "<p>Are you shuffling your mini-batch?</p>",
      "rawMarkdown": "Are you shuffling your mini-batch?",
      "votes": null
    },
    {
      "id": "761707",
      "postDate": "03/02/2020 21:18:56",
      "content": "<p>I faced this before. \nThis epoch: more predictions near 0, val loss 0.5. \nNext epoch: gradient shifts predictions to near 1, because a lot of predictions near 0 were wrong, val loss 0.35. \nNext epoch, the same problem, shifts to 0 again, val loss 0.45. \nReason: LR too high, causing fluctuation. . \nSolution: decrease LR.  </p>\n\n<p>:)</p>",
      "rawMarkdown": "I faced this before. \nThis epoch: more predictions near 0, val loss 0.5. \nNext epoch: gradient shifts predictions to near 1, because a lot of predictions near 0 were wrong, val loss 0.35. \nNext epoch, the same problem, shifts to 0 again, val loss 0.45. \nReason: LR too high, causing fluctuation. . \nSolution: decrease LR.  \n\n:)",
      "votes": null
    },
    {
      "id": "762133",
      "postDate": "03/03/2020 08:24:44",
      "content": "<p>Yep</p>",
      "rawMarkdown": "Yep",
      "votes": null
    },
    {
      "id": "762135",
      "postDate": "03/03/2020 08:26:34",
      "content": "<p>Good attempt, but I verified. The input is just fine.</p>",
      "rawMarkdown": "Good attempt, but I verified. The input is just fine.",
      "votes": null
    },
    {
      "id": "762136",
      "postDate": "03/03/2020 08:31:51",
      "content": "<p>I think I tried that and that it only delayed the issue. I also switched to optimizers  with less or no momentum since I suspected it as well. But I haven't obtained good results yet. ...I'll report if I identify the reason and manage to make it converge.</p>",
      "rawMarkdown": "I think I tried that and that it only delayed the issue. I also switched to optimizers  with less or no momentum since I suspected it as well. But I haven't obtained good results yet. ...I'll report if I identify the reason and manage to make it converge.",
      "votes": null
    },
    {
      "id": "762138",
      "postDate": "03/03/2020 08:34:36",
      "content": "<p>I also think it's related to the network itself (mobilenet v2) ...which appears to be pretty sensitive since training other networks is more stable</p>",
      "rawMarkdown": "I also think it's related to the network itself (mobilenet v2) ...which appears to be pretty sensitive since training other networks is more stable",
      "votes": null
    },
    {
      "id": "762348",
      "postDate": "03/03/2020 13:02:50",
      "content": "<p>Hi <a href=\"/dagnelies\">@dagnelies</a> \ndoes it happen in the very beginning? (1 or 2 epoch) or it works for a while then it goes crazy?\nif first case: something is very wrong, that network cannot learn anything.</p>\n\n<p>if the network keep having this prediction, the accuracy will be no good at all,\nIdeally the accuracy should arround 0.8 even when the BCE loss start increase (getting worse)</p>\n\n<p>for me, I reduce lr from 0.001 to 0.00002, then network start stable for a while (then over fit again....)</p>",
      "rawMarkdown": "Hi @dagnelies \ndoes it happen in the very beginning? (1 or 2 epoch) or it works for a while then it goes crazy?\nif first case: something is very wrong, that network cannot learn anything.\n\nif the network keep having this prediction, the accuracy will be no good at all,\nIdeally the accuracy should arround 0.8 even when the BCE loss start increase (getting worse)\n\nfor me, I reduce lr from 0.001 to 0.00002, then network start stable for a while (then over fit again....)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 760599,
      "author_name": "",
      "author_url": "",
      "post_date": "03/01/2020 14:11:23",
      "content": "<p>Equal the number of FAKE and REAL photos</p>",
      "votes": null,
      "replies": [
        {
          "id": 760617,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "03/01/2020 14:29:29",
          "content": "<p>It's already the case ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760793,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "03/01/2020 18:16:59",
      "content": "<p>This is pretty normal. labels gather around 0.0 and 1.0 means the model is pretty confident. what is NOT normal(or proves the model not good) is a even distribution from 0-1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 760849,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "03/01/2020 19:52:09",
          "content": "<p>Well, I would find it \"normal\" if <em>both</em> 0 and 1 would be present in the prediction. ...but in one epoch, everything is predicted as 0, and after the next epoch everything is suddenly predicted as 1 ...I wouldn't call it confident, I'd rather call it nonsensical. ;) :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760862,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/01/2020 20:06:28",
          "content": "<p>oh, I thought that's one picture lol. xD. my apologies. Well that is pretty abnormal. One possibility is that your training/validation generator is not right. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760945,
      "author_name": "botgreet",
      "author_url": "",
      "post_date": "03/02/2020 00:13:00",
      "content": "<p>Is there anything off in your image augmentation that might cause an overflow (e.g.if doing with uint8, going over 255) or unusable images? Just a thought, that perhaps the input is meaningless and so it settles on a meaningless back and forth since loss will be 50% at 0 and 1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 762135,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "03/03/2020 08:26:34",
          "content": "<p>Good attempt, but I verified. The input is just fine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761156,
      "author_name": "akashnandi",
      "author_url": "",
      "post_date": "03/02/2020 07:40:23",
      "content": "<p>Most likely, your network is unstable. Lower your learning rates. Put in some BatchNorm layers and check what happens. It could also be that your y_train or y_val or both is faulty.</p>",
      "votes": null,
      "replies": [
        {
          "id": 762138,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "03/03/2020 08:34:36",
          "content": "<p>I also think it's related to the network itself (mobilenet v2) ...which appears to be pretty sensitive since training other networks is more stable</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761288,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "03/02/2020 11:07:13",
      "content": "<p>Are you shuffling your mini-batch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 762133,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "03/03/2020 08:24:44",
          "content": "<p>Yep</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761707,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "03/02/2020 21:18:56",
      "content": "<p>I faced this before. \nThis epoch: more predictions near 0, val loss 0.5. \nNext epoch: gradient shifts predictions to near 1, because a lot of predictions near 0 were wrong, val loss 0.35. \nNext epoch, the same problem, shifts to 0 again, val loss 0.45. \nReason: LR too high, causing fluctuation. . \nSolution: decrease LR.  </p>\n\n<p>:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 762136,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "03/03/2020 08:31:51",
          "content": "<p>I think I tried that and that it only delayed the issue. I also switched to optimizers  with less or no momentum since I suspected it as well. But I haven't obtained good results yet. ...I'll report if I identify the reason and manage to make it converge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 762348,
      "author_name": "gody7334",
      "author_url": "",
      "post_date": "03/03/2020 13:02:50",
      "content": "<p>Hi <a href=\"/dagnelies\">@dagnelies</a> \ndoes it happen in the very beginning? (1 or 2 epoch) or it works for a while then it goes crazy?\nif first case: something is very wrong, that network cannot learn anything.</p>\n\n<p>if the network keep having this prediction, the accuracy will be no good at all,\nIdeally the accuracy should arround 0.8 even when the BCE loss start increase (getting worse)</p>\n\n<p>for me, I reduce lr from 0.001 to 0.00002, then network start stable for a while (then over fit again....)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "760428": "In many of my models, I noticed a similar behavior: the training loss continuously decreases, while the validation loss jumps around like crazy.\n\n*Overfitting*!  ...would be the obvious response. However, upon inspecting the validation predictions a little closer, I noticed something very peculiar: it's completely unbalanced! Here is the prediction distribution of two consecutive training epochs (on a balanced shuffled faces images):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2F81e270347b4af1fc1eafbfc59573e845%2Fepoch_1.png?generation=1583053969500091&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2313664%2Faf71bb5cf8b4a0c39c1eceec89f82a98%2Fepoch_2.png?generation=1583053987577979&amp;alt=media)\n\nSo, it does not look like overfitting per-se, but rather like \"unstable\" learning ...the question is of course: why is the training/prediction so unstable? (despite the training loss looks great)\n\nWhat I'm using is one of the many pre-trained models at https://keras.io/applications/, followed by a single sigmoid dense neuron. I also tried with an additional in-between dense layer and dropouts, a few other optimizers, but results always seem to have similar behavior. Predictions mostly wanders from one extreme to the other.\n\nCan you help me understand why it is so?",
    "760599": "Equal the number of FAKE and REAL photos",
    "760617": "It's already the case ;)",
    "760793": "This is pretty normal. labels gather around 0.0 and 1.0 means the model is pretty confident. what is NOT normal(or proves the model not good) is a even distribution from 0-1.",
    "760849": "Well, I would find it \"normal\" if *both* 0 and 1 would be present in the prediction. ...but in one epoch, everything is predicted as 0, and after the next epoch everything is suddenly predicted as 1 ...I wouldn't call it confident, I'd rather call it nonsensical. ;) :P",
    "760862": "oh, I thought that's one picture lol. xD. my apologies. Well that is pretty abnormal. One possibility is that your training/validation generator is not right.",
    "760945": "Is there anything off in your image augmentation that might cause an overflow (e.g.if doing with uint8, going over 255) or unusable images? Just a thought, that perhaps the input is meaningless and so it settles on a meaningless back and forth since loss will be 50% at 0 and 1.",
    "761156": "Most likely, your network is unstable. Lower your learning rates. Put in some BatchNorm layers and check what happens. It could also be that your y_train or y_val or both is faulty.",
    "761288": "Are you shuffling your mini-batch?",
    "761707": "I faced this before. \nThis epoch: more predictions near 0, val loss 0.5. \nNext epoch: gradient shifts predictions to near 1, because a lot of predictions near 0 were wrong, val loss 0.35. \nNext epoch, the same problem, shifts to 0 again, val loss 0.45. \nReason: LR too high, causing fluctuation. . \nSolution: decrease LR.  \n\n:)",
    "762133": "Yep",
    "762135": "Good attempt, but I verified. The input is just fine.",
    "762136": "I think I tried that and that it only delayed the issue. I also switched to optimizers  with less or no momentum since I suspected it as well. But I haven't obtained good results yet. ...I'll report if I identify the reason and manage to make it converge.",
    "762138": "I also think it's related to the network itself (mobilenet v2) ...which appears to be pretty sensitive since training other networks is more stable",
    "762348": "Hi @dagnelies \ndoes it happen in the very beginning? (1 or 2 epoch) or it works for a while then it goes crazy?\nif first case: something is very wrong, that network cannot learn anything.\n\nif the network keep having this prediction, the accuracy will be no good at all,\nIdeally the accuracy should arround 0.8 even when the BCE loss start increase (getting worse)\n\nfor me, I reduce lr from 0.001 to 0.00002, then network start stable for a while (then over fit again....)"
  },
  "source": "meta"
}