{
  "id": 38034,
  "title": "Model heartshock: Is this the so-called catastrophic forgetting?",
  "url": "/competitions/carvana-image-masking-challenge/discussion/38034",
  "author_name": "",
  "post_date": "2017-08-14T04:26:23.629419Z",
  "votes": -1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was training a U-network 512x512 batch-size 4 learning rate 0.01 Relu, when netwrok practically reset its training progress. Theoretically, it is possible if a change occurs in the lower layers.</p>\n\n<p>In your opinion, is a common behavior during the training of a neural network or should I suspect an implementation numerical  failure?</p>",
  "messages": [
    {
      "id": "213213",
      "postDate": "08/14/2017 04:26:23",
      "content": "<p>I was training a U-network 512x512 batch-size 4 learning rate 0.01 Relu, when netwrok practically reset its training progress. Theoretically, it is possible if a change occurs in the lower layers.</p>\n\n<p>In your opinion, is a common behavior during the training of a neural network or should I suspect an implementation numerical  failure?</p>",
      "rawMarkdown": "I was training a U-network 512x512 batch-size 4 learning rate 0.01 Relu, when netwrok practically reset its training progress. Theoretically, it is possible if a change occurs in the lower layers.\n\nIn your opinion, is a common behavior during the training of a neural network or should I suspect an implementation numerical  failure?\n\n\n\n\n  [1]: https://ibb.co/bFH8ra",
      "votes": null
    },
    {
      "id": "213615",
      "postDate": "08/15/2017 01:57:54",
      "content": "<p>Your learning rate is too large. try to reduce learning rate to 0.001 or even smaller.</p>",
      "rawMarkdown": "Your learning rate is too large. try to reduce learning rate to 0.001 or even smaller.",
      "votes": null
    },
    {
      "id": "214107",
      "postDate": "08/16/2017 01:51:17",
      "content": "<p>Follow AI.Zhag comment first. The other reason can be that the model sees a few full or partial black images due to a typo in the augmentation script.</p>",
      "rawMarkdown": "Follow AI.Zhag comment first. The other reason can be that the model sees a few full or partial black images due to a typo in the augmentation script.",
      "votes": null
    },
    {
      "id": "214227",
      "postDate": "08/16/2017 09:49:00",
      "content": "<p>I've reduced learning rate to 0.001 and it works better, thanks.</p>",
      "rawMarkdown": "I've reduced learning rate to 0.001 and it works better, thanks.",
      "votes": null
    },
    {
      "id": "214342",
      "postDate": "08/16/2017 15:34:27",
      "content": "<p>you should understand your problem as follows:</p>\n\n<ol>\n<li><p>during training, for some reason, an iteration of back propagation results in a 'very strong' gradient in some direction for the current batch</p></li>\n<li><p>And this gradient only works for the current batch.</p></li>\n<li><p>The gradient cannot work for the next batch. This cause suddenly high loss for the next batch.</p></li>\n</ol>\n\n<p>NOTE: in case of batch normalisation, a wrong running var and mean during training may also cause this problem.</p>\n\n<p>.</p>\n\n<p>what could have caused this?</p>\n\n<ul>\n<li><p>wrong data label</p></li>\n<li><p>bug, load wrong data (e.g. empty array for invalid file, etc) e.g. see <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229#214248\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229#214248</a>. Some file cannot be read by some python PIL?</p></li>\n<li><p>numerical instability in back propagation</p></li>\n<li><p>small batch size. (a batch is not the same as another batch. Or in this batch, all samples are already very correct but there is one sample that is very wrong. The learning fits only one sample)</p></li>\n<li><p>large learning rate</p></li>\n</ul>\n\n<p>.</p>\n\n<p>How to prevent this?</p>\n\n<ul>\n<li><p>check there is not problem with data  (e.g. visualise the images)</p></li>\n<li><p>check the values of the back propagation gradients ... make sure there is no very large or outlier values</p></li>\n<li><p>simply restrict the values of gradient (e.g. clipping)</p></li>\n<li><p>simply restrict the update (e.g. regularization, weight decay, small learning rate)</p></li>\n<li><p>have a larger batch size</p></li>\n</ul>",
      "rawMarkdown": "you should understand your problem as follows:\n\n1. during training, for some reason, an iteration of back propagation results in a 'very strong' gradient in some direction for the current batch\n\n2. And this gradient only works for the current batch.\n\n3. The gradient cannot work for the next batch. This cause suddenly high loss for the next batch.\n\n\nNOTE: in case of batch normalisation, a wrong running var and mean during training may also cause this problem.\n\n.\n\nwhat could have caused this?\n\n- wrong data label\n\n- bug, load wrong data (e.g. empty array for invalid file, etc) e.g. see https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229#214248. Some file cannot be read by some python PIL?\n\n- numerical instability in back propagation\n\n- small batch size. (a batch is not the same as another batch. Or in this batch, all samples are already very correct but there is one sample that is very wrong. The learning fits only one sample)\n\n- large learning rate\n\n.\n\nHow to prevent this?\n\n- check there is not problem with data  (e.g. visualise the images)\n\n- check the values of the back propagation gradients ... make sure there is no very large or outlier values\n\n- simply restrict the values of gradient (e.g. clipping)\n\n- simply restrict the update (e.g. regularization, weight decay, small learning rate)\n\n- have a larger batch size",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 213615,
      "author_name": "zhanglj",
      "author_url": "",
      "post_date": "08/15/2017 01:57:54",
      "content": "<p>Your learning rate is too large. try to reduce learning rate to 0.001 or even smaller.</p>",
      "votes": null,
      "replies": [
        {
          "id": 214227,
          "author_name": "brasnold",
          "author_url": "",
          "post_date": "08/16/2017 09:49:00",
          "content": "<p>I've reduced learning rate to 0.001 and it works better, thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214342,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/16/2017 15:34:27",
          "content": "<p>you should understand your problem as follows:</p>\n\n<ol>\n<li><p>during training, for some reason, an iteration of back propagation results in a 'very strong' gradient in some direction for the current batch</p></li>\n<li><p>And this gradient only works for the current batch.</p></li>\n<li><p>The gradient cannot work for the next batch. This cause suddenly high loss for the next batch.</p></li>\n</ol>\n\n<p>NOTE: in case of batch normalisation, a wrong running var and mean during training may also cause this problem.</p>\n\n<p>.</p>\n\n<p>what could have caused this?</p>\n\n<ul>\n<li><p>wrong data label</p></li>\n<li><p>bug, load wrong data (e.g. empty array for invalid file, etc) e.g. see <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229#214248\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229#214248</a>. Some file cannot be read by some python PIL?</p></li>\n<li><p>numerical instability in back propagation</p></li>\n<li><p>small batch size. (a batch is not the same as another batch. Or in this batch, all samples are already very correct but there is one sample that is very wrong. The learning fits only one sample)</p></li>\n<li><p>large learning rate</p></li>\n</ul>\n\n<p>.</p>\n\n<p>How to prevent this?</p>\n\n<ul>\n<li><p>check there is not problem with data  (e.g. visualise the images)</p></li>\n<li><p>check the values of the back propagation gradients ... make sure there is no very large or outlier values</p></li>\n<li><p>simply restrict the values of gradient (e.g. clipping)</p></li>\n<li><p>simply restrict the update (e.g. regularization, weight decay, small learning rate)</p></li>\n<li><p>have a larger batch size</p></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214107,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "08/16/2017 01:51:17",
      "content": "<p>Follow AI.Zhag comment first. The other reason can be that the model sees a few full or partial black images due to a typo in the augmentation script.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "213213": "I was training a U-network 512x512 batch-size 4 learning rate 0.01 Relu, when netwrok practically reset its training progress. Theoretically, it is possible if a change occurs in the lower layers.\n\nIn your opinion, is a common behavior during the training of a neural network or should I suspect an implementation numerical  failure?\n\n\n\n\n  [1]: https://ibb.co/bFH8ra",
    "213615": "Your learning rate is too large. try to reduce learning rate to 0.001 or even smaller.",
    "214107": "Follow AI.Zhag comment first. The other reason can be that the model sees a few full or partial black images due to a typo in the augmentation script.",
    "214227": "I've reduced learning rate to 0.001 and it works better, thanks.",
    "214342": "you should understand your problem as follows:\n\n1. during training, for some reason, an iteration of back propagation results in a 'very strong' gradient in some direction for the current batch\n\n2. And this gradient only works for the current batch.\n\n3. The gradient cannot work for the next batch. This cause suddenly high loss for the next batch.\n\n\nNOTE: in case of batch normalisation, a wrong running var and mean during training may also cause this problem.\n\n.\n\nwhat could have caused this?\n\n- wrong data label\n\n- bug, load wrong data (e.g. empty array for invalid file, etc) e.g. see https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229#214248. Some file cannot be read by some python PIL?\n\n- numerical instability in back propagation\n\n- small batch size. (a batch is not the same as another batch. Or in this batch, all samples are already very correct but there is one sample that is very wrong. The learning fits only one sample)\n\n- large learning rate\n\n.\n\nHow to prevent this?\n\n- check there is not problem with data  (e.g. visualise the images)\n\n- check the values of the back propagation gradients ... make sure there is no very large or outlier values\n\n- simply restrict the values of gradient (e.g. clipping)\n\n- simply restrict the update (e.g. regularization, weight decay, small learning rate)\n\n- have a larger batch size"
  },
  "source": "meta"
}