{
  "id": 22304,
  "title": "why does the cost function jump around at the beginning of an epoch?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22304",
  "author_name": "",
  "post_date": "2016-07-15T21:29:31.397Z",
  "votes": null,
  "comment_count": 2,
  "views": 421,
  "content": "<p>I have some naive observations about how the training loss evolves - can anyone give me a better understanding what the optimizer is doing? I'm currently using keras, but I've noticed the same behaviour in Lasagne/theano.</p>\n\n<ol>\n<li>At the beginning of the epoch, the training loss tends to jump around a bit before settling down. Why? This behaviour occurs with the Adam and Adagrad optimizers. I naively expected that the learning rate should just take the value of the learning rate at the end of the previous epoch, but my hypothesis is that the learning rate is reset to the initialized value at the beginning of the epoch. Any thoughts about what's really going on? Is the amount of &quot;jumping around&quot; in the training loss at the very beginning of an epoch telling me something profound about how badly I've selected the initial value of the learning rate? I would expect the training loss to jump around if the learning rate is too high, but I also wondered if the training loss would jump around if the learning rate is too low and the updates are influenced by floating point errors.</li>\n</ol>\n\n<p>.</p>\n\n<ol start=\"2\">\n<li>I've also noticed that the training loss at the beginning of a new epoch is typically much lower than the training loss at the end of the previous epoch. Any ideas about why this is so?</li>\n</ol>",
  "messages": [
    {
      "id": "127879",
      "postDate": "07/15/2016 21:29:31",
      "content": "<p>I have some naive observations about how the training loss evolves - can anyone give me a better understanding what the optimizer is doing? I'm currently using keras, but I've noticed the same behaviour in Lasagne/theano.</p>\n\n<ol>\n<li>At the beginning of the epoch, the training loss tends to jump around a bit before settling down. Why? This behaviour occurs with the Adam and Adagrad optimizers. I naively expected that the learning rate should just take the value of the learning rate at the end of the previous epoch, but my hypothesis is that the learning rate is reset to the initialized value at the beginning of the epoch. Any thoughts about what's really going on? Is the amount of &quot;jumping around&quot; in the training loss at the very beginning of an epoch telling me something profound about how badly I've selected the initial value of the learning rate? I would expect the training loss to jump around if the learning rate is too high, but I also wondered if the training loss would jump around if the learning rate is too low and the updates are influenced by floating point errors.</li>\n</ol>\n\n<p>.</p>\n\n<ol start=\"2\">\n<li>I've also noticed that the training loss at the beginning of a new epoch is typically much lower than the training loss at the end of the previous epoch. Any ideas about why this is so?</li>\n</ol>",
      "rawMarkdown": "I have some naive observations about how the training loss evolves - can anyone give me a better understanding what the optimizer is doing? I'm currently using keras, but I've noticed the same behaviour in Lasagne/theano.\r\n\r\n1. \r\nAt the beginning of the epoch, the training loss tends to jump around a bit before settling down. Why? This behaviour occurs with the Adam and Adagrad optimizers. I naively expected that the learning rate should just take the value of the learning rate at the end of the previous epoch, but my hypothesis is that the learning rate is reset to the initialized value at the beginning of the epoch. Any thoughts about what's really going on? Is the amount of \"jumping around\" in the training loss at the very beginning of an epoch telling me something profound about how badly I've selected the initial value of the learning rate? I would expect the training loss to jump around if the learning rate is too high, but I also wondered if the training loss would jump around if the learning rate is too low and the updates are influenced by floating point errors.\r\n\r\n.\r\n\r\n2. \r\nI've also noticed that the training loss at the beginning of a new epoch is typically much lower than the training loss at the end of the previous epoch. Any ideas about why this is so?",
      "votes": null
    },
    {
      "id": "128024",
      "postDate": "07/17/2016 05:25:59",
      "content": "<p>Hi duck,</p>\n\n<p>It may be simple characteristic of statistics.<br>\nIf you calculate cost with small number of images, cost is unstable.</p>",
      "rawMarkdown": "Hi duck,\r\n\r\nIt may be simple characteristic of statistics.<br>\r\nIf you calculate cost with small number of images, cost is unstable.",
      "votes": null
    },
    {
      "id": "128026",
      "postDate": "07/17/2016 05:44:29",
      "content": "<p>[quote=toshi_k;128024]\nIt may be simple characteristic of statistics.<br>\nIf you calculate cost with small number of images, cost is unstable.\n[/quote]</p>\n\n<p>Thanks for your suggestion! </p>\n\n<p>I think that if this was the case that the cost function would be about as unstable at the beginning of the epoch as at the end. However, the cost function is much more stable at the end of each epoch than at the beginning. Any further thoughts?</p>",
      "rawMarkdown": "[quote=toshi_k;128024]\r\nIt may be simple characteristic of statistics.<br>\r\nIf you calculate cost with small number of images, cost is unstable.\r\n[/quote]\r\n\r\nThanks for your suggestion! \r\n\r\nI think that if this was the case that the cost function would be about as unstable at the beginning of the epoch as at the end. However, the cost function is much more stable at the end of each epoch than at the beginning. Any further thoughts?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 128024,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "07/17/2016 05:25:59",
      "content": "<p>Hi duck,</p>\n\n<p>It may be simple characteristic of statistics.<br>\nIf you calculate cost with small number of images, cost is unstable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 128026,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "07/17/2016 05:44:29",
      "content": "<p>[quote=toshi_k;128024]\nIt may be simple characteristic of statistics.<br>\nIf you calculate cost with small number of images, cost is unstable.\n[/quote]</p>\n\n<p>Thanks for your suggestion! </p>\n\n<p>I think that if this was the case that the cost function would be about as unstable at the beginning of the epoch as at the end. However, the cost function is much more stable at the end of each epoch than at the beginning. Any further thoughts?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "127879": "I have some naive observations about how the training loss evolves - can anyone give me a better understanding what the optimizer is doing? I'm currently using keras, but I've noticed the same behaviour in Lasagne/theano.\r\n\r\n1. \r\nAt the beginning of the epoch, the training loss tends to jump around a bit before settling down. Why? This behaviour occurs with the Adam and Adagrad optimizers. I naively expected that the learning rate should just take the value of the learning rate at the end of the previous epoch, but my hypothesis is that the learning rate is reset to the initialized value at the beginning of the epoch. Any thoughts about what's really going on? Is the amount of \"jumping around\" in the training loss at the very beginning of an epoch telling me something profound about how badly I've selected the initial value of the learning rate? I would expect the training loss to jump around if the learning rate is too high, but I also wondered if the training loss would jump around if the learning rate is too low and the updates are influenced by floating point errors.\r\n\r\n.\r\n\r\n2. \r\nI've also noticed that the training loss at the beginning of a new epoch is typically much lower than the training loss at the end of the previous epoch. Any ideas about why this is so?",
    "128024": "Hi duck,\r\n\r\nIt may be simple characteristic of statistics.<br>\r\nIf you calculate cost with small number of images, cost is unstable.",
    "128026": "[quote=toshi_k;128024]\r\nIt may be simple characteristic of statistics.<br>\r\nIf you calculate cost with small number of images, cost is unstable.\r\n[/quote]\r\n\r\nThanks for your suggestion! \r\n\r\nI think that if this was the case that the cost function would be about as unstable at the beginning of the epoch as at the end. However, the cost function is much more stable at the end of each epoch than at the beginning. Any further thoughts?"
  },
  "source": "meta"
}