{
  "id": 71282,
  "title": "Loss spike",
  "url": "/competitions/airbus-ship-detection/discussion/71282",
  "author_name": "",
  "post_date": "2018-11-12T08:02:14.874996500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have this weird bug with the loss spiking, which explains my terrible score. </p>\n\n<p>I'm using Adam optimizer with initial lr=0.001. For some reason, the loss initially decreases, then suddenly explodes and remains constant. This happens to both the training and validation loss, and it always occurs at a certain epoch. I've examined the data and network architecture; there doesn't seem to be anything wrong with them.</p>",
  "messages": [
    {
      "id": "419564",
      "postDate": "11/12/2018 08:02:14",
      "content": "<p>I have this weird bug with the loss spiking, which explains my terrible score. </p>\n\n<p>I'm using Adam optimizer with initial lr=0.001. For some reason, the loss initially decreases, then suddenly explodes and remains constant. This happens to both the training and validation loss, and it always occurs at a certain epoch. I've examined the data and network architecture; there doesn't seem to be anything wrong with them.</p>",
      "rawMarkdown": "I have this weird bug with the loss spiking, which explains my terrible score. \n\nI'm using Adam optimizer with initial lr=0.001. For some reason, the loss initially decreases, then suddenly explodes and remains constant. This happens to both the training and validation loss, and it always occurs at a certain epoch. I've examined the data and network architecture; there doesn't seem to be anything wrong with them.",
      "votes": null
    },
    {
      "id": "419617",
      "postDate": "11/12/2018 09:50:28",
      "content": "<p>All bugs aside, I have had loss pikes with adam when my initial learning rate is too big. </p>",
      "rawMarkdown": "All bugs aside, I have had loss pikes with adam when my initial learning rate is too big.",
      "votes": null
    },
    {
      "id": "419663",
      "postDate": "11/12/2018 11:17:50",
      "content": "<p>Try to reduce the lr and perhaps change the optimizer to SGD. I had something similar once but it happened to be a bug</p>",
      "rawMarkdown": "Try to reduce the lr and perhaps change the optimizer to SGD. I had something similar once but it happened to be a bug",
      "votes": null
    },
    {
      "id": "419792",
      "postDate": "11/12/2018 15:06:27",
      "content": "<p>I think your model got broken at that moment. Try to load the last epoch before it happen, reduce lr, and use gradient clipping.</p>",
      "rawMarkdown": "I think your model got broken at that moment. Try to load the last epoch before it happen, reduce lr, and use gradient clipping.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 419617,
      "author_name": "mxdbld",
      "author_url": "",
      "post_date": "11/12/2018 09:50:28",
      "content": "<p>All bugs aside, I have had loss pikes with adam when my initial learning rate is too big. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419663,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "11/12/2018 11:17:50",
      "content": "<p>Try to reduce the lr and perhaps change the optimizer to SGD. I had something similar once but it happened to be a bug</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419792,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "11/12/2018 15:06:27",
      "content": "<p>I think your model got broken at that moment. Try to load the last epoch before it happen, reduce lr, and use gradient clipping.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "419564": "I have this weird bug with the loss spiking, which explains my terrible score. \n\nI'm using Adam optimizer with initial lr=0.001. For some reason, the loss initially decreases, then suddenly explodes and remains constant. This happens to both the training and validation loss, and it always occurs at a certain epoch. I've examined the data and network architecture; there doesn't seem to be anything wrong with them.",
    "419617": "All bugs aside, I have had loss pikes with adam when my initial learning rate is too big.",
    "419663": "Try to reduce the lr and perhaps change the optimizer to SGD. I had something similar once but it happened to be a bug",
    "419792": "I think your model got broken at that moment. Try to load the last epoch before it happen, reduce lr, and use gradient clipping."
  },
  "source": "meta"
}