{
  "id": 218342,
  "title": "CosineAnnealingWarmRestarts. Never reach the top accuracy again.",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/218342",
  "author_name": "",
  "post_date": "2021-02-10T07:32:04.126081600Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm using the following learning rate.</p>\n<blockquote>\n  <p>torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(optimizer, T_0=5, T_mult=1, eta_min=0.0001, last_epoch=-1)</p>\n</blockquote>\n<p>And my validation accuracy never reaches the top accuracy again.</p>\n<blockquote>\n  <p>fold 1<br>\n  epoch = 1 validation accuracy = 0.8715<br>\n  epoch = 2 validation accuracy = 0.8881<br>\n  epoch = 3 validation accuracy = 0.8853<br>\n  epoch = 4 validation accuracy = 0.8954<br>\n  epoch = 5 validation accuracy = 0.9046  &lt;- top<br>\n  epoch = 6 validation accuracy = 0.8881<br>\n  epoch = 7 validation accuracy = 0.8862<br>\n  epoch = 8 validation accuracy = 0.8823<br>\n  epoch = 9 validation accuracy = 0.8874<br>\n  epoch = 10 validation accuracy = 0.8881<br>\n  fold 2<br>\n  epoch = 1 validation accuracy = 0.8698<br>\n  epoch = 2 validation accuracy = 0.8848<br>\n  epoch = 3 validation accuracy = 0.8864<br>\n  epoch = 4 validation accuracy = 0.8902<br>\n  epoch = 5 validation accuracy = 0.8962  &lt;- top<br>\n  epoch = 6 validation accuracy = 0.8782<br>\n  epoch = 7 validation accuracy = 0.8846<br>\n  epoch = 8 validation accuracy = 0.8855<br>\n  epoch = 9 validation accuracy = 0.8850<br>\n  epoch = 10 validation accuracy = 0.8843<br>\n  …</p>\n</blockquote>\n<p>Why does this happen? This kind of cosine annealing strategy is helpful to dig into the flat minima region. So I thought at least the accuracy result of the 10th epoch should be better than that of 5th.</p>\n<p>Here are more details about my implementation:<br>\nBatch size : 6<br>\nOptimizer : Adam</p>",
  "messages": [
    {
      "id": "1194417",
      "postDate": "02/10/2021 07:32:04",
      "content": "<p>I'm using the following learning rate.</p>\n<blockquote>\n  <p>torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(optimizer, T_0=5, T_mult=1, eta_min=0.0001, last_epoch=-1)</p>\n</blockquote>\n<p>And my validation accuracy never reaches the top accuracy again.</p>\n<blockquote>\n  <p>fold 1<br>\n  epoch = 1 validation accuracy = 0.8715<br>\n  epoch = 2 validation accuracy = 0.8881<br>\n  epoch = 3 validation accuracy = 0.8853<br>\n  epoch = 4 validation accuracy = 0.8954<br>\n  epoch = 5 validation accuracy = 0.9046  &lt;- top<br>\n  epoch = 6 validation accuracy = 0.8881<br>\n  epoch = 7 validation accuracy = 0.8862<br>\n  epoch = 8 validation accuracy = 0.8823<br>\n  epoch = 9 validation accuracy = 0.8874<br>\n  epoch = 10 validation accuracy = 0.8881<br>\n  fold 2<br>\n  epoch = 1 validation accuracy = 0.8698<br>\n  epoch = 2 validation accuracy = 0.8848<br>\n  epoch = 3 validation accuracy = 0.8864<br>\n  epoch = 4 validation accuracy = 0.8902<br>\n  epoch = 5 validation accuracy = 0.8962  &lt;- top<br>\n  epoch = 6 validation accuracy = 0.8782<br>\n  epoch = 7 validation accuracy = 0.8846<br>\n  epoch = 8 validation accuracy = 0.8855<br>\n  epoch = 9 validation accuracy = 0.8850<br>\n  epoch = 10 validation accuracy = 0.8843<br>\n  …</p>\n</blockquote>\n<p>Why does this happen? This kind of cosine annealing strategy is helpful to dig into the flat minima region. So I thought at least the accuracy result of the 10th epoch should be better than that of 5th.</p>\n<p>Here are more details about my implementation:<br>\nBatch size : 6<br>\nOptimizer : Adam</p>",
      "rawMarkdown": "I'm using the following learning rate.\n> torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(optimizer, T_0=5, T_mult=1, eta_min=0.0001, last_epoch=-1)\n\nAnd my validation accuracy never reaches the top accuracy again.\n> fold 1\nepoch = 1 validation accuracy = 0.8715\nepoch = 2 validation accuracy = 0.8881\nepoch = 3 validation accuracy = 0.8853\nepoch = 4 validation accuracy = 0.8954\nepoch = 5 validation accuracy = 0.9046  <- top\nepoch = 6 validation accuracy = 0.8881\nepoch = 7 validation accuracy = 0.8862\nepoch = 8 validation accuracy = 0.8823\nepoch = 9 validation accuracy = 0.8874\nepoch = 10 validation accuracy = 0.8881\nfold 2\nepoch = 1 validation accuracy = 0.8698\nepoch = 2 validation accuracy = 0.8848\nepoch = 3 validation accuracy = 0.8864\nepoch = 4 validation accuracy = 0.8902\nepoch = 5 validation accuracy = 0.8962  <- top\nepoch = 6 validation accuracy = 0.8782\nepoch = 7 validation accuracy = 0.8846\nepoch = 8 validation accuracy = 0.8855\nepoch = 9 validation accuracy = 0.8850\nepoch = 10 validation accuracy = 0.8843\n...\n\nWhy does this happen? This kind of cosine annealing strategy is helpful to dig into the flat minima region. So I thought at least the accuracy result of the 10th epoch should be better than that of 5th.\n\nHere are more details about my implementation:\nBatch size : 6\nOptimizer : Adam",
      "votes": null
    },
    {
      "id": "1194539",
      "postDate": "02/10/2021 08:29:38",
      "content": "<p>My guessing would be after epoch 5 and epoch 6. It starts to overfit. If you monitor your training loss vs validation loss, you should be able to see it. I have something similar but after around epoch 8 and epoch 9. </p>",
      "rawMarkdown": "My guessing would be after epoch 5 and epoch 6. It starts to overfit. If you monitor your training loss vs validation loss, you should be able to see it. I have something similar but after around epoch 8 and epoch 9.",
      "votes": null
    },
    {
      "id": "1194639",
      "postDate": "02/10/2021 09:36:38",
      "content": "<p>Maybe it's already near a minimum in epoch 5 and then it's starts increasing the learning rate heavily again. You can try to reduce the learning rate the new cycle starts with if you set the T_mult to like 0.5 or even lower. </p>",
      "rawMarkdown": "Maybe it's already near a minimum in epoch 5 and then it's starts increasing the learning rate heavily again. You can try to reduce the learning rate the new cycle starts with if you set the T_mult to like 0.5 or even lower.",
      "votes": null
    },
    {
      "id": "1194781",
      "postDate": "02/10/2021 11:13:36",
      "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> yes i think it is possible.</p>",
      "rawMarkdown": "alexanderriedel yes i think it is possible.",
      "votes": null
    },
    {
      "id": "1195080",
      "postDate": "02/10/2021 14:34:06",
      "content": "<p>I noticed the same behavior, when I set T_0 = 10, I reach maximum accuracy of 0.904 or 0.906 then after that it drops like in your case :), but train and val loss are very near.</p>",
      "rawMarkdown": "I noticed the same behavior, when I set T_0 = 10, I reach maximum accuracy of 0.904 or 0.906 then after that it drops like in your case :), but train and val loss are very near.",
      "votes": null
    },
    {
      "id": "1195577",
      "postDate": "02/10/2021 23:42:06",
      "content": "<p>You could also save your model on a best epoch vs every fold to be safe.</p>",
      "rawMarkdown": "You could also save your model on a best epoch vs every fold to be safe.",
      "votes": null
    },
    {
      "id": "1195580",
      "postDate": "02/10/2021 23:46:01",
      "content": "<p>You can plot what that looks like for 10 epochs:<br>\n<a href=\"https://pasteboard.co/JNMfcxf.png\" target=\"_blank\">https://pasteboard.co/JNMfcxf.png</a></p>\n<p>You can see that after 5 epochs the LR restarts again. </p>",
      "rawMarkdown": "You can plot what that looks like for 10 epochs:\nhttps://pasteboard.co/JNMfcxf.png\n\nYou can see that after 5 epochs the LR restarts again.",
      "votes": null
    },
    {
      "id": "1195593",
      "postDate": "02/11/2021 00:13:11",
      "content": "<p>Yes I already did that. This restarting technique sometimes looks unhelpful.</p>",
      "rawMarkdown": "Yes I already did that. This restarting technique sometimes looks unhelpful.",
      "votes": null
    },
    {
      "id": "1195594",
      "postDate": "02/11/2021 00:14:33",
      "content": "<p>Oh reducing T_mult might be a solution. But I decide to enlarge my initial learning rate and reduce the number of epochs. Thanks.</p>",
      "rawMarkdown": "Oh reducing T_mult might be a solution. But I decide to enlarge my initial learning rate and reduce the number of epochs. Thanks.",
      "votes": null
    },
    {
      "id": "1195595",
      "postDate": "02/11/2021 00:15:49",
      "content": "<p>I see. So I decide to reduce the number of whole epochs. Thanks afterall.</p>",
      "rawMarkdown": "I see. So I decide to reduce the number of whole epochs. Thanks afterall.",
      "votes": null
    },
    {
      "id": "1197143",
      "postDate": "02/12/2021 00:30:02",
      "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> mate, sorry for beginner question, but 0.5 is not possible? it says T_mult &gt;=1?</p>",
      "rawMarkdown": "alexanderriedel mate, sorry for beginner question, but 0.5 is not possible? it says T_mult >=1?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1194539,
      "author_name": "tom88jerry",
      "author_url": "",
      "post_date": "02/10/2021 08:29:38",
      "content": "<p>My guessing would be after epoch 5 and epoch 6. It starts to overfit. If you monitor your training loss vs validation loss, you should be able to see it. I have something similar but after around epoch 8 and epoch 9. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1195595,
          "author_name": "hihunjin",
          "author_url": "",
          "post_date": "02/11/2021 00:15:49",
          "content": "<p>I see. So I decide to reduce the number of whole epochs. Thanks afterall.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1194639,
      "author_name": "alexanderriedel",
      "author_url": "",
      "post_date": "02/10/2021 09:36:38",
      "content": "<p>Maybe it's already near a minimum in epoch 5 and then it's starts increasing the learning rate heavily again. You can try to reduce the learning rate the new cycle starts with if you set the T_mult to like 0.5 or even lower. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1194781,
          "author_name": "tom88jerry",
          "author_url": "",
          "post_date": "02/10/2021 11:13:36",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> yes i think it is possible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1195594,
          "author_name": "hihunjin",
          "author_url": "",
          "post_date": "02/11/2021 00:14:33",
          "content": "<p>Oh reducing T_mult might be a solution. But I decide to enlarge my initial learning rate and reduce the number of epochs. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1197143,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/12/2021 00:30:02",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> mate, sorry for beginner question, but 0.5 is not possible? it says T_mult &gt;=1?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1195080,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "02/10/2021 14:34:06",
      "content": "<p>I noticed the same behavior, when I set T_0 = 10, I reach maximum accuracy of 0.904 or 0.906 then after that it drops like in your case :), but train and val loss are very near.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1195577,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/10/2021 23:42:06",
      "content": "<p>You could also save your model on a best epoch vs every fold to be safe.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1195580,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/10/2021 23:46:01",
      "content": "<p>You can plot what that looks like for 10 epochs:<br>\n<a href=\"https://pasteboard.co/JNMfcxf.png\" target=\"_blank\">https://pasteboard.co/JNMfcxf.png</a></p>\n<p>You can see that after 5 epochs the LR restarts again. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1195593,
          "author_name": "hihunjin",
          "author_url": "",
          "post_date": "02/11/2021 00:13:11",
          "content": "<p>Yes I already did that. This restarting technique sometimes looks unhelpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1194417": "I'm using the following learning rate.\n> torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(optimizer, T_0=5, T_mult=1, eta_min=0.0001, last_epoch=-1)\n\nAnd my validation accuracy never reaches the top accuracy again.\n> fold 1\nepoch = 1 validation accuracy = 0.8715\nepoch = 2 validation accuracy = 0.8881\nepoch = 3 validation accuracy = 0.8853\nepoch = 4 validation accuracy = 0.8954\nepoch = 5 validation accuracy = 0.9046  <- top\nepoch = 6 validation accuracy = 0.8881\nepoch = 7 validation accuracy = 0.8862\nepoch = 8 validation accuracy = 0.8823\nepoch = 9 validation accuracy = 0.8874\nepoch = 10 validation accuracy = 0.8881\nfold 2\nepoch = 1 validation accuracy = 0.8698\nepoch = 2 validation accuracy = 0.8848\nepoch = 3 validation accuracy = 0.8864\nepoch = 4 validation accuracy = 0.8902\nepoch = 5 validation accuracy = 0.8962  <- top\nepoch = 6 validation accuracy = 0.8782\nepoch = 7 validation accuracy = 0.8846\nepoch = 8 validation accuracy = 0.8855\nepoch = 9 validation accuracy = 0.8850\nepoch = 10 validation accuracy = 0.8843\n...\n\nWhy does this happen? This kind of cosine annealing strategy is helpful to dig into the flat minima region. So I thought at least the accuracy result of the 10th epoch should be better than that of 5th.\n\nHere are more details about my implementation:\nBatch size : 6\nOptimizer : Adam",
    "1194539": "My guessing would be after epoch 5 and epoch 6. It starts to overfit. If you monitor your training loss vs validation loss, you should be able to see it. I have something similar but after around epoch 8 and epoch 9.",
    "1194639": "Maybe it's already near a minimum in epoch 5 and then it's starts increasing the learning rate heavily again. You can try to reduce the learning rate the new cycle starts with if you set the T_mult to like 0.5 or even lower.",
    "1194781": "alexanderriedel yes i think it is possible.",
    "1195080": "I noticed the same behavior, when I set T_0 = 10, I reach maximum accuracy of 0.904 or 0.906 then after that it drops like in your case :), but train and val loss are very near.",
    "1195577": "You could also save your model on a best epoch vs every fold to be safe.",
    "1195580": "You can plot what that looks like for 10 epochs:\nhttps://pasteboard.co/JNMfcxf.png\n\nYou can see that after 5 epochs the LR restarts again.",
    "1195593": "Yes I already did that. This restarting technique sometimes looks unhelpful.",
    "1195594": "Oh reducing T_mult might be a solution. But I decide to enlarge my initial learning rate and reduce the number of epochs. Thanks.",
    "1195595": "I see. So I decide to reduce the number of whole epochs. Thanks afterall.",
    "1197143": "alexanderriedel mate, sorry for beginner question, but 0.5 is not possible? it says T_mult >=1?"
  },
  "source": "meta"
}