{
  "id": 209964,
  "title": "How to choose a better scheduler?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/209964",
  "author_name": "",
  "post_date": "2021-01-09T08:04:24.789738700Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi,how can we choose a better cheduler? I test the ReduceLROnPlateau and the CosineAnnealingWarmRestarts schedulers.I find the params of 'T_max' have a big influence in CosineAnnealingWarmRestarts scheduler.And the ReduceLROnPlateau scheduler have a factor to decrease lr,how can we choose a better params?I try factor in 0.7,0.5,0.1.Anyone have a good advice?</p>",
  "messages": [
    {
      "id": "1145532",
      "postDate": "01/09/2021 08:04:24",
      "content": "<p>Hi,how can we choose a better cheduler? I test the ReduceLROnPlateau and the CosineAnnealingWarmRestarts schedulers.I find the params of 'T_max' have a big influence in CosineAnnealingWarmRestarts scheduler.And the ReduceLROnPlateau scheduler have a factor to decrease lr,how can we choose a better params?I try factor in 0.7,0.5,0.1.Anyone have a good advice?</p>",
      "rawMarkdown": "Hi,how can we choose a better cheduler? I test the ReduceLROnPlateau and the CosineAnnealingWarmRestarts schedulers.I find the params of 'T_max' have a big influence in CosineAnnealingWarmRestarts scheduler.And the ReduceLROnPlateau scheduler have a factor to decrease lr,how can we choose a better params?I try factor in 0.7,0.5,0.1.Anyone have a good advice?",
      "votes": null
    },
    {
      "id": "1225530",
      "postDate": "03/03/2021 17:02:33",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/bcwang\" target=\"_blank\">@bcwang</a>, not expected to find this thread empty.  <br>\nI have tried CosineAnnealingWarmRestarts and CosineAnnealingLR with parameters T_max=5.0 and min LR=1e-6.<br>\nI have followed the 3 stage training strategy, thanks to <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> and <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>, CosineAnnealingWarmRestarts worked better.</p>",
      "rawMarkdown": "Hey @bcwang, not expected to find this thread empty.  \nI have tried CosineAnnealingWarmRestarts and CosineAnnealingLR with parameters T_max=5.0 and min LR=1e-6.\nI have followed the 3 stage training strategy, thanks to @yasufuminakama and @ammarali32, CosineAnnealingWarmRestarts worked better.",
      "votes": null
    },
    {
      "id": "1225906",
      "postDate": "03/04/2021 03:54:22",
      "content": "<p>This is a great question. I have noticed that the scheduler has a big influence in this competition more than other competitions.</p>",
      "rawMarkdown": "This is a great question. I have noticed that the scheduler has a big influence in this competition more than other competitions.",
      "votes": null
    },
    {
      "id": "1225968",
      "postDate": "03/04/2021 05:31:29",
      "content": "<p>Kind of surprising how big of an impact this has had on performance. Never had a problem where learning rate has had such a large impact. Typically it is just a matter of speed to convergence but for this it also greatly changes the convergence point</p>",
      "rawMarkdown": "Kind of surprising how big of an impact this has had on performance. Never had a problem where learning rate has had such a large impact. Typically it is just a matter of speed to convergence but for this it also greatly changes the convergence point",
      "votes": null
    },
    {
      "id": "1225971",
      "postDate": "03/04/2021 05:34:50",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, can you share some tips. Also, I have doubt, which scheduler is better if we're training a model continuously for 15 epochs and what parameters should we take.</p>",
      "rawMarkdown": "Hello @cdeotte, can you share some tips. Also, I have doubt, which scheduler is better if we're training a model continuously for 15 epochs and what parameters should we take.",
      "votes": null
    },
    {
      "id": "1225983",
      "postDate": "03/04/2021 05:50:34",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>, which scheduler and giving what parameters you find the scheduler impactful in training, can you give some suggestion.</p>",
      "rawMarkdown": "Hi, @ryches, which scheduler and giving what parameters you find the scheduler impactful in training, can you give some suggestion.",
      "votes": null
    },
    {
      "id": "1225990",
      "postDate": "03/04/2021 06:02:00",
      "content": "<p>I agree <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> . In this comp learning rate and scheduler has had a bigger impact than any comp i have been in.</p>",
      "rawMarkdown": "I agree @ryches . In this comp learning rate and scheduler has had a bigger impact than any comp i have been in.",
      "votes": null
    },
    {
      "id": "1225991",
      "postDate": "03/04/2021 06:04:10",
      "content": "<p>The best strategy is to try a variety of schedulers and learning rates and see which one maximizes your CV. In this comp, you need to try many different ones. In other comps many learning rates and schedulers produce the same results so it isn't important to try many.</p>",
      "rawMarkdown": "The best strategy is to try a variety of schedulers and learning rates and see which one maximizes your CV. In this comp, you need to try many different ones. In other comps many learning rates and schedulers produce the same results so it isn't important to try many.",
      "votes": null
    },
    {
      "id": "1226006",
      "postDate": "03/04/2021 06:27:40",
      "content": "<p>I started with the cosine annealing one that was in the public kernels, but swapped it for reduce lr on plateau and reducing to 1/2 every time an epoch doesnt improve. I do not think it is optimal. Have also seen rather different results between adam and sgd with momentum. Dont really have the resources to explore it all that much. Been playing with other ideas</p>",
      "rawMarkdown": "I started with the cosine annealing one that was in the public kernels, but swapped it for reduce lr on plateau and reducing to 1/2 every time an epoch doesnt improve. I do not think it is optimal. Have also seen rather different results between adam and sgd with momentum. Dont really have the resources to explore it all that much. Been playing with other ideas",
      "votes": null
    },
    {
      "id": "1226133",
      "postDate": "03/04/2021 09:12:40",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Agree, event batch size affect the final result for me. Since data are too small, It is very sensitive to the Training proceduce.</p>",
      "rawMarkdown": "cdeotte Agree, event batch size affect the final result for me. Since data are too small, It is very sensitive to the Training proceduce.",
      "votes": null
    },
    {
      "id": "1226135",
      "postDate": "03/04/2021 09:14:35",
      "content": "<p>Do u try cyclic lr and checkpoint ensemble?</p>",
      "rawMarkdown": "Do u try cyclic lr and checkpoint ensemble?",
      "votes": null
    },
    {
      "id": "1226164",
      "postDate": "03/04/2021 09:34:40",
      "content": "<p>I've tried checkpoint ensembling in the sense I average best auc, best loss, final epoch. Results are very similar on validation to the best auc weights. Have not tried submitting with it</p>",
      "rawMarkdown": "I've tried checkpoint ensembling in the sense I average best auc, best loss, final epoch. Results are very similar on validation to the best auc weights. Have not tried submitting with it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1225530,
      "author_name": "sanchitvj",
      "author_url": "",
      "post_date": "03/03/2021 17:02:33",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/bcwang\" target=\"_blank\">@bcwang</a>, not expected to find this thread empty.  <br>\nI have tried CosineAnnealingWarmRestarts and CosineAnnealingLR with parameters T_max=5.0 and min LR=1e-6.<br>\nI have followed the 3 stage training strategy, thanks to <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> and <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>, CosineAnnealingWarmRestarts worked better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1225906,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/04/2021 03:54:22",
      "content": "<p>This is a great question. I have noticed that the scheduler has a big influence in this competition more than other competitions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1225971,
          "author_name": "sanchitvj",
          "author_url": "",
          "post_date": "03/04/2021 05:34:50",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, can you share some tips. Also, I have doubt, which scheduler is better if we're training a model continuously for 15 epochs and what parameters should we take.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225991,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/04/2021 06:04:10",
          "content": "<p>The best strategy is to try a variety of schedulers and learning rates and see which one maximizes your CV. In this comp, you need to try many different ones. In other comps many learning rates and schedulers produce the same results so it isn't important to try many.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1225968,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "03/04/2021 05:31:29",
      "content": "<p>Kind of surprising how big of an impact this has had on performance. Never had a problem where learning rate has had such a large impact. Typically it is just a matter of speed to convergence but for this it also greatly changes the convergence point</p>",
      "votes": null,
      "replies": [
        {
          "id": 1225983,
          "author_name": "sanchitvj",
          "author_url": "",
          "post_date": "03/04/2021 05:50:34",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>, which scheduler and giving what parameters you find the scheduler impactful in training, can you give some suggestion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225990,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/04/2021 06:02:00",
          "content": "<p>I agree <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> . In this comp learning rate and scheduler has had a bigger impact than any comp i have been in.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1226006,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/04/2021 06:27:40",
          "content": "<p>I started with the cosine annealing one that was in the public kernels, but swapped it for reduce lr on plateau and reducing to 1/2 every time an epoch doesnt improve. I do not think it is optimal. Have also seen rather different results between adam and sgd with momentum. Dont really have the resources to explore it all that much. Been playing with other ideas</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1226133,
          "author_name": "steamedsheep",
          "author_url": "",
          "post_date": "03/04/2021 09:12:40",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Agree, event batch size affect the final result for me. Since data are too small, It is very sensitive to the Training proceduce.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1226135,
      "author_name": "steamedsheep",
      "author_url": "",
      "post_date": "03/04/2021 09:14:35",
      "content": "<p>Do u try cyclic lr and checkpoint ensemble?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1226164,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/04/2021 09:34:40",
          "content": "<p>I've tried checkpoint ensembling in the sense I average best auc, best loss, final epoch. Results are very similar on validation to the best auc weights. Have not tried submitting with it</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1145532": "Hi,how can we choose a better cheduler? I test the ReduceLROnPlateau and the CosineAnnealingWarmRestarts schedulers.I find the params of 'T_max' have a big influence in CosineAnnealingWarmRestarts scheduler.And the ReduceLROnPlateau scheduler have a factor to decrease lr,how can we choose a better params?I try factor in 0.7,0.5,0.1.Anyone have a good advice?",
    "1225530": "Hey @bcwang, not expected to find this thread empty.  \nI have tried CosineAnnealingWarmRestarts and CosineAnnealingLR with parameters T_max=5.0 and min LR=1e-6.\nI have followed the 3 stage training strategy, thanks to @yasufuminakama and @ammarali32, CosineAnnealingWarmRestarts worked better.",
    "1225906": "This is a great question. I have noticed that the scheduler has a big influence in this competition more than other competitions.",
    "1225968": "Kind of surprising how big of an impact this has had on performance. Never had a problem where learning rate has had such a large impact. Typically it is just a matter of speed to convergence but for this it also greatly changes the convergence point",
    "1225971": "Hello @cdeotte, can you share some tips. Also, I have doubt, which scheduler is better if we're training a model continuously for 15 epochs and what parameters should we take.",
    "1225983": "Hi, @ryches, which scheduler and giving what parameters you find the scheduler impactful in training, can you give some suggestion.",
    "1225990": "I agree @ryches . In this comp learning rate and scheduler has had a bigger impact than any comp i have been in.",
    "1225991": "The best strategy is to try a variety of schedulers and learning rates and see which one maximizes your CV. In this comp, you need to try many different ones. In other comps many learning rates and schedulers produce the same results so it isn't important to try many.",
    "1226006": "I started with the cosine annealing one that was in the public kernels, but swapped it for reduce lr on plateau and reducing to 1/2 every time an epoch doesnt improve. I do not think it is optimal. Have also seen rather different results between adam and sgd with momentum. Dont really have the resources to explore it all that much. Been playing with other ideas",
    "1226133": "cdeotte Agree, event batch size affect the final result for me. Since data are too small, It is very sensitive to the Training proceduce.",
    "1226135": "Do u try cyclic lr and checkpoint ensemble?",
    "1226164": "I've tried checkpoint ensembling in the sense I average best auc, best loss, final epoch. Results are very similar on validation to the best auc weights. Have not tried submitting with it"
  },
  "source": "meta"
}