{
  "id": 130007,
  "title": "Optimizer choice",
  "url": "/competitions/bengaliai-cv19/discussion/130007",
  "author_name": "",
  "post_date": "2020-02-11T16:37:31.283370500Z",
  "votes": 6,
  "comment_count": 28,
  "views": 0,
  "content": "<p>In my tests so far I only used Adam with different learning rates and ReduceOnPlateau configurations.\nDid anybody tried a different optimizer than Adam or Adam variants like RAdam ? What did you experiment ?</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": "742923",
      "postDate": "02/11/2020 16:37:31",
      "content": "<p>In my tests so far I only used Adam with different learning rates and ReduceOnPlateau configurations.\nDid anybody tried a different optimizer than Adam or Adam variants like RAdam ? What did you experiment ?</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "In my tests so far I only used Adam with different learning rates and ReduceOnPlateau configurations.\nDid anybody tried a different optimizer than Adam or Adam variants like RAdam ? What did you experiment ?\n\n\nCheers",
      "votes": null
    },
    {
      "id": "743020",
      "postDate": "02/11/2020 17:56:45",
      "content": "<p>I am currently trying RAdam with ReduceOnPlateau. With lower lr for classifier\nBut it seems to overfit. I have a big gap between LB and Val score, which is really strange </p>",
      "rawMarkdown": "I am currently trying RAdam with ReduceOnPlateau. With lower lr for classifier\nBut it seems to overfit. I have a big gap between LB and Val score, which is really strange",
      "votes": null
    },
    {
      "id": "743023",
      "postDate": "02/11/2020 18:01:30",
      "content": "<p>How big is the gap ? Maybe the problem is on the classifier, it's to complex for what you are trying to predict or you don't augment enough the images so you can generalize</p>",
      "rawMarkdown": "How big is the gap ? Maybe the problem is on the classifier, it's to complex for what you are trying to predict or you don't augment enough the images so you can generalize",
      "votes": null
    },
    {
      "id": "743034",
      "postDate": "02/11/2020 18:21:47",
      "content": "<p>In my experiments so far ReduceLROnPlateau converges quicker than OneCycle and they perform about the same. I've also been using Adam and AdamW but there's not much difference there either.</p>",
      "rawMarkdown": "In my experiments so far ReduceLROnPlateau converges quicker than OneCycle and they perform about the same. I've also been using Adam and AdamW but there's not much difference there either.",
      "votes": null
    },
    {
      "id": "743046",
      "postDate": "02/11/2020 18:36:20",
      "content": "<p>I also had better experiences with ReduceLROnPlateau. </p>",
      "rawMarkdown": "I also had better experiences with ReduceLROnPlateau.",
      "votes": null
    },
    {
      "id": "743061",
      "postDate": "02/11/2020 18:49:30",
      "content": "<p>For me Adam or Radam they perform very similar. </p>",
      "rawMarkdown": "For me Adam or Radam they perform very similar.",
      "votes": null
    },
    {
      "id": "743081",
      "postDate": "02/11/2020 19:12:41",
      "content": "<p>I'm curious how in <a href=\"/iafoss\">@iafoss</a> kernel mixup converged in 32 epochs. I'm using AdamW for 80 epochs and from my experience even after 100+ epochs score can improve.</p>\n\n<p>The other question is how to carefully choose minimal lr? I guess 1e-5 or 5e-6 should work fine.</p>",
      "rawMarkdown": "I'm curious how in @iafoss kernel mixup converged in 32 epochs. I'm using AdamW for 80 epochs and from my experience even after 100+ epochs score can improve.\n\nThe other question is how to carefully choose minimal lr? I guess 1e-5 or 5e-6 should work fine.",
      "votes": null
    },
    {
      "id": "743096",
      "postDate": "02/11/2020 19:34:57",
      "content": "<p>For me, adam generalizes better than adamw (worse in val loss at least ~0.01 lb metric)\nadam and radam perform similar\nmultistep scheduler with steps and gamma by intuition helps me generalize the best. </p>",
      "rawMarkdown": "For me, adam generalizes better than adamw (worse in val loss at least ~0.01 lb metric)\nadam and radam perform similar\nmultistep scheduler with steps and gamma by intuition helps me generalize the best.",
      "votes": null
    },
    {
      "id": "743102",
      "postDate": "02/11/2020 19:39:37",
      "content": "<p>What was your lowest validation loss? </p>",
      "rawMarkdown": "What was your lowest validation loss?",
      "votes": null
    },
    {
      "id": "743103",
      "postDate": "02/11/2020 19:41:21",
      "content": "<p>I have tried AdaBound. I found that with it the model trained bit faster than vanilla Adam. </p>",
      "rawMarkdown": "I have tried AdaBound. I found that with it the model trained bit faster than vanilla Adam.",
      "votes": null
    },
    {
      "id": "743142",
      "postDate": "02/11/2020 20:36:37",
      "content": "<p>I haven't reached below 0.095 yet</p>",
      "rawMarkdown": "I haven't reached below 0.095 yet",
      "votes": null
    },
    {
      "id": "743692",
      "postDate": "02/12/2020 07:41:43",
      "content": "<p>The main thing i have been considering for this is if the optimizer state gets loaded properly since I only have Kaggle to train my kernels. Over9000 worked best, but training for 100 epochs wasn't possible with it, since I could train for only 30 epochs in a single notebook, and loading the optimizer state next time would lead to a drop of ~1% in the CV. I'm still experimenting what to use now. Any ideas would be appreciated </p>",
      "rawMarkdown": "The main thing i have been considering for this is if the optimizer state gets loaded properly since I only have Kaggle to train my kernels. Over9000 worked best, but training for 100 epochs wasn't possible with it, since I could train for only 30 epochs in a single notebook, and loading the optimizer state next time would lead to a drop of ~1% in the CV. I'm still experimenting what to use now. Any ideas would be appreciated",
      "votes": null
    },
    {
      "id": "743872",
      "postDate": "02/12/2020 10:49:58",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> do you use ReduceOnPlateau ? </p>",
      "rawMarkdown": "drhabib do you use ReduceOnPlateau ?",
      "votes": null
    },
    {
      "id": "744218",
      "postDate": "02/12/2020 17:01:25",
      "content": "<p>So far I've only been using Adam and ReduceOnPlateau has been really helpful!\nBut then, I'm still way behind a lot of people!</p>",
      "rawMarkdown": "So far I've only been using Adam and ReduceOnPlateau has been really helpful!\nBut then, I'm still way behind a lot of people!",
      "votes": null
    },
    {
      "id": "744235",
      "postDate": "02/12/2020 17:12:48",
      "content": "<p>It's just a matter of trying different things, some people have more experience and know what knobs to tune first and for others it is just a matter of time, I am sure that you will have a breakthrough if you keep experimenting</p>",
      "rawMarkdown": "It's just a matter of trying different things, some people have more experience and know what knobs to tune first and for others it is just a matter of time, I am sure that you will have a breakthrough if you keep experimenting",
      "votes": null
    },
    {
      "id": "744247",
      "postDate": "02/12/2020 17:20:19",
      "content": "<p>I used <code>Adam</code> with <code>ReduceOnPlateau</code> and Adam with <code>OneCycle</code>. They are similar in my setup\n =) </p>",
      "rawMarkdown": "I used `Adam` with `ReduceOnPlateau` and Adam with `OneCycle`. They are similar in my setup\n =)",
      "votes": null
    },
    {
      "id": "744249",
      "postDate": "02/12/2020 17:20:57",
      "content": "<p>This is a good advice! You learn by making mistakes =) </p>",
      "rawMarkdown": "This is a good advice! You learn by making mistakes =)",
      "votes": null
    },
    {
      "id": "744516",
      "postDate": "02/12/2020 22:36:43",
      "content": "<p>I've started with Adam and gave RAdam some tries. However I didn't notice any significant difference between local CV or related LB. For now I switched back to using just Adam.</p>\n\n<p>Usually I start the model training with a learning rate of about 0.0002 - 0.00015. Lower seems to take forever to train altough I'am not sure it would lead to underfitting. Higher does give me some overfitting.</p>\n\n<p>I'am also using the ReduceOnPlateau. In some kernels a rather 'quick' reduce scheme is used (patience 3 and reduce factor of 0.5). I've experienced that a more relaxed scheme seems to give some more improvement. I'am now experimenting with a patience of 4 to 5 and factor around .7. I'am personally quitte happy with the improvements that it already gave.</p>",
      "rawMarkdown": "I've started with Adam and gave RAdam some tries. However I didn't notice any significant difference between local CV or related LB. For now I switched back to using just Adam.\n\nUsually I start the model training with a learning rate of about 0.0002 - 0.00015. Lower seems to take forever to train altough I'am not sure it would lead to underfitting. Higher does give me some overfitting.\n\nI'am also using the ReduceOnPlateau. In some kernels a rather 'quick' reduce scheme is used (patience 3 and reduce factor of 0.5). I've experienced that a more relaxed scheme seems to give some more improvement. I'am now experimenting with a patience of 4 to 5 and factor around .7. I'am personally quitte happy with the improvements that it already gave.",
      "votes": null
    },
    {
      "id": "746978",
      "postDate": "02/15/2020 19:59:42",
      "content": "<p>Same here!</p>",
      "rawMarkdown": "Same here!",
      "votes": null
    },
    {
      "id": "748619",
      "postDate": "02/17/2020 18:48:02",
      "content": "<p>Personally i've been having better results with Lookahead(SGD) and OneCycleLR by a bit margin! </p>",
      "rawMarkdown": "Personally i've been having better results with Lookahead(SGD) and OneCycleLR by a bit margin!",
      "votes": null
    },
    {
      "id": "749128",
      "postDate": "02/18/2020 10:31:27",
      "content": "<p>For me, ReduceOnPlateau converges more quickly than OneCycleCosineLR</p>",
      "rawMarkdown": "For me, ReduceOnPlateau converges more quickly than OneCycleCosineLR",
      "votes": null
    },
    {
      "id": "749870",
      "postDate": "02/19/2020 00:37:06",
      "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> A really simple question. Is the reduce LR on Plateau done when metric is no longer increasing or when val loss is no longer decreasing? I have observed cases where metric and validation loss are both increasing :)</p>",
      "rawMarkdown": "cswwp347724 A really simple question. Is the reduce LR on Plateau done when metric is no longer increasing or when val loss is no longer decreasing? I have observed cases where metric and validation loss are both increasing :)",
      "votes": null
    },
    {
      "id": "754324",
      "postDate": "02/23/2020 12:06:49",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a>  Sorry for late reply, i just set LR down when metric is no longer increase at the last patience epoch. It's possible that  metric and validation loss are both increasing, because root has more weight in final metric</p>",
      "rawMarkdown": "roguekk007  Sorry for late reply, i just set LR down when metric is no longer increase at the last patience epoch. It's possible that  metric and validation loss are both increasing, because root has more weight in final metric",
      "votes": null
    },
    {
      "id": "754333",
      "postDate": "02/23/2020 12:20:25",
      "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> Thanks for the reply! I think setting LR based on metric is a sound choice taken that the CV correlates so well with the LB</p>",
      "rawMarkdown": "cswwp347724 Thanks for the reply! I think setting LR based on metric is a sound choice taken that the CV correlates so well with the LB",
      "votes": null
    },
    {
      "id": "755148",
      "postDate": "02/24/2020 13:47:00",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Do you want team up?😁 </p>",
      "rawMarkdown": "roguekk007 Do you want team up?😁",
      "votes": null
    },
    {
      "id": "755163",
      "postDate": "02/24/2020 14:07:01",
      "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> Thank you very much for the offer! This is the first comp that looks relatively promising for me; so I’m really excited and it’s such a great opportunity for me to learn. I want to really push my single-model score and might not consider teaming up until before the merger deadline; best of luck till then👋</p>",
      "rawMarkdown": "cswwp347724 Thank you very much for the offer! This is the first comp that looks relatively promising for me; so I’m really excited and it’s such a great opportunity for me to learn. I want to really push my single-model score and might not consider teaming up until before the merger deadline; best of luck till then👋",
      "votes": null
    },
    {
      "id": "755173",
      "postDate": "02/24/2020 14:15:57",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Ok, good luck👍 </p>",
      "rawMarkdown": "roguekk007 Ok, good luck👍",
      "votes": null
    },
    {
      "id": "756461",
      "postDate": "02/25/2020 18:27:13",
      "content": "<p>I change small batchsize with small LR, it seems not too much improvement😣 <a href=\"/roguekk007\">@roguekk007</a> </p>",
      "rawMarkdown": "I change small batchsize with small LR, it seems not too much improvement😣 @roguekk007",
      "votes": null
    },
    {
      "id": "756708",
      "postDate": "02/26/2020 01:33:37",
      "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> Same for me... Stuck again. It seems like batch size and LR are not so-important details in this comp</p>",
      "rawMarkdown": "cswwp347724 Same for me... Stuck again. It seems like batch size and LR are not so-important details in this comp",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 743020,
      "author_name": "vladimirsydor",
      "author_url": "",
      "post_date": "02/11/2020 17:56:45",
      "content": "<p>I am currently trying RAdam with ReduceOnPlateau. With lower lr for classifier\nBut it seems to overfit. I have a big gap between LB and Val score, which is really strange </p>",
      "votes": null,
      "replies": [
        {
          "id": 743023,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/11/2020 18:01:30",
          "content": "<p>How big is the gap ? Maybe the problem is on the classifier, it's to complex for what you are trying to predict or you don't augment enough the images so you can generalize</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 743034,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "02/11/2020 18:21:47",
      "content": "<p>In my experiments so far ReduceLROnPlateau converges quicker than OneCycle and they perform about the same. I've also been using Adam and AdamW but there's not much difference there either.</p>",
      "votes": null,
      "replies": [
        {
          "id": 743046,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/11/2020 18:36:20",
          "content": "<p>I also had better experiences with ReduceLROnPlateau. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746978,
          "author_name": "yuanlin08",
          "author_url": "",
          "post_date": "02/15/2020 19:59:42",
          "content": "<p>Same here!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 743061,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "02/11/2020 18:49:30",
      "content": "<p>For me Adam or Radam they perform very similar. </p>",
      "votes": null,
      "replies": [
        {
          "id": 743872,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/12/2020 10:49:58",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> do you use ReduceOnPlateau ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 744247,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/12/2020 17:20:19",
          "content": "<p>I used <code>Adam</code> with <code>ReduceOnPlateau</code> and Adam with <code>OneCycle</code>. They are similar in my setup\n =) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749128,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/18/2020 10:31:27",
          "content": "<p>For me, ReduceOnPlateau converges more quickly than OneCycleCosineLR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749870,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/19/2020 00:37:06",
          "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> A really simple question. Is the reduce LR on Plateau done when metric is no longer increasing or when val loss is no longer decreasing? I have observed cases where metric and validation loss are both increasing :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754324,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/23/2020 12:06:49",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a>  Sorry for late reply, i just set LR down when metric is no longer increase at the last patience epoch. It's possible that  metric and validation loss are both increasing, because root has more weight in final metric</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754333,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/23/2020 12:20:25",
          "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> Thanks for the reply! I think setting LR based on metric is a sound choice taken that the CV correlates so well with the LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755148,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/24/2020 13:47:00",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Do you want team up?😁 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755163,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/24/2020 14:07:01",
          "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> Thank you very much for the offer! This is the first comp that looks relatively promising for me; so I’m really excited and it’s such a great opportunity for me to learn. I want to really push my single-model score and might not consider teaming up until before the merger deadline; best of luck till then👋</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755173,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/24/2020 14:15:57",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Ok, good luck👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756461,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/25/2020 18:27:13",
          "content": "<p>I change small batchsize with small LR, it seems not too much improvement😣 <a href=\"/roguekk007\">@roguekk007</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756708,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/26/2020 01:33:37",
          "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> Same for me... Stuck again. It seems like batch size and LR are not so-important details in this comp</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 743081,
      "author_name": "lightnezzofbeing",
      "author_url": "",
      "post_date": "02/11/2020 19:12:41",
      "content": "<p>I'm curious how in <a href=\"/iafoss\">@iafoss</a> kernel mixup converged in 32 epochs. I'm using AdamW for 80 epochs and from my experience even after 100+ epochs score can improve.</p>\n\n<p>The other question is how to carefully choose minimal lr? I guess 1e-5 or 5e-6 should work fine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 743096,
      "author_name": "timetraveller98",
      "author_url": "",
      "post_date": "02/11/2020 19:34:57",
      "content": "<p>For me, adam generalizes better than adamw (worse in val loss at least ~0.01 lb metric)\nadam and radam perform similar\nmultistep scheduler with steps and gamma by intuition helps me generalize the best. </p>",
      "votes": null,
      "replies": [
        {
          "id": 743102,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "02/11/2020 19:39:37",
          "content": "<p>What was your lowest validation loss? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 743142,
          "author_name": "timetraveller98",
          "author_url": "",
          "post_date": "02/11/2020 20:36:37",
          "content": "<p>I haven't reached below 0.095 yet</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 743103,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "02/11/2020 19:41:21",
      "content": "<p>I have tried AdaBound. I found that with it the model trained bit faster than vanilla Adam. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 743692,
      "author_name": "p4rallax",
      "author_url": "",
      "post_date": "02/12/2020 07:41:43",
      "content": "<p>The main thing i have been considering for this is if the optimizer state gets loaded properly since I only have Kaggle to train my kernels. Over9000 worked best, but training for 100 epochs wasn't possible with it, since I could train for only 30 epochs in a single notebook, and loading the optimizer state next time would lead to a drop of ~1% in the CV. I'm still experimenting what to use now. Any ideas would be appreciated </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 744218,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "02/12/2020 17:01:25",
      "content": "<p>So far I've only been using Adam and ReduceOnPlateau has been really helpful!\nBut then, I'm still way behind a lot of people!</p>",
      "votes": null,
      "replies": [
        {
          "id": 744235,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/12/2020 17:12:48",
          "content": "<p>It's just a matter of trying different things, some people have more experience and know what knobs to tune first and for others it is just a matter of time, I am sure that you will have a breakthrough if you keep experimenting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 744249,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/12/2020 17:20:57",
          "content": "<p>This is a good advice! You learn by making mistakes =) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 744516,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "02/12/2020 22:36:43",
      "content": "<p>I've started with Adam and gave RAdam some tries. However I didn't notice any significant difference between local CV or related LB. For now I switched back to using just Adam.</p>\n\n<p>Usually I start the model training with a learning rate of about 0.0002 - 0.00015. Lower seems to take forever to train altough I'am not sure it would lead to underfitting. Higher does give me some overfitting.</p>\n\n<p>I'am also using the ReduceOnPlateau. In some kernels a rather 'quick' reduce scheme is used (patience 3 and reduce factor of 0.5). I've experienced that a more relaxed scheme seems to give some more improvement. I'am now experimenting with a patience of 4 to 5 and factor around .7. I'am personally quitte happy with the improvements that it already gave.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 748619,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "02/17/2020 18:48:02",
      "content": "<p>Personally i've been having better results with Lookahead(SGD) and OneCycleLR by a bit margin! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "742923": "In my tests so far I only used Adam with different learning rates and ReduceOnPlateau configurations.\nDid anybody tried a different optimizer than Adam or Adam variants like RAdam ? What did you experiment ?\n\n\nCheers",
    "743020": "I am currently trying RAdam with ReduceOnPlateau. With lower lr for classifier\nBut it seems to overfit. I have a big gap between LB and Val score, which is really strange",
    "743023": "How big is the gap ? Maybe the problem is on the classifier, it's to complex for what you are trying to predict or you don't augment enough the images so you can generalize",
    "743034": "In my experiments so far ReduceLROnPlateau converges quicker than OneCycle and they perform about the same. I've also been using Adam and AdamW but there's not much difference there either.",
    "743046": "I also had better experiences with ReduceLROnPlateau.",
    "743061": "For me Adam or Radam they perform very similar.",
    "743081": "I'm curious how in @iafoss kernel mixup converged in 32 epochs. I'm using AdamW for 80 epochs and from my experience even after 100+ epochs score can improve.\n\nThe other question is how to carefully choose minimal lr? I guess 1e-5 or 5e-6 should work fine.",
    "743096": "For me, adam generalizes better than adamw (worse in val loss at least ~0.01 lb metric)\nadam and radam perform similar\nmultistep scheduler with steps and gamma by intuition helps me generalize the best.",
    "743102": "What was your lowest validation loss?",
    "743103": "I have tried AdaBound. I found that with it the model trained bit faster than vanilla Adam.",
    "743142": "I haven't reached below 0.095 yet",
    "743692": "The main thing i have been considering for this is if the optimizer state gets loaded properly since I only have Kaggle to train my kernels. Over9000 worked best, but training for 100 epochs wasn't possible with it, since I could train for only 30 epochs in a single notebook, and loading the optimizer state next time would lead to a drop of ~1% in the CV. I'm still experimenting what to use now. Any ideas would be appreciated",
    "743872": "drhabib do you use ReduceOnPlateau ?",
    "744218": "So far I've only been using Adam and ReduceOnPlateau has been really helpful!\nBut then, I'm still way behind a lot of people!",
    "744235": "It's just a matter of trying different things, some people have more experience and know what knobs to tune first and for others it is just a matter of time, I am sure that you will have a breakthrough if you keep experimenting",
    "744247": "I used `Adam` with `ReduceOnPlateau` and Adam with `OneCycle`. They are similar in my setup\n =)",
    "744249": "This is a good advice! You learn by making mistakes =)",
    "744516": "I've started with Adam and gave RAdam some tries. However I didn't notice any significant difference between local CV or related LB. For now I switched back to using just Adam.\n\nUsually I start the model training with a learning rate of about 0.0002 - 0.00015. Lower seems to take forever to train altough I'am not sure it would lead to underfitting. Higher does give me some overfitting.\n\nI'am also using the ReduceOnPlateau. In some kernels a rather 'quick' reduce scheme is used (patience 3 and reduce factor of 0.5). I've experienced that a more relaxed scheme seems to give some more improvement. I'am now experimenting with a patience of 4 to 5 and factor around .7. I'am personally quitte happy with the improvements that it already gave.",
    "746978": "Same here!",
    "748619": "Personally i've been having better results with Lookahead(SGD) and OneCycleLR by a bit margin!",
    "749128": "For me, ReduceOnPlateau converges more quickly than OneCycleCosineLR",
    "749870": "cswwp347724 A really simple question. Is the reduce LR on Plateau done when metric is no longer increasing or when val loss is no longer decreasing? I have observed cases where metric and validation loss are both increasing :)",
    "754324": "roguekk007  Sorry for late reply, i just set LR down when metric is no longer increase at the last patience epoch. It's possible that  metric and validation loss are both increasing, because root has more weight in final metric",
    "754333": "cswwp347724 Thanks for the reply! I think setting LR based on metric is a sound choice taken that the CV correlates so well with the LB",
    "755148": "roguekk007 Do you want team up?😁",
    "755163": "cswwp347724 Thank you very much for the offer! This is the first comp that looks relatively promising for me; so I’m really excited and it’s such a great opportunity for me to learn. I want to really push my single-model score and might not consider teaming up until before the merger deadline; best of luck till then👋",
    "755173": "roguekk007 Ok, good luck👍",
    "756461": "I change small batchsize with small LR, it seems not too much improvement😣 @roguekk007",
    "756708": "cswwp347724 Same for me... Stuck again. It seems like batch size and LR are not so-important details in this comp"
  },
  "source": "meta"
}