{
  "id": 212070,
  "title": "Optimizer analyses",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212070",
  "author_name": "Kirderf",
  "post_date": "2021-01-17T11:48:40.778000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Have done some analyses between common optimizers, Adam, AdamW, AdamP, Lamb, SGDP, SGD, QHAdamW, Radam, Ralamb.<br>\nDid the analyses with OneCycle+Warmup scheduler.</p>\n<p>Here my summary of the top 3 and how to use them so far in this problem:</p>\n<p>Radam – convergence fastest with best short-term score, good for e.g. tuning.<br>\nRalamb – when in convergence area stays there longer with many more checkpoints within top short-term score, good for e.g. SWA.<br>\nSGD – best long term/final score but with much longer training, ~x3, good for e.g. final training.</p>\n<p>An ensemble of these three trainings can also be a strategy.</p>",
  "messages": [
    {
      "id": 1156761,
      "postDate": "2021-01-17T11:48:40.777Z",
      "content": "<p>Have done some analyses between common optimizers, Adam, AdamW, AdamP, Lamb, SGDP, SGD, QHAdamW, Radam, Ralamb.<br>\nDid the analyses with OneCycle+Warmup scheduler.</p>\n<p>Here my summary of the top 3 and how to use them so far in this problem:</p>\n<p>Radam – convergence fastest with best short-term score, good for e.g. tuning.<br>\nRalamb – when in convergence area stays there longer with many more checkpoints within top short-term score, good for e.g. SWA.<br>\nSGD – best long term/final score but with much longer training, ~x3, good for e.g. final training.</p>\n<p>An ensemble of these three trainings can also be a strategy.</p>",
      "rawMarkdown": "Have done some analyses between common optimizers, Adam, AdamW, AdamP, Lamb, SGDP, SGD, QHAdamW, Radam, Ralamb.\nDid the analyses with OneCycle+Warmup scheduler.\n\nHere my summary of the top 3 and how to use them so far in this problem:\n\nRadam – convergence fastest with best short-term score, good for e.g. tuning.\nRalamb – when in convergence area stays there longer with many more checkpoints within top short-term score, good for e.g. SWA.\nSGD – best long term/final score but with much longer training, ~x3, good for e.g. final training.\n\nAn ensemble of these three trainings can also be a strategy.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1156761": "Have done some analyses between common optimizers, Adam, AdamW, AdamP, Lamb, SGDP, SGD, QHAdamW, Radam, Ralamb.\nDid the analyses with OneCycle+Warmup scheduler.\n\nHere my summary of the top 3 and how to use them so far in this problem:\n\nRadam – convergence fastest with best short-term score, good for e.g. tuning.\nRalamb – when in convergence area stays there longer with many more checkpoints within top short-term score, good for e.g. SWA.\nSGD – best long term/final score but with much longer training, ~x3, good for e.g. final training.\n\nAn ensemble of these three trainings can also be a strategy."
  }
}