{
  "id": 442993,
  "title": "Transformer training diverges",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/442993",
  "author_name": "",
  "post_date": "2023-09-25T05:38:46.900625Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi! I'm trying to train a transformer, but the training diverges<br>\nI'm trying to lower lr/increase batch-size, what are the others things I can do? <br>\nCan you point me to some blogs, such as <a href=\"https://liyuanlucasliu.github.io/files/slides-transformer-clinic.pdf\" target=\"_blank\">this one</a></p>",
  "messages": [
    {
      "id": "2454799",
      "postDate": "09/25/2023 05:38:46",
      "content": "<p>Hi! I'm trying to train a transformer, but the training diverges<br>\nI'm trying to lower lr/increase batch-size, what are the others things I can do? <br>\nCan you point me to some blogs, such as <a href=\"https://liyuanlucasliu.github.io/files/slides-transformer-clinic.pdf\" target=\"_blank\">this one</a></p>",
      "rawMarkdown": "Hi! I'm trying to train a transformer, but the training diverges\nI'm trying to lower lr/increase batch-size, what are the others things I can do? \nCan you point me to some blogs, such as [this one](https://liyuanlucasliu.github.io/files/slides-transformer-clinic.pdf)",
      "votes": null
    },
    {
      "id": "2457433",
      "postDate": "09/26/2023 22:11:16",
      "content": "<p>Training is diverging because you're overfitting I would imagine.</p>",
      "rawMarkdown": "Training is diverging because you're overfitting I would imagine.",
      "votes": null
    },
    {
      "id": "2461825",
      "postDate": "09/29/2023 19:34:22",
      "content": "<p>I can't say I know of any blogs to point you in the direction of, however, some things that may help are weight decay (increasing it) and/or l1/l2 normalization. </p>",
      "rawMarkdown": "I can't say I know of any blogs to point you in the direction of, however, some things that may help are weight decay (increasing it) and/or l1/l2 normalization.",
      "votes": null
    },
    {
      "id": "2467032",
      "postDate": "10/04/2023 08:59:22",
      "content": "<p>May be that of some use:<br>\n\"Google DeepMind Researchers Uncover Scalable Solutions to Combat Training Instabilities in Transformer Models: An In-depth Analysis on Smaller Scale Reproducibility and Optimization Strategies\":<br>\n<a href=\"https://www.marktechpost.com/2023/10/02/google-deepmind-researchers-uncover-scalable-solutions-to-combat-training-instabilities-in-transformer-models-an-in-depth-analysis-on-smaller-scale-reproducibility-and-optimization-strategies/\" target=\"_blank\">https://www.marktechpost.com/2023/10/02/google-deepmind-researchers-uncover-scalable-solutions-to-combat-training-instabilities-in-transformer-models-an-in-depth-analysis-on-smaller-scale-reproducibility-and-optimization-strategies/</a></p>",
      "rawMarkdown": "May be that of some use:\n\"Google DeepMind Researchers Uncover Scalable Solutions to Combat Training Instabilities in Transformer Models: An In-depth Analysis on Smaller Scale Reproducibility and Optimization Strategies\":\nhttps://www.marktechpost.com/2023/10/02/google-deepmind-researchers-uncover-scalable-solutions-to-combat-training-instabilities-in-transformer-models-an-in-depth-analysis-on-smaller-scale-reproducibility-and-optimization-strategies/",
      "votes": null
    },
    {
      "id": "2467309",
      "postDate": "10/04/2023 12:44:20",
      "content": "<p>thankyou for the link</p>",
      "rawMarkdown": "thankyou for the link",
      "votes": null
    },
    {
      "id": "2478263",
      "postDate": "10/11/2023 21:58:28",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2457433,
      "author_name": "themecheng",
      "author_url": "",
      "post_date": "09/26/2023 22:11:16",
      "content": "<p>Training is diverging because you're overfitting I would imagine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2461825,
      "author_name": "starblasters8",
      "author_url": "",
      "post_date": "09/29/2023 19:34:22",
      "content": "<p>I can't say I know of any blogs to point you in the direction of, however, some things that may help are weight decay (increasing it) and/or l1/l2 normalization. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2467032,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/04/2023 08:59:22",
      "content": "<p>May be that of some use:<br>\n\"Google DeepMind Researchers Uncover Scalable Solutions to Combat Training Instabilities in Transformer Models: An In-depth Analysis on Smaller Scale Reproducibility and Optimization Strategies\":<br>\n<a href=\"https://www.marktechpost.com/2023/10/02/google-deepmind-researchers-uncover-scalable-solutions-to-combat-training-instabilities-in-transformer-models-an-in-depth-analysis-on-smaller-scale-reproducibility-and-optimization-strategies/\" target=\"_blank\">https://www.marktechpost.com/2023/10/02/google-deepmind-researchers-uncover-scalable-solutions-to-combat-training-instabilities-in-transformer-models-an-in-depth-analysis-on-smaller-scale-reproducibility-and-optimization-strategies/</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2467309,
          "author_name": "simonnderitu",
          "author_url": "",
          "post_date": "10/04/2023 12:44:20",
          "content": "<p>thankyou for the link</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2478263,
          "author_name": "bagasshalahuddinw",
          "author_url": "",
          "post_date": "10/11/2023 21:58:28",
          "content": "<p>Thank you for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2454799": "Hi! I'm trying to train a transformer, but the training diverges\nI'm trying to lower lr/increase batch-size, what are the others things I can do? \nCan you point me to some blogs, such as [this one](https://liyuanlucasliu.github.io/files/slides-transformer-clinic.pdf)",
    "2457433": "Training is diverging because you're overfitting I would imagine.",
    "2461825": "I can't say I know of any blogs to point you in the direction of, however, some things that may help are weight decay (increasing it) and/or l1/l2 normalization.",
    "2467032": "May be that of some use:\n\"Google DeepMind Researchers Uncover Scalable Solutions to Combat Training Instabilities in Transformer Models: An In-depth Analysis on Smaller Scale Reproducibility and Optimization Strategies\":\nhttps://www.marktechpost.com/2023/10/02/google-deepmind-researchers-uncover-scalable-solutions-to-combat-training-instabilities-in-transformer-models-an-in-depth-analysis-on-smaller-scale-reproducibility-and-optimization-strategies/",
    "2467309": "thankyou for the link",
    "2478263": "Thank you for sharing"
  },
  "source": "meta"
}