{
  "id": 568298,
  "title": "About 'OOM' Issue",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568298",
  "author_name": "",
  "post_date": "2025-03-15T02:19:00.830582900Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Does anyone have a solution for the issue where Rhofold, NuFold, or other trained models take too long for inference and run into OOM errors? To avoid this, I'm currently limiting the RNA sequence length to 500, but this negatively impacts performance. Any suggestions?</p>",
  "messages": [
    {
      "id": "3150044",
      "postDate": "03/15/2025 02:19:00",
      "content": "<p>Does anyone have a solution for the issue where Rhofold, NuFold, or other trained models take too long for inference and run into OOM errors? To avoid this, I'm currently limiting the RNA sequence length to 500, but this negatively impacts performance. Any suggestions?</p>",
      "rawMarkdown": "Does anyone have a solution for the issue where Rhofold, NuFold, or other trained models take too long for inference and run into OOM errors? To avoid this, I'm currently limiting the RNA sequence length to 500, but this negatively impacts performance. Any suggestions?",
      "votes": null
    },
    {
      "id": "3150052",
      "postDate": "03/15/2025 02:42:57",
      "content": "<p>google for paper that reduce memory for alphafold2 or evoformer, etc<br>\nyou may need to recode and retrain/distill the pretrained model.</p>\n<hr>\n<p>alternatively use cpu offloading and the tricks for large LLM<br>\ne.g. deepspeed, gradient checkpoint(for droupout?)   … also not sure if quntisation would work, flash attention???</p>\n<p>you can use 2xT4 (2x12 gb ) and tensor pipline</p>",
      "rawMarkdown": "google for paper that reduce memory for alphafold2 or evoformer, etc\nyou may need to recode and retrain/distill the pretrained model.\n\n---\n\nalternatively use cpu offloading and the tricks for large LLM\ne.g. deepspeed, gradient checkpoint(for droupout?)   ... also not sure if quntisation would work, flash attention???\n\nyou can use 2xT4 (2x12 gb ) and tensor pipline",
      "votes": null
    },
    {
      "id": "3150054",
      "postDate": "03/15/2025 02:46:25",
      "content": "<p><a href=\"https://github.com/Dao-AILab/flash-attention/blob/main/usage.md\" target=\"_blank\">https://github.com/Dao-AILab/flash-attention/blob/main/usage.md</a></p>\n<p>OpenFold: a trainable, memory-efficient, and GPU-friendly PyTorch reproduction of AlphaFold 2. With FlashAttention as one of its components, it is up to 3x faster than AlphaFold2 to run inference on short sequences, and can predict 2x longer structures.</p>\n<ul>\n<li>liteformer: <a href=\"https://openreview.net/forum?id=t0m0DdCCQ2\" target=\"_blank\">https://openreview.net/forum?id=t0m0DdCCQ2</a><br>\nTo tackle this obstacle, we present Liteformer as an alternative to the original Evoformer<br>\nin AlphaFold2 to reduce memory and computational complexity from O(L^3 +sL^2) to O(L^2 +sL),<br>\nas shown in Figure 2.</li>\n</ul>\n<p>if you want to rewrite evoformer block, i suggest you might implement this as well<br>\n<a href=\"https://arxiv.org/abs/2503.10622\" target=\"_blank\">https://arxiv.org/abs/2503.10622</a><br>\nTransformers without Normalization</p>",
      "rawMarkdown": "https://github.com/Dao-AILab/flash-attention/blob/main/usage.md\n\nOpenFold: a trainable, memory-efficient, and GPU-friendly PyTorch reproduction of AlphaFold 2. With FlashAttention as one of its components, it is up to 3x faster than AlphaFold2 to run inference on short sequences, and can predict 2x longer structures.\n\n- liteformer: https://openreview.net/forum?id=t0m0DdCCQ2\nTo tackle this obstacle, we present Liteformer as an alternative to the original Evoformer\nin AlphaFold2 to reduce memory and computational complexity from O(L^3 +sL^2) to O(L^2 +sL),\nas shown in Figure 2.\n\nif you want to rewrite evoformer block, i suggest you might implement this as well\nhttps://arxiv.org/abs/2503.10622\nTransformers without Normalization",
      "votes": null
    },
    {
      "id": "3150326",
      "postDate": "03/15/2025 11:05:47",
      "content": "<p>Another solution is to load layer by layer to gpu and evaluate. </p>\n<p>Note that open source model needs also to trained for long seq too</p>",
      "rawMarkdown": "Another solution is to load layer by layer to gpu and evaluate. \n\nNote that open source model needs also to trained for long seq too",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3150052,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/15/2025 02:42:57",
      "content": "<p>google for paper that reduce memory for alphafold2 or evoformer, etc<br>\nyou may need to recode and retrain/distill the pretrained model.</p>\n<hr>\n<p>alternatively use cpu offloading and the tricks for large LLM<br>\ne.g. deepspeed, gradient checkpoint(for droupout?)   … also not sure if quntisation would work, flash attention???</p>\n<p>you can use 2xT4 (2x12 gb ) and tensor pipline</p>",
      "votes": null,
      "replies": [
        {
          "id": 3150054,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/15/2025 02:46:25",
          "content": "<p><a href=\"https://github.com/Dao-AILab/flash-attention/blob/main/usage.md\" target=\"_blank\">https://github.com/Dao-AILab/flash-attention/blob/main/usage.md</a></p>\n<p>OpenFold: a trainable, memory-efficient, and GPU-friendly PyTorch reproduction of AlphaFold 2. With FlashAttention as one of its components, it is up to 3x faster than AlphaFold2 to run inference on short sequences, and can predict 2x longer structures.</p>\n<ul>\n<li>liteformer: <a href=\"https://openreview.net/forum?id=t0m0DdCCQ2\" target=\"_blank\">https://openreview.net/forum?id=t0m0DdCCQ2</a><br>\nTo tackle this obstacle, we present Liteformer as an alternative to the original Evoformer<br>\nin AlphaFold2 to reduce memory and computational complexity from O(L^3 +sL^2) to O(L^2 +sL),<br>\nas shown in Figure 2.</li>\n</ul>\n<p>if you want to rewrite evoformer block, i suggest you might implement this as well<br>\n<a href=\"https://arxiv.org/abs/2503.10622\" target=\"_blank\">https://arxiv.org/abs/2503.10622</a><br>\nTransformers without Normalization</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3150326,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/15/2025 11:05:47",
      "content": "<p>Another solution is to load layer by layer to gpu and evaluate. </p>\n<p>Note that open source model needs also to trained for long seq too</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3150044": "Does anyone have a solution for the issue where Rhofold, NuFold, or other trained models take too long for inference and run into OOM errors? To avoid this, I'm currently limiting the RNA sequence length to 500, but this negatively impacts performance. Any suggestions?",
    "3150052": "google for paper that reduce memory for alphafold2 or evoformer, etc\nyou may need to recode and retrain/distill the pretrained model.\n\n---\n\nalternatively use cpu offloading and the tricks for large LLM\ne.g. deepspeed, gradient checkpoint(for droupout?)   ... also not sure if quntisation would work, flash attention???\n\nyou can use 2xT4 (2x12 gb ) and tensor pipline",
    "3150054": "https://github.com/Dao-AILab/flash-attention/blob/main/usage.md\n\nOpenFold: a trainable, memory-efficient, and GPU-friendly PyTorch reproduction of AlphaFold 2. With FlashAttention as one of its components, it is up to 3x faster than AlphaFold2 to run inference on short sequences, and can predict 2x longer structures.\n\n- liteformer: https://openreview.net/forum?id=t0m0DdCCQ2\nTo tackle this obstacle, we present Liteformer as an alternative to the original Evoformer\nin AlphaFold2 to reduce memory and computational complexity from O(L^3 +sL^2) to O(L^2 +sL),\nas shown in Figure 2.\n\nif you want to rewrite evoformer block, i suggest you might implement this as well\nhttps://arxiv.org/abs/2503.10622\nTransformers without Normalization",
    "3150326": "Another solution is to load layer by layer to gpu and evaluate. \n\nNote that open source model needs also to trained for long seq too"
  },
  "source": "meta"
}