{
  "id": 568823,
  "title": "Alphafold3 alternatives",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568823",
  "author_name": "",
  "post_date": "2025-03-18T06:48:44.858686700Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p>Given AlphaFold3's leading position in the field, I believe we should take full advantage of its model architecture and inference results in this competition. While it is not specifically designed for RNA, multiple studies have consistently shown that its RNA structure prediction performance remains among the best, possibly because different types of data provide general structural constraints.</p>\n<p>The main challenge, however, is that directly using AlphaFold3 comes with licensing issues, and it does not provide training scripts. Fortunately, several organizations have already replicated and even improved upon AlphaFold3, making their versions open-source. I’ve noticed that some of these models have already been mentioned here—for instance, hengck23 brought up <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a>.</p>\n<p>I recently came across a study evaluating various models for predicting protein–peptide complex structures (<a href=\"https://www.biorxiv.org/content/10.1101/2025.03.09.642277v2.full.pdf+html\" target=\"_blank\">link</a>). It introduced several models I hadn’t heard of before, all of which are also AlphaFold3 derivatives. Notably, <strong><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a></strong> and <strong><a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">Boltz-1</a></strong> provide training scripts, while the MSA-free version of <strong><a href=\"https://github.com/chaidiscovery/chai-lab\" target=\"_blank\">Chai-1</a></strong> has also shown promising results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2594502%2F20c07c11be9c90cc4bcbaf6ef1d3a1af%2FScreenshot%202025-03-18%20at%2014.31.55.png?generation=1742279564847649&amp;alt=media\" alt=\"\"></p>\n<p>I believe these resources could be highly valuable. If you have sufficient computing resources, it might be worthwhile to experiment with training/finetuning/distilling them.</p>",
  "messages": [
    {
      "id": "3152786",
      "postDate": "03/18/2025 06:48:44",
      "content": "<p>Hello everyone!</p>\n<p>Given AlphaFold3's leading position in the field, I believe we should take full advantage of its model architecture and inference results in this competition. While it is not specifically designed for RNA, multiple studies have consistently shown that its RNA structure prediction performance remains among the best, possibly because different types of data provide general structural constraints.</p>\n<p>The main challenge, however, is that directly using AlphaFold3 comes with licensing issues, and it does not provide training scripts. Fortunately, several organizations have already replicated and even improved upon AlphaFold3, making their versions open-source. I’ve noticed that some of these models have already been mentioned here—for instance, hengck23 brought up <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a>.</p>\n<p>I recently came across a study evaluating various models for predicting protein–peptide complex structures (<a href=\"https://www.biorxiv.org/content/10.1101/2025.03.09.642277v2.full.pdf+html\" target=\"_blank\">link</a>). It introduced several models I hadn’t heard of before, all of which are also AlphaFold3 derivatives. Notably, <strong><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a></strong> and <strong><a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">Boltz-1</a></strong> provide training scripts, while the MSA-free version of <strong><a href=\"https://github.com/chaidiscovery/chai-lab\" target=\"_blank\">Chai-1</a></strong> has also shown promising results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2594502%2F20c07c11be9c90cc4bcbaf6ef1d3a1af%2FScreenshot%202025-03-18%20at%2014.31.55.png?generation=1742279564847649&amp;alt=media\" alt=\"\"></p>\n<p>I believe these resources could be highly valuable. If you have sufficient computing resources, it might be worthwhile to experiment with training/finetuning/distilling them.</p>",
      "rawMarkdown": "Hello everyone!\n\nGiven AlphaFold3's leading position in the field, I believe we should take full advantage of its model architecture and inference results in this competition. While it is not specifically designed for RNA, multiple studies have consistently shown that its RNA structure prediction performance remains among the best, possibly because different types of data provide general structural constraints.\n\nThe main challenge, however, is that directly using AlphaFold3 comes with licensing issues, and it does not provide training scripts. Fortunately, several organizations have already replicated and even improved upon AlphaFold3, making their versions open-source. I’ve noticed that some of these models have already been mentioned here—for instance, hengck23 brought up [Protenix](https://github.com/bytedance/Protenix).\n\nI recently came across a study evaluating various models for predicting protein–peptide complex structures ([link](https://www.biorxiv.org/content/10.1101/2025.03.09.642277v2.full.pdf+html)). It introduced several models I hadn’t heard of before, all of which are also AlphaFold3 derivatives. Notably, **[Protenix](https://github.com/bytedance/Protenix)** and **[Boltz-1](https://github.com/jwohlwend/boltz)** provide training scripts, while the MSA-free version of **[Chai-1](https://github.com/chaidiscovery/chai-lab)** has also shown promising results.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2594502%2F20c07c11be9c90cc4bcbaf6ef1d3a1af%2FScreenshot%202025-03-18%20at%2014.31.55.png?generation=1742279564847649&alt=media)\n\nI believe these resources could be highly valuable. If you have sufficient computing resources, it might be worthwhile to experiment with training/finetuning/distilling them.",
      "votes": null
    },
    {
      "id": "3152813",
      "postDate": "03/18/2025 07:31:33",
      "content": "<p>some of this uses MMseqs2 , one may want to consider GPU version</p>\n<p><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.13.623350v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.13.623350v1</a><br>\n<a href=\"https://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/\" target=\"_blank\">https://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/</a></p>\n<p>but rna uses rMSA no sure if that has a GPU version?</p>",
      "rawMarkdown": "some of this uses MMseqs2 , one may want to consider GPU version\n\nhttps://www.biorxiv.org/content/10.1101/2024.11.13.623350v1\nhttps://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/\n\nbut rna uses rMSA no sure if that has a GPU version?",
      "votes": null
    },
    {
      "id": "3152868",
      "postDate": "03/18/2025 08:18:14",
      "content": "<p>Thanks for the information! <br>\nI just checked the Alphafold3 paper, they used mmseq2 to cluster RNA sequences and nhmmer to search MSAs against cluster representative RNAs.  Maybe we can boldly try to use mmseq2 for both clustering and searching😃</p>",
      "rawMarkdown": "Thanks for the information! \nI just checked the Alphafold3 paper, they used mmseq2 to cluster RNA sequences and nhmmer to search MSAs against cluster representative RNAs.  Maybe we can boldly try to use mmseq2 for both clustering and searching😃",
      "votes": null
    },
    {
      "id": "3178217",
      "postDate": "04/13/2025 21:32:45",
      "content": "<p>It's particularly worth noting the following in AlphaFold3's terms and conditions regarding the pretrained weights:</p>\n<blockquote>\n  <p>You must not use nor allow others to use: … AlphaFold 3 output to train machine learning models or related technology for biomolecular structure prediction similar to AlphaFold 3. </p>\n</blockquote>\n<p>So you can't use AF3 to generate targets for finetuning alternative models, unless you build and train the model yourself. </p>",
      "rawMarkdown": "It's particularly worth noting the following in AlphaFold3's terms and conditions regarding the pretrained weights:\n\n>You must not use nor allow others to use: ... AlphaFold 3 output to train machine learning models or related technology for biomolecular structure prediction similar to AlphaFold 3. \n\nSo you can't use AF3 to generate targets for finetuning alternative models, unless you build and train the model yourself.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3152813,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/18/2025 07:31:33",
      "content": "<p>some of this uses MMseqs2 , one may want to consider GPU version</p>\n<p><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.13.623350v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.13.623350v1</a><br>\n<a href=\"https://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/\" target=\"_blank\">https://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/</a></p>\n<p>but rna uses rMSA no sure if that has a GPU version?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3152868,
          "author_name": "achenge07",
          "author_url": "",
          "post_date": "03/18/2025 08:18:14",
          "content": "<p>Thanks for the information! <br>\nI just checked the Alphafold3 paper, they used mmseq2 to cluster RNA sequences and nhmmer to search MSAs against cluster representative RNAs.  Maybe we can boldly try to use mmseq2 for both clustering and searching😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3178217,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "04/13/2025 21:32:45",
      "content": "<p>It's particularly worth noting the following in AlphaFold3's terms and conditions regarding the pretrained weights:</p>\n<blockquote>\n  <p>You must not use nor allow others to use: … AlphaFold 3 output to train machine learning models or related technology for biomolecular structure prediction similar to AlphaFold 3. </p>\n</blockquote>\n<p>So you can't use AF3 to generate targets for finetuning alternative models, unless you build and train the model yourself. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3152786": "Hello everyone!\n\nGiven AlphaFold3's leading position in the field, I believe we should take full advantage of its model architecture and inference results in this competition. While it is not specifically designed for RNA, multiple studies have consistently shown that its RNA structure prediction performance remains among the best, possibly because different types of data provide general structural constraints.\n\nThe main challenge, however, is that directly using AlphaFold3 comes with licensing issues, and it does not provide training scripts. Fortunately, several organizations have already replicated and even improved upon AlphaFold3, making their versions open-source. I’ve noticed that some of these models have already been mentioned here—for instance, hengck23 brought up [Protenix](https://github.com/bytedance/Protenix).\n\nI recently came across a study evaluating various models for predicting protein–peptide complex structures ([link](https://www.biorxiv.org/content/10.1101/2025.03.09.642277v2.full.pdf+html)). It introduced several models I hadn’t heard of before, all of which are also AlphaFold3 derivatives. Notably, **[Protenix](https://github.com/bytedance/Protenix)** and **[Boltz-1](https://github.com/jwohlwend/boltz)** provide training scripts, while the MSA-free version of **[Chai-1](https://github.com/chaidiscovery/chai-lab)** has also shown promising results.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2594502%2F20c07c11be9c90cc4bcbaf6ef1d3a1af%2FScreenshot%202025-03-18%20at%2014.31.55.png?generation=1742279564847649&alt=media)\n\nI believe these resources could be highly valuable. If you have sufficient computing resources, it might be worthwhile to experiment with training/finetuning/distilling them.",
    "3152813": "some of this uses MMseqs2 , one may want to consider GPU version\n\nhttps://www.biorxiv.org/content/10.1101/2024.11.13.623350v1\nhttps://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/\n\nbut rna uses rMSA no sure if that has a GPU version?",
    "3152868": "Thanks for the information! \nI just checked the Alphafold3 paper, they used mmseq2 to cluster RNA sequences and nhmmer to search MSAs against cluster representative RNAs.  Maybe we can boldly try to use mmseq2 for both clustering and searching😃",
    "3178217": "It's particularly worth noting the following in AlphaFold3's terms and conditions regarding the pretrained weights:\n\n>You must not use nor allow others to use: ... AlphaFold 3 output to train machine learning models or related technology for biomolecular structure prediction similar to AlphaFold 3. \n\nSo you can't use AF3 to generate targets for finetuning alternative models, unless you build and train the model yourself."
  },
  "source": "meta"
}