{
  "id": 568066,
  "title": "a new paper : Drfold2 :  NO MSA required",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568066",
  "author_name": "",
  "post_date": "2025-03-13T16:09:01.660101600Z",
  "votes": 28,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://github.com/leeyang/DRfold2\" target=\"_blank\">https://github.com/leeyang/DRfold2</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F372fef9462664504099729b0f03e164a%2FSelection_999(7767).png?generation=1741882290600806&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>DRfold2: Ab initio RNA structure prediction with composite language model and denoised end-to-end learning - bioxiv2025- mar<br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2025.03.05.641632v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.03.05.641632v1.full.pdf</a><br>\n(input = seq only)</p>\n<p>QUOTE \"Detailed analyses reveal that the improvements primarily stem from the RCLM’s ability to capture co-evolutionary pattern and the effective denoising process, leading to a more than 100% increase in contact prediction precision compared to existing methods\"</p>\n<ul>\n<li>inlcude experiment results on casp15/16, RNA-Puzzles targets</li>\n<li>comparable to Alphafold3</li>\n<li>denoising to predict multiple conformation</li>\n</ul>\n<p>NOTE: one weakness<br>\nQUOTE \"Notably, large synthetic RNAs from CASP15 were excluded due to their deviation from naturally occurring RNA\"<br>\n(maybe that is not learned in RCLM’s?)</p>\n<hr>\n<p>related:<br>\n[1] DRfold : Integrating end-to-end learning with deep geometrical potentials for ab initio RNA structure prediction - 2023-Nov<br>\n<a href=\"https://www.nature.com/articles/s41467-023-41303-9\" target=\"_blank\">https://www.nature.com/articles/s41467-023-41303-9</a><br>\n(input = seq + SS (no MSA))</p>\n<p>[2] Systematic benchmarking of deep-learning methods for tertiary RNA structure prediction - 2024-nov<br>\n<a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715\" target=\"_blank\">https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715</a><br>\n(benchmark DRfold on CASP15, RNA-Puzzles)</p>",
  "messages": [
    {
      "id": "3148898",
      "postDate": "03/13/2025 16:09:01",
      "content": "<p><a href=\"https://github.com/leeyang/DRfold2\" target=\"_blank\">https://github.com/leeyang/DRfold2</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F372fef9462664504099729b0f03e164a%2FSelection_999(7767).png?generation=1741882290600806&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>DRfold2: Ab initio RNA structure prediction with composite language model and denoised end-to-end learning - bioxiv2025- mar<br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2025.03.05.641632v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.03.05.641632v1.full.pdf</a><br>\n(input = seq only)</p>\n<p>QUOTE \"Detailed analyses reveal that the improvements primarily stem from the RCLM’s ability to capture co-evolutionary pattern and the effective denoising process, leading to a more than 100% increase in contact prediction precision compared to existing methods\"</p>\n<ul>\n<li>inlcude experiment results on casp15/16, RNA-Puzzles targets</li>\n<li>comparable to Alphafold3</li>\n<li>denoising to predict multiple conformation</li>\n</ul>\n<p>NOTE: one weakness<br>\nQUOTE \"Notably, large synthetic RNAs from CASP15 were excluded due to their deviation from naturally occurring RNA\"<br>\n(maybe that is not learned in RCLM’s?)</p>\n<hr>\n<p>related:<br>\n[1] DRfold : Integrating end-to-end learning with deep geometrical potentials for ab initio RNA structure prediction - 2023-Nov<br>\n<a href=\"https://www.nature.com/articles/s41467-023-41303-9\" target=\"_blank\">https://www.nature.com/articles/s41467-023-41303-9</a><br>\n(input = seq + SS (no MSA))</p>\n<p>[2] Systematic benchmarking of deep-learning methods for tertiary RNA structure prediction - 2024-nov<br>\n<a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715\" target=\"_blank\">https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715</a><br>\n(benchmark DRfold on CASP15, RNA-Puzzles)</p>",
      "rawMarkdown": "https://github.com/leeyang/DRfold2\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F372fef9462664504099729b0f03e164a%2FSelection_999(7767).png?generation=1741882290600806&alt=media)\n\n---\n\nDRfold2: Ab initio RNA structure prediction with composite language model and denoised end-to-end learning - bioxiv2025- mar\nhttps://www.biorxiv.org/content/10.1101/2025.03.05.641632v1.full.pdf\n(input = seq only)\n\nQUOTE \"Detailed analyses reveal that the improvements primarily stem from the RCLM’s ability to capture co-evolutionary pattern and the effective denoising process, leading to a more than 100% increase in contact prediction precision compared to existing methods\"\n\n- inlcude experiment results on casp15/16, RNA-Puzzles targets\n- comparable to Alphafold3\n- denoising to predict multiple conformation\n\nNOTE: one weakness\nQUOTE \"Notably, large synthetic RNAs from CASP15 were excluded due to their deviation from naturally occurring RNA\"\n(maybe that is not learned in RCLM’s?)\n\n---\n\nrelated:\n[1] DRfold : Integrating end-to-end learning with deep geometrical potentials for ab initio RNA structure prediction - 2023-Nov\nhttps://www.nature.com/articles/s41467-023-41303-9\n(input = seq + SS (no MSA))\n\n\n[2] Systematic benchmarking of deep-learning methods for tertiary RNA structure prediction - 2024-nov\nhttps://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715\n(benchmark DRfold on CASP15, RNA-Puzzles)",
      "votes": null
    },
    {
      "id": "3148922",
      "postDate": "03/13/2025 16:25:13",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F95afd734bbe8590680a4f4166a169885%2FSelection_999(7768).png?generation=1741882960511664&amp;alt=media\" alt=\"\"></p>\n<p>The RCLM was trained on ~30M RNA sequences from the RNAcentral database (Release 22) over 67,000 batches, with a batch size of 128. The entire training process took about 15 days using a single NVIDIA A40 GPU.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F95afd734bbe8590680a4f4166a169885%2FSelection_999(7768).png?generation=1741882960511664&alt=media)\n\nThe RCLM was trained on ~30M RNA sequences from the RNAcentral database (Release 22) over 67,000 batches, with a batch size of 128. The entire training process took about 15 days using a single NVIDIA A40 GPU.",
      "votes": null
    },
    {
      "id": "3151873",
      "postDate": "03/17/2025 07:19:14",
      "content": "<p>local cv results (kaggle-casp15)</p>\n<ul>\n<li>no MAS</li>\n<li>modified from paper to use only one model (no ensemble, no refinement, no denosing)</li>\n<li>only ouput one (no conformation)</li>\n<li>takes 15 min for 12 seq</li>\n</ul>\n<p>(no limit max length, local PC)<br>\n[0.54611, 0.57469, 0.49749, 0.24299, 0.2387, 0.96262, 0.24303, 0.1964, 0.61831, 0.45845, 0.21084, 0.24965]<br>\nmean: 0.41994</p>\n<p>example notebook:<br>\n(limited to max length=400)<br>\nlink: <a href=\"https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa</a></p>\n<pre><code>ALL\n [., , , , ., , , , , , , ]\nMEAN \n</code></pre>",
      "rawMarkdown": "local cv results (kaggle-casp15)\n- no MAS\n- modified from paper to use only one model (no ensemble, no refinement, no denosing)\n- only ouput one (no conformation)\n- takes 15 min for 12 seq\n\n(no limit max length, local PC)\n[0.54611, 0.57469, 0.49749, 0.24299, 0.2387, 0.96262, 0.24303, 0.1964, 0.61831, 0.45845, 0.21084, 0.24965]\nmean: 0.41994\n\nexample notebook:\n(limited to max length=400)\nlink: https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa\n```\nALL\n [0.5461, 0.57469, 0.49749, 0.24299, 0.2387, 0.96262, 0.24303, 0.38226, 0.61831, 0.45845, 0.21084, 0.24965]\nMEAN 0.4354275\n```",
      "votes": null
    },
    {
      "id": "3163907",
      "postDate": "03/31/2025 09:40:43",
      "content": "<p>According to their supplemental material, it scored slightly better than AF3 for CASP16 monomers (20 targets). Looks promising.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff290f375ace4cdc9d3ece2bce29631f4%2FScreenshot%202025-03-31%20at%2018.25.44.png?generation=1743413988586715&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "According to their supplemental material, it scored slightly better than AF3 for CASP16 monomers (20 targets). Looks promising.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff290f375ace4cdc9d3ece2bce29631f4%2FScreenshot%202025-03-31%20at%2018.25.44.png?generation=1743413988586715&alt=media)",
      "votes": null
    },
    {
      "id": "3181750",
      "postDate": "04/18/2025 09:31:15",
      "content": "<p>Does it help if we apply the other second structure in the code instead of setting it None?</p>",
      "rawMarkdown": "Does it help if we apply the other second structure in the code instead of setting it None?",
      "votes": null
    },
    {
      "id": "3182960",
      "postDate": "04/20/2025 07:07:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , I have a question about using drfold2 in this competition. DRfold2's models are quite light so it spends less time to predict individual models, but actually we have to compare 80 models in repository so it takes lots of time to predict on all 80 models. How did you overcome this? Just removing some models or select top 20 best performing models among 80 by little experiments? What could be the solution. Thank you.</p>",
      "rawMarkdown": "Hi @hengck23 , I have a question about using drfold2 in this competition. DRfold2's models are quite light so it spends less time to predict individual models, but actually we have to compare 80 models in repository so it takes lots of time to predict on all 80 models. How did you overcome this? Just removing some models or select top 20 best performing models among 80 by little experiments? What could be the solution. Thank you.",
      "votes": null
    },
    {
      "id": "3185100",
      "postDate": "04/22/2025 20:20:25",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I got the same results as you with DRfold2, however I saw a major fall in performance in the train set compared to the validation/test set. For instance, the mean score on the train set was around 0.23, could it be that it was trained with similar data to CASP15/CASP16 ? I wander how it will perform after the data update</p>",
      "rawMarkdown": "hengck23 I got the same results as you with DRfold2, however I saw a major fall in performance in the train set compared to the validation/test set. For instance, the mean score on the train set was around 0.23, could it be that it was trained with similar data to CASP15/CASP16 ? I wander how it will perform after the data update",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3148922,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/13/2025 16:25:13",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F95afd734bbe8590680a4f4166a169885%2FSelection_999(7768).png?generation=1741882960511664&amp;alt=media\" alt=\"\"></p>\n<p>The RCLM was trained on ~30M RNA sequences from the RNAcentral database (Release 22) over 67,000 batches, with a batch size of 128. The entire training process took about 15 days using a single NVIDIA A40 GPU.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3151873,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/17/2025 07:19:14",
      "content": "<p>local cv results (kaggle-casp15)</p>\n<ul>\n<li>no MAS</li>\n<li>modified from paper to use only one model (no ensemble, no refinement, no denosing)</li>\n<li>only ouput one (no conformation)</li>\n<li>takes 15 min for 12 seq</li>\n</ul>\n<p>(no limit max length, local PC)<br>\n[0.54611, 0.57469, 0.49749, 0.24299, 0.2387, 0.96262, 0.24303, 0.1964, 0.61831, 0.45845, 0.21084, 0.24965]<br>\nmean: 0.41994</p>\n<p>example notebook:<br>\n(limited to max length=400)<br>\nlink: <a href=\"https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa</a></p>\n<pre><code>ALL\n [., , , , ., , , , , , , ]\nMEAN \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3181750,
          "author_name": "akomwins",
          "author_url": "",
          "post_date": "04/18/2025 09:31:15",
          "content": "<p>Does it help if we apply the other second structure in the code instead of setting it None?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3163907,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "03/31/2025 09:40:43",
      "content": "<p>According to their supplemental material, it scored slightly better than AF3 for CASP16 monomers (20 targets). Looks promising.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff290f375ace4cdc9d3ece2bce29631f4%2FScreenshot%202025-03-31%20at%2018.25.44.png?generation=1743413988586715&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3185100,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "04/22/2025 20:20:25",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I got the same results as you with DRfold2, however I saw a major fall in performance in the train set compared to the validation/test set. For instance, the mean score on the train set was around 0.23, could it be that it was trained with similar data to CASP15/CASP16 ? I wander how it will perform after the data update</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3182960,
      "author_name": "doheon114",
      "author_url": "",
      "post_date": "04/20/2025 07:07:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , I have a question about using drfold2 in this competition. DRfold2's models are quite light so it spends less time to predict individual models, but actually we have to compare 80 models in repository so it takes lots of time to predict on all 80 models. How did you overcome this? Just removing some models or select top 20 best performing models among 80 by little experiments? What could be the solution. Thank you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3148898": "https://github.com/leeyang/DRfold2\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F372fef9462664504099729b0f03e164a%2FSelection_999(7767).png?generation=1741882290600806&alt=media)\n\n---\n\nDRfold2: Ab initio RNA structure prediction with composite language model and denoised end-to-end learning - bioxiv2025- mar\nhttps://www.biorxiv.org/content/10.1101/2025.03.05.641632v1.full.pdf\n(input = seq only)\n\nQUOTE \"Detailed analyses reveal that the improvements primarily stem from the RCLM’s ability to capture co-evolutionary pattern and the effective denoising process, leading to a more than 100% increase in contact prediction precision compared to existing methods\"\n\n- inlcude experiment results on casp15/16, RNA-Puzzles targets\n- comparable to Alphafold3\n- denoising to predict multiple conformation\n\nNOTE: one weakness\nQUOTE \"Notably, large synthetic RNAs from CASP15 were excluded due to their deviation from naturally occurring RNA\"\n(maybe that is not learned in RCLM’s?)\n\n---\n\nrelated:\n[1] DRfold : Integrating end-to-end learning with deep geometrical potentials for ab initio RNA structure prediction - 2023-Nov\nhttps://www.nature.com/articles/s41467-023-41303-9\n(input = seq + SS (no MSA))\n\n\n[2] Systematic benchmarking of deep-learning methods for tertiary RNA structure prediction - 2024-nov\nhttps://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715\n(benchmark DRfold on CASP15, RNA-Puzzles)",
    "3148922": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F95afd734bbe8590680a4f4166a169885%2FSelection_999(7768).png?generation=1741882960511664&alt=media)\n\nThe RCLM was trained on ~30M RNA sequences from the RNAcentral database (Release 22) over 67,000 batches, with a batch size of 128. The entire training process took about 15 days using a single NVIDIA A40 GPU.",
    "3151873": "local cv results (kaggle-casp15)\n- no MAS\n- modified from paper to use only one model (no ensemble, no refinement, no denosing)\n- only ouput one (no conformation)\n- takes 15 min for 12 seq\n\n(no limit max length, local PC)\n[0.54611, 0.57469, 0.49749, 0.24299, 0.2387, 0.96262, 0.24303, 0.1964, 0.61831, 0.45845, 0.21084, 0.24965]\nmean: 0.41994\n\nexample notebook:\n(limited to max length=400)\nlink: https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa\n```\nALL\n [0.5461, 0.57469, 0.49749, 0.24299, 0.2387, 0.96262, 0.24303, 0.38226, 0.61831, 0.45845, 0.21084, 0.24965]\nMEAN 0.4354275\n```",
    "3163907": "According to their supplemental material, it scored slightly better than AF3 for CASP16 monomers (20 targets). Looks promising.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ff290f375ace4cdc9d3ece2bce29631f4%2FScreenshot%202025-03-31%20at%2018.25.44.png?generation=1743413988586715&alt=media)",
    "3181750": "Does it help if we apply the other second structure in the code instead of setting it None?",
    "3182960": "Hi @hengck23 , I have a question about using drfold2 in this competition. DRfold2's models are quite light so it spends less time to predict individual models, but actually we have to compare 80 models in repository so it takes lots of time to predict on all 80 models. How did you overcome this? Just removing some models or select top 20 best performing models among 80 by little experiments? What could be the solution. Thank you.",
    "3185100": "hengck23 I got the same results as you with DRfold2, however I saw a major fall in performance in the train set compared to the validation/test set. For instance, the mean score on the train set was around 0.23, could it be that it was trained with similar data to CASP15/CASP16 ? I wander how it will perform after the data update"
  },
  "source": "meta"
}