{
  "id": 566906,
  "title": "[placeholder lb0.321/0.500] My solution and experimental results",
  "url": "/competitions/stanford-rna-3d-folding/discussion/566906",
  "author_name": "hengck23",
  "post_date": "2025-03-07T10:34:36.855000",
  "votes": 121,
  "comment_count": 107,
  "views": 0,
  "content": "<p>… to be updated as experiment proceeds … each week, we detail how to improve your lb score …please come here often </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F60350c3075fb64f8b8e86240fdb1894a%2FSelection_096.png?generation=1742398747189706&amp;alt=media\" alt=\"\"><br>\n<strong>baseline code:</strong><br>\n1) 18-mar : lb 0.321</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/lb0-286-simple-drfold-no-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-286-simple-drfold-no-msa</a></li>\n<li>no MSA, instead a RNA language model to model evolutionary information</li>\n<li>overall, it is large RNA-lm and smaller structure downstream model</li>\n<li>use 5 out of 100 avaliable  smaller structure models (e.g. 16 mb) for current lb</li>\n<li>more structure models is better!!</li>\n<li>plan: <br>\n1.energy scoring to select best model results, clustering? (see paper)<br>\n2.train model to convert from (N,C,P)frame + distance map --&gt; C1 backbone<br>\n3.quantisation and fp16 to extend length to 800?</li>\n</ul>\n<p>i think Drfold2 alone can get to lb0.40 on public test?</p>\n<hr>\n<h2>Acknowledgement</h2>\n<p>\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"</p>\n<h2>Hardware</h2>\n<p>GPU: 2x Nvidia Ada A6000 (Ampere), each with VRAM 48 GB<br>\nCPU: Intel® Xeon(R) w7-3455 CPU @ 2.5GHz, 24 cores, 48 threads<br>\nMemory: 256 GB RAM</p>",
  "messages": [
    {
      "id": 3143558,
      "postDate": "2025-03-07T10:34:36.857Z",
      "content": "<p>… to be updated as experiment proceeds … each week, we detail how to improve your lb score …please come here often </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F60350c3075fb64f8b8e86240fdb1894a%2FSelection_096.png?generation=1742398747189706&amp;alt=media\" alt=\"\"><br>\n<strong>baseline code:</strong><br>\n1) 18-mar : lb 0.321</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/lb0-286-simple-drfold-no-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-286-simple-drfold-no-msa</a></li>\n<li>no MSA, instead a RNA language model to model evolutionary information</li>\n<li>overall, it is large RNA-lm and smaller structure downstream model</li>\n<li>use 5 out of 100 avaliable  smaller structure models (e.g. 16 mb) for current lb</li>\n<li>more structure models is better!!</li>\n<li>plan: <br>\n1.energy scoring to select best model results, clustering? (see paper)<br>\n2.train model to convert from (N,C,P)frame + distance map --&gt; C1 backbone<br>\n3.quantisation and fp16 to extend length to 800?</li>\n</ul>\n<p>i think Drfold2 alone can get to lb0.40 on public test?</p>\n<hr>\n<h2>Acknowledgement</h2>\n<p>\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"</p>\n<h2>Hardware</h2>\n<p>GPU: 2x Nvidia Ada A6000 (Ampere), each with VRAM 48 GB<br>\nCPU: Intel® Xeon(R) w7-3455 CPU @ 2.5GHz, 24 cores, 48 threads<br>\nMemory: 256 GB RAM</p>",
      "rawMarkdown": "... to be updated as experiment proceeds ... each week, we detail how to improve your lb score ...please come here often \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F60350c3075fb64f8b8e86240fdb1894a%2FSelection_096.png?generation=1742398747189706&alt=media)\n**baseline code:**\n1) 18-mar : lb 0.321\n- https://www.kaggle.com/code/hengck23/lb0-286-simple-drfold-no-msa\n- no MSA, instead a RNA language model to model evolutionary information\n- overall, it is large RNA-lm and smaller structure downstream model\n- use 5 out of 100 avaliable  smaller structure models (e.g. 16 mb) for current lb\n- more structure models is better!!\n- plan: \n1.energy scoring to select best model results, clustering? (see paper)\n2.train model to convert from (N,C,P)frame + distance map --> C1 backbone\n3.quantisation and fp16 to extend length to 800?\n\ni think Drfold2 alone can get to lb0.40 on public test?\n\n----\n\n##Acknowledgement\n\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"\n\n##Hardware\nGPU: 2x Nvidia Ada A6000 (Ampere), each with VRAM 48 GB\nCPU: Intel® Xeon(R) w7-3455 CPU @ 2.5GHz, 24 cores, 48 threads\nMemory: 256 GB RAM",
      "votes": 121
    },
    {
      "id": 3147190,
      "postDate": "2025-03-11T18:42:17.233Z",
      "content": "<p><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a><br>\nA trainable PyTorch reproduction of AlphaFold 3.</p>\n<p>curious how the diffuser work …</p>",
      "rawMarkdown": "https://github.com/bytedance/Protenix\nA trainable PyTorch reproduction of AlphaFold 3.\n\ncurious how the diffuser work ...",
      "votes": 12,
      "replies": [
        {
          "id": 3151989,
          "postDate": "2025-03-17T09:55:22.117Z",
          "content": "<p>Have you ever tried using this? I'm about to implement it on a Kaggle notebook, but even running inference seems challenging.</p>",
          "rawMarkdown": "Have you ever tried using this? I'm about to implement it on a Kaggle notebook, but even running inference seems challenging.",
          "replies": [
            {
              "id": 3151998,
              "postDate": "2025-03-17T10:12:26.730Z",
              "content": "<p>you need to simplify!!!</p>\n<p>i already about to make drfold2 work. i need a complementary model (for assembly)and AF3 is the recommended one.<br>\ni will try AAF3 from next week.</p>\n<p>drfold2 to be out in public notebook soon</p>",
              "rawMarkdown": "you need to simplify!!!\n\ni already about to make drfold2 work. i need a complementary model (for assembly)and AF3 is the recommended one.\ni will try AAF3 from next week.\n\ndrfold2 to be out in public notebook soon",
              "votes": 1
            },
            {
              "id": 3152015,
              "postDate": "2025-03-17T10:33:07.597Z",
              "content": "<p><a href=\"https://www.kaggle.com/doheon114\" target=\"_blank\">@doheon114</a> <br>\n<a href=\"https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa</a></p>",
              "rawMarkdown": "@doheon114 \nhttps://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 3155466,
      "postDate": "2025-03-21T04:16:59.447Z",
      "content": "<p>the most important benchmark, must read!!!!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffba9ee0c870e82e97dbb5c62f68655e4%2FSelection_116.png?generation=1742659558654264&amp;alt=media\" alt=\"\"></p>\n<p>notebook: <a href=\"https://www.kaggle.com/code/hengck23/vfold-baseline-offline\" target=\"_blank\">https://www.kaggle.com/code/hengck23/vfold-baseline-offline</a></p>\n<p>It's frustrating that my big local improvements on DRfold2 didn’t show up as gains on the leaderboard.<br>\nso I suspect there is some RNA that is usually difficult for drfold ONLY.<br>\ni think I am correct</p>\n<p>drfoldv2 totally breakdown at rna &gt;400<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F84b48cea4acbfaf14c3bd590d016128a%2FSelection_105.png?generation=1742534073508454&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "the most important benchmark, must read!!!!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffba9ee0c870e82e97dbb5c62f68655e4%2FSelection_116.png?generation=1742659558654264&alt=media)\n\nnotebook: https://www.kaggle.com/code/hengck23/vfold-baseline-offline\n\nIt's frustrating that my big local improvements on DRfold2 didn’t show up as gains on the leaderboard.\nso I suspect there is some RNA that is usually difficult for drfold ONLY.\ni think I am correct\n\n\ndrfoldv2 totally breakdown at rna >400\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F84b48cea4acbfaf14c3bd590d016128a%2FSelection_105.png?generation=1742534073508454&alt=media)\n",
      "votes": 10,
      "replies": [
        {
          "id": 3155521,
          "postDate": "2025-03-21T05:32:28.690Z",
          "content": "<p>Also my bottleneck… </p>\n<p>Besides, I tried different sotas, and RNAs with long sequences can lead to OOM, because the Kaggle env is limited. </p>\n<p>Can we beat the vfold baseline in the Kaggle env? I have my doubts. </p>",
          "rawMarkdown": "Also my bottleneck... \n\nBesides, I tried different sotas, and RNAs with long sequences can lead to OOM, because the Kaggle env is limited. \n\nCan we beat the vfold baseline in the Kaggle env? I have my doubts. ",
          "replies": [
            {
              "id": 3155584,
              "postDate": "2025-03-21T07:24:11.783Z",
              "content": "<p>\"Can we beat the vfold baseline in the Kaggle env?\" <br>\nyes I think we can …. but much efforts</p>\n<p>I will release a model that can run 800 rna length on kaggler notebook soon. need to patch triangular attnetion</p>",
              "rawMarkdown": "\"Can we beat the vfold baseline in the Kaggle env?\" \nyes I think we can .... but much efforts\n\nI will release a model that can run 800 rna length on kaggler notebook soon. need to patch triangular attnetion",
              "votes": 1
            }
          ]
        },
        {
          "id": 3155582,
          "postDate": "2025-03-21T07:22:45.413Z",
          "content": "<p>casp16 rna target list from casp website</p>",
          "rawMarkdown": "casp16 rna target list from casp website\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 3154714,
      "postDate": "2025-03-20T10:12:22.017Z",
      "content": "<p>vfold is the benchmark and we should understand it:<br>\n<a href=\"https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf\" target=\"_blank\">https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf</a><br>\nhint: go to casp16 website to be vfold tm score and make a rank chart</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8ff6ed5e9aef06780c5d9e6f545121bf%2FSelection_099.png?generation=1742465512170143&amp;alt=media\" alt=\"\"></p>\n<p>U must be very good at short length</p>",
      "rawMarkdown": "vfold is the benchmark and we should understand it:\nhttps://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf\nhint: go to casp16 website to be vfold tm score and make a rank chart\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8ff6ed5e9aef06780c5d9e6f545121bf%2FSelection_099.png?generation=1742465512170143&alt=media)\n\nU must be very good at short length",
      "votes": 8
    },
    {
      "id": 3148681,
      "postDate": "2025-03-13T12:21:59.103Z",
      "content": "<p>I shared a technique for experimenting with MSA on the \"open hidden\" test set for the public leaderboard.</p>\n<p><strong>Please note that this is for testing purposes only and should not be considered a legitimate solution for claiming any prizes.</strong></p>\n<ol>\n<li>go to casp16 and find the sequence fasta of rna targets</li>\n<li>you can now predict their 3d structure offline (e.g. using web mas server)</li>\n<li>the host said that some public test rna are indeed casp16 target, so in your submission code use:</li>\n</ol>\n<pre><code>OFFLINE_PREDICT={\n    ...' :  xyz  ...,\n    ...' :  xyz  ...,\n}\n\n\n i,row  test_df ...:\n     row.sequence  found  OFFLINE_PREDICT :\n             = predicted xyz \n      found:\n           xyz =  online model\n</code></pre>\n<p>in my experiments, I thin I can conclude:<br>\nif you do very well in casp15, casp16, rna puzzles you should do well in this competition in public and private hidden dataset for the three different phrases. </p>\n<p>do well doesn't mean train on these dataset but your model can generalise on these datasets.<br>\ni think the level of difficulty of this competition is similar to these datasets.</p>\n<hr>\n<p>what you can do next is to measure the alignment score (% of matched seq) on these dataset to kaggle train (or msa) and kaggle 400k diffusion synthetic data. This will tell you if these are inlier or outlier data.</p>\n<p>ML methdos tends to work (only?) on inliers data. </p>",
      "rawMarkdown": "I shared a technique for experimenting with MSA on the \"open hidden\" test set for the public leaderboard.\n\n**Please note that this is for testing purposes only and should not be considered a legitimate solution for claiming any prizes.**\n\n1. go to casp16 and find the sequence fasta of rna targets\n2. you can now predict their 3d structure offline (e.g. using web mas server)\n3. the host said that some public test rna are indeed casp16 target, so in your submission code use:\n\n```\nOFFLINE_PREDICT={\n   'AAAGGGCCC ...' : predicted xyz  ...,\n   'ACGACGACG ...' : predicted xyz  ...,\n}\n\n\nfor i,row in test_df ...:\n    if row.sequence is found in OFFLINE_PREDICT :\n            xyz = predicted xyz \n   else if not found:\n           xyz = use online model\n\n```\n\nin my experiments, I thin I can conclude:\nif you do very well in casp15, casp16, rna puzzles you should do well in this competition in public and private hidden dataset for the three different phrases. \n\ndo well doesn't mean train on these dataset but your model can generalise on these datasets.\ni think the level of difficulty of this competition is similar to these datasets.\n\n\n---\n\nwhat you can do next is to measure the alignment score (% of matched seq) on these dataset to kaggle train (or msa) and kaggle 400k diffusion synthetic data. This will tell you if these are inlier or outlier data.\n\nML methdos tends to work (only?) on inliers data. ",
      "votes": 7,
      "replies": [
        {
          "id": 3148704,
          "postDate": "2025-03-13T13:02:32.787Z",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>If I'm not mistaken, you're thinking of using data similar to the test dataset (CASP15, 16) to create pseudo-labels, right?</p>\n<blockquote>\n  <p>you can now predict their 3d structure offline (e.g. using web mas server)</p>\n</blockquote>\n<p>Why offline  (e.g. using web mas server)?  submit fasta RNA sequences then wait for the consequence, or run the sota locally to create pseudo-labels?</p>",
          "rawMarkdown": "Thank you very much, @hengck23 \n\nIf I'm not mistaken, you're thinking of using data similar to the test dataset (CASP15, 16) to create pseudo-labels, right?\n\n>you can now predict their 3d structure offline (e.g. using web mas server)\n\nWhy offline  (e.g. using web mas server)?  submit fasta RNA sequences then wait for the consequence, or run the sota locally to create pseudo-labels?",
          "replies": [
            {
              "id": 3148713,
              "postDate": "2025-03-13T13:09:49.927Z",
              "content": "<p>i give you an example.</p>\n<ul>\n<li>say i want to inlcude msa as input to ribonanaza-net.</li>\n<li>say in local experiments with kaggle 800 rna with msa train/validation split, it work very well.</li>\n<li>it is not possible to setup msa server in kaggle notebook (for now … but maybe possible after a month) so i cannot know msa lb score.</li>\n<li>but i can do local prediction on \"open hidde\" public data.</li>\n<li>i can submit local prediction to get lb score.</li>\n<li>if lb score is good, then i can focus effects on how to setup reduced msa server in kaggle notebook or replaced the msa part for legitimate solution.</li>\n</ul>\n<hr>\n<p>my style of work is</p>\n<ul>\n<li>first prove that model work (or at least you can get an upper bound estimate)</li>\n<li>then it is engineering work … how to implement</li>\n</ul>",
              "rawMarkdown": "i give you an example.\n- say i want to inlcude msa as input to ribonanaza-net.\n- say in local experiments with kaggle 800 rna with msa train/validation split, it work very well.\n- it is not possible to setup msa server in kaggle notebook (for now ... but maybe possible after a month) so i cannot know msa lb score.\n- but i can do local prediction on \"open hidde\" public data.\n- i can submit local prediction to get lb score.\n- if lb score is good, then i can focus effects on how to setup reduced msa server in kaggle notebook or replaced the msa part for legitimate solution.\n\n---\n\nmy style of work is\n- first prove that model work (or at least you can get an upper bound estimate)\n- then it is engineering work ... how to implement",
              "votes": 4
            },
            {
              "id": 3148718,
              "postDate": "2025-03-13T13:12:39.877Z",
              "content": "<p>there is no point creating pesudo label for casp15/16. this train set is too small.<br>\nwhat we should is is create pesudo label for a large dataset whose distribution is such that casp15/16 are inliers</p>",
              "rawMarkdown": "there is no point creating pesudo label for casp15/16. this train set is too small.\nwhat we should is is create pesudo label for a large dataset whose distribution is such that casp15/16 are inliers",
              "votes": 2
            },
            {
              "id": 3148729,
              "postDate": "2025-03-13T13:21:00.297Z",
              "content": "<p>Nice trick! As you say, it's not a generally applicable approach but kind of probing, but I really learn something.</p>",
              "rawMarkdown": "Nice trick! As you say, it's not a generally applicable approach but kind of probing, but I really learn something."
            },
            {
              "id": 3148733,
              "postDate": "2025-03-13T13:23:52.553Z",
              "content": "<blockquote>\n  <p>there is no point creating pesudo label for casp15/16. this train set is too small.<br>\n  what we should is is create pesudo label for a large dataset whose distribution is such that casp15/16 are inliers</p>\n</blockquote>\n<p>It's true alright. you resolve my confusion.</p>",
              "rawMarkdown": ">there is no point creating pesudo label for casp15/16. this train set is too small.\nwhat we should is is create pesudo label for a large dataset whose distribution is such that casp15/16 are inliers\n\nIt's true alright. you resolve my confusion."
            }
          ]
        }
      ]
    },
    {
      "id": 3144066,
      "postDate": "2025-03-07T22:06:53.400Z",
      "content": "<p>here i want to introduce some unconventional papers related to rna/dna/protein 3d structure prediction which may be useful for our competition.</p>\n<p>1) deepseek R1 like method …. the ah-ha moment<br>\n<a href=\"https://x.com/hla_michael/status/1898106485005336988\" target=\"_blank\">https://x.com/hla_michael/status/1898106485005336988</a><br>\n\"Introducing Pro-1, an 8b param reasoning model trained using GRPO towards a physics based reward function for protein stability.\"</p>\n<p>though process:<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13038ccdb143e2cb6b0b218b5f238ef%2FSelection_999(7748).png?generation=1741385202898702&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "here i want to introduce some unconventional papers related to rna/dna/protein 3d structure prediction which may be useful for our competition.\n\n1) deepseek R1 like method .... the ah-ha moment\nhttps://x.com/hla_michael/status/1898106485005336988\n\"Introducing Pro-1, an 8b param reasoning model trained using GRPO towards a physics based reward function for protein stability.\"\n\nthough process:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13038ccdb143e2cb6b0b218b5f238ef%2FSelection_999(7748).png?generation=1741385202898702&alt=media)",
      "votes": 8
    },
    {
      "id": 3156237,
      "postDate": "2025-03-21T23:28:35.053Z",
      "content": "<p>what i learn from CASP16:</p>\n<ul>\n<li>modeling is not enough !!! focus on pipeline (combinations of methods)</li>\n<li>AF3 is part of many solution</li>\n</ul>\n<p>for kaggle, it would be 40 test targets in 8 GPU hours, meaning 10 min per target (5 RNA structures)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ed13500816d02ea899ca5ac89c9d3cf%2FSelection_112.png?generation=1742599652340594&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0ff2bd3c3067a9aef7c937ec47e6d08a%2FSelection_113.png?generation=1742599706052300&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9d7d1f53441d58dd9f8c4e699d0debf6%2FSelection_108.png?generation=1742599676893502&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "what i learn from CASP16:\n- modeling is not enough !!! focus on pipeline (combinations of methods)\n- AF3 is part of many solution\n\nfor kaggle, it would be 40 test targets in 8 GPU hours, meaning 10 min per target (5 RNA structures)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ed13500816d02ea899ca5ac89c9d3cf%2FSelection_112.png?generation=1742599652340594&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0ff2bd3c3067a9aef7c937ec47e6d08a%2FSelection_113.png?generation=1742599706052300&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9d7d1f53441d58dd9f8c4e699d0debf6%2FSelection_108.png?generation=1742599676893502&alt=media)",
      "votes": 5
    },
    {
      "id": 3158245,
      "postDate": "2025-03-24T11:08:22.713Z",
      "content": "<p>I find that drfold2 implements alphafold2. The notation follows exactly af2 supplementary paper. The only replace msa generation with their rna lm. There is no training script or loss in the drfold2 repo  but one can just follow af2</p>",
      "rawMarkdown": "I find that drfold2 implements alphafold2. The notation follows exactly af2 supplementary paper. The only replace msa generation with their rna lm. There is no training script or loss in the drfold2 repo  but one can just follow af2",
      "votes": 6
    },
    {
      "id": 3150672,
      "postDate": "2025-03-15T18:29:02.573Z",
      "content": "<p>prelimnary results of DRfold2 is very good!<br>\nR1107 : top 3 prediction 0.55309, 0.31959, 0.30806<br>\nR1116: top3 prediction 0.61168,0.59528</p>\n<p>i am surprised that such good results is possible without MSA input</p>",
      "rawMarkdown": "prelimnary results of DRfold2 is very good!\nR1107 : top 3 prediction 0.55309, 0.31959, 0.30806\nR1116: top3 prediction 0.61168,0.59528\n\ni am surprised that such good results is possible without MSA input\n",
      "votes": 6,
      "replies": [
        {
          "id": 3152704,
          "postDate": "2025-03-18T04:48:20.150Z",
          "content": "<p>here is drfold2 magic:</p>\n<p>in the github repo, 5 set of  20 model checkpoints are provided.<br>\nif I use one full set (20 models) to predict 20 conformations, here is the tm scores for first casp15 target  R1107</p>\n<pre><code> R1107  GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU...\n### tm: [., ., , , , , , , , , , , , , , ., , , , ]\n\nMAX \n</code></pre>\n<p>it is a very high value of 0.61 if I select the best score.<br>\nthe repo code has a lot of redundant computation. (actually repeated computation) so you need to clean it up.</p>\n<p>instead of reading the code, a fast way to detect \"unless\" computation is to print the intermediate tensor output values out at critical steps of the network for 20 models on the same input seq. you will see that some steps produce identical values, aka, redundancy found!</p>",
          "rawMarkdown": "here is drfold2 magic:\n\nin the github repo, 5 set of  20 model checkpoints are provided.\nif I use one full set (20 models) to predict 20 conformations, here is the tm scores for first casp15 target  R1107\n```\n0 R1107 69 GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU...\n### tm: [0.5461, 0.1873, 0.25313, 0.23874, 0.42539, 0.29271, 0.28274, 0.29979, 0.48901, 0.27221, 0.56012, 0.49345, 0.55223, 0.40821, 0.35545, 0.233, 0.61409, 0.58367, 0.45455, 0.48951]\n \nMAX 0.61409\n```\n\nit is a very high value of 0.61 if I select the best score.\nthe repo code has a lot of redundant computation. (actually repeated computation) so you need to clean it up.\n\ninstead of reading the code, a fast way to detect \"unless\" computation is to print the intermediate tensor output values out at critical steps of the network for 20 models on the same input seq. you will see that some steps produce identical values, aka, redundancy found!",
          "replies": [
            {
              "id": 3152798,
              "postDate": "2025-03-18T07:07:40.847Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I commented out the processing after the \"Selection\" part in DRfold_infer.py and modified it to obtain the coordinates of the C4 atoms from the initially generated pdb file. I also limited the model_hub weights to just one. Additionally, I downgraded the torch version to 2.2.1 so that it runs on Kaggle Notebooks as well. Is your environment set up similarly?</p>",
              "rawMarkdown": "@hengck23 I commented out the processing after the \"Selection\" part in DRfold_infer.py and modified it to obtain the coordinates of the C4 atoms from the initially generated pdb file. I also limited the model_hub weights to just one. Additionally, I downgraded the torch version to 2.2.1 so that it runs on Kaggle Notebooks as well. Is your environment set up similarly?"
            },
            {
              "id": 3152800,
              "postDate": "2025-03-18T07:09:59.437Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Based on these results, could it be that rather than drfold2 being exceptionally superior, it’s rhofold or nufold that require particularly high-quality and extensive MSAs? If that’s the case, then drfold2’s ensemble seems to hold the highest potential.</p>",
              "rawMarkdown": "@hengck23 Based on these results, could it be that rather than drfold2 being exceptionally superior, it’s rhofold or nufold that require particularly high-quality and extensive MSAs? If that’s the case, then drfold2’s ensemble seems to hold the highest potential."
            },
            {
              "id": 3152818,
              "postDate": "2025-03-18T07:42:36.003Z",
              "content": "<p>\" … e processing after the \"Selection\" part in DRfold_infer.py …\"<br>\ni use pytorch 2.5. and the orginal version in kaggle notebook is ok with me.</p>\n<p>whatever you do, just note that:</p>\n<ul>\n<li>we jsut need output from MSA2XYZ model:</li>\n</ul>\n<pre><code>        ret = x_pre()()()\n        ret = F(self(z),dim=-)()()()(np.float16)\n        ret = F(self(z),dim=-)()()()(np.float16)\n        ret = F(self(z),dim=-)()()()(np.float16)\n        ret = plddt()()()\n        return ret\n</code></pre>\n<p>i am using only   ret['coor'], which is the (P,C4,N1/9) frame in my description below<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512</a></p>\n<p>ret['coor'] is Lx3x3 = L x(P,C4,N1/9) x (xyz)<br>\nbut kaggle needs C1 instaed of C4.</p>\n<p>we may need to make another layer to predict C1 from (P,C4,N1/9) frame, etc</p>\n<p>as a temporarily solution, the Optimzation (in potentialFold) code has a save to save Lx3x3 frame into full atom pdb. i am using that. .</p>\n<pre><code>rama=.init_quat(   your MSA2XYZ output  )..numpy()\nrama = rama  \n.outpdb(rama ,energystr=)\n</code></pre>",
              "rawMarkdown": "\" ... e processing after the \"Selection\" part in DRfold_infer.py ...\"\ni use pytorch 2.5. and the orginal version in kaggle notebook is ok with me.\n\nwhatever you do, just note that:\n- we jsut need output from MSA2XYZ model:\n\n```\n        ret['coor'] = x_pre.detach().cpu().numpy()\n        ret['dist_p'] = F.softmax(self.pdis_predor(z),dim=-1).detach().cpu().numpy().astype(np.float16)\n        ret['dist_c'] = F.softmax(self.cdis_predor(z),dim=-1).detach().cpu().numpy().astype(np.float16)\n        ret['dist_n'] = F.softmax(self.ndis_predor(z),dim=-1).detach().cpu().numpy().astype(np.float16)\n        ret['plddt'] = plddt.detach().cpu().numpy()\n        return ret\n\n\n```\ni am using only   ret['coor'], which is the (P,C4,N1/9) frame in my description below\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512\n\nret['coor'] is Lx3x3 = L x(P,C4,N1/9) x (xyz)\nbut kaggle needs C1 instaed of C4.\n\nwe may need to make another layer to predict C1 from (P,C4,N1/9) frame, etc\n\nas a temporarily solution, the Optimzation (in potentialFold) code has a save to save Lx3x3 frame into full atom pdb. i am using that. .\n\n```\nrama=self.init_quat(  ... your MSA2XYZ output ... ).data.numpy()\nrama = rama ... \nself.outpdb(rama 'save.pdb',energystr='if you want to compute engery')\n\n```"
            },
            {
              "id": 3152821,
              "postDate": "2025-03-18T07:44:17.607Z",
              "content": "<p>\"e, then drfold2’s ensemble seems to hold the highest potential.\"</p>\n<p>winning solution after 3 month would be combintions of these and non deep learning methods</p>",
              "rawMarkdown": "\"e, then drfold2’s ensemble seems to hold the highest potential.\"\n\nwinning solution after 3 month would be combintions of these and non deep learning methods",
              "votes": 1
            },
            {
              "id": 3152834,
              "postDate": "2025-03-18T07:54:52.163Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3153184,
              "postDate": "2025-03-18T13:45:57.777Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for your reply. I see… this is very insightful. I will try this as well. Your experiment notes are very educational—thank you!</p>",
              "rawMarkdown": "@hengck23 Thank you for your reply. I see... this is very insightful. I will try this as well. Your experiment notes are very educational—thank you!"
            },
            {
              "id": 3153562,
              "postDate": "2025-03-19T01:12:03.997Z",
              "content": "<p>In my understanding, optimization requires force field optimization to obtain C1, but this process takes a long time. How can we speed it up? Thank you for your help.</p>",
              "rawMarkdown": "In my understanding, optimization requires force field optimization to obtain C1, but this process takes a long time. How can we speed it up? Thank you for your help."
            },
            {
              "id": 3153581,
              "postDate": "2025-03-19T01:38:52.173Z",
              "content": "<p>the optimization optimizes (P,C4,N1/9) by computing local transformation (rotation + translation). then location of C1 and rest of the atoms obtained from transformations. </p>\n<p>we are not optimizing, rather we just get location of C1 from current transformations. </p>",
              "rawMarkdown": "the optimization optimizes (P,C4,N1/9) by computing local transformation (rotation + translation). then location of C1 and rest of the atoms obtained from transformations. \n\nwe are not optimizing, rather we just get location of C1 from current transformations. "
            },
            {
              "id": 3155490,
              "postDate": "2025-03-21T04:49:35.817Z",
              "content": "<p><a href=\"https://www.kaggle.com/shosukesuzuki\" target=\"_blank\">@shosukesuzuki</a> <br>\nnufold is very important (it is 3rd in CASP)<br>\nagain, the MSA issue needs to be solved for nufold,alphafold3 and rhofold+.</p>\n<p>need to do experiment to find the MSA pipline they are using, then request host to prepare MSA accoriding to our pipline script</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2174f217b283996332fb89518eef965a%2FSelection_999(7818).png?generation=1742532472691036&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@shosukesuzuki \nnufold is very important (it is 3rd in CASP)\nagain, the MSA issue needs to be solved for nufold,alphafold3 and rhofold+.\n\nneed to do experiment to find the MSA pipline they are using, then request host to prepare MSA accoriding to our pipline script\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2174f217b283996332fb89518eef965a%2FSelection_999(7818).png?generation=1742532472691036&alt=media)",
              "votes": 2
            },
            {
              "id": 3156874,
              "postDate": "2025-03-22T17:09:22.657Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> are you finetuning by adding some layers for c1 atoms prediction?  i tried using the local base frame landed up in 0.268LB</p>",
              "rawMarkdown": "@hengck23 are you finetuning by adding some layers for c1 atoms prediction?  i tried using the local base frame landed up in 0.268LB"
            },
            {
              "id": 3180008,
              "postDate": "2025-04-16T02:55:27.240Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/shosukesuzuki\" target=\"_blank\">@shosukesuzuki</a> are you training any additional models for C1 prediction or are there any inbuilt functions in Drfold2 to accurately predict C1? </p>",
              "rawMarkdown": "@hengck23 @shosukesuzuki are you training any additional models for C1 prediction or are there any inbuilt functions in Drfold2 to accurately predict C1? "
            }
          ]
        }
      ]
    },
    {
      "id": 3162463,
      "postDate": "2025-03-29T08:43:41.583Z",
      "content": "<p>triangle flash attnetion!!!!<br>\n<a href=\"https://arxiv.org/pdf/2412.10743v2\" target=\"_blank\">https://arxiv.org/pdf/2412.10743v2</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e5bf4c2e18b24f34f18e286691bcda3%2FSelection_166.png?generation=1743237819947581&amp;alt=media\" alt=\"\"></p>\n<p>related?<br>\n<a href=\"https://github.com/Dao-AILab/flash-attention/pull/57\" target=\"_blank\">https://github.com/Dao-AILab/flash-attention/pull/57</a></p>",
      "rawMarkdown": "triangle flash attnetion!!!!\nhttps://arxiv.org/pdf/2412.10743v2\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e5bf4c2e18b24f34f18e286691bcda3%2FSelection_166.png?generation=1743237819947581&alt=media)\n\nrelated?\nhttps://github.com/Dao-AILab/flash-attention/pull/57",
      "votes": 3
    },
    {
      "id": 3207689,
      "postDate": "2025-05-23T06:13:16.780Z",
      "content": "<p>A little complex for me, achieve score about 0.4 is pretty hard!</p>",
      "rawMarkdown": "A little complex for me, achieve score about 0.4 is pretty hard!",
      "votes": 1
    },
    {
      "id": 3154165,
      "postDate": "2025-03-19T15:40:48.393Z",
      "content": "<p>energy as an effective way to select best 5 structure for submission</p>\n<p>graph show energy vs tm score (rank-normalised)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7e5d1eacb7fd0097e25c766b094bc10%2FSelection_095.png?generation=1742398800712595&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "energy as an effective way to select best 5 structure for submission\n\ngraph show energy vs tm score (rank-normalised)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7e5d1eacb7fd0097e25c766b094bc10%2FSelection_095.png?generation=1742398800712595&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 3154253,
          "postDate": "2025-03-19T17:56:27.307Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3180723,
          "postDate": "2025-04-17T03:35:47.720Z",
          "content": "<p>Could you share how to calculate energy and how long does it take?</p>",
          "rawMarkdown": "Could you share how to calculate energy and how long does it take?"
        }
      ]
    },
    {
      "id": 3152509,
      "postDate": "2025-03-17T22:16:34.177Z",
      "content": "<p>archive:<br>\n<strong>baseline code:</strong><br>\n1) 11-mar : lb 0.196</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/baseline-droupout\" target=\"_blank\">https://www.kaggle.com/code/hengck23/baseline-droupout</a></li>\n<li>just train at constant 1e-4 for 200 epoch. follow other setup (loss, etc) from host example notebook</li>\n<li>note that better CV may not gives better LB (dropout gives less variation) </li>\n<li>long epoch is required for alignment to be stable and aligned loss to be useful<br>\n(a better alignment free loss function is required)</li>\n</ul>",
      "rawMarkdown": "archive:\n**baseline code:**\n1) 11-mar : lb 0.196\n- https://www.kaggle.com/code/hengck23/baseline-droupout\n- just train at constant 1e-4 for 200 epoch. follow other setup (loss, etc) from host example notebook\n- note that better CV may not gives better LB (dropout gives less variation) \n- long epoch is required for alignment to be stable and aligned loss to be useful\n(a better alignment free loss function is required)\n",
      "votes": 3
    },
    {
      "id": 3153162,
      "postDate": "2025-03-18T13:24:09.710Z",
      "content": "<p>current plan for the week<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f98bc059b250bf822c2577a6a90c61c%2FSelection_086.png?generation=1742304205309135&amp;alt=media\" alt=\"\"></p>\n<p>the deep net code should run 3x faster then the repo after refractoring</p>",
      "rawMarkdown": "current plan for the week\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f98bc059b250bf822c2577a6a90c61c%2FSelection_086.png?generation=1742304205309135&alt=media)\n\nthe deep net code should run 3x faster then the repo after refractoring\n",
      "votes": 4
    },
    {
      "id": 3150661,
      "postDate": "2025-03-15T18:06:54.173Z",
      "content": "<p>some lightweight model<br>\nMiniFold: Simple, Fast, and Accurate Protein Structure Prediction<br>\n<a href=\"https://openreview.net/forum?id=1p9hQTbjgo\" target=\"_blank\">https://openreview.net/forum?id=1p9hQTbjgo</a></p>\n<hr>\n<p>NOTE: I think different paper with a name:</p>\n<p><a href=\"https://github.com/hypnopump/MiniFold\" target=\"_blank\">https://github.com/hypnopump/MiniFold</a><br>\nMiniFold: a DeepLearning-based Mini Protein Folding Engine</p>",
      "rawMarkdown": "some lightweight model\nMiniFold: Simple, Fast, and Accurate Protein Structure Prediction\nhttps://openreview.net/forum?id=1p9hQTbjgo\n\n---\n\nNOTE: I think different paper with a name:\n\nhttps://github.com/hypnopump/MiniFold\nMiniFold: a DeepLearning-based Mini Protein Folding Engine",
      "votes": 4
    },
    {
      "id": 3150316,
      "postDate": "2025-03-15T10:46:44.097Z",
      "content": "<p>simple experiment on usalign parameters for casp15</p>\n<p>conclusion : kaggle \"backbone (C1') tm scorre\" is smiliar to those reported on paper for casp \"whole molecule tm score\".</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd474b8adf88db2edefd1dcfa9a822e0d%2FSelection_057.png?generation=1742040536720801&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "simple experiment on usalign parameters for casp15\n\nconclusion : kaggle \"backbone (C1') tm scorre\" is smiliar to those reported on paper for casp \"whole molecule tm score\".\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd474b8adf88db2edefd1dcfa9a822e0d%2FSelection_057.png?generation=1742040536720801&alt=media)",
      "votes": 4
    },
    {
      "id": 3149540,
      "postDate": "2025-03-14T10:56:08.923Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> You mentioned training for around 200 epochs, but I believe you could accelerate your experiments by using the optimizer from <a href=\"https://github.com/facebookresearch/schedule_free\" target=\"_blank\">Facebook Research's Schedule-Free Optimization</a>. In my tests, this optimizer not only works well for image recognition tasks as shown in the paper but also significantly speeds up convergence in fine-tuning DNA language models and training VAEs. It eliminates the need for learning rate scheduling and warm-up. Since it's implemented in PyTorch, you can easily use it by directly incorporating the source code.</p>",
      "rawMarkdown": "@hengck23 You mentioned training for around 200 epochs, but I believe you could accelerate your experiments by using the optimizer from [Facebook Research's Schedule-Free Optimization](https://github.com/facebookresearch/schedule_free). In my tests, this optimizer not only works well for image recognition tasks as shown in the paper but also significantly speeds up convergence in fine-tuning DNA language models and training VAEs. It eliminates the need for learning rate scheduling and warm-up. Since it's implemented in PyTorch, you can easily use it by directly incorporating the source code.",
      "votes": 2,
      "replies": [
        {
          "id": 3150708,
          "postDate": "2025-03-15T19:54:37.097Z",
          "content": "<p>i trained for 400 epochs reached  LB 0.205</p>",
          "rawMarkdown": "i trained for 400 epochs reached  LB 0.205",
          "votes": 1
        }
      ]
    },
    {
      "id": 3182379,
      "postDate": "2025-04-19T07:55:20.957Z",
      "content": "<p>tricks of nufold<br>\ni think the papers are out</p>\n<p>Structure Modeling Protocols for Protein Multimer and RNA in CASP16 with Enhanced MSAs, Model Ranking, and Deep Learning<br>\n<a href=\"https://d197for5662m48.cloudfront.net/documents/publicationstatus/254072/preprint_pdf/f7a0df47ad2e526a2b3baeed49617be0.pdf\" target=\"_blank\">https://d197for5662m48.cloudfront.net/documents/publicationstatus/254072/preprint_pdf/f7a0df47ad2e526a2b3baeed49617be0.pdf</a><br>\n<a href=\"https://advance.sagepub.com/doi/full/10.22541/au.174478900.09405488/v1\" target=\"_blank\">https://advance.sagepub.com/doi/full/10.22541/au.174478900.09405488/v1</a></p>\n<p>\"For RNA, we adopted an ensemble approach that incorporated multiple state-of-the-art methods,<br>\ncentered around our NuFold framework. As a result, our KiharaLab group ranked first in protein complex prediction and third<br>\nin RNA structure prediction\"</p>\n<p>\"To enhance the depth of multiple sequence alignments (MSAs), we employed a large metagenomic sequence database. Model ranking was performed with a stateof-the-art consensus ranking method, to which we added more scoring terms. \"</p>",
      "rawMarkdown": "tricks of nufold\ni think the papers are out\n\nStructure Modeling Protocols for Protein Multimer and RNA in CASP16 with Enhanced MSAs, Model Ranking, and Deep Learning\nhttps://d197for5662m48.cloudfront.net/documents/publicationstatus/254072/preprint_pdf/f7a0df47ad2e526a2b3baeed49617be0.pdf\nhttps://advance.sagepub.com/doi/full/10.22541/au.174478900.09405488/v1\n\n\"For RNA, we adopted an ensemble approach that incorporated multiple state-of-the-art methods,\ncentered around our NuFold framework. As a result, our KiharaLab group ranked first in protein complex prediction and third\nin RNA structure prediction\"\n\n\"To enhance the depth of multiple sequence alignments (MSAs), we employed a large metagenomic sequence database. Model ranking was performed with a stateof-the-art consensus ranking method, to which we added more scoring terms. \"",
      "votes": 1
    },
    {
      "id": 3143672,
      "postDate": "2025-03-07T13:11:23.233Z",
      "content": "<p>Glad to have you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 🤗</p>",
      "rawMarkdown": "Glad to have you @hengck23 🤗",
      "votes": 2
    },
    {
      "id": 3168413,
      "postDate": "2025-04-02T12:02:16.337Z",
      "content": "<p><strong>DRFOLD2 :</strong>    LB : 0.327 (no MSA) ,  0.318 (with MSA from kaggle folder).  I believe I am making a mistake in utilizing MSA -  No energy score used yet</p>",
      "rawMarkdown": "**DRFOLD2 :**    LB : 0.327 (no MSA) ,  0.318 (with MSA from kaggle folder).  I believe I am making a mistake in utilizing MSA -  No energy score used yet",
      "votes": 1,
      "replies": [
        {
          "id": 3168541,
          "postDate": "2025-04-02T14:40:40.997Z",
          "content": "<p>It seems that DRFOLD2 doesn't require a MSA file. Do you modify the model to enable the use of MSA?</p>",
          "rawMarkdown": "It seems that DRFOLD2 doesn't require a MSA file. Do you modify the model to enable the use of MSA?",
          "replies": [
            {
              "id": 3168556,
              "postDate": "2025-04-02T14:58:46.843Z",
              "content": "<p>Rhofold+  <br>\nAF3</p>",
              "rawMarkdown": "Rhofold+  \nAF3",
              "votes": 1
            },
            {
              "id": 3181763,
              "postDate": "2025-04-18T09:50:40.870Z",
              "content": "<p>DRFOLDv1 uses MSA</p>\n<p>DRFOLDv2 doesn't</p>\n<p>part of DRFOLD2 code is modified from DRFOLD1.</p>",
              "rawMarkdown": "DRFOLDv1 uses MSA\n\nDRFOLDv2 doesn't\n\npart of DRFOLD2 code is modified from DRFOLD1."
            }
          ]
        }
      ]
    },
    {
      "id": 3167131,
      "postDate": "2025-04-01T09:46:05.207Z",
      "content": "<p>Nufold casp16<br>\n<a href=\"https://x.com/kiharalab/status/1906823369372434641?s=46\" target=\"_blank\">https://x.com/kiharalab/status/1906823369372434641?s=46</a></p>",
      "rawMarkdown": "Nufold casp16\nhttps://x.com/kiharalab/status/1906823369372434641?s=46",
      "votes": 1,
      "replies": [
        {
          "id": 3168237,
          "postDate": "2025-04-02T08:58:45.703Z",
          "content": "<p>Hi, hengck23. extra rMSA needs 2TB, and how can we solve it in kaggle env?</p>",
          "rawMarkdown": "Hi, hengck23. extra rMSA needs 2TB, and how can we solve it in kaggle env?",
          "replies": [
            {
              "id": 3168239,
              "postDate": "2025-04-02T09:04:31.397Z",
              "content": "<p>I thought rmsa already given by host?</p>",
              "rawMarkdown": "I thought rmsa already given by host?",
              "votes": 1
            },
            {
              "id": 3168246,
              "postDate": "2025-04-02T09:09:25.697Z",
              "content": "<p>oh, I'm stupid.😂</p>",
              "rawMarkdown": "oh, I'm stupid.😂"
            }
          ]
        }
      ]
    },
    {
      "id": 3161488,
      "postDate": "2025-03-28T02:41:26Z",
      "content": "<p>managed to repeat CASP15 paper results for proteinX</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31cb8c0a1349eba2a8e4d80ec828ce1e%2FSelection_168.png?generation=1743254577598872&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "managed to repeat CASP15 paper results for proteinX\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31cb8c0a1349eba2a8e4d80ec828ce1e%2FSelection_168.png?generation=1743254577598872&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 3161562,
          "postDate": "2025-03-28T05:14:36.237Z",
          "content": "<p>looks good results! did you use MSA?</p>",
          "rawMarkdown": "looks good results! did you use MSA?",
          "replies": [
            {
              "id": 3161699,
              "postDate": "2025-03-28T09:32:59.350Z",
              "content": "<p>proteinx disabled msa for rna. we are trying to patch that now.</p>",
              "rawMarkdown": "proteinx disabled msa for rna. we are trying to patch that now."
            },
            {
              "id": 3162615,
              "postDate": "2025-03-29T13:24:18.863Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2695f9b6709b3893d2e98b079eee4d42%2FSelection_169.png?generation=1743254603712407&amp;alt=media\" alt=\"\"></p>\n<p>chai-1 results (also no MSA)<br>\ncomparable to protein-X.</p>\n<p>Obviously, the combination of the two is better (see r1108,r1128)</p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2695f9b6709b3893d2e98b079eee4d42%2FSelection_169.png?generation=1743254603712407&alt=media)\n\nchai-1 results (also no MSA)\ncomparable to protein-X.\n\nObviously, the combination of the two is better (see r1108,r1128)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3161351,
      "postDate": "2025-03-27T19:40:03.973Z",
      "content": "<p>there is an assembly method in the public notebook that is coring quite well.</p>\n<p><a href=\"https://www.biorxiv.org/content/10.1101/2023.12.22.573067v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2023.12.22.573067v1.full.pdf</a><br>\n\"Vfold3D (41) constructs 3D structures from fragment databases. It uses the lowest free energy secondary structures converted to known fragments. The reconstruction of fragments is coarse-grained before being converted to allatom.\"</p>\n<p>a smiliar method is <a href=\"https://github.com/Tan-group/FebRNA\" target=\"_blank\">https://github.com/Tan-group/FebRNA</a>, this is quite fast</p>",
      "rawMarkdown": "there is an assembly method in the public notebook that is coring quite well.\n\nhttps://www.biorxiv.org/content/10.1101/2023.12.22.573067v1.full.pdf\n\"Vfold3D (41) constructs 3D structures from fragment databases. It uses the lowest free energy secondary structures converted to known fragments. The reconstruction of fragments is coarse-grained before being converted to allatom.\"\n\n a smiliar method is https://github.com/Tan-group/FebRNA, this is quite fast",
      "votes": 1
    },
    {
      "id": 3159835,
      "postDate": "2025-03-26T01:13:10.753Z",
      "content": "<p>how to compare msa search<br>\n<a href=\"https://academic.oup.com/gpb/article/22/1/qzae018/7617691?login=false\" target=\"_blank\">https://academic.oup.com/gpb/article/22/1/qzae018/7617691?login=false</a></p>\n<p>Performance comparison on RNA homology search<br>\nTable 1 compares the MSAs generated by RNAcmap2, rMSA, and RNAcmap3 in terms of median Neff and average F1-score given by mfDCA for the MSAs. The distribution of F1-scores for individual RNAs is shown in Figure 2.</p>",
      "rawMarkdown": "how to compare msa search\nhttps://academic.oup.com/gpb/article/22/1/qzae018/7617691?login=false\n\nPerformance comparison on RNA homology search\nTable 1 compares the MSAs generated by RNAcmap2, rMSA, and RNAcmap3 in terms of median Neff and average F1-score given by mfDCA for the MSAs. The distribution of F1-scores for individual RNAs is shown in Figure 2.\n",
      "votes": 1
    },
    {
      "id": 3157259,
      "postDate": "2025-03-23T07:35:14.400Z",
      "content": "<p>team plan (see attachment)</p>",
      "rawMarkdown": "team plan (see attachment)",
      "votes": 1
    },
    {
      "id": 3181587,
      "postDate": "2025-04-18T04:50:53.087Z",
      "content": "<p>new paper!<br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2025.04.10.648174v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.04.10.648174v1</a><br>\nAlphaFold3 at CASP16</p>\n<p>All analysis scripts are freely available from<br>\n<a href=\"https://gitlab.com/arneelof/CASP16-predictions\" target=\"_blank\">https://gitlab.com/arneelof/CASP16-predictions</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa6b9ad0c6fd8d3e425d0bdb725e3929%2FSelection_999(8109).png?generation=1744952168393928&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa2d39c07e6b27738dceafa0b92cdd8b7%2FSelection_999(8108).png?generation=1744952036291418&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "new paper!\nhttps://www.biorxiv.org/content/10.1101/2025.04.10.648174v1\nAlphaFold3 at CASP16\n\nAll analysis scripts are freely available from\nhttps://gitlab.com/arneelof/CASP16-predictions\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa6b9ad0c6fd8d3e425d0bdb725e3929%2FSelection_999(8109).png?generation=1744952168393928&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa2d39c07e6b27738dceafa0b92cdd8b7%2FSelection_999(8108).png?generation=1744952036291418&alt=media)",
      "votes": 2
    },
    {
      "id": 3154706,
      "postDate": "2025-03-20T10:04:04.750Z",
      "content": "<p>dataset<br>\nDeep learning for RNA structure prediction <br>\n<a href=\"https://www.sciencedirect.com/science/article/pii/S0959440X25000090\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S0959440X25000090</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65b83d3f7c1ccf6d94006a9b688834d9%2FSelection_097.png?generation=1742465027201044&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F57cda013db8eb5944153b6655cbcc9c1%2FSelection_098.png?generation=1742465245565095&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "dataset\nDeep learning for RNA structure prediction \nhttps://www.sciencedirect.com/science/article/pii/S0959440X25000090\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65b83d3f7c1ccf6d94006a9b688834d9%2FSelection_097.png?generation=1742465027201044&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F57cda013db8eb5944153b6655cbcc9c1%2FSelection_098.png?generation=1742465245565095&alt=media)",
      "votes": 1
    },
    {
      "id": 3154089,
      "postDate": "2025-03-19T14:00:09.040Z",
      "content": "<p><a href=\"https://elanapearl.github.io/blog/2024/the-illustrated-alphafold/\" target=\"_blank\">https://elanapearl.github.io/blog/2024/the-illustrated-alphafold/</a><br>\ni am study af3. thi is useful</p>",
      "rawMarkdown": "https://elanapearl.github.io/blog/2024/the-illustrated-alphafold/\ni am study af3. thi is useful",
      "votes": 1
    },
    {
      "id": 3153768,
      "postDate": "2025-03-19T06:24:51.973Z",
      "content": "<p>MSA or not MSA, here is the difference!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1f7bfa214db89e1570d9c4a90ab3ccd2%2FSelection_087.png?generation=1742365489128348&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "MSA or not MSA, here is the difference!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1f7bfa214db89e1570d9c4a90ab3ccd2%2FSelection_087.png?generation=1742365489128348&alt=media)",
      "votes": 1
    },
    {
      "id": 3152205,
      "postDate": "2025-03-17T15:07:06.860Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc978760a9f37d267c34a0deba1027562%2FSelection_075.png?generation=1742224007338988&amp;alt=media\" alt=\"\"></p>\n<p>rna  foundation model</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc978760a9f37d267c34a0deba1027562%2FSelection_075.png?generation=1742224007338988&alt=media)\n\n rna  foundation model",
      "votes": 2
    },
    {
      "id": 3148692,
      "postDate": "2025-03-13T12:38:36.173Z",
      "content": "<p>😭 day two of my download<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3911f386181723cf3ef748bc856ea250%2FSelection_046.png?generation=1741869514167103&amp;alt=media\" alt=\"\"></p>\n<p>the only competition where external data is in TB … i estimate i will burn away 6 to 8 TB</p>",
      "rawMarkdown": "😭 day two of my download\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3911f386181723cf3ef748bc856ea250%2FSelection_046.png?generation=1741869514167103&alt=media)\n\nthe only competition where external data is in TB ... i estimate i will burn away 6 to 8 TB",
      "votes": 2
    },
    {
      "id": 3147849,
      "postDate": "2025-03-12T13:30:26.177Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 想问一下您的base中ribonanza_net是一个什么样的东西呢</p>",
      "rawMarkdown": "@hengck23 想问一下您的base中ribonanza_net是一个什么样的东西呢",
      "votes": -4
    },
    {
      "id": 3181877,
      "postDate": "2025-04-18T13:18:58.697Z",
      "content": "<p>it is here!!!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Feb791133a41c2501370ec4e4fc7fc186%2FSelection_999(8114).png?generation=1744982252631348&amp;alt=media\" alt=\"\"></p>\n<p>trRosettaRNA2 is a new deep learning model for RNA 3D structure prediction that incorporates a novel secondary structure prior module trained on bpRNA data. … aka yang-server</p>\n<p>do read the supplementary paper as well<br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2025.04.09.647915v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.04.09.647915v1</a></p>",
      "rawMarkdown": "it is here!!!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Feb791133a41c2501370ec4e4fc7fc186%2FSelection_999(8114).png?generation=1744982252631348&alt=media)\n\ntrRosettaRNA2 is a new deep learning model for RNA 3D structure prediction that incorporates a novel secondary structure prior module trained on bpRNA data. ... aka yang-server\n\ndo read the supplementary paper as well\nhttps://www.biorxiv.org/content/10.1101/2025.04.09.647915v1",
      "replies": [
        {
          "id": 3182454,
          "postDate": "2025-04-19T10:55:03.410Z",
          "content": "<p>It seems trRosettaRNA2 takes a long time to predict the 3d coor, how do you solve such problem?</p>",
          "rawMarkdown": "It seems trRosettaRNA2 takes a long time to predict the 3d coor, how do you solve such problem?",
          "replies": [
            {
              "id": 3182455,
              "postDate": "2025-04-19T11:00:56.337Z",
              "content": "<p>i the trRosettaRNA2 is smiliar speed as alphafold3. trRosettaRNA version 1 is the slow version.<br>\nyou can check their presentation at youtube, etc</p>",
              "rawMarkdown": "i the trRosettaRNA2 is smiliar speed as alphafold3. trRosettaRNA version 1 is the slow version.\nyou can check their presentation at youtube, etc"
            },
            {
              "id": 3182458,
              "postDate": "2025-04-19T11:05:23.343Z",
              "content": "<p>You mean trRosettaRNA2, is the trRosettaRNA_v1.1 version as proposed in their web?</p>",
              "rawMarkdown": "You mean trRosettaRNA2, is the trRosettaRNA_v1.1 version as proposed in their web?"
            },
            {
              "id": 3182472,
              "postDate": "2025-04-19T11:37:48.683Z",
              "content": "<p><a href=\"https://www.youtube.com/watch?v=yGO_G3stTBk&amp;t=157s\" target=\"_blank\">https://www.youtube.com/watch?v=yGO_G3stTBk&amp;t=157s</a><br>\n<a href=\"https://www.youtube.com/watch?v=LTp21NolEak\" target=\"_blank\">https://www.youtube.com/watch?v=LTp21NolEak</a></p>",
              "rawMarkdown": "https://www.youtube.com/watch?v=yGO_G3stTBk&t=157s\nhttps://www.youtube.com/watch?v=LTp21NolEak"
            },
            {
              "id": 3182486,
              "postDate": "2025-04-19T12:08:03.740Z",
              "content": "<p>The version I used in my experiment, trRosettaRNA_v1.1, is different from trRosettaRNA2. It runs significantly slow and often triggers a timeout error when LB scoring, maybe when predicting RNAs longer than about 200 nucleotides.</p>",
              "rawMarkdown": "The version I used in my experiment, trRosettaRNA_v1.1, is different from trRosettaRNA2. It runs significantly slow and often triggers a timeout error when LB scoring, maybe when predicting RNAs longer than about 200 nucleotides."
            }
          ]
        }
      ]
    },
    {
      "id": 3158168,
      "postDate": "2025-03-24T09:25:12.187Z",
      "content": "<p>how long on average do you run these model ?</p>",
      "rawMarkdown": "how long on average do you run these model ?"
    },
    {
      "id": 3156196,
      "postDate": "2025-03-21T21:13:21.423Z",
      "content": "<p>you are an og <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>",
      "rawMarkdown": "you are an og @hengck23"
    },
    {
      "id": 3155623,
      "postDate": "2025-03-21T08:13:49.067Z",
      "content": "<p><a href=\"https://ar5iv.labs.arxiv.org/html/2203.00854\" target=\"_blank\">https://ar5iv.labs.arxiv.org/html/2203.00854</a><br>\nfastfold</p>",
      "rawMarkdown": "https://ar5iv.labs.arxiv.org/html/2203.00854\nfastfold"
    },
    {
      "id": 3155126,
      "postDate": "2025-03-20T18:25:49.093Z",
      "content": "<p>probe of my current drfoldv2 results:</p>\n<ul>\n<li>current public notebook using 5 models is 0.321</li>\n<li>if we improve rna of length&lt;200 only, we can get about 0.334</li>\n<li>if we improve rna of 200&lt;length&lt;400 only, we can get about 0.405</li>\n</ul>\n<p>if i use energy selection for 10 models, submission shows 0.321 --&gt; 0.323<br>\nengergy compute in drfold2 is very slow , so in this submissopm, i only do selection for those less than length 200.</p>\n<p>but even if i can speed that up (just replace the cubic spline interpolation) and do selection up to 400 length, i can only end up 0.38 to 0.40 i think.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2d0d54914e9681244fe3ce6f840427e8%2FSelection_100.png?generation=1742495142254548&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "probe of my current drfoldv2 results:\n- current public notebook using 5 models is 0.321\n- if we improve rna of length<200 only, we can get about 0.334\n- if we improve rna of 200<length<400 only, we can get about 0.405\n\nif i use energy selection for 10 models, submission shows 0.321 --> 0.323\nengergy compute in drfold2 is very slow , so in this submissopm, i only do selection for those less than length 200.\n\nbut even if i can speed that up (just replace the cubic spline interpolation) and do selection up to 400 length, i can only end up 0.38 to 0.40 i think.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2d0d54914e9681244fe3ce6f840427e8%2FSelection_100.png?generation=1742495142254548&alt=media)",
      "replies": [
        {
          "id": 3166956,
          "postDate": "2025-04-01T07:07:05.393Z",
          "content": "<p>what do you mean saying if we improve rna of 200&lt;length&lt;400 only?</p>",
          "rawMarkdown": "what do you mean saying if we improve rna of 200<length<400 only?"
        },
        {
          "id": 3178336,
          "postDate": "2025-04-14T03:41:38.497Z",
          "content": "<p>Hi, can you share a hint on improving the seq of length 200&lt;length&lt;400 using drfoldv2?</p>",
          "rawMarkdown": "Hi, can you share a hint on improving the seq of length 200<length<400 using drfoldv2?"
        }
      ]
    },
    {
      "id": 3155119,
      "postDate": "2025-03-20T18:21:37.803Z",
      "content": "<p>CASP-16 benchmark (results pulled from casp16 server):</p>\n<p>rna of length less than 400:</p>\n<pre><code>: .\n: .\n: .\n+drfold2: .\n: .\n</code></pre>\n<p>all rna</p>\n<pre><code>: .\n: .\n: .\n+drfold2: .\n: .\n</code></pre>\n<p>the strategy is very clear. either alphafold3 or drfoldv2 alone cannot win vfold. we need combinations, the more the merrier.<br>\ndue to hardware resources of kaggle notebook, we need to distill top models into single one</p>\n<ol>\n<li>there are 30 millions rna seq in RNACetral, etc … but they don't have 3d structure</li>\n<li>most model like drfoldv2 are trained with about 18k rna (that is all 3d structure you can find in PDB)</li>\n<li>one solution is to use top models to puesdo label all rna in RNACetral, etc</li>\n<li>then distill to one that can run about rna length=1024 in kaggle notebook (we need to design this)</li>\n</ol>\n<p>tmr i submit alphafold3,drfold2,vfold (offline compute), etc to kaggle server.</p>\n<p>I think i can design a good network to run on kaggle notebook. I would probably release an untrained version of it in a couple of days.</p>\n<p>now i need resource to train it.</p>\n<hr>\n<p>on a side note:<br>\nyou can connect drfoldv2 RNACLM (already pretrained with evoluitionary feature) with RibonanzaNet.<br>\nRibonanzaNet in host example currently does not use triangular attention (that code is commented out). <br>\nbut if you use triangular attention, there may be resource OOM problems (same for other opensource model).</p>\n<p>liteformer claim to reduce triangular attention complexity from O(L^3) to O(L^2) without loss in performance on CASP15. this is what i am trying to do.</p>",
      "rawMarkdown": "CASP-16 benchmark (results pulled from casp16 server):\n\nrna of length less than 400:\n```\nalphafold3: 0.5271428571428571\ndrfold2: 0.5173809523809524\nvfold: 0.5907619047619047\nalphafold3+drfold2: 0.5787142857142857\nALL: 0.6080476190476191\n\n```\nall rna\n```\nalphafold3: 0.4623720930232558\ndrfold2: 0.46891891891891885\nvfold: 0.5146511627906977\nalphafold3+drfold2: 0.5038139534883721\nALL: 0.5281627906976745\n\n```\n\n\nthe strategy is very clear. either alphafold3 or drfoldv2 alone cannot win vfold. we need combinations, the more the merrier.\ndue to hardware resources of kaggle notebook, we need to distill top models into single one\n1. there are 30 millions rna seq in RNACetral, etc ... but they don't have 3d structure\n2. most model like drfoldv2 are trained with about 18k rna (that is all 3d structure you can find in PDB)\n3. one solution is to use top models to puesdo label all rna in RNACetral, etc\n4. then distill to one that can run about rna length=1024 in kaggle notebook (we need to design this)\n\ntmr i submit alphafold3,drfold2,vfold (offline compute), etc to kaggle server.\n\nI think i can design a good network to run on kaggle notebook. I would probably release an untrained version of it in a couple of days.\n\nnow i need resource to train it.\n\n---\n\non a side note:\nyou can connect drfoldv2 RNACLM (already pretrained with evoluitionary feature) with RibonanzaNet.\nRibonanzaNet in host example currently does not use triangular attention (that code is commented out). \nbut if you use triangular attention, there may be resource OOM problems (same for other opensource model).\n\nliteformer claim to reduce triangular attention complexity from O(L^3) to O(L^2) without loss in performance on CASP15. this is what i am trying to do.",
      "replies": [
        {
          "id": 3155129,
          "postDate": "2025-03-20T18:29:53.537Z",
          "content": "<p>Really cool results! One quick thing to note -- we are not licensed to use alphafold3 models for distillation. 🫥. Other methods should be OK. </p>",
          "rawMarkdown": "Really cool results! One quick thing to note -- we are not licensed to use alphafold3 models for distillation. 🫥. Other methods should be OK. ",
          "replies": [
            {
              "id": 3155135,
              "postDate": "2025-03-20T18:32:32.140Z",
              "content": "<p>thanks. noted. we should be using alphafold3 clones</p>",
              "rawMarkdown": "thanks. noted. we should be using alphafold3 clones"
            }
          ]
        },
        {
          "id": 3155130,
          "postDate": "2025-03-20T18:30:21.557Z",
          "content": "<p>For distillation, would it help to provide chemical mapping profiles (DMS and 2A3) for the Ribonanza data set of 2.1M RNA's? They may enable some filtering of the best 3D models and a cleaner distillation set.</p>",
          "rawMarkdown": "For distillation, would it help to provide chemical mapping profiles (DMS and 2A3) for the Ribonanza data set of 2.1M RNA's? They may enable some filtering of the best 3D models and a cleaner distillation set.",
          "replies": [
            {
              "id": 3155138,
              "postDate": "2025-03-20T18:35:35.357Z",
              "content": "<p>yes, i agree on that. i am reading some papers on input for rna 3d structure prediction.<br>\ni can think that the labelled rna 3d structure could be quite noisy, so we need addition input or aux output.</p>\n<p>now studying how alphafold2 actually do that for protein. i will do some experiments and feedback later.<br>\nThnaks for the comment</p>",
              "rawMarkdown": "yes, i agree on that. i am reading some papers on input for rna 3d structure prediction.\ni can think that the labelled rna 3d structure could be quite noisy, so we need addition input or aux output.\n\nnow studying how alphafold2 actually do that for protein. i will do some experiments and feedback later.\nThnaks for the comment"
            }
          ]
        },
        {
          "id": 3157133,
          "postDate": "2025-03-23T02:53:47.943Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Sorry if I missed it somewhere, but what does \"ALL\" mean?  (Among alphafold3, drfold2, etc.)</p>",
          "rawMarkdown": "@hengck23 Sorry if I missed it somewhere, but what does \"ALL\" mean?  (Among alphafold3, drfold2, etc.)"
        }
      ]
    },
    {
      "id": 3153414,
      "postDate": "2025-03-18T19:23:27.983Z",
      "content": "<p>drfold2 memory and speed is restricted by the col and row triangle attention in the evo-transformer.<br>\nignore the RNA language model, the structure model uses 10,432 train rna seq given in repo</p>\n<p>now I can speed by 2x and reduce memory by 2x by using new structure<br>\ni am wondering if I want to retrain it ….</p>\n<p><a href=\"https://github.com/Dao-AILab/flash-attention/pull/57\" target=\"_blank\">https://github.com/Dao-AILab/flash-attention/pull/57</a><br>\n<a href=\"https://openreview.net/pdf?id=brk7Ct4Tb1M\" target=\"_blank\">https://openreview.net/pdf?id=brk7Ct4Tb1M</a><br>\nliteformer: <a href=\"https://openreview.net/forum?id=t0m0DdCCQ2\" target=\"_blank\">https://openreview.net/forum?id=t0m0DdCCQ2</a><br>\n(kaggle has only 844 train rna seq)</p>",
      "rawMarkdown": "drfold2 memory and speed is restricted by the col and row triangle attention in the evo-transformer.\nignore the RNA language model, the structure model uses 10,432 train rna seq given in repo\n\nnow I can speed by 2x and reduce memory by 2x by using new structure\ni am wondering if I want to retrain it ....\n\nhttps://github.com/Dao-AILab/flash-attention/pull/57\nhttps://openreview.net/pdf?id=brk7Ct4Tb1M\nliteformer: https://openreview.net/forum?id=t0m0DdCCQ2\n(kaggle has only 844 train rna seq)"
    },
    {
      "id": 3152644,
      "postDate": "2025-03-18T03:03:10.733Z",
      "content": "<p>Do we need gu Mas search? <br>\n<a href=\"https://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/\" target=\"_blank\">https://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/</a></p>",
      "rawMarkdown": "Do we need gu Mas search? \nhttps://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/"
    },
    {
      "id": 3152490,
      "postDate": "2025-03-17T21:32:58.733Z",
      "content": "<p>hi I'm new and I feel like this competition is not newbie friendly and computational costly right ? should I find other competition that suit me more or try my best here to learn? cause after seeing your result I can see the amount of resource going in training these models are not going to be cheap  </p>",
      "rawMarkdown": "hi I'm new and I feel like this competition is not newbie friendly and computational costly right ? should I find other competition that suit me more or try my best here to learn? cause after seeing your result I can see the amount of resource going in training these models are not going to be cheap  ",
      "replies": [
        {
          "id": 3152967,
          "postDate": "2025-03-18T10:04:55.253Z",
          "content": "<p>U still have a chance, just like deepseek can do without expensive gpu in training.</p>\n<p>U can train the llm approach. Choose a good rna foundation model. Train Lora or other adapter for 3d structure prediction </p>",
          "rawMarkdown": "U still have a chance, just like deepseek can do without expensive gpu in training.\n\nU can train the llm approach. Choose a good rna foundation model. Train Lora or other adapter for 3d structure prediction ",
          "votes": 1,
          "replies": [
            {
              "id": 3153015,
              "postDate": "2025-03-18T11:05:57.267Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> when you go to different database to collect data you collect the whole nucleotide coordinate or just the C1' coordinate like what the competition did ?</p>",
              "rawMarkdown": "@hengck23 when you go to different database to collect data you collect the whole nucleotide coordinate or just the C1' coordinate like what the competition did ?"
            }
          ]
        }
      ]
    },
    {
      "id": 3151459,
      "postDate": "2025-03-16T17:55:23.037Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> LB 500 based on which model?</p>",
      "rawMarkdown": "@hengck23 LB 500 based on which model?",
      "replies": [
        {
          "id": 3151465,
          "postDate": "2025-03-16T18:04:05.883Z",
          "content": "<p>please wait, still open sourcing it</p>",
          "rawMarkdown": "please wait, still open sourcing it",
          "votes": 2
        }
      ]
    },
    {
      "id": 3148334,
      "postDate": "2025-03-13T02:41:49.710Z",
      "content": "<p>Do you plan to compete for the prestigious \"Early Sharing\" award?</p>",
      "rawMarkdown": "Do you plan to compete for the prestigious \"Early Sharing\" award?",
      "replies": [
        {
          "id": 3148339,
          "postDate": "2025-03-13T02:51:14.290Z",
          "content": "<p>i plan to get a valid solution first and then decide.<br>\nI don't think just submitting openousrce model can get a prize</p>",
          "rawMarkdown": "i plan to get a valid solution first and then decide.\nI don't think just submitting openousrce model can get a prize\n",
          "votes": 2,
          "replies": [
            {
              "id": 3148369,
              "postDate": "2025-03-13T03:35:59.510Z",
              "content": "<p>But based on official rules you had already above the threshold of eligibility to receive the prize, its only a matter of whether you want to compete for it or not.</p>",
              "rawMarkdown": "But based on official rules you had already above the threshold of eligibility to receive the prize, its only a matter of whether you want to compete for it or not."
            },
            {
              "id": 3148379,
              "postDate": "2025-03-13T03:53:58.920Z",
              "content": "<p>It's written a bit ambiguously. One could understand that the threshold needs to be exceeded in two months—when the test sample update occurs.</p>",
              "rawMarkdown": "It's written a bit ambiguously. One could understand that the threshold needs to be exceeded in two months—when the test sample update occurs.",
              "votes": 1
            },
            {
              "id": 3148388,
              "postDate": "2025-03-13T03:59:40.390Z",
              "content": "<p>Not eligible yet. Need to open source code. Also need to check, eg license, external data cut off date, etc</p>",
              "rawMarkdown": "Not eligible yet. Need to open source code. Also need to check, eg license, external data cut off date, etc",
              "votes": 3
            },
            {
              "id": 3148883,
              "postDate": "2025-03-13T15:55:37.757Z",
              "content": "<p>confirming: the solution must indeed outperform <code>vfold_human_expert</code> <em>after</em> the data refresh in 2 months.</p>",
              "rawMarkdown": "confirming: the solution must indeed outperform `vfold_human_expert` *after* the data refresh in 2 months.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3147476,
      "postDate": "2025-03-12T04:09:52.553Z",
      "content": "<p>are you using TM score? if so, how do you handle the sequences have less 41 native_cnt </p>",
      "rawMarkdown": "are you using TM score? if so, how do you handle the sequences have less 41 native_cnt ",
      "replies": [
        {
          "id": 3147481,
          "postDate": "2025-03-12T04:16:17.687Z",
          "content": "<p>it actuually doesn't matter: lb = choose best tm score for tm(j-th predict, i-th truth) where i,j can be any range</p>",
          "rawMarkdown": "it actuually doesn't matter: lb = choose best tm score for tm(j-th predict, i-th truth) where i,j can be any range",
          "votes": 1
        }
      ]
    },
    {
      "id": 3147347,
      "postDate": "2025-03-12T00:34:49.420Z",
      "content": "<p>Oops is there a data leak?</p>",
      "rawMarkdown": "Oops is there a data leak?",
      "replies": [
        {
          "id": 3147364,
          "postDate": "2025-03-12T00:53:23.440Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/data\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/data</a><br>\nhost leaked the public dataset:<br>\n\"a. During this phase the public test set sequences includes–but is not limited to–targets from the 2024 CASP16 competition whose structures have not yet been publicly released in the PDB database.\"</p>\n<p>a simple solution is to search models that score well on casp15 and casp16 (papers, github, etc).<br>\nthen distill their results.</p>\n<p>but public lb is not the final target. we are interested in bind test in the private set!</p>",
          "rawMarkdown": "https://www.kaggle.com/competitions/stanford-rna-3d-folding/data\nhost leaked the public dataset:\n\"a. During this phase the public test set sequences includes–but is not limited to–targets from the 2024 CASP16 competition whose structures have not yet been publicly released in the PDB database.\"\n\na simple solution is to search models that score well on casp15 and casp16 (papers, github, etc).\nthen distill their results.\n\nbut public lb is not the final target. we are interested in bind test in the private set!",
          "votes": 8,
          "replies": [
            {
              "id": 3150591,
              "postDate": "2025-03-15T16:32:03.430Z",
              "content": "<p>Hmm… then, if the public test set were the final private test set, we could also expect that the winner's TM-score would be around 0.500, right?</p>",
              "rawMarkdown": "Hmm... then, if the public test set were the final private test set, we could also expect that the winner's TM-score would be around 0.500, right?"
            },
            {
              "id": 3207663,
              "postDate": "2025-05-23T05:33:33.060Z",
              "content": "<p>same question</p>",
              "rawMarkdown": "same question",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3147346,
      "postDate": "2025-03-12T00:34:42.567Z",
      "content": "<blockquote>\n  <p>Hardware<br>\n  GPU: 2x Nvidia Ada A6000 (Ampere), each with VRAM 48 GB<br>\n  CPU: Intel® Xeon(R) w7-3455 CPU @ 2.5GHz, 24 cores, 48 threads<br>\n  Memory: 256 GB RAM</p>\n</blockquote>\n<p>Many well-known sotas for 3D RNA structure prediction require substantial memories and datasets. When I see you topping the leaderboard, I wonder if you had employed a similar approach. </p>\n<p>If you managed to achieve this solely by modifying the baseline, that would be truly incredible, I can't believe someone can make it only in this way.</p>\n<p>BTW, glad to see you too.</p>",
      "rawMarkdown": ">Hardware\nGPU: 2x Nvidia Ada A6000 (Ampere), each with VRAM 48 GB\nCPU: Intel® Xeon(R) w7-3455 CPU @ 2.5GHz, 24 cores, 48 threads\nMemory: 256 GB RAM\n\nMany well-known sotas for 3D RNA structure prediction require substantial memories and datasets. When I see you topping the leaderboard, I wonder if you had employed a similar approach. \n\nIf you managed to achieve this solely by modifying the baseline, that would be truly incredible, I can't believe someone can make it only in this way.\n\nBTW, glad to see you too.",
      "replies": [
        {
          "id": 3147369,
          "postDate": "2025-03-12T01:13:52.047Z",
          "content": "<p>i will update on the solution later. now i am downloading the MSA database … it takes a long time !!!!</p>",
          "rawMarkdown": "i will update on the solution later. now i am downloading the MSA database ... it takes a long time !!!!",
          "votes": 4,
          "replies": [
            {
              "id": 3147378,
              "postDate": "2025-03-12T01:25:26.723Z",
              "content": "<p>Until now, I never use any extenal dataset. I'm really looking forward to your method.</p>",
              "rawMarkdown": "Until now, I never use any extenal dataset. I'm really looking forward to your method.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3207662,
      "postDate": "2025-05-23T05:32:33.563Z",
      "content": "<p>Thanks for suggestions!</p>",
      "rawMarkdown": "Thanks for suggestions!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3147190,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-11T18:42:17.233000",
      "content": "<p><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a><br>\nA trainable PyTorch reproduction of AlphaFold 3.</p>\n<p>curious how the diffuser work …</p>",
      "votes": 12,
      "replies": [
        {
          "id": 3151989,
          "author_name": "doheon114",
          "author_url": "",
          "post_date": "2025-03-17T09:55:22.117000",
          "content": "<p>Have you ever tried using this? I'm about to implement it on a Kaggle notebook, but even running inference seems challenging.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3151998,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-17T10:12:26.730000",
              "content": "<p>you need to simplify!!!</p>\n<p>i already about to make drfold2 work. i need a complementary model (for assembly)and AF3 is the recommended one.<br>\ni will try AAF3 from next week.</p>\n<p>drfold2 to be out in public notebook soon</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3152015,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-17T10:33:07.597000",
              "content": "<p><a href=\"https://www.kaggle.com/doheon114\" target=\"_blank\">@doheon114</a> <br>\n<a href=\"https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/kaggle-casp15-0-43-simple-drfold-no-msa</a></p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3155466,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-21T04:16:59.447000",
      "content": "<p>the most important benchmark, must read!!!!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffba9ee0c870e82e97dbb5c62f68655e4%2FSelection_116.png?generation=1742659558654264&amp;alt=media\" alt=\"\"></p>\n<p>notebook: <a href=\"https://www.kaggle.com/code/hengck23/vfold-baseline-offline\" target=\"_blank\">https://www.kaggle.com/code/hengck23/vfold-baseline-offline</a></p>\n<p>It's frustrating that my big local improvements on DRfold2 didn’t show up as gains on the leaderboard.<br>\nso I suspect there is some RNA that is usually difficult for drfold ONLY.<br>\ni think I am correct</p>\n<p>drfoldv2 totally breakdown at rna &gt;400<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F84b48cea4acbfaf14c3bd590d016128a%2FSelection_105.png?generation=1742534073508454&amp;alt=media\" alt=\"\"></p>",
      "votes": 10,
      "replies": [
        {
          "id": 3155521,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2025-03-21T05:32:28.690000",
          "content": "<p>Also my bottleneck… </p>\n<p>Besides, I tried different sotas, and RNAs with long sequences can lead to OOM, because the Kaggle env is limited. </p>\n<p>Can we beat the vfold baseline in the Kaggle env? I have my doubts. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3155584,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-21T07:24:11.783000",
              "content": "<p>\"Can we beat the vfold baseline in the Kaggle env?\" <br>\nyes I think we can …. but much efforts</p>\n<p>I will release a model that can run 800 rna length on kaggler notebook soon. need to patch triangular attnetion</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3155582,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-21T07:22:45.413000",
          "content": "<p>casp16 rna target list from casp website</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3154714,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-20T10:12:22.017000",
      "content": "<p>vfold is the benchmark and we should understand it:<br>\n<a href=\"https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf\" target=\"_blank\">https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf</a><br>\nhint: go to casp16 website to be vfold tm score and make a rank chart</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8ff6ed5e9aef06780c5d9e6f545121bf%2FSelection_099.png?generation=1742465512170143&amp;alt=media\" alt=\"\"></p>\n<p>U must be very good at short length</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 3148681,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-13T12:21:59.103000",
      "content": "<p>I shared a technique for experimenting with MSA on the \"open hidden\" test set for the public leaderboard.</p>\n<p><strong>Please note that this is for testing purposes only and should not be considered a legitimate solution for claiming any prizes.</strong></p>\n<ol>\n<li>go to casp16 and find the sequence fasta of rna targets</li>\n<li>you can now predict their 3d structure offline (e.g. using web mas server)</li>\n<li>the host said that some public test rna are indeed casp16 target, so in your submission code use:</li>\n</ol>\n<pre><code>OFFLINE_PREDICT={\n    ...' :  xyz  ...,\n    ...' :  xyz  ...,\n}\n\n\n i,row  test_df ...:\n     row.sequence  found  OFFLINE_PREDICT :\n             = predicted xyz \n      found:\n           xyz =  online model\n</code></pre>\n<p>in my experiments, I thin I can conclude:<br>\nif you do very well in casp15, casp16, rna puzzles you should do well in this competition in public and private hidden dataset for the three different phrases. </p>\n<p>do well doesn't mean train on these dataset but your model can generalise on these datasets.<br>\ni think the level of difficulty of this competition is similar to these datasets.</p>\n<hr>\n<p>what you can do next is to measure the alignment score (% of matched seq) on these dataset to kaggle train (or msa) and kaggle 400k diffusion synthetic data. This will tell you if these are inlier or outlier data.</p>\n<p>ML methdos tends to work (only?) on inliers data. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 3148704,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2025-03-13T13:02:32.787000",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>If I'm not mistaken, you're thinking of using data similar to the test dataset (CASP15, 16) to create pseudo-labels, right?</p>\n<blockquote>\n  <p>you can now predict their 3d structure offline (e.g. using web mas server)</p>\n</blockquote>\n<p>Why offline  (e.g. using web mas server)?  submit fasta RNA sequences then wait for the consequence, or run the sota locally to create pseudo-labels?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3148713,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-13T13:09:49.927000",
              "content": "<p>i give you an example.</p>\n<ul>\n<li>say i want to inlcude msa as input to ribonanaza-net.</li>\n<li>say in local experiments with kaggle 800 rna with msa train/validation split, it work very well.</li>\n<li>it is not possible to setup msa server in kaggle notebook (for now … but maybe possible after a month) so i cannot know msa lb score.</li>\n<li>but i can do local prediction on \"open hidde\" public data.</li>\n<li>i can submit local prediction to get lb score.</li>\n<li>if lb score is good, then i can focus effects on how to setup reduced msa server in kaggle notebook or replaced the msa part for legitimate solution.</li>\n</ul>\n<hr>\n<p>my style of work is</p>\n<ul>\n<li>first prove that model work (or at least you can get an upper bound estimate)</li>\n<li>then it is engineering work … how to implement</li>\n</ul>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3148718,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-13T13:12:39.877000",
              "content": "<p>there is no point creating pesudo label for casp15/16. this train set is too small.<br>\nwhat we should is is create pesudo label for a large dataset whose distribution is such that casp15/16 are inliers</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3148729,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-03-13T13:21:00.297000",
              "content": "<p>Nice trick! As you say, it's not a generally applicable approach but kind of probing, but I really learn something.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3148733,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-03-13T13:23:52.553000",
              "content": "<blockquote>\n  <p>there is no point creating pesudo label for casp15/16. this train set is too small.<br>\n  what we should is is create pesudo label for a large dataset whose distribution is such that casp15/16 are inliers</p>\n</blockquote>\n<p>It's true alright. you resolve my confusion.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3144066,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-07T22:06:53.400000",
      "content": "<p>here i want to introduce some unconventional papers related to rna/dna/protein 3d structure prediction which may be useful for our competition.</p>\n<p>1) deepseek R1 like method …. the ah-ha moment<br>\n<a href=\"https://x.com/hla_michael/status/1898106485005336988\" target=\"_blank\">https://x.com/hla_michael/status/1898106485005336988</a><br>\n\"Introducing Pro-1, an 8b param reasoning model trained using GRPO towards a physics based reward function for protein stability.\"</p>\n<p>though process:<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13038ccdb143e2cb6b0b218b5f238ef%2FSelection_999(7748).png?generation=1741385202898702&amp;alt=media\" alt=\"\"></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 3156237,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-21T23:28:35.053000",
      "content": "<p>what i learn from CASP16:</p>\n<ul>\n<li>modeling is not enough !!! focus on pipeline (combinations of methods)</li>\n<li>AF3 is part of many solution</li>\n</ul>\n<p>for kaggle, it would be 40 test targets in 8 GPU hours, meaning 10 min per target (5 RNA structures)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ed13500816d02ea899ca5ac89c9d3cf%2FSelection_112.png?generation=1742599652340594&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0ff2bd3c3067a9aef7c937ec47e6d08a%2FSelection_113.png?generation=1742599706052300&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9d7d1f53441d58dd9f8c4e699d0debf6%2FSelection_108.png?generation=1742599676893502&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3158245,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-24T11:08:22.713000",
      "content": "<p>I find that drfold2 implements alphafold2. The notation follows exactly af2 supplementary paper. The only replace msa generation with their rna lm. There is no training script or loss in the drfold2 repo  but one can just follow af2</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 3150672,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-15T18:29:02.573000",
      "content": "<p>prelimnary results of DRfold2 is very good!<br>\nR1107 : top 3 prediction 0.55309, 0.31959, 0.30806<br>\nR1116: top3 prediction 0.61168,0.59528</p>\n<p>i am surprised that such good results is possible without MSA input</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3152704,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-18T04:48:20.150000",
          "content": "<p>here is drfold2 magic:</p>\n<p>in the github repo, 5 set of  20 model checkpoints are provided.<br>\nif I use one full set (20 models) to predict 20 conformations, here is the tm scores for first casp15 target  R1107</p>\n<pre><code> R1107  GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU...\n### tm: [., ., , , , , , , , , , , , , , ., , , , ]\n\nMAX \n</code></pre>\n<p>it is a very high value of 0.61 if I select the best score.<br>\nthe repo code has a lot of redundant computation. (actually repeated computation) so you need to clean it up.</p>\n<p>instead of reading the code, a fast way to detect \"unless\" computation is to print the intermediate tensor output values out at critical steps of the network for 20 models on the same input seq. you will see that some steps produce identical values, aka, redundancy found!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3152798,
              "author_name": "Shosuke Suzuki",
              "author_url": "",
              "post_date": "2025-03-18T07:07:40.847000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I commented out the processing after the \"Selection\" part in DRfold_infer.py and modified it to obtain the coordinates of the C4 atoms from the initially generated pdb file. I also limited the model_hub weights to just one. Additionally, I downgraded the torch version to 2.2.1 so that it runs on Kaggle Notebooks as well. Is your environment set up similarly?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3152800,
              "author_name": "Shosuke Suzuki",
              "author_url": "",
              "post_date": "2025-03-18T07:09:59.437000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Based on these results, could it be that rather than drfold2 being exceptionally superior, it’s rhofold or nufold that require particularly high-quality and extensive MSAs? If that’s the case, then drfold2’s ensemble seems to hold the highest potential.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3152818,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-18T07:42:36.003000",
              "content": "<p>\" … e processing after the \"Selection\" part in DRfold_infer.py …\"<br>\ni use pytorch 2.5. and the orginal version in kaggle notebook is ok with me.</p>\n<p>whatever you do, just note that:</p>\n<ul>\n<li>we jsut need output from MSA2XYZ model:</li>\n</ul>\n<pre><code>        ret = x_pre()()()\n        ret = F(self(z),dim=-)()()()(np.float16)\n        ret = F(self(z),dim=-)()()()(np.float16)\n        ret = F(self(z),dim=-)()()()(np.float16)\n        ret = plddt()()()\n        return ret\n</code></pre>\n<p>i am using only   ret['coor'], which is the (P,C4,N1/9) frame in my description below<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512</a></p>\n<p>ret['coor'] is Lx3x3 = L x(P,C4,N1/9) x (xyz)<br>\nbut kaggle needs C1 instaed of C4.</p>\n<p>we may need to make another layer to predict C1 from (P,C4,N1/9) frame, etc</p>\n<p>as a temporarily solution, the Optimzation (in potentialFold) code has a save to save Lx3x3 frame into full atom pdb. i am using that. .</p>\n<pre><code>rama=.init_quat(   your MSA2XYZ output  )..numpy()\nrama = rama  \n.outpdb(rama ,energystr=)\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3152821,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-18T07:44:17.607000",
              "content": "<p>\"e, then drfold2’s ensemble seems to hold the highest potential.\"</p>\n<p>winning solution after 3 month would be combintions of these and non deep learning methods</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3152834,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-18T07:54:52.163000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3153184,
              "author_name": "Shosuke Suzuki",
              "author_url": "",
              "post_date": "2025-03-18T13:45:57.777000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for your reply. I see… this is very insightful. I will try this as well. Your experiment notes are very educational—thank you!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3153562,
              "author_name": "lhhappy",
              "author_url": "",
              "post_date": "2025-03-19T01:12:03.997000",
              "content": "<p>In my understanding, optimization requires force field optimization to obtain C1, but this process takes a long time. How can we speed it up? Thank you for your help.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3153581,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-19T01:38:52.173000",
              "content": "<p>the optimization optimizes (P,C4,N1/9) by computing local transformation (rotation + translation). then location of C1 and rest of the atoms obtained from transformations. </p>\n<p>we are not optimizing, rather we just get location of C1 from current transformations. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3155490,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-21T04:49:35.817000",
              "content": "<p><a href=\"https://www.kaggle.com/shosukesuzuki\" target=\"_blank\">@shosukesuzuki</a> <br>\nnufold is very important (it is 3rd in CASP)<br>\nagain, the MSA issue needs to be solved for nufold,alphafold3 and rhofold+.</p>\n<p>need to do experiment to find the MSA pipline they are using, then request host to prepare MSA accoriding to our pipline script</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2174f217b283996332fb89518eef965a%2FSelection_999(7818).png?generation=1742532472691036&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3156874,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-03-22T17:09:22.657000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> are you finetuning by adding some layers for c1 atoms prediction?  i tried using the local base frame landed up in 0.268LB</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3180008,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2025-04-16T02:55:27.240000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/shosukesuzuki\" target=\"_blank\">@shosukesuzuki</a> are you training any additional models for C1 prediction or are there any inbuilt functions in Drfold2 to accurately predict C1? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3162463,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-29T08:43:41.583000",
      "content": "<p>triangle flash attnetion!!!!<br>\n<a href=\"https://arxiv.org/pdf/2412.10743v2\" target=\"_blank\">https://arxiv.org/pdf/2412.10743v2</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e5bf4c2e18b24f34f18e286691bcda3%2FSelection_166.png?generation=1743237819947581&amp;alt=media\" alt=\"\"></p>\n<p>related?<br>\n<a href=\"https://github.com/Dao-AILab/flash-attention/pull/57\" target=\"_blank\">https://github.com/Dao-AILab/flash-attention/pull/57</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3207689,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-23T06:13:16.780000",
      "content": "<p>A little complex for me, achieve score about 0.4 is pretty hard!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3154165,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-19T15:40:48.393000",
      "content": "<p>energy as an effective way to select best 5 structure for submission</p>\n<p>graph show energy vs tm score (rank-normalised)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7e5d1eacb7fd0097e25c766b094bc10%2FSelection_095.png?generation=1742398800712595&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 3154253,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-19T17:56:27.307000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3180723,
          "author_name": "Bianco Chiu",
          "author_url": "",
          "post_date": "2025-04-17T03:35:47.720000",
          "content": "<p>Could you share how to calculate energy and how long does it take?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3152509,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-17T22:16:34.177000",
      "content": "<p>archive:<br>\n<strong>baseline code:</strong><br>\n1) 11-mar : lb 0.196</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/baseline-droupout\" target=\"_blank\">https://www.kaggle.com/code/hengck23/baseline-droupout</a></li>\n<li>just train at constant 1e-4 for 200 epoch. follow other setup (loss, etc) from host example notebook</li>\n<li>note that better CV may not gives better LB (dropout gives less variation) </li>\n<li>long epoch is required for alignment to be stable and aligned loss to be useful<br>\n(a better alignment free loss function is required)</li>\n</ul>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3153162,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-18T13:24:09.710000",
      "content": "<p>current plan for the week<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f98bc059b250bf822c2577a6a90c61c%2FSelection_086.png?generation=1742304205309135&amp;alt=media\" alt=\"\"></p>\n<p>the deep net code should run 3x faster then the repo after refractoring</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3150661,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-15T18:06:54.173000",
      "content": "<p>some lightweight model<br>\nMiniFold: Simple, Fast, and Accurate Protein Structure Prediction<br>\n<a href=\"https://openreview.net/forum?id=1p9hQTbjgo\" target=\"_blank\">https://openreview.net/forum?id=1p9hQTbjgo</a></p>\n<hr>\n<p>NOTE: I think different paper with a name:</p>\n<p><a href=\"https://github.com/hypnopump/MiniFold\" target=\"_blank\">https://github.com/hypnopump/MiniFold</a><br>\nMiniFold: a DeepLearning-based Mini Protein Folding Engine</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3150316,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-15T10:46:44.097000",
      "content": "<p>simple experiment on usalign parameters for casp15</p>\n<p>conclusion : kaggle \"backbone (C1') tm scorre\" is smiliar to those reported on paper for casp \"whole molecule tm score\".</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd474b8adf88db2edefd1dcfa9a822e0d%2FSelection_057.png?generation=1742040536720801&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3149540,
      "author_name": "Shosuke Suzuki",
      "author_url": "",
      "post_date": "2025-03-14T10:56:08.923000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> You mentioned training for around 200 epochs, but I believe you could accelerate your experiments by using the optimizer from <a href=\"https://github.com/facebookresearch/schedule_free\" target=\"_blank\">Facebook Research's Schedule-Free Optimization</a>. In my tests, this optimizer not only works well for image recognition tasks as shown in the paper but also significantly speeds up convergence in fine-tuning DNA language models and training VAEs. It eliminates the need for learning rate scheduling and warm-up. Since it's implemented in PyTorch, you can easily use it by directly incorporating the source code.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3150708,
          "author_name": "Arunodhayan",
          "author_url": "",
          "post_date": "2025-03-15T19:54:37.097000",
          "content": "<p>i trained for 400 epochs reached  LB 0.205</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3182379,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-19T07:55:20.957000",
      "content": "<p>tricks of nufold<br>\ni think the papers are out</p>\n<p>Structure Modeling Protocols for Protein Multimer and RNA in CASP16 with Enhanced MSAs, Model Ranking, and Deep Learning<br>\n<a href=\"https://d197for5662m48.cloudfront.net/documents/publicationstatus/254072/preprint_pdf/f7a0df47ad2e526a2b3baeed49617be0.pdf\" target=\"_blank\">https://d197for5662m48.cloudfront.net/documents/publicationstatus/254072/preprint_pdf/f7a0df47ad2e526a2b3baeed49617be0.pdf</a><br>\n<a href=\"https://advance.sagepub.com/doi/full/10.22541/au.174478900.09405488/v1\" target=\"_blank\">https://advance.sagepub.com/doi/full/10.22541/au.174478900.09405488/v1</a></p>\n<p>\"For RNA, we adopted an ensemble approach that incorporated multiple state-of-the-art methods,<br>\ncentered around our NuFold framework. As a result, our KiharaLab group ranked first in protein complex prediction and third<br>\nin RNA structure prediction\"</p>\n<p>\"To enhance the depth of multiple sequence alignments (MSAs), we employed a large metagenomic sequence database. Model ranking was performed with a stateof-the-art consensus ranking method, to which we added more scoring terms. \"</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3143672,
      "author_name": "Simon Beck",
      "author_url": "",
      "post_date": "2025-03-07T13:11:23.233000",
      "content": "<p>Glad to have you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 🤗</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3168413,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2025-04-02T12:02:16.337000",
      "content": "<p><strong>DRFOLD2 :</strong>    LB : 0.327 (no MSA) ,  0.318 (with MSA from kaggle folder).  I believe I am making a mistake in utilizing MSA -  No energy score used yet</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3168541,
          "author_name": "Zhu Siqi",
          "author_url": "",
          "post_date": "2025-04-02T14:40:40.997000",
          "content": "<p>It seems that DRFOLD2 doesn't require a MSA file. Do you modify the model to enable the use of MSA?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3168556,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-04-02T14:58:46.843000",
              "content": "<p>Rhofold+  <br>\nAF3</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3181763,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-18T09:50:40.870000",
              "content": "<p>DRFOLDv1 uses MSA</p>\n<p>DRFOLDv2 doesn't</p>\n<p>part of DRFOLD2 code is modified from DRFOLD1.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3167131,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-01T09:46:05.207000",
      "content": "<p>Nufold casp16<br>\n<a href=\"https://x.com/kiharalab/status/1906823369372434641?s=46\" target=\"_blank\">https://x.com/kiharalab/status/1906823369372434641?s=46</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3168237,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2025-04-02T08:58:45.703000",
          "content": "<p>Hi, hengck23. extra rMSA needs 2TB, and how can we solve it in kaggle env?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3168239,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-02T09:04:31.397000",
              "content": "<p>I thought rmsa already given by host?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3168246,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-04-02T09:09:25.697000",
              "content": "<p>oh, I'm stupid.😂</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3161488,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-28T02:41:26",
      "content": "<p>managed to repeat CASP15 paper results for proteinX</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31cb8c0a1349eba2a8e4d80ec828ce1e%2FSelection_168.png?generation=1743254577598872&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3161562,
          "author_name": "Zacchaeus",
          "author_url": "",
          "post_date": "2025-03-28T05:14:36.237000",
          "content": "<p>looks good results! did you use MSA?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3161699,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-28T09:32:59.350000",
              "content": "<p>proteinx disabled msa for rna. we are trying to patch that now.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3162615,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-29T13:24:18.863000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2695f9b6709b3893d2e98b079eee4d42%2FSelection_169.png?generation=1743254603712407&amp;alt=media\" alt=\"\"></p>\n<p>chai-1 results (also no MSA)<br>\ncomparable to protein-X.</p>\n<p>Obviously, the combination of the two is better (see r1108,r1128)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3161351,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-27T19:40:03.973000",
      "content": "<p>there is an assembly method in the public notebook that is coring quite well.</p>\n<p><a href=\"https://www.biorxiv.org/content/10.1101/2023.12.22.573067v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2023.12.22.573067v1.full.pdf</a><br>\n\"Vfold3D (41) constructs 3D structures from fragment databases. It uses the lowest free energy secondary structures converted to known fragments. The reconstruction of fragments is coarse-grained before being converted to allatom.\"</p>\n<p>a smiliar method is <a href=\"https://github.com/Tan-group/FebRNA\" target=\"_blank\">https://github.com/Tan-group/FebRNA</a>, this is quite fast</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3159835,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-26T01:13:10.753000",
      "content": "<p>how to compare msa search<br>\n<a href=\"https://academic.oup.com/gpb/article/22/1/qzae018/7617691?login=false\" target=\"_blank\">https://academic.oup.com/gpb/article/22/1/qzae018/7617691?login=false</a></p>\n<p>Performance comparison on RNA homology search<br>\nTable 1 compares the MSAs generated by RNAcmap2, rMSA, and RNAcmap3 in terms of median Neff and average F1-score given by mfDCA for the MSAs. The distribution of F1-scores for individual RNAs is shown in Figure 2.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3157259,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-23T07:35:14.400000",
      "content": "<p>team plan (see attachment)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3181587,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-18T04:50:53.087000",
      "content": "<p>new paper!<br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2025.04.10.648174v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.04.10.648174v1</a><br>\nAlphaFold3 at CASP16</p>\n<p>All analysis scripts are freely available from<br>\n<a href=\"https://gitlab.com/arneelof/CASP16-predictions\" target=\"_blank\">https://gitlab.com/arneelof/CASP16-predictions</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa6b9ad0c6fd8d3e425d0bdb725e3929%2FSelection_999(8109).png?generation=1744952168393928&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa2d39c07e6b27738dceafa0b92cdd8b7%2FSelection_999(8108).png?generation=1744952036291418&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3154706,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-20T10:04:04.750000",
      "content": "<p>dataset<br>\nDeep learning for RNA structure prediction <br>\n<a href=\"https://www.sciencedirect.com/science/article/pii/S0959440X25000090\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S0959440X25000090</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65b83d3f7c1ccf6d94006a9b688834d9%2FSelection_097.png?generation=1742465027201044&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F57cda013db8eb5944153b6655cbcc9c1%2FSelection_098.png?generation=1742465245565095&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3154089,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-19T14:00:09.040000",
      "content": "<p><a href=\"https://elanapearl.github.io/blog/2024/the-illustrated-alphafold/\" target=\"_blank\">https://elanapearl.github.io/blog/2024/the-illustrated-alphafold/</a><br>\ni am study af3. thi is useful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3153768,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-19T06:24:51.973000",
      "content": "<p>MSA or not MSA, here is the difference!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1f7bfa214db89e1570d9c4a90ab3ccd2%2FSelection_087.png?generation=1742365489128348&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3152205,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-17T15:07:06.860000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc978760a9f37d267c34a0deba1027562%2FSelection_075.png?generation=1742224007338988&amp;alt=media\" alt=\"\"></p>\n<p>rna  foundation model</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3148692,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-13T12:38:36.173000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3147849,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-12T13:30:26.177000",
      "content": "",
      "votes": -4,
      "replies": []
    },
    {
      "id": 3181877,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-18T13:18:58.697000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3182454,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-19T10:55:03.410000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 3182455,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-19T11:00:56.337000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3182458,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-19T11:05:23.343000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3182472,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-19T11:37:48.683000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3182486,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-19T12:08:03.740000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3158168,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-24T09:25:12.187000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3156196,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-21T21:13:21.423000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3155623,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-21T08:13:49.067000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3155126,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-20T18:25:49.093000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3166956,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-01T07:07:05.393000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3178336,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-14T03:41:38.497000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3155119,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-20T18:21:37.803000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3155129,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-20T18:29:53.537000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 3155135,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-20T18:32:32.140000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3155130,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-20T18:30:21.557000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 3155138,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-20T18:35:35.357000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3157133,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-23T02:53:47.943000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3153414,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-18T19:23:27.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3152644,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-18T03:03:10.733000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3152490,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-17T21:32:58.733000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3152967,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-18T10:04:55.253000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 3153015,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-18T11:05:57.267000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3151459,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-16T17:55:23.037000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3151465,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-16T18:04:05.883000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3148334,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-13T02:41:49.710000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3148339,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-13T02:51:14.290000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 3148369,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-13T03:35:59.510000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3148379,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-13T03:53:58.920000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3148388,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-13T03:59:40.390000",
              "content": "",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3148883,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-13T15:55:37.757000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3147476,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-12T04:09:52.553000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3147481,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-12T04:16:17.687000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3147347,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-12T00:34:49.420000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3147364,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-12T00:53:23.440000",
          "content": "",
          "votes": 8,
          "replies": [
            {
              "id": 3150591,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-15T16:32:03.430000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3207663,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-23T05:33:33.060000",
              "content": "",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3147346,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-12T00:34:42.567000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3147369,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-12T01:13:52.047000",
          "content": "",
          "votes": 4,
          "replies": [
            {
              "id": 3147378,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-12T01:25:26.723000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3207662,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-23T05:32:33.563000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3143558": "... to be updated as experiment proceeds ... each week, we detail how to improve your lb score ...please come here often \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F60350c3075fb64f8b8e86240fdb1894a%2FSelection_096.png?generation=1742398747189706&alt=media)\n**baseline code:**\n1) 18-mar : lb 0.321\n- https://www.kaggle.com/code/hengck23/lb0-286-simple-drfold-no-msa\n- no MSA, instead a RNA language model to model evolutionary information\n- overall, it is large RNA-lm and smaller structure downstream model\n- use 5 out of 100 avaliable  smaller structure models (e.g. 16 mb) for current lb\n- more structure models is better!!\n- plan: \n1.energy scoring to select best model results, clustering? (see paper)\n2.train model to convert from (N,C,P)frame + distance map --> C1 backbone\n3.quantisation and fp16 to extend length to 800?\n\ni think Drfold2 alone can get to lb0.40 on public test?\n\n----\n\n##Acknowledgement\n\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"\n\n##Hardware\nGPU: 2x Nvidia Ada A6000 (Ampere), each with VRAM 48 GB\nCPU: Intel® Xeon(R) w7-3455 CPU @ 2.5GHz, 24 cores, 48 threads\nMemory: 256 GB RAM",
    "3147190": "https://github.com/bytedance/Protenix\nA trainable PyTorch reproduction of AlphaFold 3.\n\ncurious how the diffuser work ...",
    "3155466": "the most important benchmark, must read!!!!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffba9ee0c870e82e97dbb5c62f68655e4%2FSelection_116.png?generation=1742659558654264&alt=media)\n\nnotebook: https://www.kaggle.com/code/hengck23/vfold-baseline-offline\n\nIt's frustrating that my big local improvements on DRfold2 didn’t show up as gains on the leaderboard.\nso I suspect there is some RNA that is usually difficult for drfold ONLY.\ni think I am correct\n\n\ndrfoldv2 totally breakdown at rna >400\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F84b48cea4acbfaf14c3bd590d016128a%2FSelection_105.png?generation=1742534073508454&alt=media)\n",
    "3154714": "vfold is the benchmark and we should understand it:\nhttps://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf\nhint: go to casp16 website to be vfold tm score and make a rank chart\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8ff6ed5e9aef06780c5d9e6f545121bf%2FSelection_099.png?generation=1742465512170143&alt=media)\n\nU must be very good at short length",
    "3148681": "I shared a technique for experimenting with MSA on the \"open hidden\" test set for the public leaderboard.\n\n**Please note that this is for testing purposes only and should not be considered a legitimate solution for claiming any prizes.**\n\n1. go to casp16 and find the sequence fasta of rna targets\n2. you can now predict their 3d structure offline (e.g. using web mas server)\n3. the host said that some public test rna are indeed casp16 target, so in your submission code use:\n\n```\nOFFLINE_PREDICT={\n   'AAAGGGCCC ...' : predicted xyz  ...,\n   'ACGACGACG ...' : predicted xyz  ...,\n}\n\n\nfor i,row in test_df ...:\n    if row.sequence is found in OFFLINE_PREDICT :\n            xyz = predicted xyz \n   else if not found:\n           xyz = use online model\n\n```\n\nin my experiments, I thin I can conclude:\nif you do very well in casp15, casp16, rna puzzles you should do well in this competition in public and private hidden dataset for the three different phrases. \n\ndo well doesn't mean train on these dataset but your model can generalise on these datasets.\ni think the level of difficulty of this competition is similar to these datasets.\n\n\n---\n\nwhat you can do next is to measure the alignment score (% of matched seq) on these dataset to kaggle train (or msa) and kaggle 400k diffusion synthetic data. This will tell you if these are inlier or outlier data.\n\nML methdos tends to work (only?) on inliers data. ",
    "3144066": "here i want to introduce some unconventional papers related to rna/dna/protein 3d structure prediction which may be useful for our competition.\n\n1) deepseek R1 like method .... the ah-ha moment\nhttps://x.com/hla_michael/status/1898106485005336988\n\"Introducing Pro-1, an 8b param reasoning model trained using GRPO towards a physics based reward function for protein stability.\"\n\nthough process:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13038ccdb143e2cb6b0b218b5f238ef%2FSelection_999(7748).png?generation=1741385202898702&alt=media)",
    "3156237": "what i learn from CASP16:\n- modeling is not enough !!! focus on pipeline (combinations of methods)\n- AF3 is part of many solution\n\nfor kaggle, it would be 40 test targets in 8 GPU hours, meaning 10 min per target (5 RNA structures)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ed13500816d02ea899ca5ac89c9d3cf%2FSelection_112.png?generation=1742599652340594&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0ff2bd3c3067a9aef7c937ec47e6d08a%2FSelection_113.png?generation=1742599706052300&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9d7d1f53441d58dd9f8c4e699d0debf6%2FSelection_108.png?generation=1742599676893502&alt=media)",
    "3158245": "I find that drfold2 implements alphafold2. The notation follows exactly af2 supplementary paper. The only replace msa generation with their rna lm. There is no training script or loss in the drfold2 repo  but one can just follow af2",
    "3150672": "prelimnary results of DRfold2 is very good!\nR1107 : top 3 prediction 0.55309, 0.31959, 0.30806\nR1116: top3 prediction 0.61168,0.59528\n\ni am surprised that such good results is possible without MSA input\n",
    "3162463": "triangle flash attnetion!!!!\nhttps://arxiv.org/pdf/2412.10743v2\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0e5bf4c2e18b24f34f18e286691bcda3%2FSelection_166.png?generation=1743237819947581&alt=media)\n\nrelated?\nhttps://github.com/Dao-AILab/flash-attention/pull/57",
    "3207689": "A little complex for me, achieve score about 0.4 is pretty hard!",
    "3154165": "energy as an effective way to select best 5 structure for submission\n\ngraph show energy vs tm score (rank-normalised)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7e5d1eacb7fd0097e25c766b094bc10%2FSelection_095.png?generation=1742398800712595&alt=media)",
    "3152509": "archive:\n**baseline code:**\n1) 11-mar : lb 0.196\n- https://www.kaggle.com/code/hengck23/baseline-droupout\n- just train at constant 1e-4 for 200 epoch. follow other setup (loss, etc) from host example notebook\n- note that better CV may not gives better LB (dropout gives less variation) \n- long epoch is required for alignment to be stable and aligned loss to be useful\n(a better alignment free loss function is required)\n",
    "3153162": "current plan for the week\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f98bc059b250bf822c2577a6a90c61c%2FSelection_086.png?generation=1742304205309135&alt=media)\n\nthe deep net code should run 3x faster then the repo after refractoring\n",
    "3150661": "some lightweight model\nMiniFold: Simple, Fast, and Accurate Protein Structure Prediction\nhttps://openreview.net/forum?id=1p9hQTbjgo\n\n---\n\nNOTE: I think different paper with a name:\n\nhttps://github.com/hypnopump/MiniFold\nMiniFold: a DeepLearning-based Mini Protein Folding Engine",
    "3150316": "simple experiment on usalign parameters for casp15\n\nconclusion : kaggle \"backbone (C1') tm scorre\" is smiliar to those reported on paper for casp \"whole molecule tm score\".\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd474b8adf88db2edefd1dcfa9a822e0d%2FSelection_057.png?generation=1742040536720801&alt=media)",
    "3149540": "@hengck23 You mentioned training for around 200 epochs, but I believe you could accelerate your experiments by using the optimizer from [Facebook Research's Schedule-Free Optimization](https://github.com/facebookresearch/schedule_free). In my tests, this optimizer not only works well for image recognition tasks as shown in the paper but also significantly speeds up convergence in fine-tuning DNA language models and training VAEs. It eliminates the need for learning rate scheduling and warm-up. Since it's implemented in PyTorch, you can easily use it by directly incorporating the source code.",
    "3182379": "tricks of nufold\ni think the papers are out\n\nStructure Modeling Protocols for Protein Multimer and RNA in CASP16 with Enhanced MSAs, Model Ranking, and Deep Learning\nhttps://d197for5662m48.cloudfront.net/documents/publicationstatus/254072/preprint_pdf/f7a0df47ad2e526a2b3baeed49617be0.pdf\nhttps://advance.sagepub.com/doi/full/10.22541/au.174478900.09405488/v1\n\n\"For RNA, we adopted an ensemble approach that incorporated multiple state-of-the-art methods,\ncentered around our NuFold framework. As a result, our KiharaLab group ranked first in protein complex prediction and third\nin RNA structure prediction\"\n\n\"To enhance the depth of multiple sequence alignments (MSAs), we employed a large metagenomic sequence database. Model ranking was performed with a stateof-the-art consensus ranking method, to which we added more scoring terms. \"",
    "3143672": "Glad to have you @hengck23 🤗",
    "3168413": "**DRFOLD2 :**    LB : 0.327 (no MSA) ,  0.318 (with MSA from kaggle folder).  I believe I am making a mistake in utilizing MSA -  No energy score used yet",
    "3167131": "Nufold casp16\nhttps://x.com/kiharalab/status/1906823369372434641?s=46",
    "3161488": "managed to repeat CASP15 paper results for proteinX\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31cb8c0a1349eba2a8e4d80ec828ce1e%2FSelection_168.png?generation=1743254577598872&alt=media)",
    "3161351": "there is an assembly method in the public notebook that is coring quite well.\n\nhttps://www.biorxiv.org/content/10.1101/2023.12.22.573067v1.full.pdf\n\"Vfold3D (41) constructs 3D structures from fragment databases. It uses the lowest free energy secondary structures converted to known fragments. The reconstruction of fragments is coarse-grained before being converted to allatom.\"\n\n a smiliar method is https://github.com/Tan-group/FebRNA, this is quite fast",
    "3159835": "how to compare msa search\nhttps://academic.oup.com/gpb/article/22/1/qzae018/7617691?login=false\n\nPerformance comparison on RNA homology search\nTable 1 compares the MSAs generated by RNAcmap2, rMSA, and RNAcmap3 in terms of median Neff and average F1-score given by mfDCA for the MSAs. The distribution of F1-scores for individual RNAs is shown in Figure 2.\n",
    "3157259": "team plan (see attachment)",
    "3181587": "new paper!\nhttps://www.biorxiv.org/content/10.1101/2025.04.10.648174v1\nAlphaFold3 at CASP16\n\nAll analysis scripts are freely available from\nhttps://gitlab.com/arneelof/CASP16-predictions\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa6b9ad0c6fd8d3e425d0bdb725e3929%2FSelection_999(8109).png?generation=1744952168393928&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa2d39c07e6b27738dceafa0b92cdd8b7%2FSelection_999(8108).png?generation=1744952036291418&alt=media)",
    "3154706": "dataset\nDeep learning for RNA structure prediction \nhttps://www.sciencedirect.com/science/article/pii/S0959440X25000090\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65b83d3f7c1ccf6d94006a9b688834d9%2FSelection_097.png?generation=1742465027201044&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F57cda013db8eb5944153b6655cbcc9c1%2FSelection_098.png?generation=1742465245565095&alt=media)",
    "3154089": "https://elanapearl.github.io/blog/2024/the-illustrated-alphafold/\ni am study af3. thi is useful",
    "3153768": "MSA or not MSA, here is the difference!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1f7bfa214db89e1570d9c4a90ab3ccd2%2FSelection_087.png?generation=1742365489128348&alt=media)",
    "3152205": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc978760a9f37d267c34a0deba1027562%2FSelection_075.png?generation=1742224007338988&alt=media)\n\n rna  foundation model",
    "3148692": "😭 day two of my download\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3911f386181723cf3ef748bc856ea250%2FSelection_046.png?generation=1741869514167103&alt=media)\n\nthe only competition where external data is in TB ... i estimate i will burn away 6 to 8 TB",
    "3147849": "@hengck23 想问一下您的base中ribonanza_net是一个什么样的东西呢",
    "3181877": "it is here!!!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Feb791133a41c2501370ec4e4fc7fc186%2FSelection_999(8114).png?generation=1744982252631348&alt=media)\n\ntrRosettaRNA2 is a new deep learning model for RNA 3D structure prediction that incorporates a novel secondary structure prior module trained on bpRNA data. ... aka yang-server\n\ndo read the supplementary paper as well\nhttps://www.biorxiv.org/content/10.1101/2025.04.09.647915v1",
    "3158168": "how long on average do you run these model ?",
    "3156196": "you are an og @hengck23",
    "3155623": "https://ar5iv.labs.arxiv.org/html/2203.00854\nfastfold",
    "3155126": "probe of my current drfoldv2 results:\n- current public notebook using 5 models is 0.321\n- if we improve rna of length<200 only, we can get about 0.334\n- if we improve rna of 200<length<400 only, we can get about 0.405\n\nif i use energy selection for 10 models, submission shows 0.321 --> 0.323\nengergy compute in drfold2 is very slow , so in this submissopm, i only do selection for those less than length 200.\n\nbut even if i can speed that up (just replace the cubic spline interpolation) and do selection up to 400 length, i can only end up 0.38 to 0.40 i think.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2d0d54914e9681244fe3ce6f840427e8%2FSelection_100.png?generation=1742495142254548&alt=media)",
    "3155119": "CASP-16 benchmark (results pulled from casp16 server):\n\nrna of length less than 400:\n```\nalphafold3: 0.5271428571428571\ndrfold2: 0.5173809523809524\nvfold: 0.5907619047619047\nalphafold3+drfold2: 0.5787142857142857\nALL: 0.6080476190476191\n\n```\nall rna\n```\nalphafold3: 0.4623720930232558\ndrfold2: 0.46891891891891885\nvfold: 0.5146511627906977\nalphafold3+drfold2: 0.5038139534883721\nALL: 0.5281627906976745\n\n```\n\n\nthe strategy is very clear. either alphafold3 or drfoldv2 alone cannot win vfold. we need combinations, the more the merrier.\ndue to hardware resources of kaggle notebook, we need to distill top models into single one\n1. there are 30 millions rna seq in RNACetral, etc ... but they don't have 3d structure\n2. most model like drfoldv2 are trained with about 18k rna (that is all 3d structure you can find in PDB)\n3. one solution is to use top models to puesdo label all rna in RNACetral, etc\n4. then distill to one that can run about rna length=1024 in kaggle notebook (we need to design this)\n\ntmr i submit alphafold3,drfold2,vfold (offline compute), etc to kaggle server.\n\nI think i can design a good network to run on kaggle notebook. I would probably release an untrained version of it in a couple of days.\n\nnow i need resource to train it.\n\n---\n\non a side note:\nyou can connect drfoldv2 RNACLM (already pretrained with evoluitionary feature) with RibonanzaNet.\nRibonanzaNet in host example currently does not use triangular attention (that code is commented out). \nbut if you use triangular attention, there may be resource OOM problems (same for other opensource model).\n\nliteformer claim to reduce triangular attention complexity from O(L^3) to O(L^2) without loss in performance on CASP15. this is what i am trying to do.",
    "3153414": "drfold2 memory and speed is restricted by the col and row triangle attention in the evo-transformer.\nignore the RNA language model, the structure model uses 10,432 train rna seq given in repo\n\nnow I can speed by 2x and reduce memory by 2x by using new structure\ni am wondering if I want to retrain it ....\n\nhttps://github.com/Dao-AILab/flash-attention/pull/57\nhttps://openreview.net/pdf?id=brk7Ct4Tb1M\nliteformer: https://openreview.net/forum?id=t0m0DdCCQ2\n(kaggle has only 844 train rna seq)",
    "3152644": "Do we need gu Mas search? \nhttps://developer.nvidia.com/blog/boost-alphafold2-protein-structure-prediction-with-gpu-accelerated-mmseqs2/",
    "3152490": "hi I'm new and I feel like this competition is not newbie friendly and computational costly right ? should I find other competition that suit me more or try my best here to learn? cause after seeing your result I can see the amount of resource going in training these models are not going to be cheap  ",
    "3151459": "@hengck23 LB 500 based on which model?",
    "3148334": "Do you plan to compete for the prestigious \"Early Sharing\" award?",
    "3147476": "are you using TM score? if so, how do you handle the sequences have less 41 native_cnt ",
    "3147347": "Oops is there a data leak?",
    "3147346": ">Hardware\nGPU: 2x Nvidia Ada A6000 (Ampere), each with VRAM 48 GB\nCPU: Intel® Xeon(R) w7-3455 CPU @ 2.5GHz, 24 cores, 48 threads\nMemory: 256 GB RAM\n\nMany well-known sotas for 3D RNA structure prediction require substantial memories and datasets. When I see you topping the leaderboard, I wonder if you had employed a similar approach. \n\nIf you managed to achieve this solely by modifying the baseline, that would be truly incredible, I can't believe someone can make it only in this way.\n\nBTW, glad to see you too.",
    "3207662": "Thanks for suggestions!"
  }
}