{
  "id": 609701,
  "title": "3rd Place Solution",
  "url": "/competitions/stanford-rna-3d-folding/writeups/3rd-place-solution",
  "author_name": "",
  "post_date": "2025-09-29T04:50:21.103Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Our 3rd Place Solution – Stanford RNA 3D Folding</h1>\n<p>First of all, thanks to Kaggle and the organizers for putting together such a challenging and interesting competition.  </p>\n<p>We finished <strong>3rd on the public leaderboard</strong> before the unseen data release, and also <strong>3rd on the private leaderboard</strong>.  </p>\n<p>Earlier we shared a short overview of our solution. This post goes into more detail.  </p>\n<hr>\n<h2>Summary of Our Solution</h2>\n<ul>\n<li>Ensemble of <strong>DRfold2</strong>, <strong>Protenix</strong>, and <strong>Boltz-1</strong>  </li>\n<li>Fine-tuned <strong>Protenix</strong> with newly released RNA datasets  </li>\n</ul>\n<hr>\n<h2>Detailed Solution</h2>\n<h3>Dataset</h3>\n<ul>\n<li><p><strong>rMSA</strong>  <br>\nWe generated our own rMSA data using the <a href=\"https://github.com/pylelab/rMSA\" target=\"_blank\">official rMSA code</a> provided by the hosts.  </p>\n<ul>\n<li>The v2 rMSA released by the organizers did not cover the full set of recently published data, so we had to build our own.  </li>\n<li>This process took about <strong>14 days</strong>, even with multiprocessing and several servers.  </li>\n<li>Our rMSA data is available here: <a href=\"https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link\" target=\"_blank\">Google Drive</a>.  </li></ul></li>\n<li><p><strong>Training dataset</strong>  <br>\nWe tested two versions:  </p>\n<ol>\n<li><strong>Full RNA dataset</strong> – included all recently available data, such as <a href=\"https://www.kaggle.com/datasets/tant64/casp16\" target=\"_blank\">CASP16</a> uploaded by <a href=\"https://www.kaggle.com/tant64\" target=\"_blank\">@tant64</a>. To check quality, we compared some of the labels with the <a href=\"https://gitlab.com/arneelof/CASP16-predictions/-/tree/main/rna_results?ref_type=heads\" target=\"_blank\">CASP16 GitLab repository</a>.  </li>\n<li><strong>RNA-only dataset</strong> – excluded complexes (RNA bound to proteins/DNA), following the clarification from the hosts <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/575745\" target=\"_blank\">here</a>.  </li></ol></li>\n</ul>\n<p>Both versions of the training data and labels are available here: <a href=\"https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n\" target=\"_blank\">Google Drive</a>.  </p>\n<hr>\n<h3>Models</h3>\n<p>We built our ensemble using models from these official repositories:  </p>\n<ul>\n<li><a href=\"https://github.com/leeyang/DRfold2\" target=\"_blank\">DRfold2</a>  </li>\n<li><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a>  </li>\n<li><a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">Boltz</a>  </li>\n</ul>\n<p>We initially explored new architectures, but given the time and resource limits, it was more effective to focus on fine-tuning and combining existing models.  </p>\n<hr>\n<h4>DRfold2</h4>\n<ul>\n<li>Performed best on sequences <strong>&lt;400 nt</strong> with v1 data.  </li>\n<li>Running the full pipeline with optimization, clustering, and all 80 released checkpoints was too slow.  </li>\n<li>We found that <strong>energy selection + Arena</strong> gave most of the gain, while the other steps added little to the TM-score but took significant runtime.  </li>\n<li>So our DRfold2 setup used <strong>energy selection + Arena</strong> only for the post-processing.  </li>\n</ul>\n<hr>\n<h4>Protenix</h4>\n<ul>\n<li>Trained on an <strong>NVIDIA GH200 (96GB VRAM)</strong>.  </li>\n<li>Limited to sequences <strong>&lt;800 nt</strong> due to memory issues.  </li>\n<li>Used fine-tuning code from <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> (<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\" target=\"_blank\">discussion link</a>).  </li>\n<li>Fine-tuning <strong>without rMSA</strong> did not help, so we trained <strong>with rMSA</strong>.  </li>\n<li>Only modified <code>max_steps</code>; a full dataset run took about a day.  </li>\n<li>Early checkpoints (2–3 cycles) performed worse; performance improved with longer fine-tuning.  </li>\n<li>We also tested <strong>multiple Protenix outputs → DRfold energy selection → Arena</strong>, but it did not improve results.  </li>\n<li>For inference, we modified <a href=\"https://www.kaggle.com/code/geraseva/protenix\" target=\"_blank\">@geraseva’s code</a> to include rMSA.  </li>\n</ul>\n<hr>\n<h4>Boltz-1</h4>\n<ul>\n<li>While not the strongest model on its own, including <strong>one Boltz prediction</strong> improved ensemble diversity.  </li>\n<li>This helped because the scoring metric favored picking the best out of multiple diverse outputs.  </li>\n<li>We used <a href=\"https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission\" target=\"_blank\">@youhanlee’s inference notebook</a>.  </li>\n</ul>\n<hr>\n<h3>Final Submissions &amp; Code</h3>\n<p>We are sharing our original submission notebooks. Some commented-out code remains, which shows parts of our experiment history.  </p>\n<ol>\n<li><p><strong><a href=\"https://www.kaggle.com/code/yekim102/drfold4-pro-msa-pro-msa-base-rna\" target=\"_blank\">DRfold2 + Protenix</a></strong>  </p>\n<ul>\n<li>Public LB: <strong>0.60338</strong> | Private LB: <strong>0.52787</strong>  </li>\n<li>&lt;400 nt: 3 × DRfold2 + 2 × Protenix (RNA-only, with MSA)  </li>\n<li>&gt;400 nt: 2 × Protenix (RNA-only) + 2 × Protenix (All) + Protenix baseline  </li></ul></li>\n<li><p><strong><a href=\"https://www.kaggle.com/code/yekim102/fin-pro-all-1-rna-2-base1-boltz1\" target=\"_blank\">Protenix + Boltz</a></strong>  </p>\n<ul>\n<li>Public LB: <strong>0.61253</strong> | Private LB: <strong>0.54312</strong>  </li>\n<li>2 × Protenix (RNA-only) + Protenix (All) + Protenix baseline + Boltz baseline  </li></ul></li>\n</ol>\n<hr>\n<h2>Acknowledgements</h2>\n<p>Many thanks to the community members who shared resources and insights:  </p>\n<ul>\n<li><strong>@hengck23</strong> – for extensive discussions and guidance.  </li>\n<li><strong>@geraseva</strong> – for early Protenix inference code.  </li>\n<li><strong>@lihaoweicvch</strong> – for Protenix fine-tuning code.  </li>\n<li><strong>@youhanlee</strong> – for updated Boltz inference code.  </li>\n</ul>\n<p>Finally, thanks to <strong>Eigen Company</strong> for providing resources for this competition. </p>",
  "messages": [
    {
      "id": "3295548",
      "postDate": "09/29/2025 04:48:51",
      "content": "<h1>Our 3rd Place Solution – Stanford RNA 3D Folding</h1>\n<p>First of all, thanks to Kaggle and the organizers for putting together such a challenging and interesting competition.  </p>\n<p>We finished <strong>3rd on the public leaderboard</strong> before the unseen data release, and also <strong>3rd on the private leaderboard</strong>.  </p>\n<p>Earlier we shared a short overview of our solution. This post goes into more detail.  </p>\n<hr>\n<h2>Summary of Our Solution</h2>\n<ul>\n<li>Ensemble of <strong>DRfold2</strong>, <strong>Protenix</strong>, and <strong>Boltz-1</strong>  </li>\n<li>Fine-tuned <strong>Protenix</strong> with newly released RNA datasets  </li>\n</ul>\n<hr>\n<h2>Detailed Solution</h2>\n<h3>Dataset</h3>\n<ul>\n<li><p><strong>rMSA</strong>  <br>\nWe generated our own rMSA data using the <a href=\"https://github.com/pylelab/rMSA\" target=\"_blank\">official rMSA code</a> provided by the hosts.  </p>\n<ul>\n<li>The v2 rMSA released by the organizers did not cover the full set of recently published data, so we had to build our own.  </li>\n<li>This process took about <strong>14 days</strong>, even with multiprocessing and several servers.  </li>\n<li>Our rMSA data is available here: <a href=\"https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link\" target=\"_blank\">Google Drive</a>.  </li></ul></li>\n<li><p><strong>Training dataset</strong>  <br>\nWe tested two versions:  </p>\n<ol>\n<li><strong>Full RNA dataset</strong> – included all recently available data, such as <a href=\"https://www.kaggle.com/datasets/tant64/casp16\" target=\"_blank\">CASP16</a> uploaded by <a href=\"https://www.kaggle.com/tant64\" target=\"_blank\">@tant64</a>. To check quality, we compared some of the labels with the <a href=\"https://gitlab.com/arneelof/CASP16-predictions/-/tree/main/rna_results?ref_type=heads\" target=\"_blank\">CASP16 GitLab repository</a>.  </li>\n<li><strong>RNA-only dataset</strong> – excluded complexes (RNA bound to proteins/DNA), following the clarification from the hosts <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/575745\" target=\"_blank\">here</a>.  </li></ol></li>\n</ul>\n<p>Both versions of the training data and labels are available here: <a href=\"https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n\" target=\"_blank\">Google Drive</a>.  </p>\n<hr>\n<h3>Models</h3>\n<p>We built our ensemble using models from these official repositories:  </p>\n<ul>\n<li><a href=\"https://github.com/leeyang/DRfold2\" target=\"_blank\">DRfold2</a>  </li>\n<li><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a>  </li>\n<li><a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">Boltz</a>  </li>\n</ul>\n<p>We initially explored new architectures, but given the time and resource limits, it was more effective to focus on fine-tuning and combining existing models.  </p>\n<hr>\n<h4>DRfold2</h4>\n<ul>\n<li>Performed best on sequences <strong>&lt;400 nt</strong> with v1 data.  </li>\n<li>Running the full pipeline with optimization, clustering, and all 80 released checkpoints was too slow.  </li>\n<li>We found that <strong>energy selection + Arena</strong> gave most of the gain, while the other steps added little to the TM-score but took significant runtime.  </li>\n<li>So our DRfold2 setup used <strong>energy selection + Arena</strong> only for the post-processing.  </li>\n</ul>\n<hr>\n<h4>Protenix</h4>\n<ul>\n<li>Trained on an <strong>NVIDIA GH200 (96GB VRAM)</strong>.  </li>\n<li>Limited to sequences <strong>&lt;800 nt</strong> due to memory issues.  </li>\n<li>Used fine-tuning code from <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> (<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\" target=\"_blank\">discussion link</a>).  </li>\n<li>Fine-tuning <strong>without rMSA</strong> did not help, so we trained <strong>with rMSA</strong>.  </li>\n<li>Only modified <code>max_steps</code>; a full dataset run took about a day.  </li>\n<li>Early checkpoints (2–3 cycles) performed worse; performance improved with longer fine-tuning.  </li>\n<li>We also tested <strong>multiple Protenix outputs → DRfold energy selection → Arena</strong>, but it did not improve results.  </li>\n<li>For inference, we modified <a href=\"https://www.kaggle.com/code/geraseva/protenix\" target=\"_blank\">@geraseva’s code</a> to include rMSA.  </li>\n</ul>\n<hr>\n<h4>Boltz-1</h4>\n<ul>\n<li>While not the strongest model on its own, including <strong>one Boltz prediction</strong> improved ensemble diversity.  </li>\n<li>This helped because the scoring metric favored picking the best out of multiple diverse outputs.  </li>\n<li>We used <a href=\"https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission\" target=\"_blank\">@youhanlee’s inference notebook</a>.  </li>\n</ul>\n<hr>\n<h3>Final Submissions &amp; Code</h3>\n<p>We are sharing our original submission notebooks. Some commented-out code remains, which shows parts of our experiment history.  </p>\n<ol>\n<li><p><strong><a href=\"https://www.kaggle.com/code/yekim102/drfold4-pro-msa-pro-msa-base-rna\" target=\"_blank\">DRfold2 + Protenix</a></strong>  </p>\n<ul>\n<li>Public LB: <strong>0.60338</strong> | Private LB: <strong>0.52787</strong>  </li>\n<li>&lt;400 nt: 3 × DRfold2 + 2 × Protenix (RNA-only, with MSA)  </li>\n<li>&gt;400 nt: 2 × Protenix (RNA-only) + 2 × Protenix (All) + Protenix baseline  </li></ul></li>\n<li><p><strong><a href=\"https://www.kaggle.com/code/yekim102/fin-pro-all-1-rna-2-base1-boltz1\" target=\"_blank\">Protenix + Boltz</a></strong>  </p>\n<ul>\n<li>Public LB: <strong>0.61253</strong> | Private LB: <strong>0.54312</strong>  </li>\n<li>2 × Protenix (RNA-only) + Protenix (All) + Protenix baseline + Boltz baseline  </li></ul></li>\n</ol>\n<hr>\n<h2>Acknowledgements</h2>\n<p>Many thanks to the community members who shared resources and insights:  </p>\n<ul>\n<li><strong>@hengck23</strong> – for extensive discussions and guidance.  </li>\n<li><strong>@geraseva</strong> – for early Protenix inference code.  </li>\n<li><strong>@lihaoweicvch</strong> – for Protenix fine-tuning code.  </li>\n<li><strong>@youhanlee</strong> – for updated Boltz inference code.  </li>\n</ul>\n<p>Finally, thanks to <strong>Eigen Company</strong> for providing resources for this competition. </p>",
      "rawMarkdown": "# Our 3rd Place Solution – Stanford RNA 3D Folding\n\nFirst of all, thanks to Kaggle and the organizers for putting together such a challenging and interesting competition.  \n\nWe finished **3rd on the public leaderboard** before the unseen data release, and also **3rd on the private leaderboard**.  \n\nEarlier we shared a short overview of our solution. This post goes into more detail.  \n\n---\n\n## Summary of Our Solution\n- Ensemble of **DRfold2**, **Protenix**, and **Boltz-1**  \n- Fine-tuned **Protenix** with newly released RNA datasets  \n\n---\n\n## Detailed Solution\n\n### Dataset\n\n- **rMSA**  \n  We generated our own rMSA data using the [official rMSA code](https://github.com/pylelab/rMSA) provided by the hosts.  \n  - The v2 rMSA released by the organizers did not cover the full set of recently published data, so we had to build our own.  \n  - This process took about **14 days**, even with multiprocessing and several servers.  \n  - Our rMSA data is available here: [Google Drive](https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link).  \n\n- **Training dataset**  \n  We tested two versions:  \n  1. **Full RNA dataset** – included all recently available data, such as [CASP16](https://www.kaggle.com/datasets/tant64/casp16) uploaded by @tant64. To check quality, we compared some of the labels with the [CASP16 GitLab repository](https://gitlab.com/arneelof/CASP16-predictions/-/tree/main/rna_results?ref_type=heads).  \n  2. **RNA-only dataset** – excluded complexes (RNA bound to proteins/DNA), following the clarification from the hosts [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/575745).  \n\nBoth versions of the training data and labels are available here: [Google Drive](https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n).  \n\n---\n\n### Models\n\nWe built our ensemble using models from these official repositories:  \n- [DRfold2](https://github.com/leeyang/DRfold2)  \n- [Protenix](https://github.com/bytedance/Protenix)  \n- [Boltz](https://github.com/jwohlwend/boltz)  \n\nWe initially explored new architectures, but given the time and resource limits, it was more effective to focus on fine-tuning and combining existing models.  \n\n---\n\n#### DRfold2\n- Performed best on sequences **<400 nt** with v1 data.  \n- Running the full pipeline with optimization, clustering, and all 80 released checkpoints was too slow.  \n- We found that **energy selection + Arena** gave most of the gain, while the other steps added little to the TM-score but took significant runtime.  \n- So our DRfold2 setup used **energy selection + Arena** only for the post-processing.  \n\n---\n\n#### Protenix\n- Trained on an **NVIDIA GH200 (96GB VRAM)**.  \n- Limited to sequences **<800 nt** due to memory issues.  \n- Used fine-tuning code from @lihaoweicvch ([discussion link](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495)).  \n- Fine-tuning **without rMSA** did not help, so we trained **with rMSA**.  \n- Only modified `max_steps`; a full dataset run took about a day.  \n- Early checkpoints (2–3 cycles) performed worse; performance improved with longer fine-tuning.  \n- We also tested **multiple Protenix outputs → DRfold energy selection → Arena**, but it did not improve results.  \n- For inference, we modified [@geraseva’s code](https://www.kaggle.com/code/geraseva/protenix) to include rMSA.  \n\n---\n\n#### Boltz-1\n- While not the strongest model on its own, including **one Boltz prediction** improved ensemble diversity.  \n- This helped because the scoring metric favored picking the best out of multiple diverse outputs.  \n- We used [@youhanlee’s inference notebook](https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission).  \n\n---\n\n### Final Submissions & Code\n\nWe are sharing our original submission notebooks. Some commented-out code remains, which shows parts of our experiment history.  \n\n1. **[DRfold2 + Protenix](https://www.kaggle.com/code/yekim102/drfold4-pro-msa-pro-msa-base-rna)**  \n   - Public LB: **0.60338** | Private LB: **0.52787**  \n   - <400 nt: 3 × DRfold2 + 2 × Protenix (RNA-only, with MSA)  \n   - >400 nt: 2 × Protenix (RNA-only) + 2 × Protenix (All) + Protenix baseline  \n\n2. **[Protenix + Boltz](https://www.kaggle.com/code/yekim102/fin-pro-all-1-rna-2-base1-boltz1)**  \n   - Public LB: **0.61253** | Private LB: **0.54312**  \n   - 2 × Protenix (RNA-only) + Protenix (All) + Protenix baseline + Boltz baseline  \n\n---\n\n## Acknowledgements\n\nMany thanks to the community members who shared resources and insights:  \n- **@hengck23** – for extensive discussions and guidance.  \n- **@geraseva** – for early Protenix inference code.  \n- **@lihaoweicvch** – for Protenix fine-tuning code.  \n- **@youhanlee** – for updated Boltz inference code.  \n\nFinally, thanks to **Eigen Company** for providing resources for this competition.",
      "votes": null
    },
    {
      "id": "3295653",
      "postDate": "09/29/2025 09:47:41",
      "content": "<p>congradualation.<br>\nHow do you ensemble the outputs of different models?<br>\nBecause each model outputs coordinates.</p>",
      "rawMarkdown": "congradualation.\nHow do you ensemble the outputs of different models?\nBecause each model outputs coordinates.",
      "votes": null
    },
    {
      "id": "3295907",
      "postDate": "09/29/2025 18:44:45",
      "content": "<p><a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> The competition's <code>submission.csv</code> expects competitors to predict 5 different \"candidates\" for each input sequence. Most competitors ensembled models in 3 ways:</p>\n<ol>\n<li><strong>Length based:</strong> check the length of the input sequence, if L &gt; threshold use model A if L &lt;= threshold use model B. Example: if the sequence is shorter than 400 nucleotides use DRFold2 and if its longer use Protenix.</li>\n<li><strong>Vertical ensemble:</strong> if you have 5 possible candidates use X/5 for model A and Y/5 for model B. Example: use 2 candidates for Protenix and 3 for DRFold2.</li>\n<li><strong>Align and aggregate:</strong> align all sequences to a reference coordinate frame and then aggregate (using the mean usually). I have not seen many succesfull attempts on this approach though.</li>\n</ol>\n<p>Hope this answers your question</p>",
      "rawMarkdown": "garybios The competition's `submission.csv` expects competitors to predict 5 different \"candidates\" for each input sequence. Most competitors ensembled models in 3 ways:\n1.  **Length based:** check the length of the input sequence, if L > threshold use model A if L <= threshold use model B. Example: if the sequence is shorter than 400 nucleotides use DRFold2 and if its longer use Protenix.\n2.  **Vertical ensemble:** if you have 5 possible candidates use X/5 for model A and Y/5 for model B. Example: use 2 candidates for Protenix and 3 for DRFold2.\n3. **Align and aggregate:** align all sequences to a reference coordinate frame and then aggregate (using the mean usually). I have not seen many succesfull attempts on this approach though.\n\nHope this answers your question",
      "votes": null
    },
    {
      "id": "3297305",
      "postDate": "10/02/2025 20:16:00",
      "content": "<p><a href=\"https://www.kaggle.com/yekim102\" target=\"_blank\">@yekim102</a> <a href=\"https://www.kaggle.com/ryankim99\" target=\"_blank\">@ryankim99</a> <a href=\"https://www.kaggle.com/ouiqdmw\" target=\"_blank\">@ouiqdmw</a> thanks for a beautiful writeup and sharing your insights as early as June! Congrats on having the top AF3-like solution.</p>\n<p>A question for you. In earlier discussions, it seemed like use of CASP16 blind predictions to train Protenix led to a big increase in accuracy on public LB:</p>\n<blockquote>\n  <p>We included Casp 16 data so that bumps up the public score.<br>\n  If we don't use the casp 16 data, we get about 0.42.<br>\n  (from <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/584487\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/584487</a>) </p>\n</blockquote>\n<p>But actually… our Private LB does not include any CASP16 targets, and your model still did super well!  </p>\n<p>So I wonder if CASP16 predictions were so necessary.</p>\n<p>Do you have a checkpoint somewhere where you trained <em>without</em> CASP16 targets (but perhaps with a very up to date PDB)? I wonder how it scores on Private LB.</p>",
      "rawMarkdown": "yekim102 @ryankim99 @ouiqdmw thanks for a beautiful writeup and sharing your insights as early as June! Congrats on having the top AF3-like solution.\n\nA question for you. In earlier discussions, it seemed like use of CASP16 blind predictions to train Protenix led to a big increase in accuracy on public LB:\n\n>We included Casp 16 data so that bumps up the public score.\n> If we don't use the casp 16 data, we get about 0.42.\n(from https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/584487) \n\nBut actually... our Private LB does not include any CASP16 targets, and your model still did super well!  \n\nSo I wonder if CASP16 predictions were so necessary.\n\nDo you have a checkpoint somewhere where you trained *without* CASP16 targets (but perhaps with a very up to date PDB)? I wonder how it scores on Private LB.",
      "votes": null
    },
    {
      "id": "3297374",
      "postDate": "10/03/2025 02:30:17",
      "content": "<p>Thank you Rhiju!</p>\n<p>Unfortunately, we don’t have experiments without the CASP16 targets, as we were limited by computational resources.</p>\n<p>The RNA-only trained model performed well on the private LB, so I believe the CASP16 targets were not what boosted the score. Instead, the improvement likely came from fitting to the RNA-only targets (excluding those labeled as “complex”).<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8111559%2Fff9dfe3c79a024a4877087946c20f4ff%2F1234.png?generation=1759458540965027&amp;alt=media\" alt=\"![\"></p>",
      "rawMarkdown": "Thank you Rhiju!\n\nUnfortunately, we don’t have experiments without the CASP16 targets, as we were limited by computational resources.\n\nThe RNA-only trained model performed well on the private LB, so I believe the CASP16 targets were not what boosted the score. Instead, the improvement likely came from fitting to the RNA-only targets (excluding those labeled as “complex”).![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8111559%2Fff9dfe3c79a024a4877087946c20f4ff%2F1234.png?generation=1759458540965027&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3295653,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "09/29/2025 09:47:41",
      "content": "<p>congradualation.<br>\nHow do you ensemble the outputs of different models?<br>\nBecause each model outputs coordinates.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3295907,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "09/29/2025 18:44:45",
          "content": "<p><a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> The competition's <code>submission.csv</code> expects competitors to predict 5 different \"candidates\" for each input sequence. Most competitors ensembled models in 3 ways:</p>\n<ol>\n<li><strong>Length based:</strong> check the length of the input sequence, if L &gt; threshold use model A if L &lt;= threshold use model B. Example: if the sequence is shorter than 400 nucleotides use DRFold2 and if its longer use Protenix.</li>\n<li><strong>Vertical ensemble:</strong> if you have 5 possible candidates use X/5 for model A and Y/5 for model B. Example: use 2 candidates for Protenix and 3 for DRFold2.</li>\n<li><strong>Align and aggregate:</strong> align all sequences to a reference coordinate frame and then aggregate (using the mean usually). I have not seen many succesfull attempts on this approach though.</li>\n</ol>\n<p>Hope this answers your question</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3297305,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "10/02/2025 20:16:00",
      "content": "<p><a href=\"https://www.kaggle.com/yekim102\" target=\"_blank\">@yekim102</a> <a href=\"https://www.kaggle.com/ryankim99\" target=\"_blank\">@ryankim99</a> <a href=\"https://www.kaggle.com/ouiqdmw\" target=\"_blank\">@ouiqdmw</a> thanks for a beautiful writeup and sharing your insights as early as June! Congrats on having the top AF3-like solution.</p>\n<p>A question for you. In earlier discussions, it seemed like use of CASP16 blind predictions to train Protenix led to a big increase in accuracy on public LB:</p>\n<blockquote>\n  <p>We included Casp 16 data so that bumps up the public score.<br>\n  If we don't use the casp 16 data, we get about 0.42.<br>\n  (from <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/584487\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/584487</a>) </p>\n</blockquote>\n<p>But actually… our Private LB does not include any CASP16 targets, and your model still did super well!  </p>\n<p>So I wonder if CASP16 predictions were so necessary.</p>\n<p>Do you have a checkpoint somewhere where you trained <em>without</em> CASP16 targets (but perhaps with a very up to date PDB)? I wonder how it scores on Private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3297374,
          "author_name": "yekim102",
          "author_url": "",
          "post_date": "10/03/2025 02:30:17",
          "content": "<p>Thank you Rhiju!</p>\n<p>Unfortunately, we don’t have experiments without the CASP16 targets, as we were limited by computational resources.</p>\n<p>The RNA-only trained model performed well on the private LB, so I believe the CASP16 targets were not what boosted the score. Instead, the improvement likely came from fitting to the RNA-only targets (excluding those labeled as “complex”).<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8111559%2Fff9dfe3c79a024a4877087946c20f4ff%2F1234.png?generation=1759458540965027&amp;alt=media\" alt=\"![\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3295548": "# Our 3rd Place Solution – Stanford RNA 3D Folding\n\nFirst of all, thanks to Kaggle and the organizers for putting together such a challenging and interesting competition.  \n\nWe finished **3rd on the public leaderboard** before the unseen data release, and also **3rd on the private leaderboard**.  \n\nEarlier we shared a short overview of our solution. This post goes into more detail.  \n\n---\n\n## Summary of Our Solution\n- Ensemble of **DRfold2**, **Protenix**, and **Boltz-1**  \n- Fine-tuned **Protenix** with newly released RNA datasets  \n\n---\n\n## Detailed Solution\n\n### Dataset\n\n- **rMSA**  \n  We generated our own rMSA data using the [official rMSA code](https://github.com/pylelab/rMSA) provided by the hosts.  \n  - The v2 rMSA released by the organizers did not cover the full set of recently published data, so we had to build our own.  \n  - This process took about **14 days**, even with multiprocessing and several servers.  \n  - Our rMSA data is available here: [Google Drive](https://drive.google.com/drive/folders/15bhXWsR6QuDQo6U4j8Ii8-OE4B4SmC7i?usp=drive_link).  \n\n- **Training dataset**  \n  We tested two versions:  \n  1. **Full RNA dataset** – included all recently available data, such as [CASP16](https://www.kaggle.com/datasets/tant64/casp16) uploaded by @tant64. To check quality, we compared some of the labels with the [CASP16 GitLab repository](https://gitlab.com/arneelof/CASP16-predictions/-/tree/main/rna_results?ref_type=heads).  \n  2. **RNA-only dataset** – excluded complexes (RNA bound to proteins/DNA), following the clarification from the hosts [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/575745).  \n\nBoth versions of the training data and labels are available here: [Google Drive](https://drive.google.com/drive/folders/1XKYzk2oCcHPt6DB_wLNYL-7w3s1R793n).  \n\n---\n\n### Models\n\nWe built our ensemble using models from these official repositories:  \n- [DRfold2](https://github.com/leeyang/DRfold2)  \n- [Protenix](https://github.com/bytedance/Protenix)  \n- [Boltz](https://github.com/jwohlwend/boltz)  \n\nWe initially explored new architectures, but given the time and resource limits, it was more effective to focus on fine-tuning and combining existing models.  \n\n---\n\n#### DRfold2\n- Performed best on sequences **<400 nt** with v1 data.  \n- Running the full pipeline with optimization, clustering, and all 80 released checkpoints was too slow.  \n- We found that **energy selection + Arena** gave most of the gain, while the other steps added little to the TM-score but took significant runtime.  \n- So our DRfold2 setup used **energy selection + Arena** only for the post-processing.  \n\n---\n\n#### Protenix\n- Trained on an **NVIDIA GH200 (96GB VRAM)**.  \n- Limited to sequences **<800 nt** due to memory issues.  \n- Used fine-tuning code from @lihaoweicvch ([discussion link](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495)).  \n- Fine-tuning **without rMSA** did not help, so we trained **with rMSA**.  \n- Only modified `max_steps`; a full dataset run took about a day.  \n- Early checkpoints (2–3 cycles) performed worse; performance improved with longer fine-tuning.  \n- We also tested **multiple Protenix outputs → DRfold energy selection → Arena**, but it did not improve results.  \n- For inference, we modified [@geraseva’s code](https://www.kaggle.com/code/geraseva/protenix) to include rMSA.  \n\n---\n\n#### Boltz-1\n- While not the strongest model on its own, including **one Boltz prediction** improved ensemble diversity.  \n- This helped because the scoring metric favored picking the best out of multiple diverse outputs.  \n- We used [@youhanlee’s inference notebook](https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission).  \n\n---\n\n### Final Submissions & Code\n\nWe are sharing our original submission notebooks. Some commented-out code remains, which shows parts of our experiment history.  \n\n1. **[DRfold2 + Protenix](https://www.kaggle.com/code/yekim102/drfold4-pro-msa-pro-msa-base-rna)**  \n   - Public LB: **0.60338** | Private LB: **0.52787**  \n   - <400 nt: 3 × DRfold2 + 2 × Protenix (RNA-only, with MSA)  \n   - >400 nt: 2 × Protenix (RNA-only) + 2 × Protenix (All) + Protenix baseline  \n\n2. **[Protenix + Boltz](https://www.kaggle.com/code/yekim102/fin-pro-all-1-rna-2-base1-boltz1)**  \n   - Public LB: **0.61253** | Private LB: **0.54312**  \n   - 2 × Protenix (RNA-only) + Protenix (All) + Protenix baseline + Boltz baseline  \n\n---\n\n## Acknowledgements\n\nMany thanks to the community members who shared resources and insights:  \n- **@hengck23** – for extensive discussions and guidance.  \n- **@geraseva** – for early Protenix inference code.  \n- **@lihaoweicvch** – for Protenix fine-tuning code.  \n- **@youhanlee** – for updated Boltz inference code.  \n\nFinally, thanks to **Eigen Company** for providing resources for this competition.",
    "3295653": "congradualation.\nHow do you ensemble the outputs of different models?\nBecause each model outputs coordinates.",
    "3295907": "garybios The competition's `submission.csv` expects competitors to predict 5 different \"candidates\" for each input sequence. Most competitors ensembled models in 3 ways:\n1.  **Length based:** check the length of the input sequence, if L > threshold use model A if L <= threshold use model B. Example: if the sequence is shorter than 400 nucleotides use DRFold2 and if its longer use Protenix.\n2.  **Vertical ensemble:** if you have 5 possible candidates use X/5 for model A and Y/5 for model B. Example: use 2 candidates for Protenix and 3 for DRFold2.\n3. **Align and aggregate:** align all sequences to a reference coordinate frame and then aggregate (using the mean usually). I have not seen many succesfull attempts on this approach though.\n\nHope this answers your question",
    "3297305": "yekim102 @ryankim99 @ouiqdmw thanks for a beautiful writeup and sharing your insights as early as June! Congrats on having the top AF3-like solution.\n\nA question for you. In earlier discussions, it seemed like use of CASP16 blind predictions to train Protenix led to a big increase in accuracy on public LB:\n\n>We included Casp 16 data so that bumps up the public score.\n> If we don't use the casp 16 data, we get about 0.42.\n(from https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/584487) \n\nBut actually... our Private LB does not include any CASP16 targets, and your model still did super well!  \n\nSo I wonder if CASP16 predictions were so necessary.\n\nDo you have a checkpoint somewhere where you trained *without* CASP16 targets (but perhaps with a very up to date PDB)? I wonder how it scores on Private LB.",
    "3297374": "Thank you Rhiju!\n\nUnfortunately, we don’t have experiments without the CASP16 targets, as we were limited by computational resources.\n\nThe RNA-only trained model performed well on the private LB, so I believe the CASP16 targets were not what boosted the score. Instead, the improvement likely came from fitting to the RNA-only targets (excluding those labeled as “complex”).![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8111559%2Fff9dfe3c79a024a4877087946c20f4ff%2F1234.png?generation=1759458540965027&alt=media)"
  },
  "source": "meta"
}