{
  "id": 609515,
  "title": "10th place solution (How we ensemble)",
  "url": "/competitions/stanford-rna-3d-folding/writeups/10th-place-solution-how-we-ensemble",
  "author_name": "",
  "post_date": "2025-09-27T10:50:42.357Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I’d like to share our team’s solution for the Stanford RNA 3D Folding competition on Kaggle. Our approach focused on combining three state-of-the-art deep learning models — Protenix, DRFold2, and trRosetta2 — into a robust ensemble pipeline for RNA tertiary structure prediction.</p>\n<h1><strong>Model Structure and Workflow</strong></h1>\n<p><strong>Protenix Implementation</strong></p>\n<p>Diffusion Sampling Parameters: N_sample=10, N_step=150, N_cycle=8<br>\nCheckpoint: model_v0.2.0.pt &lt;- The most recent pretrained model weights available on the Protenix GitHub.<br>\nConfidence Scoring System: Use in the exact order sorted by the confidence score provided by Protenix.<br>\nHyperparameter optimization did not make a huge difference in my experiments.<br>\nI was unable to manage the overfitting problem during Protenix fine-tuning, so I used the pretrained weights for direct prediction. <br>\nOutputs: top 5 PDBs per sample</p>\n<p><strong>DRFold2 Implementation</strong></p>\n<p>FASTA Generation: converts each test sequence to FASTA file format<br>\nFiltering: only sequences &lt; 300 nt processed (over 300nt cause timeout failure)\nModel Ranking per sample: Reads sel_0 score files -&gt; Selects top 5 models by energy score -&gt; Uses <a href=\"https://github.com/pylelab/Arena\" target=\"_blank\">Arena</a> for PDB refinement!!<br>\nOutputs: top 5 refined PDBs for top predictions (restricted to under 300nt)</p>\n<p><strong>trRosetta2 Implementation(it was hidden in GitHub, not officially published in trRosetta2 article)</strong></p>\n<p>MSA-based prediction for higher accuracy<br>\nModel Parameters: nrows=500, refine_steps=0<br>\nFiltering: approximately ≤500–600 nt, as far as I recall, though not precise.<br>\nOutputs: top 5 refined PDBs for top predictions (restricted to under 500-600nt)</p>\n<p><strong>Ensemble Integration</strong></p>\n<p>Coordinate Extraction: parses PDBs to get C1’ atom coordinates</p>\n<p>Merging Strategy:</p>\n<p>Protenix: coordinates 1–5 (used as background coordinates)<br>\nDRFold2: coordinates 3,4 (overwrite the background coordinates as available as possible)<br>\ntrRosetta2: coordinates 5 (overwrite the background coordinates as available as possible)</p>\n<p>So, final submission.csv is  </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>1</th>\n<th>2</th>\n<th>3</th>\n<th>4</th>\n<th>5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0~300nt</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>DRFold2</td>\n<td>DRFold2</td>\n<td>trRosetta2</td>\n</tr>\n<tr>\n<td>300~600nt</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>trRosetta2</td>\n</tr>\n<tr>\n<td>600nt~</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n</tr>\n</tbody>\n</table>\n<h1> </h1>\n<h2>CONCLUSION</h2>\n<p>Ensembling was very effective with three models. However, when I added a fourth model (Boltz), the performance actually decreased. Since I was unable to fine-tune Protenix, I believe others may have succeeded, and I look forward to seeing their solutions.</p>\n<p>Last but not least, many thanks to my teammates and Competition hosts!!</p>",
  "messages": [
    {
      "id": "3295003",
      "postDate": "09/27/2025 10:48:41",
      "content": "<p>Hi everyone,</p>\n<p>I’d like to share our team’s solution for the Stanford RNA 3D Folding competition on Kaggle. Our approach focused on combining three state-of-the-art deep learning models — Protenix, DRFold2, and trRosetta2 — into a robust ensemble pipeline for RNA tertiary structure prediction.</p>\n<h1><strong>Model Structure and Workflow</strong></h1>\n<p><strong>Protenix Implementation</strong></p>\n<p>Diffusion Sampling Parameters: N_sample=10, N_step=150, N_cycle=8<br>\nCheckpoint: model_v0.2.0.pt &lt;- The most recent pretrained model weights available on the Protenix GitHub.<br>\nConfidence Scoring System: Use in the exact order sorted by the confidence score provided by Protenix.<br>\nHyperparameter optimization did not make a huge difference in my experiments.<br>\nI was unable to manage the overfitting problem during Protenix fine-tuning, so I used the pretrained weights for direct prediction. <br>\nOutputs: top 5 PDBs per sample</p>\n<p><strong>DRFold2 Implementation</strong></p>\n<p>FASTA Generation: converts each test sequence to FASTA file format<br>\nFiltering: only sequences &lt; 300 nt processed (over 300nt cause timeout failure)\nModel Ranking per sample: Reads sel_0 score files -&gt; Selects top 5 models by energy score -&gt; Uses <a href=\"https://github.com/pylelab/Arena\" target=\"_blank\">Arena</a> for PDB refinement!!<br>\nOutputs: top 5 refined PDBs for top predictions (restricted to under 300nt)</p>\n<p><strong>trRosetta2 Implementation(it was hidden in GitHub, not officially published in trRosetta2 article)</strong></p>\n<p>MSA-based prediction for higher accuracy<br>\nModel Parameters: nrows=500, refine_steps=0<br>\nFiltering: approximately ≤500–600 nt, as far as I recall, though not precise.<br>\nOutputs: top 5 refined PDBs for top predictions (restricted to under 500-600nt)</p>\n<p><strong>Ensemble Integration</strong></p>\n<p>Coordinate Extraction: parses PDBs to get C1’ atom coordinates</p>\n<p>Merging Strategy:</p>\n<p>Protenix: coordinates 1–5 (used as background coordinates)<br>\nDRFold2: coordinates 3,4 (overwrite the background coordinates as available as possible)<br>\ntrRosetta2: coordinates 5 (overwrite the background coordinates as available as possible)</p>\n<p>So, final submission.csv is  </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>1</th>\n<th>2</th>\n<th>3</th>\n<th>4</th>\n<th>5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0~300nt</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>DRFold2</td>\n<td>DRFold2</td>\n<td>trRosetta2</td>\n</tr>\n<tr>\n<td>300~600nt</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>trRosetta2</td>\n</tr>\n<tr>\n<td>600nt~</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n<td>protenix</td>\n</tr>\n</tbody>\n</table>\n<h1> </h1>\n<h2>CONCLUSION</h2>\n<p>Ensembling was very effective with three models. However, when I added a fourth model (Boltz), the performance actually decreased. Since I was unable to fine-tune Protenix, I believe others may have succeeded, and I look forward to seeing their solutions.</p>\n<p>Last but not least, many thanks to my teammates and Competition hosts!!</p>",
      "rawMarkdown": "Hi everyone,\n\nI’d like to share our team’s solution for the Stanford RNA 3D Folding competition on Kaggle. Our approach focused on combining three state-of-the-art deep learning models — Protenix, DRFold2, and trRosetta2 — into a robust ensemble pipeline for RNA tertiary structure prediction.\n\n# **Model Structure and Workflow**\n\n**Protenix Implementation**\n\nDiffusion Sampling Parameters: N_sample=10, N_step=150, N_cycle=8\nCheckpoint: model_v0.2.0.pt <- The most recent pretrained model weights available on the Protenix GitHub.\nConfidence Scoring System: Use in the exact order sorted by the confidence score provided by Protenix.\nHyperparameter optimization did not make a huge difference in my experiments.\nI was unable to manage the overfitting problem during Protenix fine-tuning, so I used the pretrained weights for direct prediction. \nOutputs: top 5 PDBs per sample\n\n**DRFold2 Implementation**\n\nFASTA Generation: converts each test sequence to FASTA file format\nFiltering: only sequences < 300 nt processed (over 300nt cause timeout failure)\nModel Ranking per sample: Reads sel_0 score files -> Selects top 5 models by energy score -> Uses [Arena](https://github.com/pylelab/Arena) for PDB refinement!!\nOutputs: top 5 refined PDBs for top predictions (restricted to under 300nt)\n\n**trRosetta2 Implementation(it was hidden in GitHub, not officially published in trRosetta2 article)**\n\nMSA-based prediction for higher accuracy\nModel Parameters: nrows=500, refine_steps=0\nFiltering: approximately ≤500–600 nt, as far as I recall, though not precise.\nOutputs: top 5 refined PDBs for top predictions (restricted to under 500-600nt)\n\n\n**Ensemble Integration**\n\nCoordinate Extraction: parses PDBs to get C1’ atom coordinates\n\nMerging Strategy:\n\nProtenix: coordinates 1–5 (used as background coordinates)\nDRFold2: coordinates 3,4 (overwrite the background coordinates as available as possible)\ntrRosetta2: coordinates 5 (overwrite the background coordinates as available as possible)\n\nSo, final submission.csv is  \n|   | 1 | 2 | 3 | 4 | 5 |  \n| --- | --- | --- | --- | --- |\n| 0~300nt | protenix | protenix | DRFold2 | DRFold2 | trRosetta2 |\n| 300~600nt | protenix | protenix | protenix | protenix | trRosetta2 |\n| 600nt~ | protenix | protenix | protenix | protenix | protenix |\n# \n## CONCLUSION\n\nEnsembling was very effective with three models. However, when I added a fourth model (Boltz), the performance actually decreased. Since I was unable to fine-tune Protenix, I believe others may have succeeded, and I look forward to seeing their solutions.\n\nLast but not least, many thanks to my teammates and Competition hosts!!",
      "votes": null
    },
    {
      "id": "3295054",
      "postDate": "09/27/2025 14:22:30",
      "content": "<p>For some of the key steps, I also used a similar approach for ensemble, but I have my own fine-tuned model. Due to the nature of diffusion models, the results inherently involve a certain degree of randomness, so the order of fusion (from 1 to 5) doesn’t really make much difference. I employed a fine-tuned version of Boltz, which performed exceptionally well on the public leaderboard. It’s highly likely that the performance drop on the private leaderboard is partly due to overfitting.</p>",
      "rawMarkdown": "For some of the key steps, I also used a similar approach for ensemble, but I have my own fine-tuned model. Due to the nature of diffusion models, the results inherently involve a certain degree of randomness, so the order of fusion (from 1 to 5) doesn’t really make much difference. I employed a fine-tuned version of Boltz, which performed exceptionally well on the public leaderboard. It’s highly likely that the performance drop on the private leaderboard is partly due to overfitting.",
      "votes": null
    },
    {
      "id": "3295215",
      "postDate": "09/28/2025 03:32:29",
      "content": "<p>I totally agree! Since diffusion-based models have inherent variability across samples, the order doesn’t really make a big difference in the end. It’s really interesting that your fine-tuned Boltz performed so well on the public leaderboard. In our case, adding Boltz actually hurt performance, which I think may come down to differences in data splits or the level of overfitting. As you pointed out, the performance gap on the private leaderboard seems to reflect that pretty clearly.</p>",
      "rawMarkdown": "I totally agree! Since diffusion-based models have inherent variability across samples, the order doesn’t really make a big difference in the end. It’s really interesting that your fine-tuned Boltz performed so well on the public leaderboard. In our case, adding Boltz actually hurt performance, which I think may come down to differences in data splits or the level of overfitting. As you pointed out, the performance gap on the private leaderboard seems to reflect that pretty clearly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3295054,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "09/27/2025 14:22:30",
      "content": "<p>For some of the key steps, I also used a similar approach for ensemble, but I have my own fine-tuned model. Due to the nature of diffusion models, the results inherently involve a certain degree of randomness, so the order of fusion (from 1 to 5) doesn’t really make much difference. I employed a fine-tuned version of Boltz, which performed exceptionally well on the public leaderboard. It’s highly likely that the performance drop on the private leaderboard is partly due to overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3295215,
          "author_name": "yeongjinlee4",
          "author_url": "",
          "post_date": "09/28/2025 03:32:29",
          "content": "<p>I totally agree! Since diffusion-based models have inherent variability across samples, the order doesn’t really make a big difference in the end. It’s really interesting that your fine-tuned Boltz performed so well on the public leaderboard. In our case, adding Boltz actually hurt performance, which I think may come down to differences in data splits or the level of overfitting. As you pointed out, the performance gap on the private leaderboard seems to reflect that pretty clearly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3295003": "Hi everyone,\n\nI’d like to share our team’s solution for the Stanford RNA 3D Folding competition on Kaggle. Our approach focused on combining three state-of-the-art deep learning models — Protenix, DRFold2, and trRosetta2 — into a robust ensemble pipeline for RNA tertiary structure prediction.\n\n# **Model Structure and Workflow**\n\n**Protenix Implementation**\n\nDiffusion Sampling Parameters: N_sample=10, N_step=150, N_cycle=8\nCheckpoint: model_v0.2.0.pt <- The most recent pretrained model weights available on the Protenix GitHub.\nConfidence Scoring System: Use in the exact order sorted by the confidence score provided by Protenix.\nHyperparameter optimization did not make a huge difference in my experiments.\nI was unable to manage the overfitting problem during Protenix fine-tuning, so I used the pretrained weights for direct prediction. \nOutputs: top 5 PDBs per sample\n\n**DRFold2 Implementation**\n\nFASTA Generation: converts each test sequence to FASTA file format\nFiltering: only sequences < 300 nt processed (over 300nt cause timeout failure)\nModel Ranking per sample: Reads sel_0 score files -> Selects top 5 models by energy score -> Uses [Arena](https://github.com/pylelab/Arena) for PDB refinement!!\nOutputs: top 5 refined PDBs for top predictions (restricted to under 300nt)\n\n**trRosetta2 Implementation(it was hidden in GitHub, not officially published in trRosetta2 article)**\n\nMSA-based prediction for higher accuracy\nModel Parameters: nrows=500, refine_steps=0\nFiltering: approximately ≤500–600 nt, as far as I recall, though not precise.\nOutputs: top 5 refined PDBs for top predictions (restricted to under 500-600nt)\n\n\n**Ensemble Integration**\n\nCoordinate Extraction: parses PDBs to get C1’ atom coordinates\n\nMerging Strategy:\n\nProtenix: coordinates 1–5 (used as background coordinates)\nDRFold2: coordinates 3,4 (overwrite the background coordinates as available as possible)\ntrRosetta2: coordinates 5 (overwrite the background coordinates as available as possible)\n\nSo, final submission.csv is  \n|   | 1 | 2 | 3 | 4 | 5 |  \n| --- | --- | --- | --- | --- |\n| 0~300nt | protenix | protenix | DRFold2 | DRFold2 | trRosetta2 |\n| 300~600nt | protenix | protenix | protenix | protenix | trRosetta2 |\n| 600nt~ | protenix | protenix | protenix | protenix | protenix |\n# \n## CONCLUSION\n\nEnsembling was very effective with three models. However, when I added a fourth model (Boltz), the performance actually decreased. Since I was unable to fine-tune Protenix, I believe others may have succeeded, and I look forward to seeing their solutions.\n\nLast but not least, many thanks to my teammates and Competition hosts!!",
    "3295054": "For some of the key steps, I also used a similar approach for ensemble, but I have my own fine-tuned model. Due to the nature of diffusion models, the results inherently involve a certain degree of randomness, so the order of fusion (from 1 to 5) doesn’t really make much difference. I employed a fine-tuned version of Boltz, which performed exceptionally well on the public leaderboard. It’s highly likely that the performance drop on the private leaderboard is partly due to overfitting.",
    "3295215": "I totally agree! Since diffusion-based models have inherent variability across samples, the order doesn’t really make a big difference in the end. It’s really interesting that your fine-tuned Boltz performed so well on the public leaderboard. In our case, adding Boltz actually hurt performance, which I think may come down to differences in data splits or the level of overfitting. As you pointed out, the performance gap on the private leaderboard seems to reflect that pretty clearly."
  },
  "source": "meta"
}