{
  "id": 609822,
  "title": "25th Place Solution",
  "url": "/competitions/stanford-rna-3d-folding/discussion/609822",
  "author_name": "k2011yi",
  "post_date": "2025-09-29T20:55:47.869000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>My 25th Place Solution – Stanford RNA 3D Folding</h1>\n<p>First of all, I would like to thank Kaggle and the competition organizers for hosting such an exciting and challenging event.<br>\nI finished 29th on the public leaderboard before the release of the hidden test set and 25th on the final private leaderboard.<br>\nIn this post, I would like to briefly share my approach.</p>\n<hr>\n<h2>Overall Strategy</h2>\n<p>Since I participated solo with limited compute resources, time, and skills, I did not perform any fine-tuning.\nInstead, my strategy focused on <strong>how to select the most accurate structures</strong> from multiple predictions generated by existing models.  </p>\n<p>When the leaderboard was refreshed midway through the competition, my ranking remained relatively stable. This gave me confidence that this selection-based approach had some potential.</p>\n<hr>\n<h2>Models Used</h2>\n<p>For the final submissions, I relied on <strong>Proteinx</strong>, <strong>DRfold2</strong>, and <strong>Boltz</strong>.<br>\nHere is how I used each model:</p>\n<hr>\n<h3>DRfold2</h3>\n<p>After the first leaderboard refresh, I identified 20 prediction sets generated with the <code>cfg_97</code> configuration that showed stable and relatively high scores.<br>\nDue to inference time constraints, I reduced this to 12 sets for further processing.  </p>\n<p>For each set:</p>\n<ul>\n<li>I performed clustering on the predicted structures.</li>\n<li>Within each cluster, I selected the structure with the lowest energy score as the representative.  </li>\n<li>Then, I calculated TM-scores between representatives.  </li>\n</ul>\n<p>Finally, I included:</p>\n<ol>\n<li>The representative with the <strong>lowest energy</strong>  </li>\n<li>The representative <strong>most structurally different</strong> from it (lowest TM-score)  </li>\n</ol>\n<p>The goal was to maintain diversity in the final predictions rather than relying solely on the lowest-energy structure.</p>\n<hr>\n<h3>Boltz</h3>\n<p>For Boltz, I set <code>diffusion_samples=3</code> to generate three candidate structures.<br>\nAmong them, I selected the final prediction based on:</p>\n<ol>\n<li><strong>complex_plddt</strong> (highest priority)</li>\n<li><strong>confidence_score</strong> (tie-breaker when pLDDT was identical)  </li>\n</ol>\n<p>I remember that using this selection approach led to a noticeable improvement in my scores.  </p>\n<p>I repeated this process with two different seeds and selected the best structure from each.</p>\n<hr>\n<h3>Proteinx</h3>\n<p>For Proteinx, I set <code>--sample_diffusion.N_sample=3</code> to generate three structures.<br>\nSince pLDDT scores are available in the JSON output, I simply selected the <strong>highest-scoring</strong> structure for the final submission.</p>\n<hr>\n<h2>Closing Remarks</h2>\n<p>I am deeply grateful to the organizers for providing such a challenging and rewarding competition.<br>\nI also appreciate all the participants who shared valuable insights in the discussion forum.  </p>\n<p>This was the Kaggle competition where I invested the most time and effort so far, and I am very happy with the results.<br>\nI look forward to joining similar challenges in the future!</p>",
  "messages": [
    {
      "id": 3295951,
      "postDate": "2025-09-29T20:55:47.870Z",
      "content": "<h1>My 25th Place Solution – Stanford RNA 3D Folding</h1>\n<p>First of all, I would like to thank Kaggle and the competition organizers for hosting such an exciting and challenging event.<br>\nI finished 29th on the public leaderboard before the release of the hidden test set and 25th on the final private leaderboard.<br>\nIn this post, I would like to briefly share my approach.</p>\n<hr>\n<h2>Overall Strategy</h2>\n<p>Since I participated solo with limited compute resources, time, and skills, I did not perform any fine-tuning.\nInstead, my strategy focused on <strong>how to select the most accurate structures</strong> from multiple predictions generated by existing models.  </p>\n<p>When the leaderboard was refreshed midway through the competition, my ranking remained relatively stable. This gave me confidence that this selection-based approach had some potential.</p>\n<hr>\n<h2>Models Used</h2>\n<p>For the final submissions, I relied on <strong>Proteinx</strong>, <strong>DRfold2</strong>, and <strong>Boltz</strong>.<br>\nHere is how I used each model:</p>\n<hr>\n<h3>DRfold2</h3>\n<p>After the first leaderboard refresh, I identified 20 prediction sets generated with the <code>cfg_97</code> configuration that showed stable and relatively high scores.<br>\nDue to inference time constraints, I reduced this to 12 sets for further processing.  </p>\n<p>For each set:</p>\n<ul>\n<li>I performed clustering on the predicted structures.</li>\n<li>Within each cluster, I selected the structure with the lowest energy score as the representative.  </li>\n<li>Then, I calculated TM-scores between representatives.  </li>\n</ul>\n<p>Finally, I included:</p>\n<ol>\n<li>The representative with the <strong>lowest energy</strong>  </li>\n<li>The representative <strong>most structurally different</strong> from it (lowest TM-score)  </li>\n</ol>\n<p>The goal was to maintain diversity in the final predictions rather than relying solely on the lowest-energy structure.</p>\n<hr>\n<h3>Boltz</h3>\n<p>For Boltz, I set <code>diffusion_samples=3</code> to generate three candidate structures.<br>\nAmong them, I selected the final prediction based on:</p>\n<ol>\n<li><strong>complex_plddt</strong> (highest priority)</li>\n<li><strong>confidence_score</strong> (tie-breaker when pLDDT was identical)  </li>\n</ol>\n<p>I remember that using this selection approach led to a noticeable improvement in my scores.  </p>\n<p>I repeated this process with two different seeds and selected the best structure from each.</p>\n<hr>\n<h3>Proteinx</h3>\n<p>For Proteinx, I set <code>--sample_diffusion.N_sample=3</code> to generate three structures.<br>\nSince pLDDT scores are available in the JSON output, I simply selected the <strong>highest-scoring</strong> structure for the final submission.</p>\n<hr>\n<h2>Closing Remarks</h2>\n<p>I am deeply grateful to the organizers for providing such a challenging and rewarding competition.<br>\nI also appreciate all the participants who shared valuable insights in the discussion forum.  </p>\n<p>This was the Kaggle competition where I invested the most time and effort so far, and I am very happy with the results.<br>\nI look forward to joining similar challenges in the future!</p>",
      "rawMarkdown": "# My 25th Place Solution – Stanford RNA 3D Folding\n\nFirst of all, I would like to thank Kaggle and the competition organizers for hosting such an exciting and challenging event.  \nI finished 29th on the public leaderboard before the release of the hidden test set and 25th on the final private leaderboard.  \nIn this post, I would like to briefly share my approach.\n\n---\n\n## Overall Strategy\nSince I participated solo with limited compute resources, time, and skills, I did not perform any fine-tuning.\nInstead, my strategy focused on **how to select the most accurate structures** from multiple predictions generated by existing models.  \n\nWhen the leaderboard was refreshed midway through the competition, my ranking remained relatively stable. This gave me confidence that this selection-based approach had some potential.\n\n---\n\n## Models Used\nFor the final submissions, I relied on **Proteinx**, **DRfold2**, and **Boltz**.  \nHere is how I used each model:\n\n---\n\n### DRfold2\nAfter the first leaderboard refresh, I identified 20 prediction sets generated with the `cfg_97` configuration that showed stable and relatively high scores.  \nDue to inference time constraints, I reduced this to 12 sets for further processing.  \n\nFor each set:\n- I performed clustering on the predicted structures.\n- Within each cluster, I selected the structure with the lowest energy score as the representative.  \n- Then, I calculated TM-scores between representatives.  \n\nFinally, I included:\n1. The representative with the **lowest energy**  \n2. The representative **most structurally different** from it (lowest TM-score)  \n\nThe goal was to maintain diversity in the final predictions rather than relying solely on the lowest-energy structure.\n\n---\n\n### Boltz\nFor Boltz, I set `diffusion_samples=3` to generate three candidate structures.  \nAmong them, I selected the final prediction based on:\n1. **complex_plddt** (highest priority)\n2. **confidence_score** (tie-breaker when pLDDT was identical)  \n\nI remember that using this selection approach led to a noticeable improvement in my scores.  \n\nI repeated this process with two different seeds and selected the best structure from each.\n\n---\n\n### Proteinx\nFor Proteinx, I set `--sample_diffusion.N_sample=3` to generate three structures.  \nSince pLDDT scores are available in the JSON output, I simply selected the **highest-scoring** structure for the final submission.\n\n---\n\n## Closing Remarks\nI am deeply grateful to the organizers for providing such a challenging and rewarding competition.  \nI also appreciate all the participants who shared valuable insights in the discussion forum.  \n\nThis was the Kaggle competition where I invested the most time and effort so far, and I am very happy with the results.  \nI look forward to joining similar challenges in the future!",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3295951": "# My 25th Place Solution – Stanford RNA 3D Folding\n\nFirst of all, I would like to thank Kaggle and the competition organizers for hosting such an exciting and challenging event.  \nI finished 29th on the public leaderboard before the release of the hidden test set and 25th on the final private leaderboard.  \nIn this post, I would like to briefly share my approach.\n\n---\n\n## Overall Strategy\nSince I participated solo with limited compute resources, time, and skills, I did not perform any fine-tuning.\nInstead, my strategy focused on **how to select the most accurate structures** from multiple predictions generated by existing models.  \n\nWhen the leaderboard was refreshed midway through the competition, my ranking remained relatively stable. This gave me confidence that this selection-based approach had some potential.\n\n---\n\n## Models Used\nFor the final submissions, I relied on **Proteinx**, **DRfold2**, and **Boltz**.  \nHere is how I used each model:\n\n---\n\n### DRfold2\nAfter the first leaderboard refresh, I identified 20 prediction sets generated with the `cfg_97` configuration that showed stable and relatively high scores.  \nDue to inference time constraints, I reduced this to 12 sets for further processing.  \n\nFor each set:\n- I performed clustering on the predicted structures.\n- Within each cluster, I selected the structure with the lowest energy score as the representative.  \n- Then, I calculated TM-scores between representatives.  \n\nFinally, I included:\n1. The representative with the **lowest energy**  \n2. The representative **most structurally different** from it (lowest TM-score)  \n\nThe goal was to maintain diversity in the final predictions rather than relying solely on the lowest-energy structure.\n\n---\n\n### Boltz\nFor Boltz, I set `diffusion_samples=3` to generate three candidate structures.  \nAmong them, I selected the final prediction based on:\n1. **complex_plddt** (highest priority)\n2. **confidence_score** (tie-breaker when pLDDT was identical)  \n\nI remember that using this selection approach led to a noticeable improvement in my scores.  \n\nI repeated this process with two different seeds and selected the best structure from each.\n\n---\n\n### Proteinx\nFor Proteinx, I set `--sample_diffusion.N_sample=3` to generate three structures.  \nSince pLDDT scores are available in the JSON output, I simply selected the **highest-scoring** structure for the final submission.\n\n---\n\n## Closing Remarks\nI am deeply grateful to the organizers for providing such a challenging and rewarding competition.  \nI also appreciate all the participants who shared valuable insights in the discussion forum.  \n\nThis was the Kaggle competition where I invested the most time and effort so far, and I am very happy with the results.  \nI look forward to joining similar challenges in the future!"
  }
}