{
  "id": 609775,
  "title": "5th Place Solution",
  "url": "/competitions/stanford-rna-3d-folding/discussion/609775",
  "author_name": "koooeo",
  "post_date": "2025-09-29T15:36:01.762000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Stanford RNA 3D Folding 5th Place Solution</h1>\n<p>First of all, I never thought we would be able to finish with such a great result, placing 5th.</p>\n<p>I'd like to thank hengck23 for his excellent discussions throughout the competition, lihaoweicvch for sharing his fine-tuning code, and everyone else who shared their insights in the discussion thread. And most importantly, I'd like to thank the hosts and Kaggle for providing us with this fantastic competition.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/overview</a></li>\n<li>Data context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>All submitted models are ensembles of Protenix fine-tuned models.<br>\nFine-tuning was carried out along two main lines.</p>\n<ol>\n<li>Curriculum learning based on Protenix's technical report[1].</li>\n<li>Fine-tuning using clustered RNA.</li>\n</ol>\n<h1>Detailed solution</h1>\n<ul>\n<li><p>Data used<br>\nData in train_sequences.csv and train_sequences.v2.csv that contain MSA, and their label data.</p></li>\n<li><p>Fine-tuning with RNA-MSA support<br>\nWe fine-tuned a forked version of the Protenix model[2], modified to accept multiple sequence alignments (MSA) as input.</p></li>\n<li><p>Cut-off date filtering<br>\nWe applied two cut-off dates to curate the training data and reduce potential data leakage:<br>\n<em>2021-09-30</em>: To filter out any data that might have been included in Protenix’s original pretraining, based on its technical documentation[1].<br>\n<em>2024-05-16</em>: To ensure only data excluded from the public leaderboard was used, following host guidance[3].</p></li>\n</ul>\n<h2>koooeo Part</h2>\n<ul>\n<li><p>Sequence length-based partitioning<br>\nTraining data was divided into groups by sequence length to enable more stable and tailored model training per group.</p></li>\n<li><p>Curriculum-style fine-tuning<br>\nWe adopted a step-wise fine-tuning strategy, first training on shorter sequences and progressively moving to longer ones.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F5ee1b38d4e633a1d3547746cdea7399d%2F2025-09-30%200.28.45.png?generation=1759159747381151&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<h2>Shun Kuraishi Part</h2>\n<ul>\n<li><p>Clustering and representative fine-tuning<br>\nTo improve training efficiency and structural diversity, we clustered the training data into 18 groups based on sequence and structural features.From each cluster, we selected 2-4 representative sequences for further fine-tuning.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F4b8256b56f751bfef9069a2d4e844d9b%2F2025-09-30%200.27.20.png?generation=1759159688356148&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Sequence length gap pattern learning<br>\nAfter segmenting the data into sequence-length bands, we trained a dedicated model for each band (e.g., 0–300, 300–500, 500–700, etc) so that each model specialized on its assigned range.</p></li>\n<li><p>Specialized model assignment by sequence range<br>\nWe trained separate models for different sequence length bands (e.g., 0–300, 300–500, 500+),and at inference time, predictions were made using the model specialized for each input's length category.</p></li>\n<li><p>Ensemble weighting based on public LB performance<br>\nWe found that the model trained on the 208–300 range performed particularly well on the public leaderboard.During inference, we performed a 2:3 weighted ensemble between this strong model and the model assigned to each sequence length band.</p></li>\n</ul>\n<h1>What went well</h1>\n<ul>\n<li>We believe that simply passing the data with MSA directly into the model did not improve performance, but dividing the data by sequence length and by clustering may have enabled more effective training. By narrowing down the data, we think it may have helped correct biases in the distribution of sequence lengths and clusters.</li>\n</ul>\n<h1>What didn't work</h1>\n<ul>\n<li>We were unable to find a method to select the optimal structure among the outputs of multiple models.<br>\nIn the end, we adopted an approach of assigning the outputs of different models to five predictions for each target.</li>\n<li>We could not successfully evaluate which model was optimal.<br>\nAlthough we evaluated model performance using data that was not used for training, we found no correlation between those local scores and the public LB scores. Likewise, there was no correlation between the public LB scores and the local scores. I am curious about how others evaluated their models.</li>\n</ul>\n<h1>Sources</h1>\n<p>[1] <a href=\"url\" target=\"_blank\">https://github.com/bytedance/Protenix/blob/main/Protenix_Technical_Report.pdf</a><br>\n[2] <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495</a><br>\n[3] <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/572096#3173463</a></p>",
  "messages": [
    {
      "id": 3295800,
      "postDate": "2025-09-29T15:36:01.763Z",
      "content": "<h1>Stanford RNA 3D Folding 5th Place Solution</h1>\n<p>First of all, I never thought we would be able to finish with such a great result, placing 5th.</p>\n<p>I'd like to thank hengck23 for his excellent discussions throughout the competition, lihaoweicvch for sharing his fine-tuning code, and everyone else who shared their insights in the discussion thread. And most importantly, I'd like to thank the hosts and Kaggle for providing us with this fantastic competition.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/overview</a></li>\n<li>Data context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>All submitted models are ensembles of Protenix fine-tuned models.<br>\nFine-tuning was carried out along two main lines.</p>\n<ol>\n<li>Curriculum learning based on Protenix's technical report[1].</li>\n<li>Fine-tuning using clustered RNA.</li>\n</ol>\n<h1>Detailed solution</h1>\n<ul>\n<li><p>Data used<br>\nData in train_sequences.csv and train_sequences.v2.csv that contain MSA, and their label data.</p></li>\n<li><p>Fine-tuning with RNA-MSA support<br>\nWe fine-tuned a forked version of the Protenix model[2], modified to accept multiple sequence alignments (MSA) as input.</p></li>\n<li><p>Cut-off date filtering<br>\nWe applied two cut-off dates to curate the training data and reduce potential data leakage:<br>\n<em>2021-09-30</em>: To filter out any data that might have been included in Protenix’s original pretraining, based on its technical documentation[1].<br>\n<em>2024-05-16</em>: To ensure only data excluded from the public leaderboard was used, following host guidance[3].</p></li>\n</ul>\n<h2>koooeo Part</h2>\n<ul>\n<li><p>Sequence length-based partitioning<br>\nTraining data was divided into groups by sequence length to enable more stable and tailored model training per group.</p></li>\n<li><p>Curriculum-style fine-tuning<br>\nWe adopted a step-wise fine-tuning strategy, first training on shorter sequences and progressively moving to longer ones.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F5ee1b38d4e633a1d3547746cdea7399d%2F2025-09-30%200.28.45.png?generation=1759159747381151&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<h2>Shun Kuraishi Part</h2>\n<ul>\n<li><p>Clustering and representative fine-tuning<br>\nTo improve training efficiency and structural diversity, we clustered the training data into 18 groups based on sequence and structural features.From each cluster, we selected 2-4 representative sequences for further fine-tuning.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F4b8256b56f751bfef9069a2d4e844d9b%2F2025-09-30%200.27.20.png?generation=1759159688356148&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Sequence length gap pattern learning<br>\nAfter segmenting the data into sequence-length bands, we trained a dedicated model for each band (e.g., 0–300, 300–500, 500–700, etc) so that each model specialized on its assigned range.</p></li>\n<li><p>Specialized model assignment by sequence range<br>\nWe trained separate models for different sequence length bands (e.g., 0–300, 300–500, 500+),and at inference time, predictions were made using the model specialized for each input's length category.</p></li>\n<li><p>Ensemble weighting based on public LB performance<br>\nWe found that the model trained on the 208–300 range performed particularly well on the public leaderboard.During inference, we performed a 2:3 weighted ensemble between this strong model and the model assigned to each sequence length band.</p></li>\n</ul>\n<h1>What went well</h1>\n<ul>\n<li>We believe that simply passing the data with MSA directly into the model did not improve performance, but dividing the data by sequence length and by clustering may have enabled more effective training. By narrowing down the data, we think it may have helped correct biases in the distribution of sequence lengths and clusters.</li>\n</ul>\n<h1>What didn't work</h1>\n<ul>\n<li>We were unable to find a method to select the optimal structure among the outputs of multiple models.<br>\nIn the end, we adopted an approach of assigning the outputs of different models to five predictions for each target.</li>\n<li>We could not successfully evaluate which model was optimal.<br>\nAlthough we evaluated model performance using data that was not used for training, we found no correlation between those local scores and the public LB scores. Likewise, there was no correlation between the public LB scores and the local scores. I am curious about how others evaluated their models.</li>\n</ul>\n<h1>Sources</h1>\n<p>[1] <a href=\"url\" target=\"_blank\">https://github.com/bytedance/Protenix/blob/main/Protenix_Technical_Report.pdf</a><br>\n[2] <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495</a><br>\n[3] <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/572096#3173463</a></p>",
      "rawMarkdown": "# Stanford RNA 3D Folding 5th Place Solution\nFirst of all, I never thought we would be able to finish with such a great result, placing 5th.\n\nI'd like to thank hengck23 for his excellent discussions throughout the competition, lihaoweicvch for sharing his fine-tuning code, and everyone else who shared their insights in the discussion thread. And most importantly, I'd like to thank the hosts and Kaggle for providing us with this fantastic competition.\n\n# Context\n- Business context: [https://www.kaggle.com/competitions/stanford-rna-3d-folding/overview](url)\n- Data context: [https://www.kaggle.com/competitions/stanford-rna-3d-folding/data](url)\n\n# Overview of the approach\nAll submitted models are ensembles of Protenix fine-tuned models.\nFine-tuning was carried out along two main lines.\n1. Curriculum learning based on Protenix's technical report[1].\n2. Fine-tuning using clustered RNA.\n\n# Detailed solution\n- Data used\nData in train_sequences.csv and train_sequences.v2.csv that contain MSA, and their label data.\n\n- Fine-tuning with RNA-MSA support\nWe fine-tuned a forked version of the Protenix model[2], modified to accept multiple sequence alignments (MSA) as input.\n\n- Cut-off date filtering\nWe applied two cut-off dates to curate the training data and reduce potential data leakage:\n*2021-09-30*: To filter out any data that might have been included in Protenix’s original pretraining, based on its technical documentation[1].\n*2024-05-16*: To ensure only data excluded from the public leaderboard was used, following host guidance[3].\n\n## koooeo Part\n- Sequence length-based partitioning\nTraining data was divided into groups by sequence length to enable more stable and tailored model training per group.\n\n- Curriculum-style fine-tuning\nWe adopted a step-wise fine-tuning strategy, first training on shorter sequences and progressively moving to longer ones.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F5ee1b38d4e633a1d3547746cdea7399d%2F2025-09-30%200.28.45.png?generation=1759159747381151&alt=media)\n\n\n## Shun Kuraishi Part\n- Clustering and representative fine-tuning\nTo improve training efficiency and structural diversity, we clustered the training data into 18 groups based on sequence and structural features.From each cluster, we selected 2-4 representative sequences for further fine-tuning.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F4b8256b56f751bfef9069a2d4e844d9b%2F2025-09-30%200.27.20.png?generation=1759159688356148&alt=media)\n\n- Sequence length gap pattern learning\nAfter segmenting the data into sequence-length bands, we trained a dedicated model for each band (e.g., 0–300, 300–500, 500–700, etc) so that each model specialized on its assigned range.\n\n- Specialized model assignment by sequence range\nWe trained separate models for different sequence length bands (e.g., 0–300, 300–500, 500+),and at inference time, predictions were made using the model specialized for each input's length category.\n\n- Ensemble weighting based on public LB performance\nWe found that the model trained on the 208–300 range performed particularly well on the public leaderboard.During inference, we performed a 2:3 weighted ensemble between this strong model and the model assigned to each sequence length band.\n\n# What went well\n- We believe that simply passing the data with MSA directly into the model did not improve performance, but dividing the data by sequence length and by clustering may have enabled more effective training. By narrowing down the data, we think it may have helped correct biases in the distribution of sequence lengths and clusters.\n\n# What didn't work\n- We were unable to find a method to select the optimal structure among the outputs of multiple models.\nIn the end, we adopted an approach of assigning the outputs of different models to five predictions for each target.\n- We could not successfully evaluate which model was optimal.\nAlthough we evaluated model performance using data that was not used for training, we found no correlation between those local scores and the public LB scores. Likewise, there was no correlation between the public LB scores and the local scores. I am curious about how others evaluated their models.\n\n# Sources\n[1] [https://github.com/bytedance/Protenix/blob/main/Protenix_Technical_Report.pdf](url)\n[2] [https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495](url)\n[3] [https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/572096#3173463](url)",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3295800": "# Stanford RNA 3D Folding 5th Place Solution\nFirst of all, I never thought we would be able to finish with such a great result, placing 5th.\n\nI'd like to thank hengck23 for his excellent discussions throughout the competition, lihaoweicvch for sharing his fine-tuning code, and everyone else who shared their insights in the discussion thread. And most importantly, I'd like to thank the hosts and Kaggle for providing us with this fantastic competition.\n\n# Context\n- Business context: [https://www.kaggle.com/competitions/stanford-rna-3d-folding/overview](url)\n- Data context: [https://www.kaggle.com/competitions/stanford-rna-3d-folding/data](url)\n\n# Overview of the approach\nAll submitted models are ensembles of Protenix fine-tuned models.\nFine-tuning was carried out along two main lines.\n1. Curriculum learning based on Protenix's technical report[1].\n2. Fine-tuning using clustered RNA.\n\n# Detailed solution\n- Data used\nData in train_sequences.csv and train_sequences.v2.csv that contain MSA, and their label data.\n\n- Fine-tuning with RNA-MSA support\nWe fine-tuned a forked version of the Protenix model[2], modified to accept multiple sequence alignments (MSA) as input.\n\n- Cut-off date filtering\nWe applied two cut-off dates to curate the training data and reduce potential data leakage:\n*2021-09-30*: To filter out any data that might have been included in Protenix’s original pretraining, based on its technical documentation[1].\n*2024-05-16*: To ensure only data excluded from the public leaderboard was used, following host guidance[3].\n\n## koooeo Part\n- Sequence length-based partitioning\nTraining data was divided into groups by sequence length to enable more stable and tailored model training per group.\n\n- Curriculum-style fine-tuning\nWe adopted a step-wise fine-tuning strategy, first training on shorter sequences and progressively moving to longer ones.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F5ee1b38d4e633a1d3547746cdea7399d%2F2025-09-30%200.28.45.png?generation=1759159747381151&alt=media)\n\n\n## Shun Kuraishi Part\n- Clustering and representative fine-tuning\nTo improve training efficiency and structural diversity, we clustered the training data into 18 groups based on sequence and structural features.From each cluster, we selected 2-4 representative sequences for further fine-tuning.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17037086%2F4b8256b56f751bfef9069a2d4e844d9b%2F2025-09-30%200.27.20.png?generation=1759159688356148&alt=media)\n\n- Sequence length gap pattern learning\nAfter segmenting the data into sequence-length bands, we trained a dedicated model for each band (e.g., 0–300, 300–500, 500–700, etc) so that each model specialized on its assigned range.\n\n- Specialized model assignment by sequence range\nWe trained separate models for different sequence length bands (e.g., 0–300, 300–500, 500+),and at inference time, predictions were made using the model specialized for each input's length category.\n\n- Ensemble weighting based on public LB performance\nWe found that the model trained on the 208–300 range performed particularly well on the public leaderboard.During inference, we performed a 2:3 weighted ensemble between this strong model and the model assigned to each sequence length band.\n\n# What went well\n- We believe that simply passing the data with MSA directly into the model did not improve performance, but dividing the data by sequence length and by clustering may have enabled more effective training. By narrowing down the data, we think it may have helped correct biases in the distribution of sequence lengths and clusters.\n\n# What didn't work\n- We were unable to find a method to select the optimal structure among the outputs of multiple models.\nIn the end, we adopted an approach of assigning the outputs of different models to five predictions for each target.\n- We could not successfully evaluate which model was optimal.\nAlthough we evaluated model performance using data that was not used for training, we found no correlation between those local scores and the public LB scores. Likewise, there was no correlation between the public LB scores and the local scores. I am curious about how others evaluated their models.\n\n# Sources\n[1] [https://github.com/bytedance/Protenix/blob/main/Protenix_Technical_Report.pdf](url)\n[2] [https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495](url)\n[3] [https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/572096#3173463](url)"
  }
}