{
  "id": 610261,
  "title": "7th Place Solution: Ensemble of Two Protenix Models",
  "url": "/competitions/stanford-rna-3d-folding/discussion/610261",
  "author_name": "AyPy",
  "post_date": "2025-10-02T21:05:51.425000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>7th Place Solution: Ensemble of Fine-tuned Protenix</h1>\n<p>First of all, I would like to express my sincere gratitude to the competition host team and the Kaggle team for organizing this interesting competition. I am also deeply thankful to all the participants who continuously shared valuable insights and information. The three months spent tackling machine learning for biomolecular 3D structure prediction have been truly enjoyable and rewarding.</p>\n<h2>Summary</h2>\n<p>My solution consists of an ensemble of two types of fine-tuned <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a> (AlphaFold3 clone), <strong>with and without MSA</strong>. In the early phase of the competition I tried <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training\" target=\"_blank\">RibonanzaNet2_DDPM</a>, <a href=\"https://github.com/leeyang/DRfold2\" target=\"_blank\">DRfold2</a>, and <a href=\"https://github.com/kiharalab/NuFold\" target=\"_blank\">NuFold</a>, but Protenix gave the best results in my experiments, so I focused exclusively on Protenix in the latter half. Protenix finetuning was conducted with the code provided by <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> and <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/10th-place-solution-d4t4-team\" target=\"_blank\">d4t4 team</a>, with minor modifications.</p>\n<ul>\n<li><a href=\"https://github.com/lhwcv/Protenix-RNA-Kaggle\" target=\"_blank\">GitHub</a> (credit: <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>)</li>\n<li><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\" target=\"_blank\">Discussion post</a></li>\n</ul>\n<p>I tried two hill-climbing strategies and selected both two models as final submissions:</p>\n<ol>\n<li><strong>Public LB score focused</strong>: Mainly trained on CASP16 VFold predicted structures</li>\n<li><strong>Local score focused</strong>: Mainly trained on host-provided PDB data (Kaggle v1/v2 data)</li>\n</ol>\n<p>TM-Score results are as follows:</p>\n<h4>1. Public LB Focused Models</h4>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>wo-MSA</td>\n<td>0.43529</td>\n<td>0.39196</td>\n</tr>\n<tr>\n<td>with-MSA</td>\n<td>0.43205</td>\n<td>0.41014</td>\n</tr>\n<tr>\n<td>Two models ensemble<br>(<strong>Final Submission 1</strong>)</td>\n<td>0.46901</td>\n<td>0.43514</td>\n</tr>\n</tbody>\n</table>\n<h4>2. Local Score Focused Models</h4>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local Validation</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>wo-MSA</td>\n<td>0.5982</td>\n<td>0.43135</td>\n<td>0.47654</td>\n</tr>\n<tr>\n<td>with-MSA</td>\n<td>0.5943</td>\n<td>0.42501</td>\n<td>0.46793</td>\n</tr>\n<tr>\n<td>Two models ensemble<br>(<strong>Final Submission 2</strong>)</td>\n<td>0.6092</td>\n<td>0.43991</td>\n<td><strong>0.48511</strong><br>(private 7th place)</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Dataset</h2>\n<p>Kaggle v1/v2 data and CASP16 VFold prediction were used for model training.</p>\n<h4>Data Filtering</h4>\n<ul>\n<li><p><strong>Quality Criteria</strong><br><br>\nOnly data that met the following criteria were used:</p>\n<ul>\n<li>Consecutive nucleotide C1'-C1' distance ≤ 15 Å</li>\n<li>Resolution ≤ 15 Å</li>\n<li>Contains all four nucleotides (A, U, G, C)</li></ul></li>\n<li><p><strong>Deduplication</strong><br><br>\nFor sequences with multiple structure data, only one structure was selected for training according to the following rules:</p>\n<ul>\n<li>Keep the structure with the lowest resolution (Å)</li>\n<li>If resolution is NaN, select one structure according to the following priority of experimental methods:<ol>\n<li>X-ray</li>\n<li>Cryo-EM</li>\n<li>NMR</li></ol></li></ul></li>\n</ul>\n<h4>Train/Validation Split</h4>\n<h5>1-a: PublicLB focused / wo-MSA model</h5>\n<p>Train and validation data is same, casp16 vfold predicted data.</p>\n<h5>1-b: PublicLB focused / with-MSA model</h5>\n<ul>\n<li>Dataset: Kaggle v1 (with MSA)</li>\n<li>Validation: the most recent 37 sequences</li>\n<li>Training: 733 sequences / Validation: 37 sequences</li>\n</ul>\n<h5>2-a: local LB focused / wo-MSA model</h5>\n<ul>\n<li>Dataset: Kaggle v1 (with MSA)</li>\n<li>Validation: the most recent 40 sequences</li>\n<li>Training: 2407 sequences / Validation: 40 sequences<br><br>\nSplitting the dataset solely based on a temporal cutoff date resulted in limited diversity within the validation set (e.g., the sequence length are very similar). To improve diversity, some sequences were manually reallocated between the training and validation sets (i.e., prioritizing the assignment of diverse types of sequences to the validation set over strict adherence to the temporal cutoff rule).</li>\n</ul>\n<h5>2-b: local LB focused / with-MSA model</h5>\n<ul>\n<li>Only the MSA data provided in the above training/validation set was used.</li>\n<li>Training: 1398 sequences / Validation: 36 sequences</li>\n</ul>\n<hr>\n<h2>Training</h2>\n<h4>Training Configuration</h4>\n<ul>\n<li><strong>Learning rate</strong>: 1e-4</li>\n<li><strong>No LR scheduler</strong></li>\n<li><strong>warmup steps</strong>: 50</li>\n<li><strong>EMA (Exponential Moving Average) decay</strong>: 0.995</li>\n<li><strong>Diffusion sampling</strong>: 20</li>\n<li><strong>Trunk recycling</strong>: 4</li>\n<li><strong>Sequence length cutoff</strong>: 416 nucleotides</li>\n<li><strong>GPU</strong>: RTX5090 or RTX4090</li>\n</ul>\n<h4>1. Public LB Focused Models</h4>\n<ul>\n<li><strong>wo-MSA model</strong><ul>\n<li>Fine-tuned with CASP16 VFold structures (44 sequences) without MSA. This <strong>CASP16 overfitting checkpoint</strong> was used not only for submission but also for further fine-tuning of other models.</li>\n<li>num_steps: 2,000</li></ul></li>\n<li><strong>with-MSA model</strong><ul>\n<li>The above <strong>CASP16 overfitting checkpoint</strong> was loaded and the model was trained with Kaggle v1 data having MSA.</li>\n<li>num_steps: 7,300</li></ul></li>\n</ul>\n<h4>2. Local Score Focused Models</h4>\n<ul>\n<li><strong>wo-MSA model</strong><ul>\n<li>Fine-tuned with Kaggle v1/v2 data without MSA.</li>\n<li>num_steps: 48,200</li></ul></li>\n<li><strong>with-MSA model</strong><ul>\n<li>The <strong>CASP16 overfitting checkpoint</strong> was loaded and the model was fine-tuned with Kaggle v1/v2 data having MSA.</li>\n<li>num_steps: 29,320</li></ul></li>\n</ul>\n<hr>\n<h2>Inference</h2>\n<h4>Inference Configuration</h4>\n<p>For final submissions:</p>\n<ul>\n<li><strong>Max sequence length</strong>: 850<ul>\n<li>Long sequences were truncated to 850 residues to fit within VRAM constraints in the Kaggle environment.</li></ul></li>\n<li><strong>Diffusion sampling</strong>: 200</li>\n<li><strong>Trunk recycling</strong>: 10</li>\n<li><strong>Number of predictions</strong>: Five structures were predicted by each of the two Protenix models (with/without MSA), for a total of ten structures per target.</li>\n</ul>\n<h3>Post-Processing (Actually not worked on private test set)</h3>\n<p>Some techniques were employed for truncated position padding and structure selection. They slightly improved both the local score and public LB score, suggesting these tricks were working properly, at least during the competition period. But in reality, they did not produce meaningful gains on the private LB, or worsened the score in some cases. Having believed they were reasonable approaches, the result was disappointing, but I'm sharing them here anyway.</p>\n<ul>\n<li><p><strong>Padding Method</strong><br><br>\nThe coordinates of the truncated position must be filled with some value. Although some public notebooks seemed to pad with (0, 0, 0), padding with the same coordinates as the truncated end resulted in a slight boost in public LB within the scope of my experiments (see figure below ). However, the private LB did not show a meaningful improvement over zero-padding. If the private set did not include any sequences exceeding 850 nucleotides, truncation would not have been triggered, and thus it may not have been an essential factor.<br><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578396#3200403\" target=\"_blank\">Related discussion post</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11204962%2Fb823fc17766e638f7e562b63b75ce1c0%2Fend-same-padding.jpg?generation=1759436134732824&amp;alt=media\" alt=\"\"></p></li>\n\n<li><p><strong>K-Medoids Clustering for Structure Selection</strong><br><br>\nFive structures needed to be selected out of ten (two Protenix output five structures each). Although selection based on confidence (pLDDT) is a common approach, I was concerned that it tends to favor similar structures, limiting the diversity. Given that the competition metric adopts the highest TM-score among five submissions, selecting structurally diverse candidates seemed to be more effective. To achieve this, k-medoids clustering (k = 5) was performed on the ten predicted structures using <code>1 - TM-score</code> as a distance metric, and the five medoids were selected for submission. This approach led to improvements in public LB scores for both the LB-focused model and the local score-focused model, compared to pLDDT-based selection. However, on the private LB there was little difference overall, and in some cases pLDDT-based selection actually performed better.</p></li>\n</ul>\n<h4>Comparison of Post-Processing Methods</h4>\n<p><strong>Submission 1 (Public LB Focused Model) Variants</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Padding</th>\n<th>Structure Selection</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Final Submission 1</strong></td>\n<td>end-same</td>\n<td>k-medoids (k=5)</td>\n<td>0.46901</td>\n<td>0.43514</td>\n</tr>\n<tr>\n<td>Variant 1-1</td>\n<td>zero</td>\n<td>k-medoids (k=5)</td>\n<td>0.46576</td>\n<td>0.43217</td>\n</tr>\n<tr>\n<td>Variant 1-2</td>\n<td>end-same</td>\n<td>pLDDT top 2+3<br>(no-MSA 2 + with-MSA 3)</td>\n<td>0.44640</td>\n<td>0.43643</td>\n</tr>\n<tr>\n<td>Variant 1-3</td>\n<td>zero</td>\n<td>pLDDT top 2+3</td>\n<td>0.44960</td>\n<td>0.43443</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p><strong>Submission 2 (Local Score Focused Model) Variants</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Padding</th>\n<th>Structure Selection</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Final Submission 2</strong></td>\n<td>end-same</td>\n<td>k-medoids (k=5)</td>\n<td>0.43991</td>\n<td>0.48511</td>\n</tr>\n<tr>\n<td>Variant 2-1</td>\n<td>zero</td>\n<td>k-medoids (k=5)</td>\n<td>0.43906</td>\n<td>0.48459</td>\n</tr>\n<tr>\n<td>Variant 2-2</td>\n<td>end-same</td>\n<td>pLDDT top 2+3</td>\n<td>0.43826</td>\n<td>0.48831</td>\n</tr>\n<tr>\n<td>Variant 2-3</td>\n<td>zero</td>\n<td>pLDDT top 2+3</td>\n<td>0.43568</td>\n<td>0.48861</td>\n</tr>\n</tbody>\n</table>\n<p><em>The 'variants' load the same weights as <strong>Final Submission 1 or 2</strong>. Only the post-processing methods in inference were changed.</em></p>",
  "messages": [
    {
      "id": 3297322,
      "postDate": "2025-10-02T21:05:51.427Z",
      "content": "<h1>7th Place Solution: Ensemble of Fine-tuned Protenix</h1>\n<p>First of all, I would like to express my sincere gratitude to the competition host team and the Kaggle team for organizing this interesting competition. I am also deeply thankful to all the participants who continuously shared valuable insights and information. The three months spent tackling machine learning for biomolecular 3D structure prediction have been truly enjoyable and rewarding.</p>\n<h2>Summary</h2>\n<p>My solution consists of an ensemble of two types of fine-tuned <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a> (AlphaFold3 clone), <strong>with and without MSA</strong>. In the early phase of the competition I tried <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training\" target=\"_blank\">RibonanzaNet2_DDPM</a>, <a href=\"https://github.com/leeyang/DRfold2\" target=\"_blank\">DRfold2</a>, and <a href=\"https://github.com/kiharalab/NuFold\" target=\"_blank\">NuFold</a>, but Protenix gave the best results in my experiments, so I focused exclusively on Protenix in the latter half. Protenix finetuning was conducted with the code provided by <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> and <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/10th-place-solution-d4t4-team\" target=\"_blank\">d4t4 team</a>, with minor modifications.</p>\n<ul>\n<li><a href=\"https://github.com/lhwcv/Protenix-RNA-Kaggle\" target=\"_blank\">GitHub</a> (credit: <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>)</li>\n<li><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495\" target=\"_blank\">Discussion post</a></li>\n</ul>\n<p>I tried two hill-climbing strategies and selected both two models as final submissions:</p>\n<ol>\n<li><strong>Public LB score focused</strong>: Mainly trained on CASP16 VFold predicted structures</li>\n<li><strong>Local score focused</strong>: Mainly trained on host-provided PDB data (Kaggle v1/v2 data)</li>\n</ol>\n<p>TM-Score results are as follows:</p>\n<h4>1. Public LB Focused Models</h4>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>wo-MSA</td>\n<td>0.43529</td>\n<td>0.39196</td>\n</tr>\n<tr>\n<td>with-MSA</td>\n<td>0.43205</td>\n<td>0.41014</td>\n</tr>\n<tr>\n<td>Two models ensemble<br>(<strong>Final Submission 1</strong>)</td>\n<td>0.46901</td>\n<td>0.43514</td>\n</tr>\n</tbody>\n</table>\n<h4>2. Local Score Focused Models</h4>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local Validation</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>wo-MSA</td>\n<td>0.5982</td>\n<td>0.43135</td>\n<td>0.47654</td>\n</tr>\n<tr>\n<td>with-MSA</td>\n<td>0.5943</td>\n<td>0.42501</td>\n<td>0.46793</td>\n</tr>\n<tr>\n<td>Two models ensemble<br>(<strong>Final Submission 2</strong>)</td>\n<td>0.6092</td>\n<td>0.43991</td>\n<td><strong>0.48511</strong><br>(private 7th place)</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Dataset</h2>\n<p>Kaggle v1/v2 data and CASP16 VFold prediction were used for model training.</p>\n<h4>Data Filtering</h4>\n<ul>\n<li><p><strong>Quality Criteria</strong><br><br>\nOnly data that met the following criteria were used:</p>\n<ul>\n<li>Consecutive nucleotide C1'-C1' distance ≤ 15 Å</li>\n<li>Resolution ≤ 15 Å</li>\n<li>Contains all four nucleotides (A, U, G, C)</li></ul></li>\n<li><p><strong>Deduplication</strong><br><br>\nFor sequences with multiple structure data, only one structure was selected for training according to the following rules:</p>\n<ul>\n<li>Keep the structure with the lowest resolution (Å)</li>\n<li>If resolution is NaN, select one structure according to the following priority of experimental methods:<ol>\n<li>X-ray</li>\n<li>Cryo-EM</li>\n<li>NMR</li></ol></li></ul></li>\n</ul>\n<h4>Train/Validation Split</h4>\n<h5>1-a: PublicLB focused / wo-MSA model</h5>\n<p>Train and validation data is same, casp16 vfold predicted data.</p>\n<h5>1-b: PublicLB focused / with-MSA model</h5>\n<ul>\n<li>Dataset: Kaggle v1 (with MSA)</li>\n<li>Validation: the most recent 37 sequences</li>\n<li>Training: 733 sequences / Validation: 37 sequences</li>\n</ul>\n<h5>2-a: local LB focused / wo-MSA model</h5>\n<ul>\n<li>Dataset: Kaggle v1 (with MSA)</li>\n<li>Validation: the most recent 40 sequences</li>\n<li>Training: 2407 sequences / Validation: 40 sequences<br><br>\nSplitting the dataset solely based on a temporal cutoff date resulted in limited diversity within the validation set (e.g., the sequence length are very similar). To improve diversity, some sequences were manually reallocated between the training and validation sets (i.e., prioritizing the assignment of diverse types of sequences to the validation set over strict adherence to the temporal cutoff rule).</li>\n</ul>\n<h5>2-b: local LB focused / with-MSA model</h5>\n<ul>\n<li>Only the MSA data provided in the above training/validation set was used.</li>\n<li>Training: 1398 sequences / Validation: 36 sequences</li>\n</ul>\n<hr>\n<h2>Training</h2>\n<h4>Training Configuration</h4>\n<ul>\n<li><strong>Learning rate</strong>: 1e-4</li>\n<li><strong>No LR scheduler</strong></li>\n<li><strong>warmup steps</strong>: 50</li>\n<li><strong>EMA (Exponential Moving Average) decay</strong>: 0.995</li>\n<li><strong>Diffusion sampling</strong>: 20</li>\n<li><strong>Trunk recycling</strong>: 4</li>\n<li><strong>Sequence length cutoff</strong>: 416 nucleotides</li>\n<li><strong>GPU</strong>: RTX5090 or RTX4090</li>\n</ul>\n<h4>1. Public LB Focused Models</h4>\n<ul>\n<li><strong>wo-MSA model</strong><ul>\n<li>Fine-tuned with CASP16 VFold structures (44 sequences) without MSA. This <strong>CASP16 overfitting checkpoint</strong> was used not only for submission but also for further fine-tuning of other models.</li>\n<li>num_steps: 2,000</li></ul></li>\n<li><strong>with-MSA model</strong><ul>\n<li>The above <strong>CASP16 overfitting checkpoint</strong> was loaded and the model was trained with Kaggle v1 data having MSA.</li>\n<li>num_steps: 7,300</li></ul></li>\n</ul>\n<h4>2. Local Score Focused Models</h4>\n<ul>\n<li><strong>wo-MSA model</strong><ul>\n<li>Fine-tuned with Kaggle v1/v2 data without MSA.</li>\n<li>num_steps: 48,200</li></ul></li>\n<li><strong>with-MSA model</strong><ul>\n<li>The <strong>CASP16 overfitting checkpoint</strong> was loaded and the model was fine-tuned with Kaggle v1/v2 data having MSA.</li>\n<li>num_steps: 29,320</li></ul></li>\n</ul>\n<hr>\n<h2>Inference</h2>\n<h4>Inference Configuration</h4>\n<p>For final submissions:</p>\n<ul>\n<li><strong>Max sequence length</strong>: 850<ul>\n<li>Long sequences were truncated to 850 residues to fit within VRAM constraints in the Kaggle environment.</li></ul></li>\n<li><strong>Diffusion sampling</strong>: 200</li>\n<li><strong>Trunk recycling</strong>: 10</li>\n<li><strong>Number of predictions</strong>: Five structures were predicted by each of the two Protenix models (with/without MSA), for a total of ten structures per target.</li>\n</ul>\n<h3>Post-Processing (Actually not worked on private test set)</h3>\n<p>Some techniques were employed for truncated position padding and structure selection. They slightly improved both the local score and public LB score, suggesting these tricks were working properly, at least during the competition period. But in reality, they did not produce meaningful gains on the private LB, or worsened the score in some cases. Having believed they were reasonable approaches, the result was disappointing, but I'm sharing them here anyway.</p>\n<ul>\n<li><p><strong>Padding Method</strong><br><br>\nThe coordinates of the truncated position must be filled with some value. Although some public notebooks seemed to pad with (0, 0, 0), padding with the same coordinates as the truncated end resulted in a slight boost in public LB within the scope of my experiments (see figure below ). However, the private LB did not show a meaningful improvement over zero-padding. If the private set did not include any sequences exceeding 850 nucleotides, truncation would not have been triggered, and thus it may not have been an essential factor.<br><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578396#3200403\" target=\"_blank\">Related discussion post</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11204962%2Fb823fc17766e638f7e562b63b75ce1c0%2Fend-same-padding.jpg?generation=1759436134732824&amp;alt=media\" alt=\"\"></p></li>\n\n<li><p><strong>K-Medoids Clustering for Structure Selection</strong><br><br>\nFive structures needed to be selected out of ten (two Protenix output five structures each). Although selection based on confidence (pLDDT) is a common approach, I was concerned that it tends to favor similar structures, limiting the diversity. Given that the competition metric adopts the highest TM-score among five submissions, selecting structurally diverse candidates seemed to be more effective. To achieve this, k-medoids clustering (k = 5) was performed on the ten predicted structures using <code>1 - TM-score</code> as a distance metric, and the five medoids were selected for submission. This approach led to improvements in public LB scores for both the LB-focused model and the local score-focused model, compared to pLDDT-based selection. However, on the private LB there was little difference overall, and in some cases pLDDT-based selection actually performed better.</p></li>\n</ul>\n<h4>Comparison of Post-Processing Methods</h4>\n<p><strong>Submission 1 (Public LB Focused Model) Variants</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Padding</th>\n<th>Structure Selection</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Final Submission 1</strong></td>\n<td>end-same</td>\n<td>k-medoids (k=5)</td>\n<td>0.46901</td>\n<td>0.43514</td>\n</tr>\n<tr>\n<td>Variant 1-1</td>\n<td>zero</td>\n<td>k-medoids (k=5)</td>\n<td>0.46576</td>\n<td>0.43217</td>\n</tr>\n<tr>\n<td>Variant 1-2</td>\n<td>end-same</td>\n<td>pLDDT top 2+3<br>(no-MSA 2 + with-MSA 3)</td>\n<td>0.44640</td>\n<td>0.43643</td>\n</tr>\n<tr>\n<td>Variant 1-3</td>\n<td>zero</td>\n<td>pLDDT top 2+3</td>\n<td>0.44960</td>\n<td>0.43443</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p><strong>Submission 2 (Local Score Focused Model) Variants</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Padding</th>\n<th>Structure Selection</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Final Submission 2</strong></td>\n<td>end-same</td>\n<td>k-medoids (k=5)</td>\n<td>0.43991</td>\n<td>0.48511</td>\n</tr>\n<tr>\n<td>Variant 2-1</td>\n<td>zero</td>\n<td>k-medoids (k=5)</td>\n<td>0.43906</td>\n<td>0.48459</td>\n</tr>\n<tr>\n<td>Variant 2-2</td>\n<td>end-same</td>\n<td>pLDDT top 2+3</td>\n<td>0.43826</td>\n<td>0.48831</td>\n</tr>\n<tr>\n<td>Variant 2-3</td>\n<td>zero</td>\n<td>pLDDT top 2+3</td>\n<td>0.43568</td>\n<td>0.48861</td>\n</tr>\n</tbody>\n</table>\n<p><em>The 'variants' load the same weights as <strong>Final Submission 1 or 2</strong>. Only the post-processing methods in inference were changed.</em></p>",
      "rawMarkdown": "# 7th Place Solution: Ensemble of Fine-tuned Protenix\n\nFirst of all, I would like to express my sincere gratitude to the competition host team and the Kaggle team for organizing this interesting competition. I am also deeply thankful to all the participants who continuously shared valuable insights and information. The three months spent tackling machine learning for biomolecular 3D structure prediction have been truly enjoyable and rewarding.\n\n## Summary\n\nMy solution consists of an ensemble of two types of fine-tuned [Protenix](https://github.com/bytedance/Protenix) (AlphaFold3 clone), **with and without MSA**. In the early phase of the competition I tried [RibonanzaNet2_DDPM](https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training), [DRfold2](https://github.com/leeyang/DRfold2), and [NuFold](https://github.com/kiharalab/NuFold), but Protenix gave the best results in my experiments, so I focused exclusively on Protenix in the latter half. Protenix finetuning was conducted with the code provided by @lihaoweicvch and [d4t4 team](https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/10th-place-solution-d4t4-team), with minor modifications.\n- [GitHub](https://github.com/lhwcv/Protenix-RNA-Kaggle) (credit: @lihaoweicvch)\n- [Discussion post]( https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495)\n\nI tried two hill-climbing strategies and selected both two models as final submissions:\n1. **Public LB score focused**: Mainly trained on CASP16 VFold predicted structures\n2. **Local score focused**: Mainly trained on host-provided PDB data (Kaggle v1/v2 data)\n\nTM-Score results are as follows:\n\n#### 1. Public LB Focused Models\n\n|                      Model                      | Public LB | Private LB |\n| :---------------------------------------------: | --------: | ---------: |\n|                     wo-MSA                      |   0.43529 |    0.39196 |\n|                    with-MSA                     |   0.43205 |    0.41014 |\n| Two models ensemble<br>(**Final Submission 1**) |   0.46901 |    0.43514 |\n\n#### 2. Local Score Focused Models\n\n|                      Model                      | Local Validation | Public LB |             Private LB             |\n| :---------------------------------------------: | ---------------: | --------: | :--------------------------------: |\n|                     wo-MSA                      |           0.5982 |   0.43135 |              0.47654               |\n|                    with-MSA                     |           0.5943 |   0.42501 |              0.46793               |\n| Two models ensemble<br>(**Final Submission 2**) |           0.6092 |   0.43991 | **0.48511**<br>(private 7th place) |\n\n---\n\n## Dataset\nKaggle v1/v2 data and CASP16 VFold prediction were used for model training.\n\n#### Data Filtering\n\n- **Quality Criteria**<br>\nOnly data that met the following criteria were used:\n    - Consecutive nucleotide C1'-C1' distance ≤ 15 Å\n    - Resolution ≤ 15 Å\n    - Contains all four nucleotides (A, U, G, C)\n\n- **Deduplication**<br>\nFor sequences with multiple structure data, only one structure was selected for training according to the following rules:\n    - Keep the structure with the lowest resolution (Å)\n    - If resolution is NaN, select one structure according to the following priority of experimental methods:\n        1. X-ray\n        2. Cryo-EM\n        3. NMR\n\n#### Train/Validation Split\n\n##### 1-a: PublicLB focused / wo-MSA model\nTrain and validation data is same, casp16 vfold predicted data.\n\n##### 1-b: PublicLB focused / with-MSA model\n- Dataset: Kaggle v1 (with MSA)\n- Validation: the most recent 37 sequences\n- Training: 733 sequences / Validation: 37 sequences\n\n##### 2-a: local LB focused / wo-MSA model\n- Dataset: Kaggle v1 (with MSA)\n- Validation: the most recent 40 sequences\n- Training: 2407 sequences / Validation: 40 sequences<br>\nSplitting the dataset solely based on a temporal cutoff date resulted in limited diversity within the validation set (e.g., the sequence length are very similar). To improve diversity, some sequences were manually reallocated between the training and validation sets (i.e., prioritizing the assignment of diverse types of sequences to the validation set over strict adherence to the temporal cutoff rule).\n\n##### 2-b: local LB focused / with-MSA model\n- Only the MSA data provided in the above training/validation set was used.\n- Training: 1398 sequences / Validation: 36 sequences\n\n---\n\n## Training\n\n#### Training Configuration\n- **Learning rate**: 1e-4\n- **No LR scheduler**\n- **warmup steps**: 50\n- **EMA (Exponential Moving Average) decay**: 0.995\n- **Diffusion sampling**: 20\n- **Trunk recycling**: 4\n- **Sequence length cutoff**: 416 nucleotides\n- **GPU**: RTX5090 or RTX4090\n\n#### 1. Public LB Focused Models\n- **wo-MSA model**\n    - Fine-tuned with CASP16 VFold structures (44 sequences) without MSA. This **CASP16 overfitting checkpoint** was used not only for submission but also for further fine-tuning of other models.\n    - num_steps: 2,000\n- **with-MSA model**\n    - The above **CASP16 overfitting checkpoint** was loaded and the model was trained with Kaggle v1 data having MSA.\n    - num_steps: 7,300\n\n#### 2. Local Score Focused Models\n- **wo-MSA model**\n    - Fine-tuned with Kaggle v1/v2 data without MSA.\n    - num_steps: 48,200\n- **with-MSA model**\n    - The **CASP16 overfitting checkpoint** was loaded and the model was fine-tuned with Kaggle v1/v2 data having MSA.\n    - num_steps: 29,320\n\n---\n\n## Inference\n\n#### Inference Configuration\nFor final submissions:\n- **Max sequence length**: 850\n    - Long sequences were truncated to 850 residues to fit within VRAM constraints in the Kaggle environment.\n- **Diffusion sampling**: 200\n- **Trunk recycling**: 10\n- **Number of predictions**: Five structures were predicted by each of the two Protenix models (with/without MSA), for a total of ten structures per target.\n\n### Post-Processing (Actually not worked on private test set)\n\nSome techniques were employed for truncated position padding and structure selection. They slightly improved both the local score and public LB score, suggesting these tricks were working properly, at least during the competition period. But in reality, they did not produce meaningful gains on the private LB, or worsened the score in some cases. Having believed they were reasonable approaches, the result was disappointing, but I'm sharing them here anyway.\n\n- **Padding Method**<br>\nThe coordinates of the truncated position must be filled with some value. Although some public notebooks seemed to pad with (0, 0, 0), padding with the same coordinates as the truncated end resulted in a slight boost in public LB within the scope of my experiments (see figure below ). However, the private LB did not show a meaningful improvement over zero-padding. If the private set did not include any sequences exceeding 850 nucleotides, truncation would not have been triggered, and thus it may not have been an essential factor.<br>[Related discussion post](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578396#3200403)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11204962%2Fb823fc17766e638f7e562b63b75ce1c0%2Fend-same-padding.jpg?generation=1759436134732824&alt=media)\n\n\n\n- **K-Medoids Clustering for Structure Selection**<br>\nFive structures needed to be selected out of ten (two Protenix output five structures each). Although selection based on confidence (pLDDT) is a common approach, I was concerned that it tends to favor similar structures, limiting the diversity. Given that the competition metric adopts the highest TM-score among five submissions, selecting structurally diverse candidates seemed to be more effective. To achieve this, k-medoids clustering (k = 5) was performed on the ten predicted structures using `1 - TM-score` as a distance metric, and the five medoids were selected for submission. This approach led to improvements in public LB scores for both the LB-focused model and the local score-focused model, compared to pLDDT-based selection. However, on the private LB there was little difference overall, and in some cases pLDDT-based selection actually performed better.\n\n#### Comparison of Post-Processing Methods\n\n**Submission 1 (Public LB Focused Model) Variants**\n\n| Model                  | Padding  | Structure Selection                      | Public LB | Private LB |\n| ---------------------- | -------- | ---------------------------------------- | --------: | ---------: |\n| **Final Submission 1** | end-same | k-medoids (k=5)                          |   0.46901 |    0.43514 |\n| Variant 1-1            | zero     | k-medoids (k=5)                          |   0.46576 |    0.43217 |\n| Variant 1-2            | end-same | pLDDT top 2+3<br>(no-MSA 2 + with-MSA 3) |   0.44640 |    0.43643 |\n| Variant 1-3            | zero     | pLDDT top 2+3                            |   0.44960 |    0.43443 |\n\n<br>\n\n**Submission 2 (Local Score Focused Model) Variants**\n\n| Model                  | Padding  | Structure Selection | Public LB | Private LB |\n| ---------------------- | -------- | ------------------- | --------: | ---------: |\n| **Final Submission 2** | end-same | k-medoids (k=5)     |   0.43991 |    0.48511 |\n| Variant 2-1            | zero     | k-medoids (k=5)     |   0.43906 |    0.48459 |\n| Variant 2-2            | end-same | pLDDT top 2+3       |   0.43826 |    0.48831 |\n| Variant 2-3            | zero     | pLDDT top 2+3       |   0.43568 |    0.48861 |\n\n*The 'variants' load the same weights as **Final Submission 1 or 2**. Only the post-processing methods in inference were changed.*\n\n",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3297322": "# 7th Place Solution: Ensemble of Fine-tuned Protenix\n\nFirst of all, I would like to express my sincere gratitude to the competition host team and the Kaggle team for organizing this interesting competition. I am also deeply thankful to all the participants who continuously shared valuable insights and information. The three months spent tackling machine learning for biomolecular 3D structure prediction have been truly enjoyable and rewarding.\n\n## Summary\n\nMy solution consists of an ensemble of two types of fine-tuned [Protenix](https://github.com/bytedance/Protenix) (AlphaFold3 clone), **with and without MSA**. In the early phase of the competition I tried [RibonanzaNet2_DDPM](https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training), [DRfold2](https://github.com/leeyang/DRfold2), and [NuFold](https://github.com/kiharalab/NuFold), but Protenix gave the best results in my experiments, so I focused exclusively on Protenix in the latter half. Protenix finetuning was conducted with the code provided by @lihaoweicvch and [d4t4 team](https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/10th-place-solution-d4t4-team), with minor modifications.\n- [GitHub](https://github.com/lhwcv/Protenix-RNA-Kaggle) (credit: @lihaoweicvch)\n- [Discussion post]( https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/573495)\n\nI tried two hill-climbing strategies and selected both two models as final submissions:\n1. **Public LB score focused**: Mainly trained on CASP16 VFold predicted structures\n2. **Local score focused**: Mainly trained on host-provided PDB data (Kaggle v1/v2 data)\n\nTM-Score results are as follows:\n\n#### 1. Public LB Focused Models\n\n|                      Model                      | Public LB | Private LB |\n| :---------------------------------------------: | --------: | ---------: |\n|                     wo-MSA                      |   0.43529 |    0.39196 |\n|                    with-MSA                     |   0.43205 |    0.41014 |\n| Two models ensemble<br>(**Final Submission 1**) |   0.46901 |    0.43514 |\n\n#### 2. Local Score Focused Models\n\n|                      Model                      | Local Validation | Public LB |             Private LB             |\n| :---------------------------------------------: | ---------------: | --------: | :--------------------------------: |\n|                     wo-MSA                      |           0.5982 |   0.43135 |              0.47654               |\n|                    with-MSA                     |           0.5943 |   0.42501 |              0.46793               |\n| Two models ensemble<br>(**Final Submission 2**) |           0.6092 |   0.43991 | **0.48511**<br>(private 7th place) |\n\n---\n\n## Dataset\nKaggle v1/v2 data and CASP16 VFold prediction were used for model training.\n\n#### Data Filtering\n\n- **Quality Criteria**<br>\nOnly data that met the following criteria were used:\n    - Consecutive nucleotide C1'-C1' distance ≤ 15 Å\n    - Resolution ≤ 15 Å\n    - Contains all four nucleotides (A, U, G, C)\n\n- **Deduplication**<br>\nFor sequences with multiple structure data, only one structure was selected for training according to the following rules:\n    - Keep the structure with the lowest resolution (Å)\n    - If resolution is NaN, select one structure according to the following priority of experimental methods:\n        1. X-ray\n        2. Cryo-EM\n        3. NMR\n\n#### Train/Validation Split\n\n##### 1-a: PublicLB focused / wo-MSA model\nTrain and validation data is same, casp16 vfold predicted data.\n\n##### 1-b: PublicLB focused / with-MSA model\n- Dataset: Kaggle v1 (with MSA)\n- Validation: the most recent 37 sequences\n- Training: 733 sequences / Validation: 37 sequences\n\n##### 2-a: local LB focused / wo-MSA model\n- Dataset: Kaggle v1 (with MSA)\n- Validation: the most recent 40 sequences\n- Training: 2407 sequences / Validation: 40 sequences<br>\nSplitting the dataset solely based on a temporal cutoff date resulted in limited diversity within the validation set (e.g., the sequence length are very similar). To improve diversity, some sequences were manually reallocated between the training and validation sets (i.e., prioritizing the assignment of diverse types of sequences to the validation set over strict adherence to the temporal cutoff rule).\n\n##### 2-b: local LB focused / with-MSA model\n- Only the MSA data provided in the above training/validation set was used.\n- Training: 1398 sequences / Validation: 36 sequences\n\n---\n\n## Training\n\n#### Training Configuration\n- **Learning rate**: 1e-4\n- **No LR scheduler**\n- **warmup steps**: 50\n- **EMA (Exponential Moving Average) decay**: 0.995\n- **Diffusion sampling**: 20\n- **Trunk recycling**: 4\n- **Sequence length cutoff**: 416 nucleotides\n- **GPU**: RTX5090 or RTX4090\n\n#### 1. Public LB Focused Models\n- **wo-MSA model**\n    - Fine-tuned with CASP16 VFold structures (44 sequences) without MSA. This **CASP16 overfitting checkpoint** was used not only for submission but also for further fine-tuning of other models.\n    - num_steps: 2,000\n- **with-MSA model**\n    - The above **CASP16 overfitting checkpoint** was loaded and the model was trained with Kaggle v1 data having MSA.\n    - num_steps: 7,300\n\n#### 2. Local Score Focused Models\n- **wo-MSA model**\n    - Fine-tuned with Kaggle v1/v2 data without MSA.\n    - num_steps: 48,200\n- **with-MSA model**\n    - The **CASP16 overfitting checkpoint** was loaded and the model was fine-tuned with Kaggle v1/v2 data having MSA.\n    - num_steps: 29,320\n\n---\n\n## Inference\n\n#### Inference Configuration\nFor final submissions:\n- **Max sequence length**: 850\n    - Long sequences were truncated to 850 residues to fit within VRAM constraints in the Kaggle environment.\n- **Diffusion sampling**: 200\n- **Trunk recycling**: 10\n- **Number of predictions**: Five structures were predicted by each of the two Protenix models (with/without MSA), for a total of ten structures per target.\n\n### Post-Processing (Actually not worked on private test set)\n\nSome techniques were employed for truncated position padding and structure selection. They slightly improved both the local score and public LB score, suggesting these tricks were working properly, at least during the competition period. But in reality, they did not produce meaningful gains on the private LB, or worsened the score in some cases. Having believed they were reasonable approaches, the result was disappointing, but I'm sharing them here anyway.\n\n- **Padding Method**<br>\nThe coordinates of the truncated position must be filled with some value. Although some public notebooks seemed to pad with (0, 0, 0), padding with the same coordinates as the truncated end resulted in a slight boost in public LB within the scope of my experiments (see figure below ). However, the private LB did not show a meaningful improvement over zero-padding. If the private set did not include any sequences exceeding 850 nucleotides, truncation would not have been triggered, and thus it may not have been an essential factor.<br>[Related discussion post](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578396#3200403)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11204962%2Fb823fc17766e638f7e562b63b75ce1c0%2Fend-same-padding.jpg?generation=1759436134732824&alt=media)\n\n\n\n- **K-Medoids Clustering for Structure Selection**<br>\nFive structures needed to be selected out of ten (two Protenix output five structures each). Although selection based on confidence (pLDDT) is a common approach, I was concerned that it tends to favor similar structures, limiting the diversity. Given that the competition metric adopts the highest TM-score among five submissions, selecting structurally diverse candidates seemed to be more effective. To achieve this, k-medoids clustering (k = 5) was performed on the ten predicted structures using `1 - TM-score` as a distance metric, and the five medoids were selected for submission. This approach led to improvements in public LB scores for both the LB-focused model and the local score-focused model, compared to pLDDT-based selection. However, on the private LB there was little difference overall, and in some cases pLDDT-based selection actually performed better.\n\n#### Comparison of Post-Processing Methods\n\n**Submission 1 (Public LB Focused Model) Variants**\n\n| Model                  | Padding  | Structure Selection                      | Public LB | Private LB |\n| ---------------------- | -------- | ---------------------------------------- | --------: | ---------: |\n| **Final Submission 1** | end-same | k-medoids (k=5)                          |   0.46901 |    0.43514 |\n| Variant 1-1            | zero     | k-medoids (k=5)                          |   0.46576 |    0.43217 |\n| Variant 1-2            | end-same | pLDDT top 2+3<br>(no-MSA 2 + with-MSA 3) |   0.44640 |    0.43643 |\n| Variant 1-3            | zero     | pLDDT top 2+3                            |   0.44960 |    0.43443 |\n\n<br>\n\n**Submission 2 (Local Score Focused Model) Variants**\n\n| Model                  | Padding  | Structure Selection | Public LB | Private LB |\n| ---------------------- | -------- | ------------------- | --------: | ---------: |\n| **Final Submission 2** | end-same | k-medoids (k=5)     |   0.43991 |    0.48511 |\n| Variant 2-1            | zero     | k-medoids (k=5)     |   0.43906 |    0.48459 |\n| Variant 2-2            | end-same | pLDDT top 2+3       |   0.43826 |    0.48831 |\n| Variant 2-3            | zero     | pLDDT top 2+3       |   0.43568 |    0.48861 |\n\n*The 'variants' load the same weights as **Final Submission 1 or 2**. Only the post-processing methods in inference were changed.*\n\n"
  }
}