{
  "id": 610934,
  "title": "public 8th/private 32nd solution",
  "url": "/competitions/stanford-rna-3d-folding/discussion/610934",
  "author_name": "Timmy Juicehouse",
  "post_date": "2025-10-07T15:09:14.716000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks very much to Das Lab and Kaggle team for hosting this wonderful competition. I used to be a bioinformatics engineer, but I have been away from the field of biology for a long time in my current work. Because of this competition, I had the opportunity to revisit the knowledge from my previous textbooks and laboratory work. Although I missed out on the gold medal, I am still very happy. It's truly refreshing to see the gold medalists, especially those who used non-deep learning approaches. I've learned a great deal from their methods.</p>\n<p>I apologize for the late submission of my writeup. The competition ended right around the Chinese National Day holiday, and my wife and I were on vacation in her hometown, spending time with her parents.</p>\n<h3>Summary</h3>\n<p>This solution utilizes a model ensemble consisting of the default pretrained Dr.fold2, the default pretrained Protenix, a Protenix model fine-tuned on Kaggle RNA and Casp16 sequences, and another fine-tuned on 10 private RNA sequences. Due to inference time constraints, 15 sets of coordinates were generated for each RNA base. The energy of these structures was then scored using a pipeline that calculates the Lennard-Jones potential and Coulomb's law with a distance-dependent dielectric. The top 5 predictions were selected as the final output. It is noted that missing segments were completed using simple structural alignment with other models, and no MSAs were used in this process.</p>\n<h4>Methods: first simple trial</h4>\n<ul>\n<li><p>This solution generated a total of 15 sets of coordinates by running an ensemble of three default pretrained models: Dr.fold2, Protenix, and Boltz-1, each producing 5 predictions. </p></li>\n<li><p>Since the maximum RNA length was set to 468, zero-padding was applied to any sequence regions exceeding this limit. </p></li>\n<li><p>An energy scoring function was then used to select the best 5 structures from the pool of 15 candidates. </p></li>\n<li><p>The inference was configured with N_ensemble=2 and N_cycle=12. </p></li>\n<li><p>This approach achieved a public score ranging between 0.413 and 0.426.</p></li>\n</ul>\n<h4>finetuning the pretrained Protenix model (based on model_v0.2.0.pt)</h4>\n<p><strong>Hardware：</strong>Intel Xeon Gold 6130 and Nvidia H800 x 2</p>\n<p><strong>Training：</strong></p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Default Value</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>lr</td>\n<td>0.001</td>\n<td>Learning rate for fine-tuning</td>\n</tr>\n<tr>\n<td>train_crop_size</td>\n<td>468</td>\n<td>Maximum sequence length to use during training</td>\n</tr>\n<tr>\n<td>diffusion_batch_size</td>\n<td>48</td>\n<td>Batch size for training</td>\n</tr>\n<tr>\n<td>max_steps</td>\n<td>100000</td>\n<td>Maximum number of fine-tuning steps</td>\n</tr>\n<tr>\n<td>warmup_steps</td>\n<td>2000</td>\n<td>Steps for learning rate warm-up</td>\n</tr>\n<tr>\n<td>ema_decay</td>\n<td>0.999</td>\n<td>Exponential moving average decay rate</td>\n</tr>\n<tr>\n<td>eval_interval</td>\n<td>400</td>\n<td>Steps between evaluation runs</td>\n</tr>\n<tr>\n<td>checkpoint_interval</td>\n<td>400</td>\n<td>Steps between saving checkpoints</td>\n</tr>\n<tr>\n<td>sample_diffusion.N_step</td>\n<td>10</td>\n<td>Number of diffusion steps for evaluation</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>No feature engineering: features are disabled by default.</li>\n<li>BF16 Mixed Precision training</li>\n<li>Monitored Metric : pLDDT, PAE, PDE.</li>\n</ul>\n<h4>Postprocess, Enhanced Energy Function</h4>\n<p><strong>Enhanced Energy Function​：</strong>\n$$\nE_{\\text{total}} = w_b E_{\\text{bond}} + w_a E_{\\text{angle}} + w_d E_{\\text{dihedral}} + w_h E_{\\text{hbond}} + w_s E_{\\text{stacking}} + w_n E_{\\text{nonbonded}}\n$$</p>\n<ul>\n<li><p>bond stretching, angle bending, and dihedral torsion for local geometry.</p></li>\n<li><p>hydrogen bonding and base stacking for secondary structure.</p></li>\n<li><p>nonbonded Lennard-Jones/electrostatics for long-range interactions.&nbsp;</p></li>\n<li><p>The postprocess distinguishes fixed and mobile atoms, applies distance cutoffs (15 Å) for efficiency, and selects the top 5 lowest-energy conformers.&nbsp;</p></li>\n<li><p>Final public score boosted: +0.045 （public score: 0.469 ~ 0.475 ）</p></li>\n<li><p>weights = { 'bond': 1.0, 'angle': 1.0, 'dihedral': 0.5, 'hbond': 1.5, 'stacking': 1.2, 'nonbonded': 0.8 }</p></li>\n</ul>\n<h3>​​Things that were not done​​:</h3>\n<h4>Why not finetune Boltz-1 and Chai-1 ？</h4>\n<p>Although with the highest quality self-estimate score (pLDDT), Boltz-1 and Chai-1 were excluded from fine-tuning candidates due to their underwhelming positional rankings on standardized model evaluation leaderboards, suggesting limited generalization capability relative to SOTAs alternatives. （I’m not sure…）</p>\n<h4>Why not use sliding window approaches for predicting RNA bases beyond maximum sequence length?</h4>\n<p>Computational overhead disproportionate to accuracy gains compared to end-to-end architectures, and inference time is very limited. </p>\n<h4>Why not finetuned Drfold2 ？</h4>\n<p>Due to constraints on company server availability and limited project timeline, certain optimizations couldn‘t be fully implemented - a somewhat regrettable but necessary compromise I feel.</p>\n<h4>finetuning RibonanzaNet？</h4>\n<p>I propose to train RibonanzaNet using a pseudo-labeling strategy, where the labels are generated from predictions of diverse models. This initiative is driven by the observation that RibonanzaNet offers significantly faster inference speed, making it suitable for deployment. </p>\n<p>While pseudo-labeling and knowledge distillation strategies presented promising avenues for model optimization, these approaches ultimately could not be implemented due to insufficient computational resource allocation—an unavoidable limitation given current infrastructure constraints.</p>",
  "messages": [
    {
      "id": 3299242,
      "postDate": "2025-10-07T15:09:14.717Z",
      "content": "<p>Thanks very much to Das Lab and Kaggle team for hosting this wonderful competition. I used to be a bioinformatics engineer, but I have been away from the field of biology for a long time in my current work. Because of this competition, I had the opportunity to revisit the knowledge from my previous textbooks and laboratory work. Although I missed out on the gold medal, I am still very happy. It's truly refreshing to see the gold medalists, especially those who used non-deep learning approaches. I've learned a great deal from their methods.</p>\n<p>I apologize for the late submission of my writeup. The competition ended right around the Chinese National Day holiday, and my wife and I were on vacation in her hometown, spending time with her parents.</p>\n<h3>Summary</h3>\n<p>This solution utilizes a model ensemble consisting of the default pretrained Dr.fold2, the default pretrained Protenix, a Protenix model fine-tuned on Kaggle RNA and Casp16 sequences, and another fine-tuned on 10 private RNA sequences. Due to inference time constraints, 15 sets of coordinates were generated for each RNA base. The energy of these structures was then scored using a pipeline that calculates the Lennard-Jones potential and Coulomb's law with a distance-dependent dielectric. The top 5 predictions were selected as the final output. It is noted that missing segments were completed using simple structural alignment with other models, and no MSAs were used in this process.</p>\n<h4>Methods: first simple trial</h4>\n<ul>\n<li><p>This solution generated a total of 15 sets of coordinates by running an ensemble of three default pretrained models: Dr.fold2, Protenix, and Boltz-1, each producing 5 predictions. </p></li>\n<li><p>Since the maximum RNA length was set to 468, zero-padding was applied to any sequence regions exceeding this limit. </p></li>\n<li><p>An energy scoring function was then used to select the best 5 structures from the pool of 15 candidates. </p></li>\n<li><p>The inference was configured with N_ensemble=2 and N_cycle=12. </p></li>\n<li><p>This approach achieved a public score ranging between 0.413 and 0.426.</p></li>\n</ul>\n<h4>finetuning the pretrained Protenix model (based on model_v0.2.0.pt)</h4>\n<p><strong>Hardware：</strong>Intel Xeon Gold 6130 and Nvidia H800 x 2</p>\n<p><strong>Training：</strong></p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Default Value</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>lr</td>\n<td>0.001</td>\n<td>Learning rate for fine-tuning</td>\n</tr>\n<tr>\n<td>train_crop_size</td>\n<td>468</td>\n<td>Maximum sequence length to use during training</td>\n</tr>\n<tr>\n<td>diffusion_batch_size</td>\n<td>48</td>\n<td>Batch size for training</td>\n</tr>\n<tr>\n<td>max_steps</td>\n<td>100000</td>\n<td>Maximum number of fine-tuning steps</td>\n</tr>\n<tr>\n<td>warmup_steps</td>\n<td>2000</td>\n<td>Steps for learning rate warm-up</td>\n</tr>\n<tr>\n<td>ema_decay</td>\n<td>0.999</td>\n<td>Exponential moving average decay rate</td>\n</tr>\n<tr>\n<td>eval_interval</td>\n<td>400</td>\n<td>Steps between evaluation runs</td>\n</tr>\n<tr>\n<td>checkpoint_interval</td>\n<td>400</td>\n<td>Steps between saving checkpoints</td>\n</tr>\n<tr>\n<td>sample_diffusion.N_step</td>\n<td>10</td>\n<td>Number of diffusion steps for evaluation</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>No feature engineering: features are disabled by default.</li>\n<li>BF16 Mixed Precision training</li>\n<li>Monitored Metric : pLDDT, PAE, PDE.</li>\n</ul>\n<h4>Postprocess, Enhanced Energy Function</h4>\n<p><strong>Enhanced Energy Function​：</strong>\n$$\nE_{\\text{total}} = w_b E_{\\text{bond}} + w_a E_{\\text{angle}} + w_d E_{\\text{dihedral}} + w_h E_{\\text{hbond}} + w_s E_{\\text{stacking}} + w_n E_{\\text{nonbonded}}\n$$</p>\n<ul>\n<li><p>bond stretching, angle bending, and dihedral torsion for local geometry.</p></li>\n<li><p>hydrogen bonding and base stacking for secondary structure.</p></li>\n<li><p>nonbonded Lennard-Jones/electrostatics for long-range interactions.&nbsp;</p></li>\n<li><p>The postprocess distinguishes fixed and mobile atoms, applies distance cutoffs (15 Å) for efficiency, and selects the top 5 lowest-energy conformers.&nbsp;</p></li>\n<li><p>Final public score boosted: +0.045 （public score: 0.469 ~ 0.475 ）</p></li>\n<li><p>weights = { 'bond': 1.0, 'angle': 1.0, 'dihedral': 0.5, 'hbond': 1.5, 'stacking': 1.2, 'nonbonded': 0.8 }</p></li>\n</ul>\n<h3>​​Things that were not done​​:</h3>\n<h4>Why not finetune Boltz-1 and Chai-1 ？</h4>\n<p>Although with the highest quality self-estimate score (pLDDT), Boltz-1 and Chai-1 were excluded from fine-tuning candidates due to their underwhelming positional rankings on standardized model evaluation leaderboards, suggesting limited generalization capability relative to SOTAs alternatives. （I’m not sure…）</p>\n<h4>Why not use sliding window approaches for predicting RNA bases beyond maximum sequence length?</h4>\n<p>Computational overhead disproportionate to accuracy gains compared to end-to-end architectures, and inference time is very limited. </p>\n<h4>Why not finetuned Drfold2 ？</h4>\n<p>Due to constraints on company server availability and limited project timeline, certain optimizations couldn‘t be fully implemented - a somewhat regrettable but necessary compromise I feel.</p>\n<h4>finetuning RibonanzaNet？</h4>\n<p>I propose to train RibonanzaNet using a pseudo-labeling strategy, where the labels are generated from predictions of diverse models. This initiative is driven by the observation that RibonanzaNet offers significantly faster inference speed, making it suitable for deployment. </p>\n<p>While pseudo-labeling and knowledge distillation strategies presented promising avenues for model optimization, these approaches ultimately could not be implemented due to insufficient computational resource allocation—an unavoidable limitation given current infrastructure constraints.</p>",
      "rawMarkdown": "Thanks very much to Das Lab and Kaggle team for hosting this wonderful competition. I used to be a bioinformatics engineer, but I have been away from the field of biology for a long time in my current work. Because of this competition, I had the opportunity to revisit the knowledge from my previous textbooks and laboratory work. Although I missed out on the gold medal, I am still very happy. It's truly refreshing to see the gold medalists, especially those who used non-deep learning approaches. I've learned a great deal from their methods.\n\n\nI apologize for the late submission of my writeup. The competition ended right around the Chinese National Day holiday, and my wife and I were on vacation in her hometown, spending time with her parents.\n\n### Summary\n\nThis solution utilizes a model ensemble consisting of the default pretrained Dr.fold2, the default pretrained Protenix, a Protenix model fine-tuned on Kaggle RNA and Casp16 sequences, and another fine-tuned on 10 private RNA sequences. Due to inference time constraints, 15 sets of coordinates were generated for each RNA base. The energy of these structures was then scored using a pipeline that calculates the Lennard-Jones potential and Coulomb's law with a distance-dependent dielectric. The top 5 predictions were selected as the final output. It is noted that missing segments were completed using simple structural alignment with other models, and no MSAs were used in this process.\n\n\n#### Methods: first simple trial\n\n- This solution generated a total of 15 sets of coordinates by running an ensemble of three default pretrained models: Dr.fold2, Protenix, and Boltz-1, each producing 5 predictions. \n\n- Since the maximum RNA length was set to 468, zero-padding was applied to any sequence regions exceeding this limit. \n\n- An energy scoring function was then used to select the best 5 structures from the pool of 15 candidates. \n\n- The inference was configured with N_ensemble=2 and N_cycle=12. \n\n- This approach achieved a public score ranging between 0.413 and 0.426.\n\n#### finetuning the pretrained Protenix model (based on model_v0.2.0.pt)\n\n**Hardware：**Intel Xeon Gold 6130 and Nvidia H800 x 2\n\n**Training：**\n\n| Parameter | Default Value | Description |\n| :--- | :--- | :--- |\n| lr | 0.001 | Learning rate for fine-tuning |\n| train_crop_size | 468 | Maximum sequence length to use during training |\n| diffusion_batch_size | 48 | Batch size for training |\n| max_steps | 100000 | Maximum number of fine-tuning steps |\n| warmup_steps | 2000 | Steps for learning rate warm-up |\n| ema_decay | 0.999 | Exponential moving average decay rate |\n| eval_interval | 400 | Steps between evaluation runs |\n| checkpoint_interval | 400 | Steps between saving checkpoints |\n| sample_diffusion.N_step | 10 | Number of diffusion steps for evaluation |\n\n- No feature engineering: features are disabled by default.\n- BF16 Mixed Precision training\n- Monitored Metric : pLDDT, PAE, PDE.\n\n\n#### Postprocess, Enhanced Energy Function\n\n**Enhanced Energy Function​：**\n$$\nE_{\\text{total}} = w_b E_{\\text{bond}} + w_a E_{\\text{angle}} + w_d E_{\\text{dihedral}} + w_h E_{\\text{hbond}} + w_s E_{\\text{stacking}} + w_n E_{\\text{nonbonded}}\n$$\n\n- bond stretching, angle bending, and dihedral torsion for local geometry.\n\n- hydrogen bonding and base stacking for secondary structure.\n\n- nonbonded Lennard-Jones/electrostatics for long-range interactions. \n\n- The postprocess distinguishes fixed and mobile atoms, applies distance cutoffs (15 Å) for efficiency, and selects the top 5 lowest-energy conformers. \n\n- Final public score boosted: +0.045 （public score: 0.469 ~ 0.475 ）\n\n- weights = { 'bond': 1.0, 'angle': 1.0, 'dihedral': 0.5, 'hbond': 1.5, 'stacking': 1.2, 'nonbonded': 0.8 }\n\n\n### ​​Things that were not done​​:\n\n#### Why not finetune Boltz-1 and Chai-1 ？\n\nAlthough with the highest quality self-estimate score (pLDDT), Boltz-1 and Chai-1 were excluded from fine-tuning candidates due to their underwhelming positional rankings on standardized model evaluation leaderboards, suggesting limited generalization capability relative to SOTAs alternatives. （I’m not sure…）\n\n#### Why not use sliding window approaches for predicting RNA bases beyond maximum sequence length?\n\nComputational overhead disproportionate to accuracy gains compared to end-to-end architectures, and inference time is very limited. \n\n#### Why not finetuned Drfold2 ？\n\nDue to constraints on company server availability and limited project timeline, certain optimizations couldn‘t be fully implemented - a somewhat regrettable but necessary compromise I feel.\n\n#### finetuning RibonanzaNet？\n\nI propose to train RibonanzaNet using a pseudo-labeling strategy, where the labels are generated from predictions of diverse models. This initiative is driven by the observation that RibonanzaNet offers significantly faster inference speed, making it suitable for deployment. \n\nWhile pseudo-labeling and knowledge distillation strategies presented promising avenues for model optimization, these approaches ultimately could not be implemented due to insufficient computational resource allocation—an unavoidable limitation given current infrastructure constraints.\n\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3299242": "Thanks very much to Das Lab and Kaggle team for hosting this wonderful competition. I used to be a bioinformatics engineer, but I have been away from the field of biology for a long time in my current work. Because of this competition, I had the opportunity to revisit the knowledge from my previous textbooks and laboratory work. Although I missed out on the gold medal, I am still very happy. It's truly refreshing to see the gold medalists, especially those who used non-deep learning approaches. I've learned a great deal from their methods.\n\n\nI apologize for the late submission of my writeup. The competition ended right around the Chinese National Day holiday, and my wife and I were on vacation in her hometown, spending time with her parents.\n\n### Summary\n\nThis solution utilizes a model ensemble consisting of the default pretrained Dr.fold2, the default pretrained Protenix, a Protenix model fine-tuned on Kaggle RNA and Casp16 sequences, and another fine-tuned on 10 private RNA sequences. Due to inference time constraints, 15 sets of coordinates were generated for each RNA base. The energy of these structures was then scored using a pipeline that calculates the Lennard-Jones potential and Coulomb's law with a distance-dependent dielectric. The top 5 predictions were selected as the final output. It is noted that missing segments were completed using simple structural alignment with other models, and no MSAs were used in this process.\n\n\n#### Methods: first simple trial\n\n- This solution generated a total of 15 sets of coordinates by running an ensemble of three default pretrained models: Dr.fold2, Protenix, and Boltz-1, each producing 5 predictions. \n\n- Since the maximum RNA length was set to 468, zero-padding was applied to any sequence regions exceeding this limit. \n\n- An energy scoring function was then used to select the best 5 structures from the pool of 15 candidates. \n\n- The inference was configured with N_ensemble=2 and N_cycle=12. \n\n- This approach achieved a public score ranging between 0.413 and 0.426.\n\n#### finetuning the pretrained Protenix model (based on model_v0.2.0.pt)\n\n**Hardware：**Intel Xeon Gold 6130 and Nvidia H800 x 2\n\n**Training：**\n\n| Parameter | Default Value | Description |\n| :--- | :--- | :--- |\n| lr | 0.001 | Learning rate for fine-tuning |\n| train_crop_size | 468 | Maximum sequence length to use during training |\n| diffusion_batch_size | 48 | Batch size for training |\n| max_steps | 100000 | Maximum number of fine-tuning steps |\n| warmup_steps | 2000 | Steps for learning rate warm-up |\n| ema_decay | 0.999 | Exponential moving average decay rate |\n| eval_interval | 400 | Steps between evaluation runs |\n| checkpoint_interval | 400 | Steps between saving checkpoints |\n| sample_diffusion.N_step | 10 | Number of diffusion steps for evaluation |\n\n- No feature engineering: features are disabled by default.\n- BF16 Mixed Precision training\n- Monitored Metric : pLDDT, PAE, PDE.\n\n\n#### Postprocess, Enhanced Energy Function\n\n**Enhanced Energy Function​：**\n$$\nE_{\\text{total}} = w_b E_{\\text{bond}} + w_a E_{\\text{angle}} + w_d E_{\\text{dihedral}} + w_h E_{\\text{hbond}} + w_s E_{\\text{stacking}} + w_n E_{\\text{nonbonded}}\n$$\n\n- bond stretching, angle bending, and dihedral torsion for local geometry.\n\n- hydrogen bonding and base stacking for secondary structure.\n\n- nonbonded Lennard-Jones/electrostatics for long-range interactions. \n\n- The postprocess distinguishes fixed and mobile atoms, applies distance cutoffs (15 Å) for efficiency, and selects the top 5 lowest-energy conformers. \n\n- Final public score boosted: +0.045 （public score: 0.469 ~ 0.475 ）\n\n- weights = { 'bond': 1.0, 'angle': 1.0, 'dihedral': 0.5, 'hbond': 1.5, 'stacking': 1.2, 'nonbonded': 0.8 }\n\n\n### ​​Things that were not done​​:\n\n#### Why not finetune Boltz-1 and Chai-1 ？\n\nAlthough with the highest quality self-estimate score (pLDDT), Boltz-1 and Chai-1 were excluded from fine-tuning candidates due to their underwhelming positional rankings on standardized model evaluation leaderboards, suggesting limited generalization capability relative to SOTAs alternatives. （I’m not sure…）\n\n#### Why not use sliding window approaches for predicting RNA bases beyond maximum sequence length?\n\nComputational overhead disproportionate to accuracy gains compared to end-to-end architectures, and inference time is very limited. \n\n#### Why not finetuned Drfold2 ？\n\nDue to constraints on company server availability and limited project timeline, certain optimizations couldn‘t be fully implemented - a somewhat regrettable but necessary compromise I feel.\n\n#### finetuning RibonanzaNet？\n\nI propose to train RibonanzaNet using a pseudo-labeling strategy, where the labels are generated from predictions of diverse models. This initiative is driven by the observation that RibonanzaNet offers significantly faster inference speed, making it suitable for deployment. \n\nWhile pseudo-labeling and knowledge distillation strategies presented promising avenues for model optimization, these approaches ultimately could not be implemented due to insufficient computational resource allocation—an unavoidable limitation given current infrastructure constraints.\n\n"
  }
}