{
  "id": 582403,
  "title": "Tryhard Boltz-1 improvements",
  "url": "/competitions/stanford-rna-3d-folding/discussion/582403",
  "author_name": "",
  "post_date": "2025-05-31T01:42:30.529009600Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi!<br>\nI want to share the approaches that worked for me — but mostly didn't! 🫠<br>\nI entered the competition relatively late, so my strategy was to squeeze everything I could out of the Alphafold3 forks, as they are the best open-source approaches. Thanks to <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a> , I chose Boltz-1 ✨. <br>\nHere is the outline of the steps I tried to implement in chronological order.<br>\nBefore going further, i want to mention one more time that the scores are noisy. I haven't conducted thorough tests, but I would say that the difference in tm score between two identical Boltz submissions can easily be at least 0.01, so it is difficult to draw any conclusions from submissions with similar scores.</p>\n<ol>\n<li><p>Experiment with the hyperparameters – the most significant one for me was setting step_scale = 1.8 (while keeping recycling_steps = 6 and sampling_steps = 200 as suggested in the paper). As I understand it, this parameter is kind of the inverse of the diffusion temperature, so increasing it to reasonable values decreases the noise in the model predictions.  This scored 0.379 on the LB.</p></li>\n<li><p>Use the Amber force field from OpenMM to refine Boltz predictions. Applying it to all five models resulted in an LB score of 0.364; applying it to the last two models resulted in an LB score of 0.377.</p></li>\n<li><p>Use the steering functionality available in version 1.0.0 of Boltz. As the authors suggested, the pose itself does not change significantly, but it is more physically plausible. LB 0.370</p></li>\n<li><p>Implement some kind of bootstrapping across multiple Boltz-1 models.<br>\n4.1 Generate 15 Boltz-1 models.<br>\n4.2 Sample 10 models from the generated set.<br>\n4.3 Randomely pick \"core\" model, align the rest with USalign (parameters -mol RNA  -atom \" C1'\" -TMscore 5)<br>\n4.4 For each C1` atom from the core model, find other models' atoms for the same residue within the threshold (tune the threshold on some train samples). Output the medoid of the group of the filtered atoms. <br>\n4.5 Repeat 4 more more times to get 5 bootstrapped models. </p></li>\n</ol>\n<p>The results of the CASP15 validation were so poor that I decided not to submit them 🥲</p>\n<p>5 Tune the Boltz model to RNA sequences only. I used the following dataset for training: <a href=\"https://github.com/marcellszi/rna3db\" target=\"_blank\">https://github.com/marcellszi/rna3db</a> (thanks to <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a> who shared it :) ), as it seems to deduplicate RNA sequences properly into clusters. The Boltz training code currently doesn't support clusters for RNAs, so I patched it to make use of clusters for unbiased training. The loss behaviour is strange; see the attached pictures. I submitted the model highlighted by the green and blue dots; they scored 0.359 and 0.347 respectively 🥲.</p>\n<p>6 Add one model from trRosettaRNA v1.1 for sequences with lengths up to 250 . (Thanks <a href=\"https://www.kaggle.com/thekog\" target=\"_blank\">@thekog</a> for sharing!)<br>\nI first used RNAformer (<a href=\"https://github.com/automl/RNAformer\" target=\"_blank\">https://github.com/automl/RNAformer</a>) to obtain the secondary structure, which was then passed to trRosettaRNA. Finally, the score improved by 0.389! 🔥</p>\n<p>7 Use all available computing power to generate as many models as possible and score them. I calculated a rough law that enabled me to generate as many models as possible for a given length within the time constraints (12 minutes per sequence). Then, I used different scoring mechanisms to filter these sequences:</p>\n<p>7.1 Native Boltz confidence scores: 0.363<br>\n7.2 Rosetta ref2015: 0.367<br>\n7.3 Rosetta ref2015 + <a href=\"https://www.science.org/doi/full/10.1126/science.abe5650\" target=\"_blank\">Ares</a>(averaging ranks based on these two scores) – 0.368.</p>\n<p>While it may seem that scores other than Boltz's natives are better, when considering the noise, they are  the same, which is confirmed by other tests. </p>\n<p>So for my final solution, I calculate many models for all sequences in roughly 4 hours  and used native confidence scores for the selection(so, 6 minutes per sequence on average). For the remaining 4 hours, I use RNAformer + trRosettaRNA for small sequences, ensuring that I do not exceed the time limit.</p>\n<p>If you are  interested in code for some steps, I can share it :) </p>",
  "messages": [
    {
      "id": "3214122",
      "postDate": "05/31/2025 01:42:30",
      "content": "<p>Hi!<br>\nI want to share the approaches that worked for me — but mostly didn't! 🫠<br>\nI entered the competition relatively late, so my strategy was to squeeze everything I could out of the Alphafold3 forks, as they are the best open-source approaches. Thanks to <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a> , I chose Boltz-1 ✨. <br>\nHere is the outline of the steps I tried to implement in chronological order.<br>\nBefore going further, i want to mention one more time that the scores are noisy. I haven't conducted thorough tests, but I would say that the difference in tm score between two identical Boltz submissions can easily be at least 0.01, so it is difficult to draw any conclusions from submissions with similar scores.</p>\n<ol>\n<li><p>Experiment with the hyperparameters – the most significant one for me was setting step_scale = 1.8 (while keeping recycling_steps = 6 and sampling_steps = 200 as suggested in the paper). As I understand it, this parameter is kind of the inverse of the diffusion temperature, so increasing it to reasonable values decreases the noise in the model predictions.  This scored 0.379 on the LB.</p></li>\n<li><p>Use the Amber force field from OpenMM to refine Boltz predictions. Applying it to all five models resulted in an LB score of 0.364; applying it to the last two models resulted in an LB score of 0.377.</p></li>\n<li><p>Use the steering functionality available in version 1.0.0 of Boltz. As the authors suggested, the pose itself does not change significantly, but it is more physically plausible. LB 0.370</p></li>\n<li><p>Implement some kind of bootstrapping across multiple Boltz-1 models.<br>\n4.1 Generate 15 Boltz-1 models.<br>\n4.2 Sample 10 models from the generated set.<br>\n4.3 Randomely pick \"core\" model, align the rest with USalign (parameters -mol RNA  -atom \" C1'\" -TMscore 5)<br>\n4.4 For each C1` atom from the core model, find other models' atoms for the same residue within the threshold (tune the threshold on some train samples). Output the medoid of the group of the filtered atoms. <br>\n4.5 Repeat 4 more more times to get 5 bootstrapped models. </p></li>\n</ol>\n<p>The results of the CASP15 validation were so poor that I decided not to submit them 🥲</p>\n<p>5 Tune the Boltz model to RNA sequences only. I used the following dataset for training: <a href=\"https://github.com/marcellszi/rna3db\" target=\"_blank\">https://github.com/marcellszi/rna3db</a> (thanks to <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a> who shared it :) ), as it seems to deduplicate RNA sequences properly into clusters. The Boltz training code currently doesn't support clusters for RNAs, so I patched it to make use of clusters for unbiased training. The loss behaviour is strange; see the attached pictures. I submitted the model highlighted by the green and blue dots; they scored 0.359 and 0.347 respectively 🥲.</p>\n<p>6 Add one model from trRosettaRNA v1.1 for sequences with lengths up to 250 . (Thanks <a href=\"https://www.kaggle.com/thekog\" target=\"_blank\">@thekog</a> for sharing!)<br>\nI first used RNAformer (<a href=\"https://github.com/automl/RNAformer\" target=\"_blank\">https://github.com/automl/RNAformer</a>) to obtain the secondary structure, which was then passed to trRosettaRNA. Finally, the score improved by 0.389! 🔥</p>\n<p>7 Use all available computing power to generate as many models as possible and score them. I calculated a rough law that enabled me to generate as many models as possible for a given length within the time constraints (12 minutes per sequence). Then, I used different scoring mechanisms to filter these sequences:</p>\n<p>7.1 Native Boltz confidence scores: 0.363<br>\n7.2 Rosetta ref2015: 0.367<br>\n7.3 Rosetta ref2015 + <a href=\"https://www.science.org/doi/full/10.1126/science.abe5650\" target=\"_blank\">Ares</a>(averaging ranks based on these two scores) – 0.368.</p>\n<p>While it may seem that scores other than Boltz's natives are better, when considering the noise, they are  the same, which is confirmed by other tests. </p>\n<p>So for my final solution, I calculate many models for all sequences in roughly 4 hours  and used native confidence scores for the selection(so, 6 minutes per sequence on average). For the remaining 4 hours, I use RNAformer + trRosettaRNA for small sequences, ensuring that I do not exceed the time limit.</p>\n<p>If you are  interested in code for some steps, I can share it :) </p>",
      "rawMarkdown": "Hi!\nI want to share the approaches that worked for me — but mostly didn't! 🫠\nI entered the competition relatively late, so my strategy was to squeeze everything I could out of the Alphafold3 forks, as they are the best open-source approaches. Thanks to @youhanlee , I chose Boltz-1 ✨. \nHere is the outline of the steps I tried to implement in chronological order.\nBefore going further, i want to mention one more time that the scores are noisy. I haven't conducted thorough tests, but I would say that the difference in tm score between two identical Boltz submissions can easily be at least 0.01, so it is difficult to draw any conclusions from submissions with similar scores.\n\n1. Experiment with the hyperparameters – the most significant one for me was setting step_scale = 1.8 (while keeping recycling_steps = 6 and sampling_steps = 200 as suggested in the paper). As I understand it, this parameter is kind of the inverse of the diffusion temperature, so increasing it to reasonable values decreases the noise in the model predictions.  This scored 0.379 on the LB.\n\n2. Use the Amber force field from OpenMM to refine Boltz predictions. Applying it to all five models resulted in an LB score of 0.364; applying it to the last two models resulted in an LB score of 0.377.\n\n3. Use the steering functionality available in version 1.0.0 of Boltz. As the authors suggested, the pose itself does not change significantly, but it is more physically plausible. LB 0.370\n\n\n4. Implement some kind of bootstrapping across multiple Boltz-1 models.\n4.1 Generate 15 Boltz-1 models.\n4.2 Sample 10 models from the generated set.\n4.3 Randomely pick \"core\" model, align the rest with USalign (parameters -mol RNA  -atom \" C1'\" -TMscore 5)\n4.4 For each C1` atom from the core model, find other models' atoms for the same residue within the threshold (tune the threshold on some train samples). Output the medoid of the group of the filtered atoms. \n4.5 Repeat 4 more more times to get 5 bootstrapped models. \n\nThe results of the CASP15 validation were so poor that I decided not to submit them 🥲\n\n5 Tune the Boltz model to RNA sequences only. I used the following dataset for training: https://github.com/marcellszi/rna3db (thanks to @ckjoshi9 who shared it :) ), as it seems to deduplicate RNA sequences properly into clusters. The Boltz training code currently doesn't support clusters for RNAs, so I patched it to make use of clusters for unbiased training. The loss behaviour is strange; see the attached pictures. I submitted the model highlighted by the green and blue dots; they scored 0.359 and 0.347 respectively 🥲.\n\n6 Add one model from trRosettaRNA v1.1 for sequences with lengths up to 250 . (Thanks @thekog for sharing!)\nI first used RNAformer (https://github.com/automl/RNAformer) to obtain the secondary structure, which was then passed to trRosettaRNA. Finally, the score improved by 0.389! 🔥\n\n7 Use all available computing power to generate as many models as possible and score them. I calculated a rough law that enabled me to generate as many models as possible for a given length within the time constraints (12 minutes per sequence). Then, I used different scoring mechanisms to filter these sequences:\n\n7.1 Native Boltz confidence scores: 0.363\n7.2 Rosetta ref2015: 0.367\n7.3 Rosetta ref2015 + [Ares](https://www.science.org/doi/full/10.1126/science.abe5650)(averaging ranks based on these two scores) – 0.368.\n\nWhile it may seem that scores other than Boltz's natives are better, when considering the noise, they are  the same, which is confirmed by other tests. \n\nSo for my final solution, I calculate many models for all sequences in roughly 4 hours  and used native confidence scores for the selection(so, 6 minutes per sequence on average). For the remaining 4 hours, I use RNAformer + trRosettaRNA for small sequences, ensuring that I do not exceed the time limit.\n\nIf you are  interested in code for some steps, I can share it :)",
      "votes": null
    },
    {
      "id": "3216158",
      "postDate": "06/03/2025 07:20:39",
      "content": "<p>did you use Boltz-1 or Boltz-1x?</p>\n<p>i think you can check the new paper on trRosettaRNA2 (version 2) paper. It uses MSA and preidct SS and 3d structure. I think his results should be good here.</p>",
      "rawMarkdown": "did you use Boltz-1 or Boltz-1x?\n\ni think you can check the new paper on trRosettaRNA2 (version 2) paper. It uses MSA and preidct SS and 3d structure. I think his results should be good here.",
      "votes": null
    },
    {
      "id": "3216183",
      "postDate": "06/03/2025 07:56:07",
      "content": "<p>By default I used Boltz-1. The third point about using the steering is Boltz-1x. So, no noticeable improvement in metrics in this competition.</p>\n<p>And yes, I also believe that trRosettaRNA2 should do well in this contest. It seems that it has all the information sources plugged in. However, I didn't dare to reimplement it 😅</p>",
      "rawMarkdown": "By default I used Boltz-1. The third point about using the steering is Boltz-1x. So, no noticeable improvement in metrics in this competition.\n\nAnd yes, I also believe that trRosettaRNA2 should do well in this contest. It seems that it has all the information sources plugged in. However, I didn't dare to reimplement it 😅",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3216158,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/03/2025 07:20:39",
      "content": "<p>did you use Boltz-1 or Boltz-1x?</p>\n<p>i think you can check the new paper on trRosettaRNA2 (version 2) paper. It uses MSA and preidct SS and 3d structure. I think his results should be good here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3216183,
          "author_name": "sergeyer",
          "author_url": "",
          "post_date": "06/03/2025 07:56:07",
          "content": "<p>By default I used Boltz-1. The third point about using the steering is Boltz-1x. So, no noticeable improvement in metrics in this competition.</p>\n<p>And yes, I also believe that trRosettaRNA2 should do well in this contest. It seems that it has all the information sources plugged in. However, I didn't dare to reimplement it 😅</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3214122": "Hi!\nI want to share the approaches that worked for me — but mostly didn't! 🫠\nI entered the competition relatively late, so my strategy was to squeeze everything I could out of the Alphafold3 forks, as they are the best open-source approaches. Thanks to @youhanlee , I chose Boltz-1 ✨. \nHere is the outline of the steps I tried to implement in chronological order.\nBefore going further, i want to mention one more time that the scores are noisy. I haven't conducted thorough tests, but I would say that the difference in tm score between two identical Boltz submissions can easily be at least 0.01, so it is difficult to draw any conclusions from submissions with similar scores.\n\n1. Experiment with the hyperparameters – the most significant one for me was setting step_scale = 1.8 (while keeping recycling_steps = 6 and sampling_steps = 200 as suggested in the paper). As I understand it, this parameter is kind of the inverse of the diffusion temperature, so increasing it to reasonable values decreases the noise in the model predictions.  This scored 0.379 on the LB.\n\n2. Use the Amber force field from OpenMM to refine Boltz predictions. Applying it to all five models resulted in an LB score of 0.364; applying it to the last two models resulted in an LB score of 0.377.\n\n3. Use the steering functionality available in version 1.0.0 of Boltz. As the authors suggested, the pose itself does not change significantly, but it is more physically plausible. LB 0.370\n\n\n4. Implement some kind of bootstrapping across multiple Boltz-1 models.\n4.1 Generate 15 Boltz-1 models.\n4.2 Sample 10 models from the generated set.\n4.3 Randomely pick \"core\" model, align the rest with USalign (parameters -mol RNA  -atom \" C1'\" -TMscore 5)\n4.4 For each C1` atom from the core model, find other models' atoms for the same residue within the threshold (tune the threshold on some train samples). Output the medoid of the group of the filtered atoms. \n4.5 Repeat 4 more more times to get 5 bootstrapped models. \n\nThe results of the CASP15 validation were so poor that I decided not to submit them 🥲\n\n5 Tune the Boltz model to RNA sequences only. I used the following dataset for training: https://github.com/marcellszi/rna3db (thanks to @ckjoshi9 who shared it :) ), as it seems to deduplicate RNA sequences properly into clusters. The Boltz training code currently doesn't support clusters for RNAs, so I patched it to make use of clusters for unbiased training. The loss behaviour is strange; see the attached pictures. I submitted the model highlighted by the green and blue dots; they scored 0.359 and 0.347 respectively 🥲.\n\n6 Add one model from trRosettaRNA v1.1 for sequences with lengths up to 250 . (Thanks @thekog for sharing!)\nI first used RNAformer (https://github.com/automl/RNAformer) to obtain the secondary structure, which was then passed to trRosettaRNA. Finally, the score improved by 0.389! 🔥\n\n7 Use all available computing power to generate as many models as possible and score them. I calculated a rough law that enabled me to generate as many models as possible for a given length within the time constraints (12 minutes per sequence). Then, I used different scoring mechanisms to filter these sequences:\n\n7.1 Native Boltz confidence scores: 0.363\n7.2 Rosetta ref2015: 0.367\n7.3 Rosetta ref2015 + [Ares](https://www.science.org/doi/full/10.1126/science.abe5650)(averaging ranks based on these two scores) – 0.368.\n\nWhile it may seem that scores other than Boltz's natives are better, when considering the noise, they are  the same, which is confirmed by other tests. \n\nSo for my final solution, I calculate many models for all sequences in roughly 4 hours  and used native confidence scores for the selection(so, 6 minutes per sequence on average). For the remaining 4 hours, I use RNAformer + trRosettaRNA for small sequences, ensuring that I do not exceed the time limit.\n\nIf you are  interested in code for some steps, I can share it :)",
    "3216158": "did you use Boltz-1 or Boltz-1x?\n\ni think you can check the new paper on trRosettaRNA2 (version 2) paper. It uses MSA and preidct SS and 3d structure. I think his results should be good here.",
    "3216183": "By default I used Boltz-1. The third point about using the steering is Boltz-1x. So, no noticeable improvement in metrics in this competition.\n\nAnd yes, I also believe that trRosettaRNA2 should do well in this contest. It seems that it has all the information sources plugged in. However, I didn't dare to reimplement it 😅"
  },
  "source": "meta"
}