{
  "id": 609843,
  "title": "2nd Place Solution",
  "url": "/competitions/stanford-rna-3d-folding/writeups/2nd-place-solution",
  "author_name": "",
  "post_date": "2025-10-02T00:25:23.033Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, thank you to the organizers and Kaggle for hosting the competition, and thank you to all the participants contributing to one of the major problems in biology.</p>\n<h1>Background</h1>\n<p>I am a bioinformatician and have participated in CASP[1] since CASP11. Hence, I am familiar with processing PDB[2] data and identifying useful data to some extent (I have mainly studied proteins, so my knowledge of RNA and DNA is more limited).</p>\n<h1>Overview of the approach</h1>\n<p>My approach is primarily Template-Based Modeling (TBM), which involves identifying similar data from a dataset of known 3D structures and transferring the atomic coordinates to the prediction target. TBM was the most effective approach for protein structure prediction prior to the emergence of AlphaFold2[3] (or residue–residue contact prediction[4]; note that AF2 also incorporates TBM-like processes).</p>\n<h1>Details of the submission</h1>\n<h2>Representation-Based Sequence Search and Alignment (RBSSA)</h2>\n<p>In modern deep learning models, not only the outputs of the final layer but also those of intermediate layers can be used for various tasks. In this document, I use the term representation (although embedding may be more common) to refer to such intermediate outputs. Recently, in protein research, using representations has become a popular technique for remote homology search[5,6,7], which is an important step in TBM.</p>\n<p>I evaluated several foundation models for RNA and DNA, but some were unsuitable due to licensing restrictions. RNAErnie[8,9] and RibonanzaNet[10] were potential candidates. In preliminary experiments, RNAErnie performed slightly better, but because it required substantially more disk space due to its larger output, I chose RibonanzaNet. Since RibonanzaNet2[11] was published very recently, I did not have sufficient time to implement and evaluate it.</p>\n<p>Single representations before the decoder layer of RibonanzaNet were generated for both target and template sequences, which were then aligned using a simple dynamic programming approach (Figure 1). A high score suggests a possible evolutionary relationship, making the pair a TBM candidate. Aligning a target sequence against all cluster representatives ( 3,991 sequences) took about 160 seconds (including RibonanzaNet inference for a target sequence) for a 300-nt target sequence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1392832%2Fde7879753f40edb23c96ab89a7829130%2Fmatrixalign.png?generation=1759275990739785&amp;alt=media\" alt=\"\"><br>\nFigure 1. Schematic of Representation-Based Sequence Search and Alignment. In Smith–Waterman alignment, bases in the target and template database sequences are compared: matches receive positive scores, mismatches receive negative scores, and the highest scoring path is the optimal alignment. For alignments using matrix-like representations, match/mismatch scores are replaced with similarity functions such as cosine similarity or Pearson correlation coefficient.</p>\n<p>For RBSSA, I used my own tool[12], which was originally designed for multiple sequence alignment construction[13]. (After the competition, I updated it to improve user experience for remote homology search[14].)</p>\n<h2>Template Dataset Preparation</h2>\n<p>I downloaded RNA-containing entries from PDB on 2025-05-21. The structures were clustered by sequence similarity using MMseqs2[15] at 95% sequence identity, with some additional miselleneous filtering. This resulted in 4,779 clusters. Cluster representatives were selected based on the number of  modeled nucleotides. RibonanzaNet representation were generated for these cluster representatives to construct a RBSSA database.</p>\n<h2>Standard TBM</h2>\n<p>Because my representation-based search tool does not provide statistical confidence, it can produce many false positives. Therefore, in addition to RBSSA, I applied a standard TBM approach: searching template sequences with blastn[16] and aligning the hits using the Smith–Waterman algorithm[17].</p>\n<h2>Deep Learning-Based Structure Prediction</h2>\n<p>Since suitable templates may not always exist, or templates may have missing regions, I also utilized Chai-1[18] and Boltz-1[19]—AlphaFold3-inspired tools and current state-of-the-art deep learning-based structure predictors.</p>\n<p>If extra time was available,<br>\n・Mutated some nucleotides and fed them into predictors to generate more diverse structures.</p>\n<p>・Predicted structures using dna or rna sequences in the all_sequences column.</p>\n<p>・Predicted structures of multimeric RNA as a concatenated monomer.</p>\n<h2>Assembly Structures</h2>\n<p>Because TBM often leaves missing regions, these gaps were patched using structural alignments with models predicted by DL-based methods or other templates.. Structural alignment was performed with BioPython's SVDSuperimposer[20]. No further refinements, such as molecular dynamics, were applied due to prioritization and time constraints of the notebook.</p>\n<h1>Results &amp; Discussions</h1>\n<p>The results are shown in Table 1.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>private</th>\n<th>public (Sep.)</th>\n<th>public (May)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>full</td>\n<td>0.56125</td>\n<td>0.59655</td>\n<td>0.605</td>\n</tr>\n<tr>\n<td>full (dot)</td>\n<td>0.56966</td>\n<td>0.66171</td>\n<td>0.672</td>\n</tr>\n<tr>\n<td>-RBSSA</td>\n<td>0.50551</td>\n<td>0.44539</td>\n<td>0.46</td>\n</tr>\n<tr>\n<td>-RBSSA -TBM</td>\n<td>0.374</td>\n<td>0.32237</td>\n<td>0.37</td>\n</tr>\n</tbody>\n</table>\n<p>Table 1. Summary of results.</p>\n<p>'Public (Sep.)' refers to the scores in the Public Score column on my current submissions page, while Public (May) refers to the scores visible on the public leaderboard at the first deadline. Note that 'Public (May)' notebooks used a development version, so the overall pipeline is not consistent. 'full (dot)' indicates RBSSA with dot product as the scoring function, while 'full' uses cosine similarity. '-RBSSA' means without RBSSA, and '-TBM' means without standard TBM.</p>\n<p>The table highlights the effectiveness of both TBM and RBSSA.<br>\nHowever, the 1st place solution <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/1st-place-solution\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/1st-place-solution</a> discussions revealed that standard sequence alignment can achieve very high scores. I will further evaluate TBM variations using different sequence alignment algorithms.</p>\n<h1>Acknowledgements</h1>\n<p>I would like to thank the competition organizers and the Kaggle team for hosting this exciting challenge. I am also grateful to the Kaggle community for sharing helpful notebooks and discussions, and to the developers and maintainers of the open-source tools and public databases that I used. Finally, I appreciate the assistance of ChatGPT and DeepL in improving my English text.</p>\n<h1>References</h1>\n<p>[1] <a href=\"https://predictioncenter.org/\" target=\"_blank\">https://predictioncenter.org/</a><br>\n[2] <a href=\"https://www.rcsb.org/\" target=\"_blank\">https://www.rcsb.org/</a><br>\n[3] <a href=\"https://github.com/google-deepmind/alphafold\" target=\"_blank\">https://github.com/google-deepmind/alphafold</a><br>\n[4] <a href=\"https://en.wikipedia.org/wiki/Protein_contact_map\" target=\"_blank\">https://en.wikipedia.org/wiki/Protein_contact_map</a><br>\n[5] Kaminski, Kamil, et al. \"pLM-BLAST: distant homology detection based on direct comparison of sequence representations from protein language models.\" Bioinformatics 39.10 (2023): btad579.<br>\n[6] Pantolini, Lorenzo, et al. \"Embedding-based alignment: combining protein language models with dynamic programming alignment to detect structural similarities in the twilight-zone.\" Bioinformatics 40.1 (2024): btad786.<br>\n[7] Liu, Wei, et al. \"PLMSearch: Protein language model powers accurate and fast sequence search for remote homology.\"&nbsp;Nature communications&nbsp;15.1 (2024): 2775.<br>\n[8] <a href=\"https://github.com/CatIIIIIIII/RNAErnie\" target=\"_blank\">https://github.com/CatIIIIIIII/RNAErnie</a><br>\n[9] <a href=\"https://huggingface.co/multimolecule/rnaernie\" target=\"_blank\">https://huggingface.co/multimolecule/rnaernie</a><br>\n[10] <a href=\"https://github.com/Shujun-He/RibonanzaNet\" target=\"_blank\">https://github.com/Shujun-He/RibonanzaNet</a><br>\n[11] <a href=\"https://www.kaggle.com/models/shujun717/ribonanzanet2\" target=\"_blank\">https://www.kaggle.com/models/shujun717/ribonanzanet2</a><br>\n[12] <a href=\"https://github.com/yamule/matrix_align/tree/523262de8fc5da274478d3d31e46a089991056a2\" target=\"_blank\">https://github.com/yamule/matrix_align/tree/523262de8fc5da274478d3d31e46a089991056a2</a><br>\n[13] <a href=\"https://en.wikipedia.org/wiki/Multiple_sequence_alignment\" target=\"_blank\">https://en.wikipedia.org/wiki/Multiple_sequence_alignment</a><br>\n[14] <a href=\"https://github.com/yamule/matrix_align/commit/ad04b715ec293b8dffa16bb74506d49bd3294e05\" target=\"_blank\">https://github.com/yamule/matrix_align/commit/ad04b715ec293b8dffa16bb74506d49bd3294e05</a><br>\n[15] <a href=\"https://github.com/soedinglab/MMseqs2/\" target=\"_blank\">https://github.com/soedinglab/MMseqs2/</a><br>\n[16] <a href=\"https://blast.ncbi.nlm.nih.gov/doc/blast-help/downloadblastdata.html\" target=\"_blank\">https://blast.ncbi.nlm.nih.gov/doc/blast-help/downloadblastdata.html</a><br>\n[17] <a href=\"https://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm\" target=\"_blank\">https://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm</a><br>\n[18] <a href=\"https://github.com/chaidiscovery/chai-lab\" target=\"_blank\">https://github.com/chaidiscovery/chai-lab</a><br>\n[19] <a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">https://github.com/jwohlwend/boltz</a><br>\n[20] <a href=\"https://biopython.org/docs/1.85/api/Bio.SVDSuperimposer.html\" target=\"_blank\">https://biopython.org/docs/1.85/api/Bio.SVDSuperimposer.html</a></p>\n<p>The URLs were accessed on 2025-10-02.</p>",
  "messages": [
    {
      "id": "3296020",
      "postDate": "09/30/2025 01:52:32",
      "content": "<p>First of all, thank you to the organizers and Kaggle for hosting the competition, and thank you to all the participants contributing to one of the major problems in biology.</p>\n<h1>Background</h1>\n<p>I am a bioinformatician and have participated in CASP[1] since CASP11. Hence, I am familiar with processing PDB[2] data and identifying useful data to some extent (I have mainly studied proteins, so my knowledge of RNA and DNA is more limited).</p>\n<h1>Overview of the approach</h1>\n<p>My approach is primarily Template-Based Modeling (TBM), which involves identifying similar data from a dataset of known 3D structures and transferring the atomic coordinates to the prediction target. TBM was the most effective approach for protein structure prediction prior to the emergence of AlphaFold2[3] (or residue–residue contact prediction[4]; note that AF2 also incorporates TBM-like processes).</p>\n<h1>Details of the submission</h1>\n<h2>Representation-Based Sequence Search and Alignment (RBSSA)</h2>\n<p>In modern deep learning models, not only the outputs of the final layer but also those of intermediate layers can be used for various tasks. In this document, I use the term representation (although embedding may be more common) to refer to such intermediate outputs. Recently, in protein research, using representations has become a popular technique for remote homology search[5,6,7], which is an important step in TBM.</p>\n<p>I evaluated several foundation models for RNA and DNA, but some were unsuitable due to licensing restrictions. RNAErnie[8,9] and RibonanzaNet[10] were potential candidates. In preliminary experiments, RNAErnie performed slightly better, but because it required substantially more disk space due to its larger output, I chose RibonanzaNet. Since RibonanzaNet2[11] was published very recently, I did not have sufficient time to implement and evaluate it.</p>\n<p>Single representations before the decoder layer of RibonanzaNet were generated for both target and template sequences, which were then aligned using a simple dynamic programming approach (Figure 1). A high score suggests a possible evolutionary relationship, making the pair a TBM candidate. Aligning a target sequence against all cluster representatives ( 3,991 sequences) took about 160 seconds (including RibonanzaNet inference for a target sequence) for a 300-nt target sequence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1392832%2Fde7879753f40edb23c96ab89a7829130%2Fmatrixalign.png?generation=1759275990739785&amp;alt=media\" alt=\"\"><br>\nFigure 1. Schematic of Representation-Based Sequence Search and Alignment. In Smith–Waterman alignment, bases in the target and template database sequences are compared: matches receive positive scores, mismatches receive negative scores, and the highest scoring path is the optimal alignment. For alignments using matrix-like representations, match/mismatch scores are replaced with similarity functions such as cosine similarity or Pearson correlation coefficient.</p>\n<p>For RBSSA, I used my own tool[12], which was originally designed for multiple sequence alignment construction[13]. (After the competition, I updated it to improve user experience for remote homology search[14].)</p>\n<h2>Template Dataset Preparation</h2>\n<p>I downloaded RNA-containing entries from PDB on 2025-05-21. The structures were clustered by sequence similarity using MMseqs2[15] at 95% sequence identity, with some additional miselleneous filtering. This resulted in 4,779 clusters. Cluster representatives were selected based on the number of  modeled nucleotides. RibonanzaNet representation were generated for these cluster representatives to construct a RBSSA database.</p>\n<h2>Standard TBM</h2>\n<p>Because my representation-based search tool does not provide statistical confidence, it can produce many false positives. Therefore, in addition to RBSSA, I applied a standard TBM approach: searching template sequences with blastn[16] and aligning the hits using the Smith–Waterman algorithm[17].</p>\n<h2>Deep Learning-Based Structure Prediction</h2>\n<p>Since suitable templates may not always exist, or templates may have missing regions, I also utilized Chai-1[18] and Boltz-1[19]—AlphaFold3-inspired tools and current state-of-the-art deep learning-based structure predictors.</p>\n<p>If extra time was available,<br>\n・Mutated some nucleotides and fed them into predictors to generate more diverse structures.</p>\n<p>・Predicted structures using dna or rna sequences in the all_sequences column.</p>\n<p>・Predicted structures of multimeric RNA as a concatenated monomer.</p>\n<h2>Assembly Structures</h2>\n<p>Because TBM often leaves missing regions, these gaps were patched using structural alignments with models predicted by DL-based methods or other templates.. Structural alignment was performed with BioPython's SVDSuperimposer[20]. No further refinements, such as molecular dynamics, were applied due to prioritization and time constraints of the notebook.</p>\n<h1>Results &amp; Discussions</h1>\n<p>The results are shown in Table 1.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>private</th>\n<th>public (Sep.)</th>\n<th>public (May)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>full</td>\n<td>0.56125</td>\n<td>0.59655</td>\n<td>0.605</td>\n</tr>\n<tr>\n<td>full (dot)</td>\n<td>0.56966</td>\n<td>0.66171</td>\n<td>0.672</td>\n</tr>\n<tr>\n<td>-RBSSA</td>\n<td>0.50551</td>\n<td>0.44539</td>\n<td>0.46</td>\n</tr>\n<tr>\n<td>-RBSSA -TBM</td>\n<td>0.374</td>\n<td>0.32237</td>\n<td>0.37</td>\n</tr>\n</tbody>\n</table>\n<p>Table 1. Summary of results.</p>\n<p>'Public (Sep.)' refers to the scores in the Public Score column on my current submissions page, while Public (May) refers to the scores visible on the public leaderboard at the first deadline. Note that 'Public (May)' notebooks used a development version, so the overall pipeline is not consistent. 'full (dot)' indicates RBSSA with dot product as the scoring function, while 'full' uses cosine similarity. '-RBSSA' means without RBSSA, and '-TBM' means without standard TBM.</p>\n<p>The table highlights the effectiveness of both TBM and RBSSA.<br>\nHowever, the 1st place solution <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/1st-place-solution\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/1st-place-solution</a> discussions revealed that standard sequence alignment can achieve very high scores. I will further evaluate TBM variations using different sequence alignment algorithms.</p>\n<h1>Acknowledgements</h1>\n<p>I would like to thank the competition organizers and the Kaggle team for hosting this exciting challenge. I am also grateful to the Kaggle community for sharing helpful notebooks and discussions, and to the developers and maintainers of the open-source tools and public databases that I used. Finally, I appreciate the assistance of ChatGPT and DeepL in improving my English text.</p>\n<h1>References</h1>\n<p>[1] <a href=\"https://predictioncenter.org/\" target=\"_blank\">https://predictioncenter.org/</a><br>\n[2] <a href=\"https://www.rcsb.org/\" target=\"_blank\">https://www.rcsb.org/</a><br>\n[3] <a href=\"https://github.com/google-deepmind/alphafold\" target=\"_blank\">https://github.com/google-deepmind/alphafold</a><br>\n[4] <a href=\"https://en.wikipedia.org/wiki/Protein_contact_map\" target=\"_blank\">https://en.wikipedia.org/wiki/Protein_contact_map</a><br>\n[5] Kaminski, Kamil, et al. \"pLM-BLAST: distant homology detection based on direct comparison of sequence representations from protein language models.\" Bioinformatics 39.10 (2023): btad579.<br>\n[6] Pantolini, Lorenzo, et al. \"Embedding-based alignment: combining protein language models with dynamic programming alignment to detect structural similarities in the twilight-zone.\" Bioinformatics 40.1 (2024): btad786.<br>\n[7] Liu, Wei, et al. \"PLMSearch: Protein language model powers accurate and fast sequence search for remote homology.\"&nbsp;Nature communications&nbsp;15.1 (2024): 2775.<br>\n[8] <a href=\"https://github.com/CatIIIIIIII/RNAErnie\" target=\"_blank\">https://github.com/CatIIIIIIII/RNAErnie</a><br>\n[9] <a href=\"https://huggingface.co/multimolecule/rnaernie\" target=\"_blank\">https://huggingface.co/multimolecule/rnaernie</a><br>\n[10] <a href=\"https://github.com/Shujun-He/RibonanzaNet\" target=\"_blank\">https://github.com/Shujun-He/RibonanzaNet</a><br>\n[11] <a href=\"https://www.kaggle.com/models/shujun717/ribonanzanet2\" target=\"_blank\">https://www.kaggle.com/models/shujun717/ribonanzanet2</a><br>\n[12] <a href=\"https://github.com/yamule/matrix_align/tree/523262de8fc5da274478d3d31e46a089991056a2\" target=\"_blank\">https://github.com/yamule/matrix_align/tree/523262de8fc5da274478d3d31e46a089991056a2</a><br>\n[13] <a href=\"https://en.wikipedia.org/wiki/Multiple_sequence_alignment\" target=\"_blank\">https://en.wikipedia.org/wiki/Multiple_sequence_alignment</a><br>\n[14] <a href=\"https://github.com/yamule/matrix_align/commit/ad04b715ec293b8dffa16bb74506d49bd3294e05\" target=\"_blank\">https://github.com/yamule/matrix_align/commit/ad04b715ec293b8dffa16bb74506d49bd3294e05</a><br>\n[15] <a href=\"https://github.com/soedinglab/MMseqs2/\" target=\"_blank\">https://github.com/soedinglab/MMseqs2/</a><br>\n[16] <a href=\"https://blast.ncbi.nlm.nih.gov/doc/blast-help/downloadblastdata.html\" target=\"_blank\">https://blast.ncbi.nlm.nih.gov/doc/blast-help/downloadblastdata.html</a><br>\n[17] <a href=\"https://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm\" target=\"_blank\">https://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm</a><br>\n[18] <a href=\"https://github.com/chaidiscovery/chai-lab\" target=\"_blank\">https://github.com/chaidiscovery/chai-lab</a><br>\n[19] <a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">https://github.com/jwohlwend/boltz</a><br>\n[20] <a href=\"https://biopython.org/docs/1.85/api/Bio.SVDSuperimposer.html\" target=\"_blank\">https://biopython.org/docs/1.85/api/Bio.SVDSuperimposer.html</a></p>\n<p>The URLs were accessed on 2025-10-02.</p>",
      "rawMarkdown": "First of all, thank you to the organizers and Kaggle for hosting the competition, and thank you to all the participants contributing to one of the major problems in biology.\n\n# Background\nI am a bioinformatician and have participated in CASP[1] since CASP11. Hence, I am familiar with processing PDB[2] data and identifying useful data to some extent (I have mainly studied proteins, so my knowledge of RNA and DNA is more limited).\n\n#  Overview of the approach \nMy approach is primarily Template-Based Modeling (TBM), which involves identifying similar data from a dataset of known 3D structures and transferring the atomic coordinates to the prediction target. TBM was the most effective approach for protein structure prediction prior to the emergence of AlphaFold2[3] (or residue–residue contact prediction[4]; note that AF2 also incorporates TBM-like processes).\n\n# Details of the submission \n## Representation-Based Sequence Search and Alignment (RBSSA)\nIn modern deep learning models, not only the outputs of the final layer but also those of intermediate layers can be used for various tasks. In this document, I use the term representation (although embedding may be more common) to refer to such intermediate outputs. Recently, in protein research, using representations has become a popular technique for remote homology search[5,6,7], which is an important step in TBM.\n\nI evaluated several foundation models for RNA and DNA, but some were unsuitable due to licensing restrictions. RNAErnie[8,9] and RibonanzaNet[10] were potential candidates. In preliminary experiments, RNAErnie performed slightly better, but because it required substantially more disk space due to its larger output, I chose RibonanzaNet. Since RibonanzaNet2[11] was published very recently, I did not have sufficient time to implement and evaluate it.\n\nSingle representations before the decoder layer of RibonanzaNet were generated for both target and template sequences, which were then aligned using a simple dynamic programming approach (Figure 1). A high score suggests a possible evolutionary relationship, making the pair a TBM candidate. Aligning a target sequence against all cluster representatives (~~4,779~~ 3,991 sequences) took about 160 seconds (including RibonanzaNet inference for a target sequence) for a 300-nt target sequence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1392832%2Fde7879753f40edb23c96ab89a7829130%2Fmatrixalign.png?generation=1759275990739785&alt=media)\nFigure 1. Schematic of Representation-Based Sequence Search and Alignment. In Smith–Waterman alignment, bases in the target and template database sequences are compared: matches receive positive scores, mismatches receive negative scores, and the highest scoring path is the optimal alignment. For alignments using matrix-like representations, match/mismatch scores are replaced with similarity functions such as cosine similarity or Pearson correlation coefficient.\n\nFor RBSSA, I used my own tool[12], which was originally designed for multiple sequence alignment construction[13]. (After the competition, I updated it to improve user experience for remote homology search[14].)\n\n\n## Template Dataset Preparation\nI downloaded RNA-containing entries from PDB on 2025-05-21. The structures were clustered by sequence similarity using MMseqs2[15] at 95% sequence identity, with some additional miselleneous filtering. This resulted in 4,779 clusters. Cluster representatives were selected based on the number of ~~C1' atoms~~ modeled nucleotides. RibonanzaNet representation were generated for these cluster representatives to construct a RBSSA database.\n\n## Standard TBM\nBecause my representation-based search tool does not provide statistical confidence, it can produce many false positives. Therefore, in addition to RBSSA, I applied a standard TBM approach: searching template sequences with blastn[16] and aligning the hits using the Smith–Waterman algorithm[17].\n\n## Deep Learning-Based Structure Prediction\nSince suitable templates may not always exist, or templates may have missing regions, I also utilized Chai-1[18] and Boltz-1[19]—AlphaFold3-inspired tools and current state-of-the-art deep learning-based structure predictors.\n\nIf extra time was available,\n・Mutated some nucleotides and fed them into predictors to generate more diverse structures.\n\n・Predicted structures using dna or rna sequences in the all_sequences column.\n\n・Predicted structures of multimeric RNA as a concatenated monomer.\n\n## Assembly Structures\nBecause TBM often leaves missing regions, these gaps were patched using structural alignments with models predicted by DL-based methods or other templates.. Structural alignment was performed with BioPython's SVDSuperimposer[20]. No further refinements, such as molecular dynamics, were applied due to prioritization and time constraints of the notebook.\n\n# Results & Discussions\nThe results are shown in Table 1.\n\n|  | private | public (Sep.) | public (May) |\n| --- | --- | --- | --- |\n| full | 0.56125 | 0.59655 | 0.605 |\n| full (dot) | 0.56966 | 0.66171 | 0.672 |\n| -RBSSA | 0.50551 | 0.44539 | 0.46 |\n| -RBSSA -TBM | 0.374 | 0.32237 | 0.37 |\n\n\nTable 1. Summary of results.\n\n'Public (Sep.)' refers to the scores in the Public Score column on my current submissions page, while Public (May) refers to the scores visible on the public leaderboard at the first deadline. Note that 'Public (May)' notebooks used a development version, so the overall pipeline is not consistent. 'full (dot)' indicates RBSSA with dot product as the scoring function, while 'full' uses cosine similarity. '-RBSSA' means without RBSSA, and '-TBM' means without standard TBM.\n\nThe table highlights the effectiveness of both TBM and RBSSA.\nHowever, the 1st place solution https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/1st-place-solution discussions revealed that standard sequence alignment can achieve very high scores. I will further evaluate TBM variations using different sequence alignment algorithms.\n\n# Acknowledgements\nI would like to thank the competition organizers and the Kaggle team for hosting this exciting challenge. I am also grateful to the Kaggle community for sharing helpful notebooks and discussions, and to the developers and maintainers of the open-source tools and public databases that I used. Finally, I appreciate the assistance of ChatGPT and DeepL in improving my English text.\n\n# References\n[1] https://predictioncenter.org/\n[2] https://www.rcsb.org/\n[3] https://github.com/google-deepmind/alphafold\n[4] https://en.wikipedia.org/wiki/Protein_contact_map\n[5] Kaminski, Kamil, et al. \"pLM-BLAST: distant homology detection based on direct comparison of sequence representations from protein language models.\" Bioinformatics 39.10 (2023): btad579.\n[6] Pantolini, Lorenzo, et al. \"Embedding-based alignment: combining protein language models with dynamic programming alignment to detect structural similarities in the twilight-zone.\" Bioinformatics 40.1 (2024): btad786.\n[7] Liu, Wei, et al. \"PLMSearch: Protein language model powers accurate and fast sequence search for remote homology.\" Nature communications 15.1 (2024): 2775.\n[8] https://github.com/CatIIIIIIII/RNAErnie\n[9] https://huggingface.co/multimolecule/rnaernie\n[10] https://github.com/Shujun-He/RibonanzaNet\n[11] https://www.kaggle.com/models/shujun717/ribonanzanet2\n[12] https://github.com/yamule/matrix_align/tree/523262de8fc5da274478d3d31e46a089991056a2\n[13] https://en.wikipedia.org/wiki/Multiple_sequence_alignment\n[14] https://github.com/yamule/matrix_align/commit/ad04b715ec293b8dffa16bb74506d49bd3294e05\n[15] https://github.com/soedinglab/MMseqs2/\n[16] https://blast.ncbi.nlm.nih.gov/doc/blast-help/downloadblastdata.html\n[17] https://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm\n[18] https://github.com/chaidiscovery/chai-lab\n[19] https://github.com/jwohlwend/boltz\n[20] https://biopython.org/docs/1.85/api/Bio.SVDSuperimposer.html\n\nThe URLs were accessed on 2025-10-02.",
      "votes": null
    },
    {
      "id": "3299127",
      "postDate": "10/07/2025 09:47:02",
      "content": "<p>Update 2025/10/07<br>\n\"4,779\" was the number of the DB sequences for the other submission (submit2). The correct number is 3,991. The representation search for 4,779  sequences took about 230 seconds.</p>",
      "rawMarkdown": "Update 2025/10/07\n\"4,779\" was the number of the DB sequences for the other submission (submit2). The correct number is 3,991. The representation search for 4,779  sequences took about 230 seconds.",
      "votes": null
    },
    {
      "id": "3299174",
      "postDate": "10/07/2025 12:10:46",
      "content": "<p>Update 2025/10/07 B<br>\nCluster representative was selected based on the number of nucleotides which can be obtained with get_list() of Biopython's Chain object.<br>\n<a href=\"https://biopython.org/docs/1.85/api/Bio.PDB.Chain.html\" target=\"_blank\">https://biopython.org/docs/1.85/api/Bio.PDB.Chain.html</a></p>",
      "rawMarkdown": "Update 2025/10/07 B\nCluster representative was selected based on the number of nucleotides which can be obtained with get_list() of Biopython's Chain object.\nhttps://biopython.org/docs/1.85/api/Bio.PDB.Chain.html",
      "votes": null
    },
    {
      "id": "3299175",
      "postDate": "10/07/2025 12:11:13",
      "content": "<p>Congrats and thank you for sharing the solution, very well deserved!</p>\n<p>The Representation-Based Sequence Search and Alignment is very interesting. You mentioned that you evaluated several foundation models for RNA and DNA, I was wondering if you happened to test AIDO.RNA, and if so, how did it perform?</p>\n<p>Previously, I tried using rMSA for MSA searches, but many cases failed. I’m now considering trying a representation-based MSA search instead. </p>",
      "rawMarkdown": "Congrats and thank you for sharing the solution, very well deserved!\n\nThe Representation-Based Sequence Search and Alignment is very interesting. You mentioned that you evaluated several foundation models for RNA and DNA, I was wondering if you happened to test AIDO.RNA, and if so, how did it perform?\n\nPreviously, I tried using rMSA for MSA searches, but many cases failed. I’m now considering trying a representation-based MSA search instead.",
      "votes": null
    },
    {
      "id": "3299379",
      "postDate": "10/07/2025 23:42:55",
      "content": "<p>No. I didn't try to use AIDO.RNA because of the license issue. The official implementation is for non-commercial <a href=\"https://huggingface.co/genbio-ai/AIDO.RNA-1.6B/blob/main/LICENSE\" target=\"_blank\">https://huggingface.co/genbio-ai/AIDO.RNA-1.6B/blob/main/LICENSE</a> &amp; multimolecule implementation is AGPL <a href=\"https://huggingface.co/multimolecule/aido.rna-1.6b-cds\" target=\"_blank\">https://huggingface.co/multimolecule/aido.rna-1.6b-cds</a> . I thought AGPL was not recommended.</p>",
      "rawMarkdown": "No. I didn't try to use AIDO.RNA because of the license issue. The official implementation is for non-commercial https://huggingface.co/genbio-ai/AIDO.RNA-1.6B/blob/main/LICENSE & multimolecule implementation is AGPL https://huggingface.co/multimolecule/aido.rna-1.6b-cds . I thought AGPL was not recommended.",
      "votes": null
    },
    {
      "id": "3299502",
      "postDate": "10/08/2025 09:05:39",
      "content": "<p>\"a representation-based MSA search\" is very interesting. One problem of representation-based search is \"storage space\". One nucleotide can be stored in 1 byte (or 2 or 3 bits) while a representation for one nucleotide may require 256*float16 = 512 bytes or larger. I don't have any good idea about this.<br>\nI wish you success.</p>",
      "rawMarkdown": "\"a representation-based MSA search\" is very interesting. One problem of representation-based search is \"storage space\". One nucleotide can be stored in 1 byte (or 2 or 3 bits) while a representation for one nucleotide may require 256*float16 = 512 bytes or larger. I don't have any good idea about this.\nI wish you success.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3299127,
      "author_name": "odat1248",
      "author_url": "",
      "post_date": "10/07/2025 09:47:02",
      "content": "<p>Update 2025/10/07<br>\n\"4,779\" was the number of the DB sequences for the other submission (submit2). The correct number is 3,991. The representation search for 4,779  sequences took about 230 seconds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3299174,
          "author_name": "odat1248",
          "author_url": "",
          "post_date": "10/07/2025 12:10:46",
          "content": "<p>Update 2025/10/07 B<br>\nCluster representative was selected based on the number of nucleotides which can be obtained with get_list() of Biopython's Chain object.<br>\n<a href=\"https://biopython.org/docs/1.85/api/Bio.PDB.Chain.html\" target=\"_blank\">https://biopython.org/docs/1.85/api/Bio.PDB.Chain.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3299175,
      "author_name": "zoushuxian",
      "author_url": "",
      "post_date": "10/07/2025 12:11:13",
      "content": "<p>Congrats and thank you for sharing the solution, very well deserved!</p>\n<p>The Representation-Based Sequence Search and Alignment is very interesting. You mentioned that you evaluated several foundation models for RNA and DNA, I was wondering if you happened to test AIDO.RNA, and if so, how did it perform?</p>\n<p>Previously, I tried using rMSA for MSA searches, but many cases failed. I’m now considering trying a representation-based MSA search instead. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3299379,
          "author_name": "odat1248",
          "author_url": "",
          "post_date": "10/07/2025 23:42:55",
          "content": "<p>No. I didn't try to use AIDO.RNA because of the license issue. The official implementation is for non-commercial <a href=\"https://huggingface.co/genbio-ai/AIDO.RNA-1.6B/blob/main/LICENSE\" target=\"_blank\">https://huggingface.co/genbio-ai/AIDO.RNA-1.6B/blob/main/LICENSE</a> &amp; multimolecule implementation is AGPL <a href=\"https://huggingface.co/multimolecule/aido.rna-1.6b-cds\" target=\"_blank\">https://huggingface.co/multimolecule/aido.rna-1.6b-cds</a> . I thought AGPL was not recommended.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3299502,
          "author_name": "odat1248",
          "author_url": "",
          "post_date": "10/08/2025 09:05:39",
          "content": "<p>\"a representation-based MSA search\" is very interesting. One problem of representation-based search is \"storage space\". One nucleotide can be stored in 1 byte (or 2 or 3 bits) while a representation for one nucleotide may require 256*float16 = 512 bytes or larger. I don't have any good idea about this.<br>\nI wish you success.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3296020": "First of all, thank you to the organizers and Kaggle for hosting the competition, and thank you to all the participants contributing to one of the major problems in biology.\n\n# Background\nI am a bioinformatician and have participated in CASP[1] since CASP11. Hence, I am familiar with processing PDB[2] data and identifying useful data to some extent (I have mainly studied proteins, so my knowledge of RNA and DNA is more limited).\n\n#  Overview of the approach \nMy approach is primarily Template-Based Modeling (TBM), which involves identifying similar data from a dataset of known 3D structures and transferring the atomic coordinates to the prediction target. TBM was the most effective approach for protein structure prediction prior to the emergence of AlphaFold2[3] (or residue–residue contact prediction[4]; note that AF2 also incorporates TBM-like processes).\n\n# Details of the submission \n## Representation-Based Sequence Search and Alignment (RBSSA)\nIn modern deep learning models, not only the outputs of the final layer but also those of intermediate layers can be used for various tasks. In this document, I use the term representation (although embedding may be more common) to refer to such intermediate outputs. Recently, in protein research, using representations has become a popular technique for remote homology search[5,6,7], which is an important step in TBM.\n\nI evaluated several foundation models for RNA and DNA, but some were unsuitable due to licensing restrictions. RNAErnie[8,9] and RibonanzaNet[10] were potential candidates. In preliminary experiments, RNAErnie performed slightly better, but because it required substantially more disk space due to its larger output, I chose RibonanzaNet. Since RibonanzaNet2[11] was published very recently, I did not have sufficient time to implement and evaluate it.\n\nSingle representations before the decoder layer of RibonanzaNet were generated for both target and template sequences, which were then aligned using a simple dynamic programming approach (Figure 1). A high score suggests a possible evolutionary relationship, making the pair a TBM candidate. Aligning a target sequence against all cluster representatives (~~4,779~~ 3,991 sequences) took about 160 seconds (including RibonanzaNet inference for a target sequence) for a 300-nt target sequence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1392832%2Fde7879753f40edb23c96ab89a7829130%2Fmatrixalign.png?generation=1759275990739785&alt=media)\nFigure 1. Schematic of Representation-Based Sequence Search and Alignment. In Smith–Waterman alignment, bases in the target and template database sequences are compared: matches receive positive scores, mismatches receive negative scores, and the highest scoring path is the optimal alignment. For alignments using matrix-like representations, match/mismatch scores are replaced with similarity functions such as cosine similarity or Pearson correlation coefficient.\n\nFor RBSSA, I used my own tool[12], which was originally designed for multiple sequence alignment construction[13]. (After the competition, I updated it to improve user experience for remote homology search[14].)\n\n\n## Template Dataset Preparation\nI downloaded RNA-containing entries from PDB on 2025-05-21. The structures were clustered by sequence similarity using MMseqs2[15] at 95% sequence identity, with some additional miselleneous filtering. This resulted in 4,779 clusters. Cluster representatives were selected based on the number of ~~C1' atoms~~ modeled nucleotides. RibonanzaNet representation were generated for these cluster representatives to construct a RBSSA database.\n\n## Standard TBM\nBecause my representation-based search tool does not provide statistical confidence, it can produce many false positives. Therefore, in addition to RBSSA, I applied a standard TBM approach: searching template sequences with blastn[16] and aligning the hits using the Smith–Waterman algorithm[17].\n\n## Deep Learning-Based Structure Prediction\nSince suitable templates may not always exist, or templates may have missing regions, I also utilized Chai-1[18] and Boltz-1[19]—AlphaFold3-inspired tools and current state-of-the-art deep learning-based structure predictors.\n\nIf extra time was available,\n・Mutated some nucleotides and fed them into predictors to generate more diverse structures.\n\n・Predicted structures using dna or rna sequences in the all_sequences column.\n\n・Predicted structures of multimeric RNA as a concatenated monomer.\n\n## Assembly Structures\nBecause TBM often leaves missing regions, these gaps were patched using structural alignments with models predicted by DL-based methods or other templates.. Structural alignment was performed with BioPython's SVDSuperimposer[20]. No further refinements, such as molecular dynamics, were applied due to prioritization and time constraints of the notebook.\n\n# Results & Discussions\nThe results are shown in Table 1.\n\n|  | private | public (Sep.) | public (May) |\n| --- | --- | --- | --- |\n| full | 0.56125 | 0.59655 | 0.605 |\n| full (dot) | 0.56966 | 0.66171 | 0.672 |\n| -RBSSA | 0.50551 | 0.44539 | 0.46 |\n| -RBSSA -TBM | 0.374 | 0.32237 | 0.37 |\n\n\nTable 1. Summary of results.\n\n'Public (Sep.)' refers to the scores in the Public Score column on my current submissions page, while Public (May) refers to the scores visible on the public leaderboard at the first deadline. Note that 'Public (May)' notebooks used a development version, so the overall pipeline is not consistent. 'full (dot)' indicates RBSSA with dot product as the scoring function, while 'full' uses cosine similarity. '-RBSSA' means without RBSSA, and '-TBM' means without standard TBM.\n\nThe table highlights the effectiveness of both TBM and RBSSA.\nHowever, the 1st place solution https://www.kaggle.com/competitions/stanford-rna-3d-folding/writeups/1st-place-solution discussions revealed that standard sequence alignment can achieve very high scores. I will further evaluate TBM variations using different sequence alignment algorithms.\n\n# Acknowledgements\nI would like to thank the competition organizers and the Kaggle team for hosting this exciting challenge. I am also grateful to the Kaggle community for sharing helpful notebooks and discussions, and to the developers and maintainers of the open-source tools and public databases that I used. Finally, I appreciate the assistance of ChatGPT and DeepL in improving my English text.\n\n# References\n[1] https://predictioncenter.org/\n[2] https://www.rcsb.org/\n[3] https://github.com/google-deepmind/alphafold\n[4] https://en.wikipedia.org/wiki/Protein_contact_map\n[5] Kaminski, Kamil, et al. \"pLM-BLAST: distant homology detection based on direct comparison of sequence representations from protein language models.\" Bioinformatics 39.10 (2023): btad579.\n[6] Pantolini, Lorenzo, et al. \"Embedding-based alignment: combining protein language models with dynamic programming alignment to detect structural similarities in the twilight-zone.\" Bioinformatics 40.1 (2024): btad786.\n[7] Liu, Wei, et al. \"PLMSearch: Protein language model powers accurate and fast sequence search for remote homology.\" Nature communications 15.1 (2024): 2775.\n[8] https://github.com/CatIIIIIIII/RNAErnie\n[9] https://huggingface.co/multimolecule/rnaernie\n[10] https://github.com/Shujun-He/RibonanzaNet\n[11] https://www.kaggle.com/models/shujun717/ribonanzanet2\n[12] https://github.com/yamule/matrix_align/tree/523262de8fc5da274478d3d31e46a089991056a2\n[13] https://en.wikipedia.org/wiki/Multiple_sequence_alignment\n[14] https://github.com/yamule/matrix_align/commit/ad04b715ec293b8dffa16bb74506d49bd3294e05\n[15] https://github.com/soedinglab/MMseqs2/\n[16] https://blast.ncbi.nlm.nih.gov/doc/blast-help/downloadblastdata.html\n[17] https://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm\n[18] https://github.com/chaidiscovery/chai-lab\n[19] https://github.com/jwohlwend/boltz\n[20] https://biopython.org/docs/1.85/api/Bio.SVDSuperimposer.html\n\nThe URLs were accessed on 2025-10-02.",
    "3299127": "Update 2025/10/07\n\"4,779\" was the number of the DB sequences for the other submission (submit2). The correct number is 3,991. The representation search for 4,779  sequences took about 230 seconds.",
    "3299174": "Update 2025/10/07 B\nCluster representative was selected based on the number of nucleotides which can be obtained with get_list() of Biopython's Chain object.\nhttps://biopython.org/docs/1.85/api/Bio.PDB.Chain.html",
    "3299175": "Congrats and thank you for sharing the solution, very well deserved!\n\nThe Representation-Based Sequence Search and Alignment is very interesting. You mentioned that you evaluated several foundation models for RNA and DNA, I was wondering if you happened to test AIDO.RNA, and if so, how did it perform?\n\nPreviously, I tried using rMSA for MSA searches, but many cases failed. I’m now considering trying a representation-based MSA search instead.",
    "3299379": "No. I didn't try to use AIDO.RNA because of the license issue. The official implementation is for non-commercial https://huggingface.co/genbio-ai/AIDO.RNA-1.6B/blob/main/LICENSE & multimolecule implementation is AGPL https://huggingface.co/multimolecule/aido.rna-1.6b-cds . I thought AGPL was not recommended.",
    "3299502": "\"a representation-based MSA search\" is very interesting. One problem of representation-based search is \"storage space\". One nucleotide can be stored in 1 byte (or 2 or 3 bits) while a representation for one nucleotide may require 256*float16 = 512 bytes or larger. I don't have any good idea about this.\nI wish you success."
  },
  "source": "meta"
}