{
  "id": 609774,
  "title": "1st Place Solution",
  "url": "/competitions/stanford-rna-3d-folding/writeups/1st-place-solution",
  "author_name": "",
  "post_date": "2025-09-29T15:29:44.213Z",
  "votes": 97,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Thank you Kaggle and the competition hosts, for this incredible competition and the opportunity to participate. Your passion for this challenge was truly infectious and served as one of my driving forces throughout the competition. </p>\n<p>Since this is my first gold and my first win on Kaggle, I would like to take the opportunity to thank <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> for the fast.ai course, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for their book which was an inspiration in my ML journey. Khan Academy for helping me rethink mathematics, and <a href=\"https://www.kaggle.com/huggingface\" target=\"_blank\">@huggingface</a> for their amazing deep learning courses. </p>\n<h2>Competition Strategy</h2>\n<p>My approach was clear from the outset. Without GPUs, training a model from scratch or fine-tuning was not viable. My early research - drawing on CASP results, literature, and conference talks, including one by host <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> - showed that Template-Based Modeling approaches consistently dominated. Based on this, I committed to TBM from day one and spent the next 90 days refining my method.</p>\n<p>Next, I focused on the evaluation metric, since understanding it determines the exploration path. TM-score has two key properties: it is normalized by structure length (so 50nt and 200nt RNAs are compared on the same 0-1 scale), and it is robust to local errors - a small number of misplaced nucleotides does not disproportionately lower the score. This insight allowed me to prioritize getting the overall fold correct over achieving atomic-level precision.</p>\n<h2>Data Strategy and Model Selection</h2>\n<p>The host-provided dataset was comprehensive. I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction. This process ensured complete coverage of the available structural data, capturing modified nucleotides that standard parsers might otherwise miss.</p>\n<p>After exploring nearly all available open-source models, I selected DRfold2 as the optimal choice due to its extensive potential for optimization. Rather than fine-tuning the model itself, I focused on enhancing its optimization and selection modules. This strategy improved prediction quality while ensuring the pipeline could execute efficiently on Kaggle GPUs.</p>\n<h2>Template-Based Modeling (TBM)</h2>\n<p>TBM follows a five-step process:</p>\n<p><strong>1. The Search - Finding Similar Structures</strong></p>\n<p>The goal is identifying database structures that resemble the target sequence through sequence alignment.</p>\n<p><strong>2. The Alignment - Sequence Mapping</strong></p>\n<p>This step creates the translation guide between query and template using global sequence alignment with gap penalties optimized for RNA. </p>\n<p><strong>3. The Transfer - Coordinate Inheritance</strong></p>\n<p>Straightforward copying of 3D coordinates for all matched positions. This leverages the evolutionary tendency for RNA to conserve 3D structure more than sequence. </p>\n<p><strong>4. The Gap Fill - Geometric Backbone Reconstruction</strong></p>\n<p>For insertions and deletions, I relied on geometric principles maintaining RNA's characteristic backbone:</p>\n<ul>\n<li>Maintains <code>C1'-C1'</code> distance (<code>~5.9Å</code> between consecutive nucleotides)</li>\n<li>For compressed gaps: extends the backbone with realistic curvature using sinusoidal perturbations perpendicular to the backbone direction</li>\n<li>For normal gaps: uses linear interpolation between flanking coordinates</li>\n<li>Terminal extensions follow the established backbone direction</li>\n</ul>\n<p><strong>5. Adaptive Refinement - Confidence-Based Optimization</strong></p>\n<p>The refinement intensity adapts to template confidence score:</p>\n<ul>\n<li>High-confidence templates (&gt;0.8 similarity): minimal constraints, preserving template geometry</li>\n<li>Medium-confidence templates: moderate sequential distance constraints (<code>5.5-6.5Å</code>)</li>\n<li>Low-confidence templates: additional steric clash prevention and light base-pairing constraints. </li>\n<li>Constraint strength scales as: <code>0.8 × (1 - min(confidence, 0.8))</code></li>\n</ul>\n<h2>DRfold2 Enhancements</h2>\n<h3>Selection Module</h3>\n<ul>\n<li><strong>Double Precision Calculations:</strong> Consistent float64 operations reduce numerical errors for more reliable model rankings</li>\n<li><strong>Vectorized Distance Calculations:</strong> GPU-accelerated pairwise distance computation via torch.cdist</li>\n<li><strong>Optimized Energy Functions:</strong> Pre-computed cubic spline coefficients enable fast structure scoring without repeated spline fitting</li>\n</ul>\n<p>These improvements were motivated by the authors' own observations that DRfold2 sometimes failed to select its best models. For example, they report cases where the 5th-ranked model significantly outperformed the top-ranked one (p. 8, lines 305–317), underscoring the need for more robust post-processing and ranking protocols (p. 9, lines 319–321). </p>\n<p>My modifications directly targeted this weakness by making scoring and ranking more accurate and consistent.</p>\n<h3>Optimization Module</h3>\n<ul>\n<li><strong>PyTorch LBFGS:</strong> Native optimizer with automatic differentiation delivers more accurate gradients and better convergence than SciPy implementations.</li>\n<li><strong>GPU Acceleration:</strong> Energy calculations and gradient computations performed on GPU where possible.</li>\n<li><strong>External Knowledge Integration:</strong> Enhanced capabilities through Boltz-1 integration (credit to <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a>) - (2nd notebook submission)</li>\n</ul>\n<p>The authors themselves highlight the flexibility of DRfold2's optimization framework, demonstrating this by integrating AlphaFold3 conformations as an additional potential term (p. 9, lines 327–329). This hybrid approach yielded significantly better results than either method alone, achieving higher TM-scores and lower RMSDs (p. 9, lines 331–334). They conclude that such integration represents a promising direction for future improvements (p. 10, lines 370–372). </p>\n<p>My own optimization experiments followed this spirit of extensibility, focusing on GPU acceleration and integration of Boltz-1.  </p>\n<h2>Hybrid Strategy</h2>\n<p>The final pipeline uses a strategic combination:</p>\n<ul>\n<li><strong>Template-based modeling:</strong> For shorter sequences and when time budget is exhausted. </li>\n<li><strong>DRfold2:</strong> For the rest of sequences where deep learning excels. </li>\n<li><strong>Graceful fallback:</strong> DRfold2 failures automatically fall back to template approach.</li>\n</ul>\n<p>Special shoutout to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for consistently sharing valuable research papers, open-source models, and insights that served as invaluable resources for the community.  </p>",
  "messages": [
    {
      "id": "3295798",
      "postDate": "09/29/2025 15:29:02",
      "content": "<p>Thank you Kaggle and the competition hosts, for this incredible competition and the opportunity to participate. Your passion for this challenge was truly infectious and served as one of my driving forces throughout the competition. </p>\n<p>Since this is my first gold and my first win on Kaggle, I would like to take the opportunity to thank <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> for the fast.ai course, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for their book which was an inspiration in my ML journey. Khan Academy for helping me rethink mathematics, and <a href=\"https://www.kaggle.com/huggingface\" target=\"_blank\">@huggingface</a> for their amazing deep learning courses. </p>\n<h2>Competition Strategy</h2>\n<p>My approach was clear from the outset. Without GPUs, training a model from scratch or fine-tuning was not viable. My early research - drawing on CASP results, literature, and conference talks, including one by host <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> - showed that Template-Based Modeling approaches consistently dominated. Based on this, I committed to TBM from day one and spent the next 90 days refining my method.</p>\n<p>Next, I focused on the evaluation metric, since understanding it determines the exploration path. TM-score has two key properties: it is normalized by structure length (so 50nt and 200nt RNAs are compared on the same 0-1 scale), and it is robust to local errors - a small number of misplaced nucleotides does not disproportionately lower the score. This insight allowed me to prioritize getting the overall fold correct over achieving atomic-level precision.</p>\n<h2>Data Strategy and Model Selection</h2>\n<p>The host-provided dataset was comprehensive. I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction. This process ensured complete coverage of the available structural data, capturing modified nucleotides that standard parsers might otherwise miss.</p>\n<p>After exploring nearly all available open-source models, I selected DRfold2 as the optimal choice due to its extensive potential for optimization. Rather than fine-tuning the model itself, I focused on enhancing its optimization and selection modules. This strategy improved prediction quality while ensuring the pipeline could execute efficiently on Kaggle GPUs.</p>\n<h2>Template-Based Modeling (TBM)</h2>\n<p>TBM follows a five-step process:</p>\n<p><strong>1. The Search - Finding Similar Structures</strong></p>\n<p>The goal is identifying database structures that resemble the target sequence through sequence alignment.</p>\n<p><strong>2. The Alignment - Sequence Mapping</strong></p>\n<p>This step creates the translation guide between query and template using global sequence alignment with gap penalties optimized for RNA. </p>\n<p><strong>3. The Transfer - Coordinate Inheritance</strong></p>\n<p>Straightforward copying of 3D coordinates for all matched positions. This leverages the evolutionary tendency for RNA to conserve 3D structure more than sequence. </p>\n<p><strong>4. The Gap Fill - Geometric Backbone Reconstruction</strong></p>\n<p>For insertions and deletions, I relied on geometric principles maintaining RNA's characteristic backbone:</p>\n<ul>\n<li>Maintains <code>C1'-C1'</code> distance (<code>~5.9Å</code> between consecutive nucleotides)</li>\n<li>For compressed gaps: extends the backbone with realistic curvature using sinusoidal perturbations perpendicular to the backbone direction</li>\n<li>For normal gaps: uses linear interpolation between flanking coordinates</li>\n<li>Terminal extensions follow the established backbone direction</li>\n</ul>\n<p><strong>5. Adaptive Refinement - Confidence-Based Optimization</strong></p>\n<p>The refinement intensity adapts to template confidence score:</p>\n<ul>\n<li>High-confidence templates (&gt;0.8 similarity): minimal constraints, preserving template geometry</li>\n<li>Medium-confidence templates: moderate sequential distance constraints (<code>5.5-6.5Å</code>)</li>\n<li>Low-confidence templates: additional steric clash prevention and light base-pairing constraints. </li>\n<li>Constraint strength scales as: <code>0.8 × (1 - min(confidence, 0.8))</code></li>\n</ul>\n<h2>DRfold2 Enhancements</h2>\n<h3>Selection Module</h3>\n<ul>\n<li><strong>Double Precision Calculations:</strong> Consistent float64 operations reduce numerical errors for more reliable model rankings</li>\n<li><strong>Vectorized Distance Calculations:</strong> GPU-accelerated pairwise distance computation via torch.cdist</li>\n<li><strong>Optimized Energy Functions:</strong> Pre-computed cubic spline coefficients enable fast structure scoring without repeated spline fitting</li>\n</ul>\n<p>These improvements were motivated by the authors' own observations that DRfold2 sometimes failed to select its best models. For example, they report cases where the 5th-ranked model significantly outperformed the top-ranked one (p. 8, lines 305–317), underscoring the need for more robust post-processing and ranking protocols (p. 9, lines 319–321). </p>\n<p>My modifications directly targeted this weakness by making scoring and ranking more accurate and consistent.</p>\n<h3>Optimization Module</h3>\n<ul>\n<li><strong>PyTorch LBFGS:</strong> Native optimizer with automatic differentiation delivers more accurate gradients and better convergence than SciPy implementations.</li>\n<li><strong>GPU Acceleration:</strong> Energy calculations and gradient computations performed on GPU where possible.</li>\n<li><strong>External Knowledge Integration:</strong> Enhanced capabilities through Boltz-1 integration (credit to <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a>) - (2nd notebook submission)</li>\n</ul>\n<p>The authors themselves highlight the flexibility of DRfold2's optimization framework, demonstrating this by integrating AlphaFold3 conformations as an additional potential term (p. 9, lines 327–329). This hybrid approach yielded significantly better results than either method alone, achieving higher TM-scores and lower RMSDs (p. 9, lines 331–334). They conclude that such integration represents a promising direction for future improvements (p. 10, lines 370–372). </p>\n<p>My own optimization experiments followed this spirit of extensibility, focusing on GPU acceleration and integration of Boltz-1.  </p>\n<h2>Hybrid Strategy</h2>\n<p>The final pipeline uses a strategic combination:</p>\n<ul>\n<li><strong>Template-based modeling:</strong> For shorter sequences and when time budget is exhausted. </li>\n<li><strong>DRfold2:</strong> For the rest of sequences where deep learning excels. </li>\n<li><strong>Graceful fallback:</strong> DRfold2 failures automatically fall back to template approach.</li>\n</ul>\n<p>Special shoutout to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for consistently sharing valuable research papers, open-source models, and insights that served as invaluable resources for the community.  </p>",
      "rawMarkdown": "Thank you Kaggle and the competition hosts, for this incredible competition and the opportunity to participate. Your passion for this challenge was truly infectious and served as one of my driving forces throughout the competition. \n\nSince this is my first gold and my first win on Kaggle, I would like to take the opportunity to thank @jhoward for the fast.ai course, @radek1 for their book which was an inspiration in my ML journey. Khan Academy for helping me rethink mathematics, and @huggingface for their amazing deep learning courses. \n\n\n## Competition Strategy\n\nMy approach was clear from the outset. Without GPUs, training a model from scratch or fine-tuning was not viable. My early research - drawing on CASP results, literature, and conference talks, including one by host @rhijudas - showed that Template-Based Modeling approaches consistently dominated. Based on this, I committed to TBM from day one and spent the next 90 days refining my method.\n\nNext, I focused on the evaluation metric, since understanding it determines the exploration path. TM-score has two key properties: it is normalized by structure length (so 50nt and 200nt RNAs are compared on the same 0-1 scale), and it is robust to local errors - a small number of misplaced nucleotides does not disproportionately lower the score. This insight allowed me to prioritize getting the overall fold correct over achieving atomic-level precision.\n\n\n## Data Strategy and Model Selection\n\nThe host-provided dataset was comprehensive. I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction. This process ensured complete coverage of the available structural data, capturing modified nucleotides that standard parsers might otherwise miss.\n\nAfter exploring nearly all available open-source models, I selected DRfold2 as the optimal choice due to its extensive potential for optimization. Rather than fine-tuning the model itself, I focused on enhancing its optimization and selection modules. This strategy improved prediction quality while ensuring the pipeline could execute efficiently on Kaggle GPUs.\n\n## Template-Based Modeling (TBM)\n\nTBM follows a five-step process:\n\n**1. The Search - Finding Similar Structures**\n\nThe goal is identifying database structures that resemble the target sequence through sequence alignment.\n\n\n**2. The Alignment - Sequence Mapping**\n\nThis step creates the translation guide between query and template using global sequence alignment with gap penalties optimized for RNA. \n\n\n**3. The Transfer - Coordinate Inheritance**\n\nStraightforward copying of 3D coordinates for all matched positions. This leverages the evolutionary tendency for RNA to conserve 3D structure more than sequence. \n\n\n**4. The Gap Fill - Geometric Backbone Reconstruction**\n\nFor insertions and deletions, I relied on geometric principles maintaining RNA's characteristic backbone:\n\n- Maintains `C1'-C1'` distance (`~5.9Å` between consecutive nucleotides)\n- For compressed gaps: extends the backbone with realistic curvature using sinusoidal perturbations perpendicular to the backbone direction\n- For normal gaps: uses linear interpolation between flanking coordinates\n- Terminal extensions follow the established backbone direction\n\n\n**5. Adaptive Refinement - Confidence-Based Optimization**\n\nThe refinement intensity adapts to template confidence score:\n\n- High-confidence templates (>0.8 similarity): minimal constraints, preserving template geometry\n- Medium-confidence templates: moderate sequential distance constraints (`5.5-6.5Å`)\n- Low-confidence templates: additional steric clash prevention and light base-pairing constraints. \n- Constraint strength scales as: `0.8 × (1 - min(confidence, 0.8))`\n\n\n\n## DRfold2 Enhancements\n\n### Selection Module\n\n- **Double Precision Calculations:** Consistent float64 operations reduce numerical errors for more reliable model rankings\n- **Vectorized Distance Calculations:** GPU-accelerated pairwise distance computation via torch.cdist\n- **Optimized Energy Functions:** Pre-computed cubic spline coefficients enable fast structure scoring without repeated spline fitting\n\nThese improvements were motivated by the authors' own observations that DRfold2 sometimes failed to select its best models. For example, they report cases where the 5th-ranked model significantly outperformed the top-ranked one (p. 8, lines 305–317), underscoring the need for more robust post-processing and ranking protocols (p. 9, lines 319–321). \n\nMy modifications directly targeted this weakness by making scoring and ranking more accurate and consistent.\n\n\n\n### Optimization Module\n\n- **PyTorch LBFGS:** Native optimizer with automatic differentiation delivers more accurate gradients and better convergence than SciPy implementations.\n- **GPU Acceleration:** Energy calculations and gradient computations performed on GPU where possible.\n- **External Knowledge Integration:** Enhanced capabilities through Boltz-1 integration (credit to @youhanlee) - (2nd notebook submission)\n\nThe authors themselves highlight the flexibility of DRfold2's optimization framework, demonstrating this by integrating AlphaFold3 conformations as an additional potential term (p. 9, lines 327–329). This hybrid approach yielded significantly better results than either method alone, achieving higher TM-scores and lower RMSDs (p. 9, lines 331–334). They conclude that such integration represents a promising direction for future improvements (p. 10, lines 370–372). \n\nMy own optimization experiments followed this spirit of extensibility, focusing on GPU acceleration and integration of Boltz-1.  \n\n\n\n## Hybrid Strategy\n\nThe final pipeline uses a strategic combination:\n\n- **Template-based modeling:** For shorter sequences and when time budget is exhausted. \n- **DRfold2:** For the rest of sequences where deep learning excels. \n- **Graceful fallback:** DRfold2 failures automatically fall back to template approach.\n\n\nSpecial shoutout to @hengck23 for consistently sharing valuable research papers, open-source models, and insights that served as invaluable resources for the community.",
      "votes": null
    },
    {
      "id": "3295897",
      "postDate": "09/29/2025 18:22:47",
      "content": "<p>Congratulations! Thanks for sharing the solution, very inspiring! </p>\n<p>What is the max sequence length for TBM in your solution?</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing the solution, very inspiring! \n\nWhat is the max sequence length for TBM in your solution?",
      "votes": null
    },
    {
      "id": "3295998",
      "postDate": "09/30/2025 00:56:06",
      "content": "<p>Hi, thanks!</p>\n<p>Instead of hardcoding the sequence length, i went with the first 14 sorted sequences for TBM. Looking back, I think a better approach would be to dynamically adapt based on the confidence score of the template search</p>",
      "rawMarkdown": "Hi, thanks!\n\nInstead of hardcoding the sequence length, i went with the first 14 sorted sequences for TBM. Looking back, I think a better approach would be to dynamically adapt based on the confidence score of the template search",
      "votes": null
    },
    {
      "id": "3296113",
      "postDate": "09/30/2025 08:14:48",
      "content": "<p>Interesting! Did you observe that as the sequence grows longer, the confidence score of the template search becomes lower, and then at some point, the deep learning method (DRfold2) starts to outperform TBM?</p>\n<blockquote>\n  <p>I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction.</p>\n</blockquote>\n<p>During data preprocessing, did you encounter many modified nucleotides? Also, could you elaborate on what you mean by \"disorder-aware coordinate extraction\"? </p>",
      "rawMarkdown": "Interesting! Did you observe that as the sequence grows longer, the confidence score of the template search becomes lower, and then at some point, the deep learning method (DRfold2) starts to outperform TBM?\n\n> I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction.\n\nDuring data preprocessing, did you encounter many modified nucleotides? Also, could you elaborate on what you mean by \"disorder-aware coordinate extraction\"?",
      "votes": null
    },
    {
      "id": "3296188",
      "postDate": "09/30/2025 10:58:00",
      "content": "<p>yeah, that was my observation and the reasoning behind sorting sequences, though I didn't analyze individual sequences, so I can't say for sure it holds across all sequences. </p>\n<p>For preprocessing, my baseline dataset (19,393 sequences) was parsed with a standard <code>A/U/G/C</code>-only parser. I later found 38 CIF files in the PDB_RNA dir that weren't captured this way. For those, I used a two-step fallback: try standard parsing first, and if that failed, switch to a comprehensive mapping (93 modified nucleotide variants) combined with disorder-aware extraction. All 38 required the fallback, and 36 were successfully processed. So in practice I only recovered 36 (and if i recall correctly, it actually improved the score by 0.02-0.03).</p>\n<p>On the disorder-handling side: the parser fails if atoms have multiple conformations (disorder issues). To handle this, BioPython represents such cases with <code>DisorderedAtom</code> wrapper objects. My method detects these (via <code>hasattr(atom, 'selected_child')</code>) and then extracts the selected atom, ensuring safe coordinate access and preventing crashes on structures with multiple recorded positions. By default, BioPython picks the conformer with the highest occupancy, or the first listed if occupancies are equal. Occupancy is a crystallography-specific concept.</p>",
      "rawMarkdown": "yeah, that was my observation and the reasoning behind sorting sequences, though I didn't analyze individual sequences, so I can't say for sure it holds across all sequences. \n\nFor preprocessing, my baseline dataset (19,393 sequences) was parsed with a standard `A/U/G/C`-only parser. I later found 38 CIF files in the PDB_RNA dir that weren't captured this way. For those, I used a two-step fallback: try standard parsing first, and if that failed, switch to a comprehensive mapping (93 modified nucleotide variants) combined with disorder-aware extraction. All 38 required the fallback, and 36 were successfully processed. So in practice I only recovered 36 (and if i recall correctly, it actually improved the score by 0.02-0.03).\n\nOn the disorder-handling side: the parser fails if atoms have multiple conformations (disorder issues). To handle this, BioPython represents such cases with `DisorderedAtom` wrapper objects. My method detects these (via `hasattr(atom, 'selected_child')`) and then extracts the selected atom, ensuring safe coordinate access and preventing crashes on structures with multiple recorded positions. By default, BioPython picks the conformer with the highest occupancy, or the first listed if occupancies are equal. Occupancy is a crystallography-specific concept.",
      "votes": null
    },
    {
      "id": "3296227",
      "postDate": "09/30/2025 13:05:30",
      "content": "<p>thank you for sharing the solution, hoping to learn from it!</p>",
      "rawMarkdown": "thank you for sharing the solution, hoping to learn from it!",
      "votes": null
    },
    {
      "id": "3296281",
      "postDate": "09/30/2025 15:02:56",
      "content": "<p><a href=\"https://www.kaggle.com/jaejohn\" target=\"_blank\">@jaejohn</a> It's amazing that the template discovery code is based on an 'off-the-shelf' aligner from Biopython. </p>\n<p>Can I ask how did you figure out the parameter settings for the aligner?  </p>\n<p>In particular, did you use an off-line validation set that you could share?</p>",
      "rawMarkdown": "jaejohn It's amazing that the template discovery code is based on an 'off-the-shelf' aligner from Biopython. \n\nCan I ask how did you figure out the parameter settings for the aligner?  \n\nIn particular, did you use an off-line validation set that you could share?",
      "votes": null
    },
    {
      "id": "3296294",
      "postDate": "09/30/2025 15:30:55",
      "content": "<p>Hi! I have joined this competition really early on, and am relatively new to RNA 3D Structure prediction. I have done alot of things using the data for my Science Fair, namely with GNNs and other related Graph Learning architectures. So, only recently I have started to look into TBM. I have seen countless examples of TBM, but never an actual notebook walking through the steps. By any chances, are any of these notebooks showing TBM? I don't really get the TBM procedure, but I am eager to learn. Do you have any resources, or even a notebook showing the steps to TBM so that I may replicate it? Thank you so much! </p>",
      "rawMarkdown": "Hi! I have joined this competition really early on, and am relatively new to RNA 3D Structure prediction. I have done alot of things using the data for my Science Fair, namely with GNNs and other related Graph Learning architectures. So, only recently I have started to look into TBM. I have seen countless examples of TBM, but never an actual notebook walking through the steps. By any chances, are any of these notebooks showing TBM? I don't really get the TBM procedure, but I am eager to learn. Do you have any resources, or even a notebook showing the steps to TBM so that I may replicate it? Thank you so much!",
      "votes": null
    },
    {
      "id": "3296295",
      "postDate": "09/30/2025 15:32:28",
      "content": "<p>Another question -- was DL even necessary? Do you have a notebook that just provides 5 models from the template search pipeline? Does it do about as well as the hybrid notebooks that also include models from DRFold2? </p>",
      "rawMarkdown": "Another question -- was DL even necessary? Do you have a notebook that just provides 5 models from the template search pipeline? Does it do about as well as the hybrid notebooks that also include models from DRFold2?",
      "votes": null
    },
    {
      "id": "3296310",
      "postDate": "09/30/2025 15:53:19",
      "content": "<p>hi, the parameter settings were determined through trial and error, using the leaderboard score as my only feedback mechanism. </p>\n<p>I did not use any offline validation set beyond the leaderboard feedback. </p>\n<p>also worth noting: as the dataset grew over time, the parameters needed retuning.</p>",
      "rawMarkdown": "hi, the parameter settings were determined through trial and error, using the leaderboard score as my only feedback mechanism. \n\nI did not use any offline validation set beyond the leaderboard feedback. \n\nalso worth noting: as the dataset grew over time, the parameters needed retuning.",
      "votes": null
    },
    {
      "id": "3296331",
      "postDate": "09/30/2025 16:42:08",
      "content": "<p>hi, i have made public my very first implementation of TBM-only approach. It is a basic implementation which was refined over 90-days. </p>\n<p>I re-submitted it for private LB scoring and it scores: 0.31098</p>\n<p>you may find the notebook here: <a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321</a></p>\n<p>The notebook is heavily commented to walk you through each step.</p>",
      "rawMarkdown": "hi, i have made public my very first implementation of TBM-only approach. It is a basic implementation which was refined over 90-days. \n\nI re-submitted it for private LB scoring and it scores: 0.31098\n\nyou may find the notebook here: https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321\n\nThe notebook is heavily commented to walk you through each step.",
      "votes": null
    },
    {
      "id": "3296340",
      "postDate": "09/30/2025 17:02:14",
      "content": "<p>wow, i just submitted the TBM-only notebook and it scored: 0.59298</p>\n<p>my winning score using hybrid approach is: 0.57773</p>\n<p>edit: i have made the notebook public:  <a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach</a></p>",
      "rawMarkdown": "wow, i just submitted the TBM-only notebook and it scored: 0.59298\n\nmy winning score using hybrid approach is: 0.57773\n\nedit: i have made the notebook public:  https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach",
      "votes": null
    },
    {
      "id": "3296399",
      "postDate": "09/30/2025 19:28:46",
      "content": "<p>Wow, indeed! </p>\n<p>What was the special sauce compared to the earlier template-only notebook (<a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321</a>) ?</p>",
      "rawMarkdown": "Wow, indeed! \n\nWhat was the special sauce compared to the earlier template-only notebook (https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321) ?",
      "votes": null
    },
    {
      "id": "3296474",
      "postDate": "10/01/2025 01:25:32",
      "content": "<p>I rescored the winning notebook with DRFold2 disabled, and it scored: 0.58487</p>\n<p>The other notebook which scored 0.59298 actually uses a different approach (slightly more complex - with clustering, diversity selection, and feature-based grouping)</p>\n<p>Going from earlier template-only notebook to the high scoring one - I think it was consistently iterating with trial and error, until I took the lead on the leaderboard, then diversified to DL approach, thinking I needed it to secure the lead. The key difference between the two notebooks would be: dataset size, enhanced template selection and composite similarity scoring.</p>",
      "rawMarkdown": "I rescored the winning notebook with DRFold2 disabled, and it scored: 0.58487\n\nThe other notebook which scored 0.59298 actually uses a different approach (slightly more complex - with clustering, diversity selection, and feature-based grouping)\n\nGoing from earlier template-only notebook to the high scoring one - I think it was consistently iterating with trial and error, until I took the lead on the leaderboard, then diversified to DL approach, thinking I needed it to secure the lead. The key difference between the two notebooks would be: dataset size, enhanced template selection and composite similarity scoring.",
      "votes": null
    },
    {
      "id": "3297235",
      "postDate": "10/02/2025 17:19:34",
      "content": "<pre><code> \n</code></pre>\n<p>It's not a question for me but I would like to share some information if you're interested in: I've processed the sequence on PDB up to March 31, 2025, around 21.7% RNA chains have modified residues. Among 23,630 RNA chains of 5,964 unique sequences, there are 5,129 RNA chains with modified residues (1,266 unique sequences).</p>",
      "rawMarkdown": "```\nDuring data preprocessing, did you encounter many modified nucleotides?\n```\n\nIt's not a question for me but I would like to share some information if you're interested in: I've processed the sequence on PDB up to March 31, 2025, around 21.7% RNA chains have modified residues. Among 23,630 RNA chains of 5,964 unique sequences, there are 5,129 RNA chains with modified residues (1,266 unique sequences).",
      "votes": null
    },
    {
      "id": "3297254",
      "postDate": "10/02/2025 18:00:23",
      "content": "<p>Congratulations and thanks for sharing your solution! It's very well-thought and I really enjoy your 2 key properties about TM-scores.</p>\n<p>I read your solution notebook and saw you choose the similarity sequence by a composite_score:</p>\n<pre><code>composite_score = (\n         * global_score + \n         * local_score + \n         * feature_similarity + \n         * kmer_similarity\n    )\n</code></pre>\n<p><code>feature_similarity</code> score really caught my attention. Here the features were extracted by <code>_extract_enhanced_rna_features</code> function. Are there anything specific why these dinucleotides are considered as features for calculating similarity?</p>\n<pre><code># . Dinucleotide frequencies (reduced  - most important  RNA)\n    important_dinucs = [, , , , , , , , , ]\n     dinuc in important_dinucs:\n         = \n         i in ((seq) - ):\n             seq[i:i+] == dinuc:\n                 += \n        freq =  / ((seq) - )  (seq) &gt;   \n        features.(freq)\n</code></pre>",
      "rawMarkdown": "Congratulations and thanks for sharing your solution! It's very well-thought and I really enjoy your 2 key properties about TM-scores.\n\nI read your solution notebook and saw you choose the similarity sequence by a composite_score:\n```python\ncomposite_score = (\n        0.4 * global_score + \n        0.3 * local_score + \n        0.2 * feature_similarity + \n        0.1 * kmer_similarity\n    )\n```\n`feature_similarity` score really caught my attention. Here the features were extracted by `_extract_enhanced_rna_features` function. Are there anything specific why these dinucleotides are considered as features for calculating similarity?\n```\n# 2. Dinucleotide frequencies (reduced set - most important for RNA)\n    important_dinucs = ['AU', 'UA', 'GC', 'CG', 'GU', 'UG', 'AA', 'UU', 'GG', 'CC']\n    for dinuc in important_dinucs:\n        count = 0\n        for i in range(len(seq) - 1):\n            if seq[i:i+2] == dinuc:\n                count += 1\n        freq = count / (len(seq) - 1) if len(seq) > 1 else 0\n        features.append(freq)\n```",
      "votes": null
    },
    {
      "id": "3297405",
      "postDate": "10/03/2025 04:23:20",
      "content": "<p>Hi, thanks! I tested it using all 16 dinucleotides and it scored: 0.5895 (decreased score vs 0.59298 with 10). It only contributes about 4% of the total similarity metric (20% to the <code>feature_similarity</code> score, which itself is only 20% of the <code>composite_score</code>). I think it is about adding a bit of sequence composition awareness, the left out 6 were probably adding noise as per my tests. </p>",
      "rawMarkdown": "Hi, thanks! I tested it using all 16 dinucleotides and it scored: 0.5895 (decreased score vs 0.59298 with 10). It only contributes about 4% of the total similarity metric (20% to the `feature_similarity` score, which itself is only 20% of the `composite_score`). I think it is about adding a bit of sequence composition awareness, the left out 6 were probably adding noise as per my tests.",
      "votes": null
    },
    {
      "id": "3297682",
      "postDate": "10/03/2025 14:52:17",
      "content": "<p>Thank you for sharing your experience!</p>",
      "rawMarkdown": "Thank you for sharing your experience!",
      "votes": null
    },
    {
      "id": "3298769",
      "postDate": "10/06/2025 10:14:08",
      "content": "<p>Thanks, John! I have another question: since one sequence can correspond to multiple structures, I was wondering — your similarity search is based only on sequences, but did you also try grouping them first and then applying a further selection step based on structural features?</p>",
      "rawMarkdown": "Thanks, John! I have another question: since one sequence can correspond to multiple structures, I was wondering — your similarity search is based only on sequences, but did you also try grouping them first and then applying a further selection step based on structural features?",
      "votes": null
    },
    {
      "id": "3303835",
      "postDate": "10/19/2025 04:18:16",
      "content": "<p>Where can I check out the winning solution’s code?</p>",
      "rawMarkdown": "Where can I check out the winning solution’s code?",
      "votes": null
    },
    {
      "id": "3304200",
      "postDate": "10/20/2025 04:46:01",
      "content": "<p>Congratulations and thank you for sharing your process 🌷🌷</p>",
      "rawMarkdown": "Congratulations and thank you for sharing your process 🌷🌷",
      "votes": null
    },
    {
      "id": "3308975",
      "postDate": "10/30/2025 16:27:29",
      "content": "<p>Congratulations! Thanks a lot for sharing your solution! I've tried your TBM-only notebook (<a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach</a>) on a recently deposit PDB 9hro, but got extended rod-like regions. Presumably these are the regions without templates. Do you see that often?<br>\nOne thing different perhaps is, I didn't find these csvs </p>\n<pre><code>train_seqs_v2 = pd.read_csv() \ntrain_labels_v2 = pd.read_csv()\n</code></pre>\n<p>(I guess you generated them yourselves)<br>\nso I used the commented part and the data from kaggle data for the template library</p>\n<pre><code>\n\n</code></pre>\n<p>and this contains only 5k instead of 18k in your notebook. Would you think this is the main reason for this, and things would be much better when used the whole template library? Or is there anything else that I am missing?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9426466%2Fa8a0a18e01dd72d359b1e086c8a45d12%2FRNA_MODEL-9HRO.png?generation=1761841530522501&amp;alt=media\" alt=\"\"> </p>",
      "rawMarkdown": "Congratulations! Thanks a lot for sharing your solution! I've tried your TBM-only notebook (https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach) on a recently deposit PDB 9hro, but got extended rod-like regions. Presumably these are the regions without templates. Do you see that often?\nOne thing different perhaps is, I didn't find these csvs \n```python\ntrain_seqs_v2 = pd.read_csv('/kaggle/input/rna-cif-to-csv/rna_sequences.csv') \ntrain_labels_v2 = pd.read_csv('/kaggle/input/rna-cif-to-csv/rna_coordinates.csv')\n``` (I guess you generated them yourselves)\nso I used the commented part and the data from kaggle data for the template library\n```python\n# train_seqs_v2 = pd.read_csv('/kaggle/input/extended-rna/train_sequences_v2.csv')\n# train_labels_v2 = pd.read_csv('/kaggle/input/extended-rna/train_labels_v2.csv')\n``` and this contains only 5k instead of 18k in your notebook. Would you think this is the main reason for this, and things would be much better when used the whole template library? Or is there anything else that I am missing?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9426466%2Fa8a0a18e01dd72d359b1e086c8a45d12%2FRNA_MODEL-9HRO.png?generation=1761841530522501&alt=media)",
      "votes": null
    },
    {
      "id": "3323192",
      "postDate": "11/14/2025 03:51:08",
      "content": "<p>Hi, sorry for the delayed response. The template library size is likely the main issue.</p>\n<p>The dataset is publicly available here: <a href=\"https://www.kaggle.com/datasets/jaejohn/rna-cif-to-csv\" target=\"_blank\">https://www.kaggle.com/datasets/jaejohn/rna-cif-to-csv</a></p>",
      "rawMarkdown": "Hi, sorry for the delayed response. The template library size is likely the main issue.\n\nThe dataset is publicly available here: https://www.kaggle.com/datasets/jaejohn/rna-cif-to-csv",
      "votes": null
    },
    {
      "id": "3323196",
      "postDate": "11/14/2025 03:56:24",
      "content": "<p>Hi! the winning notebook is located here: <a href=\"https://www.kaggle.com/code/jaejohn/sub-2-hybrid-single-model\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/sub-2-hybrid-single-model</a></p>\n<p>TBM-only notebook which scores better is located here:<a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach</a></p>",
      "rawMarkdown": "Hi! the winning notebook is located here: https://www.kaggle.com/code/jaejohn/sub-2-hybrid-single-model\n\nTBM-only notebook which scores better is located here:https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach",
      "votes": null
    },
    {
      "id": "3323263",
      "postDate": "11/14/2025 04:46:20",
      "content": "<p>Hi, sorry for the delayed response.</p>\n<p>I was focused on maximizing structural diversity, so treated them as separate templates. The rationale was that different conformations of the same sequence provide different structural information, which could be valuable for template-based modeling. The clustering step helps ensure the selected templates are diverse. (implemented this in the TBM-only notebook)</p>\n<p>I think for TBM, if a sequence is highly similar to the query, having multiple conformations of it can be more useful than forcing sequence diversity. But it's an interesting alternative perspective worth testing.</p>",
      "rawMarkdown": "Hi, sorry for the delayed response.\n\nI was focused on maximizing structural diversity, so treated them as separate templates. The rationale was that different conformations of the same sequence provide different structural information, which could be valuable for template-based modeling. The clustering step helps ensure the selected templates are diverse. (implemented this in the TBM-only notebook)\n\nI think for TBM, if a sequence is highly similar to the query, having multiple conformations of it can be more useful than forcing sequence diversity. But it's an interesting alternative perspective worth testing.",
      "votes": null
    },
    {
      "id": "3405864",
      "postDate": "02/14/2026 04:02:33",
      "content": "<p>Where did you get it?</p>",
      "rawMarkdown": "Where did you get it?",
      "votes": null
    },
    {
      "id": "3420257",
      "postDate": "03/12/2026 20:03:11",
      "content": "<p>Congrats on the win! \nI also went with a template-based approach but only handled the 4 standard bases your 93 nucleotide variant mapping with disorder-aware extraction was clearly a big differentiator. \nI'm currently participating in Part 2 of the competition and this would be incredibly useful. \nAny chance you'd be willing to share that CIF extraction notebook? \nWould love to learn from it. \nThanks! 🎉</p>",
      "rawMarkdown": "Congrats on the win! \nI also went with a template-based approach but only handled the 4 standard bases your 93 nucleotide variant mapping with disorder-aware extraction was clearly a big differentiator. \nI'm currently participating in Part 2 of the competition and this would be incredibly useful. \nAny chance you'd be willing to share that CIF extraction notebook? \nWould love to learn from it. \nThanks! 🎉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3295897,
      "author_name": "zoushuxian",
      "author_url": "",
      "post_date": "09/29/2025 18:22:47",
      "content": "<p>Congratulations! Thanks for sharing the solution, very inspiring! </p>\n<p>What is the max sequence length for TBM in your solution?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3295998,
          "author_name": "jaejohn",
          "author_url": "",
          "post_date": "09/30/2025 00:56:06",
          "content": "<p>Hi, thanks!</p>\n<p>Instead of hardcoding the sequence length, i went with the first 14 sorted sequences for TBM. Looking back, I think a better approach would be to dynamically adapt based on the confidence score of the template search</p>",
          "votes": null,
          "replies": [
            {
              "id": 3296113,
              "author_name": "zoushuxian",
              "author_url": "",
              "post_date": "09/30/2025 08:14:48",
              "content": "<p>Interesting! Did you observe that as the sequence grows longer, the confidence score of the template search becomes lower, and then at some point, the deep learning method (DRfold2) starts to outperform TBM?</p>\n<blockquote>\n  <p>I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction.</p>\n</blockquote>\n<p>During data preprocessing, did you encounter many modified nucleotides? Also, could you elaborate on what you mean by \"disorder-aware coordinate extraction\"? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3296188,
                  "author_name": "jaejohn",
                  "author_url": "",
                  "post_date": "09/30/2025 10:58:00",
                  "content": "<p>yeah, that was my observation and the reasoning behind sorting sequences, though I didn't analyze individual sequences, so I can't say for sure it holds across all sequences. </p>\n<p>For preprocessing, my baseline dataset (19,393 sequences) was parsed with a standard <code>A/U/G/C</code>-only parser. I later found 38 CIF files in the PDB_RNA dir that weren't captured this way. For those, I used a two-step fallback: try standard parsing first, and if that failed, switch to a comprehensive mapping (93 modified nucleotide variants) combined with disorder-aware extraction. All 38 required the fallback, and 36 were successfully processed. So in practice I only recovered 36 (and if i recall correctly, it actually improved the score by 0.02-0.03).</p>\n<p>On the disorder-handling side: the parser fails if atoms have multiple conformations (disorder issues). To handle this, BioPython represents such cases with <code>DisorderedAtom</code> wrapper objects. My method detects these (via <code>hasattr(atom, 'selected_child')</code>) and then extracts the selected atom, ensuring safe coordinate access and preventing crashes on structures with multiple recorded positions. By default, BioPython picks the conformer with the highest occupancy, or the first listed if occupancies are equal. Occupancy is a crystallography-specific concept.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3297235,
                  "author_name": "nguyenhoa",
                  "author_url": "",
                  "post_date": "10/02/2025 17:19:34",
                  "content": "<pre><code> \n</code></pre>\n<p>It's not a question for me but I would like to share some information if you're interested in: I've processed the sequence on PDB up to March 31, 2025, around 21.7% RNA chains have modified residues. Among 23,630 RNA chains of 5,964 unique sequences, there are 5,129 RNA chains with modified residues (1,266 unique sequences).</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3296227,
      "author_name": "irakozekelly",
      "author_url": "",
      "post_date": "09/30/2025 13:05:30",
      "content": "<p>thank you for sharing the solution, hoping to learn from it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3296281,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "09/30/2025 15:02:56",
      "content": "<p><a href=\"https://www.kaggle.com/jaejohn\" target=\"_blank\">@jaejohn</a> It's amazing that the template discovery code is based on an 'off-the-shelf' aligner from Biopython. </p>\n<p>Can I ask how did you figure out the parameter settings for the aligner?  </p>\n<p>In particular, did you use an off-line validation set that you could share?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3296310,
          "author_name": "jaejohn",
          "author_url": "",
          "post_date": "09/30/2025 15:53:19",
          "content": "<p>hi, the parameter settings were determined through trial and error, using the leaderboard score as my only feedback mechanism. </p>\n<p>I did not use any offline validation set beyond the leaderboard feedback. </p>\n<p>also worth noting: as the dataset grew over time, the parameters needed retuning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3296294,
      "author_name": "dhruvagadipathi",
      "author_url": "",
      "post_date": "09/30/2025 15:30:55",
      "content": "<p>Hi! I have joined this competition really early on, and am relatively new to RNA 3D Structure prediction. I have done alot of things using the data for my Science Fair, namely with GNNs and other related Graph Learning architectures. So, only recently I have started to look into TBM. I have seen countless examples of TBM, but never an actual notebook walking through the steps. By any chances, are any of these notebooks showing TBM? I don't really get the TBM procedure, but I am eager to learn. Do you have any resources, or even a notebook showing the steps to TBM so that I may replicate it? Thank you so much! </p>",
      "votes": null,
      "replies": [
        {
          "id": 3296331,
          "author_name": "jaejohn",
          "author_url": "",
          "post_date": "09/30/2025 16:42:08",
          "content": "<p>hi, i have made public my very first implementation of TBM-only approach. It is a basic implementation which was refined over 90-days. </p>\n<p>I re-submitted it for private LB scoring and it scores: 0.31098</p>\n<p>you may find the notebook here: <a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321</a></p>\n<p>The notebook is heavily commented to walk you through each step.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3296295,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "09/30/2025 15:32:28",
      "content": "<p>Another question -- was DL even necessary? Do you have a notebook that just provides 5 models from the template search pipeline? Does it do about as well as the hybrid notebooks that also include models from DRFold2? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3296340,
          "author_name": "jaejohn",
          "author_url": "",
          "post_date": "09/30/2025 17:02:14",
          "content": "<p>wow, i just submitted the TBM-only notebook and it scored: 0.59298</p>\n<p>my winning score using hybrid approach is: 0.57773</p>\n<p>edit: i have made the notebook public:  <a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3296399,
              "author_name": "rhijudas",
              "author_url": "",
              "post_date": "09/30/2025 19:28:46",
              "content": "<p>Wow, indeed! </p>\n<p>What was the special sauce compared to the earlier template-only notebook (<a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321</a>) ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3296474,
                  "author_name": "jaejohn",
                  "author_url": "",
                  "post_date": "10/01/2025 01:25:32",
                  "content": "<p>I rescored the winning notebook with DRFold2 disabled, and it scored: 0.58487</p>\n<p>The other notebook which scored 0.59298 actually uses a different approach (slightly more complex - with clustering, diversity selection, and feature-based grouping)</p>\n<p>Going from earlier template-only notebook to the high scoring one - I think it was consistently iterating with trial and error, until I took the lead on the leaderboard, then diversified to DL approach, thinking I needed it to secure the lead. The key difference between the two notebooks would be: dataset size, enhanced template selection and composite similarity scoring.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3297254,
      "author_name": "nguyenhoa",
      "author_url": "",
      "post_date": "10/02/2025 18:00:23",
      "content": "<p>Congratulations and thanks for sharing your solution! It's very well-thought and I really enjoy your 2 key properties about TM-scores.</p>\n<p>I read your solution notebook and saw you choose the similarity sequence by a composite_score:</p>\n<pre><code>composite_score = (\n         * global_score + \n         * local_score + \n         * feature_similarity + \n         * kmer_similarity\n    )\n</code></pre>\n<p><code>feature_similarity</code> score really caught my attention. Here the features were extracted by <code>_extract_enhanced_rna_features</code> function. Are there anything specific why these dinucleotides are considered as features for calculating similarity?</p>\n<pre><code># . Dinucleotide frequencies (reduced  - most important  RNA)\n    important_dinucs = [, , , , , , , , , ]\n     dinuc in important_dinucs:\n         = \n         i in ((seq) - ):\n             seq[i:i+] == dinuc:\n                 += \n        freq =  / ((seq) - )  (seq) &gt;   \n        features.(freq)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3297405,
          "author_name": "jaejohn",
          "author_url": "",
          "post_date": "10/03/2025 04:23:20",
          "content": "<p>Hi, thanks! I tested it using all 16 dinucleotides and it scored: 0.5895 (decreased score vs 0.59298 with 10). It only contributes about 4% of the total similarity metric (20% to the <code>feature_similarity</code> score, which itself is only 20% of the <code>composite_score</code>). I think it is about adding a bit of sequence composition awareness, the left out 6 were probably adding noise as per my tests. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3298769,
              "author_name": "nguyenhoa",
              "author_url": "",
              "post_date": "10/06/2025 10:14:08",
              "content": "<p>Thanks, John! I have another question: since one sequence can correspond to multiple structures, I was wondering — your similarity search is based only on sequences, but did you also try grouping them first and then applying a further selection step based on structural features?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3323263,
                  "author_name": "jaejohn",
                  "author_url": "",
                  "post_date": "11/14/2025 04:46:20",
                  "content": "<p>Hi, sorry for the delayed response.</p>\n<p>I was focused on maximizing structural diversity, so treated them as separate templates. The rationale was that different conformations of the same sequence provide different structural information, which could be valuable for template-based modeling. The clustering step helps ensure the selected templates are diverse. (implemented this in the TBM-only notebook)</p>\n<p>I think for TBM, if a sequence is highly similar to the query, having multiple conformations of it can be more useful than forcing sequence diversity. But it's an interesting alternative perspective worth testing.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3297682,
      "author_name": "irakozekelly",
      "author_url": "",
      "post_date": "10/03/2025 14:52:17",
      "content": "<p>Thank you for sharing your experience!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3303835,
      "author_name": "ds123456",
      "author_url": "",
      "post_date": "10/19/2025 04:18:16",
      "content": "<p>Where can I check out the winning solution’s code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3323196,
          "author_name": "jaejohn",
          "author_url": "",
          "post_date": "11/14/2025 03:56:24",
          "content": "<p>Hi! the winning notebook is located here: <a href=\"https://www.kaggle.com/code/jaejohn/sub-2-hybrid-single-model\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/sub-2-hybrid-single-model</a></p>\n<p>TBM-only notebook which scores better is located here:<a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3304200,
      "author_name": "divyapancholi",
      "author_url": "",
      "post_date": "10/20/2025 04:46:01",
      "content": "<p>Congratulations and thank you for sharing your process 🌷🌷</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3308975,
      "author_name": "yihengwukp",
      "author_url": "",
      "post_date": "10/30/2025 16:27:29",
      "content": "<p>Congratulations! Thanks a lot for sharing your solution! I've tried your TBM-only notebook (<a href=\"https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach</a>) on a recently deposit PDB 9hro, but got extended rod-like regions. Presumably these are the regions without templates. Do you see that often?<br>\nOne thing different perhaps is, I didn't find these csvs </p>\n<pre><code>train_seqs_v2 = pd.read_csv() \ntrain_labels_v2 = pd.read_csv()\n</code></pre>\n<p>(I guess you generated them yourselves)<br>\nso I used the commented part and the data from kaggle data for the template library</p>\n<pre><code>\n\n</code></pre>\n<p>and this contains only 5k instead of 18k in your notebook. Would you think this is the main reason for this, and things would be much better when used the whole template library? Or is there anything else that I am missing?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9426466%2Fa8a0a18e01dd72d359b1e086c8a45d12%2FRNA_MODEL-9HRO.png?generation=1761841530522501&amp;alt=media\" alt=\"\"> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3323192,
          "author_name": "jaejohn",
          "author_url": "",
          "post_date": "11/14/2025 03:51:08",
          "content": "<p>Hi, sorry for the delayed response. The template library size is likely the main issue.</p>\n<p>The dataset is publicly available here: <a href=\"https://www.kaggle.com/datasets/jaejohn/rna-cif-to-csv\" target=\"_blank\">https://www.kaggle.com/datasets/jaejohn/rna-cif-to-csv</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3405864,
              "author_name": "agrin88",
              "author_url": "",
              "post_date": "02/14/2026 04:02:33",
              "content": "<p>Where did you get it?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3420257,
      "author_name": "vyankteshdwivedi",
      "author_url": "",
      "post_date": "03/12/2026 20:03:11",
      "content": "<p>Congrats on the win! \nI also went with a template-based approach but only handled the 4 standard bases your 93 nucleotide variant mapping with disorder-aware extraction was clearly a big differentiator. \nI'm currently participating in Part 2 of the competition and this would be incredibly useful. \nAny chance you'd be willing to share that CIF extraction notebook? \nWould love to learn from it. \nThanks! 🎉</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3295798": "Thank you Kaggle and the competition hosts, for this incredible competition and the opportunity to participate. Your passion for this challenge was truly infectious and served as one of my driving forces throughout the competition. \n\nSince this is my first gold and my first win on Kaggle, I would like to take the opportunity to thank @jhoward for the fast.ai course, @radek1 for their book which was an inspiration in my ML journey. Khan Academy for helping me rethink mathematics, and @huggingface for their amazing deep learning courses. \n\n\n## Competition Strategy\n\nMy approach was clear from the outset. Without GPUs, training a model from scratch or fine-tuning was not viable. My early research - drawing on CASP results, literature, and conference talks, including one by host @rhijudas - showed that Template-Based Modeling approaches consistently dominated. Based on this, I committed to TBM from day one and spent the next 90 days refining my method.\n\nNext, I focused on the evaluation metric, since understanding it determines the exploration path. TM-score has two key properties: it is normalized by structure length (so 50nt and 200nt RNAs are compared on the same 0-1 scale), and it is robust to local errors - a small number of misplaced nucleotides does not disproportionately lower the score. This insight allowed me to prioritize getting the overall fold correct over achieving atomic-level precision.\n\n\n## Data Strategy and Model Selection\n\nThe host-provided dataset was comprehensive. I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction. This process ensured complete coverage of the available structural data, capturing modified nucleotides that standard parsers might otherwise miss.\n\nAfter exploring nearly all available open-source models, I selected DRfold2 as the optimal choice due to its extensive potential for optimization. Rather than fine-tuning the model itself, I focused on enhancing its optimization and selection modules. This strategy improved prediction quality while ensuring the pipeline could execute efficiently on Kaggle GPUs.\n\n## Template-Based Modeling (TBM)\n\nTBM follows a five-step process:\n\n**1. The Search - Finding Similar Structures**\n\nThe goal is identifying database structures that resemble the target sequence through sequence alignment.\n\n\n**2. The Alignment - Sequence Mapping**\n\nThis step creates the translation guide between query and template using global sequence alignment with gap penalties optimized for RNA. \n\n\n**3. The Transfer - Coordinate Inheritance**\n\nStraightforward copying of 3D coordinates for all matched positions. This leverages the evolutionary tendency for RNA to conserve 3D structure more than sequence. \n\n\n**4. The Gap Fill - Geometric Backbone Reconstruction**\n\nFor insertions and deletions, I relied on geometric principles maintaining RNA's characteristic backbone:\n\n- Maintains `C1'-C1'` distance (`~5.9Å` between consecutive nucleotides)\n- For compressed gaps: extends the backbone with realistic curvature using sinusoidal perturbations perpendicular to the backbone direction\n- For normal gaps: uses linear interpolation between flanking coordinates\n- Terminal extensions follow the established backbone direction\n\n\n**5. Adaptive Refinement - Confidence-Based Optimization**\n\nThe refinement intensity adapts to template confidence score:\n\n- High-confidence templates (>0.8 similarity): minimal constraints, preserving template geometry\n- Medium-confidence templates: moderate sequential distance constraints (`5.5-6.5Å`)\n- Low-confidence templates: additional steric clash prevention and light base-pairing constraints. \n- Constraint strength scales as: `0.8 × (1 - min(confidence, 0.8))`\n\n\n\n## DRfold2 Enhancements\n\n### Selection Module\n\n- **Double Precision Calculations:** Consistent float64 operations reduce numerical errors for more reliable model rankings\n- **Vectorized Distance Calculations:** GPU-accelerated pairwise distance computation via torch.cdist\n- **Optimized Energy Functions:** Pre-computed cubic spline coefficients enable fast structure scoring without repeated spline fitting\n\nThese improvements were motivated by the authors' own observations that DRfold2 sometimes failed to select its best models. For example, they report cases where the 5th-ranked model significantly outperformed the top-ranked one (p. 8, lines 305–317), underscoring the need for more robust post-processing and ranking protocols (p. 9, lines 319–321). \n\nMy modifications directly targeted this weakness by making scoring and ranking more accurate and consistent.\n\n\n\n### Optimization Module\n\n- **PyTorch LBFGS:** Native optimizer with automatic differentiation delivers more accurate gradients and better convergence than SciPy implementations.\n- **GPU Acceleration:** Energy calculations and gradient computations performed on GPU where possible.\n- **External Knowledge Integration:** Enhanced capabilities through Boltz-1 integration (credit to @youhanlee) - (2nd notebook submission)\n\nThe authors themselves highlight the flexibility of DRfold2's optimization framework, demonstrating this by integrating AlphaFold3 conformations as an additional potential term (p. 9, lines 327–329). This hybrid approach yielded significantly better results than either method alone, achieving higher TM-scores and lower RMSDs (p. 9, lines 331–334). They conclude that such integration represents a promising direction for future improvements (p. 10, lines 370–372). \n\nMy own optimization experiments followed this spirit of extensibility, focusing on GPU acceleration and integration of Boltz-1.  \n\n\n\n## Hybrid Strategy\n\nThe final pipeline uses a strategic combination:\n\n- **Template-based modeling:** For shorter sequences and when time budget is exhausted. \n- **DRfold2:** For the rest of sequences where deep learning excels. \n- **Graceful fallback:** DRfold2 failures automatically fall back to template approach.\n\n\nSpecial shoutout to @hengck23 for consistently sharing valuable research papers, open-source models, and insights that served as invaluable resources for the community.",
    "3295897": "Congratulations! Thanks for sharing the solution, very inspiring! \n\nWhat is the max sequence length for TBM in your solution?",
    "3295998": "Hi, thanks!\n\nInstead of hardcoding the sequence length, i went with the first 14 sorted sequences for TBM. Looking back, I think a better approach would be to dynamically adapt based on the confidence score of the template search",
    "3296113": "Interesting! Did you observe that as the sequence grows longer, the confidence score of the template search becomes lower, and then at some point, the deep learning method (DRfold2) starts to outperform TBM?\n\n> I systematically processed all CIF files in the provided PDB_RNA directory with comprehensive nucleotide mapping (93 variants including modified bases) and disorder-aware coordinate extraction.\n\nDuring data preprocessing, did you encounter many modified nucleotides? Also, could you elaborate on what you mean by \"disorder-aware coordinate extraction\"?",
    "3296188": "yeah, that was my observation and the reasoning behind sorting sequences, though I didn't analyze individual sequences, so I can't say for sure it holds across all sequences. \n\nFor preprocessing, my baseline dataset (19,393 sequences) was parsed with a standard `A/U/G/C`-only parser. I later found 38 CIF files in the PDB_RNA dir that weren't captured this way. For those, I used a two-step fallback: try standard parsing first, and if that failed, switch to a comprehensive mapping (93 modified nucleotide variants) combined with disorder-aware extraction. All 38 required the fallback, and 36 were successfully processed. So in practice I only recovered 36 (and if i recall correctly, it actually improved the score by 0.02-0.03).\n\nOn the disorder-handling side: the parser fails if atoms have multiple conformations (disorder issues). To handle this, BioPython represents such cases with `DisorderedAtom` wrapper objects. My method detects these (via `hasattr(atom, 'selected_child')`) and then extracts the selected atom, ensuring safe coordinate access and preventing crashes on structures with multiple recorded positions. By default, BioPython picks the conformer with the highest occupancy, or the first listed if occupancies are equal. Occupancy is a crystallography-specific concept.",
    "3296227": "thank you for sharing the solution, hoping to learn from it!",
    "3296281": "jaejohn It's amazing that the template discovery code is based on an 'off-the-shelf' aligner from Biopython. \n\nCan I ask how did you figure out the parameter settings for the aligner?  \n\nIn particular, did you use an off-line validation set that you could share?",
    "3296294": "Hi! I have joined this competition really early on, and am relatively new to RNA 3D Structure prediction. I have done alot of things using the data for my Science Fair, namely with GNNs and other related Graph Learning architectures. So, only recently I have started to look into TBM. I have seen countless examples of TBM, but never an actual notebook walking through the steps. By any chances, are any of these notebooks showing TBM? I don't really get the TBM procedure, but I am eager to learn. Do you have any resources, or even a notebook showing the steps to TBM so that I may replicate it? Thank you so much!",
    "3296295": "Another question -- was DL even necessary? Do you have a notebook that just provides 5 models from the template search pipeline? Does it do about as well as the hybrid notebooks that also include models from DRFold2?",
    "3296310": "hi, the parameter settings were determined through trial and error, using the leaderboard score as my only feedback mechanism. \n\nI did not use any offline validation set beyond the leaderboard feedback. \n\nalso worth noting: as the dataset grew over time, the parameters needed retuning.",
    "3296331": "hi, i have made public my very first implementation of TBM-only approach. It is a basic implementation which was refined over 90-days. \n\nI re-submitted it for private LB scoring and it scores: 0.31098\n\nyou may find the notebook here: https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321\n\nThe notebook is heavily commented to walk you through each step.",
    "3296340": "wow, i just submitted the TBM-only notebook and it scored: 0.59298\n\nmy winning score using hybrid approach is: 0.57773\n\nedit: i have made the notebook public:  https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach",
    "3296399": "Wow, indeed! \n\nWhat was the special sauce compared to the earlier template-only notebook (https://www.kaggle.com/code/jaejohn/rna-3d-folding-template-based-0-321) ?",
    "3296474": "I rescored the winning notebook with DRFold2 disabled, and it scored: 0.58487\n\nThe other notebook which scored 0.59298 actually uses a different approach (slightly more complex - with clustering, diversity selection, and feature-based grouping)\n\nGoing from earlier template-only notebook to the high scoring one - I think it was consistently iterating with trial and error, until I took the lead on the leaderboard, then diversified to DL approach, thinking I needed it to secure the lead. The key difference between the two notebooks would be: dataset size, enhanced template selection and composite similarity scoring.",
    "3297235": "```\nDuring data preprocessing, did you encounter many modified nucleotides?\n```\n\nIt's not a question for me but I would like to share some information if you're interested in: I've processed the sequence on PDB up to March 31, 2025, around 21.7% RNA chains have modified residues. Among 23,630 RNA chains of 5,964 unique sequences, there are 5,129 RNA chains with modified residues (1,266 unique sequences).",
    "3297254": "Congratulations and thanks for sharing your solution! It's very well-thought and I really enjoy your 2 key properties about TM-scores.\n\nI read your solution notebook and saw you choose the similarity sequence by a composite_score:\n```python\ncomposite_score = (\n        0.4 * global_score + \n        0.3 * local_score + \n        0.2 * feature_similarity + \n        0.1 * kmer_similarity\n    )\n```\n`feature_similarity` score really caught my attention. Here the features were extracted by `_extract_enhanced_rna_features` function. Are there anything specific why these dinucleotides are considered as features for calculating similarity?\n```\n# 2. Dinucleotide frequencies (reduced set - most important for RNA)\n    important_dinucs = ['AU', 'UA', 'GC', 'CG', 'GU', 'UG', 'AA', 'UU', 'GG', 'CC']\n    for dinuc in important_dinucs:\n        count = 0\n        for i in range(len(seq) - 1):\n            if seq[i:i+2] == dinuc:\n                count += 1\n        freq = count / (len(seq) - 1) if len(seq) > 1 else 0\n        features.append(freq)\n```",
    "3297405": "Hi, thanks! I tested it using all 16 dinucleotides and it scored: 0.5895 (decreased score vs 0.59298 with 10). It only contributes about 4% of the total similarity metric (20% to the `feature_similarity` score, which itself is only 20% of the `composite_score`). I think it is about adding a bit of sequence composition awareness, the left out 6 were probably adding noise as per my tests.",
    "3297682": "Thank you for sharing your experience!",
    "3298769": "Thanks, John! I have another question: since one sequence can correspond to multiple structures, I was wondering — your similarity search is based only on sequences, but did you also try grouping them first and then applying a further selection step based on structural features?",
    "3303835": "Where can I check out the winning solution’s code?",
    "3304200": "Congratulations and thank you for sharing your process 🌷🌷",
    "3308975": "Congratulations! Thanks a lot for sharing your solution! I've tried your TBM-only notebook (https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach) on a recently deposit PDB 9hro, but got extended rod-like regions. Presumably these are the regions without templates. Do you see that often?\nOne thing different perhaps is, I didn't find these csvs \n```python\ntrain_seqs_v2 = pd.read_csv('/kaggle/input/rna-cif-to-csv/rna_sequences.csv') \ntrain_labels_v2 = pd.read_csv('/kaggle/input/rna-cif-to-csv/rna_coordinates.csv')\n``` (I guess you generated them yourselves)\nso I used the commented part and the data from kaggle data for the template library\n```python\n# train_seqs_v2 = pd.read_csv('/kaggle/input/extended-rna/train_sequences_v2.csv')\n# train_labels_v2 = pd.read_csv('/kaggle/input/extended-rna/train_labels_v2.csv')\n``` and this contains only 5k instead of 18k in your notebook. Would you think this is the main reason for this, and things would be much better when used the whole template library? Or is there anything else that I am missing?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9426466%2Fa8a0a18e01dd72d359b1e086c8a45d12%2FRNA_MODEL-9HRO.png?generation=1761841530522501&alt=media)",
    "3323192": "Hi, sorry for the delayed response. The template library size is likely the main issue.\n\nThe dataset is publicly available here: https://www.kaggle.com/datasets/jaejohn/rna-cif-to-csv",
    "3323196": "Hi! the winning notebook is located here: https://www.kaggle.com/code/jaejohn/sub-2-hybrid-single-model\n\nTBM-only notebook which scores better is located here:https://www.kaggle.com/code/jaejohn/rna-3d-folds-tbm-only-approach",
    "3323263": "Hi, sorry for the delayed response.\n\nI was focused on maximizing structural diversity, so treated them as separate templates. The rationale was that different conformations of the same sequence provide different structural information, which could be valuable for template-based modeling. The clustering step helps ensure the selected templates are diverse. (implemented this in the TBM-only notebook)\n\nI think for TBM, if a sequence is highly similar to the query, having multiple conformations of it can be more useful than forcing sequence diversity. But it's an interesting alternative perspective worth testing.",
    "3405864": "Where did you get it?",
    "3420257": "Congrats on the win! \nI also went with a template-based approach but only handled the 4 standard bases your 93 nucleotide variant mapping with disorder-aware extraction was clearly a big differentiator. \nI'm currently participating in Part 2 of the competition and this would be incredibly useful. \nAny chance you'd be willing to share that CIF extraction notebook? \nWould love to learn from it. \nThanks! 🎉"
  },
  "source": "meta"
}