{
  "id": 609921,
  "title": "9th place solution  - d4t4 team ",
  "url": "/competitions/stanford-rna-3d-folding/writeups/10th-place-solution-d4t4-team",
  "author_name": "",
  "post_date": "2025-09-30T13:41:08.703Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>Many thanks to Kaggle and the competition hosts for organizing this great challenge and giving us the chance to take part.</strong></p>\n<h3>📂 Dataset Creation</h3>\n<p>We built our training dataset by merging several complementary sources:</p>\n<ul>\n<li><p><strong>Stanford RNA 3D Folding Competition Data</strong></p>\n<ul>\n<li>Versions: v1 and v2 (v2 used as the primary training set).</li></ul></li>\n<li><p><strong>CASP16 Top Predictions (Pseudo-labels)</strong></p>\n<ul>\n<li>We included Top-1 predicted structures per target as pseudo-labels.</li></ul></li>\n<li><p><strong>RNA PDB Database (cutoff: March 31, 2025)</strong></p>\n<ul>\n<li>Only entries released on or before March 31, 2025 were included.</li></ul></li>\n</ul>\n<hr>\n<h3>🧩 Handling Missing Residues</h3>\n<ul>\n<li>Residues with unresolved atoms in PDB were kept in the sequence but masked in coordinates.</li>\n<li>Missing atoms were zero-padded or ignored in loss computation.</li>\n<li>This ensures reliable training without losing sequence context.</li>\n</ul>\n<hr>\n<h3>🔁 Multiple Conformations</h3>\n<ul>\n<li>For some datasets (PDB and CASP16), we generated up to 5 conformations per RNA sequence.</li>\n<li>These conformations were stored as (num_conf, seq_len, 3) arrays in the dataset.</li>\n<li>During training, we supported two strategies:<ol>\n<li>Top-1 only → using the first conformation as a deterministic baseline.</li>\n<li>Multi-conf sampling → exposing the model to all 5 conformations, either:</li></ol><ul>\n<li>Returned as coordinate_multi (all conformations available at once), or</li>\n<li>Randomly selecting one conformation per epoch, so the model sees different structures across training.</li></ul></li>\n<li>This approach increased structural diversity and improved generalization without changing the overall dataset size.</li>\n</ul>\n<hr>\n<h3>🔬 MSA Pipeline</h3>\n<ul>\n<li>For sequences where evolutionary information was available, we incorporated Multiple Sequence Alignments (MSA) into training.</li>\n<li>We used MMseqs2 to build RNA MSAs against a curated RNA sequence database.</li>\n<li>Precomputed MSAs were stored in a dedicated directory and linked during training.</li>\n<li>For sequences without MSA coverage, the pipeline fell back to dummy/no-MSA features, ensuring consistent input formats.</li>\n<li>This hybrid strategy enriched structural context for many targets and improved accuracy on conserved RNAs.</li>\n</ul>\n<hr>\n<h3>📊 Final Outputs</h3>\n<ul>\n<li>Consolidated merged sequence dataset across competition, CASP16, and PDB.</li>\n<li>Aligned atom-level labels with missing residues masked and up to 5 conformations per sequence.</li>\n<li>Precomputed MSAs (via MMseqs2) for a subset of structures.</li>\n</ul>\n<p>===========================================================</p>\n<h2>⚙️ Training Configuration</h2>\n<ul>\n<li><p><strong>Hardware</strong></p>\n<ul>\n<li>GPU: NVIDIA H100 (96 GB)</li>\n<li>Mixed precision: bfloat16 (bf16)</li></ul></li>\n<li><p><strong>Optimization</strong></p>\n<ul>\n<li>Batch size: 8</li>\n<li>Max steps: 12,000</li>\n<li>Warmup steps: 50</li>\n<li>Learning rate: 1e-4</li>\n<li>Crop size: 800 nucleotides</li></ul></li>\n<li><p><strong>Sampling</strong></p>\n<ul>\n<li>Diffusion steps: 20</li></ul></li>\n<li><p><strong>Evaluation</strong></p>\n<ul>\n<li>Checkpoint interval: every 2,000 steps</li>\n<li>Evaluation interval: every 50,000 steps</li></ul></li>\n<li><p><strong>Features</strong></p>\n<ul>\n<li>MSA pipeline: precomputed with MMseqs2 for sequences where alignments were available.</li>\n<li>Fallback mode: dummy/no-MSA features used when MSA coverage was not available.</li>\n<li>Multi-conformations: up to 5 conformations per RNA sequence included.</li>\n<li>Masked residues: unresolved residues in PDB masked to avoid noisy supervision.</li></ul></li>\n</ul>\n<p>===========================================================</p>\n<h2>Conformation Selection Strategy</h2>\n<ul>\n<li>Generate multiple conformations per sequence.  </li>\n<li>Step 1: Select Top-1 structure by pLDDT score (highest confidence).  </li>\n<li>Step 2: From the remaining predictions, iteratively select conformations that maximize RMSD diversity using the Kabsch RMSD algorithm.  </li>\n<li>Final output: Top-5 diverse conformations per sequence.  </li>\n</ul>\n<p>===========================================================</p>\n<h2>📊 Results Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Model / Strategy</th>\n<th>Sequence Length</th>\n<th>Training Steps</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Protenix</strong> (standard)</td>\n<td>≤ 800</td>\n<td>8000</td>\n<td>0.46388</td>\n</tr>\n<tr>\n<td><strong>Protenix + RibonanzaNet</strong> (hybrid, &gt;800 handled)</td>\n<td>≤ 800 + &gt;800</td>\n<td>8000</td>\n<td>0.478</td>\n</tr>\n<tr>\n<td><strong>Protenix + Nufold</strong> (hybrid, &gt;800 handled)</td>\n<td>≤ 800 + &gt;800</td>\n<td>8000</td>\n<td>0.479</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Submission</strong>: Single-model, no ensemble → <strong>0.46388</strong> on leaderboard.  </li>\n</ul>",
  "messages": [
    {
      "id": "3296239",
      "postDate": "09/30/2025 13:40:58",
      "content": "<p><strong>Many thanks to Kaggle and the competition hosts for organizing this great challenge and giving us the chance to take part.</strong></p>\n<h3>📂 Dataset Creation</h3>\n<p>We built our training dataset by merging several complementary sources:</p>\n<ul>\n<li><p><strong>Stanford RNA 3D Folding Competition Data</strong></p>\n<ul>\n<li>Versions: v1 and v2 (v2 used as the primary training set).</li></ul></li>\n<li><p><strong>CASP16 Top Predictions (Pseudo-labels)</strong></p>\n<ul>\n<li>We included Top-1 predicted structures per target as pseudo-labels.</li></ul></li>\n<li><p><strong>RNA PDB Database (cutoff: March 31, 2025)</strong></p>\n<ul>\n<li>Only entries released on or before March 31, 2025 were included.</li></ul></li>\n</ul>\n<hr>\n<h3>🧩 Handling Missing Residues</h3>\n<ul>\n<li>Residues with unresolved atoms in PDB were kept in the sequence but masked in coordinates.</li>\n<li>Missing atoms were zero-padded or ignored in loss computation.</li>\n<li>This ensures reliable training without losing sequence context.</li>\n</ul>\n<hr>\n<h3>🔁 Multiple Conformations</h3>\n<ul>\n<li>For some datasets (PDB and CASP16), we generated up to 5 conformations per RNA sequence.</li>\n<li>These conformations were stored as (num_conf, seq_len, 3) arrays in the dataset.</li>\n<li>During training, we supported two strategies:<ol>\n<li>Top-1 only → using the first conformation as a deterministic baseline.</li>\n<li>Multi-conf sampling → exposing the model to all 5 conformations, either:</li></ol><ul>\n<li>Returned as coordinate_multi (all conformations available at once), or</li>\n<li>Randomly selecting one conformation per epoch, so the model sees different structures across training.</li></ul></li>\n<li>This approach increased structural diversity and improved generalization without changing the overall dataset size.</li>\n</ul>\n<hr>\n<h3>🔬 MSA Pipeline</h3>\n<ul>\n<li>For sequences where evolutionary information was available, we incorporated Multiple Sequence Alignments (MSA) into training.</li>\n<li>We used MMseqs2 to build RNA MSAs against a curated RNA sequence database.</li>\n<li>Precomputed MSAs were stored in a dedicated directory and linked during training.</li>\n<li>For sequences without MSA coverage, the pipeline fell back to dummy/no-MSA features, ensuring consistent input formats.</li>\n<li>This hybrid strategy enriched structural context for many targets and improved accuracy on conserved RNAs.</li>\n</ul>\n<hr>\n<h3>📊 Final Outputs</h3>\n<ul>\n<li>Consolidated merged sequence dataset across competition, CASP16, and PDB.</li>\n<li>Aligned atom-level labels with missing residues masked and up to 5 conformations per sequence.</li>\n<li>Precomputed MSAs (via MMseqs2) for a subset of structures.</li>\n</ul>\n<p>===========================================================</p>\n<h2>⚙️ Training Configuration</h2>\n<ul>\n<li><p><strong>Hardware</strong></p>\n<ul>\n<li>GPU: NVIDIA H100 (96 GB)</li>\n<li>Mixed precision: bfloat16 (bf16)</li></ul></li>\n<li><p><strong>Optimization</strong></p>\n<ul>\n<li>Batch size: 8</li>\n<li>Max steps: 12,000</li>\n<li>Warmup steps: 50</li>\n<li>Learning rate: 1e-4</li>\n<li>Crop size: 800 nucleotides</li></ul></li>\n<li><p><strong>Sampling</strong></p>\n<ul>\n<li>Diffusion steps: 20</li></ul></li>\n<li><p><strong>Evaluation</strong></p>\n<ul>\n<li>Checkpoint interval: every 2,000 steps</li>\n<li>Evaluation interval: every 50,000 steps</li></ul></li>\n<li><p><strong>Features</strong></p>\n<ul>\n<li>MSA pipeline: precomputed with MMseqs2 for sequences where alignments were available.</li>\n<li>Fallback mode: dummy/no-MSA features used when MSA coverage was not available.</li>\n<li>Multi-conformations: up to 5 conformations per RNA sequence included.</li>\n<li>Masked residues: unresolved residues in PDB masked to avoid noisy supervision.</li></ul></li>\n</ul>\n<p>===========================================================</p>\n<h2>Conformation Selection Strategy</h2>\n<ul>\n<li>Generate multiple conformations per sequence.  </li>\n<li>Step 1: Select Top-1 structure by pLDDT score (highest confidence).  </li>\n<li>Step 2: From the remaining predictions, iteratively select conformations that maximize RMSD diversity using the Kabsch RMSD algorithm.  </li>\n<li>Final output: Top-5 diverse conformations per sequence.  </li>\n</ul>\n<p>===========================================================</p>\n<h2>📊 Results Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Model / Strategy</th>\n<th>Sequence Length</th>\n<th>Training Steps</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Protenix</strong> (standard)</td>\n<td>≤ 800</td>\n<td>8000</td>\n<td>0.46388</td>\n</tr>\n<tr>\n<td><strong>Protenix + RibonanzaNet</strong> (hybrid, &gt;800 handled)</td>\n<td>≤ 800 + &gt;800</td>\n<td>8000</td>\n<td>0.478</td>\n</tr>\n<tr>\n<td><strong>Protenix + Nufold</strong> (hybrid, &gt;800 handled)</td>\n<td>≤ 800 + &gt;800</td>\n<td>8000</td>\n<td>0.479</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Submission</strong>: Single-model, no ensemble → <strong>0.46388</strong> on leaderboard.  </li>\n</ul>",
      "rawMarkdown": "**Many thanks to Kaggle and the competition hosts for organizing this great challenge and giving us the chance to take part.**\n### 📂 Dataset Creation\n\nWe built our training dataset by merging several complementary sources:\n\n- **Stanford RNA 3D Folding Competition Data**\n  - Versions: v1 and v2 (v2 used as the primary training set).\n\n- **CASP16 Top Predictions (Pseudo-labels)**\n  - We included Top-1 predicted structures per target as pseudo-labels.\n\n- **RNA PDB Database (cutoff: March 31, 2025)**\n  - Only entries released on or before March 31, 2025 were included.\n\n---\n\n### 🧩 Handling Missing Residues\n- Residues with unresolved atoms in PDB were kept in the sequence but masked in coordinates.\n- Missing atoms were zero-padded or ignored in loss computation.\n- This ensures reliable training without losing sequence context.\n\n---\n\n### 🔁 Multiple Conformations\n- For some datasets (PDB and CASP16), we generated up to 5 conformations per RNA sequence.\n- These conformations were stored as (num_conf, seq_len, 3) arrays in the dataset.\n- During training, we supported two strategies:\n  1. Top-1 only → using the first conformation as a deterministic baseline.\n  2. Multi-conf sampling → exposing the model to all 5 conformations, either:\n     - Returned as coordinate_multi (all conformations available at once), or\n     - Randomly selecting one conformation per epoch, so the model sees different structures across training.\n- This approach increased structural diversity and improved generalization without changing the overall dataset size.\n\n---\n\n### 🔬 MSA Pipeline\n- For sequences where evolutionary information was available, we incorporated Multiple Sequence Alignments (MSA) into training.\n- We used MMseqs2 to build RNA MSAs against a curated RNA sequence database.\n- Precomputed MSAs were stored in a dedicated directory and linked during training.\n- For sequences without MSA coverage, the pipeline fell back to dummy/no-MSA features, ensuring consistent input formats.\n- This hybrid strategy enriched structural context for many targets and improved accuracy on conserved RNAs.\n\n---\n\n### 📊 Final Outputs\n- Consolidated merged sequence dataset across competition, CASP16, and PDB.\n- Aligned atom-level labels with missing residues masked and up to 5 conformations per sequence.\n- Precomputed MSAs (via MMseqs2) for a subset of structures.\n\n===========================================================\n\n## ⚙️ Training Configuration\n\n- **Hardware**\n  - GPU: NVIDIA H100 (96 GB)\n  - Mixed precision: bfloat16 (bf16)\n\n- **Optimization**\n  - Batch size: 8\n  - Max steps: 12,000\n  - Warmup steps: 50\n  - Learning rate: 1e-4\n  - Crop size: 800 nucleotides\n\n- **Sampling**\n  - Diffusion steps: 20\n\n- **Evaluation**\n  - Checkpoint interval: every 2,000 steps\n  - Evaluation interval: every 50,000 steps\n\n- **Features**\n  - MSA pipeline: precomputed with MMseqs2 for sequences where alignments were available.\n  - Fallback mode: dummy/no-MSA features used when MSA coverage was not available.\n  - Multi-conformations: up to 5 conformations per RNA sequence included.\n  - Masked residues: unresolved residues in PDB masked to avoid noisy supervision.\n\n===========================================================\n\n##  Conformation Selection Strategy\n\n- Generate multiple conformations per sequence.  \n- Step 1: Select Top-1 structure by pLDDT score (highest confidence).  \n- Step 2: From the remaining predictions, iteratively select conformations that maximize RMSD diversity using the Kabsch RMSD algorithm.  \n- Final output: Top-5 diverse conformations per sequence.  \n\n===========================================================\n\n## 📊 Results Summary\n\n| Model / Strategy                                     | Sequence Length | Training Steps | Score   |\n|------------------------------------------------------|-----------------|----------------|---------|\n| **Protenix** (standard)                              | ≤ 800           | 8000           | 0.46388 |\n| **Protenix + RibonanzaNet** (hybrid, >800 handled)   | ≤ 800 + >800    | 8000           | 0.478   |\n| **Protenix + Nufold** (hybrid, >800 handled)   | ≤ 800 + >800    | 8000           | 0.479   |\n\n- **Submission**: Single-model, no ensemble → **0.46388** on leaderboard.",
      "votes": null
    },
    {
      "id": "3296278",
      "postDate": "09/30/2025 14:59:39",
      "content": "<p>Great writeup, and congrats on your performance as well as on the wide impact of the protenix fine-tuning code that you released during the training phase!</p>\n<p>Have you tried any re-training with a PDB cutoff date after March 31, 2025? E.g., up to the training phase close near May 29, 2025?</p>",
      "rawMarkdown": "Great writeup, and congrats on your performance as well as on the wide impact of the protenix fine-tuning code that you released during the training phase!\n\nHave you tried any re-training with a PDB cutoff date after March 31, 2025? E.g., up to the training phase close near May 29, 2025?",
      "votes": null
    },
    {
      "id": "3296292",
      "postDate": "09/30/2025 15:28:49",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> Not yet — I’m planning to retrain the model with the May 29, 2025, cutoff data in a couple of days and then test it.”</p>",
      "rawMarkdown": "rhijudas Not yet — I’m planning to retrain the model with the May 29, 2025, cutoff data in a couple of days and then test it.”",
      "votes": null
    },
    {
      "id": "3297898",
      "postDate": "10/04/2025 01:09:56",
      "content": "<p>kaggle presentation slides as attached<br>\nStanfordRNA_HN-local-copy.pdf</p>",
      "rawMarkdown": "kaggle presentation slides as attached\nStanfordRNA_HN-local-copy.pdf",
      "votes": null
    },
    {
      "id": "3414305",
      "postDate": "02/26/2026 13:39:23",
      "content": "<p>Hi, I'm looking to adapt your excellent training pipeline for this year's competition. To help me determine the right instance specifications to rent, could you share the GPU requirements and disk storage used during your training run?</p>",
      "rawMarkdown": "Hi, I'm looking to adapt your excellent training pipeline for this year's competition. To help me determine the right instance specifications to rent, could you share the GPU requirements and disk storage used during your training run?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3296278,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "09/30/2025 14:59:39",
      "content": "<p>Great writeup, and congrats on your performance as well as on the wide impact of the protenix fine-tuning code that you released during the training phase!</p>\n<p>Have you tried any re-training with a PDB cutoff date after March 31, 2025? E.g., up to the training phase close near May 29, 2025?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3296292,
          "author_name": "arunodhayan",
          "author_url": "",
          "post_date": "09/30/2025 15:28:49",
          "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> Not yet — I’m planning to retrain the model with the May 29, 2025, cutoff data in a couple of days and then test it.”</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3297898,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/04/2025 01:09:56",
      "content": "<p>kaggle presentation slides as attached<br>\nStanfordRNA_HN-local-copy.pdf</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3414305,
      "author_name": "yutongzhang20080108",
      "author_url": "",
      "post_date": "02/26/2026 13:39:23",
      "content": "<p>Hi, I'm looking to adapt your excellent training pipeline for this year's competition. To help me determine the right instance specifications to rent, could you share the GPU requirements and disk storage used during your training run?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3296239": "**Many thanks to Kaggle and the competition hosts for organizing this great challenge and giving us the chance to take part.**\n### 📂 Dataset Creation\n\nWe built our training dataset by merging several complementary sources:\n\n- **Stanford RNA 3D Folding Competition Data**\n  - Versions: v1 and v2 (v2 used as the primary training set).\n\n- **CASP16 Top Predictions (Pseudo-labels)**\n  - We included Top-1 predicted structures per target as pseudo-labels.\n\n- **RNA PDB Database (cutoff: March 31, 2025)**\n  - Only entries released on or before March 31, 2025 were included.\n\n---\n\n### 🧩 Handling Missing Residues\n- Residues with unresolved atoms in PDB were kept in the sequence but masked in coordinates.\n- Missing atoms were zero-padded or ignored in loss computation.\n- This ensures reliable training without losing sequence context.\n\n---\n\n### 🔁 Multiple Conformations\n- For some datasets (PDB and CASP16), we generated up to 5 conformations per RNA sequence.\n- These conformations were stored as (num_conf, seq_len, 3) arrays in the dataset.\n- During training, we supported two strategies:\n  1. Top-1 only → using the first conformation as a deterministic baseline.\n  2. Multi-conf sampling → exposing the model to all 5 conformations, either:\n     - Returned as coordinate_multi (all conformations available at once), or\n     - Randomly selecting one conformation per epoch, so the model sees different structures across training.\n- This approach increased structural diversity and improved generalization without changing the overall dataset size.\n\n---\n\n### 🔬 MSA Pipeline\n- For sequences where evolutionary information was available, we incorporated Multiple Sequence Alignments (MSA) into training.\n- We used MMseqs2 to build RNA MSAs against a curated RNA sequence database.\n- Precomputed MSAs were stored in a dedicated directory and linked during training.\n- For sequences without MSA coverage, the pipeline fell back to dummy/no-MSA features, ensuring consistent input formats.\n- This hybrid strategy enriched structural context for many targets and improved accuracy on conserved RNAs.\n\n---\n\n### 📊 Final Outputs\n- Consolidated merged sequence dataset across competition, CASP16, and PDB.\n- Aligned atom-level labels with missing residues masked and up to 5 conformations per sequence.\n- Precomputed MSAs (via MMseqs2) for a subset of structures.\n\n===========================================================\n\n## ⚙️ Training Configuration\n\n- **Hardware**\n  - GPU: NVIDIA H100 (96 GB)\n  - Mixed precision: bfloat16 (bf16)\n\n- **Optimization**\n  - Batch size: 8\n  - Max steps: 12,000\n  - Warmup steps: 50\n  - Learning rate: 1e-4\n  - Crop size: 800 nucleotides\n\n- **Sampling**\n  - Diffusion steps: 20\n\n- **Evaluation**\n  - Checkpoint interval: every 2,000 steps\n  - Evaluation interval: every 50,000 steps\n\n- **Features**\n  - MSA pipeline: precomputed with MMseqs2 for sequences where alignments were available.\n  - Fallback mode: dummy/no-MSA features used when MSA coverage was not available.\n  - Multi-conformations: up to 5 conformations per RNA sequence included.\n  - Masked residues: unresolved residues in PDB masked to avoid noisy supervision.\n\n===========================================================\n\n##  Conformation Selection Strategy\n\n- Generate multiple conformations per sequence.  \n- Step 1: Select Top-1 structure by pLDDT score (highest confidence).  \n- Step 2: From the remaining predictions, iteratively select conformations that maximize RMSD diversity using the Kabsch RMSD algorithm.  \n- Final output: Top-5 diverse conformations per sequence.  \n\n===========================================================\n\n## 📊 Results Summary\n\n| Model / Strategy                                     | Sequence Length | Training Steps | Score   |\n|------------------------------------------------------|-----------------|----------------|---------|\n| **Protenix** (standard)                              | ≤ 800           | 8000           | 0.46388 |\n| **Protenix + RibonanzaNet** (hybrid, >800 handled)   | ≤ 800 + >800    | 8000           | 0.478   |\n| **Protenix + Nufold** (hybrid, >800 handled)   | ≤ 800 + >800    | 8000           | 0.479   |\n\n- **Submission**: Single-model, no ensemble → **0.46388** on leaderboard.",
    "3296278": "Great writeup, and congrats on your performance as well as on the wide impact of the protenix fine-tuning code that you released during the training phase!\n\nHave you tried any re-training with a PDB cutoff date after March 31, 2025? E.g., up to the training phase close near May 29, 2025?",
    "3296292": "rhijudas Not yet — I’m planning to retrain the model with the May 29, 2025, cutoff data in a couple of days and then test it.”",
    "3297898": "kaggle presentation slides as attached\nStanfordRNA_HN-local-copy.pdf",
    "3414305": "Hi, I'm looking to adapt your excellent training pipeline for this year's competition. To help me determine the right instance specifications to rent, could you share the GPU requirements and disk storage used during your training run?"
  },
  "source": "meta"
}