{
  "id": 574214,
  "title": "Potential Approaches & Ideas",
  "url": "/competitions/stanford-rna-3d-folding/discussion/574214",
  "author_name": "",
  "post_date": "2025-04-20T17:50:31.519074400Z",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey everyone,<br>\nThis RNA folding competition is definitely one of the more exciting and challenging ones! Predicting 3D structure from sequence is tough, especially with the known limitations in available 3D training data compared to proteins. Add the memory constraints with potentially long sequences and the need for 5 diverse predictions, and we've got a proper puzzle on our hands.<br>\nBased on recent progress (like AlphaFold, RoseTTAFold, and the previous RibonanzaNet comp), just feeding the sequence into a standard Transformer encoder, while a good starting point, might hit a performance ceiling/time consuming. RNA structure relies heavily on pairwise interactions (base pairing, stacking) and global topology.<br>\nHere are a few key ideas and architectural directions that seem promising, focusing on incorporating more structural bias while managing resources:</p>\n<ol>\n<li><p>Beyond 1D: Embracing Pairwise Information (Evoformer-Inspired)<br>\nWhy: A simple sequence encoder struggles to explicitly model which distant residues interact. Base pairing is fundamental! We need to represent and reason about residue pairs.<br>\nHow: Consider architectures that process both a 1D sequence representation (MSA features + sequence embeddings + positional encodings) and a 2D pairwise representation (L x L x Channels tensor).<br>\nInformation should flow between these tracks (e.g., use pair info to bias sequence attention, use sequence info to update pair features).<br>\nThe pair representation can be updated using techniques like:<br>\n2D Convolutions: Simpler to implement, captures local pairwise patterns.<br>\nSimplified Attention: Maybe attention mechanisms operating on the pair matrix (less complex than full triangular attention initially).<br>\nGoal: Explicitly model residue-residue interactions and geometric constraints.</p></li>\n<li><p>Leveraging Auxiliary Data<br>\nMSAs (Multiple Sequence Alignments): These are provided and are gold! Co-evolving residues in an MSA often imply spatial proximity or functional interaction.<br>\nHow: Generate profile features (e.g., frequency of A/C/G/U/- at each position) from the MSAs. Project these features and add them to the initial sequence embeddings. This significantly enriches the input.<br>\nChemical Mapping Data (RibonanzaNet Context): The previous competition generated RibonanzaNet based on 2D chemical mapping data (DMS/SHAPE), which probes local RNA structure.</p></li>\n<li><p>Efficient Training &amp; Refinement<br>\nRecycling: Inspired by AlphaFold, feeding the output representations back into the main processing blocks for a few iterations (with shared weights) can allow for iterative refinement within a single forward pass.<br>\nWhy: Can lead to better convergence with fewer epochs and acts as a form of regularization.<br>\nAMP (Automatic Mixed Precision): Use torch.cuda.amp (autocast and GradScaler).<br>\nWhy: Significantly reduces memory usage (especially for large pair representations) and speeds up training, crucial for fitting within VRAM limits.<br>\nGradient Accumulation: If batch size 1 still causes OOM, increase accumulation steps.<br>\nWhy: Simulates a larger batch size for stable gradients while keeping peak memory low.</p></li>\n<li><p>Coordinate Prediction &amp; Output<br>\nInitial Goal (C1' atoms): Start simple. A linear layer projecting the final sequence representation (after Pairformer blocks + Recycling) to 3D coordinates might be sufficient.<br>\nAdvanced: A dedicated Structure Module (like AlphaFold's IPA) could be explored later if coordinate generation quality is a bottleneck, but adds complexity.<br>\nGenerating 5 Predictions:<br>\nDropout Inference: Keep dropout active during inference and run the model 5 times.<br>\nEnsembling: Train multiple models (different seeds, slightly different hyperparameters) and average predictions (more resource-intensive).<br>\nSampling: If the model outputs a distribution, sample from it (less common for direct coordinate regression).</p></li>\n</ol>\n<p><strong>MAJOR CONSTRAINTS</strong>:</p>\n<ul>\n<li>Memory (OOM): This is the big one with pair representations. Aggressively use AMP, gradient accumulation, and systematically reduce model dimensions.</li>\n<li>NaN Loss Masking: Essential for training on the provided labels. Ensure your loss function only considers residues with valid coordinates.</li>\n<li>Temporal Splits: Strictly adhere to the <code>temporal_cutoff</code> for validation and training to get meaningful evaluation.</li>\n</ul>\n<p>What are your thoughts? What other approaches are you guys considering?</p>",
  "messages": [
    {
      "id": "3183305",
      "postDate": "04/20/2025 17:50:31",
      "content": "<p>Hey everyone,<br>\nThis RNA folding competition is definitely one of the more exciting and challenging ones! Predicting 3D structure from sequence is tough, especially with the known limitations in available 3D training data compared to proteins. Add the memory constraints with potentially long sequences and the need for 5 diverse predictions, and we've got a proper puzzle on our hands.<br>\nBased on recent progress (like AlphaFold, RoseTTAFold, and the previous RibonanzaNet comp), just feeding the sequence into a standard Transformer encoder, while a good starting point, might hit a performance ceiling/time consuming. RNA structure relies heavily on pairwise interactions (base pairing, stacking) and global topology.<br>\nHere are a few key ideas and architectural directions that seem promising, focusing on incorporating more structural bias while managing resources:</p>\n<ol>\n<li><p>Beyond 1D: Embracing Pairwise Information (Evoformer-Inspired)<br>\nWhy: A simple sequence encoder struggles to explicitly model which distant residues interact. Base pairing is fundamental! We need to represent and reason about residue pairs.<br>\nHow: Consider architectures that process both a 1D sequence representation (MSA features + sequence embeddings + positional encodings) and a 2D pairwise representation (L x L x Channels tensor).<br>\nInformation should flow between these tracks (e.g., use pair info to bias sequence attention, use sequence info to update pair features).<br>\nThe pair representation can be updated using techniques like:<br>\n2D Convolutions: Simpler to implement, captures local pairwise patterns.<br>\nSimplified Attention: Maybe attention mechanisms operating on the pair matrix (less complex than full triangular attention initially).<br>\nGoal: Explicitly model residue-residue interactions and geometric constraints.</p></li>\n<li><p>Leveraging Auxiliary Data<br>\nMSAs (Multiple Sequence Alignments): These are provided and are gold! Co-evolving residues in an MSA often imply spatial proximity or functional interaction.<br>\nHow: Generate profile features (e.g., frequency of A/C/G/U/- at each position) from the MSAs. Project these features and add them to the initial sequence embeddings. This significantly enriches the input.<br>\nChemical Mapping Data (RibonanzaNet Context): The previous competition generated RibonanzaNet based on 2D chemical mapping data (DMS/SHAPE), which probes local RNA structure.</p></li>\n<li><p>Efficient Training &amp; Refinement<br>\nRecycling: Inspired by AlphaFold, feeding the output representations back into the main processing blocks for a few iterations (with shared weights) can allow for iterative refinement within a single forward pass.<br>\nWhy: Can lead to better convergence with fewer epochs and acts as a form of regularization.<br>\nAMP (Automatic Mixed Precision): Use torch.cuda.amp (autocast and GradScaler).<br>\nWhy: Significantly reduces memory usage (especially for large pair representations) and speeds up training, crucial for fitting within VRAM limits.<br>\nGradient Accumulation: If batch size 1 still causes OOM, increase accumulation steps.<br>\nWhy: Simulates a larger batch size for stable gradients while keeping peak memory low.</p></li>\n<li><p>Coordinate Prediction &amp; Output<br>\nInitial Goal (C1' atoms): Start simple. A linear layer projecting the final sequence representation (after Pairformer blocks + Recycling) to 3D coordinates might be sufficient.<br>\nAdvanced: A dedicated Structure Module (like AlphaFold's IPA) could be explored later if coordinate generation quality is a bottleneck, but adds complexity.<br>\nGenerating 5 Predictions:<br>\nDropout Inference: Keep dropout active during inference and run the model 5 times.<br>\nEnsembling: Train multiple models (different seeds, slightly different hyperparameters) and average predictions (more resource-intensive).<br>\nSampling: If the model outputs a distribution, sample from it (less common for direct coordinate regression).</p></li>\n</ol>\n<p><strong>MAJOR CONSTRAINTS</strong>:</p>\n<ul>\n<li>Memory (OOM): This is the big one with pair representations. Aggressively use AMP, gradient accumulation, and systematically reduce model dimensions.</li>\n<li>NaN Loss Masking: Essential for training on the provided labels. Ensure your loss function only considers residues with valid coordinates.</li>\n<li>Temporal Splits: Strictly adhere to the <code>temporal_cutoff</code> for validation and training to get meaningful evaluation.</li>\n</ul>\n<p>What are your thoughts? What other approaches are you guys considering?</p>",
      "rawMarkdown": "Hey everyone,\nThis RNA folding competition is definitely one of the more exciting and challenging ones! Predicting 3D structure from sequence is tough, especially with the known limitations in available 3D training data compared to proteins. Add the memory constraints with potentially long sequences and the need for 5 diverse predictions, and we've got a proper puzzle on our hands.\nBased on recent progress (like AlphaFold, RoseTTAFold, and the previous RibonanzaNet comp), just feeding the sequence into a standard Transformer encoder, while a good starting point, might hit a performance ceiling/time consuming. RNA structure relies heavily on pairwise interactions (base pairing, stacking) and global topology.\nHere are a few key ideas and architectural directions that seem promising, focusing on incorporating more structural bias while managing resources:\n1. Beyond 1D: Embracing Pairwise Information (Evoformer-Inspired)\nWhy: A simple sequence encoder struggles to explicitly model which distant residues interact. Base pairing is fundamental! We need to represent and reason about residue pairs.\nHow: Consider architectures that process both a 1D sequence representation (MSA features + sequence embeddings + positional encodings) and a 2D pairwise representation (L x L x Channels tensor).\nInformation should flow between these tracks (e.g., use pair info to bias sequence attention, use sequence info to update pair features).\nThe pair representation can be updated using techniques like:\n2D Convolutions: Simpler to implement, captures local pairwise patterns.\nSimplified Attention: Maybe attention mechanisms operating on the pair matrix (less complex than full triangular attention initially).\nGoal: Explicitly model residue-residue interactions and geometric constraints.\n2. Leveraging Auxiliary Data\nMSAs (Multiple Sequence Alignments): These are provided and are gold! Co-evolving residues in an MSA often imply spatial proximity or functional interaction.\nHow: Generate profile features (e.g., frequency of A/C/G/U/- at each position) from the MSAs. Project these features and add them to the initial sequence embeddings. This significantly enriches the input.\nChemical Mapping Data (RibonanzaNet Context): The previous competition generated RibonanzaNet based on 2D chemical mapping data (DMS/SHAPE), which probes local RNA structure.\n\n3. Efficient Training & Refinement\nRecycling: Inspired by AlphaFold, feeding the output representations back into the main processing blocks for a few iterations (with shared weights) can allow for iterative refinement within a single forward pass.\nWhy: Can lead to better convergence with fewer epochs and acts as a form of regularization.\nAMP (Automatic Mixed Precision): Use torch.cuda.amp (autocast and GradScaler).\nWhy: Significantly reduces memory usage (especially for large pair representations) and speeds up training, crucial for fitting within VRAM limits.\nGradient Accumulation: If batch size 1 still causes OOM, increase accumulation steps.\nWhy: Simulates a larger batch size for stable gradients while keeping peak memory low.\n4. Coordinate Prediction & Output\nInitial Goal (C1' atoms): Start simple. A linear layer projecting the final sequence representation (after Pairformer blocks + Recycling) to 3D coordinates might be sufficient.\nAdvanced: A dedicated Structure Module (like AlphaFold's IPA) could be explored later if coordinate generation quality is a bottleneck, but adds complexity.\nGenerating 5 Predictions:\nDropout Inference: Keep dropout active during inference and run the model 5 times.\nEnsembling: Train multiple models (different seeds, slightly different hyperparameters) and average predictions (more resource-intensive).\nSampling: If the model outputs a distribution, sample from it (less common for direct coordinate regression).\n\n**MAJOR CONSTRAINTS**:\n* Memory (OOM): This is the big one with pair representations. Aggressively use AMP, gradient accumulation, and systematically reduce model dimensions.\n* NaN Loss Masking: Essential for training on the provided labels. Ensure your loss function only considers residues with valid coordinates.\n* Temporal Splits: Strictly adhere to the `temporal_cutoff` for validation and training to get meaningful evaluation.\n\nWhat are your thoughts? What other approaches are you guys considering?",
      "votes": null
    },
    {
      "id": "3184251",
      "postDate": "04/21/2025 20:48:23",
      "content": "<p>is it fierst rna competition ?</p>",
      "rawMarkdown": "is it fierst rna competition ?",
      "votes": null
    },
    {
      "id": "3184717",
      "postDate": "04/22/2025 12:01:11",
      "content": "<p>Hey can you point where is the MSA data provided by the organizers?</p>",
      "rawMarkdown": "Hey can you point where is the MSA data provided by the organizers?",
      "votes": null
    },
    {
      "id": "3185034",
      "postDate": "04/22/2025 18:21:18",
      "content": "<p>Do you chek it?</p>",
      "rawMarkdown": "Do you chek it?",
      "votes": null
    },
    {
      "id": "3185485",
      "postDate": "04/23/2025 11:29:21",
      "content": "<p>As for ensemble, I tried extracting embeddings from Rhofold→fused with embeddings from Ribonanza using linear layer→passed it to ribonanza xyz_predictor. See notebook here↓</p>\n<p><a href=\"https://www.kaggle.com/code/nanacat0520/extract-embeddings-from-rhofold-for-ensemble\" target=\"_blank\">https://www.kaggle.com/code/nanacat0520/extract-embeddings-from-rhofold-for-ensemble</a></p>\n<p>As a result, scores were slightly better than Ribonanza's but far below Rhofold's. I'm planning to USalign &amp; ensemble predicted coordinates later.</p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "As for ensemble, I tried extracting embeddings from Rhofold→fused with embeddings from Ribonanza using linear layer→passed it to ribonanza xyz_predictor. See notebook here↓\n\nhttps://www.kaggle.com/code/nanacat0520/extract-embeddings-from-rhofold-for-ensemble\n\nAs a result, scores were slightly better than Ribonanza's but far below Rhofold's. I'm planning to USalign & ensemble predicted coordinates later.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "3185627",
      "postDate": "04/23/2025 15:48:53",
      "content": "<p>Cool approach! \nNaN values bug me of so much. \nI'm trying to deal with them since they account for ~5% of data. But nothing seems to come close to getting it. \nI don't like Z scaling in loss function but without it I have seen values explode/vanish.</p>",
      "rawMarkdown": "Cool approach! \nNaN values bug me of so much. \nI'm trying to deal with them since they account for ~5% of data. But nothing seems to come close to getting it. \nI don't like Z scaling in loss function but without it I have seen values explode/vanish.",
      "votes": null
    },
    {
      "id": "3200127",
      "postDate": "05/12/2025 05:33:45",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA</a><br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA_v2\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA_v2</a></p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA_v2",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3184251,
      "author_name": "",
      "author_url": "",
      "post_date": "04/21/2025 20:48:23",
      "content": "<p>is it fierst rna competition ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3184717,
      "author_name": "ananyapathak",
      "author_url": "",
      "post_date": "04/22/2025 12:01:11",
      "content": "<p>Hey can you point where is the MSA data provided by the organizers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3200127,
          "author_name": "christie",
          "author_url": "",
          "post_date": "05/12/2025 05:33:45",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA</a><br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA_v2\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA_v2</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3185034,
      "author_name": "",
      "author_url": "",
      "post_date": "04/22/2025 18:21:18",
      "content": "<p>Do you chek it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3185485,
      "author_name": "nanacat0520",
      "author_url": "",
      "post_date": "04/23/2025 11:29:21",
      "content": "<p>As for ensemble, I tried extracting embeddings from Rhofold→fused with embeddings from Ribonanza using linear layer→passed it to ribonanza xyz_predictor. See notebook here↓</p>\n<p><a href=\"https://www.kaggle.com/code/nanacat0520/extract-embeddings-from-rhofold-for-ensemble\" target=\"_blank\">https://www.kaggle.com/code/nanacat0520/extract-embeddings-from-rhofold-for-ensemble</a></p>\n<p>As a result, scores were slightly better than Ribonanza's but far below Rhofold's. I'm planning to USalign &amp; ensemble predicted coordinates later.</p>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3185627,
          "author_name": "adhoppin",
          "author_url": "",
          "post_date": "04/23/2025 15:48:53",
          "content": "<p>Cool approach! \nNaN values bug me of so much. \nI'm trying to deal with them since they account for ~5% of data. But nothing seems to come close to getting it. \nI don't like Z scaling in loss function but without it I have seen values explode/vanish.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3183305": "Hey everyone,\nThis RNA folding competition is definitely one of the more exciting and challenging ones! Predicting 3D structure from sequence is tough, especially with the known limitations in available 3D training data compared to proteins. Add the memory constraints with potentially long sequences and the need for 5 diverse predictions, and we've got a proper puzzle on our hands.\nBased on recent progress (like AlphaFold, RoseTTAFold, and the previous RibonanzaNet comp), just feeding the sequence into a standard Transformer encoder, while a good starting point, might hit a performance ceiling/time consuming. RNA structure relies heavily on pairwise interactions (base pairing, stacking) and global topology.\nHere are a few key ideas and architectural directions that seem promising, focusing on incorporating more structural bias while managing resources:\n1. Beyond 1D: Embracing Pairwise Information (Evoformer-Inspired)\nWhy: A simple sequence encoder struggles to explicitly model which distant residues interact. Base pairing is fundamental! We need to represent and reason about residue pairs.\nHow: Consider architectures that process both a 1D sequence representation (MSA features + sequence embeddings + positional encodings) and a 2D pairwise representation (L x L x Channels tensor).\nInformation should flow between these tracks (e.g., use pair info to bias sequence attention, use sequence info to update pair features).\nThe pair representation can be updated using techniques like:\n2D Convolutions: Simpler to implement, captures local pairwise patterns.\nSimplified Attention: Maybe attention mechanisms operating on the pair matrix (less complex than full triangular attention initially).\nGoal: Explicitly model residue-residue interactions and geometric constraints.\n2. Leveraging Auxiliary Data\nMSAs (Multiple Sequence Alignments): These are provided and are gold! Co-evolving residues in an MSA often imply spatial proximity or functional interaction.\nHow: Generate profile features (e.g., frequency of A/C/G/U/- at each position) from the MSAs. Project these features and add them to the initial sequence embeddings. This significantly enriches the input.\nChemical Mapping Data (RibonanzaNet Context): The previous competition generated RibonanzaNet based on 2D chemical mapping data (DMS/SHAPE), which probes local RNA structure.\n\n3. Efficient Training & Refinement\nRecycling: Inspired by AlphaFold, feeding the output representations back into the main processing blocks for a few iterations (with shared weights) can allow for iterative refinement within a single forward pass.\nWhy: Can lead to better convergence with fewer epochs and acts as a form of regularization.\nAMP (Automatic Mixed Precision): Use torch.cuda.amp (autocast and GradScaler).\nWhy: Significantly reduces memory usage (especially for large pair representations) and speeds up training, crucial for fitting within VRAM limits.\nGradient Accumulation: If batch size 1 still causes OOM, increase accumulation steps.\nWhy: Simulates a larger batch size for stable gradients while keeping peak memory low.\n4. Coordinate Prediction & Output\nInitial Goal (C1' atoms): Start simple. A linear layer projecting the final sequence representation (after Pairformer blocks + Recycling) to 3D coordinates might be sufficient.\nAdvanced: A dedicated Structure Module (like AlphaFold's IPA) could be explored later if coordinate generation quality is a bottleneck, but adds complexity.\nGenerating 5 Predictions:\nDropout Inference: Keep dropout active during inference and run the model 5 times.\nEnsembling: Train multiple models (different seeds, slightly different hyperparameters) and average predictions (more resource-intensive).\nSampling: If the model outputs a distribution, sample from it (less common for direct coordinate regression).\n\n**MAJOR CONSTRAINTS**:\n* Memory (OOM): This is the big one with pair representations. Aggressively use AMP, gradient accumulation, and systematically reduce model dimensions.\n* NaN Loss Masking: Essential for training on the provided labels. Ensure your loss function only considers residues with valid coordinates.\n* Temporal Splits: Strictly adhere to the `temporal_cutoff` for validation and training to get meaningful evaluation.\n\nWhat are your thoughts? What other approaches are you guys considering?",
    "3184251": "is it fierst rna competition ?",
    "3184717": "Hey can you point where is the MSA data provided by the organizers?",
    "3185034": "Do you chek it?",
    "3185485": "As for ensemble, I tried extracting embeddings from Rhofold→fused with embeddings from Ribonanza using linear layer→passed it to ribonanza xyz_predictor. See notebook here↓\n\nhttps://www.kaggle.com/code/nanacat0520/extract-embeddings-from-rhofold-for-ensemble\n\nAs a result, scores were slightly better than Ribonanza's but far below Rhofold's. I'm planning to USalign & ensemble predicted coordinates later.\n\nHope this helps!",
    "3185627": "Cool approach! \nNaN values bug me of so much. \nI'm trying to deal with them since they account for ~5% of data. But nothing seems to come close to getting it. \nI don't like Z scaling in loss function but without it I have seen values explode/vanish.",
    "3200127": "https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=MSA_v2"
  },
  "source": "meta"
}