{
  "id": 572850,
  "title": " 🧬 Why RNA 3D Structure Prediction is Hard — And How We're Tackling It",
  "url": "/competitions/stanford-rna-3d-folding/discussion/572850",
  "author_name": "",
  "post_date": "2025-04-11T17:41:19.973470500Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Understanding RNA 3D structure is one of the grand challenges in computational biology—and it's what this competition is all about. But what makes this problem so hard? Why isn't it as straightforward as protein structure prediction?</p>\n<h2>In this post, I’ll break down some of the core challenges, and outline strategies researchers (and Kagglers!) use to overcome them.</h2>\n<h2>🔄 1. RNA Is Flexible and Dynamic</h2>\n<p>Unlike proteins, RNA molecules often adopt <strong>multiple stable conformations</strong>. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> One RNA sequence ≠ one fixed structure  </li>\n</ul>\n<h2>- 🧠 <strong>Strategy:</strong> Use ensemble learning or train models to predict <em>multiple</em> conformations (like the 5 structures per target in this comp!).</h2>\n<h2>🧩 2. Sparse High-Quality Labels</h2>\n<p>There are <strong>far fewer RNA structures</strong> available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Not enough labeled data  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Use pretraining on unlabeled sequences (self-supervised learning)  </li>\n<li>Use data augmentation and structure perturbations  </li>\n</ul>\n<h2>- Leverage <strong>multiple conformations</strong> per sample to improve robustness</h2>\n<h2>🧠 3. RNA Structure Is Hierarchical</h2>\n<p>RNA doesn’t just form a 3D shape from the sequence directly. It forms <strong>secondary structures</strong> (like stems, loops, bulges) which then fold into 3D <strong>tertiary structures</strong>.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Modeling complex intermediate dependencies  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Encode secondary structure explicitly  </li>\n<li>Use graph-based models or attention layers to capture long-range interactions  </li>\n</ul>\n<h2>- Add <strong>domain knowledge</strong> (e.g., Watson-Crick base pairing rules) into the model</h2>\n<h2>📐 4. Geometry Matters – A Lot</h2>\n<p>Structure prediction isn’t just about finding the <em>right</em> atoms, but placing them in the <em>right 3D space</em>. Models must learn to output <strong>rotationally and translationally invariant</strong> predictions, while capturing 3D distances and angles precisely.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Predicting correct atomic coordinates in 3D space  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Use geometric deep learning or equivariant networks (e.g., SE(3)-transformers)  </li>\n</ul>\n<h2>- Add loss functions based on <strong>distances</strong>, <strong>angles</strong>, and <strong>torsions</strong></h2>\n<h2>🔗 5. It's a Graph Problem in Disguise</h2>\n<p>RNA structures aren’t just linear sequences—they’re better represented as <strong>graphs</strong> where nodes are nucleotides and edges represent bonds or proximity.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Standard models like RNNs/CNNs aren’t enough  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Use Graph Neural Networks (GNNs)  </li>\n<li>Construct graphs using base-pairing info or spatial distance cutoffs  </li>\n</ul>\n<h2>- Incorporate 3D geometry directly into the graph edges/features</h2>\n<h2>✅ Final Thoughts</h2>\n<p>RNA 3D prediction combines the complexity of biology, physics, and deep learning in one task. This competition pushes us to think creatively—how do we model structure, uncertainty, geometry, and data scarcity all at once?<br>\nFeel free to share your thoughts below — what’s been the hardest part of this problem for you so far?</p>",
  "messages": [
    {
      "id": "3176773",
      "postDate": "04/11/2025 17:41:19",
      "content": "<p>Understanding RNA 3D structure is one of the grand challenges in computational biology—and it's what this competition is all about. But what makes this problem so hard? Why isn't it as straightforward as protein structure prediction?</p>\n<h2>In this post, I’ll break down some of the core challenges, and outline strategies researchers (and Kagglers!) use to overcome them.</h2>\n<h2>🔄 1. RNA Is Flexible and Dynamic</h2>\n<p>Unlike proteins, RNA molecules often adopt <strong>multiple stable conformations</strong>. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> One RNA sequence ≠ one fixed structure  </li>\n</ul>\n<h2>- 🧠 <strong>Strategy:</strong> Use ensemble learning or train models to predict <em>multiple</em> conformations (like the 5 structures per target in this comp!).</h2>\n<h2>🧩 2. Sparse High-Quality Labels</h2>\n<p>There are <strong>far fewer RNA structures</strong> available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Not enough labeled data  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Use pretraining on unlabeled sequences (self-supervised learning)  </li>\n<li>Use data augmentation and structure perturbations  </li>\n</ul>\n<h2>- Leverage <strong>multiple conformations</strong> per sample to improve robustness</h2>\n<h2>🧠 3. RNA Structure Is Hierarchical</h2>\n<p>RNA doesn’t just form a 3D shape from the sequence directly. It forms <strong>secondary structures</strong> (like stems, loops, bulges) which then fold into 3D <strong>tertiary structures</strong>.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Modeling complex intermediate dependencies  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Encode secondary structure explicitly  </li>\n<li>Use graph-based models or attention layers to capture long-range interactions  </li>\n</ul>\n<h2>- Add <strong>domain knowledge</strong> (e.g., Watson-Crick base pairing rules) into the model</h2>\n<h2>📐 4. Geometry Matters – A Lot</h2>\n<p>Structure prediction isn’t just about finding the <em>right</em> atoms, but placing them in the <em>right 3D space</em>. Models must learn to output <strong>rotationally and translationally invariant</strong> predictions, while capturing 3D distances and angles precisely.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Predicting correct atomic coordinates in 3D space  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Use geometric deep learning or equivariant networks (e.g., SE(3)-transformers)  </li>\n</ul>\n<h2>- Add loss functions based on <strong>distances</strong>, <strong>angles</strong>, and <strong>torsions</strong></h2>\n<h2>🔗 5. It's a Graph Problem in Disguise</h2>\n<p>RNA structures aren’t just linear sequences—they’re better represented as <strong>graphs</strong> where nodes are nucleotides and edges represent bonds or proximity.</p>\n<ul>\n<li>📉 <strong>Challenge:</strong> Standard models like RNNs/CNNs aren’t enough  </li>\n<li>🧠 <strong>Strategy:</strong>  </li>\n<li>Use Graph Neural Networks (GNNs)  </li>\n<li>Construct graphs using base-pairing info or spatial distance cutoffs  </li>\n</ul>\n<h2>- Incorporate 3D geometry directly into the graph edges/features</h2>\n<h2>✅ Final Thoughts</h2>\n<p>RNA 3D prediction combines the complexity of biology, physics, and deep learning in one task. This competition pushes us to think creatively—how do we model structure, uncertainty, geometry, and data scarcity all at once?<br>\nFeel free to share your thoughts below — what’s been the hardest part of this problem for you so far?</p>",
      "rawMarkdown": "Understanding RNA 3D structure is one of the grand challenges in computational biology—and it's what this competition is all about. But what makes this problem so hard? Why isn't it as straightforward as protein structure prediction?\n\nIn this post, I’ll break down some of the core challenges, and outline strategies researchers (and Kagglers!) use to overcome them.\n\n---\n\n## 🔄 1. RNA Is Flexible and Dynamic\n\nUnlike proteins, RNA molecules often adopt **multiple stable conformations**. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.\n\n- 📉 **Challenge:** One RNA sequence ≠ one fixed structure  \n- 🧠 **Strategy:** Use ensemble learning or train models to predict *multiple* conformations (like the 5 structures per target in this comp!).\n\n---\n\n## 🧩 2. Sparse High-Quality Labels\n\nThere are **far fewer RNA structures** available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.\n\n- 📉 **Challenge:** Not enough labeled data  \n- 🧠 **Strategy:**  \n  - Use pretraining on unlabeled sequences (self-supervised learning)  \n  - Use data augmentation and structure perturbations  \n  - Leverage **multiple conformations** per sample to improve robustness\n\n---\n\n## 🧠 3. RNA Structure Is Hierarchical\n\nRNA doesn’t just form a 3D shape from the sequence directly. It forms **secondary structures** (like stems, loops, bulges) which then fold into 3D **tertiary structures**.\n\n- 📉 **Challenge:** Modeling complex intermediate dependencies  \n- 🧠 **Strategy:**  \n  - Encode secondary structure explicitly  \n  - Use graph-based models or attention layers to capture long-range interactions  \n  - Add **domain knowledge** (e.g., Watson-Crick base pairing rules) into the model\n\n---\n\n## 📐 4. Geometry Matters – A Lot\n\nStructure prediction isn’t just about finding the *right* atoms, but placing them in the *right 3D space*. Models must learn to output **rotationally and translationally invariant** predictions, while capturing 3D distances and angles precisely.\n\n- 📉 **Challenge:** Predicting correct atomic coordinates in 3D space  \n- 🧠 **Strategy:**  \n  - Use geometric deep learning or equivariant networks (e.g., SE(3)-transformers)  \n  - Add loss functions based on **distances**, **angles**, and **torsions**\n\n---\n\n## 🔗 5. It's a Graph Problem in Disguise\n\nRNA structures aren’t just linear sequences—they’re better represented as **graphs** where nodes are nucleotides and edges represent bonds or proximity.\n\n- 📉 **Challenge:** Standard models like RNNs/CNNs aren’t enough  \n- 🧠 **Strategy:**  \n  - Use Graph Neural Networks (GNNs)  \n  - Construct graphs using base-pairing info or spatial distance cutoffs  \n  - Incorporate 3D geometry directly into the graph edges/features\n\n---\n\n## ✅ Final Thoughts\n\nRNA 3D prediction combines the complexity of biology, physics, and deep learning in one task. This competition pushes us to think creatively—how do we model structure, uncertainty, geometry, and data scarcity all at once?\n\nFeel free to share your thoughts below — what’s been the hardest part of this problem for you so far?",
      "votes": null
    },
    {
      "id": "3176923",
      "postDate": "04/11/2025 23:48:49",
      "content": "<p>I applaud your attempt to provide some perspective on how to approach this problem. Yet there are too many sweeping generalizations in your writing that simply aren't true.</p>\n<blockquote>\n  <p>Unlike proteins, RNA molecules often adopt multiple stable conformations. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.</p>\n</blockquote>\n<p>All biological molecules are flexible and adopt different conformations depending on their environments. It simply isn't true that structured RNAs are more flexible than proteins. Proteins often have domains connected by flexible linkers that can swing around by large distances.</p>\n<blockquote>\n  <p>There are far fewer RNA structures available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.</p>\n</blockquote>\n<p>Very true, and this is actually the main reason RNAs are difficult to model. RNA structure should be easier to predict than protein structure because: 1) There are fewer building blocks (4 versus 20); 2) RNA molecules are generally smaller which translates into a fewer degrees of freedom. There are many proteins &gt; 1000 residues, while most biologically relevant structured RNAs are smaller than that; 3) There are defined pairings of RNA building blocks (A with U, G mostly with C and sometimes with U), which again reduces the degrees of freedom. There is no such thing in proteins that lysine always interacts with aspartate, or any other pair of amino acids. Yet the reason we can predict proteins better than RNA is because we have at least an order of magnitude more protein structures to learn from than RNA structures.</p>\n<blockquote>\n  <p>RNA Structure Is Hierarchical</p>\n</blockquote>\n<p>So is protein's structure.</p>\n<blockquote>\n  <p>Structure prediction isn’t just about finding the right atoms, but placing them in the right 3D space.</p>\n</blockquote>\n<p>Structure prediction is never about finding the right atoms. We know ahead of time, based on RNA sequence alone, what atoms should be present. We just have to put them in correct spatial positions.</p>\n<blockquote>\n  <p>Models must learn to output rotationally and translationally invariant predictions, while capturing 3D distances and angles precisely.</p>\n</blockquote>\n<p>Also not the case. It is not a problem if models are rotationally and translationally variable as long as they can be brought to similar conformations by simple rotations and translations. One would have to have a huge number of extremely accurate distance restraints, either generated experimentally or by predictions from MSAs, to create \"rotationally and translationally invariant\" models.</p>",
      "rawMarkdown": "I applaud your attempt to provide some perspective on how to approach this problem. Yet there are too many sweeping generalizations in your writing that simply aren't true.\n\n> Unlike proteins, RNA molecules often adopt multiple stable conformations. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.\n\nAll biological molecules are flexible and adopt different conformations depending on their environments. It simply isn't true that structured RNAs are more flexible than proteins. Proteins often have domains connected by flexible linkers that can swing around by large distances.\n\n> There are far fewer RNA structures available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.\n\nVery true, and this is actually the main reason RNAs are difficult to model. RNA structure should be easier to predict than protein structure because: 1) There are fewer building blocks (4 versus 20); 2) RNA molecules are generally smaller which translates into a fewer degrees of freedom. There are many proteins > 1000 residues, while most biologically relevant structured RNAs are smaller than that; 3) There are defined pairings of RNA building blocks (A with U, G mostly with C and sometimes with U), which again reduces the degrees of freedom. There is no such thing in proteins that lysine always interacts with aspartate, or any other pair of amino acids. Yet the reason we can predict proteins better than RNA is because we have at least an order of magnitude more protein structures to learn from than RNA structures.\n\n> RNA Structure Is Hierarchical\n\nSo is protein's structure.\n\n> Structure prediction isn’t just about finding the right atoms, but placing them in the right 3D space.\n\nStructure prediction is never about finding the right atoms. We know ahead of time, based on RNA sequence alone, what atoms should be present. We just have to put them in correct spatial positions.\n\n> Models must learn to output rotationally and translationally invariant predictions, while capturing 3D distances and angles precisely.\n\nAlso not the case. It is not a problem if models are rotationally and translationally variable as long as they can be brought to similar conformations by simple rotations and translations. One would have to have a huge number of extremely accurate distance restraints, either generated experimentally or by predictions from MSAs, to create \"rotationally and translationally invariant\" models.",
      "votes": null
    },
    {
      "id": "3177070",
      "postDate": "04/12/2025 07:40:33",
      "content": "<p>I am very thankful for your insights on this ! </p>",
      "rawMarkdown": "I am very thankful for your insights on this !",
      "votes": null
    },
    {
      "id": "3191611",
      "postDate": "05/01/2025 23:12:45",
      "content": "<p>Hi, sorry to bother you, but could you please tell us, are we limited to  single RNA or there can be multi-RNA complexes, i.e. quaternary RNA structures involved?</p>",
      "rawMarkdown": "Hi, sorry to bother you, but could you please tell us, are we limited to  single RNA or there can be multi-RNA complexes, i.e. quaternary RNA structures involved?",
      "votes": null
    },
    {
      "id": "3191638",
      "postDate": "05/02/2025 00:16:43",
      "content": "<p>It is a question for competition organizers.</p>",
      "rawMarkdown": "It is a question for competition organizers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3176923,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/11/2025 23:48:49",
      "content": "<p>I applaud your attempt to provide some perspective on how to approach this problem. Yet there are too many sweeping generalizations in your writing that simply aren't true.</p>\n<blockquote>\n  <p>Unlike proteins, RNA molecules often adopt multiple stable conformations. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.</p>\n</blockquote>\n<p>All biological molecules are flexible and adopt different conformations depending on their environments. It simply isn't true that structured RNAs are more flexible than proteins. Proteins often have domains connected by flexible linkers that can swing around by large distances.</p>\n<blockquote>\n  <p>There are far fewer RNA structures available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.</p>\n</blockquote>\n<p>Very true, and this is actually the main reason RNAs are difficult to model. RNA structure should be easier to predict than protein structure because: 1) There are fewer building blocks (4 versus 20); 2) RNA molecules are generally smaller which translates into a fewer degrees of freedom. There are many proteins &gt; 1000 residues, while most biologically relevant structured RNAs are smaller than that; 3) There are defined pairings of RNA building blocks (A with U, G mostly with C and sometimes with U), which again reduces the degrees of freedom. There is no such thing in proteins that lysine always interacts with aspartate, or any other pair of amino acids. Yet the reason we can predict proteins better than RNA is because we have at least an order of magnitude more protein structures to learn from than RNA structures.</p>\n<blockquote>\n  <p>RNA Structure Is Hierarchical</p>\n</blockquote>\n<p>So is protein's structure.</p>\n<blockquote>\n  <p>Structure prediction isn’t just about finding the right atoms, but placing them in the right 3D space.</p>\n</blockquote>\n<p>Structure prediction is never about finding the right atoms. We know ahead of time, based on RNA sequence alone, what atoms should be present. We just have to put them in correct spatial positions.</p>\n<blockquote>\n  <p>Models must learn to output rotationally and translationally invariant predictions, while capturing 3D distances and angles precisely.</p>\n</blockquote>\n<p>Also not the case. It is not a problem if models are rotationally and translationally variable as long as they can be brought to similar conformations by simple rotations and translations. One would have to have a huge number of extremely accurate distance restraints, either generated experimentally or by predictions from MSAs, to create \"rotationally and translationally invariant\" models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3177070,
          "author_name": "vaishnavimudaliar",
          "author_url": "",
          "post_date": "04/12/2025 07:40:33",
          "content": "<p>I am very thankful for your insights on this ! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3191611,
          "author_name": "ilyakupchenko",
          "author_url": "",
          "post_date": "05/01/2025 23:12:45",
          "content": "<p>Hi, sorry to bother you, but could you please tell us, are we limited to  single RNA or there can be multi-RNA complexes, i.e. quaternary RNA structures involved?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3191638,
              "author_name": "tilii7",
              "author_url": "",
              "post_date": "05/02/2025 00:16:43",
              "content": "<p>It is a question for competition organizers.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3176773": "Understanding RNA 3D structure is one of the grand challenges in computational biology—and it's what this competition is all about. But what makes this problem so hard? Why isn't it as straightforward as protein structure prediction?\n\nIn this post, I’ll break down some of the core challenges, and outline strategies researchers (and Kagglers!) use to overcome them.\n\n---\n\n## 🔄 1. RNA Is Flexible and Dynamic\n\nUnlike proteins, RNA molecules often adopt **multiple stable conformations**. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.\n\n- 📉 **Challenge:** One RNA sequence ≠ one fixed structure  \n- 🧠 **Strategy:** Use ensemble learning or train models to predict *multiple* conformations (like the 5 structures per target in this comp!).\n\n---\n\n## 🧩 2. Sparse High-Quality Labels\n\nThere are **far fewer RNA structures** available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.\n\n- 📉 **Challenge:** Not enough labeled data  \n- 🧠 **Strategy:**  \n  - Use pretraining on unlabeled sequences (self-supervised learning)  \n  - Use data augmentation and structure perturbations  \n  - Leverage **multiple conformations** per sample to improve robustness\n\n---\n\n## 🧠 3. RNA Structure Is Hierarchical\n\nRNA doesn’t just form a 3D shape from the sequence directly. It forms **secondary structures** (like stems, loops, bulges) which then fold into 3D **tertiary structures**.\n\n- 📉 **Challenge:** Modeling complex intermediate dependencies  \n- 🧠 **Strategy:**  \n  - Encode secondary structure explicitly  \n  - Use graph-based models or attention layers to capture long-range interactions  \n  - Add **domain knowledge** (e.g., Watson-Crick base pairing rules) into the model\n\n---\n\n## 📐 4. Geometry Matters – A Lot\n\nStructure prediction isn’t just about finding the *right* atoms, but placing them in the *right 3D space*. Models must learn to output **rotationally and translationally invariant** predictions, while capturing 3D distances and angles precisely.\n\n- 📉 **Challenge:** Predicting correct atomic coordinates in 3D space  \n- 🧠 **Strategy:**  \n  - Use geometric deep learning or equivariant networks (e.g., SE(3)-transformers)  \n  - Add loss functions based on **distances**, **angles**, and **torsions**\n\n---\n\n## 🔗 5. It's a Graph Problem in Disguise\n\nRNA structures aren’t just linear sequences—they’re better represented as **graphs** where nodes are nucleotides and edges represent bonds or proximity.\n\n- 📉 **Challenge:** Standard models like RNNs/CNNs aren’t enough  \n- 🧠 **Strategy:**  \n  - Use Graph Neural Networks (GNNs)  \n  - Construct graphs using base-pairing info or spatial distance cutoffs  \n  - Incorporate 3D geometry directly into the graph edges/features\n\n---\n\n## ✅ Final Thoughts\n\nRNA 3D prediction combines the complexity of biology, physics, and deep learning in one task. This competition pushes us to think creatively—how do we model structure, uncertainty, geometry, and data scarcity all at once?\n\nFeel free to share your thoughts below — what’s been the hardest part of this problem for you so far?",
    "3176923": "I applaud your attempt to provide some perspective on how to approach this problem. Yet there are too many sweeping generalizations in your writing that simply aren't true.\n\n> Unlike proteins, RNA molecules often adopt multiple stable conformations. A single RNA sequence can fold differently depending on environmental conditions, interactions, or even time.\n\nAll biological molecules are flexible and adopt different conformations depending on their environments. It simply isn't true that structured RNAs are more flexible than proteins. Proteins often have domains connected by flexible linkers that can swing around by large distances.\n\n> There are far fewer RNA structures available in the Protein Data Bank (PDB) compared to proteins. Most RNA structures are small, incomplete, or noisy due to experimental challenges in resolving them.\n\nVery true, and this is actually the main reason RNAs are difficult to model. RNA structure should be easier to predict than protein structure because: 1) There are fewer building blocks (4 versus 20); 2) RNA molecules are generally smaller which translates into a fewer degrees of freedom. There are many proteins > 1000 residues, while most biologically relevant structured RNAs are smaller than that; 3) There are defined pairings of RNA building blocks (A with U, G mostly with C and sometimes with U), which again reduces the degrees of freedom. There is no such thing in proteins that lysine always interacts with aspartate, or any other pair of amino acids. Yet the reason we can predict proteins better than RNA is because we have at least an order of magnitude more protein structures to learn from than RNA structures.\n\n> RNA Structure Is Hierarchical\n\nSo is protein's structure.\n\n> Structure prediction isn’t just about finding the right atoms, but placing them in the right 3D space.\n\nStructure prediction is never about finding the right atoms. We know ahead of time, based on RNA sequence alone, what atoms should be present. We just have to put them in correct spatial positions.\n\n> Models must learn to output rotationally and translationally invariant predictions, while capturing 3D distances and angles precisely.\n\nAlso not the case. It is not a problem if models are rotationally and translationally variable as long as they can be brought to similar conformations by simple rotations and translations. One would have to have a huge number of extremely accurate distance restraints, either generated experimentally or by predictions from MSAs, to create \"rotationally and translationally invariant\" models.",
    "3177070": "I am very thankful for your insights on this !",
    "3191611": "Hi, sorry to bother you, but could you please tell us, are we limited to  single RNA or there can be multi-RNA complexes, i.e. quaternary RNA structures involved?",
    "3191638": "It is a question for competition organizers."
  },
  "source": "meta"
}