{
  "id": 568633,
  "title": "Introduction to RNA folding, part 2",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568633",
  "author_name": "",
  "post_date": "2025-03-17T04:51:10.174867100Z",
  "votes": 56,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This is a continuation of <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568445\" target=\"_blank\"><strong>an earlier post</strong></a>.</p>\n<p>First, let's understand what is it exactly that we are predicting. We will look at a G-C base pair:</p>\n<p><img src=\"https://i.ibb.co/rWd5j6b/g-c-pair.png\" alt=\"G-C pairing\"></p>\n<p>Imagine you are Superman and your right eye is shooting a laser through a thick, wavy line connected to the N atom of the C nucleotide (right next to 54o). Likewise, your left eye is shooting a laser through a thick, wavy line connected to the N atom of the G nucleotide (also next to 54o). Those two laser trajectories would become side rails of a ladder, and a G-C base pair laying flat on the screen surface would be a rung connecting those two rails. If we stacked other base pairs on top of the screen towards our face, and some under the screen away from us, we'd get a complete ladder that schematically represents a double-helical RNA. The same would be true for DNA, except that opposite A we'd have a T (Thymine) rather than U. To complete this analogy we'd have to twist that ladder to resemble a spiral staircase.</p>\n<p>Earlier we ignored the phosphate and sugar components of nucleotides, but hopefully their role is now obvious: they form the side rails and connect RNA bases vertically. Those spots where each base connects to a sugar molecule - where we shot the lasers - are C1 atoms for which we are trying to predict 3D coordinates. Let's add some facts: 1) individual bases are rigid; 2) they lay flat in the plane; 3) hydrogen bonds are always of the same length (with small angle deviations). Knowing that two RNA bases form a base pair and given near-fixed covalent bond lengths and hydrogen bond distances, it is fairly straightforward to predict C1 positions in such a base pair. A take-home message: it is much easier to predict relative positions of C1 atoms for the two bases that form a base pair (where RNA is double-stranded) than a random two bases. Any time we correctly predict that two RNA bases are paired to each other, we create a distance restraint that defines relative distances of their C1 atoms. For example, if you look into PDB structure 1SCL (it has been discussed several times), you will see that for RNA bases that pair each other the average C1-C1 distance is 10.46 angstroms (stdev 0.27). That is in addition to a well-defined rise between bases - a distance between the rungs of the ladder, if you will. For the same structure as above, average C1-C1 distance between consecutive bases is 5.49 (stdev 0.61).  Therefore, correctly predicting base-pairing patterns will make it easier to predict RNA folding in double-helical regions.</p>\n<p>Hopefully I convinced you in the paragraph above how important it is to correctly predict base pairings. Let's look again at how base-pairing happens in a simple RNA molecule.</p>\n<pre><code>&gt;RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n&gt;RNA_secondary_structure\n((((((((....))))))))\n</code></pre>\n<p>This molecule forms a hairpin, which is schematically shown below. Hydrogen bonds in canonical bases are shown as horizontal lines connecting the letters, while a non-canonical base pairing is shown as <code>GoU</code>.</p>\n<pre><code>    ___\n   /   \\\n   \\   /\n    U-\n    GoU\n    -C\n    -U\n    -U\n    -C\n    -C\n    C-\n</code></pre>\n<p>If all RNA molecules were this short, or if we had to look only 20 bases ahead to find base pairing partners, there would be no difficulty to predict RNA folding. In a general case, however, the first 3 nucleotides in the above example (<code>CGG</code>) could base-pair any <code>CCG</code> bases ahead of them, at any distance from them. Because of \"wobble\" GoU pairings, they could even match with any <code>UCG</code> or <code>CUG</code> at any distance ahead of them. For RNA molecules that are hundreds or thousands of nucleotides long, this quickly becomes a serious combinatorial problem.</p>\n<p>Now, let's talk about mutations. Generally speaking, a mutation in any RNA that codes for a protein is either bad, or in the best case neutral. The same is true for a single mutation in structured RNAs. Let's say we mutate the third base from G to C (marked by an asterisk below).</p>\n<pre><code>    ___\n   /   \\\n   \\   /\n    U-\n    GoU\n    -C\n    -U\n    -U\n   *C C\n    -C\n    C-\n</code></pre>\n<p>The \"ladder\" part of the stem will lose one of its rungs as C can't pair with another C. That makes our molecule less stable and compromises its structural integrity.</p>\n<p>Now, what happens if we have two mutations rather than one? If we are talking about coding RNAs, that's again pretty much guaranteed to be bad. For structured RNA, however, the second mutation could be bad, or it could fix the effect of the first mutation. This is known as correlated variation, or co-variation. Let's say that our second mutation is in base #18, where we go from C to a G. Those two mutations are marked by asterisks below.</p>\n<pre><code>    ___\n   /   \\\n   \\   /\n    U-\n    GoU\n    -C\n    -U\n    -U\n   *C-*\n    -C\n    C-\n</code></pre>\n<p>The co-variation means that mutations in one RNA nucleotide will make no difference if compensatory mutations are made in the other nucleotide. To generalize that further, I have mutated the first 8 and the last 8 residues of our hairpin-forming sequence so that it is nothing like our starting RNA, yet we end up with an identical RNA structure.</p>\n<pre><code>&gt;Original_RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n&gt;Mutated_RNA_sequence\nUACGCCUGGCGACAGGCGUA\n\n    ___        ___\n   /   \\      /   \\\n   \\   /      \\   /\n    U-        -C\n    GoU        U-\n    -C        C-\n    -U        C-\n    -U        -C\n    -C        C-\n    -C        -U\n    C-        U-\n</code></pre>\n<p>Co-variated mutations may result in a small loss of stability, say when a G-C pair is replaced by an A-U pair. That's because there are 3 hydrogen bonds in G-C (3 molecular staples) versus 2 hydrogen bonds in A-U. However, replacing any base pair with any other will not affect the overall molecular architecture.</p>\n<p>Co-variated mutation patterns can't be observed from a single sequence. We need at least two sequences as I have shown above, but ideally we need an alignment of many RNA sequences. These are usually equivalent RNA molecules from different species. When co-variation patterns are observed in large multiple sequence alignments (MSAs), it becomes very likely that the two nucleotides showing these correlated mutations are base-paired to each other, meaning their C1 atoms are ~10.5 angstroms from each other in 3D space.</p>\n<p>Here is an example of co-variation that is found in structured RNAs next to ribosomal proteins of bacterial species.</p>\n<p><img src=\"https://i.ibb.co/gb6KmQJX/co-variance.png\" alt=\"Co-variation\"></p>\n<p>The two highlighted columns contain nucleotides that form base pairs with each other. Whenever there is a <code>C</code> on the left, the matching column on the right will have a <code>G</code>. If Cs on the left are mutated into Us, Gs on the right will be mutated into As so that a base pair can still be made between them. This is a prediction how that RNA above folds:</p>\n<p><img src=\"https://i.ibb.co/vC7T6TMR/RNA-folding.gif\" alt=\"RNA folding animation\"></p>\n<p>To sum: base-pairing patterns can be figured out even from a single sequence when dealing with relatively short RNAs. Yet MSAs are much more informative when it comes to detecting base pairs in larger RNAs, where some of the double-stranded patterns are created by long-range interactions.</p>\n<p>A couple of random thoughts that I didn't share so far. Not only is it double-helical RNA easier to predict, but it is more energetically stable as hydrogen bonds act like molecular staples that hold parts of the molecule together. In thermodynamics parlance, double-helical RNA has lower energy. That means  that a more energetically stable molecule is more likely to be the correctly folded structure. A problem with that is that many predicted RNA folds will have similar yet non-identical structures, and because of the inaccuracy of our energy calculations it is not guaranteed that what we calculate as the lowest energy will actually be the correct fold. Proper energy calculations can be done but are very time-consuming.</p>\n<p>Two challenges still remain: 1) predicting RNA folding in single-stranded regions; 2) predicting relative orientations of multiple double-helical RNA regions. There isn't much that can be done about the first problem when it comes to long single-helical stretches. Luckily, most of the targets in this competition are structured RNA molecules, and they don't have long single-stranded regions. Some aspects of the second problem can be solved by correctly predicting long-distance base-pairing patterns. The rest of it we still need to learn.</p>",
  "messages": [
    {
      "id": "3151773",
      "postDate": "03/17/2025 04:51:10",
      "content": "<p>This is a continuation of <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568445\" target=\"_blank\"><strong>an earlier post</strong></a>.</p>\n<p>First, let's understand what is it exactly that we are predicting. We will look at a G-C base pair:</p>\n<p><img src=\"https://i.ibb.co/rWd5j6b/g-c-pair.png\" alt=\"G-C pairing\"></p>\n<p>Imagine you are Superman and your right eye is shooting a laser through a thick, wavy line connected to the N atom of the C nucleotide (right next to 54o). Likewise, your left eye is shooting a laser through a thick, wavy line connected to the N atom of the G nucleotide (also next to 54o). Those two laser trajectories would become side rails of a ladder, and a G-C base pair laying flat on the screen surface would be a rung connecting those two rails. If we stacked other base pairs on top of the screen towards our face, and some under the screen away from us, we'd get a complete ladder that schematically represents a double-helical RNA. The same would be true for DNA, except that opposite A we'd have a T (Thymine) rather than U. To complete this analogy we'd have to twist that ladder to resemble a spiral staircase.</p>\n<p>Earlier we ignored the phosphate and sugar components of nucleotides, but hopefully their role is now obvious: they form the side rails and connect RNA bases vertically. Those spots where each base connects to a sugar molecule - where we shot the lasers - are C1 atoms for which we are trying to predict 3D coordinates. Let's add some facts: 1) individual bases are rigid; 2) they lay flat in the plane; 3) hydrogen bonds are always of the same length (with small angle deviations). Knowing that two RNA bases form a base pair and given near-fixed covalent bond lengths and hydrogen bond distances, it is fairly straightforward to predict C1 positions in such a base pair. A take-home message: it is much easier to predict relative positions of C1 atoms for the two bases that form a base pair (where RNA is double-stranded) than a random two bases. Any time we correctly predict that two RNA bases are paired to each other, we create a distance restraint that defines relative distances of their C1 atoms. For example, if you look into PDB structure 1SCL (it has been discussed several times), you will see that for RNA bases that pair each other the average C1-C1 distance is 10.46 angstroms (stdev 0.27). That is in addition to a well-defined rise between bases - a distance between the rungs of the ladder, if you will. For the same structure as above, average C1-C1 distance between consecutive bases is 5.49 (stdev 0.61).  Therefore, correctly predicting base-pairing patterns will make it easier to predict RNA folding in double-helical regions.</p>\n<p>Hopefully I convinced you in the paragraph above how important it is to correctly predict base pairings. Let's look again at how base-pairing happens in a simple RNA molecule.</p>\n<pre><code>&gt;RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n&gt;RNA_secondary_structure\n((((((((....))))))))\n</code></pre>\n<p>This molecule forms a hairpin, which is schematically shown below. Hydrogen bonds in canonical bases are shown as horizontal lines connecting the letters, while a non-canonical base pairing is shown as <code>GoU</code>.</p>\n<pre><code>    ___\n   /   \\\n   \\   /\n    U-\n    GoU\n    -C\n    -U\n    -U\n    -C\n    -C\n    C-\n</code></pre>\n<p>If all RNA molecules were this short, or if we had to look only 20 bases ahead to find base pairing partners, there would be no difficulty to predict RNA folding. In a general case, however, the first 3 nucleotides in the above example (<code>CGG</code>) could base-pair any <code>CCG</code> bases ahead of them, at any distance from them. Because of \"wobble\" GoU pairings, they could even match with any <code>UCG</code> or <code>CUG</code> at any distance ahead of them. For RNA molecules that are hundreds or thousands of nucleotides long, this quickly becomes a serious combinatorial problem.</p>\n<p>Now, let's talk about mutations. Generally speaking, a mutation in any RNA that codes for a protein is either bad, or in the best case neutral. The same is true for a single mutation in structured RNAs. Let's say we mutate the third base from G to C (marked by an asterisk below).</p>\n<pre><code>    ___\n   /   \\\n   \\   /\n    U-\n    GoU\n    -C\n    -U\n    -U\n   *C C\n    -C\n    C-\n</code></pre>\n<p>The \"ladder\" part of the stem will lose one of its rungs as C can't pair with another C. That makes our molecule less stable and compromises its structural integrity.</p>\n<p>Now, what happens if we have two mutations rather than one? If we are talking about coding RNAs, that's again pretty much guaranteed to be bad. For structured RNA, however, the second mutation could be bad, or it could fix the effect of the first mutation. This is known as correlated variation, or co-variation. Let's say that our second mutation is in base #18, where we go from C to a G. Those two mutations are marked by asterisks below.</p>\n<pre><code>    ___\n   /   \\\n   \\   /\n    U-\n    GoU\n    -C\n    -U\n    -U\n   *C-*\n    -C\n    C-\n</code></pre>\n<p>The co-variation means that mutations in one RNA nucleotide will make no difference if compensatory mutations are made in the other nucleotide. To generalize that further, I have mutated the first 8 and the last 8 residues of our hairpin-forming sequence so that it is nothing like our starting RNA, yet we end up with an identical RNA structure.</p>\n<pre><code>&gt;Original_RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n&gt;Mutated_RNA_sequence\nUACGCCUGGCGACAGGCGUA\n\n    ___        ___\n   /   \\      /   \\\n   \\   /      \\   /\n    U-        -C\n    GoU        U-\n    -C        C-\n    -U        C-\n    -U        -C\n    -C        C-\n    -C        -U\n    C-        U-\n</code></pre>\n<p>Co-variated mutations may result in a small loss of stability, say when a G-C pair is replaced by an A-U pair. That's because there are 3 hydrogen bonds in G-C (3 molecular staples) versus 2 hydrogen bonds in A-U. However, replacing any base pair with any other will not affect the overall molecular architecture.</p>\n<p>Co-variated mutation patterns can't be observed from a single sequence. We need at least two sequences as I have shown above, but ideally we need an alignment of many RNA sequences. These are usually equivalent RNA molecules from different species. When co-variation patterns are observed in large multiple sequence alignments (MSAs), it becomes very likely that the two nucleotides showing these correlated mutations are base-paired to each other, meaning their C1 atoms are ~10.5 angstroms from each other in 3D space.</p>\n<p>Here is an example of co-variation that is found in structured RNAs next to ribosomal proteins of bacterial species.</p>\n<p><img src=\"https://i.ibb.co/gb6KmQJX/co-variance.png\" alt=\"Co-variation\"></p>\n<p>The two highlighted columns contain nucleotides that form base pairs with each other. Whenever there is a <code>C</code> on the left, the matching column on the right will have a <code>G</code>. If Cs on the left are mutated into Us, Gs on the right will be mutated into As so that a base pair can still be made between them. This is a prediction how that RNA above folds:</p>\n<p><img src=\"https://i.ibb.co/vC7T6TMR/RNA-folding.gif\" alt=\"RNA folding animation\"></p>\n<p>To sum: base-pairing patterns can be figured out even from a single sequence when dealing with relatively short RNAs. Yet MSAs are much more informative when it comes to detecting base pairs in larger RNAs, where some of the double-stranded patterns are created by long-range interactions.</p>\n<p>A couple of random thoughts that I didn't share so far. Not only is it double-helical RNA easier to predict, but it is more energetically stable as hydrogen bonds act like molecular staples that hold parts of the molecule together. In thermodynamics parlance, double-helical RNA has lower energy. That means  that a more energetically stable molecule is more likely to be the correctly folded structure. A problem with that is that many predicted RNA folds will have similar yet non-identical structures, and because of the inaccuracy of our energy calculations it is not guaranteed that what we calculate as the lowest energy will actually be the correct fold. Proper energy calculations can be done but are very time-consuming.</p>\n<p>Two challenges still remain: 1) predicting RNA folding in single-stranded regions; 2) predicting relative orientations of multiple double-helical RNA regions. There isn't much that can be done about the first problem when it comes to long single-helical stretches. Luckily, most of the targets in this competition are structured RNA molecules, and they don't have long single-stranded regions. Some aspects of the second problem can be solved by correctly predicting long-distance base-pairing patterns. The rest of it we still need to learn.</p>",
      "rawMarkdown": "This is a continuation of [**an earlier post**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568445).\n\nFirst, let's understand what is it exactly that we are predicting. We will look at a G-C base pair:\n\n![G-C pairing](https://i.ibb.co/rWd5j6b/g-c-pair.png)\n\nImagine you are Superman and your right eye is shooting a laser through a thick, wavy line connected to the N atom of the C nucleotide (right next to 54o). Likewise, your left eye is shooting a laser through a thick, wavy line connected to the N atom of the G nucleotide (also next to 54o). Those two laser trajectories would become side rails of a ladder, and a G-C base pair laying flat on the screen surface would be a rung connecting those two rails. If we stacked other base pairs on top of the screen towards our face, and some under the screen away from us, we'd get a complete ladder that schematically represents a double-helical RNA. The same would be true for DNA, except that opposite A we'd have a T (Thymine) rather than U. To complete this analogy we'd have to twist that ladder to resemble a spiral staircase.\n\nEarlier we ignored the phosphate and sugar components of nucleotides, but hopefully their role is now obvious: they form the side rails and connect RNA bases vertically. Those spots where each base connects to a sugar molecule - where we shot the lasers - are C1 atoms for which we are trying to predict 3D coordinates. Let's add some facts: 1) individual bases are rigid; 2) they lay flat in the plane; 3) hydrogen bonds are always of the same length (with small angle deviations). Knowing that two RNA bases form a base pair and given near-fixed covalent bond lengths and hydrogen bond distances, it is fairly straightforward to predict C1 positions in such a base pair. A take-home message: it is much easier to predict relative positions of C1 atoms for the two bases that form a base pair (where RNA is double-stranded) than a random two bases. Any time we correctly predict that two RNA bases are paired to each other, we create a distance restraint that defines relative distances of their C1 atoms. For example, if you look into PDB structure 1SCL (it has been discussed several times), you will see that for RNA bases that pair each other the average C1-C1 distance is 10.46 angstroms (stdev 0.27). That is in addition to a well-defined rise between bases - a distance between the rungs of the ladder, if you will. For the same structure as above, average C1-C1 distance between consecutive bases is 5.49 (stdev 0.61).  Therefore, correctly predicting base-pairing patterns will make it easier to predict RNA folding in double-helical regions.\n\nHopefully I convinced you in the paragraph above how important it is to correctly predict base pairings. Let's look again at how base-pairing happens in a simple RNA molecule.\n\n```\n>RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n>RNA_secondary_structure\n((((((((....))))))))\n```\n\nThis molecule forms a hairpin, which is schematically shown below. Hydrogen bonds in canonical bases are shown as horizontal lines connecting the letters, while a non-canonical base pairing is shown as `GoU`.\n\n```\n    ___\n   /   \\\n   \\   /\n    U-A\n    GoU\n    G-C\n    A-U\n    A-U\n    G-C\n    G-C\n    C-G\n```\n\nIf all RNA molecules were this short, or if we had to look only 20 bases ahead to find base pairing partners, there would be no difficulty to predict RNA folding. In a general case, however, the first 3 nucleotides in the above example (`CGG`) could base-pair any `CCG` bases ahead of them, at any distance from them. Because of \"wobble\" GoU pairings, they could even match with any `UCG` or `CUG` at any distance ahead of them. For RNA molecules that are hundreds or thousands of nucleotides long, this quickly becomes a serious combinatorial problem.\n\nNow, let's talk about mutations. Generally speaking, a mutation in any RNA that codes for a protein is either bad, or in the best case neutral. The same is true for a single mutation in structured RNAs. Let's say we mutate the third base from G to C (marked by an asterisk below).\n\n```\n    ___\n   /   \\\n   \\   /\n    U-A\n    GoU\n    G-C\n    A-U\n    A-U\n   *C C\n    G-C\n    C-G\n```\n\nThe \"ladder\" part of the stem will lose one of its rungs as C can't pair with another C. That makes our molecule less stable and compromises its structural integrity.\n\nNow, what happens if we have two mutations rather than one? If we are talking about coding RNAs, that's again pretty much guaranteed to be bad. For structured RNA, however, the second mutation could be bad, or it could fix the effect of the first mutation. This is known as correlated variation, or co-variation. Let's say that our second mutation is in base #18, where we go from C to a G. Those two mutations are marked by asterisks below.\n\n```\n    ___\n   /   \\\n   \\   /\n    U-A\n    GoU\n    G-C\n    A-U\n    A-U\n   *C-G*\n    G-C\n    C-G\n```\n\nThe co-variation means that mutations in one RNA nucleotide will make no difference if compensatory mutations are made in the other nucleotide. To generalize that further, I have mutated the first 8 and the last 8 residues of our hairpin-forming sequence so that it is nothing like our starting RNA, yet we end up with an identical RNA structure.\n\n```\n>Original_RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n>Mutated_RNA_sequence\nUACGCCUGGCGACAGGCGUA\n\n    ___        ___\n   /   \\      /   \\\n   \\   /      \\   /\n    U-A        G-C\n    GoU        U-A\n    G-C        C-G\n    A-U        C-G\n    A-U        G-C\n    G-C        C-G\n    G-C        A-U\n    C-G        U-A\n```\n\nCo-variated mutations may result in a small loss of stability, say when a G-C pair is replaced by an A-U pair. That's because there are 3 hydrogen bonds in G-C (3 molecular staples) versus 2 hydrogen bonds in A-U. However, replacing any base pair with any other will not affect the overall molecular architecture.\n\nCo-variated mutation patterns can't be observed from a single sequence. We need at least two sequences as I have shown above, but ideally we need an alignment of many RNA sequences. These are usually equivalent RNA molecules from different species. When co-variation patterns are observed in large multiple sequence alignments (MSAs), it becomes very likely that the two nucleotides showing these correlated mutations are base-paired to each other, meaning their C1 atoms are ~10.5 angstroms from each other in 3D space.\n\nHere is an example of co-variation that is found in structured RNAs next to ribosomal proteins of bacterial species.\n\n![Co-variation](https://i.ibb.co/gb6KmQJX/co-variance.png)\n\nThe two highlighted columns contain nucleotides that form base pairs with each other. Whenever there is a `C` on the left, the matching column on the right will have a `G`. If Cs on the left are mutated into Us, Gs on the right will be mutated into As so that a base pair can still be made between them. This is a prediction how that RNA above folds:\n\n![RNA folding animation](https://i.ibb.co/vC7T6TMR/RNA-folding.gif)\n\nTo sum: base-pairing patterns can be figured out even from a single sequence when dealing with relatively short RNAs. Yet MSAs are much more informative when it comes to detecting base pairs in larger RNAs, where some of the double-stranded patterns are created by long-range interactions.\n\nA couple of random thoughts that I didn't share so far. Not only is it double-helical RNA easier to predict, but it is more energetically stable as hydrogen bonds act like molecular staples that hold parts of the molecule together. In thermodynamics parlance, double-helical RNA has lower energy. That means  that a more energetically stable molecule is more likely to be the correctly folded structure. A problem with that is that many predicted RNA folds will have similar yet non-identical structures, and because of the inaccuracy of our energy calculations it is not guaranteed that what we calculate as the lowest energy will actually be the correct fold. Proper energy calculations can be done but are very time-consuming.\n\nTwo challenges still remain: 1) predicting RNA folding in single-stranded regions; 2) predicting relative orientations of multiple double-helical RNA regions. There isn't much that can be done about the first problem when it comes to long single-helical stretches. Luckily, most of the targets in this competition are structured RNA molecules, and they don't have long single-stranded regions. Some aspects of the second problem can be solved by correctly predicting long-distance base-pairing patterns. The rest of it we still need to learn.",
      "votes": null
    },
    {
      "id": "3152042",
      "postDate": "03/17/2025 11:02:24",
      "content": "<p>Hi, thank you for the explanation it was very helpful for someone like me with limited background in biochemistry. I’d also like to ask are there any intuitive principles or ideas for estimating the probability that a given base will pair with another in longer RNA sequences?</p>",
      "rawMarkdown": "Hi, thank you for the explanation it was very helpful for someone like me with limited background in biochemistry. I’d also like to ask are there any intuitive principles or ideas for estimating the probability that a given base will pair with another in longer RNA sequences?",
      "votes": null
    },
    {
      "id": "3152348",
      "postDate": "03/17/2025 18:10:20",
      "content": "<p>Several recently developed models make distograms, which predict distances for all possible pairs of nucleotides.</p>\n<ul>\n<li><a href=\"https://www.nature.com/articles/s41467-025-56261-7\" target=\"_blank\">https://www.nature.com/articles/s41467-025-56261-7</a></li>\n<li><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.28.625345v1.full\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.28.625345v1.full</a></li>\n</ul>\n<p>In terms of general principles this is similar to what AlphaFold does for proteins. </p>",
      "rawMarkdown": "Several recently developed models make distograms, which predict distances for all possible pairs of nucleotides.\n\n- https://www.nature.com/articles/s41467-025-56261-7\n- https://www.biorxiv.org/content/10.1101/2024.11.28.625345v1.full\n\nIn terms of general principles this is similar to what AlphaFold does for proteins.",
      "votes": null
    },
    {
      "id": "3152442",
      "postDate": "03/17/2025 19:51:40",
      "content": "<p>these 2 parts are amazing however I have already been doing them from them from chatgpt recommend. I add Secondary structure features using ViennaRNA and MSAs feature but my score not going up as much. one more thing, I noticed how complex Alphafold and other model they worked amazingly. however, whenever I try to increase the complexity of my model the score drop. I'm kinda stuck right now. I tried NN, CNN, GNN and even transformer. no good result, is the problem come from the data or am I just a noob? XD </p>\n<p>do you have any recommendations for me ? thank in advance</p>",
      "rawMarkdown": "these 2 parts are amazing however I have already been doing them from them from chatgpt recommend. I add Secondary structure features using ViennaRNA and MSAs feature but my score not going up as much. one more thing, I noticed how complex Alphafold and other model they worked amazingly. however, whenever I try to increase the complexity of my model the score drop. I'm kinda stuck right now. I tried NN, CNN, GNN and even transformer. no good result, is the problem come from the data or am I just a noob? XD \n\ndo you have any recommendations for me ? thank in advance",
      "votes": null
    },
    {
      "id": "3152543",
      "postDate": "03/17/2025 23:37:17",
      "content": "<p>I haven't started modeling yet - hopefully in a couple of days - so it is difficult to answer your question. Modeling structured RNA is not easy, and it is not by accident that humans still do a better job than automated procedures. Knowing that principles of co-variation and distance restraints that I mentioned here have been successfully applied by AlphaFold, it is likely that we have to do the same. Some heuristics will be needed, and that is presumably what we are supposed to discover in the next 2.5 months.</p>",
      "rawMarkdown": "I haven't started modeling yet - hopefully in a couple of days - so it is difficult to answer your question. Modeling structured RNA is not easy, and it is not by accident that humans still do a better job than automated procedures. Knowing that principles of co-variation and distance restraints that I mentioned here have been successfully applied by AlphaFold, it is likely that we have to do the same. Some heuristics will be needed, and that is presumably what we are supposed to discover in the next 2.5 months.",
      "votes": null
    },
    {
      "id": "3160823",
      "postDate": "03/27/2025 07:22:00",
      "content": "<p>Cool, more cases and explain, thank you!</p>",
      "rawMarkdown": "Cool, more cases and explain, thank you!",
      "votes": null
    },
    {
      "id": "3182487",
      "postDate": "04/19/2025 12:09:16",
      "content": "<p>Thanks for both the posts!</p>",
      "rawMarkdown": "Thanks for both the posts!",
      "votes": null
    },
    {
      "id": "3201684",
      "postDate": "05/14/2025 08:50:41",
      "content": "<p>thanks for sharing, very helpfull and interesting. Appreciate it.</p>",
      "rawMarkdown": "thanks for sharing, very helpfull and interesting. Appreciate it.",
      "votes": null
    },
    {
      "id": "3207691",
      "postDate": "05/23/2025 06:16:01",
      "content": "<p>Very useful work, thanks a lot!</p>",
      "rawMarkdown": "Very useful work, thanks a lot!",
      "votes": null
    },
    {
      "id": "3208956",
      "postDate": "05/25/2025 02:56:44",
      "content": "<p>awesome …hardwork</p>",
      "rawMarkdown": "awesome ...hardwork",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3152042,
      "author_name": "shirshakghatak",
      "author_url": "",
      "post_date": "03/17/2025 11:02:24",
      "content": "<p>Hi, thank you for the explanation it was very helpful for someone like me with limited background in biochemistry. I’d also like to ask are there any intuitive principles or ideas for estimating the probability that a given base will pair with another in longer RNA sequences?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3152348,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/17/2025 18:10:20",
          "content": "<p>Several recently developed models make distograms, which predict distances for all possible pairs of nucleotides.</p>\n<ul>\n<li><a href=\"https://www.nature.com/articles/s41467-025-56261-7\" target=\"_blank\">https://www.nature.com/articles/s41467-025-56261-7</a></li>\n<li><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.28.625345v1.full\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.28.625345v1.full</a></li>\n</ul>\n<p>In terms of general principles this is similar to what AlphaFold does for proteins. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3152442,
      "author_name": "theonetheonly",
      "author_url": "",
      "post_date": "03/17/2025 19:51:40",
      "content": "<p>these 2 parts are amazing however I have already been doing them from them from chatgpt recommend. I add Secondary structure features using ViennaRNA and MSAs feature but my score not going up as much. one more thing, I noticed how complex Alphafold and other model they worked amazingly. however, whenever I try to increase the complexity of my model the score drop. I'm kinda stuck right now. I tried NN, CNN, GNN and even transformer. no good result, is the problem come from the data or am I just a noob? XD </p>\n<p>do you have any recommendations for me ? thank in advance</p>",
      "votes": null,
      "replies": [
        {
          "id": 3152543,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/17/2025 23:37:17",
          "content": "<p>I haven't started modeling yet - hopefully in a couple of days - so it is difficult to answer your question. Modeling structured RNA is not easy, and it is not by accident that humans still do a better job than automated procedures. Knowing that principles of co-variation and distance restraints that I mentioned here have been successfully applied by AlphaFold, it is likely that we have to do the same. Some heuristics will be needed, and that is presumably what we are supposed to discover in the next 2.5 months.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3160823,
      "author_name": "dedquoc",
      "author_url": "",
      "post_date": "03/27/2025 07:22:00",
      "content": "<p>Cool, more cases and explain, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3182487,
      "author_name": "vineetkreddy",
      "author_url": "",
      "post_date": "04/19/2025 12:09:16",
      "content": "<p>Thanks for both the posts!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3201684,
      "author_name": "xaocreal",
      "author_url": "",
      "post_date": "05/14/2025 08:50:41",
      "content": "<p>thanks for sharing, very helpfull and interesting. Appreciate it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3207691,
      "author_name": "",
      "author_url": "",
      "post_date": "05/23/2025 06:16:01",
      "content": "<p>Very useful work, thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3208956,
      "author_name": "sonmou",
      "author_url": "",
      "post_date": "05/25/2025 02:56:44",
      "content": "<p>awesome …hardwork</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3151773": "This is a continuation of [**an earlier post**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568445).\n\nFirst, let's understand what is it exactly that we are predicting. We will look at a G-C base pair:\n\n![G-C pairing](https://i.ibb.co/rWd5j6b/g-c-pair.png)\n\nImagine you are Superman and your right eye is shooting a laser through a thick, wavy line connected to the N atom of the C nucleotide (right next to 54o). Likewise, your left eye is shooting a laser through a thick, wavy line connected to the N atom of the G nucleotide (also next to 54o). Those two laser trajectories would become side rails of a ladder, and a G-C base pair laying flat on the screen surface would be a rung connecting those two rails. If we stacked other base pairs on top of the screen towards our face, and some under the screen away from us, we'd get a complete ladder that schematically represents a double-helical RNA. The same would be true for DNA, except that opposite A we'd have a T (Thymine) rather than U. To complete this analogy we'd have to twist that ladder to resemble a spiral staircase.\n\nEarlier we ignored the phosphate and sugar components of nucleotides, but hopefully their role is now obvious: they form the side rails and connect RNA bases vertically. Those spots where each base connects to a sugar molecule - where we shot the lasers - are C1 atoms for which we are trying to predict 3D coordinates. Let's add some facts: 1) individual bases are rigid; 2) they lay flat in the plane; 3) hydrogen bonds are always of the same length (with small angle deviations). Knowing that two RNA bases form a base pair and given near-fixed covalent bond lengths and hydrogen bond distances, it is fairly straightforward to predict C1 positions in such a base pair. A take-home message: it is much easier to predict relative positions of C1 atoms for the two bases that form a base pair (where RNA is double-stranded) than a random two bases. Any time we correctly predict that two RNA bases are paired to each other, we create a distance restraint that defines relative distances of their C1 atoms. For example, if you look into PDB structure 1SCL (it has been discussed several times), you will see that for RNA bases that pair each other the average C1-C1 distance is 10.46 angstroms (stdev 0.27). That is in addition to a well-defined rise between bases - a distance between the rungs of the ladder, if you will. For the same structure as above, average C1-C1 distance between consecutive bases is 5.49 (stdev 0.61).  Therefore, correctly predicting base-pairing patterns will make it easier to predict RNA folding in double-helical regions.\n\nHopefully I convinced you in the paragraph above how important it is to correctly predict base pairings. Let's look again at how base-pairing happens in a simple RNA molecule.\n\n```\n>RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n>RNA_secondary_structure\n((((((((....))))))))\n```\n\nThis molecule forms a hairpin, which is schematically shown below. Hydrogen bonds in canonical bases are shown as horizontal lines connecting the letters, while a non-canonical base pairing is shown as `GoU`.\n\n```\n    ___\n   /   \\\n   \\   /\n    U-A\n    GoU\n    G-C\n    A-U\n    A-U\n    G-C\n    G-C\n    C-G\n```\n\nIf all RNA molecules were this short, or if we had to look only 20 bases ahead to find base pairing partners, there would be no difficulty to predict RNA folding. In a general case, however, the first 3 nucleotides in the above example (`CGG`) could base-pair any `CCG` bases ahead of them, at any distance from them. Because of \"wobble\" GoU pairings, they could even match with any `UCG` or `CUG` at any distance ahead of them. For RNA molecules that are hundreds or thousands of nucleotides long, this quickly becomes a serious combinatorial problem.\n\nNow, let's talk about mutations. Generally speaking, a mutation in any RNA that codes for a protein is either bad, or in the best case neutral. The same is true for a single mutation in structured RNAs. Let's say we mutate the third base from G to C (marked by an asterisk below).\n\n```\n    ___\n   /   \\\n   \\   /\n    U-A\n    GoU\n    G-C\n    A-U\n    A-U\n   *C C\n    G-C\n    C-G\n```\n\nThe \"ladder\" part of the stem will lose one of its rungs as C can't pair with another C. That makes our molecule less stable and compromises its structural integrity.\n\nNow, what happens if we have two mutations rather than one? If we are talking about coding RNAs, that's again pretty much guaranteed to be bad. For structured RNA, however, the second mutation could be bad, or it could fix the effect of the first mutation. This is known as correlated variation, or co-variation. Let's say that our second mutation is in base #18, where we go from C to a G. Those two mutations are marked by asterisks below.\n\n```\n    ___\n   /   \\\n   \\   /\n    U-A\n    GoU\n    G-C\n    A-U\n    A-U\n   *C-G*\n    G-C\n    C-G\n```\n\nThe co-variation means that mutations in one RNA nucleotide will make no difference if compensatory mutations are made in the other nucleotide. To generalize that further, I have mutated the first 8 and the last 8 residues of our hairpin-forming sequence so that it is nothing like our starting RNA, yet we end up with an identical RNA structure.\n\n```\n>Original_RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n>Mutated_RNA_sequence\nUACGCCUGGCGACAGGCGUA\n\n    ___        ___\n   /   \\      /   \\\n   \\   /      \\   /\n    U-A        G-C\n    GoU        U-A\n    G-C        C-G\n    A-U        C-G\n    A-U        G-C\n    G-C        C-G\n    G-C        A-U\n    C-G        U-A\n```\n\nCo-variated mutations may result in a small loss of stability, say when a G-C pair is replaced by an A-U pair. That's because there are 3 hydrogen bonds in G-C (3 molecular staples) versus 2 hydrogen bonds in A-U. However, replacing any base pair with any other will not affect the overall molecular architecture.\n\nCo-variated mutation patterns can't be observed from a single sequence. We need at least two sequences as I have shown above, but ideally we need an alignment of many RNA sequences. These are usually equivalent RNA molecules from different species. When co-variation patterns are observed in large multiple sequence alignments (MSAs), it becomes very likely that the two nucleotides showing these correlated mutations are base-paired to each other, meaning their C1 atoms are ~10.5 angstroms from each other in 3D space.\n\nHere is an example of co-variation that is found in structured RNAs next to ribosomal proteins of bacterial species.\n\n![Co-variation](https://i.ibb.co/gb6KmQJX/co-variance.png)\n\nThe two highlighted columns contain nucleotides that form base pairs with each other. Whenever there is a `C` on the left, the matching column on the right will have a `G`. If Cs on the left are mutated into Us, Gs on the right will be mutated into As so that a base pair can still be made between them. This is a prediction how that RNA above folds:\n\n![RNA folding animation](https://i.ibb.co/vC7T6TMR/RNA-folding.gif)\n\nTo sum: base-pairing patterns can be figured out even from a single sequence when dealing with relatively short RNAs. Yet MSAs are much more informative when it comes to detecting base pairs in larger RNAs, where some of the double-stranded patterns are created by long-range interactions.\n\nA couple of random thoughts that I didn't share so far. Not only is it double-helical RNA easier to predict, but it is more energetically stable as hydrogen bonds act like molecular staples that hold parts of the molecule together. In thermodynamics parlance, double-helical RNA has lower energy. That means  that a more energetically stable molecule is more likely to be the correctly folded structure. A problem with that is that many predicted RNA folds will have similar yet non-identical structures, and because of the inaccuracy of our energy calculations it is not guaranteed that what we calculate as the lowest energy will actually be the correct fold. Proper energy calculations can be done but are very time-consuming.\n\nTwo challenges still remain: 1) predicting RNA folding in single-stranded regions; 2) predicting relative orientations of multiple double-helical RNA regions. There isn't much that can be done about the first problem when it comes to long single-helical stretches. Luckily, most of the targets in this competition are structured RNA molecules, and they don't have long single-stranded regions. Some aspects of the second problem can be solved by correctly predicting long-distance base-pairing patterns. The rest of it we still need to learn.",
    "3152042": "Hi, thank you for the explanation it was very helpful for someone like me with limited background in biochemistry. I’d also like to ask are there any intuitive principles or ideas for estimating the probability that a given base will pair with another in longer RNA sequences?",
    "3152348": "Several recently developed models make distograms, which predict distances for all possible pairs of nucleotides.\n\n- https://www.nature.com/articles/s41467-025-56261-7\n- https://www.biorxiv.org/content/10.1101/2024.11.28.625345v1.full\n\nIn terms of general principles this is similar to what AlphaFold does for proteins.",
    "3152442": "these 2 parts are amazing however I have already been doing them from them from chatgpt recommend. I add Secondary structure features using ViennaRNA and MSAs feature but my score not going up as much. one more thing, I noticed how complex Alphafold and other model they worked amazingly. however, whenever I try to increase the complexity of my model the score drop. I'm kinda stuck right now. I tried NN, CNN, GNN and even transformer. no good result, is the problem come from the data or am I just a noob? XD \n\ndo you have any recommendations for me ? thank in advance",
    "3152543": "I haven't started modeling yet - hopefully in a couple of days - so it is difficult to answer your question. Modeling structured RNA is not easy, and it is not by accident that humans still do a better job than automated procedures. Knowing that principles of co-variation and distance restraints that I mentioned here have been successfully applied by AlphaFold, it is likely that we have to do the same. Some heuristics will be needed, and that is presumably what we are supposed to discover in the next 2.5 months.",
    "3160823": "Cool, more cases and explain, thank you!",
    "3182487": "Thanks for both the posts!",
    "3201684": "thanks for sharing, very helpfull and interesting. Appreciate it.",
    "3207691": "Very useful work, thanks a lot!",
    "3208956": "awesome ...hardwork"
  },
  "source": "meta"
}