{
  "id": 568445,
  "title": "A medium-level introduction to RNA folding, part 1",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568445",
  "author_name": "",
  "post_date": "2025-03-16T00:07:26.685211Z",
  "votes": 80,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I have been following this competition from the start, and it seems that the main problem is in understanding the driving forces for RNA folding. Organizers have given a link to <a href=\"https://www.pnas.org/doi/10.1073/pnas.2112677119\" target=\"_blank\"><strong>a paper</strong></a> that explains RNA folding at a high level, and that paper is definitely worth reading. However, I suspect that not everyone will understand RNA folding that is explained at that level and written with formalism that is required in scientific publications.</p>\n<p>I am sure many intuitively understand the goal of the challenge, which is how a string of RNA nucleotides assumes a certain 3D shape. Yet there are some weak molecular interactions that both simplify and complicate the calculation of that shape. This previous statement will likely remain contradictory until we get to the end of this story.</p>\n<p>We are taught in school from the early days that DNA is double-stranded, while RNA is single-stranded. That is a true statement, but somewhat imprecise for the purposes of this competition. RNA is single-stranded <strong>most of the time</strong>, and especially when we are talking about <strong>coding, non-structured RNA molecules</strong>. RNA molecules in this competition are of the <strong>non-coding, structured kind</strong>, and they are double-stranded more often than not. It may sound like that's complicating our lives, but in reality it will make the predictions easier.</p>\n<p>I will start explaining this by using a spaghetti analogy, as most people like pasta or at least have a general understanding of spaghetti shape(s). When we pull a handful of spaghetti from the box and before we dump them into a pot of boiling water, all the individual strings are of the same shape and length. After we are done cooking all of them still have the same length, but shapes are completely different. Predicting shapes of long, single-stranded RNAs is like predicting spaghetti shapes. One could reasonably predict a shape of a 2-inch / 5-cm spaghetti because it is not very long, and similarly for a single-stranded RNA molecule that is maybe 5-10 nucleotides long. For anything longer than that this becomes an impossible exercise because there are too many degrees of freedom, and there isn't a single shape that is more likely than the others.</p>\n<p>This is where double-stranded nature of structured RNAs will help us. Surely you have noticed that cooked spaghetti strings stick to each other, and sometimes the sticking happens within different parts of the same string. For spaghetti this is completely random, but not so for RNA. Imagine if a string of spaghetti goes up, makes a U-turn, then sticks back to itself in a reproducible way. Then goes to the right, makes another U-turn, and again sticks back to itself. If we understood what governs these sticking rules, we'd have much easier time predicting the final shape. For RNA, this comes down to base-pairing, which is highly reproducible.</p>\n<p>RNA nucleotides (I will interchangeably call them RNA bases) are made of 3 components: a phosphate, a pentose sugar and a nitrogenous base. We will ignore the first two components for a moment and only talk about bases. Those are Adenine (A), Cytosine (C), Guanine (G) and Uracil (U). A canonical base-pairing is A with U and G with C. This is how those pairings look. Note that sugar-base angles are 54 degrees in all cases.</p>\n<p><img src=\"https://i.ibb.co/ZpnGNmxR/canonical-pairing.png\" alt=\"RNA canonical base-pairing\"></p>\n<p>You may ask: why doesn't A pair with G, or A with C, or U with C? Notice first that A and G are larger bases (known as purines) because they have two connected rings, while C and U are single-ring bases (known as pyrimidines). Neither two large nor two small bases opposite each other will work, which means we always have to pair a large base with a small one. In addition to canonical pairings shown above, that still leaves a possibility of A pairing with C and G pairing with U. This is where those dotted lines come into play. They represent hydrogen bonds, and you can see that there are 3 of them in a G-C pair and 2 of them in an A-U pair. A hydrogen bond must have a donor group (either NH or NH2 for nucleic acids) and an acceptor group (either a solo O or a solo N atom). We can't form a hydrogen bond with two donors across from each other, nor with two acceptor groups across from each other. If you mentally replace a U base above with a C, you will notice that in a A-C pair two donors and two acceptors will be lined up across each other, which means that A can't pair with C. Similarly, if you slide a U base up and mentally position it where C is, two acceptors and two donors will be lined up, which means no base-pairing. If you do the same exercise but rotate the U base slightly clockwise and a G base slightly counter-clockwise, you will see that a pattern of donors and acceptors will change, so G can pair with U in what is known as a \"wobble\" pair (see below). It is not as stable as either A-U or G-C, but it is stable enough that it can be made. Note different sugar-base angles for the G-U pair compared to others.</p>\n<p><img src=\"https://i.ibb.co/YBLrGVJm/g-u-pairing.png\" alt=\"RNA wobble base-pairing\"></p>\n<p>If thinking in terms of donors and acceptors is too abstract, you can substitute a donor for a key, and an acceptor for a lock. For a G base the pattern from the top would be <code>lock-key-key</code>, while for a C base it would be <code>key-lock-lock</code>. When you put them together and slide towards each other, locks and keys match. For an A base the pattern from the top would be <code>key-lock</code>, which matches the U’s pattern of <code>lock-key</code>. Normally A-C and G-U patterns are incompatible, but if we slide U up and rotate it slightly it will match the part of G’s pattern, which will enable it to make a “wobble” base pair.</p>\n<p>Understanding molecular reasons for A-U, G-C and G-U base-pairing is not really required. What matters is that it will always be reproducible. Now that we know that our goal is to find a solution that maximizes the number of base pairs, folding RNA will be a breeze. Right? Well, yes and no. For relatively short RNA molecules there will be only so many base-pairing combinations, so folding them is relatively easy. Consider the 20-nucleotide RNA molecule below:</p>\n<pre><code>&gt;RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n&gt;RNA_secondary_structure\n((((((((....))))))))\n</code></pre>\n<p>That molecule forms a stable hairpin structure. Note that for our example there should be only 8 base pairs in a stem rather than 10 as shown below.</p>\n<p><img src=\"https://i.ibb.co/WrXtDcW/RNA-Hairpin.png\" alt=\"RNA hairpin\"></p>\n<p>The meaning of secondary structure representation below the sequence is that each left rounded bracket must pair with a right bracket. Sometimes people use squared (<code>[]</code>), curly (<code>{}</code>) or angle (<code>&lt;&gt;</code>) brackets, but their meaning is the same. Specifically, that bracket notation above means that base #1 pairs with base #20, #2 with #19, #3 with #18 and so forth until #8 pairs with #13. Middle 4 nucleotides are unpaired and form a stem loop. For that particular sequence there is no other base-pairing pattern that would give us 8 pairs in a double-helical conformation. Keep in mind that consecutive bases can't be paired because it is not stereochemically possible - there has to be at least 3-4 bases separating base-pairing partners. Note that in this case the greatest separation between the two bases in the same pair is 18 (#1 with #20), and the shortest separation is 4 (#8 with #13). All of these are relatively short distances, and it would be easy to fold RNA if all base-pairing patterns were similar to this. In many cases, however, the separation between bases that end up in the same pair can be in hundreds. That complicates our optimization problem from a local to a global scale. So even though we know the general \"sticking\" rules for our molecular spaghetti, longer RNA molecules create a large number of combinatorial possibilities. Going back to my statement that <code>weak molecular interactions that both simplify and complicate the calculation of that shape</code>: hydrogen bonds between RNA nucleotides that have relatively similar sequence numbers make the folding easier, but complications arise from the fact that base-pairing works just as well between RNA nucleotides that are very far from each other in terms of sequence numbers.</p>\n<p>I will explain that last statement by using an example of a more complicated base-pairing pattern (borrowed from <a href=\"http://rna.tbi.univie.ac.at/forna/\" target=\"_blank\"><strong>here</strong></a>):</p>\n<pre><code>&gt;RNA_sequence\nCGCUUCAUAUAAUCCUAAUGAUAUGGUUUGGGAGUUUCUACCAAGAGCCUUAAACUCUUGAUUAUGAAGUG\n&gt;RNA_secondary_structure\n(((((((..(((((())))))).((((((.))))))..))))))\n</code></pre>\n<p>This RNA sequence has 71 nucleotides, and it features two stems (double-stranded helical regions) formed by local interactions, and another one formed by long-distance interactions. That may be easier to see in the image below.</p>\n<div>\n    <img src=\"https://i.ibb.co/zHL0tfvL/structured-rna.png\">\n</div>\n<p>The sequence starts from <code>CGC</code> at the bottom and you can follow it clockwise. Each 10th base is labeled. For example, you can see that A at #18 pairs with U at #28, or that U at #13 pairs with A at #33. That would be an example of a local stem, and the same is true for the other one at 2 o’clock. Now, notice that U#10 pairs with A#40, and that U#4 pairs with A#68. Basically, those are nucleotides from the start and the end pairing with each other. This is still “only” a ~60 nucleotides separation because the molecule’s length is 71, but something like this could have happened in a molecule that is 1000+ nucleotides long. These long-range interactions are most difficult to predict as the number of potential pairing combinations increases dramatically with sequence length.</p>\n<p>In the <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568633\" target=\"_blank\"><strong>second part</strong></a> I explain the principle of co-variation and why MSAs help RNA folding predictions.</p>",
  "messages": [
    {
      "id": "3150812",
      "postDate": "03/16/2025 00:07:26",
      "content": "<p>I have been following this competition from the start, and it seems that the main problem is in understanding the driving forces for RNA folding. Organizers have given a link to <a href=\"https://www.pnas.org/doi/10.1073/pnas.2112677119\" target=\"_blank\"><strong>a paper</strong></a> that explains RNA folding at a high level, and that paper is definitely worth reading. However, I suspect that not everyone will understand RNA folding that is explained at that level and written with formalism that is required in scientific publications.</p>\n<p>I am sure many intuitively understand the goal of the challenge, which is how a string of RNA nucleotides assumes a certain 3D shape. Yet there are some weak molecular interactions that both simplify and complicate the calculation of that shape. This previous statement will likely remain contradictory until we get to the end of this story.</p>\n<p>We are taught in school from the early days that DNA is double-stranded, while RNA is single-stranded. That is a true statement, but somewhat imprecise for the purposes of this competition. RNA is single-stranded <strong>most of the time</strong>, and especially when we are talking about <strong>coding, non-structured RNA molecules</strong>. RNA molecules in this competition are of the <strong>non-coding, structured kind</strong>, and they are double-stranded more often than not. It may sound like that's complicating our lives, but in reality it will make the predictions easier.</p>\n<p>I will start explaining this by using a spaghetti analogy, as most people like pasta or at least have a general understanding of spaghetti shape(s). When we pull a handful of spaghetti from the box and before we dump them into a pot of boiling water, all the individual strings are of the same shape and length. After we are done cooking all of them still have the same length, but shapes are completely different. Predicting shapes of long, single-stranded RNAs is like predicting spaghetti shapes. One could reasonably predict a shape of a 2-inch / 5-cm spaghetti because it is not very long, and similarly for a single-stranded RNA molecule that is maybe 5-10 nucleotides long. For anything longer than that this becomes an impossible exercise because there are too many degrees of freedom, and there isn't a single shape that is more likely than the others.</p>\n<p>This is where double-stranded nature of structured RNAs will help us. Surely you have noticed that cooked spaghetti strings stick to each other, and sometimes the sticking happens within different parts of the same string. For spaghetti this is completely random, but not so for RNA. Imagine if a string of spaghetti goes up, makes a U-turn, then sticks back to itself in a reproducible way. Then goes to the right, makes another U-turn, and again sticks back to itself. If we understood what governs these sticking rules, we'd have much easier time predicting the final shape. For RNA, this comes down to base-pairing, which is highly reproducible.</p>\n<p>RNA nucleotides (I will interchangeably call them RNA bases) are made of 3 components: a phosphate, a pentose sugar and a nitrogenous base. We will ignore the first two components for a moment and only talk about bases. Those are Adenine (A), Cytosine (C), Guanine (G) and Uracil (U). A canonical base-pairing is A with U and G with C. This is how those pairings look. Note that sugar-base angles are 54 degrees in all cases.</p>\n<p><img src=\"https://i.ibb.co/ZpnGNmxR/canonical-pairing.png\" alt=\"RNA canonical base-pairing\"></p>\n<p>You may ask: why doesn't A pair with G, or A with C, or U with C? Notice first that A and G are larger bases (known as purines) because they have two connected rings, while C and U are single-ring bases (known as pyrimidines). Neither two large nor two small bases opposite each other will work, which means we always have to pair a large base with a small one. In addition to canonical pairings shown above, that still leaves a possibility of A pairing with C and G pairing with U. This is where those dotted lines come into play. They represent hydrogen bonds, and you can see that there are 3 of them in a G-C pair and 2 of them in an A-U pair. A hydrogen bond must have a donor group (either NH or NH2 for nucleic acids) and an acceptor group (either a solo O or a solo N atom). We can't form a hydrogen bond with two donors across from each other, nor with two acceptor groups across from each other. If you mentally replace a U base above with a C, you will notice that in a A-C pair two donors and two acceptors will be lined up across each other, which means that A can't pair with C. Similarly, if you slide a U base up and mentally position it where C is, two acceptors and two donors will be lined up, which means no base-pairing. If you do the same exercise but rotate the U base slightly clockwise and a G base slightly counter-clockwise, you will see that a pattern of donors and acceptors will change, so G can pair with U in what is known as a \"wobble\" pair (see below). It is not as stable as either A-U or G-C, but it is stable enough that it can be made. Note different sugar-base angles for the G-U pair compared to others.</p>\n<p><img src=\"https://i.ibb.co/YBLrGVJm/g-u-pairing.png\" alt=\"RNA wobble base-pairing\"></p>\n<p>If thinking in terms of donors and acceptors is too abstract, you can substitute a donor for a key, and an acceptor for a lock. For a G base the pattern from the top would be <code>lock-key-key</code>, while for a C base it would be <code>key-lock-lock</code>. When you put them together and slide towards each other, locks and keys match. For an A base the pattern from the top would be <code>key-lock</code>, which matches the U’s pattern of <code>lock-key</code>. Normally A-C and G-U patterns are incompatible, but if we slide U up and rotate it slightly it will match the part of G’s pattern, which will enable it to make a “wobble” base pair.</p>\n<p>Understanding molecular reasons for A-U, G-C and G-U base-pairing is not really required. What matters is that it will always be reproducible. Now that we know that our goal is to find a solution that maximizes the number of base pairs, folding RNA will be a breeze. Right? Well, yes and no. For relatively short RNA molecules there will be only so many base-pairing combinations, so folding them is relatively easy. Consider the 20-nucleotide RNA molecule below:</p>\n<pre><code>&gt;RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n&gt;RNA_secondary_structure\n((((((((....))))))))\n</code></pre>\n<p>That molecule forms a stable hairpin structure. Note that for our example there should be only 8 base pairs in a stem rather than 10 as shown below.</p>\n<p><img src=\"https://i.ibb.co/WrXtDcW/RNA-Hairpin.png\" alt=\"RNA hairpin\"></p>\n<p>The meaning of secondary structure representation below the sequence is that each left rounded bracket must pair with a right bracket. Sometimes people use squared (<code>[]</code>), curly (<code>{}</code>) or angle (<code>&lt;&gt;</code>) brackets, but their meaning is the same. Specifically, that bracket notation above means that base #1 pairs with base #20, #2 with #19, #3 with #18 and so forth until #8 pairs with #13. Middle 4 nucleotides are unpaired and form a stem loop. For that particular sequence there is no other base-pairing pattern that would give us 8 pairs in a double-helical conformation. Keep in mind that consecutive bases can't be paired because it is not stereochemically possible - there has to be at least 3-4 bases separating base-pairing partners. Note that in this case the greatest separation between the two bases in the same pair is 18 (#1 with #20), and the shortest separation is 4 (#8 with #13). All of these are relatively short distances, and it would be easy to fold RNA if all base-pairing patterns were similar to this. In many cases, however, the separation between bases that end up in the same pair can be in hundreds. That complicates our optimization problem from a local to a global scale. So even though we know the general \"sticking\" rules for our molecular spaghetti, longer RNA molecules create a large number of combinatorial possibilities. Going back to my statement that <code>weak molecular interactions that both simplify and complicate the calculation of that shape</code>: hydrogen bonds between RNA nucleotides that have relatively similar sequence numbers make the folding easier, but complications arise from the fact that base-pairing works just as well between RNA nucleotides that are very far from each other in terms of sequence numbers.</p>\n<p>I will explain that last statement by using an example of a more complicated base-pairing pattern (borrowed from <a href=\"http://rna.tbi.univie.ac.at/forna/\" target=\"_blank\"><strong>here</strong></a>):</p>\n<pre><code>&gt;RNA_sequence\nCGCUUCAUAUAAUCCUAAUGAUAUGGUUUGGGAGUUUCUACCAAGAGCCUUAAACUCUUGAUUAUGAAGUG\n&gt;RNA_secondary_structure\n(((((((..(((((())))))).((((((.))))))..))))))\n</code></pre>\n<p>This RNA sequence has 71 nucleotides, and it features two stems (double-stranded helical regions) formed by local interactions, and another one formed by long-distance interactions. That may be easier to see in the image below.</p>\n<div>\n    <img src=\"https://i.ibb.co/zHL0tfvL/structured-rna.png\">\n</div>\n<p>The sequence starts from <code>CGC</code> at the bottom and you can follow it clockwise. Each 10th base is labeled. For example, you can see that A at #18 pairs with U at #28, or that U at #13 pairs with A at #33. That would be an example of a local stem, and the same is true for the other one at 2 o’clock. Now, notice that U#10 pairs with A#40, and that U#4 pairs with A#68. Basically, those are nucleotides from the start and the end pairing with each other. This is still “only” a ~60 nucleotides separation because the molecule’s length is 71, but something like this could have happened in a molecule that is 1000+ nucleotides long. These long-range interactions are most difficult to predict as the number of potential pairing combinations increases dramatically with sequence length.</p>\n<p>In the <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568633\" target=\"_blank\"><strong>second part</strong></a> I explain the principle of co-variation and why MSAs help RNA folding predictions.</p>",
      "rawMarkdown": "I have been following this competition from the start, and it seems that the main problem is in understanding the driving forces for RNA folding. Organizers have given a link to [**a paper**](https://www.pnas.org/doi/10.1073/pnas.2112677119) that explains RNA folding at a high level, and that paper is definitely worth reading. However, I suspect that not everyone will understand RNA folding that is explained at that level and written with formalism that is required in scientific publications.\n\nI am sure many intuitively understand the goal of the challenge, which is how a string of RNA nucleotides assumes a certain 3D shape. Yet there are some weak molecular interactions that both simplify and complicate the calculation of that shape. This previous statement will likely remain contradictory until we get to the end of this story.\n\nWe are taught in school from the early days that DNA is double-stranded, while RNA is single-stranded. That is a true statement, but somewhat imprecise for the purposes of this competition. RNA is single-stranded **most of the time**, and especially when we are talking about **coding, non-structured RNA molecules**. RNA molecules in this competition are of the **non-coding, structured kind**, and they are double-stranded more often than not. It may sound like that's complicating our lives, but in reality it will make the predictions easier.\n\nI will start explaining this by using a spaghetti analogy, as most people like pasta or at least have a general understanding of spaghetti shape(s). When we pull a handful of spaghetti from the box and before we dump them into a pot of boiling water, all the individual strings are of the same shape and length. After we are done cooking all of them still have the same length, but shapes are completely different. Predicting shapes of long, single-stranded RNAs is like predicting spaghetti shapes. One could reasonably predict a shape of a 2-inch / 5-cm spaghetti because it is not very long, and similarly for a single-stranded RNA molecule that is maybe 5-10 nucleotides long. For anything longer than that this becomes an impossible exercise because there are too many degrees of freedom, and there isn't a single shape that is more likely than the others.\n\nThis is where double-stranded nature of structured RNAs will help us. Surely you have noticed that cooked spaghetti strings stick to each other, and sometimes the sticking happens within different parts of the same string. For spaghetti this is completely random, but not so for RNA. Imagine if a string of spaghetti goes up, makes a U-turn, then sticks back to itself in a reproducible way. Then goes to the right, makes another U-turn, and again sticks back to itself. If we understood what governs these sticking rules, we'd have much easier time predicting the final shape. For RNA, this comes down to base-pairing, which is highly reproducible.\n\nRNA nucleotides (I will interchangeably call them RNA bases) are made of 3 components: a phosphate, a pentose sugar and a nitrogenous base. We will ignore the first two components for a moment and only talk about bases. Those are Adenine (A), Cytosine (C), Guanine (G) and Uracil (U). A canonical base-pairing is A with U and G with C. This is how those pairings look. Note that sugar-base angles are 54 degrees in all cases.\n\n![RNA canonical base-pairing](https://i.ibb.co/ZpnGNmxR/canonical-pairing.png)\n\nYou may ask: why doesn't A pair with G, or A with C, or U with C? Notice first that A and G are larger bases (known as purines) because they have two connected rings, while C and U are single-ring bases (known as pyrimidines). Neither two large nor two small bases opposite each other will work, which means we always have to pair a large base with a small one. In addition to canonical pairings shown above, that still leaves a possibility of A pairing with C and G pairing with U. This is where those dotted lines come into play. They represent hydrogen bonds, and you can see that there are 3 of them in a G-C pair and 2 of them in an A-U pair. A hydrogen bond must have a donor group (either NH or NH2 for nucleic acids) and an acceptor group (either a solo O or a solo N atom). We can't form a hydrogen bond with two donors across from each other, nor with two acceptor groups across from each other. If you mentally replace a U base above with a C, you will notice that in a A-C pair two donors and two acceptors will be lined up across each other, which means that A can't pair with C. Similarly, if you slide a U base up and mentally position it where C is, two acceptors and two donors will be lined up, which means no base-pairing. If you do the same exercise but rotate the U base slightly clockwise and a G base slightly counter-clockwise, you will see that a pattern of donors and acceptors will change, so G can pair with U in what is known as a \"wobble\" pair (see below). It is not as stable as either A-U or G-C, but it is stable enough that it can be made. Note different sugar-base angles for the G-U pair compared to others.\n\n![RNA wobble base-pairing](https://i.ibb.co/YBLrGVJm/g-u-pairing.png)\n\nIf thinking in terms of donors and acceptors is too abstract, you can substitute a donor for a key, and an acceptor for a lock. For a G base the pattern from the top would be `lock-key-key`, while for a C base it would be `key-lock-lock`. When you put them together and slide towards each other, locks and keys match. For an A base the pattern from the top would be `key-lock`, which matches the U’s pattern of `lock-key`. Normally A-C and G-U patterns are incompatible, but if we slide U up and rotate it slightly it will match the part of G’s pattern, which will enable it to make a “wobble” base pair.\n\nUnderstanding molecular reasons for A-U, G-C and G-U base-pairing is not really required. What matters is that it will always be reproducible. Now that we know that our goal is to find a solution that maximizes the number of base pairs, folding RNA will be a breeze. Right? Well, yes and no. For relatively short RNA molecules there will be only so many base-pairing combinations, so folding them is relatively easy. Consider the 20-nucleotide RNA molecule below:\n\n```\n>RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n>RNA_secondary_structure\n((((((((....))))))))\n```\n\nThat molecule forms a stable hairpin structure. Note that for our example there should be only 8 base pairs in a stem rather than 10 as shown below.\n\n![RNA hairpin](https://i.ibb.co/WrXtDcW/RNA-Hairpin.png)\n\nThe meaning of secondary structure representation below the sequence is that each left rounded bracket must pair with a right bracket. Sometimes people use squared (`[]`), curly (`{}`) or angle (`<>`) brackets, but their meaning is the same. Specifically, that bracket notation above means that base #1 pairs with base #20, #2 with #19, #3 with #18 and so forth until #8 pairs with #13. Middle 4 nucleotides are unpaired and form a stem loop. For that particular sequence there is no other base-pairing pattern that would give us 8 pairs in a double-helical conformation. Keep in mind that consecutive bases can't be paired because it is not stereochemically possible - there has to be at least 3-4 bases separating base-pairing partners. Note that in this case the greatest separation between the two bases in the same pair is 18 (#1 with #20), and the shortest separation is 4 (#8 with #13). All of these are relatively short distances, and it would be easy to fold RNA if all base-pairing patterns were similar to this. In many cases, however, the separation between bases that end up in the same pair can be in hundreds. That complicates our optimization problem from a local to a global scale. So even though we know the general \"sticking\" rules for our molecular spaghetti, longer RNA molecules create a large number of combinatorial possibilities. Going back to my statement that `weak molecular interactions that both simplify and complicate the calculation of that shape`: hydrogen bonds between RNA nucleotides that have relatively similar sequence numbers make the folding easier, but complications arise from the fact that base-pairing works just as well between RNA nucleotides that are very far from each other in terms of sequence numbers.\n\nI will explain that last statement by using an example of a more complicated base-pairing pattern (borrowed from [**here**](http://rna.tbi.univie.ac.at/forna/)):\n\n```\n>RNA_sequence\nCGCUUCAUAUAAUCCUAAUGAUAUGGUUUGGGAGUUUCUACCAAGAGCCUUAAACUCUUGAUUAUGAAGUG\n>RNA_secondary_structure\n...(((((((..((((((.........))))))......).((((((.......))))))..))))))...\n```\n\nThis RNA sequence has 71 nucleotides, and it features two stems (double-stranded helical regions) formed by local interactions, and another one formed by long-distance interactions. That may be easier to see in the image below.\n\n<div>\n    <img src=\"https://i.ibb.co/zHL0tfvL/structured-rna.png\" width=\"600\"/>\n</div>\n\nThe sequence starts from `CGC` at the bottom and you can follow it clockwise. Each 10th base is labeled. For example, you can see that A at #18 pairs with U at #28, or that U at #13 pairs with A at #33. That would be an example of a local stem, and the same is true for the other one at 2 o’clock. Now, notice that U#10 pairs with A#40, and that U#4 pairs with A#68. Basically, those are nucleotides from the start and the end pairing with each other. This is still “only” a ~60 nucleotides separation because the molecule’s length is 71, but something like this could have happened in a molecule that is 1000+ nucleotides long. These long-range interactions are most difficult to predict as the number of potential pairing combinations increases dramatically with sequence length.\n\nIn the [**second part**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568633) I explain the principle of co-variation and why MSAs help RNA folding predictions.",
      "votes": null
    },
    {
      "id": "3151338",
      "postDate": "03/16/2025 15:25:18",
      "content": "<blockquote>\n  <p>In the second part I will explain the principle of co-variation and why MSAs help RNA folding predictions.</p>\n</blockquote>\n<p>I just can't wait.  Thank you for the lecture.  Super-duper appreciated.  I finally feel that I'm learning some knowledge instead of just seeing transformer or DL being thrown at every single problem out there.</p>\n<p>It is knowledge like this that will push the boundary of not only biology but also ML/AI itself, not DL, not transformer. Mark my words for this.</p>",
      "rawMarkdown": ">In the second part I will explain the principle of co-variation and why MSAs help RNA folding predictions.\n\nI just can't wait.  Thank you for the lecture.  Super-duper appreciated.  I finally feel that I'm learning some knowledge instead of just seeing transformer or DL being thrown at every single problem out there.\n\nIt is knowledge like this that will push the boundary of not only biology but also ML/AI itself, not DL, not transformer. Mark my words for this.",
      "votes": null
    },
    {
      "id": "3151784",
      "postDate": "03/17/2025 05:02:33",
      "content": "<p>Thank you for your support. I learn from others on Kaggle all the time, and it is good to know that I have helped as well.</p>",
      "rawMarkdown": "Thank you for your support. I learn from others on Kaggle all the time, and it is good to know that I have helped as well.",
      "votes": null
    },
    {
      "id": "3151929",
      "postDate": "03/17/2025 08:53:55",
      "content": "<p>You explained the details completely and accurately thanks for sharing it.</p>",
      "rawMarkdown": "You explained the details completely and accurately thanks for sharing it.",
      "votes": null
    },
    {
      "id": "3152823",
      "postDate": "03/18/2025 07:45:18",
      "content": "<p>Could you please   rephrase it in simple word. I do not understand it well.I mean the above explanation</p>",
      "rawMarkdown": "Could you please   rephrase it in simple word. I do not understand it well.I mean the above explanation",
      "votes": null
    },
    {
      "id": "3152869",
      "postDate": "03/18/2025 08:18:40",
      "content": "<p>Which part? This is a difficult problem that doesn't lend itself to super simple explanations.</p>",
      "rawMarkdown": "Which part? This is a difficult problem that doesn't lend itself to super simple explanations.",
      "votes": null
    },
    {
      "id": "3152896",
      "postDate": "03/18/2025 08:43:17",
      "content": "<p>If I have to say then , the explanation of base-pairing rules, especially why certain bases cannot pair (like A with C or U with C), was a bit complex for me. Also, the part about the separation between paired bases and how that impacts RNA folding predictions felt challenging to follow. Could you simplify these sections or provide a more intuitive explanation? That would help a lot!  </p>",
      "rawMarkdown": "If I have to say then , the explanation of base-pairing rules, especially why certain bases cannot pair (like A with C or U with C), was a bit complex for me. Also, the part about the separation between paired bases and how that impacts RNA folding predictions felt challenging to follow. Could you simplify these sections or provide a more intuitive explanation? That would help a lot!",
      "votes": null
    },
    {
      "id": "3153193",
      "postDate": "03/18/2025 13:55:02",
      "content": "<p>if you want to consider base paring, then you need to consider the 3d location and rotation of the base (residue).<br>\njust close + wrong angle will not lead to pairing.</p>\n<p>Hence I am thinking that if we predict all atoms or at least 3 per residue (3 is needed to define the angle of residue), the results will be more accurate?</p>",
      "rawMarkdown": "if you want to consider base paring, then you need to consider the 3d location and rotation of the base (residue).\njust close + wrong angle will not lead to pairing.\n\nHence I am thinking that if we predict all atoms or at least 3 per residue (3 is needed to define the angle of residue), the results will be more accurate?",
      "votes": null
    },
    {
      "id": "3153279",
      "postDate": "03/18/2025 15:46:12",
      "content": "<p>Thanks for the great explanation. How can you tell the coding ones from the non-coding ones? Can you tell from the sequence (RNA) itself? Or are they somehow marked in databases like rcsb.org?</p>",
      "rawMarkdown": "Thanks for the great explanation. How can you tell the coding ones from the non-coding ones? Can you tell from the sequence (RNA) itself? Or are they somehow marked in databases like rcsb.org?",
      "votes": null
    },
    {
      "id": "3153500",
      "postDate": "03/18/2025 22:48:31",
      "content": "<p>I added an explanation about the base-pairing rules using the lock-key analogy - let's hope that works better. If not, it really is not that important to understand the molecular details as long as one knows about G-C, A-U and GoU pairings.</p>\n<p>I also added a section about long-distance interactions.</p>\n<p>It may help if you read the companion post, which is in a separate thread.</p>",
      "rawMarkdown": "I added an explanation about the base-pairing rules using the lock-key analogy - let's hope that works better. If not, it really is not that important to understand the molecular details as long as one knows about G-C, A-U and GoU pairings.\n\nI also added a section about long-distance interactions.\n\nIt may help if you read the companion post, which is in a separate thread.",
      "votes": null
    },
    {
      "id": "3153505",
      "postDate": "03/18/2025 22:55:17",
      "content": "<p>Generally speaking, coding RNAs have long open reading frames that specify an order of amino-acids in a protein, while non-coding RNAs are shorter and usually have more stop codons inside of them. There are computer programs that use hidden Markov models to discriminate between these types.</p>\n<p>Coding RNAs usually don't have a defined structure, so they are very difficult to study in ways of high-resolution molecular techniques. For practical purposes that means there are no coding RNAs in PDB databases. There may be a short RNA molecule here and there, usually bound by a protein, that is stiff enough so it can be resolved by crystallography or cryo-EM. In a huge majority of cases structures are known only for non-coding RNAs, which is what we are dealing with in this competition.</p>",
      "rawMarkdown": "Generally speaking, coding RNAs have long open reading frames that specify an order of amino-acids in a protein, while non-coding RNAs are shorter and usually have more stop codons inside of them. There are computer programs that use hidden Markov models to discriminate between these types.\n\nCoding RNAs usually don't have a defined structure, so they are very difficult to study in ways of high-resolution molecular techniques. For practical purposes that means there are no coding RNAs in PDB databases. There may be a short RNA molecule here and there, usually bound by a protein, that is stiff enough so it can be resolved by crystallography or cryo-EM. In a huge majority of cases structures are known only for non-coding RNAs, which is what we are dealing with in this competition.",
      "votes": null
    },
    {
      "id": "3153517",
      "postDate": "03/18/2025 23:08:57",
      "content": "<p>You are absolutely right that many distance restraints are required beyond C1-C1 separation. Modeling programs create many of them automatically, as nitrogenous bases are stiff and their bond lengths and angles are almost fixed. The angle of bases with regard to each other also allows little variation if hydrogen bonds are to be formed. Distances between donor and acceptor atoms engaging in hydrogen bonding likewise do not vary appreciably.</p>\n<p>This is all to say that knowing the identity of each nucleotide already provides a set of local distance restraints within the base. There are some other simple restraints between bases, such as C1 of base #3 and of base #4 should be separated with a minimum of 4.5 and a maximum of 7 angstroms. If base #3 is paired with #18, their C1-C1 restraint would be a minimum of 9.5 and a maximum of 11 angstroms. If you want to learn how this works in greater detail, I suggest you check out <a href=\"https://genesilico.pl/software/stand-alone/simrna\" target=\"_blank\"><strong>SimRNA</strong></a>.</p>\n<p>I do not recommend that you develop your own modeling procedure based on distance restraints, as you'd be reinventing the wheel. Also, it is unlikely that anyone could come up with a better way of doing it in the 2.5 months we have left. Instead, I think it is more productive to find ways to correctly predict the base-pairing and other interaction patterns, and let the energy-based programs such as <a href=\"https://genesilico.pl/software/stand-alone/simrna\" target=\"_blank\"><strong>SimRNA</strong></a> or one of many DL-based programs (you wrote about <a href=\"https://github.com/ml4bio/RhoFold\" target=\"_blank\"><strong>RhoFold+</strong></a>) figure out how to create a 3D model.</p>",
      "rawMarkdown": "You are absolutely right that many distance restraints are required beyond C1-C1 separation. Modeling programs create many of them automatically, as nitrogenous bases are stiff and their bond lengths and angles are almost fixed. The angle of bases with regard to each other also allows little variation if hydrogen bonds are to be formed. Distances between donor and acceptor atoms engaging in hydrogen bonding likewise do not vary appreciably.\n\nThis is all to say that knowing the identity of each nucleotide already provides a set of local distance restraints within the base. There are some other simple restraints between bases, such as C1 of base #3 and of base #4 should be separated with a minimum of 4.5 and a maximum of 7 angstroms. If base #3 is paired with #18, their C1-C1 restraint would be a minimum of 9.5 and a maximum of 11 angstroms. If you want to learn how this works in greater detail, I suggest you check out [**SimRNA**](https://genesilico.pl/software/stand-alone/simrna).\n\nI do not recommend that you develop your own modeling procedure based on distance restraints, as you'd be reinventing the wheel. Also, it is unlikely that anyone could come up with a better way of doing it in the 2.5 months we have left. Instead, I think it is more productive to find ways to correctly predict the base-pairing and other interaction patterns, and let the energy-based programs such as [**SimRNA**](https://genesilico.pl/software/stand-alone/simrna) or one of many DL-based programs (you wrote about [**RhoFold+**](https://github.com/ml4bio/RhoFold)) figure out how to create a 3D model.",
      "votes": null
    },
    {
      "id": "3153592",
      "postDate": "03/19/2025 02:03:38",
      "content": "<p>thanks for introducing simRNA, i will take a look.<br>\njust a correction: the post <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512</a> is based on drfoldv1/2.</p>\n<p>however,  RhoFold also have its relaxation based on openMM and AMBER.</p>\n<hr>\n<p>It seems that raw results from deep learning is the best for tm-score. any form of relaxation or energy optimisation worsen t the score. (but not sure about other metrics like rmsd or inf)</p>",
      "rawMarkdown": "thanks for introducing simRNA, i will take a look.\njust a correction: the post https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512 is based on drfoldv1/2.\n\nhowever,  RhoFold also have its relaxation based on openMM and AMBER.\n\n---\n\nIt seems that raw results from deep learning is the best for tm-score. any form of relaxation or energy optimisation worsen t the score. (but not sure about other metrics like rmsd or inf)",
      "votes": null
    },
    {
      "id": "3154266",
      "postDate": "03/19/2025 18:20:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Thanks for the illustration.  It there a reason that these bases (three A-U and two C-G) are not paired up?  Surely, they're more than 4 bases apart from each other in the primary structure, right?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F7bc9faceb1c158839ce2771860ba3072%2Fsample_pairing.png?generation=1742409245515417&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi @tilii7 Thanks for the illustration.  It there a reason that these bases (three A-U and two C-G) are not paired up?  Surely, they're more than 4 bases apart from each other in the primary structure, right?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F7bc9faceb1c158839ce2771860ba3072%2Fsample_pairing.png?generation=1742409245515417&alt=media)",
      "votes": null
    },
    {
      "id": "3154335",
      "postDate": "03/19/2025 20:18:26",
      "content": "<p>Good catch. Keep in mind that this is a prediction that could be wrong, and the drawing program will do whatever it is told to do. That explains why C#1 and G#71 are not paired, even though they should be able to do it. Same for the other two bases in that region.</p>\n<p>More importantly, we have to account for the vertical rise with each base. If we look at the base pair A#18-U#28, A#21 is 12+ angstroms removed from #18, while U#27 is only about 5 angstroms away from #28. They are not in the same plane and can't form a hydrogen bond. The same reasoning applies to other examples in the top half of the image. One can even argue that A#11 and U#39 should pair, but that wouldn't leave enough room for all the bases in the upper rim of that that internal loop.</p>",
      "rawMarkdown": "Good catch. Keep in mind that this is a prediction that could be wrong, and the drawing program will do whatever it is told to do. That explains why C#1 and G#71 are not paired, even though they should be able to do it. Same for the other two bases in that region.\n\nMore importantly, we have to account for the vertical rise with each base. If we look at the base pair A#18-U#28, A#21 is 12+ angstroms removed from #18, while U#27 is only about 5 angstroms away from #28. They are not in the same plane and can't form a hydrogen bond. The same reasoning applies to other examples in the top half of the image. One can even argue that A#11 and U#39 should pair, but that wouldn't leave enough room for all the bases in the upper rim of that that internal loop.",
      "votes": null
    },
    {
      "id": "3154724",
      "postDate": "03/20/2025 10:20:24",
      "content": "<p>this show the benchmark vfold performance on coding and non coding<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3154714\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3154714</a></p>\n<p>\"Long non-coding RNAs (lncRNAs) are RNA molecules, longer than 200 nucleotides, that do not code for proteins but play crucial regulatory roles in various cellular processes, including gene expression and development. \"<br>\n<a href=\"https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf\" target=\"_blank\">https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf</a></p>",
      "rawMarkdown": "this show the benchmark vfold performance on coding and non coding\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3154714\n\n\"Long non-coding RNAs (lncRNAs) are RNA molecules, longer than 200 nucleotides, that do not code for proteins but play crucial regulatory roles in various cellular processes, including gene expression and development. \"\nhttps://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf",
      "votes": null
    },
    {
      "id": "3160820",
      "postDate": "03/27/2025 07:20:34",
      "content": "<p>thank for your clear explain</p>",
      "rawMarkdown": "thank for your clear explain",
      "votes": null
    },
    {
      "id": "3201682",
      "postDate": "05/14/2025 08:48:47",
      "content": "<p>Thanks for sharing, I'm novice in that so I found that very helpfull.</p>",
      "rawMarkdown": "Thanks for sharing, I'm novice in that so I found that very helpfull.",
      "votes": null
    },
    {
      "id": "3207739",
      "postDate": "05/23/2025 07:20:01",
      "content": "<p>Wonderful sharing! </p>",
      "rawMarkdown": "Wonderful sharing!",
      "votes": null
    },
    {
      "id": "3208955",
      "postDate": "05/25/2025 02:53:54",
      "content": "<p>very intriguing</p>",
      "rawMarkdown": "very intriguing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3151338,
      "author_name": "revealer",
      "author_url": "",
      "post_date": "03/16/2025 15:25:18",
      "content": "<blockquote>\n  <p>In the second part I will explain the principle of co-variation and why MSAs help RNA folding predictions.</p>\n</blockquote>\n<p>I just can't wait.  Thank you for the lecture.  Super-duper appreciated.  I finally feel that I'm learning some knowledge instead of just seeing transformer or DL being thrown at every single problem out there.</p>\n<p>It is knowledge like this that will push the boundary of not only biology but also ML/AI itself, not DL, not transformer. Mark my words for this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3151784,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/17/2025 05:02:33",
          "content": "<p>Thank you for your support. I learn from others on Kaggle all the time, and it is good to know that I have helped as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3152823,
          "author_name": "",
          "author_url": "",
          "post_date": "03/18/2025 07:45:18",
          "content": "<p>Could you please   rephrase it in simple word. I do not understand it well.I mean the above explanation</p>",
          "votes": null,
          "replies": [
            {
              "id": 3152869,
              "author_name": "tilii7",
              "author_url": "",
              "post_date": "03/18/2025 08:18:40",
              "content": "<p>Which part? This is a difficult problem that doesn't lend itself to super simple explanations.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3152896,
                  "author_name": "",
                  "author_url": "",
                  "post_date": "03/18/2025 08:43:17",
                  "content": "<p>If I have to say then , the explanation of base-pairing rules, especially why certain bases cannot pair (like A with C or U with C), was a bit complex for me. Also, the part about the separation between paired bases and how that impacts RNA folding predictions felt challenging to follow. Could you simplify these sections or provide a more intuitive explanation? That would help a lot!  </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3153500,
                      "author_name": "tilii7",
                      "author_url": "",
                      "post_date": "03/18/2025 22:48:31",
                      "content": "<p>I added an explanation about the base-pairing rules using the lock-key analogy - let's hope that works better. If not, it really is not that important to understand the molecular details as long as one knows about G-C, A-U and GoU pairings.</p>\n<p>I also added a section about long-distance interactions.</p>\n<p>It may help if you read the companion post, which is in a separate thread.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3151929,
      "author_name": "aidaabd96",
      "author_url": "",
      "post_date": "03/17/2025 08:53:55",
      "content": "<p>You explained the details completely and accurately thanks for sharing it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3153193,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/18/2025 13:55:02",
      "content": "<p>if you want to consider base paring, then you need to consider the 3d location and rotation of the base (residue).<br>\njust close + wrong angle will not lead to pairing.</p>\n<p>Hence I am thinking that if we predict all atoms or at least 3 per residue (3 is needed to define the angle of residue), the results will be more accurate?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3153517,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/18/2025 23:08:57",
          "content": "<p>You are absolutely right that many distance restraints are required beyond C1-C1 separation. Modeling programs create many of them automatically, as nitrogenous bases are stiff and their bond lengths and angles are almost fixed. The angle of bases with regard to each other also allows little variation if hydrogen bonds are to be formed. Distances between donor and acceptor atoms engaging in hydrogen bonding likewise do not vary appreciably.</p>\n<p>This is all to say that knowing the identity of each nucleotide already provides a set of local distance restraints within the base. There are some other simple restraints between bases, such as C1 of base #3 and of base #4 should be separated with a minimum of 4.5 and a maximum of 7 angstroms. If base #3 is paired with #18, their C1-C1 restraint would be a minimum of 9.5 and a maximum of 11 angstroms. If you want to learn how this works in greater detail, I suggest you check out <a href=\"https://genesilico.pl/software/stand-alone/simrna\" target=\"_blank\"><strong>SimRNA</strong></a>.</p>\n<p>I do not recommend that you develop your own modeling procedure based on distance restraints, as you'd be reinventing the wheel. Also, it is unlikely that anyone could come up with a better way of doing it in the 2.5 months we have left. Instead, I think it is more productive to find ways to correctly predict the base-pairing and other interaction patterns, and let the energy-based programs such as <a href=\"https://genesilico.pl/software/stand-alone/simrna\" target=\"_blank\"><strong>SimRNA</strong></a> or one of many DL-based programs (you wrote about <a href=\"https://github.com/ml4bio/RhoFold\" target=\"_blank\"><strong>RhoFold+</strong></a>) figure out how to create a 3D model.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3153592,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/19/2025 02:03:38",
              "content": "<p>thanks for introducing simRNA, i will take a look.<br>\njust a correction: the post <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512</a> is based on drfoldv1/2.</p>\n<p>however,  RhoFold also have its relaxation based on openMM and AMBER.</p>\n<hr>\n<p>It seems that raw results from deep learning is the best for tm-score. any form of relaxation or energy optimisation worsen t the score. (but not sure about other metrics like rmsd or inf)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3153279,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "03/18/2025 15:46:12",
      "content": "<p>Thanks for the great explanation. How can you tell the coding ones from the non-coding ones? Can you tell from the sequence (RNA) itself? Or are they somehow marked in databases like rcsb.org?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3153505,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/18/2025 22:55:17",
          "content": "<p>Generally speaking, coding RNAs have long open reading frames that specify an order of amino-acids in a protein, while non-coding RNAs are shorter and usually have more stop codons inside of them. There are computer programs that use hidden Markov models to discriminate between these types.</p>\n<p>Coding RNAs usually don't have a defined structure, so they are very difficult to study in ways of high-resolution molecular techniques. For practical purposes that means there are no coding RNAs in PDB databases. There may be a short RNA molecule here and there, usually bound by a protein, that is stiff enough so it can be resolved by crystallography or cryo-EM. In a huge majority of cases structures are known only for non-coding RNAs, which is what we are dealing with in this competition.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3154724,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/20/2025 10:20:24",
              "content": "<p>this show the benchmark vfold performance on coding and non coding<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3154714\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3154714</a></p>\n<p>\"Long non-coding RNAs (lncRNAs) are RNA molecules, longer than 200 nucleotides, that do not code for proteins but play crucial regulatory roles in various cellular processes, including gene expression and development. \"<br>\n<a href=\"https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf\" target=\"_blank\">https://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3154266,
      "author_name": "revealer",
      "author_url": "",
      "post_date": "03/19/2025 18:20:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Thanks for the illustration.  It there a reason that these bases (three A-U and two C-G) are not paired up?  Surely, they're more than 4 bases apart from each other in the primary structure, right?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F7bc9faceb1c158839ce2771860ba3072%2Fsample_pairing.png?generation=1742409245515417&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3154335,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/19/2025 20:18:26",
          "content": "<p>Good catch. Keep in mind that this is a prediction that could be wrong, and the drawing program will do whatever it is told to do. That explains why C#1 and G#71 are not paired, even though they should be able to do it. Same for the other two bases in that region.</p>\n<p>More importantly, we have to account for the vertical rise with each base. If we look at the base pair A#18-U#28, A#21 is 12+ angstroms removed from #18, while U#27 is only about 5 angstroms away from #28. They are not in the same plane and can't form a hydrogen bond. The same reasoning applies to other examples in the top half of the image. One can even argue that A#11 and U#39 should pair, but that wouldn't leave enough room for all the bases in the upper rim of that that internal loop.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3160820,
      "author_name": "dedquoc",
      "author_url": "",
      "post_date": "03/27/2025 07:20:34",
      "content": "<p>thank for your clear explain</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3201682,
      "author_name": "xaocreal",
      "author_url": "",
      "post_date": "05/14/2025 08:48:47",
      "content": "<p>Thanks for sharing, I'm novice in that so I found that very helpfull.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3207739,
      "author_name": "",
      "author_url": "",
      "post_date": "05/23/2025 07:20:01",
      "content": "<p>Wonderful sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3208955,
      "author_name": "sonmou",
      "author_url": "",
      "post_date": "05/25/2025 02:53:54",
      "content": "<p>very intriguing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3150812": "I have been following this competition from the start, and it seems that the main problem is in understanding the driving forces for RNA folding. Organizers have given a link to [**a paper**](https://www.pnas.org/doi/10.1073/pnas.2112677119) that explains RNA folding at a high level, and that paper is definitely worth reading. However, I suspect that not everyone will understand RNA folding that is explained at that level and written with formalism that is required in scientific publications.\n\nI am sure many intuitively understand the goal of the challenge, which is how a string of RNA nucleotides assumes a certain 3D shape. Yet there are some weak molecular interactions that both simplify and complicate the calculation of that shape. This previous statement will likely remain contradictory until we get to the end of this story.\n\nWe are taught in school from the early days that DNA is double-stranded, while RNA is single-stranded. That is a true statement, but somewhat imprecise for the purposes of this competition. RNA is single-stranded **most of the time**, and especially when we are talking about **coding, non-structured RNA molecules**. RNA molecules in this competition are of the **non-coding, structured kind**, and they are double-stranded more often than not. It may sound like that's complicating our lives, but in reality it will make the predictions easier.\n\nI will start explaining this by using a spaghetti analogy, as most people like pasta or at least have a general understanding of spaghetti shape(s). When we pull a handful of spaghetti from the box and before we dump them into a pot of boiling water, all the individual strings are of the same shape and length. After we are done cooking all of them still have the same length, but shapes are completely different. Predicting shapes of long, single-stranded RNAs is like predicting spaghetti shapes. One could reasonably predict a shape of a 2-inch / 5-cm spaghetti because it is not very long, and similarly for a single-stranded RNA molecule that is maybe 5-10 nucleotides long. For anything longer than that this becomes an impossible exercise because there are too many degrees of freedom, and there isn't a single shape that is more likely than the others.\n\nThis is where double-stranded nature of structured RNAs will help us. Surely you have noticed that cooked spaghetti strings stick to each other, and sometimes the sticking happens within different parts of the same string. For spaghetti this is completely random, but not so for RNA. Imagine if a string of spaghetti goes up, makes a U-turn, then sticks back to itself in a reproducible way. Then goes to the right, makes another U-turn, and again sticks back to itself. If we understood what governs these sticking rules, we'd have much easier time predicting the final shape. For RNA, this comes down to base-pairing, which is highly reproducible.\n\nRNA nucleotides (I will interchangeably call them RNA bases) are made of 3 components: a phosphate, a pentose sugar and a nitrogenous base. We will ignore the first two components for a moment and only talk about bases. Those are Adenine (A), Cytosine (C), Guanine (G) and Uracil (U). A canonical base-pairing is A with U and G with C. This is how those pairings look. Note that sugar-base angles are 54 degrees in all cases.\n\n![RNA canonical base-pairing](https://i.ibb.co/ZpnGNmxR/canonical-pairing.png)\n\nYou may ask: why doesn't A pair with G, or A with C, or U with C? Notice first that A and G are larger bases (known as purines) because they have two connected rings, while C and U are single-ring bases (known as pyrimidines). Neither two large nor two small bases opposite each other will work, which means we always have to pair a large base with a small one. In addition to canonical pairings shown above, that still leaves a possibility of A pairing with C and G pairing with U. This is where those dotted lines come into play. They represent hydrogen bonds, and you can see that there are 3 of them in a G-C pair and 2 of them in an A-U pair. A hydrogen bond must have a donor group (either NH or NH2 for nucleic acids) and an acceptor group (either a solo O or a solo N atom). We can't form a hydrogen bond with two donors across from each other, nor with two acceptor groups across from each other. If you mentally replace a U base above with a C, you will notice that in a A-C pair two donors and two acceptors will be lined up across each other, which means that A can't pair with C. Similarly, if you slide a U base up and mentally position it where C is, two acceptors and two donors will be lined up, which means no base-pairing. If you do the same exercise but rotate the U base slightly clockwise and a G base slightly counter-clockwise, you will see that a pattern of donors and acceptors will change, so G can pair with U in what is known as a \"wobble\" pair (see below). It is not as stable as either A-U or G-C, but it is stable enough that it can be made. Note different sugar-base angles for the G-U pair compared to others.\n\n![RNA wobble base-pairing](https://i.ibb.co/YBLrGVJm/g-u-pairing.png)\n\nIf thinking in terms of donors and acceptors is too abstract, you can substitute a donor for a key, and an acceptor for a lock. For a G base the pattern from the top would be `lock-key-key`, while for a C base it would be `key-lock-lock`. When you put them together and slide towards each other, locks and keys match. For an A base the pattern from the top would be `key-lock`, which matches the U’s pattern of `lock-key`. Normally A-C and G-U patterns are incompatible, but if we slide U up and rotate it slightly it will match the part of G’s pattern, which will enable it to make a “wobble” base pair.\n\nUnderstanding molecular reasons for A-U, G-C and G-U base-pairing is not really required. What matters is that it will always be reproducible. Now that we know that our goal is to find a solution that maximizes the number of base pairs, folding RNA will be a breeze. Right? Well, yes and no. For relatively short RNA molecules there will be only so many base-pairing combinations, so folding them is relatively easy. Consider the 20-nucleotide RNA molecule below:\n\n```\n>RNA_sequence\nCGGAAGGUGCGAAUCUUCCG\n>RNA_secondary_structure\n((((((((....))))))))\n```\n\nThat molecule forms a stable hairpin structure. Note that for our example there should be only 8 base pairs in a stem rather than 10 as shown below.\n\n![RNA hairpin](https://i.ibb.co/WrXtDcW/RNA-Hairpin.png)\n\nThe meaning of secondary structure representation below the sequence is that each left rounded bracket must pair with a right bracket. Sometimes people use squared (`[]`), curly (`{}`) or angle (`<>`) brackets, but their meaning is the same. Specifically, that bracket notation above means that base #1 pairs with base #20, #2 with #19, #3 with #18 and so forth until #8 pairs with #13. Middle 4 nucleotides are unpaired and form a stem loop. For that particular sequence there is no other base-pairing pattern that would give us 8 pairs in a double-helical conformation. Keep in mind that consecutive bases can't be paired because it is not stereochemically possible - there has to be at least 3-4 bases separating base-pairing partners. Note that in this case the greatest separation between the two bases in the same pair is 18 (#1 with #20), and the shortest separation is 4 (#8 with #13). All of these are relatively short distances, and it would be easy to fold RNA if all base-pairing patterns were similar to this. In many cases, however, the separation between bases that end up in the same pair can be in hundreds. That complicates our optimization problem from a local to a global scale. So even though we know the general \"sticking\" rules for our molecular spaghetti, longer RNA molecules create a large number of combinatorial possibilities. Going back to my statement that `weak molecular interactions that both simplify and complicate the calculation of that shape`: hydrogen bonds between RNA nucleotides that have relatively similar sequence numbers make the folding easier, but complications arise from the fact that base-pairing works just as well between RNA nucleotides that are very far from each other in terms of sequence numbers.\n\nI will explain that last statement by using an example of a more complicated base-pairing pattern (borrowed from [**here**](http://rna.tbi.univie.ac.at/forna/)):\n\n```\n>RNA_sequence\nCGCUUCAUAUAAUCCUAAUGAUAUGGUUUGGGAGUUUCUACCAAGAGCCUUAAACUCUUGAUUAUGAAGUG\n>RNA_secondary_structure\n...(((((((..((((((.........))))))......).((((((.......))))))..))))))...\n```\n\nThis RNA sequence has 71 nucleotides, and it features two stems (double-stranded helical regions) formed by local interactions, and another one formed by long-distance interactions. That may be easier to see in the image below.\n\n<div>\n    <img src=\"https://i.ibb.co/zHL0tfvL/structured-rna.png\" width=\"600\"/>\n</div>\n\nThe sequence starts from `CGC` at the bottom and you can follow it clockwise. Each 10th base is labeled. For example, you can see that A at #18 pairs with U at #28, or that U at #13 pairs with A at #33. That would be an example of a local stem, and the same is true for the other one at 2 o’clock. Now, notice that U#10 pairs with A#40, and that U#4 pairs with A#68. Basically, those are nucleotides from the start and the end pairing with each other. This is still “only” a ~60 nucleotides separation because the molecule’s length is 71, but something like this could have happened in a molecule that is 1000+ nucleotides long. These long-range interactions are most difficult to predict as the number of potential pairing combinations increases dramatically with sequence length.\n\nIn the [**second part**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568633) I explain the principle of co-variation and why MSAs help RNA folding predictions.",
    "3151338": ">In the second part I will explain the principle of co-variation and why MSAs help RNA folding predictions.\n\nI just can't wait.  Thank you for the lecture.  Super-duper appreciated.  I finally feel that I'm learning some knowledge instead of just seeing transformer or DL being thrown at every single problem out there.\n\nIt is knowledge like this that will push the boundary of not only biology but also ML/AI itself, not DL, not transformer. Mark my words for this.",
    "3151784": "Thank you for your support. I learn from others on Kaggle all the time, and it is good to know that I have helped as well.",
    "3151929": "You explained the details completely and accurately thanks for sharing it.",
    "3152823": "Could you please   rephrase it in simple word. I do not understand it well.I mean the above explanation",
    "3152869": "Which part? This is a difficult problem that doesn't lend itself to super simple explanations.",
    "3152896": "If I have to say then , the explanation of base-pairing rules, especially why certain bases cannot pair (like A with C or U with C), was a bit complex for me. Also, the part about the separation between paired bases and how that impacts RNA folding predictions felt challenging to follow. Could you simplify these sections or provide a more intuitive explanation? That would help a lot!",
    "3153193": "if you want to consider base paring, then you need to consider the 3d location and rotation of the base (residue).\njust close + wrong angle will not lead to pairing.\n\nHence I am thinking that if we predict all atoms or at least 3 per residue (3 is needed to define the angle of residue), the results will be more accurate?",
    "3153279": "Thanks for the great explanation. How can you tell the coding ones from the non-coding ones? Can you tell from the sequence (RNA) itself? Or are they somehow marked in databases like rcsb.org?",
    "3153500": "I added an explanation about the base-pairing rules using the lock-key analogy - let's hope that works better. If not, it really is not that important to understand the molecular details as long as one knows about G-C, A-U and GoU pairings.\n\nI also added a section about long-distance interactions.\n\nIt may help if you read the companion post, which is in a separate thread.",
    "3153505": "Generally speaking, coding RNAs have long open reading frames that specify an order of amino-acids in a protein, while non-coding RNAs are shorter and usually have more stop codons inside of them. There are computer programs that use hidden Markov models to discriminate between these types.\n\nCoding RNAs usually don't have a defined structure, so they are very difficult to study in ways of high-resolution molecular techniques. For practical purposes that means there are no coding RNAs in PDB databases. There may be a short RNA molecule here and there, usually bound by a protein, that is stiff enough so it can be resolved by crystallography or cryo-EM. In a huge majority of cases structures are known only for non-coding RNAs, which is what we are dealing with in this competition.",
    "3153517": "You are absolutely right that many distance restraints are required beyond C1-C1 separation. Modeling programs create many of them automatically, as nitrogenous bases are stiff and their bond lengths and angles are almost fixed. The angle of bases with regard to each other also allows little variation if hydrogen bonds are to be formed. Distances between donor and acceptor atoms engaging in hydrogen bonding likewise do not vary appreciably.\n\nThis is all to say that knowing the identity of each nucleotide already provides a set of local distance restraints within the base. There are some other simple restraints between bases, such as C1 of base #3 and of base #4 should be separated with a minimum of 4.5 and a maximum of 7 angstroms. If base #3 is paired with #18, their C1-C1 restraint would be a minimum of 9.5 and a maximum of 11 angstroms. If you want to learn how this works in greater detail, I suggest you check out [**SimRNA**](https://genesilico.pl/software/stand-alone/simrna).\n\nI do not recommend that you develop your own modeling procedure based on distance restraints, as you'd be reinventing the wheel. Also, it is unlikely that anyone could come up with a better way of doing it in the 2.5 months we have left. Instead, I think it is more productive to find ways to correctly predict the base-pairing and other interaction patterns, and let the energy-based programs such as [**SimRNA**](https://genesilico.pl/software/stand-alone/simrna) or one of many DL-based programs (you wrote about [**RhoFold+**](https://github.com/ml4bio/RhoFold)) figure out how to create a 3D model.",
    "3153592": "thanks for introducing simRNA, i will take a look.\njust a correction: the post https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568512 is based on drfoldv1/2.\n\nhowever,  RhoFold also have its relaxation based on openMM and AMBER.\n\n---\n\nIt seems that raw results from deep learning is the best for tm-score. any form of relaxation or energy optimisation worsen t the score. (but not sure about other metrics like rmsd or inf)",
    "3154266": "Hi @tilii7 Thanks for the illustration.  It there a reason that these bases (three A-U and two C-G) are not paired up?  Surely, they're more than 4 bases apart from each other in the primary structure, right?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F7bc9faceb1c158839ce2771860ba3072%2Fsample_pairing.png?generation=1742409245515417&alt=media)",
    "3154335": "Good catch. Keep in mind that this is a prediction that could be wrong, and the drawing program will do whatever it is told to do. That explains why C#1 and G#71 are not paired, even though they should be able to do it. Same for the other two bases in that region.\n\nMore importantly, we have to account for the vertical rise with each base. If we look at the base pair A#18-U#28, A#21 is 12+ angstroms removed from #18, while U#27 is only about 5 angstroms away from #28. They are not in the same plane and can't form a hydrogen bond. The same reasoning applies to other examples in the top half of the image. One can even argue that A#11 and U#39 should pair, but that wouldn't leave enough room for all the bases in the upper rim of that that internal loop.",
    "3154724": "this show the benchmark vfold performance on coding and non coding\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3154714\n\n\"Long non-coding RNAs (lncRNAs) are RNA molecules, longer than 200 nucleotides, that do not code for proteins but play crucial regulatory roles in various cellular processes, including gene expression and development. \"\nhttps://predictioncenter.org/casp16/doc/presentations/Day-3/Day3-03-Chen-Vfold-RNA-Predictor-Talk1_Redacted.pdf",
    "3160820": "thank for your clear explain",
    "3201682": "Thanks for sharing, I'm novice in that so I found that very helpfull.",
    "3207739": "Wonderful sharing!",
    "3208955": "very intriguing"
  },
  "source": "meta"
}