{
  "id": 446900,
  "title": "RNA Ensembles",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/446900",
  "author_name": "DigitalEmbrace",
  "post_date": "2023-10-13T14:32:54.957000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>One reason the reactivity value for most positions is not precisely 0 or 1 is that RNA molecules constantly move among two or more conformations. Some of the movements are minor but the shifts also can be significant. This is one of the aspects that makes RNA folding prediction challenging and where we believe deep learning can help.</p>\n<p>RNA researchers refer to the collection of conformations as the ensemble. The ensemble is \"read\" from the thousands of copies tested for each sequence. In a similar vein, the predicted ensemble is reflected in the bpp (base pair probability) additional files with matrices.</p>\n<p>To answer a question about the bpp files that has been raised, yes, bpp files have been generated in <a href=\"https://github.com/eternagame/EternaFold\" target=\"_blank\">Eternafold</a> for the test set. (Note that Eternafold does not model pseudoknots when calculating bpp.)</p>\n<p>Thomas's explanation of the bpp additional files from <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2468751\" target=\"_blank\">another thread</a>:</p>\n<blockquote>\n  <p>The column names reflect the software packages utilized, but there are some nuances to be aware of. Generally speaking, the name of the column references the software used to make the prediction. For example,vienna2_mfe is a prediction made by the ViennaRNA package. Here is the list of packages:</p>\n  <p>ViennaRNA<br>\n  Contrafold<br>\n  Nupack<br>\n  Eternafold<br>\n  e2efold<br>\n  hotknots<br>\n  ipknots<br>\n  iterative-hfold<br>\n  knotty<br>\n  pknots<br>\n  spotrna<br>\n  shapify</p>\n  <p>Many of these prediction algorithms produce dot-bracket notation strings representing the predicted structure, but some can also produce a base pair probability matrix detailing the likelihood of each base in an input sequence pairing to every other base in the sequence. These base pair probability matrices can be used as inputs to heuristic algorithms that can predict a dot-bracket structure. These heuristic algorithms can be used to predict pseudoknots, even when paired with input algorithms that don't consider pseudoknot presence. We use two heuristic algorithms here, Threshknot and an implementation of the Hungarian algorithm. Columns of the form eternafold[hungarian]_mfe used Eternafold to predict a base pair probability matrix, then used the hungarian algorithm to generate a dot-bracket structure.</p>\n  <p>One more wrinkle to these data; most of these predictions were made running on a Stanford supercomputer cluster. If the algorithm is prefixed with eterna_ as in <a href=\"https://forum.eternagame.org/t/new-folding-engine-discussion-eternafoldthreshknots/4540\" target=\"_blank\">eterna_eternafold_threshknot</a>, the predictor was run using the version of the associated algorithm available in the <a href=\"https://eternagame.org/labs/11627606\" target=\"_blank\">Eterna game</a>. These should largely be the same, but it is possible that differences in software versions or heuristic implementations could lead to differences. So in your specific example, eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by threshknot, run on a Stanford supercomputer, while eterna_eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by Eterna's implementation of threshknot, run in the Eterna game.</p>\n</blockquote>",
  "messages": [
    {
      "id": 2480754,
      "postDate": "2023-10-13T14:32:54.957Z",
      "content": "<p>One reason the reactivity value for most positions is not precisely 0 or 1 is that RNA molecules constantly move among two or more conformations. Some of the movements are minor but the shifts also can be significant. This is one of the aspects that makes RNA folding prediction challenging and where we believe deep learning can help.</p>\n<p>RNA researchers refer to the collection of conformations as the ensemble. The ensemble is \"read\" from the thousands of copies tested for each sequence. In a similar vein, the predicted ensemble is reflected in the bpp (base pair probability) additional files with matrices.</p>\n<p>To answer a question about the bpp files that has been raised, yes, bpp files have been generated in <a href=\"https://github.com/eternagame/EternaFold\" target=\"_blank\">Eternafold</a> for the test set. (Note that Eternafold does not model pseudoknots when calculating bpp.)</p>\n<p>Thomas's explanation of the bpp additional files from <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2468751\" target=\"_blank\">another thread</a>:</p>\n<blockquote>\n  <p>The column names reflect the software packages utilized, but there are some nuances to be aware of. Generally speaking, the name of the column references the software used to make the prediction. For example,vienna2_mfe is a prediction made by the ViennaRNA package. Here is the list of packages:</p>\n  <p>ViennaRNA<br>\n  Contrafold<br>\n  Nupack<br>\n  Eternafold<br>\n  e2efold<br>\n  hotknots<br>\n  ipknots<br>\n  iterative-hfold<br>\n  knotty<br>\n  pknots<br>\n  spotrna<br>\n  shapify</p>\n  <p>Many of these prediction algorithms produce dot-bracket notation strings representing the predicted structure, but some can also produce a base pair probability matrix detailing the likelihood of each base in an input sequence pairing to every other base in the sequence. These base pair probability matrices can be used as inputs to heuristic algorithms that can predict a dot-bracket structure. These heuristic algorithms can be used to predict pseudoknots, even when paired with input algorithms that don't consider pseudoknot presence. We use two heuristic algorithms here, Threshknot and an implementation of the Hungarian algorithm. Columns of the form eternafold[hungarian]_mfe used Eternafold to predict a base pair probability matrix, then used the hungarian algorithm to generate a dot-bracket structure.</p>\n  <p>One more wrinkle to these data; most of these predictions were made running on a Stanford supercomputer cluster. If the algorithm is prefixed with eterna_ as in <a href=\"https://forum.eternagame.org/t/new-folding-engine-discussion-eternafoldthreshknots/4540\" target=\"_blank\">eterna_eternafold_threshknot</a>, the predictor was run using the version of the associated algorithm available in the <a href=\"https://eternagame.org/labs/11627606\" target=\"_blank\">Eterna game</a>. These should largely be the same, but it is possible that differences in software versions or heuristic implementations could lead to differences. So in your specific example, eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by threshknot, run on a Stanford supercomputer, while eterna_eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by Eterna's implementation of threshknot, run in the Eterna game.</p>\n</blockquote>",
      "rawMarkdown": "One reason the reactivity value for most positions is not precisely 0 or 1 is that RNA molecules constantly move among two or more conformations. Some of the movements are minor but the shifts also can be significant. This is one of the aspects that makes RNA folding prediction challenging and where we believe deep learning can help.\n\nRNA researchers refer to the collection of conformations as the ensemble. The ensemble is \"read\" from the thousands of copies tested for each sequence. In a similar vein, the predicted ensemble is reflected in the bpp (base pair probability) additional files with matrices.\n\nTo answer a question about the bpp files that has been raised, yes, bpp files have been generated in [Eternafold](https://github.com/eternagame/EternaFold) for the test set. (Note that Eternafold does not model pseudoknots when calculating bpp.)\n\nThomas's explanation of the bpp additional files from [another thread](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2468751):\n\n>The column names reflect the software packages utilized, but there are some nuances to be aware of. Generally speaking, the name of the column references the software used to make the prediction. For example,vienna2_mfe is a prediction made by the ViennaRNA package. Here is the list of packages:\n>\n>ViennaRNA\n>Contrafold\n>Nupack\n>Eternafold\n>e2efold\n>hotknots\n>ipknots\n>iterative-hfold\n>knotty\n>pknots\n>spotrna\n>shapify\n>\n>Many of these prediction algorithms produce dot-bracket notation strings representing the predicted structure, but some can also produce a base pair probability matrix detailing the likelihood of each base in an input sequence pairing to every other base in the sequence. These base pair probability matrices can be used as inputs to heuristic algorithms that can predict a dot-bracket structure. These heuristic algorithms can be used to predict pseudoknots, even when paired with input algorithms that don't consider pseudoknot presence. We use two heuristic algorithms here, Threshknot and an implementation of the Hungarian algorithm. Columns of the form eternafold[hungarian]_mfe used Eternafold to predict a base pair probability matrix, then used the hungarian algorithm to generate a dot-bracket structure.\n>\n>One more wrinkle to these data; most of these predictions were made running on a Stanford supercomputer cluster. If the algorithm is prefixed with eterna_ as in [eterna_eternafold_threshknot](https://forum.eternagame.org/t/new-folding-engine-discussion-eternafoldthreshknots/4540), the predictor was run using the version of the associated algorithm available in the [Eterna game](https://eternagame.org/labs/11627606). These should largely be the same, but it is possible that differences in software versions or heuristic implementations could lead to differences. So in your specific example, eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by threshknot, run on a Stanford supercomputer, while eterna_eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by Eterna's implementation of threshknot, run in the Eterna game.\n",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2480754": "One reason the reactivity value for most positions is not precisely 0 or 1 is that RNA molecules constantly move among two or more conformations. Some of the movements are minor but the shifts also can be significant. This is one of the aspects that makes RNA folding prediction challenging and where we believe deep learning can help.\n\nRNA researchers refer to the collection of conformations as the ensemble. The ensemble is \"read\" from the thousands of copies tested for each sequence. In a similar vein, the predicted ensemble is reflected in the bpp (base pair probability) additional files with matrices.\n\nTo answer a question about the bpp files that has been raised, yes, bpp files have been generated in [Eternafold](https://github.com/eternagame/EternaFold) for the test set. (Note that Eternafold does not model pseudoknots when calculating bpp.)\n\nThomas's explanation of the bpp additional files from [another thread](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2468751):\n\n>The column names reflect the software packages utilized, but there are some nuances to be aware of. Generally speaking, the name of the column references the software used to make the prediction. For example,vienna2_mfe is a prediction made by the ViennaRNA package. Here is the list of packages:\n>\n>ViennaRNA\n>Contrafold\n>Nupack\n>Eternafold\n>e2efold\n>hotknots\n>ipknots\n>iterative-hfold\n>knotty\n>pknots\n>spotrna\n>shapify\n>\n>Many of these prediction algorithms produce dot-bracket notation strings representing the predicted structure, but some can also produce a base pair probability matrix detailing the likelihood of each base in an input sequence pairing to every other base in the sequence. These base pair probability matrices can be used as inputs to heuristic algorithms that can predict a dot-bracket structure. These heuristic algorithms can be used to predict pseudoknots, even when paired with input algorithms that don't consider pseudoknot presence. We use two heuristic algorithms here, Threshknot and an implementation of the Hungarian algorithm. Columns of the form eternafold[hungarian]_mfe used Eternafold to predict a base pair probability matrix, then used the hungarian algorithm to generate a dot-bracket structure.\n>\n>One more wrinkle to these data; most of these predictions were made running on a Stanford supercomputer cluster. If the algorithm is prefixed with eterna_ as in [eterna_eternafold_threshknot](https://forum.eternagame.org/t/new-folding-engine-discussion-eternafoldthreshknots/4540), the predictor was run using the version of the associated algorithm available in the [Eterna game](https://eternagame.org/labs/11627606). These should largely be the same, but it is possible that differences in software versions or heuristic implementations could lead to differences. So in your specific example, eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by threshknot, run on a Stanford supercomputer, while eterna_eternafold_threshknot is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by Eterna's implementation of threshknot, run in the Eterna game.\n"
  }
}