{
  "id": 437481,
  "title": "Welcome to the Ribonanza Challenge!",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/437481",
  "author_name": "Rhiju Das",
  "post_date": "2023-09-06T22:23:46.843000",
  "votes": 37,
  "comment_count": 55,
  "views": 0,
  "content": "<p><strong>Hello RNA world</strong></p>\n<p>RNA is the basis for new medicines and the oldest forms of life —  and yet is one of the most poorly understood molecules in biology. </p>\n<p>The other major kinds of macromolecule of biology — DNA and protein —  have had predictable structures since the seminal work of Watson, Crick and Franklin in the 1950s and the AlphaFold 2 breakthrough of 2021. </p>\n<p>But we're still bad at modeling RNA.</p>\n<p>You might ask: </p>\n<p><em>How can we be so bad at RNA, especially if there are so many more different kinds of RNA’s than there are proteins in our bodies and throughout biology?</em> </p>\n<p>The answer is that the RNA’s have historically been difficult to experimentally characterize — they can form multiple 3D structures and that has rendered RNA difficult to see with conventional visualization approaches. </p>\n<p><em>What’s the prospect of RNA having an AlphaFold moment when we have 3D coordinate information on only thousands of RNA molecules, compared to hundreds of thousands of protein structure?</em></p>\n<p>Enter the new <strong>Ribonanza</strong> data set — our attempt to get enough data to finally crack the problem of RNA structure prediction.</p>\n<p>We’ve scaled up a way to get rich information of RNA sequences based on a chemical mapping approach that reads out one number per position. </p>\n<p>This number is highly sensitive to how structured the RNA is at each position, averaged over the ensemble of conformations.   </p>\n<p>And we’ve been synthesizing a diverse assortment of over 1M RNA’s, some from the expert RNA databases and some from the fabulous internet open science project Eterna!</p>\n<p>Now we need to find a model that can learn from the Ribonanza data.</p>\n<p>RNA scientists could then use the model as an oracle to understand the trillions of existing molecules for which we are missing structural information and to guide the design of new RNAs for the future of medicine and biology.</p>\n<p>That’s where you come in: <em>Help us find the oracle for RNA structure!</em></p>\n<p>We need each and every idea in data science and machine learning to be brought to this data set RNA structure prediction -- and we know that Kaggle will deliver. See you in the competition!</p>\n<p><strong>Hosts</strong></p>\n<p>Rhiju Das <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a></p>\n<p>Shujun He <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a></p>\n<p>Thomas Karagianes <a href=\"https://www.kaggle.com/brainbowrna\" target=\"_blank\">@brainbowrna</a></p>\n<p>Jill Townley <a href=\"https://www.kaggle.com/digitalembrace\" target=\"_blank\">@digitalembrace</a></p>\n<p>Rachael Kretsch <a href=\"https://www.kaggle.com/rkretsch\" target=\"_blank\">@rkretsch</a></p>\n<p><em>Special thanks to Rui Huang for developing the experimental protocols for Ribonanza!</em></p>\n<p>Thanks to members of the Eterna community and the Das laboratory, including John Nichol, Grace Nye, Christian Choe, and Jonathan Romano, for key contributions in software development and library design. </p>\n<p>And our collaborators at Kaggle:</p>\n<p>Maggie Demkin <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a><br>\nInversion <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a></p>",
  "messages": [
    {
      "id": 2426885,
      "postDate": "2023-09-06T22:23:46.843Z",
      "content": "<p><strong>Hello RNA world</strong></p>\n<p>RNA is the basis for new medicines and the oldest forms of life —  and yet is one of the most poorly understood molecules in biology. </p>\n<p>The other major kinds of macromolecule of biology — DNA and protein —  have had predictable structures since the seminal work of Watson, Crick and Franklin in the 1950s and the AlphaFold 2 breakthrough of 2021. </p>\n<p>But we're still bad at modeling RNA.</p>\n<p>You might ask: </p>\n<p><em>How can we be so bad at RNA, especially if there are so many more different kinds of RNA’s than there are proteins in our bodies and throughout biology?</em> </p>\n<p>The answer is that the RNA’s have historically been difficult to experimentally characterize — they can form multiple 3D structures and that has rendered RNA difficult to see with conventional visualization approaches. </p>\n<p><em>What’s the prospect of RNA having an AlphaFold moment when we have 3D coordinate information on only thousands of RNA molecules, compared to hundreds of thousands of protein structure?</em></p>\n<p>Enter the new <strong>Ribonanza</strong> data set — our attempt to get enough data to finally crack the problem of RNA structure prediction.</p>\n<p>We’ve scaled up a way to get rich information of RNA sequences based on a chemical mapping approach that reads out one number per position. </p>\n<p>This number is highly sensitive to how structured the RNA is at each position, averaged over the ensemble of conformations.   </p>\n<p>And we’ve been synthesizing a diverse assortment of over 1M RNA’s, some from the expert RNA databases and some from the fabulous internet open science project Eterna!</p>\n<p>Now we need to find a model that can learn from the Ribonanza data.</p>\n<p>RNA scientists could then use the model as an oracle to understand the trillions of existing molecules for which we are missing structural information and to guide the design of new RNAs for the future of medicine and biology.</p>\n<p>That’s where you come in: <em>Help us find the oracle for RNA structure!</em></p>\n<p>We need each and every idea in data science and machine learning to be brought to this data set RNA structure prediction -- and we know that Kaggle will deliver. See you in the competition!</p>\n<p><strong>Hosts</strong></p>\n<p>Rhiju Das <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a></p>\n<p>Shujun He <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a></p>\n<p>Thomas Karagianes <a href=\"https://www.kaggle.com/brainbowrna\" target=\"_blank\">@brainbowrna</a></p>\n<p>Jill Townley <a href=\"https://www.kaggle.com/digitalembrace\" target=\"_blank\">@digitalembrace</a></p>\n<p>Rachael Kretsch <a href=\"https://www.kaggle.com/rkretsch\" target=\"_blank\">@rkretsch</a></p>\n<p><em>Special thanks to Rui Huang for developing the experimental protocols for Ribonanza!</em></p>\n<p>Thanks to members of the Eterna community and the Das laboratory, including John Nichol, Grace Nye, Christian Choe, and Jonathan Romano, for key contributions in software development and library design. </p>\n<p>And our collaborators at Kaggle:</p>\n<p>Maggie Demkin <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a><br>\nInversion <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a></p>",
      "rawMarkdown": "**Hello RNA world**\n\nRNA is the basis for new medicines and the oldest forms of life —  and yet is one of the most poorly understood molecules in biology. \n\nThe other major kinds of macromolecule of biology — DNA and protein —  have had predictable structures since the seminal work of Watson, Crick and Franklin in the 1950s and the AlphaFold 2 breakthrough of 2021. \n\nBut we're still bad at modeling RNA.\n\nYou might ask: \n\n*How can we be so bad at RNA, especially if there are so many more different kinds of RNA’s than there are proteins in our bodies and throughout biology?* \n\nThe answer is that the RNA’s have historically been difficult to experimentally characterize — they can form multiple 3D structures and that has rendered RNA difficult to see with conventional visualization approaches. \n\n*What’s the prospect of RNA having an AlphaFold moment when we have 3D coordinate information on only thousands of RNA molecules, compared to hundreds of thousands of protein structure?*\n\nEnter the new **Ribonanza** data set — our attempt to get enough data to finally crack the problem of RNA structure prediction.\n\nWe’ve scaled up a way to get rich information of RNA sequences based on a chemical mapping approach that reads out one number per position. \n\nThis number is highly sensitive to how structured the RNA is at each position, averaged over the ensemble of conformations.   \n\nAnd we’ve been synthesizing a diverse assortment of over 1M RNA’s, some from the expert RNA databases and some from the fabulous internet open science project Eterna!\n\nNow we need to find a model that can learn from the Ribonanza data.\n\nRNA scientists could then use the model as an oracle to understand the trillions of existing molecules for which we are missing structural information and to guide the design of new RNAs for the future of medicine and biology.\n\nThat’s where you come in: *Help us find the oracle for RNA structure!*\n\nWe need each and every idea in data science and machine learning to be brought to this data set RNA structure prediction -- and we know that Kaggle will deliver. See you in the competition!\n\n**Hosts**\n\nRhiju Das @rhijudas\n\nShujun He @shujun717\n\nThomas Karagianes @brainbowrna\n\nJill Townley @digitalembrace\n\nRachael Kretsch @rkretsch\n\n*Special thanks to Rui Huang for developing the experimental protocols for Ribonanza!*\n\nThanks to members of the Eterna community and the Das laboratory, including John Nichol, Grace Nye, Christian Choe, and Jonathan Romano, for key contributions in software development and library design. \n\nAnd our collaborators at Kaggle:\n\nMaggie Demkin @maggiemd\nInversion @inversion\n",
      "votes": 37
    },
    {
      "id": 2479156,
      "postDate": "2023-10-12T13:08:45.403Z",
      "content": "<p>How were the reactivity errors calculated?</p>",
      "rawMarkdown": "How were the reactivity errors calculated?",
      "votes": 1,
      "replies": [
        {
          "id": 2497738,
          "postDate": "2023-10-24T20:58:07.393Z",
          "content": "<p>if you look at the <a href=\"https://github.com/Weeks-UNC/shapemapper2/blob/master/docs/analysis_steps.md\" target=\"_blank\">ShapeMapper2 </a> Under the section <a href=\"https://github.com/Weeks-UNC/shapemapper2/blob/master/docs/analysis_steps.md#reactivity-profile-calculation-and-normalization\" target=\"_blank\">Reactivity profile calculation and normalization</a><br>\nThis is where the <a href=\"https://academic.oup.com/nar/article/51/16/8744/7201944\" target=\"_blank\">DMS Paper</a> references its data gathering process.</p>",
          "rawMarkdown": "if you look at the [ShapeMapper2 ](https://github.com/Weeks-UNC/shapemapper2/blob/master/docs/analysis_steps.md) Under the section [Reactivity profile calculation and normalization](https://github.com/Weeks-UNC/shapemapper2/blob/master/docs/analysis_steps.md#reactivity-profile-calculation-and-normalization)\nThis is where the [DMS Paper](https://academic.oup.com/nar/article/51/16/8744/7201944) references its data gathering process."
        }
      ]
    },
    {
      "id": 2468230,
      "postDate": "2023-10-05T10:19:15.663Z",
      "content": "<p>Could you kindly provide additional details about the supplementary files? Specifically, I am interested in the CSV files within the supplementary_silico_predictions folder. It appears that these files have varying headers (columns). Could you please specify the software package utilized for their generation? Are the column names consistent for each algorithm used during the generation process? For instance, is 'eternafold_threshknot' the same as 'eterna_eternafold_threshknot'?</p>",
      "rawMarkdown": "Could you kindly provide additional details about the supplementary files? Specifically, I am interested in the CSV files within the supplementary_silico_predictions folder. It appears that these files have varying headers (columns). Could you please specify the software package utilized for their generation? Are the column names consistent for each algorithm used during the generation process? For instance, is 'eternafold_threshknot' the same as 'eterna_eternafold_threshknot'?",
      "votes": 1,
      "replies": [
        {
          "id": 2468751,
          "postDate": "2023-10-05T19:17:05.550Z",
          "content": "<p>This is a bit of a complicated question. The column names do reflect the software packages utilized, but there are some nuances to be aware of. Generally speaking, the name of the column references the software used to make the prediction. For example,<code>vienna2_mfe</code> is a prediction made by the <a href=\"https://www.tbi.univie.ac.at/RNA/index.html\" target=\"_blank\">ViennaRNA package</a>. Here is the list of packages:</p>\n<ul>\n<li><a href=\"https://www.tbi.univie.ac.at/RNA/index.html\" target=\"_blank\">ViennaRNA</a></li>\n<li><a href=\"http://contra.stanford.edu/contrafold/\" target=\"_blank\">Contrafold</a></li>\n<li><a href=\"https://nupack.org/\" target=\"_blank\">Nupack</a></li>\n<li><a href=\"https://github.com/eternagame/EternaFold\" target=\"_blank\">Eternafold</a></li>\n<li><a href=\"https://github.com/ml4bio/e2efold\" target=\"_blank\">e2efold</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/16199760/\" target=\"_blank\">hotknots</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/21685106/\" target=\"_blank\">ipknots</a></li>\n<li><a href=\"https://github.com/HosnaJabbari/Iterative-HFold/\" target=\"_blank\">iterative-hfold</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/29868872/\" target=\"_blank\">knotty</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/9925784/\" target=\"_blank\">pknots</a></li>\n<li><a href=\"https://github.com/jaswindersingh2/SPOT-RNA\" target=\"_blank\">spotrna</a></li>\n<li><a href=\"https://github.com/ltrinity/Shapify\" target=\"_blank\">shapify</a></li>\n</ul>\n<p>Many of these prediction algorithms produce dot-bracket notation strings representing the predicted structure, but some can also produce a base pair probability matrix detailing the likelihood of each base in an input sequence pairing to every other base in the sequence. These base pair probability matrices can be used as inputs to heuristic algorithms that can predict a dot-bracket structure. These heuristic algorithms can be used to predict pseudoknots, even when paired with input algorithms that don't consider pseudoknot presence. We use two heuristic algorithms here, Threshknot and an implementation of the Hungarian algorithm. Columns of the form <code>eternafold[hungarian]_mfe</code> used Eternafold to predict a base pair probability matrix, then used the hungarian algorithm to generate a dot-bracket structure.</p>\n<p>One more wrinkle to these data; most of these predictions were made running on a Stanford supercomputer cluster. If the algorithm is prefixed with <code>eterna_</code> as in <code>eterna_eternafold_threshknot</code>, the predictor was run using the version of the associated algorithm available in the Eterna game. These should largely be the same, but it is possible that differences in software versions or heuristic implementations could lead to differences. So in your specific example, <code>eternafold_threshknot</code> is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by threshknot, run on a Stanford supercomputer, while <code>eterna_eternafold_threshknot</code> is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by Eterna's implementation of threshknot, run in the Eterna game.</p>\n<p>Let me know if you have any other questions about those data.</p>",
          "rawMarkdown": "This is a bit of a complicated question. The column names do reflect the software packages utilized, but there are some nuances to be aware of. Generally speaking, the name of the column references the software used to make the prediction. For example,`vienna2_mfe` is a prediction made by the [ViennaRNA package](https://www.tbi.univie.ac.at/RNA/index.html). Here is the list of packages:\n- [ViennaRNA](https://www.tbi.univie.ac.at/RNA/index.html)\n- [Contrafold](http://contra.stanford.edu/contrafold/)\n- [Nupack](https://nupack.org/)\n- [Eternafold](https://github.com/eternagame/EternaFold)\n- [e2efold](https://github.com/ml4bio/e2efold)\n- [hotknots](https://pubmed.ncbi.nlm.nih.gov/16199760/)\n- [ipknots](https://pubmed.ncbi.nlm.nih.gov/21685106/)\n- [iterative-hfold](https://github.com/HosnaJabbari/Iterative-HFold/)\n- [knotty](https://pubmed.ncbi.nlm.nih.gov/29868872/)\n- [pknots](https://pubmed.ncbi.nlm.nih.gov/9925784/)\n- [spotrna](https://github.com/jaswindersingh2/SPOT-RNA)\n- [shapify](https://github.com/ltrinity/Shapify)\n\nMany of these prediction algorithms produce dot-bracket notation strings representing the predicted structure, but some can also produce a base pair probability matrix detailing the likelihood of each base in an input sequence pairing to every other base in the sequence. These base pair probability matrices can be used as inputs to heuristic algorithms that can predict a dot-bracket structure. These heuristic algorithms can be used to predict pseudoknots, even when paired with input algorithms that don't consider pseudoknot presence. We use two heuristic algorithms here, Threshknot and an implementation of the Hungarian algorithm. Columns of the form `eternafold[hungarian]_mfe` used Eternafold to predict a base pair probability matrix, then used the hungarian algorithm to generate a dot-bracket structure.\n\nOne more wrinkle to these data; most of these predictions were made running on a Stanford supercomputer cluster. If the algorithm is prefixed with `eterna_` as in `eterna_eternafold_threshknot`, the predictor was run using the version of the associated algorithm available in the Eterna game. These should largely be the same, but it is possible that differences in software versions or heuristic implementations could lead to differences. So in your specific example, `eternafold_threshknot` is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by threshknot, run on a Stanford supercomputer, while `eterna_eternafold_threshknot` is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by Eterna's implementation of threshknot, run in the Eterna game.\n\nLet me know if you have any other questions about those data.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2465344,
      "postDate": "2023-10-02T21:19:08.863Z",
      "content": "<p>We would like to invite you to a <strong>seminar series</strong> focused on <strong>RNA 3D structure</strong>. In particular, <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> will be giving the next talk on October 10th at 8am Pacific Time titled <strong>Ribonanza: big data for RNA structure prediction</strong> on <a href=\"https://tinyurl.com/casp-rna-zoom\" target=\"_blank\">zoom</a>.</p>\n<p>You can receive messages about future seminars by adding yourself to the <a href=\"https://groups.google.com/g/casp-rna-sig\" target=\"_blank\">google group</a>, check <a href=\"https://tinyurl.com/rna-sig-schedule\" target=\"_blank\">past and upcoming talks</a>, watch <a href=\"https://tinyurl.com/rna-sig-playlist\" target=\"_blank\">past talks</a>, and add events to your <a href=\"https://tinyurl.com/rna-sig-calendar\" target=\"_blank\">google</a> or <a href=\"https://tinyurl.com/rna-sig-cal-ics\" target=\"_blank\">outlook</a> calendar.         </p>\n<p>Hope you are enjoying working with chemical mapping data!<br>\nRachael</p>",
      "rawMarkdown": "We would like to invite you to a **seminar series** focused on **RNA 3D structure**. In particular, @rhijudas will be giving the next talk on October 10th at 8am Pacific Time titled **Ribonanza: big data for RNA structure prediction** on [zoom](https://tinyurl.com/casp-rna-zoom).\n\nYou can receive messages about future seminars by adding yourself to the [google group](https://groups.google.com/g/casp-rna-sig), check [past and upcoming talks](https://tinyurl.com/rna-sig-schedule), watch [past talks](https://tinyurl.com/rna-sig-playlist), and add events to your [google](https://tinyurl.com/rna-sig-calendar) or [outlook](https://tinyurl.com/rna-sig-cal-ics) calendar. \t\t\n\nHope you are enjoying working with chemical mapping data!\nRachael",
      "votes": 1
    },
    {
      "id": 2449924,
      "postDate": "2023-09-21T14:19:54.350Z",
      "content": "<p>Could you please help to resolve the following inconsistency:<br>\nIn the data description we have</p>\n<blockquote>\n  <p>We have split out 311,935 of these 1,118,513 sequences for a public test set to allow for continuous evaluation through the competition, on the Public Leaderboard. This set has been additionally filtered to ensure high signal-to-noise and read coverage (see note on SN_filter above).<br>\n  Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. </p>\n</blockquote>\n<p>i.e. the real evaluation set may be expected to be 1/4 of 311,935 = 78k, assuming the same fraction of samples passed the SN_filter criterion as in the train data, 1/4. So it gives approximately 7% of the test data given that 1,031,888 are going to be synthesized and are expected to pass SN criterion. Meanwhile, the description of the LB says </p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 2% of the test data.</p>\n</blockquote>\n<p>Is 2% a misprint, or do you either use a different SN criterion than SN_filter (which filters out more samples), or do your 311,935 test samples for some reason have a much higher fraction of low SN samples than train data? Thanks.</p>",
      "rawMarkdown": "Could you please help to resolve the following inconsistency:\nIn the data description we have\n>We have split out 311,935 of these 1,118,513 sequences for a public test set to allow for continuous evaluation through the competition, on the Public Leaderboard. This set has been additionally filtered to ensure high signal-to-noise and read coverage (see note on SN_filter above).\n>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. \n\ni.e. the real evaluation set may be expected to be 1/4 of 311,935 = 78k, assuming the same fraction of samples passed the SN_filter criterion as in the train data, 1/4. So it gives approximately 7% of the test data given that 1,031,888 are going to be synthesized and are expected to pass SN criterion. Meanwhile, the description of the LB says \n>This leaderboard is calculated with approximately 2% of the test data.\n\nIs 2% a misprint, or do you either use a different SN criterion than SN_filter (which filters out more samples), or do your 311,935 test samples for some reason have a much higher fraction of low SN samples than train data? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 2452219,
          "postDate": "2023-09-23T07:01:37.857Z",
          "content": "<p>Not sure if this helps or adds to the confusion,  I had asked a question below about the scoring of Private LB in ref to -</p>\n<blockquote>\n  <p>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins, to help ensure rigor. Furthermore, these test sequences will have lengths ranging from 207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models. As those profiles are collected, they will again be filtered for acceptable signal to noise before computing your final Private Leaderboard scores.</p>\n</blockquote>\n<p>all Test has 1343823 sequences (from the data for test sequences)<br>\nprivate Test 1031888 sequences  of these  &gt;= 207 1008000  (&gt; 207 8000 = 207 1000000)   so seems 23888 &lt; 207<br>\ndiff all-priv  311935  sequences <br>\nall test &lt; 207 335823</p>\n<blockquote>\n  <p>At the beginning of the competition, Stanford scientists have data on 1,118,513 RNA sequences of lengths ranging from 115 to 206.<br>\n  We have split out 311,935 of these 1,118,513 sequences for a public test set …</p>\n</blockquote>\n<p>So the 1,118,513 only refers to train or public test  and in train there are 806573 unique sequences  (806578 is the difference of 1,118,513 and 311,935) .</p>\n<blockquote>\n  <p>The remaining 806,578 sequences for which we have data are in train_data.csv. We note that 37,828 of the test sequences, derived from the RFAM database, are identical or near-identical to the train set; for these cases the leaderboard test set contains higher signal-to-noise measurements than the train_data and serve as a test of model ability to 'denoise' chemical mapping data.</p>\n</blockquote>\n<p>Given private test appears to have 23888 sequences not newly synthesized and &lt; 207 they could be part of the 37828 derived from RFAM database with higher signal to noise than their train counterparts and presumably not out of the 311,935 not in train.</p>\n<p>If 2% of test data is correct, it would seem like 2% of the total in all of test 1343823 maybe, but can really only be from the set &lt; 207 and the original split 311,935 of 1,118,513 sequences and possibly from the 37828 derived from RFAM database with higher signal to noise than train.  </p>\n<p>It may be that public LB is looking at a test of model ability to 'denoise' chemical mapping data so higher signal to noise, and Private LB is looking at a test of the generality of models to different length distribution.</p>",
          "rawMarkdown": "Not sure if this helps or adds to the confusion,  I had asked a question below about the scoring of Private LB in ref to -\n\n>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins, to help ensure rigor. Furthermore, these test sequences will have lengths ranging from 207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models. As those profiles are collected, they will again be filtered for acceptable signal to noise before computing your final Private Leaderboard scores.\n\nall Test has 1343823 sequences (from the data for test sequences)\nprivate Test 1031888 sequences  of these  >= 207 1008000  (> 207 8000 = 207 1000000)   so seems 23888 < 207\ndiff all-priv  311935  sequences \nall test < 207 335823\n\n> At the beginning of the competition, Stanford scientists have data on 1,118,513 RNA sequences of lengths ranging from 115 to 206.\nWe have split out 311,935 of these 1,118,513 sequences for a public test set ...\n \nSo the 1,118,513 only refers to train or public test  and in train there are 806573 unique sequences  (806578 is the difference of 1,118,513 and 311,935) .\n\n> The remaining 806,578 sequences for which we have data are in train_data.csv. We note that 37,828 of the test sequences, derived from the RFAM database, are identical or near-identical to the train set; for these cases the leaderboard test set contains higher signal-to-noise measurements than the train_data and serve as a test of model ability to 'denoise' chemical mapping data.\n\nGiven private test appears to have 23888 sequences not newly synthesized and < 207 they could be part of the 37828 derived from RFAM database with higher signal to noise than their train counterparts and presumably not out of the 311,935 not in train.\n\nIf 2% of test data is correct, it would seem like 2% of the total in all of test 1343823 maybe, but can really only be from the set < 207 and the original split 311,935 of 1,118,513 sequences and possibly from the 37828 derived from RFAM database with higher signal to noise than train.  \n\nIt may be that public LB is looking at a test of model ability to 'denoise' chemical mapping data so higher signal to noise, and Private LB is looking at a test of the generality of models to different length distribution."
        },
        {
          "id": 2454543,
          "postDate": "2023-09-24T22:49:22.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> thanks for the discussion (and for your posts in other threads!).</p>\n<p>The 2% number is confusing. </p>\n<p>It is a small number here not just because there are so many 'future' data, approximately 1M sequences. It also is influenced by the fact that we do not have specific positions (e.g., the first 26 and last ~51) for any of the available test sequences. Kaggle's MAE scoring scripts conflate the those unavailable positions in the public LB  with all of the future data positions.</p>\n<p>Then there is also the fact that you point out -- something like 4/5 of the available test sequences do not have data with acceptable signal-to-noise for inclusion in the MAE evaluation score.      </p>\n<p>For the actual private LB sequences, we don't suddenly expect to have another 98% of the data to actually score models, as seems to be implied by the 2% number. </p>\n<p>Instead, with our current protocols, we are probing 1M sequences, and conservatively expect to have roughly another 200,000 sequences with acceptable signal-to-noise for the private LB, and we'll be missing the initial and final positions.  The number of rows that will be used for scoring in the private LB will still end up being bigger than the public LB, but by a factor of 4x, not by a factor of 50x.  It would probably be better for the \"approximately 2%\" number to be set as \"approximately 20%\". </p>\n<p>Just one more caveat:  we are working on an experimental advance that may significantly enhance signal-to-noise, by taking advantage of an upgrade in Illumina sequencing technology that is literally becoming available in October  2023. So the above numbers may change – in the direction of allowing even more robust scoring of your submissions.</p>",
          "rawMarkdown": "@Iafoss @something4kag thanks for the discussion (and for your posts in other threads!).\n \nThe 2% number is confusing. \n\nIt is a small number here not just because there are so many 'future' data, approximately 1M sequences. It also is influenced by the fact that we do not have specific positions (e.g., the first 26 and last ~51) for any of the available test sequences. Kaggle's MAE scoring scripts conflate the those unavailable positions in the public LB  with all of the future data positions.\n\nThen there is also the fact that you point out -- something like 4/5 of the available test sequences do not have data with acceptable signal-to-noise for inclusion in the MAE evaluation score.      \n\nFor the actual private LB sequences, we don't suddenly expect to have another 98% of the data to actually score models, as seems to be implied by the 2% number. \n\nInstead, with our current protocols, we are probing 1M sequences, and conservatively expect to have roughly another 200,000 sequences with acceptable signal-to-noise for the private LB, and we'll be missing the initial and final positions.  The number of rows that will be used for scoring in the private LB will still end up being bigger than the public LB, but by a factor of 4x, not by a factor of 50x.  It would probably be better for the \"approximately 2%\" number to be set as \"approximately 20%\". \n\nJust one more caveat:  we are working on an experimental advance that may significantly enhance signal-to-noise, by taking advantage of an upgrade in Illumina sequencing technology that is literally becoming available in October ~~2024~~ 2023. So the above numbers may change – in the direction of allowing even more robust scoring of your submissions.\n\n\n",
          "votes": 2,
          "replies": [
            {
              "id": 2454570,
              "postDate": "2023-09-24T23:04:23.767Z",
              "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> for your detailed explanation. It is clear now.</p>",
              "rawMarkdown": "Thank you so much @rhijudas for your detailed explanation. It is clear now."
            },
            {
              "id": 2454926,
              "postDate": "2023-09-25T07:20:00.727Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> for this reply.  </p>\n<blockquote>\n  <p>an upgrade in Illumina sequencing technology that is literally becoming available in October 2024</p>\n</blockquote>\n<p>is that October 2023?  Hopefully the experimental advance is successful and \"more robust scoring\" is a good thing!</p>",
              "rawMarkdown": "Thanks @rhijudas for this reply.  \n>an upgrade in Illumina sequencing technology that is literally becoming available in October 2024\n\nis that October 2023?  Hopefully the experimental advance is successful and \"more robust scoring\" is a good thing!\n"
            },
            {
              "id": 2455683,
              "postDate": "2023-09-25T16:36:12.333Z",
              "content": "<p>Oops, fixed. 😊</p>",
              "rawMarkdown": "Oops, fixed. 😊",
              "votes": 1
            },
            {
              "id": 2455992,
              "postDate": "2023-09-25T21:06:21.960Z",
              "content": "<p>Update: we've updated the blurb on leaderboard to say:</p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.</p>\n</blockquote>\n<p>Thanks for the discussion!</p>",
              "rawMarkdown": "Update: we've updated the blurb on leaderboard to say:\n\n>This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.\n\nThanks for the discussion!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2440109,
      "postDate": "2023-09-15T09:15:53.977Z",
      "content": "<p>Hello, I'm a postdoc in Germany, working in an academic research institute. I am considering participating to the challenge, but would be keen to publishing my method if proven successful (in the context of this competition, but also considering other aspects). Would this infringe the rules of the competition? </p>",
      "rawMarkdown": "Hello, I'm a postdoc in Germany, working in an academic research institute. I am considering participating to the challenge, but would be keen to publishing my method if proven successful (in the context of this competition, but also considering other aspects). Would this infringe the rules of the competition? ",
      "votes": 1,
      "replies": [
        {
          "id": 2441024,
          "postDate": "2023-09-16T00:50:30.313Z",
          "content": "<p>That should be fine. The winning algorithms and the dataset will be released as open source after the competition ends. The goal of this competition is better models and robust research!</p>",
          "rawMarkdown": "That should be fine. The winning algorithms and the dataset will be released as open source after the competition ends. The goal of this competition is better models and robust research!",
          "votes": 1,
          "replies": [
            {
              "id": 2499312,
              "postDate": "2023-10-26T01:05:15.943Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2431149,
      "postDate": "2023-09-09T20:44:33.680Z",
      "content": "<p>Thanks for putting this competition together. I'm excited to learn.</p>\n<p>Small note. When describing the columns within the train/test files. You provide a reference to DMS/2A3, however the 2A3 link returns 404. The section is below:</p>\n<blockquote>\n  <p>experiment_type - (string) Either DMS_MaP or 2A3_MaP to describe the type of chemical mapping experiment that was used to generate each profile. References: DMS, 2A3.</p>\n</blockquote>\n<p>The link for 2A3 points to --&gt; <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/10.1093/nar/gkaa1255\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/10.1093/nar/gkaa1255</a></p>\n<p>I assume it is supposed to point here: <a href=\"https://academic.oup.com/nar/article/49/6/e34/6062772\" target=\"_blank\">https://academic.oup.com/nar/article/49/6/e34/6062772</a></p>\n<p><strong>A novel SHAPE reagent enables the analysis of RNA structure in living cells with unprecedented accuracy</strong></p>\n<blockquote>\n  <p><strong>Abstract:</strong><br>\n  Due to the mounting evidence that RNA structure plays a critical role in regulating almost any physiological as well as pathological process, being able to accurately define the folding of RNA molecules within living cells has become a crucial need. We introduce here 2-aminopyridine-3-carboxylic acid imidazolide (2A3), as a general probe for the interrogation of RNA structures in vivo. 2A3 shows moderate improvements with respect to the state-of-the-art selective 2′-hydroxyl acylation analyzed by primer extension (SHAPE) reagent NAI on naked RNA under in vitro conditions, but it significantly outperforms NAI when probing RNA structure in vivo, particularly in bacteria, underlining its increased ability to permeate biological membranes. When used as a restraint to drive RNA structure prediction, data derived by SHAPE-MaP with 2A3 yields more accurate predictions than NAI-derived data. Due to its extreme efficiency and accuracy, we can anticipate that 2A3 will rapidly take over conventional SHAPE reagents for probing RNA structures both in vitro and in vivo.</p>\n</blockquote>\n<p>Thanks again!</p>",
      "rawMarkdown": "Thanks for putting this competition together. I'm excited to learn.\n\nSmall note. When describing the columns within the train/test files. You provide a reference to DMS/2A3, however the 2A3 link returns 404. The section is below:\n\n> experiment_type - (string) Either DMS_MaP or 2A3_MaP to describe the type of chemical mapping experiment that was used to generate each profile. References: DMS, 2A3.\n\nThe link for 2A3 points to --> https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/10.1093/nar/gkaa1255\n\nI assume it is supposed to point here: https://academic.oup.com/nar/article/49/6/e34/6062772\n\n**A novel SHAPE reagent enables the analysis of RNA structure in living cells with unprecedented accuracy**\n> **Abstract:**\n> Due to the mounting evidence that RNA structure plays a critical role in regulating almost any physiological as well as pathological process, being able to accurately define the folding of RNA molecules within living cells has become a crucial need. We introduce here 2-aminopyridine-3-carboxylic acid imidazolide (2A3), as a general probe for the interrogation of RNA structures in vivo. 2A3 shows moderate improvements with respect to the state-of-the-art selective 2′-hydroxyl acylation analyzed by primer extension (SHAPE) reagent NAI on naked RNA under in vitro conditions, but it significantly outperforms NAI when probing RNA structure in vivo, particularly in bacteria, underlining its increased ability to permeate biological membranes. When used as a restraint to drive RNA structure prediction, data derived by SHAPE-MaP with 2A3 yields more accurate predictions than NAI-derived data. Due to its extreme efficiency and accuracy, we can anticipate that 2A3 will rapidly take over conventional SHAPE reagents for probing RNA structures both in vitro and in vivo.\n\nThanks again!",
      "votes": 2,
      "replies": [
        {
          "id": 2432153,
          "postDate": "2023-09-10T16:05:37.397Z",
          "content": "<p>Fixed the reference. Thanks for catching!</p>",
          "rawMarkdown": "Fixed the reference. Thanks for catching!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2539329,
      "postDate": "2023-11-27T00:58:57.330Z",
      "content": "<p>I just finished my data science diploma and, though ambitious, I want to use this competition as a personal project to improve my skills. Therefore I wanted to ask if the data will remain available after the final deadline</p>",
      "rawMarkdown": "I just finished my data science diploma and, though ambitious, I want to use this competition as a personal project to improve my skills. Therefore I wanted to ask if the data will remain available after the final deadline"
    },
    {
      "id": 2506798,
      "postDate": "2023-10-31T14:53:47.187Z",
      "content": "<p>Is the submitted submissions final with all the necessary information, includes all the test data, or is it a re-run of the kernels/models after the deadline? </p>",
      "rawMarkdown": "Is the submitted submissions final with all the necessary information, includes all the test data, or is it a re-run of the kernels/models after the deadline? "
    },
    {
      "id": 2505734,
      "postDate": "2023-10-30T18:52:04.010Z",
      "content": "<p>Hello! My understanding is correct that we have a set of RNA molecules with a known primary and secondary structure. We must predict the shape of a molecule in three dimensions (tertiary structure), which is determined by the degree of response to certain chemical influences (reactivity_0001, reactivity_0002,…)?</p>",
      "rawMarkdown": "Hello! My understanding is correct that we have a set of RNA molecules with a known primary and secondary structure. We must predict the shape of a molecule in three dimensions (tertiary structure), which is determined by the degree of response to certain chemical influences (reactivity_0001, reactivity_0002,...)?",
      "replies": [
        {
          "id": 2505774,
          "postDate": "2023-10-30T19:52:40.100Z",
          "content": "<p>Not quite - you have a known sequence (primary structure) and must predict the reactivity. The reactivity prediction can later be used to more accurately infer secondary and even tertiary structure from the sequence, but that process is outside the scope of this competition. There are existing methods (with mixed accuracy) for predicting a secondary structure from the sequence which you may wish to use as part of your solution.</p>",
          "rawMarkdown": "Not quite - you have a known sequence (primary structure) and must predict the reactivity. The reactivity prediction can later be used to more accurately infer secondary and even tertiary structure from the sequence, but that process is outside the scope of this competition. There are existing methods (with mixed accuracy) for predicting a secondary structure from the sequence which you may wish to use as part of your solution.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2500676,
      "postDate": "2023-10-26T23:07:18.717Z",
      "content": "<p>I have a question about the rules. I would like to participate in this competition and make a submission as part of a project I'm doing for a class. In this class, at the end of the semester (by December 4, which is before the competition ends), I will submit a project report to the instructors of my class. Does this (submitting my report privately to the instructors of my class) constitute the infringement of the rules of the competition or not?<br>\nThe report won’t contain any code nor data (it will contain only ideas and results).<br>\nI read the rules, and it seemed to me that I’m okay so long as I’m not sharing privately my code or data. But I would like to confirm. (I’m new to Kaggle.)</p>",
      "rawMarkdown": "I have a question about the rules. I would like to participate in this competition and make a submission as part of a project I'm doing for a class. In this class, at the end of the semester (by December 4, which is before the competition ends), I will submit a project report to the instructors of my class. Does this (submitting my report privately to the instructors of my class) constitute the infringement of the rules of the competition or not?\nThe report won’t contain any code nor data (it will contain only ideas and results).\nI read the rules, and it seemed to me that I’m okay so long as I’m not sharing privately my code or data. But I would like to confirm. (I’m new to Kaggle.)",
      "replies": [
        {
          "id": 2500740,
          "postDate": "2023-10-27T01:42:10.600Z",
          "content": "<p>Yes, that will be fine. Good luck on your project!</p>",
          "rawMarkdown": "Yes, that will be fine. Good luck on your project!",
          "votes": 1,
          "replies": [
            {
              "id": 2500796,
              "postDate": "2023-10-27T02:57:45.477Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            },
            {
              "id": 2554154,
              "postDate": "2023-12-08T20:29:05.513Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2496176,
      "postDate": "2023-10-23T19:32:38.493Z",
      "content": "<p>Hi, i'm new to the competition. Can we use the test_sequences for pretraining or is that against the competition rules? </p>",
      "rawMarkdown": "Hi, i'm new to the competition. Can we use the test_sequences for pretraining or is that against the competition rules? ",
      "replies": [
        {
          "id": 2497656,
          "postDate": "2023-10-24T18:38:57.580Z",
          "content": "<p>You are welcome to use test_sequences  for pre training. Good luck!</p>",
          "rawMarkdown": "You are welcome to use test_sequences  for pre training. Good luck!"
        }
      ]
    },
    {
      "id": 2486485,
      "postDate": "2023-10-18T01:05:49Z",
      "content": "<p>Regarding the following statement on the overview</p>\n<blockquote>\n  <p>From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.</p>\n</blockquote>\n<p>What are those two ground truth values. Is this the reactivity columns? Does it include the error? Is the ground truth the two mapping experiments?</p>\n<p>If so the wording is off because the first 26 values are not available due to technical reasons. Also not sure if same sequence has both mapping experiments.</p>",
      "rawMarkdown": "Regarding the following statement on the overview\n\n\n>From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.\n\nWhat are those two ground truth values. Is this the reactivity columns? Does it include the error? Is the ground truth the two mapping experiments?\n\nIf so the wording is off because the first 26 values are not available due to technical reasons. Also not sure if same sequence has both mapping experiments.\n",
      "replies": [
        {
          "id": 2497661,
          "postDate": "2023-10-24T18:43:35.647Z",
          "content": "<p>Yes the reactivity columns are the values you are trying to predict. </p>\n<p>In the train data, there are separate error columns that give estimates of the experimental uncertainties  (precision) on the measured reactivity values - you are welcome to ignore those estimated errors or to try to take them into account to down weight high error positions or sequences during training.</p>\n<p>For the first 26 values they are not in the train data but may be available in private leaderboard test data - it will be an interesting challenge for models to generalize to this part of the sequence or to longer lengths - you can look through some of the other discussion threads to look for ideas and tests of generalization. Thanks for the questions!</p>",
          "rawMarkdown": "Yes the reactivity columns are the values you are trying to predict. \n\nIn the train data, there are separate error columns that give estimates of the experimental uncertainties  (precision) on the measured reactivity values - you are welcome to ignore those estimated errors or to try to take them into account to down weight high error positions or sequences during training.\n\nFor the first 26 values they are not in the train data but may be available in private leaderboard test data - it will be an interesting challenge for models to generalize to this part of the sequence or to longer lengths - you can look through some of the other discussion threads to look for ideas and tests of generalization. Thanks for the questions!"
        }
      ]
    },
    {
      "id": 2479981,
      "postDate": "2023-10-13T03:40:58.393Z",
      "content": "<p>Just want to know what 'nts' means in the signal column \"mean( measurement value over probed nts )/mean( statistical error in measurement value over probed nts)\"</p>",
      "rawMarkdown": "Just want to know what 'nts' means in the signal column \"mean( measurement value over probed nts )/mean( statistical error in measurement value over probed nts)\"",
      "replies": [
        {
          "id": 2480680,
          "postDate": "2023-10-13T13:26:50.037Z",
          "content": "<p>'nts' is an abbreviation for nucleotides.</p>",
          "rawMarkdown": "'nts' is an abbreviation for nucleotides.",
          "votes": 1,
          "replies": [
            {
              "id": 2483764,
              "postDate": "2023-10-16T01:10:11.087Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2468821,
      "postDate": "2023-10-05T21:00:32.260Z",
      "content": "<p>This is great project to predict RNA molecular structure! I can't believe that the RNA molecular structures are still not well understood even though there are tons of RNA drugs! Thank you for y'all work and I will try my best to make a good prediction model! DDS forever!!</p>",
      "rawMarkdown": "This is great project to predict RNA molecular structure! I can't believe that the RNA molecular structures are still not well understood even though there are tons of RNA drugs! Thank you for y'all work and I will try my best to make a good prediction model! DDS forever!!"
    },
    {
      "id": 2467430,
      "postDate": "2023-10-04T14:48:07.203Z",
      "content": "<p>For file input and path setup, it takes several minutes in kaggle notebook. what should I do?</p>",
      "rawMarkdown": "For file input and path setup, it takes several minutes in kaggle notebook. what should I do?\n"
    },
    {
      "id": 2453698,
      "postDate": "2023-09-24T09:03:28.507Z",
      "content": "<p>Some of chemical softwares are free but have the following licenses:</p>\n<ul>\n<li>GPL / LGPL</li>\n<li>Non-Commercial Academic Use Only</li>\n</ul>\n<p>Is it possible to check the availability of these software?</p>",
      "rawMarkdown": "Some of chemical softwares are free but have the following licenses:\n\n- GPL / LGPL\n- Non-Commercial Academic Use Only\n\nIs it possible to check the availability of these software?",
      "replies": [
        {
          "id": 2454557,
          "postDate": "2023-09-24T22:58:10.733Z",
          "content": "<p>You are welcome to use those software packages to prepare your submissions!</p>\n<p>If your submission becomes prize eligible, you will need to make your <em>own</em> code available under a fully open source license. </p>\n<p>It will be fine at that time to simply include links to any codebases that you used that are licensed as GPL/LGPL or for non-commercial use.  If you have had to make updates to those prior codebases, we might ask you to provide a diff of your updated version to the prior codebase along with your software. </p>",
          "rawMarkdown": "You are welcome to use those software packages to prepare your submissions!\n\nIf your submission becomes prize eligible, you will need to make your *own* code available under a fully open source license. \n\nIt will be fine at that time to simply include links to any codebases that you used that are licensed as GPL/LGPL or for non-commercial use.  If you have had to make updates to those prior codebases, we might ask you to provide a diff of your updated version to the prior codebase along with your software. ",
          "votes": 1,
          "replies": [
            {
              "id": 2454601,
              "postDate": "2023-09-24T23:52:51.607Z",
              "content": "<p>Thanks for confirmation!</p>",
              "rawMarkdown": "Thanks for confirmation!"
            }
          ]
        }
      ]
    },
    {
      "id": 2452594,
      "postDate": "2023-09-23T12:35:24.403Z",
      "content": "<p>Hello Ribonanza Challenge Team,</p>\n<p>I'm excited to see this initiative aiming to tackle the complex problem of RNA structure prediction. RNA's significance in biology, especially its role in medicine and understanding the origins of life, cannot be overstated.</p>\n<p>It's impressive to see the scale of data collection efforts and the collaboration with both expert RNA databases and open science projects like Eterna. This approach promises to provide a wealth of valuable information to researchers.</p>\n<p>I'm looking forward to the competition and the innovative solutions that the data science and machine learning community will bring to the table. This is an opportunity to make a significant impact on our understanding of RNA and its applications in medicine and biology.</p>\n<p>Count me in, and let's work together to find the oracle for RNA structure! 🧬💻 #RibonanzaChallenge</p>\n<p>Best regards,<br>\nBhavesh Padharia</p>",
      "rawMarkdown": "Hello Ribonanza Challenge Team,\n\nI'm excited to see this initiative aiming to tackle the complex problem of RNA structure prediction. RNA's significance in biology, especially its role in medicine and understanding the origins of life, cannot be overstated.\n\nIt's impressive to see the scale of data collection efforts and the collaboration with both expert RNA databases and open science projects like Eterna. This approach promises to provide a wealth of valuable information to researchers.\n\nI'm looking forward to the competition and the innovative solutions that the data science and machine learning community will bring to the table. This is an opportunity to make a significant impact on our understanding of RNA and its applications in medicine and biology.\n\nCount me in, and let's work together to find the oracle for RNA structure! 🧬💻 #RibonanzaChallenge\n\nBest regards,\nBhavesh Padharia"
    },
    {
      "id": 2448513,
      "postDate": "2023-09-20T16:19:08.327Z",
      "content": "<p>In relation to the <strong>base-pairing probabilities</strong>, what is the difference between these two things?</p>\n<ul>\n<li>The files total 27 Go. <strong>Ribonanza_bpp_files</strong> (provided as data), which give the BPPs in txt format for the train and test sequences/ this format: position of nucleotide a, position of nucleotide b, probability</li>\n<li>And the possibility of calculating them (these BPPs) using your pinned notebook: <a href=\"https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment\" target=\"_blank\">RNA Science Computational Environment</a>, I mean arnie library (and package = \"<strong>eternafold</strong>\")<br>\n<code>from arnie.bpps import bpps</code><br>\n<code>bpps(sequence,package=\"eternafold\")</code></li>\n</ul>\n<p>In my understanding, with the exception of their format, <strong>which is not a matrix</strong>, the txt files provided as data (Ribonanza_bpp_files) are supposed to accomplish the same thing (in the data section: you define them : \"a TXT file of base pair probabilities from the LinearPartition-<strong>EternaFold</strong> package\").&nbsp;&nbsp;</p>",
      "rawMarkdown": "In relation to the **base-pairing probabilities**, what is the difference between these two things?\n- The files total 27 Go. **Ribonanza_bpp_files** (provided as data), which give the BPPs in txt format for the train and test sequences/ this format: position of nucleotide a, position of nucleotide b, probability\n- And the possibility of calculating them (these BPPs) using your pinned notebook: [RNA Science Computational Environment](https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment), I mean arnie library (and package = \"**eternafold**\")\n`from arnie.bpps import bpps`\n`bpps(sequence,package=\"eternafold\") `\n\nIn my understanding, with the exception of their format, **which is not a matrix**, the txt files provided as data (Ribonanza_bpp_files) are supposed to accomplish the same thing (in the data section: you define them : \"a TXT file of base pair probabilities from the LinearPartition-**EternaFold** package\").  ",
      "replies": [
        {
          "id": 2454576,
          "postDate": "2023-09-24T23:05:48.540Z",
          "content": "<p>Good question! The <strong>Ribonanza_bpp_files</strong> are not matrices. For conciseness, these files only list pairs of nucleotides which have non-zero base pair probabilities. I have updated the data description. Thanks for pointing this out.</p>",
          "rawMarkdown": "Good question! The **Ribonanza_bpp_files** are not matrices. For conciseness, these files only list pairs of nucleotides which have non-zero base pair probabilities. I have updated the data description. Thanks for pointing this out.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2438618,
      "postDate": "2023-09-14T11:43:16.927Z",
      "content": "<p>Hey,<br>\nMy name is Miss Mailiana, I am an entry level professional in the field of Data Analytics. I am a recent graduate with BA in Statistics with Economics and Mathematics, I’m ready to team up to share my knowledge and build up more to win the prize.<br>\nPlease reach out </p>",
      "rawMarkdown": "Hey,\nMy name is Miss Mailiana, I am an entry level professional in the field of Data Analytics. I am a recent graduate with BA in Statistics with Economics and Mathematics, I’m ready to team up to share my knowledge and build up more to win the prize.\nPlease reach out ",
      "replies": [
        {
          "id": 2439922,
          "postDate": "2023-09-15T06:52:01.470Z",
          "content": "<p>You may want to post this in <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437707\" target=\"_blank\">Looking for a Team Megathread</a>  Use this thread to find a teammate if you're interested in finding others to work with!</p>",
          "rawMarkdown": "You may want to post this in [Looking for a Team Megathread](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437707)  Use this thread to find a teammate if you're interested in finding others to work with!"
        }
      ]
    },
    {
      "id": 2436090,
      "postDate": "2023-09-13T10:40:49.553Z",
      "content": "<p>Noticed this does not seem to be a code competition and wondered how the Private LB and private test set will work - </p>\n<blockquote>\n  <p>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins, to help ensure rigor. Furthermore, these test sequences will have lengths ranging from 207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models. As those profiles are collected, they will again be filtered for acceptable signal to noise before computing your final Private Leaderboard scores.</p>\n</blockquote>\n<p>Know it is early in the competition so maybe still being determined?  e.g., will there be a date when it  is released so that new submissions will need to be done, or is it that the selected submission notebooks will be rerun after competition end?  </p>",
      "rawMarkdown": "Noticed this does not seem to be a code competition and wondered how the Private LB and private test set will work - \n\n>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins, to help ensure rigor. Furthermore, these test sequences will have lengths ranging from 207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models. As those profiles are collected, they will again be filtered for acceptable signal to noise before computing your final Private Leaderboard scores.\n\nKnow it is early in the competition so maybe still being determined?  e.g., will there be a date when it  is released so that new submissions will need to be done, or is it that the selected submission notebooks will be rerun after competition end?  ",
      "replies": [
        {
          "id": 2441007,
          "postDate": "2023-09-15T23:31:41.867Z",
          "content": "<p>After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for <strong>private leaderboard</strong>* test sequences for which the submissions already include predictions, so need for you to prepare new submissions or provide notebooks. Thanks for asking the question!</p>\n<p>*Edited after discussion below.</p>",
          "rawMarkdown": "After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for **private leaderboard*** test sequences for which the submissions already include predictions, so need for you to prepare new submissions or provide notebooks. Thanks for asking the question!\n\n*Edited after discussion below.",
          "votes": 1,
          "replies": [
            {
              "id": 2441425,
              "postDate": "2023-09-16T07:28:02.527Z",
              "content": "<p>Thanks for the reply!  </p>\n<blockquote>\n  <p>these test sequences will have lengths ranging from 207 to 457 bases</p>\n</blockquote>\n<p>Although there is new data, that refers to what will affect the solution for scoring, and the data in the sequence column for test_sequences will not change. Good to know!</p>",
              "rawMarkdown": "Thanks for the reply!  \n\n>these test sequences will have lengths ranging from 207 to 457 bases\n\nAlthough there is new data, that refers to what will affect the solution for scoring, and the data in the sequence column for test_sequences will not change. Good to know!\n"
            },
            {
              "id": 2452579,
              "postDate": "2023-09-23T12:24:25.263Z",
              "content": "<blockquote>\n  <p>After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for test sequences for which the submissions already include predictions, so need for you to prepare new submissions or provide notebooks. </p>\n</blockquote>\n<p>My understanding from the context and also according to chatGPT there is \"no need to prepare new submissions.\"  Based on this, do I follow correctly?</p>\n<ol>\n<li>A sequence X with  N bases will be re-synthesized, but M bases will be added, thus if N = 200 and M = 100, the new Y seq will have 300 bases.</li>\n<li>Reactivity for Y will be measured</li>\n<li>Data will be clipped back to the first N bases/reactivities.</li>\n</ol>\n<blockquote>\n  <p>207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models.</p>\n</blockquote>\n<p>In the original seq X, let's say X_reactivity_0060 = 0.95. But base on the position 260 in the seq Y may affect Y_reactivity_0060 in Y such that it is only 0.05. </p>\n<p>Therefore, I don't understand how it tests the generality of the model, nor how it makes sense at all, as any model would need the knowledge about a base in position 260 to correctly predict 160.</p>\n<p>I guess I badly misunderstood something.</p>",
              "rawMarkdown": ">After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for test sequences for which the submissions already include predictions, so need for you to prepare new submissions or provide notebooks. \n\nMy understanding from the context and also according to chatGPT there is \"no need to prepare new submissions.\"  Based on this, do I follow correctly?\n\n1. A sequence X with  N bases will be re-synthesized, but M bases will be added, thus if N = 200 and M = 100, the new Y seq will have 300 bases.\n2. Reactivity for Y will be measured\n3. Data will be clipped back to the first N bases/reactivities.\n\n>207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models.\n\nIn the original seq X, let's say X_reactivity_0060 = 0.95. But base on the position 260 in the seq Y may affect Y_reactivity_0060 in Y such that it is only 0.05. \n\nTherefore, I don't understand how it tests the generality of the model, nor how it makes sense at all, as any model would need the knowledge about a base in position 260 to correctly predict 160.\n\nI guess I badly misunderstood something."
            },
            {
              "id": 2454596,
              "postDate": "2023-09-24T23:45:50.950Z",
              "content": "<p>Thanks for the question. </p>\n<p>There are numerous sequences in <code>test_sequences.csv</code> for which we are asking for submissions. But we don't have the data to evaluate your submissions for those sequences at the start of the competition.  </p>\n<p>These are literally sequences that we have not synthesized before, so there won't be any re-synthesis -- we'll synthesize the RNA for the first time in the upcoming months.  </p>\n<p>Also there will not be any clipping back to the first N bases. </p>\n<p>What might be confusing here is that the <code>train_data.csv</code> only has columns going to <code>reactivity_0206</code>. However, we are asking for predictions for test sequences with lengths up to 457. If you look in <code>test_sequences.csv</code>, you'll see that those long sequences will actually correspond to 457 rows that you need to fill in <code>sample_submission.csv</code>. </p>\n<p>(In retrospect, we could have included columns up to <code>reactivity_0457</code> in <code>train_data.csv</code>, filled with blanks!)</p>\n<p>Good luck with your submissions!</p>",
              "rawMarkdown": "Thanks for the question. \n\nThere are numerous sequences in `test_sequences.csv` for which we are asking for submissions. But we don't have the data to evaluate your submissions for those sequences at the start of the competition.  \n\nThese are literally sequences that we have not synthesized before, so there won't be any re-synthesis -- we'll synthesize the RNA for the first time in the upcoming months.  \n\nAlso there will not be any clipping back to the first N bases. \n\n What might be confusing here is that the `train_data.csv` only has columns going to `reactivity_0206`. However, we are asking for predictions for test sequences with lengths up to 457. If you look in `test_sequences.csv`, you'll see that those long sequences will actually correspond to 457 rows that you need to fill in `sample_submission.csv`. \n\n(In retrospect, we could have included columns up to `reactivity_0457` in `train_data.csv`, filled with blanks!)\n\nGood luck with your submissions!\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2432476,
      "postDate": "2023-09-10T21:26:42.423Z",
      "content": "<p>Hello, this might be an obvious question, but from the provided files I don't seem to find one that provides actual results, i.e the DMS_MaP and 2A3_MaP for any sequence. I understand that for the final score we can not have access to these, but without some training data with the target how can we train models?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hello, this might be an obvious question, but from the provided files I don't seem to find one that provides actual results, i.e the DMS_MaP and 2A3_MaP for any sequence. I understand that for the final score we can not have access to these, but without some training data with the target how can we train models?\n\nThanks",
      "replies": [
        {
          "id": 2433586,
          "postDate": "2023-09-11T17:05:40.523Z",
          "content": "<p>The measured values are found in the reactivity_0001 through reactivity_0170 columns in the training data (as well as reactivity error columns which provide experimental error). Each row of training data will either be providing results for either DMS or 2A3, as specified by the experiment_type column. When you construct your submission file, it will have one line per sequence position and both DMS and 2A3 data rather than per sequence with only one or the other</p>",
          "rawMarkdown": "The measured values are found in the reactivity_0001 through reactivity_0170 columns in the training data (as well as reactivity error columns which provide experimental error). Each row of training data will either be providing results for either DMS or 2A3, as specified by the experiment_type column. When you construct your submission file, it will have one line per sequence position and both DMS and 2A3 data rather than per sequence with only one or the other",
          "replies": [
            {
              "id": 2434994,
              "postDate": "2023-09-12T16:25:40.843Z",
              "content": "<p>CORRECTION: reactivity_0001 through reactivity_0206 (not 170) - I misread the available columns when I was looking at the training data</p>",
              "rawMarkdown": "CORRECTION: reactivity_0001 through reactivity_0206 (not 170) - I misread the available columns when I was looking at the training data"
            },
            {
              "id": 2450331,
              "postDate": "2023-09-21T19:01:34.173Z",
              "content": "<p>I think you haven't misread, they were mis-displayed. I was just puzzled a minute ago as to why in the viewer (Data tab) there are only 170 reactivity but still 206 error columns. It seems some columns might not be displayed sometimes, worth keeping at the back of mind.</p>",
              "rawMarkdown": "I think you haven't misread, they were mis-displayed. I was just puzzled a minute ago as to why in the viewer (Data tab) there are only 170 reactivity but still 206 error columns. It seems some columns might not be displayed sometimes, worth keeping at the back of mind."
            },
            {
              "id": 2450434,
              "postDate": "2023-09-21T21:03:14.983Z",
              "content": "<p>I think my issue was reactivity 1-170 were shown, then it switched to error, then it went back to reactivity 171-206</p>",
              "rawMarkdown": "I think my issue was reactivity 1-170 were shown, then it switched to error, then it went back to reactivity 171-206"
            }
          ]
        }
      ]
    },
    {
      "id": 2431548,
      "postDate": "2023-09-10T07:34:24.597Z",
      "content": "<p>Can you please provide more clarity on the below sentence  in bold fromthe data description in simple words for non-domain users like me?<br>\n\"In this competition, you will be predicting the reactivity of an RNA sequence to two chemical modifiers DMS and 2A3. <br>\n<strong>These data can be measured efficiently through a mutational profiling (MaP) experiment read out by high-throughput sequencing and positions that are protected from chemical modification are likely to be forming base pairs or other kinds of RNA structure</strong>.\"</p>\n<p>Thanks,<br>\nKiran</p>",
      "rawMarkdown": "Can you please provide more clarity on the below sentence  in bold fromthe data description in simple words for non-domain users like me?\n\"In this competition, you will be predicting the reactivity of an RNA sequence to two chemical modifiers DMS and 2A3. \n**These data can be measured efficiently through a mutational profiling (MaP) experiment read out by high-throughput sequencing and positions that are protected from chemical modification are likely to be forming base pairs or other kinds of RNA structure**.\"\n\nThanks,\nKiran",
      "replies": [
        {
          "id": 2472093,
          "postDate": "2023-10-07T01:49:25.097Z",
          "content": "<p>Hi! For discussion of the experimental protocol you might want to check out this thread: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/445415\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/445415</a></p>",
          "rawMarkdown": "Hi! For discussion of the experimental protocol you might want to check out this thread: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/445415"
        }
      ]
    },
    {
      "id": 2455993,
      "postDate": "2023-09-25T21:09:34.977Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2479156,
      "author_name": "Ramon Viñas",
      "author_url": "",
      "post_date": "2023-10-12T13:08:45.403000",
      "content": "<p>How were the reactivity errors calculated?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2497738,
          "author_name": "NathanTuttle",
          "author_url": "",
          "post_date": "2023-10-24T20:58:07.393000",
          "content": "<p>if you look at the <a href=\"https://github.com/Weeks-UNC/shapemapper2/blob/master/docs/analysis_steps.md\" target=\"_blank\">ShapeMapper2 </a> Under the section <a href=\"https://github.com/Weeks-UNC/shapemapper2/blob/master/docs/analysis_steps.md#reactivity-profile-calculation-and-normalization\" target=\"_blank\">Reactivity profile calculation and normalization</a><br>\nThis is where the <a href=\"https://academic.oup.com/nar/article/51/16/8744/7201944\" target=\"_blank\">DMS Paper</a> references its data gathering process.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2468230,
      "author_name": "Andrij David",
      "author_url": "",
      "post_date": "2023-10-05T10:19:15.663000",
      "content": "<p>Could you kindly provide additional details about the supplementary files? Specifically, I am interested in the CSV files within the supplementary_silico_predictions folder. It appears that these files have varying headers (columns). Could you please specify the software package utilized for their generation? Are the column names consistent for each algorithm used during the generation process? For instance, is 'eternafold_threshknot' the same as 'eterna_eternafold_threshknot'?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2468751,
          "author_name": "Thomas",
          "author_url": "",
          "post_date": "2023-10-05T19:17:05.550000",
          "content": "<p>This is a bit of a complicated question. The column names do reflect the software packages utilized, but there are some nuances to be aware of. Generally speaking, the name of the column references the software used to make the prediction. For example,<code>vienna2_mfe</code> is a prediction made by the <a href=\"https://www.tbi.univie.ac.at/RNA/index.html\" target=\"_blank\">ViennaRNA package</a>. Here is the list of packages:</p>\n<ul>\n<li><a href=\"https://www.tbi.univie.ac.at/RNA/index.html\" target=\"_blank\">ViennaRNA</a></li>\n<li><a href=\"http://contra.stanford.edu/contrafold/\" target=\"_blank\">Contrafold</a></li>\n<li><a href=\"https://nupack.org/\" target=\"_blank\">Nupack</a></li>\n<li><a href=\"https://github.com/eternagame/EternaFold\" target=\"_blank\">Eternafold</a></li>\n<li><a href=\"https://github.com/ml4bio/e2efold\" target=\"_blank\">e2efold</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/16199760/\" target=\"_blank\">hotknots</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/21685106/\" target=\"_blank\">ipknots</a></li>\n<li><a href=\"https://github.com/HosnaJabbari/Iterative-HFold/\" target=\"_blank\">iterative-hfold</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/29868872/\" target=\"_blank\">knotty</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/9925784/\" target=\"_blank\">pknots</a></li>\n<li><a href=\"https://github.com/jaswindersingh2/SPOT-RNA\" target=\"_blank\">spotrna</a></li>\n<li><a href=\"https://github.com/ltrinity/Shapify\" target=\"_blank\">shapify</a></li>\n</ul>\n<p>Many of these prediction algorithms produce dot-bracket notation strings representing the predicted structure, but some can also produce a base pair probability matrix detailing the likelihood of each base in an input sequence pairing to every other base in the sequence. These base pair probability matrices can be used as inputs to heuristic algorithms that can predict a dot-bracket structure. These heuristic algorithms can be used to predict pseudoknots, even when paired with input algorithms that don't consider pseudoknot presence. We use two heuristic algorithms here, Threshknot and an implementation of the Hungarian algorithm. Columns of the form <code>eternafold[hungarian]_mfe</code> used Eternafold to predict a base pair probability matrix, then used the hungarian algorithm to generate a dot-bracket structure.</p>\n<p>One more wrinkle to these data; most of these predictions were made running on a Stanford supercomputer cluster. If the algorithm is prefixed with <code>eterna_</code> as in <code>eterna_eternafold_threshknot</code>, the predictor was run using the version of the associated algorithm available in the Eterna game. These should largely be the same, but it is possible that differences in software versions or heuristic implementations could lead to differences. So in your specific example, <code>eternafold_threshknot</code> is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by threshknot, run on a Stanford supercomputer, while <code>eterna_eternafold_threshknot</code> is a base pair probability matrix generated by eternafold, processed into a dot bracket structure by Eterna's implementation of threshknot, run in the Eterna game.</p>\n<p>Let me know if you have any other questions about those data.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2465344,
      "author_name": "Rachael Christine Kretsch",
      "author_url": "",
      "post_date": "2023-10-02T21:19:08.863000",
      "content": "<p>We would like to invite you to a <strong>seminar series</strong> focused on <strong>RNA 3D structure</strong>. In particular, <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> will be giving the next talk on October 10th at 8am Pacific Time titled <strong>Ribonanza: big data for RNA structure prediction</strong> on <a href=\"https://tinyurl.com/casp-rna-zoom\" target=\"_blank\">zoom</a>.</p>\n<p>You can receive messages about future seminars by adding yourself to the <a href=\"https://groups.google.com/g/casp-rna-sig\" target=\"_blank\">google group</a>, check <a href=\"https://tinyurl.com/rna-sig-schedule\" target=\"_blank\">past and upcoming talks</a>, watch <a href=\"https://tinyurl.com/rna-sig-playlist\" target=\"_blank\">past talks</a>, and add events to your <a href=\"https://tinyurl.com/rna-sig-calendar\" target=\"_blank\">google</a> or <a href=\"https://tinyurl.com/rna-sig-cal-ics\" target=\"_blank\">outlook</a> calendar.         </p>\n<p>Hope you are enjoying working with chemical mapping data!<br>\nRachael</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2449924,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2023-09-21T14:19:54.350000",
      "content": "<p>Could you please help to resolve the following inconsistency:<br>\nIn the data description we have</p>\n<blockquote>\n  <p>We have split out 311,935 of these 1,118,513 sequences for a public test set to allow for continuous evaluation through the competition, on the Public Leaderboard. This set has been additionally filtered to ensure high signal-to-noise and read coverage (see note on SN_filter above).<br>\n  Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. </p>\n</blockquote>\n<p>i.e. the real evaluation set may be expected to be 1/4 of 311,935 = 78k, assuming the same fraction of samples passed the SN_filter criterion as in the train data, 1/4. So it gives approximately 7% of the test data given that 1,031,888 are going to be synthesized and are expected to pass SN criterion. Meanwhile, the description of the LB says </p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 2% of the test data.</p>\n</blockquote>\n<p>Is 2% a misprint, or do you either use a different SN criterion than SN_filter (which filters out more samples), or do your 311,935 test samples for some reason have a much higher fraction of low SN samples than train data? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2452219,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2023-09-23T07:01:37.857000",
          "content": "<p>Not sure if this helps or adds to the confusion,  I had asked a question below about the scoring of Private LB in ref to -</p>\n<blockquote>\n  <p>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins, to help ensure rigor. Furthermore, these test sequences will have lengths ranging from 207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models. As those profiles are collected, they will again be filtered for acceptable signal to noise before computing your final Private Leaderboard scores.</p>\n</blockquote>\n<p>all Test has 1343823 sequences (from the data for test sequences)<br>\nprivate Test 1031888 sequences  of these  &gt;= 207 1008000  (&gt; 207 8000 = 207 1000000)   so seems 23888 &lt; 207<br>\ndiff all-priv  311935  sequences <br>\nall test &lt; 207 335823</p>\n<blockquote>\n  <p>At the beginning of the competition, Stanford scientists have data on 1,118,513 RNA sequences of lengths ranging from 115 to 206.<br>\n  We have split out 311,935 of these 1,118,513 sequences for a public test set …</p>\n</blockquote>\n<p>So the 1,118,513 only refers to train or public test  and in train there are 806573 unique sequences  (806578 is the difference of 1,118,513 and 311,935) .</p>\n<blockquote>\n  <p>The remaining 806,578 sequences for which we have data are in train_data.csv. We note that 37,828 of the test sequences, derived from the RFAM database, are identical or near-identical to the train set; for these cases the leaderboard test set contains higher signal-to-noise measurements than the train_data and serve as a test of model ability to 'denoise' chemical mapping data.</p>\n</blockquote>\n<p>Given private test appears to have 23888 sequences not newly synthesized and &lt; 207 they could be part of the 37828 derived from RFAM database with higher signal to noise than their train counterparts and presumably not out of the 311,935 not in train.</p>\n<p>If 2% of test data is correct, it would seem like 2% of the total in all of test 1343823 maybe, but can really only be from the set &lt; 207 and the original split 311,935 of 1,118,513 sequences and possibly from the 37828 derived from RFAM database with higher signal to noise than train.  </p>\n<p>It may be that public LB is looking at a test of model ability to 'denoise' chemical mapping data so higher signal to noise, and Private LB is looking at a test of the generality of models to different length distribution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2454543,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-09-24T22:49:22.350000",
          "content": "<p><a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> thanks for the discussion (and for your posts in other threads!).</p>\n<p>The 2% number is confusing. </p>\n<p>It is a small number here not just because there are so many 'future' data, approximately 1M sequences. It also is influenced by the fact that we do not have specific positions (e.g., the first 26 and last ~51) for any of the available test sequences. Kaggle's MAE scoring scripts conflate the those unavailable positions in the public LB  with all of the future data positions.</p>\n<p>Then there is also the fact that you point out -- something like 4/5 of the available test sequences do not have data with acceptable signal-to-noise for inclusion in the MAE evaluation score.      </p>\n<p>For the actual private LB sequences, we don't suddenly expect to have another 98% of the data to actually score models, as seems to be implied by the 2% number. </p>\n<p>Instead, with our current protocols, we are probing 1M sequences, and conservatively expect to have roughly another 200,000 sequences with acceptable signal-to-noise for the private LB, and we'll be missing the initial and final positions.  The number of rows that will be used for scoring in the private LB will still end up being bigger than the public LB, but by a factor of 4x, not by a factor of 50x.  It would probably be better for the \"approximately 2%\" number to be set as \"approximately 20%\". </p>\n<p>Just one more caveat:  we are working on an experimental advance that may significantly enhance signal-to-noise, by taking advantage of an upgrade in Illumina sequencing technology that is literally becoming available in October  2023. So the above numbers may change – in the direction of allowing even more robust scoring of your submissions.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2454570,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-09-24T23:04:23.767000",
              "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> for your detailed explanation. It is clear now.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2454926,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2023-09-25T07:20:00.727000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> for this reply.  </p>\n<blockquote>\n  <p>an upgrade in Illumina sequencing technology that is literally becoming available in October 2024</p>\n</blockquote>\n<p>is that October 2023?  Hopefully the experimental advance is successful and \"more robust scoring\" is a good thing!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2455683,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2023-09-25T16:36:12.333000",
              "content": "<p>Oops, fixed. 😊</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2455992,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2023-09-25T21:06:21.960000",
              "content": "<p>Update: we've updated the blurb on leaderboard to say:</p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.</p>\n</blockquote>\n<p>Thanks for the discussion!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2440109,
      "author_name": "LambertMoyon",
      "author_url": "",
      "post_date": "2023-09-15T09:15:53.977000",
      "content": "<p>Hello, I'm a postdoc in Germany, working in an academic research institute. I am considering participating to the challenge, but would be keen to publishing my method if proven successful (in the context of this competition, but also considering other aspects). Would this infringe the rules of the competition? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2441024,
          "author_name": "DigitalEmbrace",
          "author_url": "",
          "post_date": "2023-09-16T00:50:30.313000",
          "content": "<p>That should be fine. The winning algorithms and the dataset will be released as open source after the competition ends. The goal of this competition is better models and robust research!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2499312,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-26T01:05:15.943000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2431149,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2023-09-09T20:44:33.680000",
      "content": "<p>Thanks for putting this competition together. I'm excited to learn.</p>\n<p>Small note. When describing the columns within the train/test files. You provide a reference to DMS/2A3, however the 2A3 link returns 404. The section is below:</p>\n<blockquote>\n  <p>experiment_type - (string) Either DMS_MaP or 2A3_MaP to describe the type of chemical mapping experiment that was used to generate each profile. References: DMS, 2A3.</p>\n</blockquote>\n<p>The link for 2A3 points to --&gt; <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/10.1093/nar/gkaa1255\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/10.1093/nar/gkaa1255</a></p>\n<p>I assume it is supposed to point here: <a href=\"https://academic.oup.com/nar/article/49/6/e34/6062772\" target=\"_blank\">https://academic.oup.com/nar/article/49/6/e34/6062772</a></p>\n<p><strong>A novel SHAPE reagent enables the analysis of RNA structure in living cells with unprecedented accuracy</strong></p>\n<blockquote>\n  <p><strong>Abstract:</strong><br>\n  Due to the mounting evidence that RNA structure plays a critical role in regulating almost any physiological as well as pathological process, being able to accurately define the folding of RNA molecules within living cells has become a crucial need. We introduce here 2-aminopyridine-3-carboxylic acid imidazolide (2A3), as a general probe for the interrogation of RNA structures in vivo. 2A3 shows moderate improvements with respect to the state-of-the-art selective 2′-hydroxyl acylation analyzed by primer extension (SHAPE) reagent NAI on naked RNA under in vitro conditions, but it significantly outperforms NAI when probing RNA structure in vivo, particularly in bacteria, underlining its increased ability to permeate biological membranes. When used as a restraint to drive RNA structure prediction, data derived by SHAPE-MaP with 2A3 yields more accurate predictions than NAI-derived data. Due to its extreme efficiency and accuracy, we can anticipate that 2A3 will rapidly take over conventional SHAPE reagents for probing RNA structures both in vitro and in vivo.</p>\n</blockquote>\n<p>Thanks again!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2432153,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-09-10T16:05:37.397000",
          "content": "<p>Fixed the reference. Thanks for catching!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2539329,
      "author_name": "Unranked",
      "author_url": "",
      "post_date": "2023-11-27T00:58:57.330000",
      "content": "<p>I just finished my data science diploma and, though ambitious, I want to use this competition as a personal project to improve my skills. Therefore I wanted to ask if the data will remain available after the final deadline</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2506798,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2023-10-31T14:53:47.187000",
      "content": "<p>Is the submitted submissions final with all the necessary information, includes all the test data, or is it a re-run of the kernels/models after the deadline? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2505734,
      "author_name": "BorisNiels",
      "author_url": "",
      "post_date": "2023-10-30T18:52:04.010000",
      "content": "<p>Hello! My understanding is correct that we have a set of RNA molecules with a known primary and secondary structure. We must predict the shape of a molecule in three dimensions (tertiary structure), which is determined by the degree of response to certain chemical influences (reactivity_0001, reactivity_0002,…)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2505774,
          "author_name": "Jonathan Romano",
          "author_url": "",
          "post_date": "2023-10-30T19:52:40.100000",
          "content": "<p>Not quite - you have a known sequence (primary structure) and must predict the reactivity. The reactivity prediction can later be used to more accurately infer secondary and even tertiary structure from the sequence, but that process is outside the scope of this competition. There are existing methods (with mixed accuracy) for predicting a secondary structure from the sequence which you may wish to use as part of your solution.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2500676,
      "author_name": "Ilyana Anderson",
      "author_url": "",
      "post_date": "2023-10-26T23:07:18.717000",
      "content": "<p>I have a question about the rules. I would like to participate in this competition and make a submission as part of a project I'm doing for a class. In this class, at the end of the semester (by December 4, which is before the competition ends), I will submit a project report to the instructors of my class. Does this (submitting my report privately to the instructors of my class) constitute the infringement of the rules of the competition or not?<br>\nThe report won’t contain any code nor data (it will contain only ideas and results).<br>\nI read the rules, and it seemed to me that I’m okay so long as I’m not sharing privately my code or data. But I would like to confirm. (I’m new to Kaggle.)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2500740,
          "author_name": "DigitalEmbrace",
          "author_url": "",
          "post_date": "2023-10-27T01:42:10.600000",
          "content": "<p>Yes, that will be fine. Good luck on your project!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2500796,
              "author_name": "Ilyana Anderson",
              "author_url": "",
              "post_date": "2023-10-27T02:57:45.477000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2554154,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-12-08T20:29:05.513000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2496176,
      "author_name": "Ashwin Ramesh",
      "author_url": "",
      "post_date": "2023-10-23T19:32:38.493000",
      "content": "<p>Hi, i'm new to the competition. Can we use the test_sequences for pretraining or is that against the competition rules? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2497656,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-10-24T18:38:57.580000",
          "content": "<p>You are welcome to use test_sequences  for pre training. Good luck!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2486485,
      "author_name": "NathanTuttle",
      "author_url": "",
      "post_date": "2023-10-18T01:05:49",
      "content": "<p>Regarding the following statement on the overview</p>\n<blockquote>\n  <p>From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.</p>\n</blockquote>\n<p>What are those two ground truth values. Is this the reactivity columns? Does it include the error? Is the ground truth the two mapping experiments?</p>\n<p>If so the wording is off because the first 26 values are not available due to technical reasons. Also not sure if same sequence has both mapping experiments.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2497661,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-10-24T18:43:35.647000",
          "content": "<p>Yes the reactivity columns are the values you are trying to predict. </p>\n<p>In the train data, there are separate error columns that give estimates of the experimental uncertainties  (precision) on the measured reactivity values - you are welcome to ignore those estimated errors or to try to take them into account to down weight high error positions or sequences during training.</p>\n<p>For the first 26 values they are not in the train data but may be available in private leaderboard test data - it will be an interesting challenge for models to generalize to this part of the sequence or to longer lengths - you can look through some of the other discussion threads to look for ideas and tests of generalization. Thanks for the questions!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2479981,
      "author_name": "NathanTuttle",
      "author_url": "",
      "post_date": "2023-10-13T03:40:58.393000",
      "content": "<p>Just want to know what 'nts' means in the signal column \"mean( measurement value over probed nts )/mean( statistical error in measurement value over probed nts)\"</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2480680,
          "author_name": "DigitalEmbrace",
          "author_url": "",
          "post_date": "2023-10-13T13:26:50.037000",
          "content": "<p>'nts' is an abbreviation for nucleotides.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2483764,
              "author_name": "NathanTuttle",
              "author_url": "",
              "post_date": "2023-10-16T01:10:11.087000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2468821,
      "author_name": "Jason Yoon",
      "author_url": "",
      "post_date": "2023-10-05T21:00:32.260000",
      "content": "<p>This is great project to predict RNA molecular structure! I can't believe that the RNA molecular structures are still not well understood even though there are tons of RNA drugs! Thank you for y'all work and I will try my best to make a good prediction model! DDS forever!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2467430,
      "author_name": "Bharath_011",
      "author_url": "",
      "post_date": "2023-10-04T14:48:07.203000",
      "content": "<p>For file input and path setup, it takes several minutes in kaggle notebook. what should I do?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2453698,
      "author_name": "KF",
      "author_url": "",
      "post_date": "2023-09-24T09:03:28.507000",
      "content": "<p>Some of chemical softwares are free but have the following licenses:</p>\n<ul>\n<li>GPL / LGPL</li>\n<li>Non-Commercial Academic Use Only</li>\n</ul>\n<p>Is it possible to check the availability of these software?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2454557,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-09-24T22:58:10.733000",
          "content": "<p>You are welcome to use those software packages to prepare your submissions!</p>\n<p>If your submission becomes prize eligible, you will need to make your <em>own</em> code available under a fully open source license. </p>\n<p>It will be fine at that time to simply include links to any codebases that you used that are licensed as GPL/LGPL or for non-commercial use.  If you have had to make updates to those prior codebases, we might ask you to provide a diff of your updated version to the prior codebase along with your software. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2454601,
              "author_name": "KF",
              "author_url": "",
              "post_date": "2023-09-24T23:52:51.607000",
              "content": "<p>Thanks for confirmation!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2452594,
      "author_name": "Bhavesh Padharia",
      "author_url": "",
      "post_date": "2023-09-23T12:35:24.403000",
      "content": "<p>Hello Ribonanza Challenge Team,</p>\n<p>I'm excited to see this initiative aiming to tackle the complex problem of RNA structure prediction. RNA's significance in biology, especially its role in medicine and understanding the origins of life, cannot be overstated.</p>\n<p>It's impressive to see the scale of data collection efforts and the collaboration with both expert RNA databases and open science projects like Eterna. This approach promises to provide a wealth of valuable information to researchers.</p>\n<p>I'm looking forward to the competition and the innovative solutions that the data science and machine learning community will bring to the table. This is an opportunity to make a significant impact on our understanding of RNA and its applications in medicine and biology.</p>\n<p>Count me in, and let's work together to find the oracle for RNA structure! 🧬💻 #RibonanzaChallenge</p>\n<p>Best regards,<br>\nBhavesh Padharia</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2448513,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-20T16:19:08.327000",
      "content": "<p>In relation to the <strong>base-pairing probabilities</strong>, what is the difference between these two things?</p>\n<ul>\n<li>The files total 27 Go. <strong>Ribonanza_bpp_files</strong> (provided as data), which give the BPPs in txt format for the train and test sequences/ this format: position of nucleotide a, position of nucleotide b, probability</li>\n<li>And the possibility of calculating them (these BPPs) using your pinned notebook: <a href=\"https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment\" target=\"_blank\">RNA Science Computational Environment</a>, I mean arnie library (and package = \"<strong>eternafold</strong>\")<br>\n<code>from arnie.bpps import bpps</code><br>\n<code>bpps(sequence,package=\"eternafold\")</code></li>\n</ul>\n<p>In my understanding, with the exception of their format, <strong>which is not a matrix</strong>, the txt files provided as data (Ribonanza_bpp_files) are supposed to accomplish the same thing (in the data section: you define them : \"a TXT file of base pair probabilities from the LinearPartition-<strong>EternaFold</strong> package\").&nbsp;&nbsp;</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2454576,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-09-24T23:05:48.540000",
          "content": "<p>Good question! The <strong>Ribonanza_bpp_files</strong> are not matrices. For conciseness, these files only list pairs of nucleotides which have non-zero base pair probabilities. I have updated the data description. Thanks for pointing this out.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2438618,
      "author_name": "MISS MAILIANA",
      "author_url": "",
      "post_date": "2023-09-14T11:43:16.927000",
      "content": "<p>Hey,<br>\nMy name is Miss Mailiana, I am an entry level professional in the field of Data Analytics. I am a recent graduate with BA in Statistics with Economics and Mathematics, I’m ready to team up to share my knowledge and build up more to win the prize.<br>\nPlease reach out </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2439922,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2023-09-15T06:52:01.470000",
          "content": "<p>You may want to post this in <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437707\" target=\"_blank\">Looking for a Team Megathread</a>  Use this thread to find a teammate if you're interested in finding others to work with!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2436090,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-09-13T10:40:49.553000",
      "content": "<p>Noticed this does not seem to be a code competition and wondered how the Private LB and private test set will work - </p>\n<blockquote>\n  <p>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins, to help ensure rigor. Furthermore, these test sequences will have lengths ranging from 207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models. As those profiles are collected, they will again be filtered for acceptable signal to noise before computing your final Private Leaderboard scores.</p>\n</blockquote>\n<p>Know it is early in the competition so maybe still being determined?  e.g., will there be a date when it  is released so that new submissions will need to be done, or is it that the selected submission notebooks will be rerun after competition end?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2441007,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-09-15T23:31:41.867000",
          "content": "<p>After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for <strong>private leaderboard</strong>* test sequences for which the submissions already include predictions, so need for you to prepare new submissions or provide notebooks. Thanks for asking the question!</p>\n<p>*Edited after discussion below.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2441425,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2023-09-16T07:28:02.527000",
              "content": "<p>Thanks for the reply!  </p>\n<blockquote>\n  <p>these test sequences will have lengths ranging from 207 to 457 bases</p>\n</blockquote>\n<p>Although there is new data, that refers to what will affect the solution for scoring, and the data in the sequence column for test_sequences will not change. Good to know!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2452579,
              "author_name": "hotbit",
              "author_url": "",
              "post_date": "2023-09-23T12:24:25.263000",
              "content": "<blockquote>\n  <p>After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for test sequences for which the submissions already include predictions, so need for you to prepare new submissions or provide notebooks. </p>\n</blockquote>\n<p>My understanding from the context and also according to chatGPT there is \"no need to prepare new submissions.\"  Based on this, do I follow correctly?</p>\n<ol>\n<li>A sequence X with  N bases will be re-synthesized, but M bases will be added, thus if N = 200 and M = 100, the new Y seq will have 300 bases.</li>\n<li>Reactivity for Y will be measured</li>\n<li>Data will be clipped back to the first N bases/reactivities.</li>\n</ol>\n<blockquote>\n  <p>207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models.</p>\n</blockquote>\n<p>In the original seq X, let's say X_reactivity_0060 = 0.95. But base on the position 260 in the seq Y may affect Y_reactivity_0060 in Y such that it is only 0.05. </p>\n<p>Therefore, I don't understand how it tests the generality of the model, nor how it makes sense at all, as any model would need the knowledge about a base in position 260 to correctly predict 160.</p>\n<p>I guess I badly misunderstood something.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2454596,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2023-09-24T23:45:50.950000",
              "content": "<p>Thanks for the question. </p>\n<p>There are numerous sequences in <code>test_sequences.csv</code> for which we are asking for submissions. But we don't have the data to evaluate your submissions for those sequences at the start of the competition.  </p>\n<p>These are literally sequences that we have not synthesized before, so there won't be any re-synthesis -- we'll synthesize the RNA for the first time in the upcoming months.  </p>\n<p>Also there will not be any clipping back to the first N bases. </p>\n<p>What might be confusing here is that the <code>train_data.csv</code> only has columns going to <code>reactivity_0206</code>. However, we are asking for predictions for test sequences with lengths up to 457. If you look in <code>test_sequences.csv</code>, you'll see that those long sequences will actually correspond to 457 rows that you need to fill in <code>sample_submission.csv</code>. </p>\n<p>(In retrospect, we could have included columns up to <code>reactivity_0457</code> in <code>train_data.csv</code>, filled with blanks!)</p>\n<p>Good luck with your submissions!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2432476,
      "author_name": "josjesus",
      "author_url": "",
      "post_date": "2023-09-10T21:26:42.423000",
      "content": "<p>Hello, this might be an obvious question, but from the provided files I don't seem to find one that provides actual results, i.e the DMS_MaP and 2A3_MaP for any sequence. I understand that for the final score we can not have access to these, but without some training data with the target how can we train models?</p>\n<p>Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2433586,
          "author_name": "Jonathan Romano",
          "author_url": "",
          "post_date": "2023-09-11T17:05:40.523000",
          "content": "<p>The measured values are found in the reactivity_0001 through reactivity_0170 columns in the training data (as well as reactivity error columns which provide experimental error). Each row of training data will either be providing results for either DMS or 2A3, as specified by the experiment_type column. When you construct your submission file, it will have one line per sequence position and both DMS and 2A3 data rather than per sequence with only one or the other</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2434994,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-09-12T16:25:40.843000",
              "content": "<p>CORRECTION: reactivity_0001 through reactivity_0206 (not 170) - I misread the available columns when I was looking at the training data</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2450331,
              "author_name": "hotbit",
              "author_url": "",
              "post_date": "2023-09-21T19:01:34.173000",
              "content": "<p>I think you haven't misread, they were mis-displayed. I was just puzzled a minute ago as to why in the viewer (Data tab) there are only 170 reactivity but still 206 error columns. It seems some columns might not be displayed sometimes, worth keeping at the back of mind.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2450434,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-09-21T21:03:14.983000",
              "content": "<p>I think my issue was reactivity 1-170 were shown, then it switched to error, then it went back to reactivity 171-206</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2431548,
      "author_name": "Sai Kiran Varma",
      "author_url": "",
      "post_date": "2023-09-10T07:34:24.597000",
      "content": "<p>Can you please provide more clarity on the below sentence  in bold fromthe data description in simple words for non-domain users like me?<br>\n\"In this competition, you will be predicting the reactivity of an RNA sequence to two chemical modifiers DMS and 2A3. <br>\n<strong>These data can be measured efficiently through a mutational profiling (MaP) experiment read out by high-throughput sequencing and positions that are protected from chemical modification are likely to be forming base pairs or other kinds of RNA structure</strong>.\"</p>\n<p>Thanks,<br>\nKiran</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2472093,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-10-07T01:49:25.097000",
          "content": "<p>Hi! For discussion of the experimental protocol you might want to check out this thread: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/445415\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/445415</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2455993,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-25T21:09:34.977000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2426885": "**Hello RNA world**\n\nRNA is the basis for new medicines and the oldest forms of life —  and yet is one of the most poorly understood molecules in biology. \n\nThe other major kinds of macromolecule of biology — DNA and protein —  have had predictable structures since the seminal work of Watson, Crick and Franklin in the 1950s and the AlphaFold 2 breakthrough of 2021. \n\nBut we're still bad at modeling RNA.\n\nYou might ask: \n\n*How can we be so bad at RNA, especially if there are so many more different kinds of RNA’s than there are proteins in our bodies and throughout biology?* \n\nThe answer is that the RNA’s have historically been difficult to experimentally characterize — they can form multiple 3D structures and that has rendered RNA difficult to see with conventional visualization approaches. \n\n*What’s the prospect of RNA having an AlphaFold moment when we have 3D coordinate information on only thousands of RNA molecules, compared to hundreds of thousands of protein structure?*\n\nEnter the new **Ribonanza** data set — our attempt to get enough data to finally crack the problem of RNA structure prediction.\n\nWe’ve scaled up a way to get rich information of RNA sequences based on a chemical mapping approach that reads out one number per position. \n\nThis number is highly sensitive to how structured the RNA is at each position, averaged over the ensemble of conformations.   \n\nAnd we’ve been synthesizing a diverse assortment of over 1M RNA’s, some from the expert RNA databases and some from the fabulous internet open science project Eterna!\n\nNow we need to find a model that can learn from the Ribonanza data.\n\nRNA scientists could then use the model as an oracle to understand the trillions of existing molecules for which we are missing structural information and to guide the design of new RNAs for the future of medicine and biology.\n\nThat’s where you come in: *Help us find the oracle for RNA structure!*\n\nWe need each and every idea in data science and machine learning to be brought to this data set RNA structure prediction -- and we know that Kaggle will deliver. See you in the competition!\n\n**Hosts**\n\nRhiju Das @rhijudas\n\nShujun He @shujun717\n\nThomas Karagianes @brainbowrna\n\nJill Townley @digitalembrace\n\nRachael Kretsch @rkretsch\n\n*Special thanks to Rui Huang for developing the experimental protocols for Ribonanza!*\n\nThanks to members of the Eterna community and the Das laboratory, including John Nichol, Grace Nye, Christian Choe, and Jonathan Romano, for key contributions in software development and library design. \n\nAnd our collaborators at Kaggle:\n\nMaggie Demkin @maggiemd\nInversion @inversion\n",
    "2479156": "How were the reactivity errors calculated?",
    "2468230": "Could you kindly provide additional details about the supplementary files? Specifically, I am interested in the CSV files within the supplementary_silico_predictions folder. It appears that these files have varying headers (columns). Could you please specify the software package utilized for their generation? Are the column names consistent for each algorithm used during the generation process? For instance, is 'eternafold_threshknot' the same as 'eterna_eternafold_threshknot'?",
    "2465344": "We would like to invite you to a **seminar series** focused on **RNA 3D structure**. In particular, @rhijudas will be giving the next talk on October 10th at 8am Pacific Time titled **Ribonanza: big data for RNA structure prediction** on [zoom](https://tinyurl.com/casp-rna-zoom).\n\nYou can receive messages about future seminars by adding yourself to the [google group](https://groups.google.com/g/casp-rna-sig), check [past and upcoming talks](https://tinyurl.com/rna-sig-schedule), watch [past talks](https://tinyurl.com/rna-sig-playlist), and add events to your [google](https://tinyurl.com/rna-sig-calendar) or [outlook](https://tinyurl.com/rna-sig-cal-ics) calendar. \t\t\n\nHope you are enjoying working with chemical mapping data!\nRachael",
    "2449924": "Could you please help to resolve the following inconsistency:\nIn the data description we have\n>We have split out 311,935 of these 1,118,513 sequences for a public test set to allow for continuous evaluation through the competition, on the Public Leaderboard. This set has been additionally filtered to ensure high signal-to-noise and read coverage (see note on SN_filter above).\n>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. \n\ni.e. the real evaluation set may be expected to be 1/4 of 311,935 = 78k, assuming the same fraction of samples passed the SN_filter criterion as in the train data, 1/4. So it gives approximately 7% of the test data given that 1,031,888 are going to be synthesized and are expected to pass SN criterion. Meanwhile, the description of the LB says \n>This leaderboard is calculated with approximately 2% of the test data.\n\nIs 2% a misprint, or do you either use a different SN criterion than SN_filter (which filters out more samples), or do your 311,935 test samples for some reason have a much higher fraction of low SN samples than train data? Thanks.",
    "2440109": "Hello, I'm a postdoc in Germany, working in an academic research institute. I am considering participating to the challenge, but would be keen to publishing my method if proven successful (in the context of this competition, but also considering other aspects). Would this infringe the rules of the competition? ",
    "2431149": "Thanks for putting this competition together. I'm excited to learn.\n\nSmall note. When describing the columns within the train/test files. You provide a reference to DMS/2A3, however the 2A3 link returns 404. The section is below:\n\n> experiment_type - (string) Either DMS_MaP or 2A3_MaP to describe the type of chemical mapping experiment that was used to generate each profile. References: DMS, 2A3.\n\nThe link for 2A3 points to --> https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/10.1093/nar/gkaa1255\n\nI assume it is supposed to point here: https://academic.oup.com/nar/article/49/6/e34/6062772\n\n**A novel SHAPE reagent enables the analysis of RNA structure in living cells with unprecedented accuracy**\n> **Abstract:**\n> Due to the mounting evidence that RNA structure plays a critical role in regulating almost any physiological as well as pathological process, being able to accurately define the folding of RNA molecules within living cells has become a crucial need. We introduce here 2-aminopyridine-3-carboxylic acid imidazolide (2A3), as a general probe for the interrogation of RNA structures in vivo. 2A3 shows moderate improvements with respect to the state-of-the-art selective 2′-hydroxyl acylation analyzed by primer extension (SHAPE) reagent NAI on naked RNA under in vitro conditions, but it significantly outperforms NAI when probing RNA structure in vivo, particularly in bacteria, underlining its increased ability to permeate biological membranes. When used as a restraint to drive RNA structure prediction, data derived by SHAPE-MaP with 2A3 yields more accurate predictions than NAI-derived data. Due to its extreme efficiency and accuracy, we can anticipate that 2A3 will rapidly take over conventional SHAPE reagents for probing RNA structures both in vitro and in vivo.\n\nThanks again!",
    "2539329": "I just finished my data science diploma and, though ambitious, I want to use this competition as a personal project to improve my skills. Therefore I wanted to ask if the data will remain available after the final deadline",
    "2506798": "Is the submitted submissions final with all the necessary information, includes all the test data, or is it a re-run of the kernels/models after the deadline? ",
    "2505734": "Hello! My understanding is correct that we have a set of RNA molecules with a known primary and secondary structure. We must predict the shape of a molecule in three dimensions (tertiary structure), which is determined by the degree of response to certain chemical influences (reactivity_0001, reactivity_0002,...)?",
    "2500676": "I have a question about the rules. I would like to participate in this competition and make a submission as part of a project I'm doing for a class. In this class, at the end of the semester (by December 4, which is before the competition ends), I will submit a project report to the instructors of my class. Does this (submitting my report privately to the instructors of my class) constitute the infringement of the rules of the competition or not?\nThe report won’t contain any code nor data (it will contain only ideas and results).\nI read the rules, and it seemed to me that I’m okay so long as I’m not sharing privately my code or data. But I would like to confirm. (I’m new to Kaggle.)",
    "2496176": "Hi, i'm new to the competition. Can we use the test_sequences for pretraining or is that against the competition rules? ",
    "2486485": "Regarding the following statement on the overview\n\n\n>From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.\n\nWhat are those two ground truth values. Is this the reactivity columns? Does it include the error? Is the ground truth the two mapping experiments?\n\nIf so the wording is off because the first 26 values are not available due to technical reasons. Also not sure if same sequence has both mapping experiments.\n",
    "2479981": "Just want to know what 'nts' means in the signal column \"mean( measurement value over probed nts )/mean( statistical error in measurement value over probed nts)\"",
    "2468821": "This is great project to predict RNA molecular structure! I can't believe that the RNA molecular structures are still not well understood even though there are tons of RNA drugs! Thank you for y'all work and I will try my best to make a good prediction model! DDS forever!!",
    "2467430": "For file input and path setup, it takes several minutes in kaggle notebook. what should I do?\n",
    "2453698": "Some of chemical softwares are free but have the following licenses:\n\n- GPL / LGPL\n- Non-Commercial Academic Use Only\n\nIs it possible to check the availability of these software?",
    "2452594": "Hello Ribonanza Challenge Team,\n\nI'm excited to see this initiative aiming to tackle the complex problem of RNA structure prediction. RNA's significance in biology, especially its role in medicine and understanding the origins of life, cannot be overstated.\n\nIt's impressive to see the scale of data collection efforts and the collaboration with both expert RNA databases and open science projects like Eterna. This approach promises to provide a wealth of valuable information to researchers.\n\nI'm looking forward to the competition and the innovative solutions that the data science and machine learning community will bring to the table. This is an opportunity to make a significant impact on our understanding of RNA and its applications in medicine and biology.\n\nCount me in, and let's work together to find the oracle for RNA structure! 🧬💻 #RibonanzaChallenge\n\nBest regards,\nBhavesh Padharia",
    "2448513": "In relation to the **base-pairing probabilities**, what is the difference between these two things?\n- The files total 27 Go. **Ribonanza_bpp_files** (provided as data), which give the BPPs in txt format for the train and test sequences/ this format: position of nucleotide a, position of nucleotide b, probability\n- And the possibility of calculating them (these BPPs) using your pinned notebook: [RNA Science Computational Environment](https://www.kaggle.com/code/brainbowrna/rna-science-computational-environment), I mean arnie library (and package = \"**eternafold**\")\n`from arnie.bpps import bpps`\n`bpps(sequence,package=\"eternafold\") `\n\nIn my understanding, with the exception of their format, **which is not a matrix**, the txt files provided as data (Ribonanza_bpp_files) are supposed to accomplish the same thing (in the data section: you define them : \"a TXT file of base pair probabilities from the LinearPartition-**EternaFold** package\").  ",
    "2438618": "Hey,\nMy name is Miss Mailiana, I am an entry level professional in the field of Data Analytics. I am a recent graduate with BA in Statistics with Economics and Mathematics, I’m ready to team up to share my knowledge and build up more to win the prize.\nPlease reach out ",
    "2436090": "Noticed this does not seem to be a code competition and wondered how the Private LB and private test set will work - \n\n>Our final and most important scoring (the Private Leaderboard) involves 1,031,888 sequences. Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins, to help ensure rigor. Furthermore, these test sequences will have lengths ranging from 207 to 457 bases -- we intentionally chose a different length distribution compared to the train data and Public Leaderboard, to help test the generality of your models. As those profiles are collected, they will again be filtered for acceptable signal to noise before computing your final Private Leaderboard scores.\n\nKnow it is early in the competition so maybe still being determined?  e.g., will there be a date when it  is released so that new submissions will need to be done, or is it that the selected submission notebooks will be rerun after competition end?  ",
    "2432476": "Hello, this might be an obvious question, but from the provided files I don't seem to find one that provides actual results, i.e the DMS_MaP and 2A3_MaP for any sequence. I understand that for the final score we can not have access to these, but without some training data with the target how can we train models?\n\nThanks",
    "2431548": "Can you please provide more clarity on the below sentence  in bold fromthe data description in simple words for non-domain users like me?\n\"In this competition, you will be predicting the reactivity of an RNA sequence to two chemical modifiers DMS and 2A3. \n**These data can be measured efficiently through a mutational profiling (MaP) experiment read out by high-throughput sequencing and positions that are protected from chemical modification are likely to be forming base pairs or other kinds of RNA structure**.\"\n\nThanks,\nKiran",
    "2455993": ""
  }
}