{
  "id": 454397,
  "title": "External data from RNA Mapping DataBase",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/454397",
  "author_name": "",
  "post_date": "2023-11-10T04:29:34.310943500Z",
  "votes": 19,
  "comment_count": 7,
  "views": 0,
  "content": "<p>With about a month to go, we hope you are enjoying the Ribonanza competition and are beginning to look for and test new innovations!</p>\n<p>For folks who are curious about external data sources, the <a href=\"https://rmdb.stanford.edu\" target=\"_blank\">RNA Mapping DataBase</a> holds numerous chemical mapping profiles for a wide range of RNA molecules, collected by several laboratories.  </p>\n<p>These public data have been available since the beginning of the Ribonanza competition through RMDB’s <a href=\"https://rmdb.stanford.edu/site_data/published_rdat.zip\" target=\"_blank\">Download All link</a>,  but the data are in a specialized <code>RDAT</code> format. To make it easier to test out use of these diverse data, we’ve wrangled the majority of these data into a CSV file with identical formatting to the Ribonanza <code>train_data.csv</code> format.</p>\n<p>You can find the <code>rmdb_data.csv</code> data set, involving 142,300 profiles, here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data</a></p>\n<p>Some additional notes:</p>\n<ul>\n<li>The <code>experiment_type</code> field has 12 values for the different kinds of chemical mapping experiments in the dataset.<ul>\n<li>Importantly, most of the RMDB data are based on experimental read outs different from the mutational profiling (‘MaP’) experiments in the Ribonanza data set. While a technical detail, this means that the RMDB data will not exactly  agree with the Ribonanza data for any molecule. If you are using neural networks, you may wish to create new ‘heads’ of the network for each of the reaction types. For the aficionados: most RMDB measurements measure the termination of reverse transcription at nucleotides that have been modified or cleaved by chemical reactions, rather than the special mutational read through of chemical modifications that is used in Ribonanza.</li>\n<li><code>1M7</code>,<code>NMIA</code>,<code>BzCN</code> are all SHAPE modifiers that selectively tag parts of the RNA molecules that are unstructured whether they are A, C, G, or U. These are closest to the <code>2A3_MaP</code> <code>experiment_type</code> in the Ribonanza data set, which is also based on a SHAPE chemical reaction, although with a different chemical that gives higher signal to noise and a different readout, as noted above. </li>\n<li><code>DMS</code> data in RMDB involve the same dimethyl sulfate reaction as the <code>DMS_MaP</code> measurements in Ribonanza, but with the difference in readout noted above.</li>\n<li><code>DMS_M2_seq</code> are DMS experiments in which a sequence and an ensemble of mutants are separately probed (so-called ‘mutate-and-map-seq’ or <a href=\"https://doi.org/10.1073/pnas.1619897114\" target=\"_blank\">M2-seq</a> experiments).  In these data sets, the presence of a mutation at a given position is recorded, and the average reactivity at all other mutations is inferred through mutational profiling. In this <code>rmdb_data.csv</code>, the sequences of the different profiles are adjusted so that the site of one mutation (recorded as ‘X’ in the original RMDB data) has been mutated to its complement (<code>A</code> to <code>U</code>, <code>U</code> to <code>A</code>, <code>G</code> to <code>C</code>, <code>C</code> to <code>G</code>). You may wish to mask the positions around this mutation site; and you may want to try the other two possible mutations as sequences and average results. In addition, the very beginning and last 20 nucleotides or so are not properly probed in this method, so you may also want to mask those out.</li>\n<li><code>CMCT</code> is a chemical modifier that hits unstructured G’s and U’s.</li>\n<li><code>BzCN_cotx</code> and <code>DMS_cotx</code> involve so-called <a href=\"https://doi.org/10.1038/nsmb.3316\" target=\"_blank\">‘cotranscriptional’</a> SHAPE or DMS experiments  where RNA molecules are probed after they are transcribed but without any heating/cooling to try to put them in their thermodynamically most favored state.</li>\n<li>`deg_Mg_50C','deg_Mg_pH10','deg_50C', and ‘deg_pH10' are the data for RNA degradation in four separate conditions used as train and test set in the <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine\" target=\"_blank\">OpenVaccine Kaggle challenge</a>.</li>\n<li>Some other <code>experiment_type</code> present in the RMDB, notably the <a href=\"https://doi.org/10.7554/eLife.07600\" target=\"_blank\">MOHCA-seq</a> protocol, have not been included in here because the data are not really profiles of chemical reactivity like in Ribonanza. If there is interest, please convey in this thread and it may be possible to curate these into a separate data set.</li></ul></li>\n<li>The <code>reads</code> column is empty in this dataset, since many of the experiments either do not involve high-throughput DNA sequencers or the RMDB-deposited files did not record the number of reads.</li>\n<li><code>SN_filter</code> is set based on whether <code>signal_to_noise</code>&gt;1.0. Unlike Ribonanza data, the filter does not take into account <code>reads</code>, since those values are unknown or not recorded in most cases.</li>\n<li>The <code>reactivity_*</code> columns are the data you could try to use for training. These have been scaled compared to the original RMDB data sets in the same way that the Ribonanza data have been scaled, i.e. so that the 90th percentile values across each data set is 1.0.</li>\n<li>The <code>reactivity_error_*</code> columns are not always present in the original RMDB entries but have been inferred here based on either raw data present or a heuristic formula that approximates experimental errors well for older data that use a low-throughput experimental readout based on capillary electrophoresis (<code>reactivity_error = 0.1 x mean(reactivity) + 0.2 x reactivity</code>).</li>\n<li>There are some longer RNA’s in RMDB which do not contribute many data but make the CSV’s unwieldy. Only sequences with length of 512 nucleotides or less have been included in this data set. </li>\n</ul>\n<p>Thanks to <a href=\"https://www.kaggle.com/brainbow\" target=\"_blank\">@brainbow</a> and <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> for helping with curation and testing.</p>\n<p>Please post comments here with questions or if you see bugs.</p>",
  "messages": [
    {
      "id": "2519469",
      "postDate": "11/10/2023 04:29:34",
      "content": "<p>With about a month to go, we hope you are enjoying the Ribonanza competition and are beginning to look for and test new innovations!</p>\n<p>For folks who are curious about external data sources, the <a href=\"https://rmdb.stanford.edu\" target=\"_blank\">RNA Mapping DataBase</a> holds numerous chemical mapping profiles for a wide range of RNA molecules, collected by several laboratories.  </p>\n<p>These public data have been available since the beginning of the Ribonanza competition through RMDB’s <a href=\"https://rmdb.stanford.edu/site_data/published_rdat.zip\" target=\"_blank\">Download All link</a>,  but the data are in a specialized <code>RDAT</code> format. To make it easier to test out use of these diverse data, we’ve wrangled the majority of these data into a CSV file with identical formatting to the Ribonanza <code>train_data.csv</code> format.</p>\n<p>You can find the <code>rmdb_data.csv</code> data set, involving 142,300 profiles, here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data</a></p>\n<p>Some additional notes:</p>\n<ul>\n<li>The <code>experiment_type</code> field has 12 values for the different kinds of chemical mapping experiments in the dataset.<ul>\n<li>Importantly, most of the RMDB data are based on experimental read outs different from the mutational profiling (‘MaP’) experiments in the Ribonanza data set. While a technical detail, this means that the RMDB data will not exactly  agree with the Ribonanza data for any molecule. If you are using neural networks, you may wish to create new ‘heads’ of the network for each of the reaction types. For the aficionados: most RMDB measurements measure the termination of reverse transcription at nucleotides that have been modified or cleaved by chemical reactions, rather than the special mutational read through of chemical modifications that is used in Ribonanza.</li>\n<li><code>1M7</code>,<code>NMIA</code>,<code>BzCN</code> are all SHAPE modifiers that selectively tag parts of the RNA molecules that are unstructured whether they are A, C, G, or U. These are closest to the <code>2A3_MaP</code> <code>experiment_type</code> in the Ribonanza data set, which is also based on a SHAPE chemical reaction, although with a different chemical that gives higher signal to noise and a different readout, as noted above. </li>\n<li><code>DMS</code> data in RMDB involve the same dimethyl sulfate reaction as the <code>DMS_MaP</code> measurements in Ribonanza, but with the difference in readout noted above.</li>\n<li><code>DMS_M2_seq</code> are DMS experiments in which a sequence and an ensemble of mutants are separately probed (so-called ‘mutate-and-map-seq’ or <a href=\"https://doi.org/10.1073/pnas.1619897114\" target=\"_blank\">M2-seq</a> experiments).  In these data sets, the presence of a mutation at a given position is recorded, and the average reactivity at all other mutations is inferred through mutational profiling. In this <code>rmdb_data.csv</code>, the sequences of the different profiles are adjusted so that the site of one mutation (recorded as ‘X’ in the original RMDB data) has been mutated to its complement (<code>A</code> to <code>U</code>, <code>U</code> to <code>A</code>, <code>G</code> to <code>C</code>, <code>C</code> to <code>G</code>). You may wish to mask the positions around this mutation site; and you may want to try the other two possible mutations as sequences and average results. In addition, the very beginning and last 20 nucleotides or so are not properly probed in this method, so you may also want to mask those out.</li>\n<li><code>CMCT</code> is a chemical modifier that hits unstructured G’s and U’s.</li>\n<li><code>BzCN_cotx</code> and <code>DMS_cotx</code> involve so-called <a href=\"https://doi.org/10.1038/nsmb.3316\" target=\"_blank\">‘cotranscriptional’</a> SHAPE or DMS experiments  where RNA molecules are probed after they are transcribed but without any heating/cooling to try to put them in their thermodynamically most favored state.</li>\n<li>`deg_Mg_50C','deg_Mg_pH10','deg_50C', and ‘deg_pH10' are the data for RNA degradation in four separate conditions used as train and test set in the <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine\" target=\"_blank\">OpenVaccine Kaggle challenge</a>.</li>\n<li>Some other <code>experiment_type</code> present in the RMDB, notably the <a href=\"https://doi.org/10.7554/eLife.07600\" target=\"_blank\">MOHCA-seq</a> protocol, have not been included in here because the data are not really profiles of chemical reactivity like in Ribonanza. If there is interest, please convey in this thread and it may be possible to curate these into a separate data set.</li></ul></li>\n<li>The <code>reads</code> column is empty in this dataset, since many of the experiments either do not involve high-throughput DNA sequencers or the RMDB-deposited files did not record the number of reads.</li>\n<li><code>SN_filter</code> is set based on whether <code>signal_to_noise</code>&gt;1.0. Unlike Ribonanza data, the filter does not take into account <code>reads</code>, since those values are unknown or not recorded in most cases.</li>\n<li>The <code>reactivity_*</code> columns are the data you could try to use for training. These have been scaled compared to the original RMDB data sets in the same way that the Ribonanza data have been scaled, i.e. so that the 90th percentile values across each data set is 1.0.</li>\n<li>The <code>reactivity_error_*</code> columns are not always present in the original RMDB entries but have been inferred here based on either raw data present or a heuristic formula that approximates experimental errors well for older data that use a low-throughput experimental readout based on capillary electrophoresis (<code>reactivity_error = 0.1 x mean(reactivity) + 0.2 x reactivity</code>).</li>\n<li>There are some longer RNA’s in RMDB which do not contribute many data but make the CSV’s unwieldy. Only sequences with length of 512 nucleotides or less have been included in this data set. </li>\n</ul>\n<p>Thanks to <a href=\"https://www.kaggle.com/brainbow\" target=\"_blank\">@brainbow</a> and <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> for helping with curation and testing.</p>\n<p>Please post comments here with questions or if you see bugs.</p>",
      "rawMarkdown": "With about a month to go, we hope you are enjoying the Ribonanza competition and are beginning to look for and test new innovations!\n\nFor folks who are curious about external data sources, the [RNA Mapping DataBase](https://rmdb.stanford.edu) holds numerous chemical mapping profiles for a wide range of RNA molecules, collected by several laboratories.  \n\nThese public data have been available since the beginning of the Ribonanza competition through RMDB’s [Download All link](https://rmdb.stanford.edu/site_data/published_rdat.zip),  but the data are in a specialized `RDAT` format. To make it easier to test out use of these diverse data, we’ve wrangled the majority of these data into a CSV file with identical formatting to the Ribonanza `train_data.csv` format.\n\nYou can find the `rmdb_data.csv` data set, involving 142,300 profiles, here:\n\nhttps://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data\n\nSome additional notes:\n- The `experiment_type` field has 12 values for the different kinds of chemical mapping experiments in the dataset.\n    - Importantly, most of the RMDB data are based on experimental read outs different from the mutational profiling (‘MaP’) experiments in the Ribonanza data set. While a technical detail, this means that the RMDB data will not exactly  agree with the Ribonanza data for any molecule. If you are using neural networks, you may wish to create new ‘heads’ of the network for each of the reaction types. For the aficionados: most RMDB measurements measure the termination of reverse transcription at nucleotides that have been modified or cleaved by chemical reactions, rather than the special mutational read through of chemical modifications that is used in Ribonanza.\n    - `1M7`,`NMIA`,`BzCN` are all SHAPE modifiers that selectively tag parts of the RNA molecules that are unstructured whether they are A, C, G, or U. These are closest to the `2A3_MaP` `experiment_type` in the Ribonanza data set, which is also based on a SHAPE chemical reaction, although with a different chemical that gives higher signal to noise and a different readout, as noted above. \n    - `DMS` data in RMDB involve the same dimethyl sulfate reaction as the `DMS_MaP` measurements in Ribonanza, but with the difference in readout noted above.\n    - `DMS_M2_seq` are DMS experiments in which a sequence and an ensemble of mutants are separately probed (so-called ‘mutate-and-map-seq’ or [M2-seq](https://doi.org/10.1073/pnas.1619897114) experiments).  In these data sets, the presence of a mutation at a given position is recorded, and the average reactivity at all other mutations is inferred through mutational profiling. In this `rmdb_data.csv`, the sequences of the different profiles are adjusted so that the site of one mutation (recorded as ‘X’ in the original RMDB data) has been mutated to its complement (`A` to `U`, `U` to `A`, `G` to `C`, `C` to `G`). You may wish to mask the positions around this mutation site; and you may want to try the other two possible mutations as sequences and average results. In addition, the very beginning and last 20 nucleotides or so are not properly probed in this method, so you may also want to mask those out.\n    - `CMCT` is a chemical modifier that hits unstructured G’s and U’s.\n    -  `BzCN_cotx` and `DMS_cotx` involve so-called [‘cotranscriptional’](https://doi.org/10.1038/nsmb.3316) SHAPE or DMS experiments  where RNA molecules are probed after they are transcribed but without any heating/cooling to try to put them in their thermodynamically most favored state.\n    - `deg_Mg_50C','deg_Mg_pH10','deg_50C', and ‘deg_pH10' are the data for RNA degradation in four separate conditions used as train and test set in the [OpenVaccine Kaggle challenge](https://www.kaggle.com/competitions/stanford-covid-vaccine).\n    - Some other `experiment_type` present in the RMDB, notably the [MOHCA-seq](https://doi.org/10.7554/eLife.07600) protocol, have not been included in here because the data are not really profiles of chemical reactivity like in Ribonanza. If there is interest, please convey in this thread and it may be possible to curate these into a separate data set.\n- The `reads` column is empty in this dataset, since many of the experiments either do not involve high-throughput DNA sequencers or the RMDB-deposited files did not record the number of reads.\n- `SN_filter` is set based on whether `signal_to_noise`>1.0. Unlike Ribonanza data, the filter does not take into account `reads`, since those values are unknown or not recorded in most cases.\n- The `reactivity_*` columns are the data you could try to use for training. These have been scaled compared to the original RMDB data sets in the same way that the Ribonanza data have been scaled, i.e. so that the 90th percentile values across each data set is 1.0.\n- The `reactivity_error_*` columns are not always present in the original RMDB entries but have been inferred here based on either raw data present or a heuristic formula that approximates experimental errors well for older data that use a low-throughput experimental readout based on capillary electrophoresis (`reactivity_error = 0.1 x mean(reactivity) + 0.2 x reactivity`).\n- There are some longer RNA’s in RMDB which do not contribute many data but make the CSV’s unwieldy. Only sequences with length of 512 nucleotides or less have been included in this data set. \n\nThanks to @brainbow and @shujun717 for helping with curation and testing.\n\nPlease post comments here with questions or if you see bugs.",
      "votes": null
    },
    {
      "id": "2520167",
      "postDate": "11/10/2023 16:14:10",
      "content": "<p>Thanks for sharing.</p>\n<p>Pretty excited to get some sequences with reactivity values for the initial positions, the longer sequence length also pretty big.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff35b56aa8c98a8c144bbc14a4cdafca8%2FScreenshot%20from%202023-11-10%2016-45-45.png?generation=1699631559890912&amp;alt=media\" alt=\"\"></p>\n<p>433 max sequence length for individual experiments, 342 max for sequences that are in both.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Fe05ed62656f8d4b5f399634a6c0bb37a%2FScreenshot%20from%202023-11-10%2017-09-16.png?generation=1699632606187247&amp;alt=media\" alt=\"\"></p>\n<p>4403 unique sequences with length over 207.</p>\n<hr>\n<p>Out of curiosity, any explanation to why are the error values so high in the first position in some experiments?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0a4758c6393b0682c9bab35d0f0f5365%2FScreenshot%20from%202023-11-10%2017-25-58.png?generation=1699633653615570&amp;alt=media\" alt=\"\"></p>\n<p>Thanks.</p>",
      "rawMarkdown": "Thanks for sharing.\n\nPretty excited to get some sequences with reactivity values for the initial positions, the longer sequence length also pretty big.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff35b56aa8c98a8c144bbc14a4cdafca8%2FScreenshot%20from%202023-11-10%2016-45-45.png?generation=1699631559890912&alt=media)\n\n433 max sequence length for individual experiments, 342 max for sequences that are in both.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Fe05ed62656f8d4b5f399634a6c0bb37a%2FScreenshot%20from%202023-11-10%2017-09-16.png?generation=1699632606187247&alt=media)\n\n4403 unique sequences with length over 207.\n\n---------------------------------------------------------------------------------------------------------------------\n\nOut of curiosity, any explanation to why are the error values so high in the first position in some experiments?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0a4758c6393b0682c9bab35d0f0f5365%2FScreenshot%20from%202023-11-10%2017-25-58.png?generation=1699633653615570&alt=media)\n\nThanks.",
      "votes": null
    },
    {
      "id": "2520234",
      "postDate": "11/10/2023 17:13:41",
      "content": "<p>Also here's a very simple example notebook to visualize some subsets of the data:</p>\n<p><a href=\"https://www.kaggle.com/rhijudas/example-rmdb-data-readin\" target=\"_blank\">https://www.kaggle.com/rhijudas/example-rmdb-data-readin</a></p>",
      "rawMarkdown": "Also here's a very simple example notebook to visualize some subsets of the data:\n\nhttps://www.kaggle.com/rhijudas/example-rmdb-data-readin",
      "votes": null
    },
    {
      "id": "2520278",
      "postDate": "11/10/2023 17:39:27",
      "content": "<p>I'm actually not sure about those <code>DMS_cotx</code> experiments, but you may want to mask those out just in case. 😉</p>",
      "rawMarkdown": "I'm actually not sure about those `DMS_cotx` experiments, but you may want to mask those out just in case. 😉",
      "votes": null
    },
    {
      "id": "2520558",
      "postDate": "11/10/2023 23:49:15",
      "content": "<p>Thanks Rhiju for sharing. I have gone through the process of converting Ubr to python and I was wondering. Would I be able to find the experimental profiles -condition sets- (the values bowtie needs)  to reproduce the reactivity for these as well as the training data? Are these sequences bar coded to identify them?</p>",
      "rawMarkdown": "Thanks Rhiju for sharing. I have gone through the process of converting Ubr to python and I was wondering. Would I be able to find the experimental profiles -condition sets- (the values bowtie needs)  to reproduce the reactivity for these as well as the training data? Are these sequences bar coded to identify them?",
      "votes": null
    },
    {
      "id": "2520651",
      "postDate": "11/11/2023 03:20:24",
      "content": "<p>Yes, for the Ribonanza data sets, we were originally thinking about making the Ribonanza raw data available in the hopes that someone  might try to dig deeply into them! </p>\n<p>In the end, however, we ran into an issue where we acquired some data sets that contained information relevant to both the Ribonanza train data and the test (leaderboard) data. We couldn't risk leakage by making the raw data available during competition.</p>\n<p>We'd love to see the Python code you put together if you have time to do a PR. </p>\n<p>And if you or others are interested in moving the science forward post-competition, we'd love to interact regarding the raw data underlying the competition when we release them in December (likely to SRA, the sequencing read archive). </p>\n<p>For the RMDB data sets linked in this thread, a few of them have raw data uploaded to SRA, but not many; and the ones that do have mainly not been processed with the UBR pipeline.</p>",
      "rawMarkdown": "Yes, for the Ribonanza data sets, we were originally thinking about making the Ribonanza raw data available in the hopes that someone  might try to dig deeply into them! \n\nIn the end, however, we ran into an issue where we acquired some data sets that contained information relevant to both the Ribonanza train data and the test (leaderboard) data. We couldn't risk leakage by making the raw data available during competition.\n\nWe'd love to see the Python code you put together if you have time to do a PR. \n\nAnd if you or others are interested in moving the science forward post-competition, we'd love to interact regarding the raw data underlying the competition when we release them in December (likely to SRA, the sequencing read archive). \n\nFor the RMDB data sets linked in this thread, a few of them have raw data uploaded to SRA, but not many; and the ones that do have mainly not been processed with the UBR pipeline.",
      "votes": null
    },
    {
      "id": "2523938",
      "postDate": "11/13/2023 21:20:08",
      "content": "<p>Thanks for the reply. I am working on documented the code conversion and willing to share. Feel free to propose a channel (github, notebook, etc). I am very interested in moving this forward and I have a lot of ideas that will take longer then 17 days to complete. </p>",
      "rawMarkdown": "Thanks for the reply. I am working on documented the code conversion and willing to share. Feel free to propose a channel (github, notebook, etc). I am very interested in moving this forward and I have a lot of ideas that will take longer then 17 days to complete.",
      "votes": null
    },
    {
      "id": "2524725",
      "postDate": "11/14/2023 13:41:04",
      "content": "<p>We really enjoy the competition and of course want to make the results accessible and reproducible to the scientific community. The ability to test whether we can build a model that works on raw data (possibly better then when using preprocessed one) is also amazing, so you can rely on us.</p>",
      "rawMarkdown": "We really enjoy the competition and of course want to make the results accessible and reproducible to the scientific community. The ability to test whether we can build a model that works on raw data (possibly better then when using preprocessed one) is also amazing, so you can rely on us.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2520167,
      "author_name": "enriquezaf",
      "author_url": "",
      "post_date": "11/10/2023 16:14:10",
      "content": "<p>Thanks for sharing.</p>\n<p>Pretty excited to get some sequences with reactivity values for the initial positions, the longer sequence length also pretty big.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff35b56aa8c98a8c144bbc14a4cdafca8%2FScreenshot%20from%202023-11-10%2016-45-45.png?generation=1699631559890912&amp;alt=media\" alt=\"\"></p>\n<p>433 max sequence length for individual experiments, 342 max for sequences that are in both.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Fe05ed62656f8d4b5f399634a6c0bb37a%2FScreenshot%20from%202023-11-10%2017-09-16.png?generation=1699632606187247&amp;alt=media\" alt=\"\"></p>\n<p>4403 unique sequences with length over 207.</p>\n<hr>\n<p>Out of curiosity, any explanation to why are the error values so high in the first position in some experiments?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0a4758c6393b0682c9bab35d0f0f5365%2FScreenshot%20from%202023-11-10%2017-25-58.png?generation=1699633653615570&amp;alt=media\" alt=\"\"></p>\n<p>Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2520278,
          "author_name": "rhijudas",
          "author_url": "",
          "post_date": "11/10/2023 17:39:27",
          "content": "<p>I'm actually not sure about those <code>DMS_cotx</code> experiments, but you may want to mask those out just in case. 😉</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2520234,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "11/10/2023 17:13:41",
      "content": "<p>Also here's a very simple example notebook to visualize some subsets of the data:</p>\n<p><a href=\"https://www.kaggle.com/rhijudas/example-rmdb-data-readin\" target=\"_blank\">https://www.kaggle.com/rhijudas/example-rmdb-data-readin</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2520558,
      "author_name": "tuttlen",
      "author_url": "",
      "post_date": "11/10/2023 23:49:15",
      "content": "<p>Thanks Rhiju for sharing. I have gone through the process of converting Ubr to python and I was wondering. Would I be able to find the experimental profiles -condition sets- (the values bowtie needs)  to reproduce the reactivity for these as well as the training data? Are these sequences bar coded to identify them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2520651,
          "author_name": "rhijudas",
          "author_url": "",
          "post_date": "11/11/2023 03:20:24",
          "content": "<p>Yes, for the Ribonanza data sets, we were originally thinking about making the Ribonanza raw data available in the hopes that someone  might try to dig deeply into them! </p>\n<p>In the end, however, we ran into an issue where we acquired some data sets that contained information relevant to both the Ribonanza train data and the test (leaderboard) data. We couldn't risk leakage by making the raw data available during competition.</p>\n<p>We'd love to see the Python code you put together if you have time to do a PR. </p>\n<p>And if you or others are interested in moving the science forward post-competition, we'd love to interact regarding the raw data underlying the competition when we release them in December (likely to SRA, the sequencing read archive). </p>\n<p>For the RMDB data sets linked in this thread, a few of them have raw data uploaded to SRA, but not many; and the ones that do have mainly not been processed with the UBR pipeline.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2523938,
              "author_name": "tuttlen",
              "author_url": "",
              "post_date": "11/13/2023 21:20:08",
              "content": "<p>Thanks for the reply. I am working on documented the code conversion and willing to share. Feel free to propose a channel (github, notebook, etc). I am very interested in moving this forward and I have a lot of ideas that will take longer then 17 days to complete. </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2524725,
              "author_name": "dmitrypenzar1996",
              "author_url": "",
              "post_date": "11/14/2023 13:41:04",
              "content": "<p>We really enjoy the competition and of course want to make the results accessible and reproducible to the scientific community. The ability to test whether we can build a model that works on raw data (possibly better then when using preprocessed one) is also amazing, so you can rely on us.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2519469": "With about a month to go, we hope you are enjoying the Ribonanza competition and are beginning to look for and test new innovations!\n\nFor folks who are curious about external data sources, the [RNA Mapping DataBase](https://rmdb.stanford.edu) holds numerous chemical mapping profiles for a wide range of RNA molecules, collected by several laboratories.  \n\nThese public data have been available since the beginning of the Ribonanza competition through RMDB’s [Download All link](https://rmdb.stanford.edu/site_data/published_rdat.zip),  but the data are in a specialized `RDAT` format. To make it easier to test out use of these diverse data, we’ve wrangled the majority of these data into a CSV file with identical formatting to the Ribonanza `train_data.csv` format.\n\nYou can find the `rmdb_data.csv` data set, involving 142,300 profiles, here:\n\nhttps://www.kaggle.com/datasets/rhijudas/rmdb-rna-mapping-database-2023-data/data\n\nSome additional notes:\n- The `experiment_type` field has 12 values for the different kinds of chemical mapping experiments in the dataset.\n    - Importantly, most of the RMDB data are based on experimental read outs different from the mutational profiling (‘MaP’) experiments in the Ribonanza data set. While a technical detail, this means that the RMDB data will not exactly  agree with the Ribonanza data for any molecule. If you are using neural networks, you may wish to create new ‘heads’ of the network for each of the reaction types. For the aficionados: most RMDB measurements measure the termination of reverse transcription at nucleotides that have been modified or cleaved by chemical reactions, rather than the special mutational read through of chemical modifications that is used in Ribonanza.\n    - `1M7`,`NMIA`,`BzCN` are all SHAPE modifiers that selectively tag parts of the RNA molecules that are unstructured whether they are A, C, G, or U. These are closest to the `2A3_MaP` `experiment_type` in the Ribonanza data set, which is also based on a SHAPE chemical reaction, although with a different chemical that gives higher signal to noise and a different readout, as noted above. \n    - `DMS` data in RMDB involve the same dimethyl sulfate reaction as the `DMS_MaP` measurements in Ribonanza, but with the difference in readout noted above.\n    - `DMS_M2_seq` are DMS experiments in which a sequence and an ensemble of mutants are separately probed (so-called ‘mutate-and-map-seq’ or [M2-seq](https://doi.org/10.1073/pnas.1619897114) experiments).  In these data sets, the presence of a mutation at a given position is recorded, and the average reactivity at all other mutations is inferred through mutational profiling. In this `rmdb_data.csv`, the sequences of the different profiles are adjusted so that the site of one mutation (recorded as ‘X’ in the original RMDB data) has been mutated to its complement (`A` to `U`, `U` to `A`, `G` to `C`, `C` to `G`). You may wish to mask the positions around this mutation site; and you may want to try the other two possible mutations as sequences and average results. In addition, the very beginning and last 20 nucleotides or so are not properly probed in this method, so you may also want to mask those out.\n    - `CMCT` is a chemical modifier that hits unstructured G’s and U’s.\n    -  `BzCN_cotx` and `DMS_cotx` involve so-called [‘cotranscriptional’](https://doi.org/10.1038/nsmb.3316) SHAPE or DMS experiments  where RNA molecules are probed after they are transcribed but without any heating/cooling to try to put them in their thermodynamically most favored state.\n    - `deg_Mg_50C','deg_Mg_pH10','deg_50C', and ‘deg_pH10' are the data for RNA degradation in four separate conditions used as train and test set in the [OpenVaccine Kaggle challenge](https://www.kaggle.com/competitions/stanford-covid-vaccine).\n    - Some other `experiment_type` present in the RMDB, notably the [MOHCA-seq](https://doi.org/10.7554/eLife.07600) protocol, have not been included in here because the data are not really profiles of chemical reactivity like in Ribonanza. If there is interest, please convey in this thread and it may be possible to curate these into a separate data set.\n- The `reads` column is empty in this dataset, since many of the experiments either do not involve high-throughput DNA sequencers or the RMDB-deposited files did not record the number of reads.\n- `SN_filter` is set based on whether `signal_to_noise`>1.0. Unlike Ribonanza data, the filter does not take into account `reads`, since those values are unknown or not recorded in most cases.\n- The `reactivity_*` columns are the data you could try to use for training. These have been scaled compared to the original RMDB data sets in the same way that the Ribonanza data have been scaled, i.e. so that the 90th percentile values across each data set is 1.0.\n- The `reactivity_error_*` columns are not always present in the original RMDB entries but have been inferred here based on either raw data present or a heuristic formula that approximates experimental errors well for older data that use a low-throughput experimental readout based on capillary electrophoresis (`reactivity_error = 0.1 x mean(reactivity) + 0.2 x reactivity`).\n- There are some longer RNA’s in RMDB which do not contribute many data but make the CSV’s unwieldy. Only sequences with length of 512 nucleotides or less have been included in this data set. \n\nThanks to @brainbow and @shujun717 for helping with curation and testing.\n\nPlease post comments here with questions or if you see bugs.",
    "2520167": "Thanks for sharing.\n\nPretty excited to get some sequences with reactivity values for the initial positions, the longer sequence length also pretty big.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff35b56aa8c98a8c144bbc14a4cdafca8%2FScreenshot%20from%202023-11-10%2016-45-45.png?generation=1699631559890912&alt=media)\n\n433 max sequence length for individual experiments, 342 max for sequences that are in both.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Fe05ed62656f8d4b5f399634a6c0bb37a%2FScreenshot%20from%202023-11-10%2017-09-16.png?generation=1699632606187247&alt=media)\n\n4403 unique sequences with length over 207.\n\n---------------------------------------------------------------------------------------------------------------------\n\nOut of curiosity, any explanation to why are the error values so high in the first position in some experiments?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0a4758c6393b0682c9bab35d0f0f5365%2FScreenshot%20from%202023-11-10%2017-25-58.png?generation=1699633653615570&alt=media)\n\nThanks.",
    "2520234": "Also here's a very simple example notebook to visualize some subsets of the data:\n\nhttps://www.kaggle.com/rhijudas/example-rmdb-data-readin",
    "2520278": "I'm actually not sure about those `DMS_cotx` experiments, but you may want to mask those out just in case. 😉",
    "2520558": "Thanks Rhiju for sharing. I have gone through the process of converting Ubr to python and I was wondering. Would I be able to find the experimental profiles -condition sets- (the values bowtie needs)  to reproduce the reactivity for these as well as the training data? Are these sequences bar coded to identify them?",
    "2520651": "Yes, for the Ribonanza data sets, we were originally thinking about making the Ribonanza raw data available in the hopes that someone  might try to dig deeply into them! \n\nIn the end, however, we ran into an issue where we acquired some data sets that contained information relevant to both the Ribonanza train data and the test (leaderboard) data. We couldn't risk leakage by making the raw data available during competition.\n\nWe'd love to see the Python code you put together if you have time to do a PR. \n\nAnd if you or others are interested in moving the science forward post-competition, we'd love to interact regarding the raw data underlying the competition when we release them in December (likely to SRA, the sequencing read archive). \n\nFor the RMDB data sets linked in this thread, a few of them have raw data uploaded to SRA, but not many; and the ones that do have mainly not been processed with the UBR pipeline.",
    "2523938": "Thanks for the reply. I am working on documented the code conversion and willing to share. Feel free to propose a channel (github, notebook, etc). I am very interested in moving this forward and I have a lot of ideas that will take longer then 17 days to complete.",
    "2524725": "We really enjoy the competition and of course want to make the results accessible and reproducible to the scientific community. The ability to test whether we can build a model that works on raw data (possibly better then when using preprocessed one) is also amazing, so you can rely on us."
  },
  "source": "meta"
}