{
  "id": 444368,
  "title": "Sequence Reliability",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/444368",
  "author_name": "",
  "post_date": "2023-10-01T14:30:45.804149900Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Is every single sequence itself absolutely 100% reliable regardless whether it is in the test set or training set, and regardless what its values (SN_filter, signal-to-noise, reactivity scores, anything) are?</p>\n<p>By \"every single sequence itself is absolutely 100% reliable\" I mean, for example if a data point from our .csv file reads \"GGGAACGACUCGAGUAGAGUCGAAAACAUUCCCAAAUUCCACCUUGGUGAUGGCACCCGGAGAGGAGCCAUCACCACACAAAUUUCGAUUUGUGAAAAGAAACAACAACAACAAC\", then that data point literally and physically consists of that exact sequence of nucleotide when going through your experiment and data preparation, as opposed to say some of the nucleotides shown above are wrong (I.E. different from what was physically present) due to some inherent measurement error/noise?</p>",
  "messages": [
    {
      "id": "2463767",
      "postDate": "10/01/2023 14:30:45",
      "content": "<p>Is every single sequence itself absolutely 100% reliable regardless whether it is in the test set or training set, and regardless what its values (SN_filter, signal-to-noise, reactivity scores, anything) are?</p>\n<p>By \"every single sequence itself is absolutely 100% reliable\" I mean, for example if a data point from our .csv file reads \"GGGAACGACUCGAGUAGAGUCGAAAACAUUCCCAAAUUCCACCUUGGUGAUGGCACCCGGAGAGGAGCCAUCACCACACAAAUUUCGAUUUGUGAAAAGAAACAACAACAACAAC\", then that data point literally and physically consists of that exact sequence of nucleotide when going through your experiment and data preparation, as opposed to say some of the nucleotides shown above are wrong (I.E. different from what was physically present) due to some inherent measurement error/noise?</p>",
      "rawMarkdown": "Is every single sequence itself absolutely 100% reliable regardless whether it is in the test set or training set, and regardless what its values (SN_filter, signal-to-noise, reactivity scores, anything) are?\n\nBy \"every single sequence itself is absolutely 100% reliable\" I mean, for example if a data point from our .csv file reads \"GGGAACGACUCGAGUAGAGUCGAAAACAUUCCCAAAUUCCACCUUGGUGAUGGCACCCGGAGAGGAGCCAUCACCACACAAAUUUCGAUUUGUGAAAAGAAACAACAACAACAAC\", then that data point literally and physically consists of that exact sequence of nucleotide when going through your experiment and data preparation, as opposed to say some of the nucleotides shown above are wrong (I.E. different from what was physically present) due to some inherent measurement error/noise?",
      "votes": null
    },
    {
      "id": "2471872",
      "postDate": "10/06/2023 17:38:52",
      "content": "<p>The error rate of chemical DNA synthesis is &lt;0.1% per nucleotide, which means that for a 200-nucleotide RNA, &gt;98% of the sequences are exactly what is listed in the test and training set .csv's. =) </p>\n<p>For a few positions (probably &lt;1% of positions), there are some 'hotspots' for which T7 RNA polymerase is prone to make mistakes at greater than that error rate of 0.1%. (T7 RNA polymerase is the little machine we use to take DNA templates purchased from a company to 'transcribe' RNA.) But again even with these errors, we expect the vast majority of RNA's to be exactly what you see in the .csv's. </p>",
      "rawMarkdown": "The error rate of chemical DNA synthesis is <0.1% per nucleotide, which means that for a 200-nucleotide RNA, >98% of the sequences are exactly what is listed in the test and training set .csv's. =) \n\nFor a few positions (probably <1% of positions), there are some 'hotspots' for which T7 RNA polymerase is prone to make mistakes at greater than that error rate of 0.1%. (T7 RNA polymerase is the little machine we use to take DNA templates purchased from a company to 'transcribe' RNA.) But again even with these errors, we expect the vast majority of RNA's to be exactly what you see in the .csv's.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2471872,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "10/06/2023 17:38:52",
      "content": "<p>The error rate of chemical DNA synthesis is &lt;0.1% per nucleotide, which means that for a 200-nucleotide RNA, &gt;98% of the sequences are exactly what is listed in the test and training set .csv's. =) </p>\n<p>For a few positions (probably &lt;1% of positions), there are some 'hotspots' for which T7 RNA polymerase is prone to make mistakes at greater than that error rate of 0.1%. (T7 RNA polymerase is the little machine we use to take DNA templates purchased from a company to 'transcribe' RNA.) But again even with these errors, we expect the vast majority of RNA's to be exactly what you see in the .csv's. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2463767": "Is every single sequence itself absolutely 100% reliable regardless whether it is in the test set or training set, and regardless what its values (SN_filter, signal-to-noise, reactivity scores, anything) are?\n\nBy \"every single sequence itself is absolutely 100% reliable\" I mean, for example if a data point from our .csv file reads \"GGGAACGACUCGAGUAGAGUCGAAAACAUUCCCAAAUUCCACCUUGGUGAUGGCACCCGGAGAGGAGCCAUCACCACACAAAUUUCGAUUUGUGAAAAGAAACAACAACAACAAC\", then that data point literally and physically consists of that exact sequence of nucleotide when going through your experiment and data preparation, as opposed to say some of the nucleotides shown above are wrong (I.E. different from what was physically present) due to some inherent measurement error/noise?",
    "2471872": "The error rate of chemical DNA synthesis is <0.1% per nucleotide, which means that for a 200-nucleotide RNA, >98% of the sequences are exactly what is listed in the test and training set .csv's. =) \n\nFor a few positions (probably <1% of positions), there are some 'hotspots' for which T7 RNA polymerase is prone to make mistakes at greater than that error rate of 0.1%. (T7 RNA polymerase is the little machine we use to take DNA templates purchased from a company to 'transcribe' RNA.) But again even with these errors, we expect the vast majority of RNA's to be exactly what you see in the .csv's."
  },
  "source": "meta"
}