{
  "id": 441720,
  "title": "a long list of questions",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/441720",
  "author_name": "",
  "post_date": "2023-09-19T21:50:13.149015100Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>General Questions</p>\n<ol>\n<li>Are the Ribonanza RNA sequences designed by humans or are some real sequences like nature produce them? </li>\n<li>Are all of the sequences produced in the same environment?</li>\n<li>Are all sequences produced synthetic or did some get extracted from \"nature\"?</li>\n<li>Were different machines used to collect the data?</li>\n<li>If 4. == \"yes\": Different types of machines? like different manufacturers or so?<br>\nIm just wondering about secondary errors that might have influenced the data, even if it's unlikely to figure those out. </li>\n</ol>\n<p>Data Questions</p>\n<ol>\n<li>Why are there ROWS with all reactivity null?<br>\nI mean, why would i want sequences with no data in the train data?</li>\n<li>How did the data get collected if <code>reads</code> are 0?</li>\n<li>How did the data get collected if <code>signal_to_noise</code> is 0?</li>\n</ol>\n<p>Submission<br>\nWhats the influence of the values for the first and last few of the sequence to the scoring?<br>\nShould those be null or can they be any predicted number?</p>",
  "messages": [
    {
      "id": "2447152",
      "postDate": "09/19/2023 21:50:13",
      "content": "<p>General Questions</p>\n<ol>\n<li>Are the Ribonanza RNA sequences designed by humans or are some real sequences like nature produce them? </li>\n<li>Are all of the sequences produced in the same environment?</li>\n<li>Are all sequences produced synthetic or did some get extracted from \"nature\"?</li>\n<li>Were different machines used to collect the data?</li>\n<li>If 4. == \"yes\": Different types of machines? like different manufacturers or so?<br>\nIm just wondering about secondary errors that might have influenced the data, even if it's unlikely to figure those out. </li>\n</ol>\n<p>Data Questions</p>\n<ol>\n<li>Why are there ROWS with all reactivity null?<br>\nI mean, why would i want sequences with no data in the train data?</li>\n<li>How did the data get collected if <code>reads</code> are 0?</li>\n<li>How did the data get collected if <code>signal_to_noise</code> is 0?</li>\n</ol>\n<p>Submission<br>\nWhats the influence of the values for the first and last few of the sequence to the scoring?<br>\nShould those be null or can they be any predicted number?</p>",
      "rawMarkdown": "General Questions\n1. Are the Ribonanza RNA sequences designed by humans or are some real sequences like nature produce them? \n2. Are all of the sequences produced in the same environment?\n3. Are all sequences produced synthetic or did some get extracted from \"nature\"?\n4. Were different machines used to collect the data?\n5. If 4. == \"yes\": Different types of machines? like different manufacturers or so?\nIm just wondering about secondary errors that might have influenced the data, even if it's unlikely to figure those out. \n\nData Questions\n6. Why are there ROWS with all reactivity null?\nI mean, why would i want sequences with no data in the train data?\n7. How did the data get collected if `reads` are 0?\n8. How did the data get collected if `signal_to_noise` is 0?\n\nSubmission\nWhats the influence of the values for the first and last few of the sequence to the scoring?\nShould those be null or can they be any predicted number?",
      "votes": null
    },
    {
      "id": "2448447",
      "postDate": "09/20/2023 15:41:14",
      "content": "<p>Good questions.</p>\n<blockquote>\n  <p>Whats the influence of the values for the first and last few of the sequence to the scoring? Should those be null or can they be any predicted number?</p>\n</blockquote>\n<p>Do you mean 26 first characters and if I remember up to 39 at the end of each sequence? <br>\nIn Overview it's explained those values do not count. Tested and it seems to be true. 0.25 and 0.0 for those \"head and tail\" characters gave the same score having values for middle characters the same in both submissions.</p>",
      "rawMarkdown": "Good questions.\n> Whats the influence of the values for the first and last few of the sequence to the scoring? Should those be null or can they be any predicted number?\n\nDo you mean 26 first characters and if I remember up to 39 at the end of each sequence? \nIn Overview it's explained those values do not count. Tested and it seems to be true. 0.25 and 0.0 for those \"head and tail\" characters gave the same score having values for middle characters the same in both submissions.",
      "votes": null
    },
    {
      "id": "2448501",
      "postDate": "09/20/2023 16:11:26",
      "content": "<p>Yes, i mean those.<br>\nThank you for trying it out and sharing. :D</p>\n<p>and yea, some questions are more out of curiosity and for later, if i run into anomalies. At one of the videos linked here, they pointed out, that it makes little differences if the RNA is designed by humans or natural. My first guess would be, lower errors on designed once. But im not an expert, and got almost 0 domain knowledge.</p>",
      "rawMarkdown": "Yes, i mean those.\nThank you for trying it out and sharing. :D\n\nand yea, some questions are more out of curiosity and for later, if i run into anomalies. At one of the videos linked here, they pointed out, that it makes little differences if the RNA is designed by humans or natural. My first guess would be, lower errors on designed once. But im not an expert, and got almost 0 domain knowledge.",
      "votes": null
    },
    {
      "id": "2451584",
      "postDate": "09/22/2023 15:45:11",
      "content": "<p>Let me see if I can answer these:</p>\n<ol>\n<li>The dataset has synthetically designed as well as naturally sourced sequences.</li>\n<li>The sequences are all produced via the same experimental procedure, although the data was gathered in  several real-world experiments.</li>\n<li>All sequences are synthetic; none of the sequences tested came from a natural source (e.g., extracted directly from a cell).</li>\n<li>Not sure what you mean by different machines here. The experimental procedure is a little complicated and we've had a lot of questions on how the data was gathered, so we're prepping an explanation thread.</li>\n<li>Sources of error: the main errors that we provide are based on statistical errors from sequencing experiments. Other sources of errors include things like source of the DNA from different vendors, human errors in doing the experiments, temperature variations, etc. We have previous data on pilot libraries that indicate that those sources of error are smaller than statistical errors; unfortunately not all those data are publicly available yet, though some are described in this tweet: <a href=\"https://twitter.com/RDasLab/status/1684643146016382976\" target=\"_blank\">https://twitter.com/RDasLab/status/1684643146016382976</a></li>\n</ol>\n<p>Data Questions: </p>\n<ol>\n<li>We included purely null data for ‘dropout’ sequences to help us do cross checks, e.g., reconciling numbers of sequences in the sequence library FASTA files and number of sequences in train_data.csv.</li>\n<li>We'll address the data preparation in our experimental explanation thread.</li>\n<li>We'll address the data preparation in our experimental explanation thread.</li>\n</ol>\n<p>Submission:<br>\nAs <a href=\"https://www.kaggle.com/hotbit\" target=\"_blank\">@hotbit</a> notes, the first and last bases have no reactivity data because of some details of the experimental procedure. They are not evaluated in the scoring metric.</p>",
      "rawMarkdown": "Let me see if I can answer these:\n1. The dataset has synthetically designed as well as naturally sourced sequences.\n2. The sequences are all produced via the same experimental procedure, although the data was gathered in  several real-world experiments.\n3. All sequences are synthetic; none of the sequences tested came from a natural source (e.g., extracted directly from a cell).\n4. Not sure what you mean by different machines here. The experimental procedure is a little complicated and we've had a lot of questions on how the data was gathered, so we're prepping an explanation thread.\n5. Sources of error: the main errors that we provide are based on statistical errors from sequencing experiments. Other sources of errors include things like source of the DNA from different vendors, human errors in doing the experiments, temperature variations, etc. We have previous data on pilot libraries that indicate that those sources of error are smaller than statistical errors; unfortunately not all those data are publicly available yet, though some are described in this tweet: https://twitter.com/RDasLab/status/1684643146016382976\n\nData Questions: \n1. We included purely null data for ‘dropout’ sequences to help us do cross checks, e.g., reconciling numbers of sequences in the sequence library FASTA files and number of sequences in train_data.csv.\n2. We'll address the data preparation in our experimental explanation thread.\n3. We'll address the data preparation in our experimental explanation thread.\n\nSubmission:\nAs @hotbit notes, the first and last bases have no reactivity data because of some details of the experimental procedure. They are not evaluated in the scoring metric.",
      "votes": null
    },
    {
      "id": "2451638",
      "postDate": "09/22/2023 16:39:25",
      "content": "<p>Thank you very much for the detailed answers!</p>\n<p>Answer 2 kinda answering my 4. Different machines used is most likely different labs, persons…. enough different things, that could go wrong, so i don't need more details. 😁</p>\n<p>I'm looking forward to the data preparation in our experimental explanation thread!</p>\n<p>Maybe i do understand than, why there are some sequences in train AND test data, but Sample_submission + values from Train.csv of those sequence gives even less score than the empty Sample_submission.csv. 😉</p>",
      "rawMarkdown": "Thank you very much for the detailed answers!\n\nAnswer 2 kinda answering my 4. Different machines used is most likely different labs, persons.... enough different things, that could go wrong, so i don't need more details. 😁\n\nI'm looking forward to the data preparation in our experimental explanation thread!\n\nMaybe i do understand than, why there are some sequences in train AND test data, but Sample_submission + values from Train.csv of those sequence gives even less score than the empty Sample_submission.csv. 😉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2448447,
      "author_name": "hotbit",
      "author_url": "",
      "post_date": "09/20/2023 15:41:14",
      "content": "<p>Good questions.</p>\n<blockquote>\n  <p>Whats the influence of the values for the first and last few of the sequence to the scoring? Should those be null or can they be any predicted number?</p>\n</blockquote>\n<p>Do you mean 26 first characters and if I remember up to 39 at the end of each sequence? <br>\nIn Overview it's explained those values do not count. Tested and it seems to be true. 0.25 and 0.0 for those \"head and tail\" characters gave the same score having values for middle characters the same in both submissions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2448501,
          "author_name": "jimmyknox",
          "author_url": "",
          "post_date": "09/20/2023 16:11:26",
          "content": "<p>Yes, i mean those.<br>\nThank you for trying it out and sharing. :D</p>\n<p>and yea, some questions are more out of curiosity and for later, if i run into anomalies. At one of the videos linked here, they pointed out, that it makes little differences if the RNA is designed by humans or natural. My first guess would be, lower errors on designed once. But im not an expert, and got almost 0 domain knowledge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2451584,
      "author_name": "brainbowrna",
      "author_url": "",
      "post_date": "09/22/2023 15:45:11",
      "content": "<p>Let me see if I can answer these:</p>\n<ol>\n<li>The dataset has synthetically designed as well as naturally sourced sequences.</li>\n<li>The sequences are all produced via the same experimental procedure, although the data was gathered in  several real-world experiments.</li>\n<li>All sequences are synthetic; none of the sequences tested came from a natural source (e.g., extracted directly from a cell).</li>\n<li>Not sure what you mean by different machines here. The experimental procedure is a little complicated and we've had a lot of questions on how the data was gathered, so we're prepping an explanation thread.</li>\n<li>Sources of error: the main errors that we provide are based on statistical errors from sequencing experiments. Other sources of errors include things like source of the DNA from different vendors, human errors in doing the experiments, temperature variations, etc. We have previous data on pilot libraries that indicate that those sources of error are smaller than statistical errors; unfortunately not all those data are publicly available yet, though some are described in this tweet: <a href=\"https://twitter.com/RDasLab/status/1684643146016382976\" target=\"_blank\">https://twitter.com/RDasLab/status/1684643146016382976</a></li>\n</ol>\n<p>Data Questions: </p>\n<ol>\n<li>We included purely null data for ‘dropout’ sequences to help us do cross checks, e.g., reconciling numbers of sequences in the sequence library FASTA files and number of sequences in train_data.csv.</li>\n<li>We'll address the data preparation in our experimental explanation thread.</li>\n<li>We'll address the data preparation in our experimental explanation thread.</li>\n</ol>\n<p>Submission:<br>\nAs <a href=\"https://www.kaggle.com/hotbit\" target=\"_blank\">@hotbit</a> notes, the first and last bases have no reactivity data because of some details of the experimental procedure. They are not evaluated in the scoring metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2451638,
          "author_name": "jimmyknox",
          "author_url": "",
          "post_date": "09/22/2023 16:39:25",
          "content": "<p>Thank you very much for the detailed answers!</p>\n<p>Answer 2 kinda answering my 4. Different machines used is most likely different labs, persons…. enough different things, that could go wrong, so i don't need more details. 😁</p>\n<p>I'm looking forward to the data preparation in our experimental explanation thread!</p>\n<p>Maybe i do understand than, why there are some sequences in train AND test data, but Sample_submission + values from Train.csv of those sequence gives even less score than the empty Sample_submission.csv. 😉</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2447152": "General Questions\n1. Are the Ribonanza RNA sequences designed by humans or are some real sequences like nature produce them? \n2. Are all of the sequences produced in the same environment?\n3. Are all sequences produced synthetic or did some get extracted from \"nature\"?\n4. Were different machines used to collect the data?\n5. If 4. == \"yes\": Different types of machines? like different manufacturers or so?\nIm just wondering about secondary errors that might have influenced the data, even if it's unlikely to figure those out. \n\nData Questions\n6. Why are there ROWS with all reactivity null?\nI mean, why would i want sequences with no data in the train data?\n7. How did the data get collected if `reads` are 0?\n8. How did the data get collected if `signal_to_noise` is 0?\n\nSubmission\nWhats the influence of the values for the first and last few of the sequence to the scoring?\nShould those be null or can they be any predicted number?",
    "2448447": "Good questions.\n> Whats the influence of the values for the first and last few of the sequence to the scoring? Should those be null or can they be any predicted number?\n\nDo you mean 26 first characters and if I remember up to 39 at the end of each sequence? \nIn Overview it's explained those values do not count. Tested and it seems to be true. 0.25 and 0.0 for those \"head and tail\" characters gave the same score having values for middle characters the same in both submissions.",
    "2448501": "Yes, i mean those.\nThank you for trying it out and sharing. :D\n\nand yea, some questions are more out of curiosity and for later, if i run into anomalies. At one of the videos linked here, they pointed out, that it makes little differences if the RNA is designed by humans or natural. My first guess would be, lower errors on designed once. But im not an expert, and got almost 0 domain knowledge.",
    "2451584": "Let me see if I can answer these:\n1. The dataset has synthetically designed as well as naturally sourced sequences.\n2. The sequences are all produced via the same experimental procedure, although the data was gathered in  several real-world experiments.\n3. All sequences are synthetic; none of the sequences tested came from a natural source (e.g., extracted directly from a cell).\n4. Not sure what you mean by different machines here. The experimental procedure is a little complicated and we've had a lot of questions on how the data was gathered, so we're prepping an explanation thread.\n5. Sources of error: the main errors that we provide are based on statistical errors from sequencing experiments. Other sources of errors include things like source of the DNA from different vendors, human errors in doing the experiments, temperature variations, etc. We have previous data on pilot libraries that indicate that those sources of error are smaller than statistical errors; unfortunately not all those data are publicly available yet, though some are described in this tweet: https://twitter.com/RDasLab/status/1684643146016382976\n\nData Questions: \n1. We included purely null data for ‘dropout’ sequences to help us do cross checks, e.g., reconciling numbers of sequences in the sequence library FASTA files and number of sequences in train_data.csv.\n2. We'll address the data preparation in our experimental explanation thread.\n3. We'll address the data preparation in our experimental explanation thread.\n\nSubmission:\nAs @hotbit notes, the first and last bases have no reactivity data because of some details of the experimental procedure. They are not evaluated in the scoring metric.",
    "2451638": "Thank you very much for the detailed answers!\n\nAnswer 2 kinda answering my 4. Different machines used is most likely different labs, persons.... enough different things, that could go wrong, so i don't need more details. 😁\n\nI'm looking forward to the data preparation in our experimental explanation thread!\n\nMaybe i do understand than, why there are some sequences in train AND test data, but Sample_submission + values from Train.csv of those sequence gives even less score than the empty Sample_submission.csv. 😉"
  },
  "source": "meta"
}