{
  "id": 445415,
  "title": "Experimental protocol behind the data",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/445415",
  "author_name": "",
  "post_date": "2023-10-06T20:42:55.577248900Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We've had a lot of questions about the experimental protocol that generated this dataset, so we'll try to give a more in-depth explanation for how this data got to Kagglers. </p>\n<p>We start with a large set of RNA sequences. We got these sequences from a range of sources, including players of the <a href=\"https://eternagame.org/\" target=\"_blank\">Eterna</a> online citizen science game. We generate DNA sequences that code for those RNA sequences, and have a DNA synthesis company make it for us. When we receive the single-stranded DNA from our vendor, we run through several steps in the lab:</p>\n<ol>\n<li><strong>Emulsion PCR</strong>: This step converts the single stranded DNA into the double-stranded DNA necessary for step 2.</li>\n<li><strong>In-vitro transcription</strong>: We use the same process your body uses to create RNA, transcription, but we do it in a test tube. After our transcription step, we end up with the actual RNA of the sequences in the dataset. But how do we figure out what structure the RNA strands are forming?</li>\n<li><strong>Chemical probing</strong>: We add certain chemical reagents that will selectively bind to the RNA molecule. If a base is already paired to another base, it won't react to the chemical modifier. Unpaired bases will react, resulting in a chemical \"tag\".</li>\n<li><strong>Reverse transcription</strong>: Next, we use this modified RNA as a template for reverse transcription. The reverse transcription process allows us to make the DNA that codes for a given RNA. The chemical tags we added in step 3 cause the reverse transcription process to make mistakes and introduce mutations.</li>\n<li><strong>PAGE purification</strong>: Sometimes the chemical tags cause the reverse transcription process to stop entirely, rather than mutate the DNA. So after step 4, we have a lot of DNA of varying lengths. We only want full-length DNA, so we purify via polyacrylimide gel electrophoresis (PAGE). It separates the DNA by length, so we can pick only the DNA we want.</li>\n<li><strong>Amplification PCR</strong>: We have the right DNA, but we usually don't have many DNA molecules at this point. We use PCR to make lots of copies of each DNA which we can analyze in step 7.</li>\n<li><strong>DNA sequencing</strong>: We read the contents of the DNA with a DNA sequencing machine that can tell us what bases are in each DNA molecule. We can use the sequencing data to figure out which bases were mutated and thus which RNA bases reacted to our chemical reagents.</li>\n</ol>\n<p>Here's a walkthrough in our lab by Rui, one of our lab technicians.</p>",
  "messages": [
    {
      "id": "2472002",
      "postDate": "10/06/2023 20:42:55",
      "content": "<p>We've had a lot of questions about the experimental protocol that generated this dataset, so we'll try to give a more in-depth explanation for how this data got to Kagglers. </p>\n<p>We start with a large set of RNA sequences. We got these sequences from a range of sources, including players of the <a href=\"https://eternagame.org/\" target=\"_blank\">Eterna</a> online citizen science game. We generate DNA sequences that code for those RNA sequences, and have a DNA synthesis company make it for us. When we receive the single-stranded DNA from our vendor, we run through several steps in the lab:</p>\n<ol>\n<li><strong>Emulsion PCR</strong>: This step converts the single stranded DNA into the double-stranded DNA necessary for step 2.</li>\n<li><strong>In-vitro transcription</strong>: We use the same process your body uses to create RNA, transcription, but we do it in a test tube. After our transcription step, we end up with the actual RNA of the sequences in the dataset. But how do we figure out what structure the RNA strands are forming?</li>\n<li><strong>Chemical probing</strong>: We add certain chemical reagents that will selectively bind to the RNA molecule. If a base is already paired to another base, it won't react to the chemical modifier. Unpaired bases will react, resulting in a chemical \"tag\".</li>\n<li><strong>Reverse transcription</strong>: Next, we use this modified RNA as a template for reverse transcription. The reverse transcription process allows us to make the DNA that codes for a given RNA. The chemical tags we added in step 3 cause the reverse transcription process to make mistakes and introduce mutations.</li>\n<li><strong>PAGE purification</strong>: Sometimes the chemical tags cause the reverse transcription process to stop entirely, rather than mutate the DNA. So after step 4, we have a lot of DNA of varying lengths. We only want full-length DNA, so we purify via polyacrylimide gel electrophoresis (PAGE). It separates the DNA by length, so we can pick only the DNA we want.</li>\n<li><strong>Amplification PCR</strong>: We have the right DNA, but we usually don't have many DNA molecules at this point. We use PCR to make lots of copies of each DNA which we can analyze in step 7.</li>\n<li><strong>DNA sequencing</strong>: We read the contents of the DNA with a DNA sequencing machine that can tell us what bases are in each DNA molecule. We can use the sequencing data to figure out which bases were mutated and thus which RNA bases reacted to our chemical reagents.</li>\n</ol>\n<p>Here's a walkthrough in our lab by Rui, one of our lab technicians.</p>",
      "rawMarkdown": "We've had a lot of questions about the experimental protocol that generated this dataset, so we'll try to give a more in-depth explanation for how this data got to Kagglers. \n\nWe start with a large set of RNA sequences. We got these sequences from a range of sources, including players of the [Eterna](https://eternagame.org/) online citizen science game. We generate DNA sequences that code for those RNA sequences, and have a DNA synthesis company make it for us. When we receive the single-stranded DNA from our vendor, we run through several steps in the lab:\n\n1.  **Emulsion PCR**: This step converts the single stranded DNA into the double-stranded DNA necessary for step 2.\n2. **In-vitro transcription**: We use the same process your body uses to create RNA, transcription, but we do it in a test tube. After our transcription step, we end up with the actual RNA of the sequences in the dataset. But how do we figure out what structure the RNA strands are forming?\n3. **Chemical probing**: We add certain chemical reagents that will selectively bind to the RNA molecule. If a base is already paired to another base, it won't react to the chemical modifier. Unpaired bases will react, resulting in a chemical \"tag\".\n4. **Reverse transcription**: Next, we use this modified RNA as a template for reverse transcription. The reverse transcription process allows us to make the DNA that codes for a given RNA. The chemical tags we added in step 3 cause the reverse transcription process to make mistakes and introduce mutations.\n5. **PAGE purification**: Sometimes the chemical tags cause the reverse transcription process to stop entirely, rather than mutate the DNA. So after step 4, we have a lot of DNA of varying lengths. We only want full-length DNA, so we purify via polyacrylimide gel electrophoresis (PAGE). It separates the DNA by length, so we can pick only the DNA we want.\n6. **Amplification PCR**: We have the right DNA, but we usually don't have many DNA molecules at this point. We use PCR to make lots of copies of each DNA which we can analyze in step 7.\n7. **DNA sequencing**: We read the contents of the DNA with a DNA sequencing machine that can tell us what bases are in each DNA molecule. We can use the sequencing data to figure out which bases were mutated and thus which RNA bases reacted to our chemical reagents.\n\nHere's a walkthrough in our lab by Rui, one of our lab technicians.\n<iframe width=\"560\" height=\"315\" src=\"https://www.youtube.com/embed/ohqxF_9H-_k?si=cd1tYhKlbIUPKhVk\" title=\"YouTube video player\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" allowfullscreen></iframe>",
      "votes": null
    },
    {
      "id": "2488114",
      "postDate": "10/19/2023 03:37:47",
      "content": "<p>Would be nice to know the what the RNA structure information actually is. Where does the reactivity factor come into play?</p>",
      "rawMarkdown": "Would be nice to know the what the RNA structure information actually is. Where does the reactivity factor come into play?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2488114,
      "author_name": "tuttlen",
      "author_url": "",
      "post_date": "10/19/2023 03:37:47",
      "content": "<p>Would be nice to know the what the RNA structure information actually is. Where does the reactivity factor come into play?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2472002": "We've had a lot of questions about the experimental protocol that generated this dataset, so we'll try to give a more in-depth explanation for how this data got to Kagglers. \n\nWe start with a large set of RNA sequences. We got these sequences from a range of sources, including players of the [Eterna](https://eternagame.org/) online citizen science game. We generate DNA sequences that code for those RNA sequences, and have a DNA synthesis company make it for us. When we receive the single-stranded DNA from our vendor, we run through several steps in the lab:\n\n1.  **Emulsion PCR**: This step converts the single stranded DNA into the double-stranded DNA necessary for step 2.\n2. **In-vitro transcription**: We use the same process your body uses to create RNA, transcription, but we do it in a test tube. After our transcription step, we end up with the actual RNA of the sequences in the dataset. But how do we figure out what structure the RNA strands are forming?\n3. **Chemical probing**: We add certain chemical reagents that will selectively bind to the RNA molecule. If a base is already paired to another base, it won't react to the chemical modifier. Unpaired bases will react, resulting in a chemical \"tag\".\n4. **Reverse transcription**: Next, we use this modified RNA as a template for reverse transcription. The reverse transcription process allows us to make the DNA that codes for a given RNA. The chemical tags we added in step 3 cause the reverse transcription process to make mistakes and introduce mutations.\n5. **PAGE purification**: Sometimes the chemical tags cause the reverse transcription process to stop entirely, rather than mutate the DNA. So after step 4, we have a lot of DNA of varying lengths. We only want full-length DNA, so we purify via polyacrylimide gel electrophoresis (PAGE). It separates the DNA by length, so we can pick only the DNA we want.\n6. **Amplification PCR**: We have the right DNA, but we usually don't have many DNA molecules at this point. We use PCR to make lots of copies of each DNA which we can analyze in step 7.\n7. **DNA sequencing**: We read the contents of the DNA with a DNA sequencing machine that can tell us what bases are in each DNA molecule. We can use the sequencing data to figure out which bases were mutated and thus which RNA bases reacted to our chemical reagents.\n\nHere's a walkthrough in our lab by Rui, one of our lab technicians.\n<iframe width=\"560\" height=\"315\" src=\"https://www.youtube.com/embed/ohqxF_9H-_k?si=cd1tYhKlbIUPKhVk\" title=\"YouTube video player\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" allowfullscreen></iframe>",
    "2488114": "Would be nice to know the what the RNA structure information actually is. Where does the reactivity factor come into play?"
  },
  "source": "meta"
}