{
  "id": 438345,
  "title": "How To Map Reactivity to the Sequence",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/438345",
  "author_name": "",
  "post_date": "2023-09-10T17:32:56.599922200Z",
  "votes": 13,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This may be a dumb question but I'll put it out there…</p>\n<p>I would have assumed the <strong><code>sequence</code></strong> column would have the same length as the non-null reactivity values in the reactivity_0*** columns… however, this is not true. The quote from the data page indicates they <strong><em>should</em></strong> have the same length… however, it also explains that some positions may have null values due to difficulty in probing. My experimentation in the training data shows the following columns are always null:</p>\n<ul>\n<li>reactivity 1 through 26 inclusive</li>\n<li>reactivity 166 through 206 inclusive</li>\n</ul>\n<p>I may have messed something up but I didn't think so…</p>\n<blockquote>\n  <p>reactivity_0001, reactivity_0002,… - (float) An array of floating point numbers of the train data, <strong>should have the same length as the RNA sequence</strong>, which defines the reactivity profile for the RNA. For sequences shorter than the maximum RNA length, positions that go beyond the sequence length have null. Several positions early and late in the sequence also cannot be probed due to technical reasons, and their reactivity values are null.</p>\n</blockquote>\n<hr>\n<p>My best guess at how it works based on the language from above and what I've seen in the data is…</p>\n<ul>\n<li>Based on the 'for sequences shorter …' part above, I assume the first base in the sequence should align to reactivity_0001 and then it would increment forward (with null padding at the end)… this would mean that the first 26 bases of the sequence don't have a reactivity measure, and we would then just map accordingly… </li>\n<li>i.e. if we get non-null values at reactivity_0050 through reactivity_0100 and our sequence has a length of 120 bases. This would mean that the bases from 0-50 are unmeasurable and the bases at 100-120 are unmeasurable and we should have reactivity that is measurable between 50-100.</li>\n</ul>\n<hr>\n<p>Am I on the right track? Thanks in advance!</p>",
  "messages": [
    {
      "id": "2432249",
      "postDate": "09/10/2023 17:32:56",
      "content": "<p>This may be a dumb question but I'll put it out there…</p>\n<p>I would have assumed the <strong><code>sequence</code></strong> column would have the same length as the non-null reactivity values in the reactivity_0*** columns… however, this is not true. The quote from the data page indicates they <strong><em>should</em></strong> have the same length… however, it also explains that some positions may have null values due to difficulty in probing. My experimentation in the training data shows the following columns are always null:</p>\n<ul>\n<li>reactivity 1 through 26 inclusive</li>\n<li>reactivity 166 through 206 inclusive</li>\n</ul>\n<p>I may have messed something up but I didn't think so…</p>\n<blockquote>\n  <p>reactivity_0001, reactivity_0002,… - (float) An array of floating point numbers of the train data, <strong>should have the same length as the RNA sequence</strong>, which defines the reactivity profile for the RNA. For sequences shorter than the maximum RNA length, positions that go beyond the sequence length have null. Several positions early and late in the sequence also cannot be probed due to technical reasons, and their reactivity values are null.</p>\n</blockquote>\n<hr>\n<p>My best guess at how it works based on the language from above and what I've seen in the data is…</p>\n<ul>\n<li>Based on the 'for sequences shorter …' part above, I assume the first base in the sequence should align to reactivity_0001 and then it would increment forward (with null padding at the end)… this would mean that the first 26 bases of the sequence don't have a reactivity measure, and we would then just map accordingly… </li>\n<li>i.e. if we get non-null values at reactivity_0050 through reactivity_0100 and our sequence has a length of 120 bases. This would mean that the bases from 0-50 are unmeasurable and the bases at 100-120 are unmeasurable and we should have reactivity that is measurable between 50-100.</li>\n</ul>\n<hr>\n<p>Am I on the right track? Thanks in advance!</p>",
      "rawMarkdown": "This may be a dumb question but I'll put it out there...\n\nI would have assumed the **`sequence`** column would have the same length as the non-null reactivity values in the reactivity_0*** columns... however, this is not true. The quote from the data page indicates they ***should*** have the same length... however, it also explains that some positions may have null values due to difficulty in probing. My experimentation in the training data shows the following columns are always null:\n* reactivity 1 through 26 inclusive\n* reactivity 166 through 206 inclusive\n\nI may have messed something up but I didn't think so...\n \n> reactivity_0001, reactivity_0002,… - (float) An array of floating point numbers of the train data, **should have the same length as the RNA sequence**, which defines the reactivity profile for the RNA. For sequences shorter than the maximum RNA length, positions that go beyond the sequence length have null. Several positions early and late in the sequence also cannot be probed due to technical reasons, and their reactivity values are null.\n\n--- \n\nMy best guess at how it works based on the language from above and what I've seen in the data is...\n* Based on the 'for sequences shorter ...' part above, I assume the first base in the sequence should align to reactivity_0001 and then it would increment forward (with null padding at the end)... this would mean that the first 26 bases of the sequence don't have a reactivity measure, and we would then just map accordingly... \n* i.e. if we get non-null values at reactivity_0050 through reactivity_0100 and our sequence has a length of 120 bases. This would mean that the bases from 0-50 are unmeasurable and the bases at 100-120 are unmeasurable and we should have reactivity that is measurable between 50-100.\n\n---\n\nAm I on the right track? Thanks in advance!",
      "votes": null
    },
    {
      "id": "2432494",
      "postDate": "09/10/2023 22:20:14",
      "content": "<p>I struggled with this aswell.</p>\n<p>The sum of sequence lengths in test is equal to the length of the submission file. In other words, the first 177 rows in your submission should correspond to the reactivities predicted for each of the nucleotides/bases/sites in the first row of the test file (which has a sequence length of 177).</p>\n<p>Good luck, this is going to be a tricky one!</p>",
      "rawMarkdown": "I struggled with this aswell.\n\nThe sum of sequence lengths in test is equal to the length of the submission file. In other words, the first 177 rows in your submission should correspond to the reactivities predicted for each of the nucleotides/bases/sites in the first row of the test file (which has a sequence length of 177).\n\nGood luck, this is going to be a tricky one!",
      "votes": null
    },
    {
      "id": "2433295",
      "postDate": "09/11/2023 13:26:31",
      "content": "<p>Hi Sean, this makes sense for the test/sample-submission, however, my question above refers to the training data. </p>\n<p>I believe they handle it differently (test v. train) correct?</p>",
      "rawMarkdown": "Hi Sean, this makes sense for the test/sample-submission, however, my question above refers to the training data. \n\nI believe they handle it differently (test v. train) correct?",
      "votes": null
    },
    {
      "id": "2433331",
      "postDate": "09/11/2023 13:47:52",
      "content": "<p>Yes, i'm not an expert in the domain, but my research field was clinical biochemistry and to some extent bioinformatics; anyway, I discussed this w/a friend yesterday who did their PhD in a directly related research field, his input: '…it's weird, but probably a detail to explain it…all the test sequences start w/a GGG, this is probably so that they can measure it'.</p>\n<p>It may also be worth playing w/a penalty function for bases within a certain window of the 5' and 3' end (or vice versa, maybe they're more exposed and more reactive??)..</p>",
      "rawMarkdown": "Yes, i'm not an expert in the domain, but my research field was clinical biochemistry and to some extent bioinformatics; anyway, I discussed this w/a friend yesterday who did their PhD in a directly related research field, his input: '...it's weird, but probably a detail to explain it...all the test sequences start w/a GGG, this is probably so that they can measure it'.\n\nIt may also be worth playing w/a penalty function for bases within a certain window of the 5' and 3' end (or vice versa, maybe they're more exposed and more reactive??)..",
      "votes": null
    },
    {
      "id": "2433580",
      "postDate": "09/11/2023 17:00:26",
      "content": "<p>You are correct, as far as I understand. Data should always start at reactivity_0001 and last for as many columns as the length of the sequence, with null columns in that region being unmeasurable and null columns after that region being padding. Positions at either end of a sequence are typically unmeasurable (those regions play a specific role in the experimental process).</p>",
      "rawMarkdown": "You are correct, as far as I understand. Data should always start at reactivity_0001 and last for as many columns as the length of the sequence, with null columns in that region being unmeasurable and null columns after that region being padding. Positions at either end of a sequence are typically unmeasurable (those regions play a specific role in the experimental process).",
      "votes": null
    },
    {
      "id": "2433776",
      "postDate": "09/11/2023 20:23:41",
      "content": "<p>The first 26 positions in the sequence are the leader on the 5' end that is necessary for the experiment. The first 26 positions in the sequence are identical across the set.</p>\n<p>The last 39-51 positions (it varies) are the barcode hairpin and tail necessary for the experiment. The barcode hairpin is a unique identifier within each batch of experimental testing that was performed.</p>\n<p>Perhaps viewing an <a href=\"https://eternagame.org/labs/11311397\" target=\"_blank\">Eterna puzzle</a> is a helpful illustration.</p>",
      "rawMarkdown": "The first 26 positions in the sequence are the leader on the 5' end that is necessary for the experiment. The first 26 positions in the sequence are identical across the set.\n\nThe last 39-51 positions (it varies) are the barcode hairpin and tail necessary for the experiment. The barcode hairpin is a unique identifier within each batch of experimental testing that was performed.\n\nPerhaps viewing an [Eterna puzzle](https://eternagame.org/labs/11311397) is a helpful illustration.",
      "votes": null
    },
    {
      "id": "2434817",
      "postDate": "09/12/2023 14:23:00",
      "content": "<p>This definitely helps. I'll poke around on the Eterna page and learn what I can before I go deeper.</p>",
      "rawMarkdown": "This definitely helps. I'll poke around on the Eterna page and learn what I can before I go deeper.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2432494,
      "author_name": "seanmacmaghnusa",
      "author_url": "",
      "post_date": "09/10/2023 22:20:14",
      "content": "<p>I struggled with this aswell.</p>\n<p>The sum of sequence lengths in test is equal to the length of the submission file. In other words, the first 177 rows in your submission should correspond to the reactivities predicted for each of the nucleotides/bases/sites in the first row of the test file (which has a sequence length of 177).</p>\n<p>Good luck, this is going to be a tricky one!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2433295,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "09/11/2023 13:26:31",
          "content": "<p>Hi Sean, this makes sense for the test/sample-submission, however, my question above refers to the training data. </p>\n<p>I believe they handle it differently (test v. train) correct?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2433331,
              "author_name": "seanmacmaghnusa",
              "author_url": "",
              "post_date": "09/11/2023 13:47:52",
              "content": "<p>Yes, i'm not an expert in the domain, but my research field was clinical biochemistry and to some extent bioinformatics; anyway, I discussed this w/a friend yesterday who did their PhD in a directly related research field, his input: '…it's weird, but probably a detail to explain it…all the test sequences start w/a GGG, this is probably so that they can measure it'.</p>\n<p>It may also be worth playing w/a penalty function for bases within a certain window of the 5' and 3' end (or vice versa, maybe they're more exposed and more reactive??)..</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2433776,
                  "author_name": "digitalembrace",
                  "author_url": "",
                  "post_date": "09/11/2023 20:23:41",
                  "content": "<p>The first 26 positions in the sequence are the leader on the 5' end that is necessary for the experiment. The first 26 positions in the sequence are identical across the set.</p>\n<p>The last 39-51 positions (it varies) are the barcode hairpin and tail necessary for the experiment. The barcode hairpin is a unique identifier within each batch of experimental testing that was performed.</p>\n<p>Perhaps viewing an <a href=\"https://eternagame.org/labs/11311397\" target=\"_blank\">Eterna puzzle</a> is a helpful illustration.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2434817,
                      "author_name": "dschettler8845",
                      "author_url": "",
                      "post_date": "09/12/2023 14:23:00",
                      "content": "<p>This definitely helps. I'll poke around on the Eterna page and learn what I can before I go deeper.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2433580,
      "author_name": "jonathanromano",
      "author_url": "",
      "post_date": "09/11/2023 17:00:26",
      "content": "<p>You are correct, as far as I understand. Data should always start at reactivity_0001 and last for as many columns as the length of the sequence, with null columns in that region being unmeasurable and null columns after that region being padding. Positions at either end of a sequence are typically unmeasurable (those regions play a specific role in the experimental process).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2432249": "This may be a dumb question but I'll put it out there...\n\nI would have assumed the **`sequence`** column would have the same length as the non-null reactivity values in the reactivity_0*** columns... however, this is not true. The quote from the data page indicates they ***should*** have the same length... however, it also explains that some positions may have null values due to difficulty in probing. My experimentation in the training data shows the following columns are always null:\n* reactivity 1 through 26 inclusive\n* reactivity 166 through 206 inclusive\n\nI may have messed something up but I didn't think so...\n \n> reactivity_0001, reactivity_0002,… - (float) An array of floating point numbers of the train data, **should have the same length as the RNA sequence**, which defines the reactivity profile for the RNA. For sequences shorter than the maximum RNA length, positions that go beyond the sequence length have null. Several positions early and late in the sequence also cannot be probed due to technical reasons, and their reactivity values are null.\n\n--- \n\nMy best guess at how it works based on the language from above and what I've seen in the data is...\n* Based on the 'for sequences shorter ...' part above, I assume the first base in the sequence should align to reactivity_0001 and then it would increment forward (with null padding at the end)... this would mean that the first 26 bases of the sequence don't have a reactivity measure, and we would then just map accordingly... \n* i.e. if we get non-null values at reactivity_0050 through reactivity_0100 and our sequence has a length of 120 bases. This would mean that the bases from 0-50 are unmeasurable and the bases at 100-120 are unmeasurable and we should have reactivity that is measurable between 50-100.\n\n---\n\nAm I on the right track? Thanks in advance!",
    "2432494": "I struggled with this aswell.\n\nThe sum of sequence lengths in test is equal to the length of the submission file. In other words, the first 177 rows in your submission should correspond to the reactivities predicted for each of the nucleotides/bases/sites in the first row of the test file (which has a sequence length of 177).\n\nGood luck, this is going to be a tricky one!",
    "2433295": "Hi Sean, this makes sense for the test/sample-submission, however, my question above refers to the training data. \n\nI believe they handle it differently (test v. train) correct?",
    "2433331": "Yes, i'm not an expert in the domain, but my research field was clinical biochemistry and to some extent bioinformatics; anyway, I discussed this w/a friend yesterday who did their PhD in a directly related research field, his input: '...it's weird, but probably a detail to explain it...all the test sequences start w/a GGG, this is probably so that they can measure it'.\n\nIt may also be worth playing w/a penalty function for bases within a certain window of the 5' and 3' end (or vice versa, maybe they're more exposed and more reactive??)..",
    "2433580": "You are correct, as far as I understand. Data should always start at reactivity_0001 and last for as many columns as the length of the sequence, with null columns in that region being unmeasurable and null columns after that region being padding. Positions at either end of a sequence are typically unmeasurable (those regions play a specific role in the experimental process).",
    "2433776": "The first 26 positions in the sequence are the leader on the 5' end that is necessary for the experiment. The first 26 positions in the sequence are identical across the set.\n\nThe last 39-51 positions (it varies) are the barcode hairpin and tail necessary for the experiment. The barcode hairpin is a unique identifier within each batch of experimental testing that was performed.\n\nPerhaps viewing an [Eterna puzzle](https://eternagame.org/labs/11311397) is a helpful illustration.",
    "2434817": "This definitely helps. I'll poke around on the Eterna page and learn what I can before I go deeper."
  },
  "source": "meta"
}