{
  "id": 451448,
  "title": "Reactivity value and errors",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/451448",
  "author_name": "",
  "post_date": "2023-10-28T19:33:27.377253500Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Q1. Most reactivity values in data vary from (-3 to 3). What is the significance of values being negative for both profiles (DMS &amp; 2A3)? While making predictions, should the model only give values between (0 to 1) or make predictions in negative value also?</p>",
  "messages": [
    {
      "id": "2503147",
      "postDate": "10/28/2023 19:33:27",
      "content": "<p>Q1. Most reactivity values in data vary from (-3 to 3). What is the significance of values being negative for both profiles (DMS &amp; 2A3)? While making predictions, should the model only give values between (0 to 1) or make predictions in negative value also?</p>",
      "rawMarkdown": "Q1. Most reactivity values in data vary from (-3 to 3). What is the significance of values being negative for both profiles (DMS & 2A3)? While making predictions, should the model only give values between (0 to 1) or make predictions in negative value also?",
      "votes": null
    },
    {
      "id": "2503311",
      "postDate": "10/29/2023 00:14:11",
      "content": "<p>Thanks for the question -- other folks have wondered too. I've added to the data description the explanation: \"The values should be greater than or equal to zero, but due to experimental errors can become negative. The values are normalized so that the 90th percentile value within each dataset is 1.0.\"</p>",
      "rawMarkdown": "Thanks for the question -- other folks have wondered too. I've added to the data description the explanation: \"The values should be greater than or equal to zero, but due to experimental errors can become negative. The values are normalized so that the 90th percentile value within each dataset is 1.0.\"",
      "votes": null
    },
    {
      "id": "2503369",
      "postDate": "10/29/2023 03:09:13",
      "content": "<p>Interesting, that seems to slightly conflict with this comment where <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> says that the values should lie within [0, 1]. Can you clarify? <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/439094#2436956\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/439094#2436956</a></p>",
      "rawMarkdown": "Interesting, that seems to slightly conflict with this comment where @shujun717 says that the values should lie within [0, 1]. Can you clarify? https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/439094#2436956",
      "votes": null
    },
    {
      "id": "2503395",
      "postDate": "10/29/2023 04:01:23",
      "content": "<p>what is the contradiction? </p>\n<pre><code>array = np.linspace(, , )\narray[] = \n()\n</code></pre>",
      "rawMarkdown": "what is the contradiction? \n```Python\narray = np.linspace(0, 1, 10)\narray[8] = 1\nprint(f\"the 90th percentile: {np.percentile(array, 90)}\")\n```",
      "votes": null
    },
    {
      "id": "2503586",
      "postDate": "10/29/2023 08:27:08",
      "content": "<p>Thank you for the answer. One more question: <br>\nIs there any relation between reactivity values and reactivity errors (i.e. value ± errors)? If not, then it would be hard to incorporate the error values in the model training.</p>",
      "rawMarkdown": "Thank you for the answer. One more question: \nIs there any relation between reactivity values and reactivity errors (i.e. value ± errors)? If not, then it would be hard to incorporate the error values in the model training.",
      "votes": null
    },
    {
      "id": "2503591",
      "postDate": "10/29/2023 08:29:34",
      "content": "<p>So if it helps the model perform better, we can set all negative reactivity values to <strong>0</strong> while data pre-processing? </p>",
      "rawMarkdown": "So if it helps the model perform better, we can set all negative reactivity values to **0** while data pre-processing?",
      "votes": null
    },
    {
      "id": "2503971",
      "postDate": "10/29/2023 15:09:09",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> is right -- the clamping between [0,1] occurs with both your models and the data before MAE computation for scoring. </p>\n<p>We've chosen however to give you the raw data values in case that provides any kind of additional useful signal during training.</p>",
      "rawMarkdown": "shujun717 is right -- the clamping between [0,1] occurs with both your models and the data before MAE computation for scoring. \n\nWe've chosen however to give you the raw data values in case that provides any kind of additional useful signal during training.",
      "votes": null
    },
    {
      "id": "2503988",
      "postDate": "10/29/2023 15:14:25",
      "content": "<p>Yes, you are welcome to clamp the data values between 0 and 1 during training.</p>",
      "rawMarkdown": "Yes, you are welcome to clamp the data values between 0 and 1 during training.",
      "votes": null
    },
    {
      "id": "2503991",
      "postDate": "10/29/2023 15:15:14",
      "content": "<p>That makes sense, thank you!</p>",
      "rawMarkdown": "That makes sense, thank you!",
      "votes": null
    },
    {
      "id": "2503996",
      "postDate": "10/29/2023 15:18:37",
      "content": "<p>You are welcome to ignore the reactivity errors. </p>\n<p>We do not know if there is any useful signal in the experimental errors in this competition for auxiliary learning tasks,  for data augmentation, or for definition of your training loss function, though we would be curious to find out! </p>\n<p>You may be interested to look at how the error values were used in the previous Kaggle OpenVaccine competition, e.g., by first place solution <a href=\"https://www.kaggle.com/nullrecurrent\" target=\"_blank\">@nullrecurrent</a>, as described in her <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">summary post</a>. </p>",
      "rawMarkdown": "You are welcome to ignore the reactivity errors. \n\nWe do not know if there is any useful signal in the experimental errors in this competition for auxiliary learning tasks,  for data augmentation, or for definition of your training loss function, though we would be curious to find out! \n\nYou may be interested to look at how the error values were used in the previous Kaggle OpenVaccine competition, e.g., by first place solution @nullrecurrent, as described in her [summary post](https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189620).",
      "votes": null
    },
    {
      "id": "2504599",
      "postDate": "10/30/2023 03:22:55",
      "content": "<p>That maybe true but I was working on the hypothesis that the negative values came from the Shannon entropy aggregation. The values of multiple reads are reduced to a probability function 0.5 being equal probability and 1 being perfect information.  And it reads that values in the top percentiles are taken and the max of the lower quartile or r = 0.001. Is there any credibility to my assumption?</p>",
      "rawMarkdown": "That maybe true but I was working on the hypothesis that the negative values came from the Shannon entropy aggregation. The values of multiple reads are reduced to a probability function 0.5 being equal probability and 1 being perfect information.  And it reads that values in the top percentiles are taken and the max of the lower quartile or r = 0.001. Is there any credibility to my assumption?",
      "votes": null
    },
    {
      "id": "2505296",
      "postDate": "10/30/2023 14:02:16",
      "content": "<p>Good question. </p>\n<p>The <code>reactivity_error_*</code> values are not Shannon entropy.  These values are our best estimate of what the standard deviation of measurements would be if we repeated the measurements (and we have actually checked this expectation in some pilot experiments!).</p>\n<p>For the Ribonanza data, the error values are computed based on:</p>\n<ol>\n<li><p>The error is assumed to be Poisson and is estimated as <code>sqrt(counts)</code>, where <code>counts</code> is an integer that literally counts the number of mutational events seen in sequencers. These events are the signature stored in the sequenced DNA that there was a chemical modification in the RNA during the chemical mapping. The mutational signature is placed in the DNA by the enzymatic machine, called reverse transcriptase, which we use to convert the RNA to DNA for measurements on Illumina DNA sequencers. </p></li>\n<li><p>The errors (along with the raw counts) are divided by the total number of sequences captured by the sequencer ('coverage'). So the reactivity estimates are, to start with, numbers between 0 and 1.</p></li>\n<li><p>Our profiles rely on background subtraction of measurements made on the RNA with the chemical mapping and measurements made without chemical mapping. The errors of the two measurements are summed in quadrature, along with a 'pseudocount' of 1/coverage which ensures that errors do not go below zero.  Note that this background subtraction step can lead to negative values for the reactivity, although chemically such values should not go below zero -- when they do, that's a sign of statistical or systematic error. The values of the estimated error on the reactivity, however, are positive.</p></li>\n<li><p>The reactivity values are normalized to the 90th percentile value seen across all sequence positions across all sequences probed in a given experiment. The same normalization constant used to normalize the reactivity is also applied to scale the reactivity error.</p></li>\n</ol>\n<p>If you want to get into the details, the data processing code and examples are <a href=\"https://github.com/DasLab/ubr\" target=\"_blank\">publicly available</a>, with the error estimation largely happening in the <a href=\"https://github.com/DasLab/ubr/blob/main/matlab/data/get_reactivity.m\" target=\"_blank\">get_reactivity</a> MATLAB function. (Note that there are some additional steps that classify different mutation types --deletions, insertions, mutations-- and correct for some ambiguities that occur when a deletion occurs in a stretch of the same nucleotides.)</p>",
      "rawMarkdown": "Good question. \n\nThe `reactivity_error_*` values are not Shannon entropy.  These values are our best estimate of what the standard deviation of measurements would be if we repeated the measurements (and we have actually checked this expectation in some pilot experiments!).\n\nFor the Ribonanza data, the error values are computed based on:\n\n1.  The error is assumed to be Poisson and is estimated as `sqrt(counts)`, where `counts` is an integer that literally counts the number of mutational events seen in sequencers. These events are the signature stored in the sequenced DNA that there was a chemical modification in the RNA during the chemical mapping. The mutational signature is placed in the DNA by the enzymatic machine, called reverse transcriptase, which we use to convert the RNA to DNA for measurements on Illumina DNA sequencers. \n\n2. The errors (along with the raw counts) are divided by the total number of sequences captured by the sequencer ('coverage'). So the reactivity estimates are, to start with, numbers between 0 and 1.\n\n3. Our profiles rely on background subtraction of measurements made on the RNA with the chemical mapping and measurements made without chemical mapping. The errors of the two measurements are summed in quadrature, along with a 'pseudocount' of 1/coverage which ensures that errors do not go below zero.  Note that this background subtraction step can lead to negative values for the reactivity, although chemically such values should not go below zero -- when they do, that's a sign of statistical or systematic error. The values of the estimated error on the reactivity, however, are positive.\n\n4. The reactivity values are normalized to the 90th percentile value seen across all sequence positions across all sequences probed in a given experiment. The same normalization constant used to normalize the reactivity is also applied to scale the reactivity error.\n\nIf you want to get into the details, the data processing code and examples are [publicly available](https://github.com/DasLab/ubr), with the error estimation largely happening in the [get_reactivity](https://github.com/DasLab/ubr/blob/main/matlab/data/get_reactivity.m) MATLAB function. (Note that there are some additional steps that classify different mutation types --deletions, insertions, mutations-- and correct for some ambiguities that occur when a deletion occurs in a stretch of the same nucleotides.)",
      "votes": null
    },
    {
      "id": "2505951",
      "postDate": "10/31/2023 02:15:02",
      "content": "<p>We were discussing on discord. Turns out <a href=\"https://github.com/DasLab/ubr\" target=\"_blank\">Ubr</a> is the culprit for this data generation. I am still working out how it actually ends up negative.</p>",
      "rawMarkdown": "We were discussing on discord. Turns out [Ubr](https://github.com/DasLab/ubr) is the culprit for this data generation. I am still working out how it actually ends up negative.",
      "votes": null
    },
    {
      "id": "2510436",
      "postDate": "11/03/2023 02:02:45",
      "content": "<p>Great! Nicely explained.</p>",
      "rawMarkdown": "Great! Nicely explained.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2503311,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "10/29/2023 00:14:11",
      "content": "<p>Thanks for the question -- other folks have wondered too. I've added to the data description the explanation: \"The values should be greater than or equal to zero, but due to experimental errors can become negative. The values are normalized so that the 90th percentile value within each dataset is 1.0.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 2503369,
          "author_name": "marktenenholtz",
          "author_url": "",
          "post_date": "10/29/2023 03:09:13",
          "content": "<p>Interesting, that seems to slightly conflict with this comment where <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> says that the values should lie within [0, 1]. Can you clarify? <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/439094#2436956\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/439094#2436956</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2503395,
              "author_name": "sapr3s",
              "author_url": "",
              "post_date": "10/29/2023 04:01:23",
              "content": "<p>what is the contradiction? </p>\n<pre><code>array = np.linspace(, , )\narray[] = \n()\n</code></pre>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2503971,
              "author_name": "rhijudas",
              "author_url": "",
              "post_date": "10/29/2023 15:09:09",
              "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> is right -- the clamping between [0,1] occurs with both your models and the data before MAE computation for scoring. </p>\n<p>We've chosen however to give you the raw data values in case that provides any kind of additional useful signal during training.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2503991,
                  "author_name": "marktenenholtz",
                  "author_url": "",
                  "post_date": "10/29/2023 15:15:14",
                  "content": "<p>That makes sense, thank you!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2503591,
          "author_name": "satishgaurav",
          "author_url": "",
          "post_date": "10/29/2023 08:29:34",
          "content": "<p>So if it helps the model perform better, we can set all negative reactivity values to <strong>0</strong> while data pre-processing? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2503988,
              "author_name": "rhijudas",
              "author_url": "",
              "post_date": "10/29/2023 15:14:25",
              "content": "<p>Yes, you are welcome to clamp the data values between 0 and 1 during training.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2504599,
          "author_name": "tuttlen",
          "author_url": "",
          "post_date": "10/30/2023 03:22:55",
          "content": "<p>That maybe true but I was working on the hypothesis that the negative values came from the Shannon entropy aggregation. The values of multiple reads are reduced to a probability function 0.5 being equal probability and 1 being perfect information.  And it reads that values in the top percentiles are taken and the max of the lower quartile or r = 0.001. Is there any credibility to my assumption?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2505296,
              "author_name": "rhijudas",
              "author_url": "",
              "post_date": "10/30/2023 14:02:16",
              "content": "<p>Good question. </p>\n<p>The <code>reactivity_error_*</code> values are not Shannon entropy.  These values are our best estimate of what the standard deviation of measurements would be if we repeated the measurements (and we have actually checked this expectation in some pilot experiments!).</p>\n<p>For the Ribonanza data, the error values are computed based on:</p>\n<ol>\n<li><p>The error is assumed to be Poisson and is estimated as <code>sqrt(counts)</code>, where <code>counts</code> is an integer that literally counts the number of mutational events seen in sequencers. These events are the signature stored in the sequenced DNA that there was a chemical modification in the RNA during the chemical mapping. The mutational signature is placed in the DNA by the enzymatic machine, called reverse transcriptase, which we use to convert the RNA to DNA for measurements on Illumina DNA sequencers. </p></li>\n<li><p>The errors (along with the raw counts) are divided by the total number of sequences captured by the sequencer ('coverage'). So the reactivity estimates are, to start with, numbers between 0 and 1.</p></li>\n<li><p>Our profiles rely on background subtraction of measurements made on the RNA with the chemical mapping and measurements made without chemical mapping. The errors of the two measurements are summed in quadrature, along with a 'pseudocount' of 1/coverage which ensures that errors do not go below zero.  Note that this background subtraction step can lead to negative values for the reactivity, although chemically such values should not go below zero -- when they do, that's a sign of statistical or systematic error. The values of the estimated error on the reactivity, however, are positive.</p></li>\n<li><p>The reactivity values are normalized to the 90th percentile value seen across all sequence positions across all sequences probed in a given experiment. The same normalization constant used to normalize the reactivity is also applied to scale the reactivity error.</p></li>\n</ol>\n<p>If you want to get into the details, the data processing code and examples are <a href=\"https://github.com/DasLab/ubr\" target=\"_blank\">publicly available</a>, with the error estimation largely happening in the <a href=\"https://github.com/DasLab/ubr/blob/main/matlab/data/get_reactivity.m\" target=\"_blank\">get_reactivity</a> MATLAB function. (Note that there are some additional steps that classify different mutation types --deletions, insertions, mutations-- and correct for some ambiguities that occur when a deletion occurs in a stretch of the same nucleotides.)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2510436,
                  "author_name": "mdrakibtrofder",
                  "author_url": "",
                  "post_date": "11/03/2023 02:02:45",
                  "content": "<p>Great! Nicely explained.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2503586,
      "author_name": "satishgaurav",
      "author_url": "",
      "post_date": "10/29/2023 08:27:08",
      "content": "<p>Thank you for the answer. One more question: <br>\nIs there any relation between reactivity values and reactivity errors (i.e. value ± errors)? If not, then it would be hard to incorporate the error values in the model training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2503996,
          "author_name": "rhijudas",
          "author_url": "",
          "post_date": "10/29/2023 15:18:37",
          "content": "<p>You are welcome to ignore the reactivity errors. </p>\n<p>We do not know if there is any useful signal in the experimental errors in this competition for auxiliary learning tasks,  for data augmentation, or for definition of your training loss function, though we would be curious to find out! </p>\n<p>You may be interested to look at how the error values were used in the previous Kaggle OpenVaccine competition, e.g., by first place solution <a href=\"https://www.kaggle.com/nullrecurrent\" target=\"_blank\">@nullrecurrent</a>, as described in her <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">summary post</a>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2505951,
      "author_name": "tuttlen",
      "author_url": "",
      "post_date": "10/31/2023 02:15:02",
      "content": "<p>We were discussing on discord. Turns out <a href=\"https://github.com/DasLab/ubr\" target=\"_blank\">Ubr</a> is the culprit for this data generation. I am still working out how it actually ends up negative.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2503147": "Q1. Most reactivity values in data vary from (-3 to 3). What is the significance of values being negative for both profiles (DMS & 2A3)? While making predictions, should the model only give values between (0 to 1) or make predictions in negative value also?",
    "2503311": "Thanks for the question -- other folks have wondered too. I've added to the data description the explanation: \"The values should be greater than or equal to zero, but due to experimental errors can become negative. The values are normalized so that the 90th percentile value within each dataset is 1.0.\"",
    "2503369": "Interesting, that seems to slightly conflict with this comment where @shujun717 says that the values should lie within [0, 1]. Can you clarify? https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/439094#2436956",
    "2503395": "what is the contradiction? \n```Python\narray = np.linspace(0, 1, 10)\narray[8] = 1\nprint(f\"the 90th percentile: {np.percentile(array, 90)}\")\n```",
    "2503586": "Thank you for the answer. One more question: \nIs there any relation between reactivity values and reactivity errors (i.e. value ± errors)? If not, then it would be hard to incorporate the error values in the model training.",
    "2503591": "So if it helps the model perform better, we can set all negative reactivity values to **0** while data pre-processing?",
    "2503971": "shujun717 is right -- the clamping between [0,1] occurs with both your models and the data before MAE computation for scoring. \n\nWe've chosen however to give you the raw data values in case that provides any kind of additional useful signal during training.",
    "2503988": "Yes, you are welcome to clamp the data values between 0 and 1 during training.",
    "2503991": "That makes sense, thank you!",
    "2503996": "You are welcome to ignore the reactivity errors. \n\nWe do not know if there is any useful signal in the experimental errors in this competition for auxiliary learning tasks,  for data augmentation, or for definition of your training loss function, though we would be curious to find out! \n\nYou may be interested to look at how the error values were used in the previous Kaggle OpenVaccine competition, e.g., by first place solution @nullrecurrent, as described in her [summary post](https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189620).",
    "2504599": "That maybe true but I was working on the hypothesis that the negative values came from the Shannon entropy aggregation. The values of multiple reads are reduced to a probability function 0.5 being equal probability and 1 being perfect information.  And it reads that values in the top percentiles are taken and the max of the lower quartile or r = 0.001. Is there any credibility to my assumption?",
    "2505296": "Good question. \n\nThe `reactivity_error_*` values are not Shannon entropy.  These values are our best estimate of what the standard deviation of measurements would be if we repeated the measurements (and we have actually checked this expectation in some pilot experiments!).\n\nFor the Ribonanza data, the error values are computed based on:\n\n1.  The error is assumed to be Poisson and is estimated as `sqrt(counts)`, where `counts` is an integer that literally counts the number of mutational events seen in sequencers. These events are the signature stored in the sequenced DNA that there was a chemical modification in the RNA during the chemical mapping. The mutational signature is placed in the DNA by the enzymatic machine, called reverse transcriptase, which we use to convert the RNA to DNA for measurements on Illumina DNA sequencers. \n\n2. The errors (along with the raw counts) are divided by the total number of sequences captured by the sequencer ('coverage'). So the reactivity estimates are, to start with, numbers between 0 and 1.\n\n3. Our profiles rely on background subtraction of measurements made on the RNA with the chemical mapping and measurements made without chemical mapping. The errors of the two measurements are summed in quadrature, along with a 'pseudocount' of 1/coverage which ensures that errors do not go below zero.  Note that this background subtraction step can lead to negative values for the reactivity, although chemically such values should not go below zero -- when they do, that's a sign of statistical or systematic error. The values of the estimated error on the reactivity, however, are positive.\n\n4. The reactivity values are normalized to the 90th percentile value seen across all sequence positions across all sequences probed in a given experiment. The same normalization constant used to normalize the reactivity is also applied to scale the reactivity error.\n\nIf you want to get into the details, the data processing code and examples are [publicly available](https://github.com/DasLab/ubr), with the error estimation largely happening in the [get_reactivity](https://github.com/DasLab/ubr/blob/main/matlab/data/get_reactivity.m) MATLAB function. (Note that there are some additional steps that classify different mutation types --deletions, insertions, mutations-- and correct for some ambiguities that occur when a deletion occurs in a stretch of the same nucleotides.)",
    "2505951": "We were discussing on discord. Turns out [Ubr](https://github.com/DasLab/ubr) is the culprit for this data generation. I am still working out how it actually ends up negative.",
    "2510436": "Great! Nicely explained."
  },
  "source": "meta"
}