{
  "id": 456357,
  "title": "How is \"signal_to_noise\" calculated?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/456357",
  "author_name": "",
  "post_date": "2023-11-19T16:48:13.813242600Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The data column says <code>mean(measurement value over probed positions)/mean( statistical error in measurement value over probed positions)</code> as the calculation method of <code>signal_to_noise</code>. </p>\n<p>However, the following calculation gives a close result, but I can't reproduce it.</p>\n<pre><code>df_2A3 = df.loc[df.experiment_type == ]\nreact_2A3 = df_2A3[\n    [c  c  df_2A3.columns    c]\n].values\nreact_err_2A3 = df_2A3[\n    [c  c  df_2A3.columns    c]\n].values\nreact_mean = np.nanmean(react_2A3, axis=)\nreact_err_mean = np.nanmean(react_err_2A3, axis=)\n(react_mean / react_err_mean)\n\n&gt; [         ...       ]\n\n(df_2A3[].values)\n\n&gt; [      ...    ]\n</code></pre>\n<p>How to reproduce it completely?</p>",
  "messages": [
    {
      "id": "2530922",
      "postDate": "11/19/2023 16:48:13",
      "content": "<p>The data column says <code>mean(measurement value over probed positions)/mean( statistical error in measurement value over probed positions)</code> as the calculation method of <code>signal_to_noise</code>. </p>\n<p>However, the following calculation gives a close result, but I can't reproduce it.</p>\n<pre><code>df_2A3 = df.loc[df.experiment_type == ]\nreact_2A3 = df_2A3[\n    [c  c  df_2A3.columns    c]\n].values\nreact_err_2A3 = df_2A3[\n    [c  c  df_2A3.columns    c]\n].values\nreact_mean = np.nanmean(react_2A3, axis=)\nreact_err_mean = np.nanmean(react_err_2A3, axis=)\n(react_mean / react_err_mean)\n\n&gt; [         ...       ]\n\n(df_2A3[].values)\n\n&gt; [      ...    ]\n</code></pre>\n<p>How to reproduce it completely?</p>",
      "rawMarkdown": "The data column says `mean(measurement value over probed positions)/mean( statistical error in measurement value over probed positions)` as the calculation method of `signal_to_noise`. \n\nHowever, the following calculation gives a close result, but I can't reproduce it.\n```python\ndf_2A3 = df.loc[df.experiment_type == \"2A3_MaP\"]\nreact_2A3 = df_2A3[\n    [c for c in df_2A3.columns if \"reactivity_0\" in c]\n].values\nreact_err_2A3 = df_2A3[\n    [c for c in df_2A3.columns if \"reactivity_error_0\" in c]\n].values\nreact_mean = np.nanmean(react_2A3, axis=1)\nreact_err_mean = np.nanmean(react_err_2A3, axis=1)\nprint(react_mean / react_err_mean)\n\n> [ 0.986319   1.9689928  2.07819   ...  9.472149  15.853026  13.717828 ]\n\nprint(df_2A3[\"signal_to_noise\"].values)\n\n> [ 0.944  1.933  2.347 ...  9.554 15.995 13.799]\n```\n\nHow to reproduce it completely?",
      "votes": null
    },
    {
      "id": "2531132",
      "postDate": "11/19/2023 22:41:15",
      "content": "<p>I guess they are, but have you checked if the NaN positions on reactivity are the same than the ones on errors? Your calculation suggest just a small change.</p>",
      "rawMarkdown": "I guess they are, but have you checked if the NaN positions on reactivity are the same than the ones on errors? Your calculation suggest just a small change.",
      "votes": null
    },
    {
      "id": "2531134",
      "postDate": "11/19/2023 22:49:35",
      "content": "<p>The host explained <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451158\" target=\"_blank\">here</a>. There are several intricacies, like ignoring the first and last no-nan positions and ignoring sequences with less than four valid positions. See the exact function <a href=\"https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "The host explained [here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451158). There are several intricacies, like ignoring the first and last no-nan positions and ignoring sequences with less than four valid positions. See the exact function [here](https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m).",
      "votes": null
    },
    {
      "id": "2531186",
      "postDate": "11/20/2023 02:24:09",
      "content": "<p>Thank you, I didn't take care of that.</p>",
      "rawMarkdown": "Thank you, I didn't take care of that.",
      "votes": null
    },
    {
      "id": "2531187",
      "postDate": "11/20/2023 02:24:50",
      "content": "<p>I missed that post, will check it out.</p>",
      "rawMarkdown": "I missed that post, will check it out.",
      "votes": null
    },
    {
      "id": "2532388",
      "postDate": "11/21/2023 03:00:52",
      "content": "<p>According to the code. a signal is a positive value and a noise is the reactivity error. However these values are computed before normalization which is a value computed in a batch. We are unsure of what and where they sit in the batch. Some of these values may be from a large batch that were not given to us. But I think you got a good approximation</p>",
      "rawMarkdown": "According to the code. a signal is a positive value and a noise is the reactivity error. However these values are computed before normalization which is a value computed in a batch. We are unsure of what and where they sit in the batch. Some of these values may be from a large batch that were not given to us. But I think you got a good approximation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2531132,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "11/19/2023 22:41:15",
      "content": "<p>I guess they are, but have you checked if the NaN positions on reactivity are the same than the ones on errors? Your calculation suggest just a small change.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2531186,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "11/20/2023 02:24:09",
          "content": "<p>Thank you, I didn't take care of that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2531134,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "11/19/2023 22:49:35",
      "content": "<p>The host explained <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451158\" target=\"_blank\">here</a>. There are several intricacies, like ignoring the first and last no-nan positions and ignoring sequences with less than four valid positions. See the exact function <a href=\"https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2531187,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "11/20/2023 02:24:50",
          "content": "<p>I missed that post, will check it out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2532388,
      "author_name": "tuttlen",
      "author_url": "",
      "post_date": "11/21/2023 03:00:52",
      "content": "<p>According to the code. a signal is a positive value and a noise is the reactivity error. However these values are computed before normalization which is a value computed in a batch. We are unsure of what and where they sit in the batch. Some of these values may be from a large batch that were not given to us. But I think you got a good approximation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2530922": "The data column says `mean(measurement value over probed positions)/mean( statistical error in measurement value over probed positions)` as the calculation method of `signal_to_noise`. \n\nHowever, the following calculation gives a close result, but I can't reproduce it.\n```python\ndf_2A3 = df.loc[df.experiment_type == \"2A3_MaP\"]\nreact_2A3 = df_2A3[\n    [c for c in df_2A3.columns if \"reactivity_0\" in c]\n].values\nreact_err_2A3 = df_2A3[\n    [c for c in df_2A3.columns if \"reactivity_error_0\" in c]\n].values\nreact_mean = np.nanmean(react_2A3, axis=1)\nreact_err_mean = np.nanmean(react_err_2A3, axis=1)\nprint(react_mean / react_err_mean)\n\n> [ 0.986319   1.9689928  2.07819   ...  9.472149  15.853026  13.717828 ]\n\nprint(df_2A3[\"signal_to_noise\"].values)\n\n> [ 0.944  1.933  2.347 ...  9.554 15.995 13.799]\n```\n\nHow to reproduce it completely?",
    "2531132": "I guess they are, but have you checked if the NaN positions on reactivity are the same than the ones on errors? Your calculation suggest just a small change.",
    "2531134": "The host explained [here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/451158). There are several intricacies, like ignoring the first and last no-nan positions and ignoring sequences with less than four valid positions. See the exact function [here](https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m).",
    "2531186": "Thank you, I didn't take care of that.",
    "2531187": "I missed that post, will check it out.",
    "2532388": "According to the code. a signal is a positive value and a noise is the reactivity error. However these values are computed before normalization which is a value computed in a batch. We are unsure of what and where they sit in the batch. Some of these values may be from a large batch that were not given to us. But I think you got a good approximation"
  },
  "source": "meta"
}