{
  "id": 451158,
  "title": "Reactivity errors of 2A3 equals to those of DMS",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/451158",
  "author_name": "",
  "post_date": "2023-10-27T11:46:18.953403700Z",
  "votes": 25,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Pretty much the title. They are exactly equals. I show it <a href=\"https://www.kaggle.com/code/shlomoron/srrf-reactivity-errors\" target=\"_blank\">in this notebook</a>.<br>\nNow, I don't know exactly how the errors were calculated, but I'm pretty sure they are not supposed to be the same for both 2A3 and DMS(?)<br>\nI would like to get a clarification from the host, if possible. Thank you.</p>",
  "messages": [
    {
      "id": "2501363",
      "postDate": "10/27/2023 11:46:18",
      "content": "<p>Pretty much the title. They are exactly equals. I show it <a href=\"https://www.kaggle.com/code/shlomoron/srrf-reactivity-errors\" target=\"_blank\">in this notebook</a>.<br>\nNow, I don't know exactly how the errors were calculated, but I'm pretty sure they are not supposed to be the same for both 2A3 and DMS(?)<br>\nI would like to get a clarification from the host, if possible. Thank you.</p>",
      "rawMarkdown": "Pretty much the title. They are exactly equals. I show it [in this notebook](https://www.kaggle.com/code/shlomoron/srrf-reactivity-errors).\nNow, I don't know exactly how the errors were calculated, but I'm pretty sure they are not supposed to be the same for both 2A3 and DMS(?)\nI would like to get a clarification from the host, if possible. Thank you.",
      "votes": null
    },
    {
      "id": "2501510",
      "postDate": "10/27/2023 14:04:40",
      "content": "<p>Thanks for catching! There was indeed a bug in our output pipeline for <code>train_data.csv</code>, which we are addressing now -- stay tuned.</p>",
      "rawMarkdown": "Thanks for catching! There was indeed a bug in our output pipeline for `train_data.csv`, which we are addressing now -- stay tuned.",
      "votes": null
    },
    {
      "id": "2501533",
      "postDate": "10/27/2023 14:16:15",
      "content": "<p>We are ready to retrain a lot of stuff 🥸 Better now than a few days from the the end of the comp! </p>",
      "rawMarkdown": "We are ready to retrain a lot of stuff 🥸 Better now than a few days from the the end of the comp!",
      "votes": null
    },
    {
      "id": "2501563",
      "postDate": "10/27/2023 14:26:27",
      "content": "<p>Thank you. I eagerly wait for the repaired data.</p>",
      "rawMarkdown": "Thank you. I eagerly wait for the repaired data.",
      "votes": null
    },
    {
      "id": "2502292",
      "postDate": "10/28/2023 05:10:16",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> and I have tracked down the bug. Again, thanks to greySnow for bringing to our attention!</p>\n<p>We have replaced <code>train_data.csv</code> in the data. </p>\n<p>The original <code>train_data.csv</code> file is being moved to <code>OLD/train_data.csv</code>. </p>\n<p>As a cross-check, we have confirmed that the <code>signal_to_noise</code> values can be recovered from the <code>reactivity_*</code> and <code>reactivity_error*</code> values, up to small floating point errors. (If you're interested in checking yourself, it may help to look at the actual <code>signal_to_noise</code> computation in <a href=\"https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m\" target=\"_blank\">this function</a> from our data processing pipeline.)</p>\n<p>Please let us know if the values look OK or if you see further issues. </p>\n<p>Hope these fixed error estimates help your modeling!</p>",
      "rawMarkdown": "shujun717 and I have tracked down the bug. Again, thanks to greySnow for bringing to our attention!\n\nWe have replaced `train_data.csv` in the data. \n\nThe original `train_data.csv` file is being moved to `OLD/train_data.csv`. \n\nAs a cross-check, we have confirmed that the `signal_to_noise` values can be recovered from the `reactivity_*` and `reactivity_error*` values, up to small floating point errors. (If you're interested in checking yourself, it may help to look at the actual `signal_to_noise` computation in [this function](https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m) from our data processing pipeline.)\n\nPlease let us know if the values look OK or if you see further issues. \n\nHope these fixed error estimates help your modeling!",
      "votes": null
    },
    {
      "id": "2502337",
      "postDate": "10/28/2023 05:59:13",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "2502421",
      "postDate": "10/28/2023 07:24:54",
      "content": "<p>For some reason, kaggle download now returns \"404 - not found\"</p>",
      "rawMarkdown": "For some reason, kaggle download now returns \"404 - not found\"",
      "votes": null
    },
    {
      "id": "2502884",
      "postDate": "10/28/2023 14:39:00",
      "content": "<p>There may have been a temporary glitch during the data upload process -- but there may also be a permissions error, which will require some help from Kaggle admin. Would you mind posting the link that is giving you  '404 not found'?</p>",
      "rawMarkdown": "There may have been a temporary glitch during the data upload process -- but there may also be a permissions error, which will require some help from Kaggle admin. Would you mind posting the link that is giving you  '404 not found'?",
      "votes": null
    },
    {
      "id": "2504304",
      "postDate": "10/29/2023 18:47:31",
      "content": "<p>Yes, it was a glitch) Thanks</p>",
      "rawMarkdown": "Yes, it was a glitch) Thanks",
      "votes": null
    },
    {
      "id": "2504477",
      "postDate": "10/29/2023 21:56:46",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> I noted that SN_filter changed for a few samples. This may affect also the test set evaluation due to &gt; 100 reads and &gt; 1 signal_to_noise can you double check?</p>\n<p><a href=\"https://www.kaggle.com/code/callmeb/signal-to-noise-changes\" target=\"_blank\">https://www.kaggle.com/code/callmeb/signal-to-noise-changes</a></p>",
      "rawMarkdown": "rhijudas I noted that SN_filter changed for a few samples. This may affect also the test set evaluation due to > 100 reads and > 1 signal_to_noise can you double check?\n\nhttps://www.kaggle.com/code/callmeb/signal-to-noise-changes",
      "votes": null
    },
    {
      "id": "2504488",
      "postDate": "10/29/2023 22:44:49",
      "content": "<p>Yes, that's right -- there was a very minor error in <code>SN_filter</code> computation for a few sequences (&lt;0.1%), which we took the opportunity to correct here in the updated <code>train_data.csv</code>. </p>\n<p>The same fix has been implemented for the private LB, so we should be OK for competition evaluation. </p>\n<p>There's a brief note in the data_description: \"A few <code>SN_filter</code> values have also been updated after a bug fix; no other columns changed.\".</p>",
      "rawMarkdown": "Yes, that's right -- there was a very minor error in `SN_filter` computation for a few sequences (<0.1%), which we took the opportunity to correct here in the updated `train_data.csv`. \n\nThe same fix has been implemented for the private LB, so we should be OK for competition evaluation. \n\nThere's a brief note in the data_description: \"A few `SN_filter` values have also been updated after a bug fix; no other columns changed.\".",
      "votes": null
    },
    {
      "id": "2504571",
      "postDate": "10/30/2023 02:45:39",
      "content": "<p>I don't think so. The last line doesn't make sense. Do you do a mean of the value of comparison? I checked that last line and I got a mean of 9.30207601709638 for 2a3 and a mean of <br>\n3.9844365445499506 for DMS. </p>\n<p>I am still pretty sure that both reactivity's are aggregated shannon entropy sums and the error is the standard error for the aggregate. It is the only thing that makes sense. The 2a3 paper makes more sense of it</p>",
      "rawMarkdown": "I don't think so. The last line doesn't make sense. Do you do a mean of the value of comparison? I checked that last line and I got a mean of 9.30207601709638 for 2a3 and a mean of \n3.9844365445499506 for DMS. \n\nI am still pretty sure that both reactivity's are aggregated shannon entropy sums and the error is the standard error for the aggregate. It is the only thing that makes sense. The 2a3 paper makes more sense of it",
      "votes": null
    },
    {
      "id": "2504640",
      "postDate": "10/30/2023 04:23:39",
      "content": "<p>Read the comments lol. The host already confirmed there was a bug. You get different means now because the data is already repaired, try it with OLD/train_data.csv.</p>",
      "rawMarkdown": "Read the comments lol. The host already confirmed there was a bug. You get different means now because the data is already repaired, try it with OLD/train_data.csv.",
      "votes": null
    },
    {
      "id": "2504644",
      "postDate": "10/30/2023 04:37:20",
      "content": "<p>Wow. That answered my other post too.</p>",
      "rawMarkdown": "Wow. That answered my other post too.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2501510,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "10/27/2023 14:04:40",
      "content": "<p>Thanks for catching! There was indeed a bug in our output pipeline for <code>train_data.csv</code>, which we are addressing now -- stay tuned.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2501533,
          "author_name": "callmeb",
          "author_url": "",
          "post_date": "10/27/2023 14:16:15",
          "content": "<p>We are ready to retrain a lot of stuff 🥸 Better now than a few days from the the end of the comp! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2501563,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "10/27/2023 14:26:27",
          "content": "<p>Thank you. I eagerly wait for the repaired data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2502292,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "10/28/2023 05:10:16",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> and I have tracked down the bug. Again, thanks to greySnow for bringing to our attention!</p>\n<p>We have replaced <code>train_data.csv</code> in the data. </p>\n<p>The original <code>train_data.csv</code> file is being moved to <code>OLD/train_data.csv</code>. </p>\n<p>As a cross-check, we have confirmed that the <code>signal_to_noise</code> values can be recovered from the <code>reactivity_*</code> and <code>reactivity_error*</code> values, up to small floating point errors. (If you're interested in checking yourself, it may help to look at the actual <code>signal_to_noise</code> computation in <a href=\"https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m\" target=\"_blank\">this function</a> from our data processing pipeline.)</p>\n<p>Please let us know if the values look OK or if you see further issues. </p>\n<p>Hope these fixed error estimates help your modeling!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2502337,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "10/28/2023 05:59:13",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2502421,
          "author_name": "dmitrypenzar1996",
          "author_url": "",
          "post_date": "10/28/2023 07:24:54",
          "content": "<p>For some reason, kaggle download now returns \"404 - not found\"</p>",
          "votes": null,
          "replies": [
            {
              "id": 2502884,
              "author_name": "rhijudas",
              "author_url": "",
              "post_date": "10/28/2023 14:39:00",
              "content": "<p>There may have been a temporary glitch during the data upload process -- but there may also be a permissions error, which will require some help from Kaggle admin. Would you mind posting the link that is giving you  '404 not found'?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2504304,
                  "author_name": "dmitrypenzar1996",
                  "author_url": "",
                  "post_date": "10/29/2023 18:47:31",
                  "content": "<p>Yes, it was a glitch) Thanks</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2504477,
      "author_name": "callmeb",
      "author_url": "",
      "post_date": "10/29/2023 21:56:46",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> I noted that SN_filter changed for a few samples. This may affect also the test set evaluation due to &gt; 100 reads and &gt; 1 signal_to_noise can you double check?</p>\n<p><a href=\"https://www.kaggle.com/code/callmeb/signal-to-noise-changes\" target=\"_blank\">https://www.kaggle.com/code/callmeb/signal-to-noise-changes</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2504488,
          "author_name": "rhijudas",
          "author_url": "",
          "post_date": "10/29/2023 22:44:49",
          "content": "<p>Yes, that's right -- there was a very minor error in <code>SN_filter</code> computation for a few sequences (&lt;0.1%), which we took the opportunity to correct here in the updated <code>train_data.csv</code>. </p>\n<p>The same fix has been implemented for the private LB, so we should be OK for competition evaluation. </p>\n<p>There's a brief note in the data_description: \"A few <code>SN_filter</code> values have also been updated after a bug fix; no other columns changed.\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2504571,
      "author_name": "tuttlen",
      "author_url": "",
      "post_date": "10/30/2023 02:45:39",
      "content": "<p>I don't think so. The last line doesn't make sense. Do you do a mean of the value of comparison? I checked that last line and I got a mean of 9.30207601709638 for 2a3 and a mean of <br>\n3.9844365445499506 for DMS. </p>\n<p>I am still pretty sure that both reactivity's are aggregated shannon entropy sums and the error is the standard error for the aggregate. It is the only thing that makes sense. The 2a3 paper makes more sense of it</p>",
      "votes": null,
      "replies": [
        {
          "id": 2504640,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "10/30/2023 04:23:39",
          "content": "<p>Read the comments lol. The host already confirmed there was a bug. You get different means now because the data is already repaired, try it with OLD/train_data.csv.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2504644,
              "author_name": "tuttlen",
              "author_url": "",
              "post_date": "10/30/2023 04:37:20",
              "content": "<p>Wow. That answered my other post too.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2501363": "Pretty much the title. They are exactly equals. I show it [in this notebook](https://www.kaggle.com/code/shlomoron/srrf-reactivity-errors).\nNow, I don't know exactly how the errors were calculated, but I'm pretty sure they are not supposed to be the same for both 2A3 and DMS(?)\nI would like to get a clarification from the host, if possible. Thank you.",
    "2501510": "Thanks for catching! There was indeed a bug in our output pipeline for `train_data.csv`, which we are addressing now -- stay tuned.",
    "2501533": "We are ready to retrain a lot of stuff 🥸 Better now than a few days from the the end of the comp!",
    "2501563": "Thank you. I eagerly wait for the repaired data.",
    "2502292": "shujun717 and I have tracked down the bug. Again, thanks to greySnow for bringing to our attention!\n\nWe have replaced `train_data.csv` in the data. \n\nThe original `train_data.csv` file is being moved to `OLD/train_data.csv`. \n\nAs a cross-check, we have confirmed that the `signal_to_noise` values can be recovered from the `reactivity_*` and `reactivity_error*` values, up to small floating point errors. (If you're interested in checking yourself, it may help to look at the actual `signal_to_noise` computation in [this function](https://github.com/DasLab/ubr/blob/main/matlab/data/ubr_estimate_signal_to_noise_ratio.m) from our data processing pipeline.)\n\nPlease let us know if the values look OK or if you see further issues. \n\nHope these fixed error estimates help your modeling!",
    "2502337": "Thank you!",
    "2502421": "For some reason, kaggle download now returns \"404 - not found\"",
    "2502884": "There may have been a temporary glitch during the data upload process -- but there may also be a permissions error, which will require some help from Kaggle admin. Would you mind posting the link that is giving you  '404 not found'?",
    "2504304": "Yes, it was a glitch) Thanks",
    "2504477": "rhijudas I noted that SN_filter changed for a few samples. This may affect also the test set evaluation due to > 100 reads and > 1 signal_to_noise can you double check?\n\nhttps://www.kaggle.com/code/callmeb/signal-to-noise-changes",
    "2504488": "Yes, that's right -- there was a very minor error in `SN_filter` computation for a few sequences (<0.1%), which we took the opportunity to correct here in the updated `train_data.csv`. \n\nThe same fix has been implemented for the private LB, so we should be OK for competition evaluation. \n\nThere's a brief note in the data_description: \"A few `SN_filter` values have also been updated after a bug fix; no other columns changed.\".",
    "2504571": "I don't think so. The last line doesn't make sense. Do you do a mean of the value of comparison? I checked that last line and I got a mean of 9.30207601709638 for 2a3 and a mean of \n3.9844365445499506 for DMS. \n\nI am still pretty sure that both reactivity's are aggregated shannon entropy sums and the error is the standard error for the aggregate. It is the only thing that makes sense. The 2a3 paper makes more sense of it",
    "2504640": "Read the comments lol. The host already confirmed there was a bug. You get different means now because the data is already repaired, try it with OLD/train_data.csv.",
    "2504644": "Wow. That answered my other post too."
  },
  "source": "meta"
}