{
  "id": 456180,
  "title": "Scoring of flanking regions",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/456180",
  "author_name": "",
  "post_date": "2023-11-18T14:40:58.149590600Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, I might have missed this discussion (sorry), but are the predictions on the flanking regions (including the barcode bits) scored in the final test set? Or are they simply ignored, or even removed before prediction+scoring?</p>",
  "messages": [
    {
      "id": "2529788",
      "postDate": "11/18/2023 14:40:58",
      "content": "<p>Hi, I might have missed this discussion (sorry), but are the predictions on the flanking regions (including the barcode bits) scored in the final test set? Or are they simply ignored, or even removed before prediction+scoring?</p>",
      "rawMarkdown": "Hi, I might have missed this discussion (sorry), but are the predictions on the flanking regions (including the barcode bits) scored in the final test set? Or are they simply ignored, or even removed before prediction+scoring?",
      "votes": null
    },
    {
      "id": "2530662",
      "postDate": "11/19/2023 11:45:38",
      "content": "<p>I'd like to know it too. And I'd like to ask also if the labels of test have known errors. As far I understand the score is calculated without taking into account any label error. So the test have been chosen with small enough errors to be ignored?</p>",
      "rawMarkdown": "I'd like to know it too. And I'd like to ask also if the labels of test have known errors. As far I understand the score is calculated without taking into account any label error. So the test have been chosen with small enough errors to be ignored?",
      "votes": null
    },
    {
      "id": "2531040",
      "postDate": "11/19/2023 19:50:06",
      "content": "<p>I would assume that the flanking regions aren't known in the test set since they aren't known in the training set, and that the same procedure is used to collect these two datasets.</p>",
      "rawMarkdown": "I would assume that the flanking regions aren't known in the test set since they aren't known in the training set, and that the same procedure is used to collect these two datasets.",
      "votes": null
    },
    {
      "id": "2531053",
      "postDate": "11/19/2023 20:04:09",
      "content": "<p>They may be scored and may be ignored. The host referred to it <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">here</a>:</p>\n<p>\"Additionally, you also want to make sure that your model is outputting reasonable predictions for the beginning and end regions where there is no training data. Currently, due to some technical reasons, we cannot measure these positions but we might be able to resolve these limitations in the final private test set. We have noticed internally that some of our models under certain conditions like to output entirely 0s for the beginning and end regions, which should not be happening.\"</p>",
      "rawMarkdown": "They may be scored and may be ignored. The host referred to it [here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653):\n\n\"Additionally, you also want to make sure that your model is outputting reasonable predictions for the beginning and end regions where there is no training data. Currently, due to some technical reasons, we cannot measure these positions but we might be able to resolve these limitations in the final private test set. We have noticed internally that some of our models under certain conditions like to output entirely 0s for the beginning and end regions, which should not be happening.\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2530662,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "11/19/2023 11:45:38",
      "content": "<p>I'd like to know it too. And I'd like to ask also if the labels of test have known errors. As far I understand the score is calculated without taking into account any label error. So the test have been chosen with small enough errors to be ignored?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2531040,
      "author_name": "yvesmartindt",
      "author_url": "",
      "post_date": "11/19/2023 19:50:06",
      "content": "<p>I would assume that the flanking regions aren't known in the test set since they aren't known in the training set, and that the same procedure is used to collect these two datasets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2531053,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "11/19/2023 20:04:09",
      "content": "<p>They may be scored and may be ignored. The host referred to it <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653\" target=\"_blank\">here</a>:</p>\n<p>\"Additionally, you also want to make sure that your model is outputting reasonable predictions for the beginning and end regions where there is no training data. Currently, due to some technical reasons, we cannot measure these positions but we might be able to resolve these limitations in the final private test set. We have noticed internally that some of our models under certain conditions like to output entirely 0s for the beginning and end regions, which should not be happening.\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2529788": "Hi, I might have missed this discussion (sorry), but are the predictions on the flanking regions (including the barcode bits) scored in the final test set? Or are they simply ignored, or even removed before prediction+scoring?",
    "2530662": "I'd like to know it too. And I'd like to ask also if the labels of test have known errors. As far I understand the score is calculated without taking into account any label error. So the test have been chosen with small enough errors to be ignored?",
    "2531040": "I would assume that the flanking regions aren't known in the test set since they aren't known in the training set, and that the same procedure is used to collect these two datasets.",
    "2531053": "They may be scored and may be ignored. The host referred to it [here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/444653):\n\n\"Additionally, you also want to make sure that your model is outputting reasonable predictions for the beginning and end regions where there is no training data. Currently, due to some technical reasons, we cannot measure these positions but we might be able to resolve these limitations in the final private test set. We have noticed internally that some of our models under certain conditions like to output entirely 0s for the beginning and end regions, which should not be happening.\""
  },
  "source": "meta"
}