{
  "id": 442828,
  "title": "How the start-end positions with UNKNOWN groundtruth are scored on public LB ? (Excluded or fill by zero or ...? )",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/442828",
  "author_name": "",
  "post_date": "2023-09-24T10:57:20.324567200Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First 26 positions have NAN target values in train and many positions at sequence end also - that because current experiments cannot capture them - so NO ground-truth is known. </p>\n<p>However the submission file requires ALL positions to be filled by predictions. </p>\n<p><strong>Question:</strong> I wonder how start and end positions are processed now for public LB scoring - because seems  NO experimental ground-truth for them  ?  Are they just excluded from score computation, or filled by zeros, or by 0.5 or whatever ? Can the detailed scheme be provided please ? </p>\n<p>As organizers explain - it might be possible to have experimental results for these positions in future - so we are asked to submit full predictions. But  public LB is made of currently (not a future data) available data  - so seems to be NO experimental results available. </p>\n<blockquote>\n  <p>It is because our current experimental method cannot capture the reactivity at the beginning and end of each RNA sequence.<br>\n  However, we may be able to make measurements that capture beginning and end of the RNA for some of the test sequences.</p>\n</blockquote>\n<p>( <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/441122#2443959\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/441122#2443959</a> )</p>\n<p>PS</p>\n<p>As kindly shared by Wei Lu in comment below it is suggested that positions outside 26:126 are probably excluded. </p>\n<p>Then, I wonder would it be the same for private ? </p>",
  "messages": [
    {
      "id": "2453826",
      "postDate": "09/24/2023 10:57:20",
      "content": "<p>First 26 positions have NAN target values in train and many positions at sequence end also - that because current experiments cannot capture them - so NO ground-truth is known. </p>\n<p>However the submission file requires ALL positions to be filled by predictions. </p>\n<p><strong>Question:</strong> I wonder how start and end positions are processed now for public LB scoring - because seems  NO experimental ground-truth for them  ?  Are they just excluded from score computation, or filled by zeros, or by 0.5 or whatever ? Can the detailed scheme be provided please ? </p>\n<p>As organizers explain - it might be possible to have experimental results for these positions in future - so we are asked to submit full predictions. But  public LB is made of currently (not a future data) available data  - so seems to be NO experimental results available. </p>\n<blockquote>\n  <p>It is because our current experimental method cannot capture the reactivity at the beginning and end of each RNA sequence.<br>\n  However, we may be able to make measurements that capture beginning and end of the RNA for some of the test sequences.</p>\n</blockquote>\n<p>( <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/441122#2443959\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/441122#2443959</a> )</p>\n<p>PS</p>\n<p>As kindly shared by Wei Lu in comment below it is suggested that positions outside 26:126 are probably excluded. </p>\n<p>Then, I wonder would it be the same for private ? </p>",
      "rawMarkdown": "First 26 positions have NAN target values in train and many positions at sequence end also - that because current experiments cannot capture them - so NO ground-truth is known. \n\nHowever the submission file requires ALL positions to be filled by predictions. \n\n**Question:** I wonder how start and end positions are processed now for public LB scoring - because seems  NO experimental ground-truth for them  ?  Are they just excluded from score computation, or filled by zeros, or by 0.5 or whatever ? Can the detailed scheme be provided please ? \n\n\n As organizers explain - it might be possible to have experimental results for these positions in future - so we are asked to submit full predictions. But  public LB is made of currently (not a future data) available data  - so seems to be NO experimental results available. \n\n>It is because our current experimental method cannot capture the reactivity at the beginning and end of each RNA sequence.\nHowever, we may be able to make measurements that capture beginning and end of the RNA for some of the test sequences.\n\n( https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/441122#2443959 )\n\nPS\n\nAs kindly shared by Wei Lu in comment below it is suggested that positions outside 26:126 are probably excluded. \n\nThen, I wonder would it be the same for private ?",
      "votes": null
    },
    {
      "id": "2453974",
      "postDate": "09/24/2023 13:09:46",
      "content": "<p>I think they are excluded from score computation. I changed all prediction for index 1 to 26 and after 126 from 0 to 1. The final score is the same. </p>",
      "rawMarkdown": "I think they are excluded from score computation. I changed all prediction for index 1 to 26 and after 126 from 0 to 1. The final score is the same.",
      "votes": null
    },
    {
      "id": "2454027",
      "postDate": "09/24/2023 13:59:41",
      "content": "<p>Thanks for sharing ! </p>\n<p>Then, I wonder would it be the same for private - updated the question.</p>",
      "rawMarkdown": "Thanks for sharing ! \n\nThen, I wonder would it be the same for private - updated the question.",
      "votes": null
    },
    {
      "id": "2454599",
      "postDate": "09/24/2023 23:52:29",
      "content": "<p>Thanks for the questions. </p>\n<p>For private LB, the initial and last positions for each sequence may indeed be scored. </p>\n<p>Note in particular that the test sequences include several RNA's with lengths up to 457, so it would not be a good idea to restrict inference to, say 26:126.  </p>\n<p>It's an interesting challenge to get models to make predictions if the training data are for shorter lengths or are all missing the first positions, but we've seen success in this kind of generalization in the past, e.g., with the <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/data\" target=\"_blank\">OpenVaccine kaggle competition</a>. Good luck!</p>",
      "rawMarkdown": "Thanks for the questions. \n\nFor private LB, the initial and last positions for each sequence may indeed be scored. \n\nNote in particular that the test sequences include several RNA's with lengths up to 457, so it would not be a good idea to restrict inference to, say 26:126.  \n\nIt's an interesting challenge to get models to make predictions if the training data are for shorter lengths or are all missing the first positions, but we've seen success in this kind of generalization in the past, e.g., with the [OpenVaccine kaggle competition](https://www.kaggle.com/c/stanford-covid-vaccine/data). Good luck!",
      "votes": null
    },
    {
      "id": "2454939",
      "postDate": "09/25/2023 07:28:31",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/438345#2433776\" target=\"_blank\">this post</a> by <a href=\"https://www.kaggle.com/digitalembrace\" target=\"_blank\">@digitalembrace</a> a competition host may be helpful (if you have not already seen it).</p>",
      "rawMarkdown": "[this post](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/438345#2433776) by @digitalembrace a competition host may be helpful (if you have not already seen it).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2453974,
      "author_name": "riceweilu",
      "author_url": "",
      "post_date": "09/24/2023 13:09:46",
      "content": "<p>I think they are excluded from score computation. I changed all prediction for index 1 to 26 and after 126 from 0 to 1. The final score is the same. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2454027,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/24/2023 13:59:41",
          "content": "<p>Thanks for sharing ! </p>\n<p>Then, I wonder would it be the same for private - updated the question.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2454599,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "09/24/2023 23:52:29",
      "content": "<p>Thanks for the questions. </p>\n<p>For private LB, the initial and last positions for each sequence may indeed be scored. </p>\n<p>Note in particular that the test sequences include several RNA's with lengths up to 457, so it would not be a good idea to restrict inference to, say 26:126.  </p>\n<p>It's an interesting challenge to get models to make predictions if the training data are for shorter lengths or are all missing the first positions, but we've seen success in this kind of generalization in the past, e.g., with the <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/data\" target=\"_blank\">OpenVaccine kaggle competition</a>. Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2454939,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "09/25/2023 07:28:31",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/438345#2433776\" target=\"_blank\">this post</a> by <a href=\"https://www.kaggle.com/digitalembrace\" target=\"_blank\">@digitalembrace</a> a competition host may be helpful (if you have not already seen it).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2453826": "First 26 positions have NAN target values in train and many positions at sequence end also - that because current experiments cannot capture them - so NO ground-truth is known. \n\nHowever the submission file requires ALL positions to be filled by predictions. \n\n**Question:** I wonder how start and end positions are processed now for public LB scoring - because seems  NO experimental ground-truth for them  ?  Are they just excluded from score computation, or filled by zeros, or by 0.5 or whatever ? Can the detailed scheme be provided please ? \n\n\n As organizers explain - it might be possible to have experimental results for these positions in future - so we are asked to submit full predictions. But  public LB is made of currently (not a future data) available data  - so seems to be NO experimental results available. \n\n>It is because our current experimental method cannot capture the reactivity at the beginning and end of each RNA sequence.\nHowever, we may be able to make measurements that capture beginning and end of the RNA for some of the test sequences.\n\n( https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/441122#2443959 )\n\nPS\n\nAs kindly shared by Wei Lu in comment below it is suggested that positions outside 26:126 are probably excluded. \n\nThen, I wonder would it be the same for private ?",
    "2453974": "I think they are excluded from score computation. I changed all prediction for index 1 to 26 and after 126 from 0 to 1. The final score is the same.",
    "2454027": "Thanks for sharing ! \n\nThen, I wonder would it be the same for private - updated the question.",
    "2454599": "Thanks for the questions. \n\nFor private LB, the initial and last positions for each sequence may indeed be scored. \n\nNote in particular that the test sequences include several RNA's with lengths up to 457, so it would not be a good idea to restrict inference to, say 26:126.  \n\nIt's an interesting challenge to get models to make predictions if the training data are for shorter lengths or are all missing the first positions, but we've seen success in this kind of generalization in the past, e.g., with the [OpenVaccine kaggle competition](https://www.kaggle.com/c/stanford-covid-vaccine/data). Good luck!",
    "2454939": "[this post](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/438345#2433776) by @digitalembrace a competition host may be helpful (if you have not already seen it)."
  },
  "source": "meta"
}