{
  "id": 567735,
  "title": "Very Short RNA Sequences",
  "url": "/competitions/stanford-rna-3d-folding/discussion/567735",
  "author_name": "",
  "post_date": "2025-03-11T20:09:01.690131900Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In the training set:</p>\n<p>2OM3_R is just GAA.<br>\n1Y1Y_P is just AGGC.<br>\n5ELS_I is just AAAUAA.<br>\n6HYU_D is just AAAAAA.</p>\n<p>I'm not an RNA guru by a far stretch, but these sequences seem too short to be of valid use.  Can they even exist in nature?  Should we just exclude them from training?</p>",
  "messages": [
    {
      "id": "3147253",
      "postDate": "03/11/2025 20:09:01",
      "content": "<p>In the training set:</p>\n<p>2OM3_R is just GAA.<br>\n1Y1Y_P is just AGGC.<br>\n5ELS_I is just AAAUAA.<br>\n6HYU_D is just AAAAAA.</p>\n<p>I'm not an RNA guru by a far stretch, but these sequences seem too short to be of valid use.  Can they even exist in nature?  Should we just exclude them from training?</p>",
      "rawMarkdown": "In the training set:\n\n2OM3_R is just GAA.\n1Y1Y_P is just AGGC.\n5ELS_I is just AAAUAA.\n6HYU_D is just AAAAAA.\n\nI'm not an RNA guru by a far stretch, but these sequences seem too short to be of valid use.  Can they even exist in nature?  Should we just exclude them from training?",
      "votes": null
    },
    {
      "id": "3147593",
      "postDate": "03/12/2025 07:09:45",
      "content": "<p>If these short sequences didn't exist in PDB structure files, they wouldn't be included in train data. See for an example:</p>\n<p><a href=\"https://www.rcsb.org/structure/6HYU\" target=\"_blank\">https://www.rcsb.org/structure/6HYU</a></p>\n<p>I think the question is whether we know for sure that these types of sequences would not be present in the final test dataset. Because if we remove them from training and there are some sequences of this type in test data, it is a safe bet that they wouldn't be modelled properly.</p>",
      "rawMarkdown": "If these short sequences didn't exist in PDB structure files, they wouldn't be included in train data. See for an example:\n\nhttps://www.rcsb.org/structure/6HYU\n\nI think the question is whether we know for sure that these types of sequences would not be present in the final test dataset. Because if we remove them from training and there are some sequences of this type in test data, it is a safe bet that they wouldn't be modelled properly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3147593,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "03/12/2025 07:09:45",
      "content": "<p>If these short sequences didn't exist in PDB structure files, they wouldn't be included in train data. See for an example:</p>\n<p><a href=\"https://www.rcsb.org/structure/6HYU\" target=\"_blank\">https://www.rcsb.org/structure/6HYU</a></p>\n<p>I think the question is whether we know for sure that these types of sequences would not be present in the final test dataset. Because if we remove them from training and there are some sequences of this type in test data, it is a safe bet that they wouldn't be modelled properly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3147253": "In the training set:\n\n2OM3_R is just GAA.\n1Y1Y_P is just AGGC.\n5ELS_I is just AAAUAA.\n6HYU_D is just AAAAAA.\n\nI'm not an RNA guru by a far stretch, but these sequences seem too short to be of valid use.  Can they even exist in nature?  Should we just exclude them from training?",
    "3147593": "If these short sequences didn't exist in PDB structure files, they wouldn't be included in train data. See for an example:\n\nhttps://www.rcsb.org/structure/6HYU\n\nI think the question is whether we know for sure that these types of sequences would not be present in the final test dataset. Because if we remove them from training and there are some sequences of this type in test data, it is a safe bet that they wouldn't be modelled properly."
  },
  "source": "meta"
}