{
  "id": 579844,
  "title": "Clarification on Non-Standard Characters (\"X\" and \"-\") in RNA Sequences",
  "url": "/competitions/stanford-rna-3d-folding/discussion/579844",
  "author_name": "",
  "post_date": "2025-05-20T18:55:19.445659300Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>While exploring the sequence data, I noticed that in addition to the standard RNA bases <strong>(\"A\", \"G\", \"C\", and \"U\")</strong>, some sequences also contain the characters <strong>\"X\" and \"-\".</strong></p>\n<p>Could anyone please clarify the following:</p>\n<ul>\n<li>Are \"X\" and \"-\" valid characters in the context of this dataset, or should they be considered as outliers or data errors?</li>\n<li>I noticed that the TRAIN labels also include coordinates for these characters. What does that imply about their role or meaning in the structure?</li>\n</ul>\n<p>Any insights or guidance from the community or organizers would be greatly appreciated!</p>\n<p>Thanks in advance.</p>",
  "messages": [
    {
      "id": "3206040",
      "postDate": "05/20/2025 18:55:19",
      "content": "<p>Hi everyone,</p>\n<p>While exploring the sequence data, I noticed that in addition to the standard RNA bases <strong>(\"A\", \"G\", \"C\", and \"U\")</strong>, some sequences also contain the characters <strong>\"X\" and \"-\".</strong></p>\n<p>Could anyone please clarify the following:</p>\n<ul>\n<li>Are \"X\" and \"-\" valid characters in the context of this dataset, or should they be considered as outliers or data errors?</li>\n<li>I noticed that the TRAIN labels also include coordinates for these characters. What does that imply about their role or meaning in the structure?</li>\n</ul>\n<p>Any insights or guidance from the community or organizers would be greatly appreciated!</p>\n<p>Thanks in advance.</p>",
      "rawMarkdown": "Hi everyone,\n\nWhile exploring the sequence data, I noticed that in addition to the standard RNA bases **(\"A\", \"G\", \"C\", and \"U\")**, some sequences also contain the characters **\"X\" and \"-\".**\n\nCould anyone please clarify the following:\n\n- Are \"X\" and \"-\" valid characters in the context of this dataset, or should they be considered as outliers or data errors?\n- I noticed that the TRAIN labels also include coordinates for these characters. What does that imply about their role or meaning in the structure?\n\nAny insights or guidance from the community or organizers would be greatly appreciated!\n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "3206086",
      "postDate": "05/20/2025 19:40:46",
      "content": "<p>The only valid RNA bases are A, G, C, and U.  I believe an \"X\" or \"-\" indicates that the base at that location couldn't be reliably determined, possibly due to experimental imprecision or such.</p>\n<p>However, it's still an RNA base (even though we don't know what type it is), and so it's totally valid to have its own coordinates just like any other residue in the sequence.</p>",
      "rawMarkdown": "The only valid RNA bases are A, G, C, and U.  I believe an \"X\" or \"-\" indicates that the base at that location couldn't be reliably determined, possibly due to experimental imprecision or such.\n\nHowever, it's still an RNA base (even though we don't know what type it is), and so it's totally valid to have its own coordinates just like any other residue in the sequence.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3206086,
      "author_name": "revealer",
      "author_url": "",
      "post_date": "05/20/2025 19:40:46",
      "content": "<p>The only valid RNA bases are A, G, C, and U.  I believe an \"X\" or \"-\" indicates that the base at that location couldn't be reliably determined, possibly due to experimental imprecision or such.</p>\n<p>However, it's still an RNA base (even though we don't know what type it is), and so it's totally valid to have its own coordinates just like any other residue in the sequence.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3206040": "Hi everyone,\n\nWhile exploring the sequence data, I noticed that in addition to the standard RNA bases **(\"A\", \"G\", \"C\", and \"U\")**, some sequences also contain the characters **\"X\" and \"-\".**\n\nCould anyone please clarify the following:\n\n- Are \"X\" and \"-\" valid characters in the context of this dataset, or should they be considered as outliers or data errors?\n- I noticed that the TRAIN labels also include coordinates for these characters. What does that imply about their role or meaning in the structure?\n\nAny insights or guidance from the community or organizers would be greatly appreciated!\n\nThanks in advance.",
    "3206086": "The only valid RNA bases are A, G, C, and U.  I believe an \"X\" or \"-\" indicates that the base at that location couldn't be reliably determined, possibly due to experimental imprecision or such.\n\nHowever, it's still an RNA base (even though we don't know what type it is), and so it's totally valid to have its own coordinates just like any other residue in the sequence."
  },
  "source": "meta"
}