{
  "id": 574188,
  "title": "Issues Regarding EDA and Feature Extraction/Engineering",
  "url": "/competitions/stanford-rna-3d-folding/discussion/574188",
  "author_name": "AyushGharat17",
  "post_date": "2025-04-20T14:49:07.981000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone 👋,<br>\nI’ve been exploring the dataset and ran into a few issues during EDA and feature extraction that I’d love some input or validation on from the community:</p>\n<blockquote>\n  <h5>Label–Sequence Alignment Issue</h5>\n  <ul>\n  <li>I noticed that some label entries (train_labels.csv) reference nucleotide positions that don’t exist in the associated sequence (train_sequences.csv).</li>\n  <li>For example, for a given ID, the position value is 133, but the sequence length is only 133 — so if indexing starts from 0, the max valid index would be 132.</li>\n  <li>This results in NaN nucleotide values after merging sequence and labels. I’ve identified 844 such entries.<br>\n  👉 Current plan: Drop these rows, but I’m curious to know how others are handling this.</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h5>Missing 3D Coordinates</h5>\n  <ul>\n  <li>In train_coordinates.csv, there are about 6,145 residues with completely missing x, y, z, and atom_type values.<br>\n  Interestingly, most of these are G, C, or U residues — often near the ends of the sequence.<br>\n  👉 Question: Do these correspond to unmodeled parts of the structure, or is this a data issue? How are you treating these in your pipeline?</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h5>Invalid Characters in Sequences</h5>\n  <ul>\n  <li>Some sequences contain characters like X, -, or others that are not part of the standard nucleotide set (A, C, G, U).</li>\n  <li>This causes parsing/tokenization issues when extracting features.<br>\n  👉 Plan: I'm thinking of cleaning these by removing or replacing non-standard characters. Would love to know if someone has a more nuanced approach.</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h5>Structural Label Merge Strategy</h5>\n  <ul>\n  <li>If anyone has a clean and scalable method for merging sequences, coordinates, and labels — especially ensuring only residues with valid coordinate + label + sequence info are used — I’d love to collaborate or learn from your approach!</li>\n  </ul>\n</blockquote>\n<p>Would love to hear how others are tackling these issues 🙌<br>\nHappy to share my notebook too if you’re encountering the same!</p>\n<p>Thanks!<br>\n– Ayush</p>",
  "messages": [
    {
      "id": 3183206,
      "postDate": "2025-04-20T14:49:07.980Z",
      "content": "<p>Hi everyone 👋,<br>\nI’ve been exploring the dataset and ran into a few issues during EDA and feature extraction that I’d love some input or validation on from the community:</p>\n<blockquote>\n  <h5>Label–Sequence Alignment Issue</h5>\n  <ul>\n  <li>I noticed that some label entries (train_labels.csv) reference nucleotide positions that don’t exist in the associated sequence (train_sequences.csv).</li>\n  <li>For example, for a given ID, the position value is 133, but the sequence length is only 133 — so if indexing starts from 0, the max valid index would be 132.</li>\n  <li>This results in NaN nucleotide values after merging sequence and labels. I’ve identified 844 such entries.<br>\n  👉 Current plan: Drop these rows, but I’m curious to know how others are handling this.</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h5>Missing 3D Coordinates</h5>\n  <ul>\n  <li>In train_coordinates.csv, there are about 6,145 residues with completely missing x, y, z, and atom_type values.<br>\n  Interestingly, most of these are G, C, or U residues — often near the ends of the sequence.<br>\n  👉 Question: Do these correspond to unmodeled parts of the structure, or is this a data issue? How are you treating these in your pipeline?</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h5>Invalid Characters in Sequences</h5>\n  <ul>\n  <li>Some sequences contain characters like X, -, or others that are not part of the standard nucleotide set (A, C, G, U).</li>\n  <li>This causes parsing/tokenization issues when extracting features.<br>\n  👉 Plan: I'm thinking of cleaning these by removing or replacing non-standard characters. Would love to know if someone has a more nuanced approach.</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h5>Structural Label Merge Strategy</h5>\n  <ul>\n  <li>If anyone has a clean and scalable method for merging sequences, coordinates, and labels — especially ensuring only residues with valid coordinate + label + sequence info are used — I’d love to collaborate or learn from your approach!</li>\n  </ul>\n</blockquote>\n<p>Would love to hear how others are tackling these issues 🙌<br>\nHappy to share my notebook too if you’re encountering the same!</p>\n<p>Thanks!<br>\n– Ayush</p>",
      "rawMarkdown": "Hi everyone 👋,\nI’ve been exploring the dataset and ran into a few issues during EDA and feature extraction that I’d love some input or validation on from the community:\n\n> ##### Label–Sequence Alignment Issue\n- I noticed that some label entries (train_labels.csv) reference nucleotide positions that don’t exist in the associated sequence (train_sequences.csv).\n- For example, for a given ID, the position value is 133, but the sequence length is only 133 — so if indexing starts from 0, the max valid index would be 132.\n- This results in NaN nucleotide values after merging sequence and labels. I’ve identified 844 such entries.\n👉 Current plan: Drop these rows, but I’m curious to know how others are handling this.\n\n---\n\n> ##### Missing 3D Coordinates\n- In train_coordinates.csv, there are about 6,145 residues with completely missing x, y, z, and atom_type values.\nInterestingly, most of these are G, C, or U residues — often near the ends of the sequence.\n👉 Question: Do these correspond to unmodeled parts of the structure, or is this a data issue? How are you treating these in your pipeline?\n\n---\n\n> ##### Invalid Characters in Sequences\n- Some sequences contain characters like X, -, or others that are not part of the standard nucleotide set (A, C, G, U).\n- This causes parsing/tokenization issues when extracting features.\n👉 Plan: I'm thinking of cleaning these by removing or replacing non-standard characters. Would love to know if someone has a more nuanced approach.\n\n---\n\n> ##### Structural Label Merge Strategy\n- If anyone has a clean and scalable method for merging sequences, coordinates, and labels — especially ensuring only residues with valid coordinate + label + sequence info are used — I’d love to collaborate or learn from your approach!\n\nWould love to hear how others are tackling these issues 🙌\nHappy to share my notebook too if you’re encountering the same!\n\nThanks!\n– Ayush"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3183206": "Hi everyone 👋,\nI’ve been exploring the dataset and ran into a few issues during EDA and feature extraction that I’d love some input or validation on from the community:\n\n> ##### Label–Sequence Alignment Issue\n- I noticed that some label entries (train_labels.csv) reference nucleotide positions that don’t exist in the associated sequence (train_sequences.csv).\n- For example, for a given ID, the position value is 133, but the sequence length is only 133 — so if indexing starts from 0, the max valid index would be 132.\n- This results in NaN nucleotide values after merging sequence and labels. I’ve identified 844 such entries.\n👉 Current plan: Drop these rows, but I’m curious to know how others are handling this.\n\n---\n\n> ##### Missing 3D Coordinates\n- In train_coordinates.csv, there are about 6,145 residues with completely missing x, y, z, and atom_type values.\nInterestingly, most of these are G, C, or U residues — often near the ends of the sequence.\n👉 Question: Do these correspond to unmodeled parts of the structure, or is this a data issue? How are you treating these in your pipeline?\n\n---\n\n> ##### Invalid Characters in Sequences\n- Some sequences contain characters like X, -, or others that are not part of the standard nucleotide set (A, C, G, U).\n- This causes parsing/tokenization issues when extracting features.\n👉 Plan: I'm thinking of cleaning these by removing or replacing non-standard characters. Would love to know if someone has a more nuanced approach.\n\n---\n\n> ##### Structural Label Merge Strategy\n- If anyone has a clean and scalable method for merging sequences, coordinates, and labels — especially ensuring only residues with valid coordinate + label + sequence info are used — I’d love to collaborate or learn from your approach!\n\nWould love to hear how others are tackling these issues 🙌\nHappy to share my notebook too if you’re encountering the same!\n\nThanks!\n– Ayush"
  }
}