{
  "id": 568763,
  "title": "Missing segments of structurally unresolved residues in train data",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568763",
  "author_name": "",
  "post_date": "2025-03-17T20:37:34.218859300Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Some of the RNA sequences in the train data have successive residues that are impossibly far apart. See for example 2ZJQ_X_248 and 2ZJQ_X_249 which are &gt;50Å apart. When I check the sequence on <a href=\"https://www.rcsb.org/3d-view/2ZJQ\" target=\"_blank\">PDB</a>, there appears to be more sequence between these, but it's greyed out which I believe means its structure could not be resolved. The sequence should be in the train data but with NA in the coordinates.</p>\n<p>I detected 18 unique chains that have successive residues with distance &gt; 15. Any suggestions on how to deal with this and how to guarantee that all of these unresolved segment cases are dealt with properly and aren't hiding in the data?</p>",
  "messages": [
    {
      "id": "3152469",
      "postDate": "03/17/2025 20:37:34",
      "content": "<p>Some of the RNA sequences in the train data have successive residues that are impossibly far apart. See for example 2ZJQ_X_248 and 2ZJQ_X_249 which are &gt;50Å apart. When I check the sequence on <a href=\"https://www.rcsb.org/3d-view/2ZJQ\" target=\"_blank\">PDB</a>, there appears to be more sequence between these, but it's greyed out which I believe means its structure could not be resolved. The sequence should be in the train data but with NA in the coordinates.</p>\n<p>I detected 18 unique chains that have successive residues with distance &gt; 15. Any suggestions on how to deal with this and how to guarantee that all of these unresolved segment cases are dealt with properly and aren't hiding in the data?</p>",
      "rawMarkdown": "Some of the RNA sequences in the train data have successive residues that are impossibly far apart. See for example 2ZJQ_X_248 and 2ZJQ_X_249 which are >50Å apart. When I check the sequence on [PDB](https://www.rcsb.org/3d-view/2ZJQ), there appears to be more sequence between these, but it's greyed out which I believe means its structure could not be resolved. The sequence should be in the train data but with NA in the coordinates.\n\nI detected 18 unique chains that have successive residues with distance > 15. Any suggestions on how to deal with this and how to guarantee that all of these unresolved segment cases are dealt with properly and aren't hiding in the data?",
      "votes": null
    },
    {
      "id": "3152478",
      "postDate": "03/17/2025 20:53:31",
      "content": "<p>Good catch. Indeed, those residues seem to be missing. Any time consecutive residues are further apart than 5-7 angstroms, something is likely to be wrong.</p>\n<p>I think the organizers need to add the missing residues and fix the numbering. <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>",
      "rawMarkdown": "Good catch. Indeed, those residues seem to be missing. Any time consecutive residues are further apart than 5-7 angstroms, something is likely to be wrong.\n\nI think the organizers need to add the missing residues and fix the numbering. @rhijudas @shujun717",
      "votes": null
    },
    {
      "id": "3154393",
      "postDate": "03/19/2025 22:07:49",
      "content": "<p>Those residues are missing but it's due to limitations of automated processing. For each pdb and its corresponding fasta sequences, i did pairwise alignments to figure out chain/fasta sequence correspondence. However, &gt;1000 sequences were too slow to align so they were kept as they were with missing residues. If you'd like to try to align them, check out the official data processing pipeline here: <a href=\"https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing\" target=\"_blank\">https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing</a></p>",
      "rawMarkdown": "Those residues are missing but it's due to limitations of automated processing. For each pdb and its corresponding fasta sequences, i did pairwise alignments to figure out chain/fasta sequence correspondence. However, >1000 sequences were too slow to align so they were kept as they were with missing residues. If you'd like to try to align them, check out the official data processing pipeline here: https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3152478,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "03/17/2025 20:53:31",
      "content": "<p>Good catch. Indeed, those residues seem to be missing. Any time consecutive residues are further apart than 5-7 angstroms, something is likely to be wrong.</p>\n<p>I think the organizers need to add the missing residues and fix the numbering. <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3154393,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "03/19/2025 22:07:49",
          "content": "<p>Those residues are missing but it's due to limitations of automated processing. For each pdb and its corresponding fasta sequences, i did pairwise alignments to figure out chain/fasta sequence correspondence. However, &gt;1000 sequences were too slow to align so they were kept as they were with missing residues. If you'd like to try to align them, check out the official data processing pipeline here: <a href=\"https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing\" target=\"_blank\">https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3152469": "Some of the RNA sequences in the train data have successive residues that are impossibly far apart. See for example 2ZJQ_X_248 and 2ZJQ_X_249 which are >50Å apart. When I check the sequence on [PDB](https://www.rcsb.org/3d-view/2ZJQ), there appears to be more sequence between these, but it's greyed out which I believe means its structure could not be resolved. The sequence should be in the train data but with NA in the coordinates.\n\nI detected 18 unique chains that have successive residues with distance > 15. Any suggestions on how to deal with this and how to guarantee that all of these unresolved segment cases are dealt with properly and aren't hiding in the data?",
    "3152478": "Good catch. Indeed, those residues seem to be missing. Any time consecutive residues are further apart than 5-7 angstroms, something is likely to be wrong.\n\nI think the organizers need to add the missing residues and fix the numbering. @rhijudas @shujun717",
    "3154393": "Those residues are missing but it's due to limitations of automated processing. For each pdb and its corresponding fasta sequences, i did pairwise alignments to figure out chain/fasta sequence correspondence. However, >1000 sequences were too slow to align so they were kept as they were with missing residues. If you'd like to try to align them, check out the official data processing pipeline here: https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing"
  },
  "source": "meta"
}