{
  "id": 573097,
  "title": "[Regarding the Reliability of Training Data]❓ Here is critical case that induce no pair feature. ",
  "url": "/competitions/stanford-rna-3d-folding/discussion/573097",
  "author_name": "Doug",
  "post_date": "2025-04-13T14:15:50.318000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>While reviewing the organizer’s pipeline(I reviesed little bit) and extracting data, I came across an interesting point and decided to write this.</p>\n<p>First of all, I acknowledge that I lack expertise in life science-related fields.</p>\n<p>In RNA PDB files, there is something called a Chain ID. From what I’ve understood so far, this Chain seems to indicate a division based on segments of protein or RNA sequences. I believe the reason for this distinction is that PDB files don’t contain pure RNA only, but rather RNA along with amino acids and other components, so the chains are separated accordingly.</p>\n<p>Please let me know if anything I’ve said above is incorrect.</p>\n<p>At first, I thought RNA was divided by Chain ID because certain segments can exist independently. Of course, there are many cases where an RNA structure only has a single Chain ID. For example, in 1SCL_A, “A” is the Chain ID.</p>\n<p>However, while extracting new data, I discovered a very unusual RNA sequence: 1H1K.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2Fe28381a3eb3acbeda1c3fa56a6a90a77%2F2025-04-13%20230948.png?generation=1744553504211418&amp;alt=media\" alt=\"\"><br>\n(The image was too long to display and was cut off.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2F78785b1c1276ae84c70a4bbc52d09c70%2F2025-04-13%20231202.png?generation=1744553544735253&amp;alt=media\" alt=\"\"><br>\nPDB link: <a href=\"https://www.rcsb.org/structure/1H1K\" target=\"_blank\">https://www.rcsb.org/structure/1H1K</a></p>\n<p>If you look at this RNA, you’ll see that 3 chains contain only A, and another 3 chains contain only U, with equal counts. Therefore, this forms a secondary AU helical structure.</p>\n<p>Not all data are like this, and I’m not sure whether such a special case is included in the provided training data. But this raises concerns about the reliability of a dataset that’s simply divided by Chain.</p>\n<p>If there is any part in the Pipe Line code that accounts for such Chain-related issues, please let me know.<br>\n<a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 3177946,
      "postDate": "2025-04-13T14:15:50.317Z",
      "content": "<p>While reviewing the organizer’s pipeline(I reviesed little bit) and extracting data, I came across an interesting point and decided to write this.</p>\n<p>First of all, I acknowledge that I lack expertise in life science-related fields.</p>\n<p>In RNA PDB files, there is something called a Chain ID. From what I’ve understood so far, this Chain seems to indicate a division based on segments of protein or RNA sequences. I believe the reason for this distinction is that PDB files don’t contain pure RNA only, but rather RNA along with amino acids and other components, so the chains are separated accordingly.</p>\n<p>Please let me know if anything I’ve said above is incorrect.</p>\n<p>At first, I thought RNA was divided by Chain ID because certain segments can exist independently. Of course, there are many cases where an RNA structure only has a single Chain ID. For example, in 1SCL_A, “A” is the Chain ID.</p>\n<p>However, while extracting new data, I discovered a very unusual RNA sequence: 1H1K.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2Fe28381a3eb3acbeda1c3fa56a6a90a77%2F2025-04-13%20230948.png?generation=1744553504211418&amp;alt=media\" alt=\"\"><br>\n(The image was too long to display and was cut off.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2F78785b1c1276ae84c70a4bbc52d09c70%2F2025-04-13%20231202.png?generation=1744553544735253&amp;alt=media\" alt=\"\"><br>\nPDB link: <a href=\"https://www.rcsb.org/structure/1H1K\" target=\"_blank\">https://www.rcsb.org/structure/1H1K</a></p>\n<p>If you look at this RNA, you’ll see that 3 chains contain only A, and another 3 chains contain only U, with equal counts. Therefore, this forms a secondary AU helical structure.</p>\n<p>Not all data are like this, and I’m not sure whether such a special case is included in the provided training data. But this raises concerns about the reliability of a dataset that’s simply divided by Chain.</p>\n<p>If there is any part in the Pipe Line code that accounts for such Chain-related issues, please let me know.<br>\n<a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>Thank you.</p>",
      "rawMarkdown": "While reviewing the organizer’s pipeline(I reviesed little bit) and extracting data, I came across an interesting point and decided to write this.\n\nFirst of all, I acknowledge that I lack expertise in life science-related fields.\n\nIn RNA PDB files, there is something called a Chain ID. From what I’ve understood so far, this Chain seems to indicate a division based on segments of protein or RNA sequences. I believe the reason for this distinction is that PDB files don’t contain pure RNA only, but rather RNA along with amino acids and other components, so the chains are separated accordingly.\n\nPlease let me know if anything I’ve said above is incorrect.\n\nAt first, I thought RNA was divided by Chain ID because certain segments can exist independently. Of course, there are many cases where an RNA structure only has a single Chain ID. For example, in 1SCL_A, “A” is the Chain ID.\n\nHowever, while extracting new data, I discovered a very unusual RNA sequence: 1H1K.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2Fe28381a3eb3acbeda1c3fa56a6a90a77%2F2025-04-13%20230948.png?generation=1744553504211418&alt=media)\n(The image was too long to display and was cut off.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2F78785b1c1276ae84c70a4bbc52d09c70%2F2025-04-13%20231202.png?generation=1744553544735253&alt=media)\nPDB link: https://www.rcsb.org/structure/1H1K\n\nIf you look at this RNA, you’ll see that 3 chains contain only A, and another 3 chains contain only U, with equal counts. Therefore, this forms a secondary AU helical structure.\n\nNot all data are like this, and I’m not sure whether such a special case is included in the provided training data. But this raises concerns about the reliability of a dataset that’s simply divided by Chain.\n\nIf there is any part in the Pipe Line code that accounts for such Chain-related issues, please let me know.\n@shujun717 \n\nThank you.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3177946": "While reviewing the organizer’s pipeline(I reviesed little bit) and extracting data, I came across an interesting point and decided to write this.\n\nFirst of all, I acknowledge that I lack expertise in life science-related fields.\n\nIn RNA PDB files, there is something called a Chain ID. From what I’ve understood so far, this Chain seems to indicate a division based on segments of protein or RNA sequences. I believe the reason for this distinction is that PDB files don’t contain pure RNA only, but rather RNA along with amino acids and other components, so the chains are separated accordingly.\n\nPlease let me know if anything I’ve said above is incorrect.\n\nAt first, I thought RNA was divided by Chain ID because certain segments can exist independently. Of course, there are many cases where an RNA structure only has a single Chain ID. For example, in 1SCL_A, “A” is the Chain ID.\n\nHowever, while extracting new data, I discovered a very unusual RNA sequence: 1H1K.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2Fe28381a3eb3acbeda1c3fa56a6a90a77%2F2025-04-13%20230948.png?generation=1744553504211418&alt=media)\n(The image was too long to display and was cut off.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25267245%2F78785b1c1276ae84c70a4bbc52d09c70%2F2025-04-13%20231202.png?generation=1744553544735253&alt=media)\nPDB link: https://www.rcsb.org/structure/1H1K\n\nIf you look at this RNA, you’ll see that 3 chains contain only A, and another 3 chains contain only U, with equal counts. Therefore, this forms a secondary AU helical structure.\n\nNot all data are like this, and I’m not sure whether such a special case is included in the provided training data. But this raises concerns about the reliability of a dataset that’s simply divided by Chain.\n\nIf there is any part in the Pipe Line code that accounts for such Chain-related issues, please let me know.\n@shujun717 \n\nThank you."
  }
}