{
  "id": 571324,
  "title": "Question on RNA structure variability and 40-ground-truth consideration",
  "url": "/competitions/stanford-rna-3d-folding/discussion/571324",
  "author_name": "",
  "post_date": "2025-04-02T16:04:20.977057Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> and <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>I’m working on data augmentation for training by processing all RNA-containing structures in the PDB. During my analysis, I noticed significant structural variations between RNA molecules of the same sequence, both within a single PDB entry and across different PDB IDs.</p>\n<p>Note: <em>Below I calculated TM score based on the code provided in <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">Ribonanza TM-score</a></em></p>\n<p><strong>1. Within the same PDB ID</strong> </p>\n<p>For example, in PDB ID: 1MME, which contains four RNA chains, chains B and D correspond to the same RNA sequence (Hammerhead ribozyme):</p>\n<p>Sequence: <code>GGCCGAAACUCGUAAGAGUCACCAC</code> (length: 25)<br>\nBoth chains are fully built with all 25 residues.</p>\n<ul>\n<li>ground truth: 1MME_D (blue)</li>\n<li>prediction: 1MME_B (orange)</li>\n<li>TM-score: 0.60975</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2Fbde1d289e31841651f7e654b660adf8a%2F1MME_D-1MME_B.png?generation=1743609049725642&amp;alt=media\" alt=\"\"></p>\n<p><strong>2. Across different PDB IDs:</strong></p>\n<p>Comparing RNA of the same sequence across different PDB entries also reveals structural variation:</p>\n<ul>\n<li>379D: 2 RNA chains + colbalt (II) ion</li>\n<li>1MME: 4 RNA chains</li>\n<li>TM-score:  0.43091</li>\n</ul>\n<p>379D_B (25 residues) in blue and 1MME_B (25 residues) in orange<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2F47042ff56c635683242cea834a949c3a%2F379D_B-1MME_B.png?generation=1743609059852400&amp;alt=media\" alt=\"\"></p>\n<p><strong>Question</strong>: When defining the 40 ground truths per sequence, do you consider only native RNA structures (as in the first case, where RNA is unbound) or do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc. (as in the second case)?</p>\n<p>Thanks for your clarification!</p>",
  "messages": [
    {
      "id": "3168625",
      "postDate": "04/02/2025 16:04:20",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> and <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>I’m working on data augmentation for training by processing all RNA-containing structures in the PDB. During my analysis, I noticed significant structural variations between RNA molecules of the same sequence, both within a single PDB entry and across different PDB IDs.</p>\n<p>Note: <em>Below I calculated TM score based on the code provided in <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">Ribonanza TM-score</a></em></p>\n<p><strong>1. Within the same PDB ID</strong> </p>\n<p>For example, in PDB ID: 1MME, which contains four RNA chains, chains B and D correspond to the same RNA sequence (Hammerhead ribozyme):</p>\n<p>Sequence: <code>GGCCGAAACUCGUAAGAGUCACCAC</code> (length: 25)<br>\nBoth chains are fully built with all 25 residues.</p>\n<ul>\n<li>ground truth: 1MME_D (blue)</li>\n<li>prediction: 1MME_B (orange)</li>\n<li>TM-score: 0.60975</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2Fbde1d289e31841651f7e654b660adf8a%2F1MME_D-1MME_B.png?generation=1743609049725642&amp;alt=media\" alt=\"\"></p>\n<p><strong>2. Across different PDB IDs:</strong></p>\n<p>Comparing RNA of the same sequence across different PDB entries also reveals structural variation:</p>\n<ul>\n<li>379D: 2 RNA chains + colbalt (II) ion</li>\n<li>1MME: 4 RNA chains</li>\n<li>TM-score:  0.43091</li>\n</ul>\n<p>379D_B (25 residues) in blue and 1MME_B (25 residues) in orange<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2F47042ff56c635683242cea834a949c3a%2F379D_B-1MME_B.png?generation=1743609059852400&amp;alt=media\" alt=\"\"></p>\n<p><strong>Question</strong>: When defining the 40 ground truths per sequence, do you consider only native RNA structures (as in the first case, where RNA is unbound) or do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc. (as in the second case)?</p>\n<p>Thanks for your clarification!</p>",
      "rawMarkdown": "rhijudas and @shujun717 \n\nI’m working on data augmentation for training by processing all RNA-containing structures in the PDB. During my analysis, I noticed significant structural variations between RNA molecules of the same sequence, both within a single PDB entry and across different PDB IDs.\n\nNote: *Below I calculated TM score based on the code provided in [Ribonanza TM-score](https://www.kaggle.com/code/metric/ribonanza-tm-score)*\n\n**1. Within the same PDB ID** \n\nFor example, in PDB ID: 1MME, which contains four RNA chains, chains B and D correspond to the same RNA sequence (Hammerhead ribozyme):\n\nSequence: `GGCCGAAACUCGUAAGAGUCACCAC` (length: 25)\nBoth chains are fully built with all 25 residues.\n\n* ground truth: 1MME_D (blue)\n* prediction: 1MME_B (orange)\n* TM-score: 0.60975\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2Fbde1d289e31841651f7e654b660adf8a%2F1MME_D-1MME_B.png?generation=1743609049725642&alt=media)\n\n**2. Across different PDB IDs:**\n\nComparing RNA of the same sequence across different PDB entries also reveals structural variation:\n- 379D: 2 RNA chains + colbalt (II) ion\n- 1MME: 4 RNA chains\n- TM-score:  0.43091\n\n379D_B (25 residues) in blue and 1MME_B (25 residues) in orange\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2F47042ff56c635683242cea834a949c3a%2F379D_B-1MME_B.png?generation=1743609059852400&alt=media)\n\n**Question**: When defining the 40 ground truths per sequence, do you consider only native RNA structures (as in the first case, where RNA is unbound) or do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc. (as in the second case)?\n\nThanks for your clarification!",
      "votes": null
    },
    {
      "id": "3169162",
      "postDate": "04/03/2025 07:16:28",
      "content": "<p>I also observed the same pattern. In many cases, one single RNA sequence shows up with different PDB IDs. And it's quite confusing which one to keep. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1929445%2F1eaff8e6ae7900ee024fafcb4935949b%2FScreenshot%202025-04-03%20at%2011.12.46.png?generation=1743664407050934&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I also observed the same pattern. In many cases, one single RNA sequence shows up with different PDB IDs. And it's quite confusing which one to keep. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1929445%2F1eaff8e6ae7900ee024fafcb4935949b%2FScreenshot%202025-04-03%20at%2011.12.46.png?generation=1743664407050934&alt=media)",
      "votes": null
    },
    {
      "id": "3169507",
      "postDate": "04/03/2025 14:49:32",
      "content": "<h2>\" do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc.\"</h2>\n<p>From the \"all_sequences\" column we can find the chains of the entire complex. I guess that the test set is extracted from complexes as well.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11578898%2F3107459396fd5af3a59415a25edb43ac%2FCapture.PNG?generation=1743691625652637&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "\" do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc.\"\n---\nFrom the \"all_sequences\" column we can find the chains of the entire complex. I guess that the test set is extracted from complexes as well.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11578898%2F3107459396fd5af3a59415a25edb43ac%2FCapture.PNG?generation=1743691625652637&alt=media)",
      "votes": null
    },
    {
      "id": "3169539",
      "postDate": "04/03/2025 15:19:30",
      "content": "<p>In this particular case you can just choose the 1st one as it has better resolution.</p>",
      "rawMarkdown": "In this particular case you can just choose the 1st one as it has better resolution.",
      "votes": null
    },
    {
      "id": "3170003",
      "postDate": "04/04/2025 07:18:07",
      "content": "<p>Yes, resolution is a good point. But I think the structures corresponding to this sequence are not the same except for the measurement error (resolution). They could be in different shapes depending on which molecules they are interacting with. So they could all be ground truths. </p>",
      "rawMarkdown": "Yes, resolution is a good point. But I think the structures corresponding to this sequence are not the same except for the measurement error (resolution). They could be in different shapes depending on which molecules they are interacting with. So they could all be ground truths.",
      "votes": null
    },
    {
      "id": "3170078",
      "postDate": "04/04/2025 09:19:01",
      "content": "<p>yes, 40 ground truths for each sequence should consider these factors. Even two RNA chains with same sequence in the same complex have TM-score around 0.609:</p>\n<ul>\n<li>ground truth: 1MME_D</li>\n<li>prediction: 1MME_B</li>\n<li>TM-score: 0.60975</li>\n</ul>\n<p>With Cobalt (II) ion: TM-score 0.43091.</p>",
      "rawMarkdown": "yes, 40 ground truths for each sequence should consider these factors. Even two RNA chains with same sequence in the same complex have TM-score around 0.609:\n- ground truth: 1MME_D\n- prediction: 1MME_B\n- TM-score: 0.60975\n\nWith Cobalt (II) ion: TM-score 0.43091.",
      "votes": null
    },
    {
      "id": "3170230",
      "postDate": "04/04/2025 13:09:58",
      "content": "<p>Thanks for your suggestion! It's true, there are:</p>\n<ul>\n<li>197 target_ids containing ligands such as ZN, PO4, PAR (Paromomycin), etc.</li>\n<li>479 containing protein</li>\n<li>66 containing DNA</li>\n</ul>",
      "rawMarkdown": "Thanks for your suggestion! It's true, there are:\n* 197 target_ids containing ligands such as ZN, PO4, PAR (Paromomycin), etc.\n* 479 containing protein\n* 66 containing DNA",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3169162,
      "author_name": "zoushuxian",
      "author_url": "",
      "post_date": "04/03/2025 07:16:28",
      "content": "<p>I also observed the same pattern. In many cases, one single RNA sequence shows up with different PDB IDs. And it's quite confusing which one to keep. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1929445%2F1eaff8e6ae7900ee024fafcb4935949b%2FScreenshot%202025-04-03%20at%2011.12.46.png?generation=1743664407050934&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3169539,
          "author_name": "ogurtsov",
          "author_url": "",
          "post_date": "04/03/2025 15:19:30",
          "content": "<p>In this particular case you can just choose the 1st one as it has better resolution.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3170003,
              "author_name": "zoushuxian",
              "author_url": "",
              "post_date": "04/04/2025 07:18:07",
              "content": "<p>Yes, resolution is a good point. But I think the structures corresponding to this sequence are not the same except for the measurement error (resolution). They could be in different shapes depending on which molecules they are interacting with. So they could all be ground truths. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3170078,
                  "author_name": "nguyenhoa",
                  "author_url": "",
                  "post_date": "04/04/2025 09:19:01",
                  "content": "<p>yes, 40 ground truths for each sequence should consider these factors. Even two RNA chains with same sequence in the same complex have TM-score around 0.609:</p>\n<ul>\n<li>ground truth: 1MME_D</li>\n<li>prediction: 1MME_B</li>\n<li>TM-score: 0.60975</li>\n</ul>\n<p>With Cobalt (II) ion: TM-score 0.43091.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3169507,
      "author_name": "deepauxilary",
      "author_url": "",
      "post_date": "04/03/2025 14:49:32",
      "content": "<h2>\" do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc.\"</h2>\n<p>From the \"all_sequences\" column we can find the chains of the entire complex. I guess that the test set is extracted from complexes as well.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11578898%2F3107459396fd5af3a59415a25edb43ac%2FCapture.PNG?generation=1743691625652637&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3170230,
          "author_name": "nguyenhoa",
          "author_url": "",
          "post_date": "04/04/2025 13:09:58",
          "content": "<p>Thanks for your suggestion! It's true, there are:</p>\n<ul>\n<li>197 target_ids containing ligands such as ZN, PO4, PAR (Paromomycin), etc.</li>\n<li>479 containing protein</li>\n<li>66 containing DNA</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3168625": "rhijudas and @shujun717 \n\nI’m working on data augmentation for training by processing all RNA-containing structures in the PDB. During my analysis, I noticed significant structural variations between RNA molecules of the same sequence, both within a single PDB entry and across different PDB IDs.\n\nNote: *Below I calculated TM score based on the code provided in [Ribonanza TM-score](https://www.kaggle.com/code/metric/ribonanza-tm-score)*\n\n**1. Within the same PDB ID** \n\nFor example, in PDB ID: 1MME, which contains four RNA chains, chains B and D correspond to the same RNA sequence (Hammerhead ribozyme):\n\nSequence: `GGCCGAAACUCGUAAGAGUCACCAC` (length: 25)\nBoth chains are fully built with all 25 residues.\n\n* ground truth: 1MME_D (blue)\n* prediction: 1MME_B (orange)\n* TM-score: 0.60975\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2Fbde1d289e31841651f7e654b660adf8a%2F1MME_D-1MME_B.png?generation=1743609049725642&alt=media)\n\n**2. Across different PDB IDs:**\n\nComparing RNA of the same sequence across different PDB entries also reveals structural variation:\n- 379D: 2 RNA chains + colbalt (II) ion\n- 1MME: 4 RNA chains\n- TM-score:  0.43091\n\n379D_B (25 residues) in blue and 1MME_B (25 residues) in orange\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2F47042ff56c635683242cea834a949c3a%2F379D_B-1MME_B.png?generation=1743609059852400&alt=media)\n\n**Question**: When defining the 40 ground truths per sequence, do you consider only native RNA structures (as in the first case, where RNA is unbound) or do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc. (as in the second case)?\n\nThanks for your clarification!",
    "3169162": "I also observed the same pattern. In many cases, one single RNA sequence shows up with different PDB IDs. And it's quite confusing which one to keep. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1929445%2F1eaff8e6ae7900ee024fafcb4935949b%2FScreenshot%202025-04-03%20at%2011.12.46.png?generation=1743664407050934&alt=media)",
    "3169507": "\" do you also include RNA structures in complexes with proteins, DNA, other RNA, ligands, ions, etc.\"\n---\nFrom the \"all_sequences\" column we can find the chains of the entire complex. I guess that the test set is extracted from complexes as well.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11578898%2F3107459396fd5af3a59415a25edb43ac%2FCapture.PNG?generation=1743691625652637&alt=media)",
    "3169539": "In this particular case you can just choose the 1st one as it has better resolution.",
    "3170003": "Yes, resolution is a good point. But I think the structures corresponding to this sequence are not the same except for the measurement error (resolution). They could be in different shapes depending on which molecules they are interacting with. So they could all be ground truths.",
    "3170078": "yes, 40 ground truths for each sequence should consider these factors. Even two RNA chains with same sequence in the same complex have TM-score around 0.609:\n- ground truth: 1MME_D\n- prediction: 1MME_B\n- TM-score: 0.60975\n\nWith Cobalt (II) ion: TM-score 0.43091.",
    "3170230": "Thanks for your suggestion! It's true, there are:\n* 197 target_ids containing ligands such as ZN, PO4, PAR (Paromomycin), etc.\n* 479 containing protein\n* 66 containing DNA"
  },
  "source": "meta"
}