{
  "id": 571261,
  "title": "Clarification Request: Missing Biological Assembly Data",
  "url": "/competitions/stanford-rna-3d-folding/discussion/571261",
  "author_name": "Bilzard",
  "post_date": "2025-04-02T09:16:20.234000",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Dear organizers,<br>\n<a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a></p>\n<p>It appears that biological assembly information is missing from the competition data.</p>\n<p>For example, target R1107 (PDB ID: 7QR4) is known to form a biological assembly consisting of two copies of chain A and two copies of chain B (four chains in total, as shown below). However, this is not reflected in the provided FASTA data (under the <code>all_sequences</code> column):</p>\n<pre><code>&gt;QR4_1| A| small nuclear ribonucleoprotein A| sapiens ()\nRPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM\n&gt;QR4_2| B| CPEB3 ribozyme| sapiens ()\nGGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\n</code></pre>\n<p>I believe this lack of assembly-level information may impact model inference accuracy.  <br>\nCould you kindly clarify whether this omission was intentional?</p>\n<p>Best regards,  <br>\nBilzard</p>\n<hr>\n<p><strong>EDIT</strong>:</p>\n<p>As a follow-up question, I would also like to ask:</p>\n<p>Would it be possible for the organizers to provide chain and biological assembly information as part of the structural data (e.g., provided as a separate JSON/YAML file), rather than only through the FASTA headers?</p>\n<p>For example, in the following FASTA entry:</p>\n<blockquote>\n  <p>7QR3_1|Chains A, B|U1 small nuclear ribonucleoprotein A|Homo sapiens (9606)  <br>\n  RPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM  <br>\n  7QR3_2|Chains C, D|chimpanzee CPEB3 ribozyme|Pan troglodytes (9598)  <br>\n  GGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU</p>\n</blockquote>\n<p>To avoid ambiguity and improve clarity, a more structured format could be used instead.  <br>\nI suggest the following AF3-like format:</p>\n<pre><code>\n  \n     \n     \n      \n         \n           \n           \n           \n        \n      \n      \n         \n           \n           \n           \n        \n      \n    \n  \n\n</code></pre>\n<p>It is difficult to programmatically parse and separate multiple chains when they are listed together in a single FASTA header (e.g., \"Chains A, B\"), especially since the FASTA format allows free-form annotations.</p>\n<p>Moreover, in many cases, the non-RNA chains—such as protein components—play a critical role in stabilizing or determining the structure of the target complex. Having accurate chain and assembly mappings is therefore essential for reproducible modeling and evaluation.</p>\n<p>We would greatly appreciate it if this information could be made available as part of the official structural templates, if feasible.</p>\n<p>Thank you again for your support.</p>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0c8b83b7a6095c0d409d6da8eb97a68d%2FScreenshot%202025-04-02%20at%2018.08.48.png?generation=1743585350194741&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Faefba64880228958a61c0a5110dfd7e1%2FScreenshot%202025-04-02%20at%2018.08.53.png?generation=1743585365098813&amp;alt=media\" alt=\"\"></p>\n<p>source: <a href=\"https://www.rcsb.org/structure/7QR4\" target=\"_blank\">https://www.rcsb.org/structure/7QR4</a></p>",
  "messages": [
    {
      "id": 3168252,
      "postDate": "2025-04-02T09:16:20.233Z",
      "content": "<p>Dear organizers,<br>\n<a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a></p>\n<p>It appears that biological assembly information is missing from the competition data.</p>\n<p>For example, target R1107 (PDB ID: 7QR4) is known to form a biological assembly consisting of two copies of chain A and two copies of chain B (four chains in total, as shown below). However, this is not reflected in the provided FASTA data (under the <code>all_sequences</code> column):</p>\n<pre><code>&gt;QR4_1| A| small nuclear ribonucleoprotein A| sapiens ()\nRPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM\n&gt;QR4_2| B| CPEB3 ribozyme| sapiens ()\nGGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\n</code></pre>\n<p>I believe this lack of assembly-level information may impact model inference accuracy.  <br>\nCould you kindly clarify whether this omission was intentional?</p>\n<p>Best regards,  <br>\nBilzard</p>\n<hr>\n<p><strong>EDIT</strong>:</p>\n<p>As a follow-up question, I would also like to ask:</p>\n<p>Would it be possible for the organizers to provide chain and biological assembly information as part of the structural data (e.g., provided as a separate JSON/YAML file), rather than only through the FASTA headers?</p>\n<p>For example, in the following FASTA entry:</p>\n<blockquote>\n  <p>7QR3_1|Chains A, B|U1 small nuclear ribonucleoprotein A|Homo sapiens (9606)  <br>\n  RPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM  <br>\n  7QR3_2|Chains C, D|chimpanzee CPEB3 ribozyme|Pan troglodytes (9598)  <br>\n  GGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU</p>\n</blockquote>\n<p>To avoid ambiguity and improve clarity, a more structured format could be used instead.  <br>\nI suggest the following AF3-like format:</p>\n<pre><code>\n  \n     \n     \n      \n         \n           \n           \n           \n        \n      \n      \n         \n           \n           \n           \n        \n      \n    \n  \n\n</code></pre>\n<p>It is difficult to programmatically parse and separate multiple chains when they are listed together in a single FASTA header (e.g., \"Chains A, B\"), especially since the FASTA format allows free-form annotations.</p>\n<p>Moreover, in many cases, the non-RNA chains—such as protein components—play a critical role in stabilizing or determining the structure of the target complex. Having accurate chain and assembly mappings is therefore essential for reproducible modeling and evaluation.</p>\n<p>We would greatly appreciate it if this information could be made available as part of the official structural templates, if feasible.</p>\n<p>Thank you again for your support.</p>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0c8b83b7a6095c0d409d6da8eb97a68d%2FScreenshot%202025-04-02%20at%2018.08.48.png?generation=1743585350194741&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Faefba64880228958a61c0a5110dfd7e1%2FScreenshot%202025-04-02%20at%2018.08.53.png?generation=1743585365098813&amp;alt=media\" alt=\"\"></p>\n<p>source: <a href=\"https://www.rcsb.org/structure/7QR4\" target=\"_blank\">https://www.rcsb.org/structure/7QR4</a></p>",
      "rawMarkdown": "Dear organizers,\n@rhijudas @shujun717\n\nIt appears that biological assembly information is missing from the competition data.\n\nFor example, target R1107 (PDB ID: 7QR4) is known to form a biological assembly consisting of two copies of chain A and two copies of chain B (four chains in total, as shown below). However, this is not reflected in the provided FASTA data (under the `all_sequences` column):\n\n```\n>7QR4_1|Chain A|U1 small nuclear ribonucleoprotein A|Homo sapiens (9606)\nRPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM\n>7QR4_2|Chain B|RNA CPEB3 ribozyme|Homo sapiens (9606)\nGGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\n```\n\nI believe this lack of assembly-level information may impact model inference accuracy.  \nCould you kindly clarify whether this omission was intentional?\n\nBest regards,  \nBilzard\n\n---\n\n**EDIT**:\n\nAs a follow-up question, I would also like to ask:\n\nWould it be possible for the organizers to provide chain and biological assembly information as part of the structural data (e.g., provided as a separate JSON/YAML file), rather than only through the FASTA headers?\n\nFor example, in the following FASTA entry:\n\n>7QR3_1|Chains A, B|U1 small nuclear ribonucleoprotein A|Homo sapiens (9606)  \nRPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM  \n>7QR3_2|Chains C, D|chimpanzee CPEB3 ribozyme|Pan troglodytes (9598)  \nGGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\n\nTo avoid ambiguity and improve clarity, a more structured format could be used instead.  \nI suggest the following AF3-like format:\n\n```json\n[\n  {\n    \"name\": \"7QR3_complex\",\n    \"sequences\": [\n      {\n        \"proteinChain\": {\n          \"sequence\": \"RPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM\",\n          \"count\": 2,\n          \"description\": \"U1 small nuclear ribonucleoprotein A | Homo sapiens\"\n        }\n      },\n      {\n        \"rnaSequence\": {\n          \"sequence\": \"GGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\",\n          \"count\": 2,\n          \"description\": \"CPEB3 ribozyme | Pan troglodytes\"\n        }\n      }\n    ]\n  }\n]\n```\n\nIt is difficult to programmatically parse and separate multiple chains when they are listed together in a single FASTA header (e.g., \"Chains A, B\"), especially since the FASTA format allows free-form annotations.\n\nMoreover, in many cases, the non-RNA chains—such as protein components—play a critical role in stabilizing or determining the structure of the target complex. Having accurate chain and assembly mappings is therefore essential for reproducible modeling and evaluation.\n\nWe would greatly appreciate it if this information could be made available as part of the official structural templates, if feasible.\n\nThank you again for your support.\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0c8b83b7a6095c0d409d6da8eb97a68d%2FScreenshot%202025-04-02%20at%2018.08.48.png?generation=1743585350194741&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Faefba64880228958a61c0a5110dfd7e1%2FScreenshot%202025-04-02%20at%2018.08.53.png?generation=1743585365098813&alt=media)\n\nsource: https://www.rcsb.org/structure/7QR4",
      "votes": 6
    },
    {
      "id": 3170747,
      "postDate": "2025-04-05T01:01:46.743Z",
      "content": "<p>Good catch! </p>\n<p>For the actual LB test sets, we will try to provide the full biological assembly information as curated by experimentalists (including ourselves) in the <code>all_sequences</code> column. </p>\n<p>For training, competitors should consider the publicly viewable competition data sets as potentially incomplete. Some competitiors may well gain an advantage by going into the PDB and extracting the biological assembly information or curating a better training data set! If you see a notable improvement with such curation, please do report in forums and share your data set, so we can all get better!</p>",
      "rawMarkdown": "Good catch! \n\nFor the actual LB test sets, we will try to provide the full biological assembly information as curated by experimentalists (including ourselves) in the `all_sequences` column. \n\nFor training, competitors should consider the publicly viewable competition data sets as potentially incomplete. Some competitiors may well gain an advantage by going into the PDB and extracting the biological assembly information or curating a better training data set! If you see a notable improvement with such curation, please do report in forums and share your data set, so we can all get better!",
      "votes": 1,
      "replies": [
        {
          "id": 3170772,
          "postDate": "2025-04-05T01:51:39.193Z",
          "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> I agree that training examples need not be normalized (though it’s preferred). However, would it be possible to ensure that <strong>the format of validation and test examples is consistent and normalized?</strong></p>\n<p>The FASTA sequence format (in the <code>all_sequence</code> column) is quite flexible and loosely defined—for example, the same sequence may appear as “Chains A,B”, while a single chain might be expressed as “Chain A”. These notations are not strictly standardized, which makes programmatic parsing quite difficult. To avoid potential issues in prediction models, it would be helpful if such conformations were clearly and consistently defined—at least in the validation and test sets.</p>\n<p>Alternatively, can we assume that all test sequences will be monomers with a single molecule?</p>\n<p>c.f.) <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567262#3147215\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567262#3147215</a></p>",
          "rawMarkdown": "@rhijudas I agree that training examples need not be normalized (though it’s preferred). However, would it be possible to ensure that **the format of validation and test examples is consistent and normalized?**\n\nThe FASTA sequence format (in the `all_sequence` column) is quite flexible and loosely defined—for example, the same sequence may appear as “Chains A,B”, while a single chain might be expressed as “Chain A”. These notations are not strictly standardized, which makes programmatic parsing quite difficult. To avoid potential issues in prediction models, it would be helpful if such conformations were clearly and consistently defined—at least in the validation and test sets.\n\nAlternatively, can we assume that all test sequences will be monomers with a single molecule?\n\nc.f.) https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567262#3147215",
          "votes": 2,
          "replies": [
            {
              "id": 3170773,
              "postDate": "2025-04-05T02:02:25.213Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3168438,
      "postDate": "2025-04-02T12:43:02.650Z",
      "content": "<p>there is so much that goes into the formation of 3D structure, pH levels, temperature, Ion concentrations (so critical), biological context - presence of chaperones that assist folding, co-transcriptional folding, different cellular environments, membrane association effect. </p>\n<p>but the experts say, it is possible to infer the 3D structure via it's sequence alone (it's so hard to get my head around this). </p>\n<p>but yeah, any additional info is always relevant to incorporate into the workflow</p>",
      "rawMarkdown": "there is so much that goes into the formation of 3D structure, pH levels, temperature, Ion concentrations (so critical), biological context - presence of chaperones that assist folding, co-transcriptional folding, different cellular environments, membrane association effect. \n\nbut the experts say, it is possible to infer the 3D structure via it's sequence alone (it's so hard to get my head around this). \n\nbut yeah, any additional info is always relevant to incorporate into the workflow",
      "votes": 1,
      "replies": [
        {
          "id": 3168567,
          "postDate": "2025-04-02T15:05:33.263Z",
          "content": "<p>It's true that there are a lot of factors that make the RNAs fold differently. I observed two PDB IDs that containing the same RNA (Hammerhead ribozyme):</p>\n<ul>\n<li>379D: 2 chains of RNA + colbalt (II) ion</li>\n<li>1MME: 4 chains of RNA</li>\n</ul>\n<p>The sequence of RNA of interest: GGCCGAAACUCGUAAGAGUCACCAC (sequence length: 25)</p>\n<ul>\n<li>TM-score (C1' atoms, calculated by USalign command line tool):  0.51368</li>\n<li>TM-score (C1' atoms, calculated with the <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">Kaggle TM-score</a>: 0.43091</li>\n</ul>\n<p>Here how they look like when I aligned them together: 379D_B in blue and 1MME_B in orange<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2F0519174a99f32bb4bddb1a42b4530478%2F379D_B-1MME_B.png?generation=1743605948449347&amp;alt=media\" alt=\"\"></p>\n<p>I think the host allows us to submit 5 structures and they have 40 possible results to evaluate to compensate these factors.</p>",
          "rawMarkdown": "It's true that there are a lot of factors that make the RNAs fold differently. I observed two PDB IDs that containing the same RNA (Hammerhead ribozyme):\n* 379D: 2 chains of RNA + colbalt (II) ion\n* 1MME: 4 chains of RNA\n\nThe sequence of RNA of interest: GGCCGAAACUCGUAAGAGUCACCAC (sequence length: 25)\n* TM-score (C1' atoms, calculated by USalign command line tool):  0.51368\n* TM-score (C1' atoms, calculated with the [Kaggle TM-score](https://www.kaggle.com/code/metric/ribonanza-tm-score): 0.43091\n\nHere how they look like when I aligned them together: 379D_B in blue and 1MME_B in orange\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2F0519174a99f32bb4bddb1a42b4530478%2F379D_B-1MME_B.png?generation=1743605948449347&alt=media)\n\nI think the host allows us to submit 5 structures and they have 40 possible results to evaluate to compensate these factors.",
          "votes": 1,
          "replies": [
            {
              "id": 3168605,
              "postDate": "2025-04-02T15:43:27.003Z",
              "content": "<p>In the validation set, most targets have only a few reference 3D structures (only one target, <code>R1156</code>, has 40 reference structures). Given the complexity and diversity of RNA folding, I'm not sure it's feasible for the host to cover all possible conformations.</p>",
              "rawMarkdown": "In the validation set, most targets have only a few reference 3D structures (only one target, `R1156`, has 40 reference structures). Given the complexity and diversity of RNA folding, I'm not sure it's feasible for the host to cover all possible conformations.",
              "votes": 1
            },
            {
              "id": 3174159,
              "postDate": "2025-04-08T19:09:12.977Z",
              "content": "<p>How are you visualizing this.</p>",
              "rawMarkdown": "How are you visualizing this."
            }
          ]
        }
      ]
    },
    {
      "id": 3168597,
      "postDate": "2025-04-02T15:30:20.463Z",
      "content": "<p><strong>Another concern</strong>:<br>\nGiven that only the RNA component has an available MSA in this competition, while AF3 utilizes MSAs for all components, could this limitation be contributing to reduced prediction performance?</p>",
      "rawMarkdown": "**Another concern**:\nGiven that only the RNA component has an available MSA in this competition, while AF3 utilizes MSAs for all components, could this limitation be contributing to reduced prediction performance?"
    }
  ],
  "comments": [
    {
      "id": 3170747,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2025-04-05T01:01:46.743000",
      "content": "<p>Good catch! </p>\n<p>For the actual LB test sets, we will try to provide the full biological assembly information as curated by experimentalists (including ourselves) in the <code>all_sequences</code> column. </p>\n<p>For training, competitors should consider the publicly viewable competition data sets as potentially incomplete. Some competitiors may well gain an advantage by going into the PDB and extracting the biological assembly information or curating a better training data set! If you see a notable improvement with such curation, please do report in forums and share your data set, so we can all get better!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3170772,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2025-04-05T01:51:39.193000",
          "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> I agree that training examples need not be normalized (though it’s preferred). However, would it be possible to ensure that <strong>the format of validation and test examples is consistent and normalized?</strong></p>\n<p>The FASTA sequence format (in the <code>all_sequence</code> column) is quite flexible and loosely defined—for example, the same sequence may appear as “Chains A,B”, while a single chain might be expressed as “Chain A”. These notations are not strictly standardized, which makes programmatic parsing quite difficult. To avoid potential issues in prediction models, it would be helpful if such conformations were clearly and consistently defined—at least in the validation and test sets.</p>\n<p>Alternatively, can we assume that all test sequences will be monomers with a single molecule?</p>\n<p>c.f.) <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567262#3147215\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/567262#3147215</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 3170773,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-05T02:02:25.213000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3168438,
      "author_name": "g john rao",
      "author_url": "",
      "post_date": "2025-04-02T12:43:02.650000",
      "content": "<p>there is so much that goes into the formation of 3D structure, pH levels, temperature, Ion concentrations (so critical), biological context - presence of chaperones that assist folding, co-transcriptional folding, different cellular environments, membrane association effect. </p>\n<p>but the experts say, it is possible to infer the 3D structure via it's sequence alone (it's so hard to get my head around this). </p>\n<p>but yeah, any additional info is always relevant to incorporate into the workflow</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3168567,
          "author_name": "Pi",
          "author_url": "",
          "post_date": "2025-04-02T15:05:33.263000",
          "content": "<p>It's true that there are a lot of factors that make the RNAs fold differently. I observed two PDB IDs that containing the same RNA (Hammerhead ribozyme):</p>\n<ul>\n<li>379D: 2 chains of RNA + colbalt (II) ion</li>\n<li>1MME: 4 chains of RNA</li>\n</ul>\n<p>The sequence of RNA of interest: GGCCGAAACUCGUAAGAGUCACCAC (sequence length: 25)</p>\n<ul>\n<li>TM-score (C1' atoms, calculated by USalign command line tool):  0.51368</li>\n<li>TM-score (C1' atoms, calculated with the <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">Kaggle TM-score</a>: 0.43091</li>\n</ul>\n<p>Here how they look like when I aligned them together: 379D_B in blue and 1MME_B in orange<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F878726%2F0519174a99f32bb4bddb1a42b4530478%2F379D_B-1MME_B.png?generation=1743605948449347&amp;alt=media\" alt=\"\"></p>\n<p>I think the host allows us to submit 5 structures and they have 40 possible results to evaluate to compensate these factors.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3168605,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2025-04-02T15:43:27.003000",
              "content": "<p>In the validation set, most targets have only a few reference 3D structures (only one target, <code>R1156</code>, has 40 reference structures). Given the complexity and diversity of RNA folding, I'm not sure it's feasible for the host to cover all possible conformations.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3174159,
              "author_name": "Siddhantoon",
              "author_url": "",
              "post_date": "2025-04-08T19:09:12.977000",
              "content": "<p>How are you visualizing this.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3168597,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2025-04-02T15:30:20.463000",
      "content": "<p><strong>Another concern</strong>:<br>\nGiven that only the RNA component has an available MSA in this competition, while AF3 utilizes MSAs for all components, could this limitation be contributing to reduced prediction performance?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3168252": "Dear organizers,\n@rhijudas @shujun717\n\nIt appears that biological assembly information is missing from the competition data.\n\nFor example, target R1107 (PDB ID: 7QR4) is known to form a biological assembly consisting of two copies of chain A and two copies of chain B (four chains in total, as shown below). However, this is not reflected in the provided FASTA data (under the `all_sequences` column):\n\n```\n>7QR4_1|Chain A|U1 small nuclear ribonucleoprotein A|Homo sapiens (9606)\nRPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM\n>7QR4_2|Chain B|RNA CPEB3 ribozyme|Homo sapiens (9606)\nGGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\n```\n\nI believe this lack of assembly-level information may impact model inference accuracy.  \nCould you kindly clarify whether this omission was intentional?\n\nBest regards,  \nBilzard\n\n---\n\n**EDIT**:\n\nAs a follow-up question, I would also like to ask:\n\nWould it be possible for the organizers to provide chain and biological assembly information as part of the structural data (e.g., provided as a separate JSON/YAML file), rather than only through the FASTA headers?\n\nFor example, in the following FASTA entry:\n\n>7QR3_1|Chains A, B|U1 small nuclear ribonucleoprotein A|Homo sapiens (9606)  \nRPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM  \n>7QR3_2|Chains C, D|chimpanzee CPEB3 ribozyme|Pan troglodytes (9598)  \nGGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\n\nTo avoid ambiguity and improve clarity, a more structured format could be used instead.  \nI suggest the following AF3-like format:\n\n```json\n[\n  {\n    \"name\": \"7QR3_complex\",\n    \"sequences\": [\n      {\n        \"proteinChain\": {\n          \"sequence\": \"RPNHTIYINNLNEKIKKDELKKSLHAIFSRFGQILDILVSRSLKMRGQAFVIFKEVSSATNALRSMQGFPFYDKPMRIQYAKTDSDIIAKM\",\n          \"count\": 2,\n          \"description\": \"U1 small nuclear ribonucleoprotein A | Homo sapiens\"\n        }\n      },\n      {\n        \"rnaSequence\": {\n          \"sequence\": \"GGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\",\n          \"count\": 2,\n          \"description\": \"CPEB3 ribozyme | Pan troglodytes\"\n        }\n      }\n    ]\n  }\n]\n```\n\nIt is difficult to programmatically parse and separate multiple chains when they are listed together in a single FASTA header (e.g., \"Chains A, B\"), especially since the FASTA format allows free-form annotations.\n\nMoreover, in many cases, the non-RNA chains—such as protein components—play a critical role in stabilizing or determining the structure of the target complex. Having accurate chain and assembly mappings is therefore essential for reproducible modeling and evaluation.\n\nWe would greatly appreciate it if this information could be made available as part of the official structural templates, if feasible.\n\nThank you again for your support.\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0c8b83b7a6095c0d409d6da8eb97a68d%2FScreenshot%202025-04-02%20at%2018.08.48.png?generation=1743585350194741&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Faefba64880228958a61c0a5110dfd7e1%2FScreenshot%202025-04-02%20at%2018.08.53.png?generation=1743585365098813&alt=media)\n\nsource: https://www.rcsb.org/structure/7QR4",
    "3170747": "Good catch! \n\nFor the actual LB test sets, we will try to provide the full biological assembly information as curated by experimentalists (including ourselves) in the `all_sequences` column. \n\nFor training, competitors should consider the publicly viewable competition data sets as potentially incomplete. Some competitiors may well gain an advantage by going into the PDB and extracting the biological assembly information or curating a better training data set! If you see a notable improvement with such curation, please do report in forums and share your data set, so we can all get better!",
    "3168438": "there is so much that goes into the formation of 3D structure, pH levels, temperature, Ion concentrations (so critical), biological context - presence of chaperones that assist folding, co-transcriptional folding, different cellular environments, membrane association effect. \n\nbut the experts say, it is possible to infer the 3D structure via it's sequence alone (it's so hard to get my head around this). \n\nbut yeah, any additional info is always relevant to incorporate into the workflow",
    "3168597": "**Another concern**:\nGiven that only the RNA component has an available MSA in this competition, while AF3 utilizes MSAs for all components, could this limitation be contributing to reduced prediction performance?"
  }
}