{
  "id": 576766,
  "title": "Why do you consider the same sequence as different target ids?",
  "url": "/competitions/stanford-rna-3d-folding/discussion/576766",
  "author_name": "",
  "post_date": "2025-05-07T00:45:12.100484500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Dear <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> das and <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a>,</p>\n<p>Thank you again for organizing this exciting challenge. I’d like to revisit <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324\" target=\"_blank\">a question I asked earlier</a> regarding validation, which I believe is still unclear and important for the fairness and consistency of the evaluation.</p>\n<p>In the <code>validation_sequences.csv</code> and <code>validation_labels.csv</code> files, I noticed that 2 target IDs, R1189 and R1190, have an identical RNA sequence but are treated as separate targets:</p>\n<ul>\n<li><p>R1189 (7YR7_1) is associated with 6 chains including CsrA protein,</p></li>\n<li><p>R1190 (7YR6_1) is associated with 4 chains of CsrA.</p></li>\n</ul>\n<p>My questions:</p>\n<ul>\n<li><p>Why were these considered distinct target IDs, rather than treating them as the same target with multiple possible 3D conformations (i.e., multiple valid ground truths for the same sequence)?</p></li>\n<li><p>If the rationale is based on the accompanying <code>all_sequences</code> (which lists macromolecule chains), how can we distinguish cases where RNA is co-crystallized with small molecules, ions, or ligands, since these are not included in <code>all_sequences</code>?</p></li>\n</ul>\n<p>As previously raised in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324\" target=\"_blank\">this discussion</a> , small molecules and ions can significantly influence RNA conformation, yet we currently lack the input features needed to account for them.</p>\n<p><strong>Suggestion</strong>:<br>\nWould it be possible to merge all structures with the same RNA sequence into a single target ID, and consider all of their 3D conformations as valid ground truths?<br>\nThis could reduce ambiguity during testing and better reflect the biological reality where a single RNA sequence may adopt multiple conformations depending on context (e.g., protein, ion binding).</p>\n<p>Thank you for considering this. Looking forward to your thoughts!</p>",
  "messages": [
    {
      "id": "3195371",
      "postDate": "05/07/2025 00:45:12",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> das and <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a>,</p>\n<p>Thank you again for organizing this exciting challenge. I’d like to revisit <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324\" target=\"_blank\">a question I asked earlier</a> regarding validation, which I believe is still unclear and important for the fairness and consistency of the evaluation.</p>\n<p>In the <code>validation_sequences.csv</code> and <code>validation_labels.csv</code> files, I noticed that 2 target IDs, R1189 and R1190, have an identical RNA sequence but are treated as separate targets:</p>\n<ul>\n<li><p>R1189 (7YR7_1) is associated with 6 chains including CsrA protein,</p></li>\n<li><p>R1190 (7YR6_1) is associated with 4 chains of CsrA.</p></li>\n</ul>\n<p>My questions:</p>\n<ul>\n<li><p>Why were these considered distinct target IDs, rather than treating them as the same target with multiple possible 3D conformations (i.e., multiple valid ground truths for the same sequence)?</p></li>\n<li><p>If the rationale is based on the accompanying <code>all_sequences</code> (which lists macromolecule chains), how can we distinguish cases where RNA is co-crystallized with small molecules, ions, or ligands, since these are not included in <code>all_sequences</code>?</p></li>\n</ul>\n<p>As previously raised in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324\" target=\"_blank\">this discussion</a> , small molecules and ions can significantly influence RNA conformation, yet we currently lack the input features needed to account for them.</p>\n<p><strong>Suggestion</strong>:<br>\nWould it be possible to merge all structures with the same RNA sequence into a single target ID, and consider all of their 3D conformations as valid ground truths?<br>\nThis could reduce ambiguity during testing and better reflect the biological reality where a single RNA sequence may adopt multiple conformations depending on context (e.g., protein, ion binding).</p>\n<p>Thank you for considering this. Looking forward to your thoughts!</p>",
      "rawMarkdown": "Dear @rhijudas das and @shujun717,\n\nThank you again for organizing this exciting challenge. I’d like to revisit [a question I asked earlier](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324) regarding validation, which I believe is still unclear and important for the fairness and consistency of the evaluation.\n\nIn the `validation_sequences.csv` and `validation_labels.csv` files, I noticed that 2 target IDs, R1189 and R1190, have an identical RNA sequence but are treated as separate targets:\n\n* R1189 (7YR7_1) is associated with 6 chains including CsrA protein,\n\n* R1190 (7YR6_1) is associated with 4 chains of CsrA.\n\nMy questions:\n* Why were these considered distinct target IDs, rather than treating them as the same target with multiple possible 3D conformations (i.e., multiple valid ground truths for the same sequence)?\n\n* If the rationale is based on the accompanying `all_sequences` (which lists macromolecule chains), how can we distinguish cases where RNA is co-crystallized with small molecules, ions, or ligands, since these are not included in `all_sequences`?\n\nAs previously raised in [this discussion](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324) , small molecules and ions can significantly influence RNA conformation, yet we currently lack the input features needed to account for them.\n\n**Suggestion**:\nWould it be possible to merge all structures with the same RNA sequence into a single target ID, and consider all of their 3D conformations as valid ground truths?\nThis could reduce ambiguity during testing and better reflect the biological reality where a single RNA sequence may adopt multiple conformations depending on context (e.g., protein, ion binding).\n\nThank you for considering this. Looking forward to your thoughts!",
      "votes": null
    },
    {
      "id": "3196838",
      "postDate": "05/07/2025 13:39:33",
      "content": "<p>Thanks for the post! </p>\n<p>We included these duplicate sequences TR1189 and TR1190 in the host provided validation because they were indeed separate targets in CASP15 and have become a widely used benchmark in the field, including for development of approaches that can simulate the protein interactions that are described in <code>all_sequences</code> (or ligand information or human readable<br>\nInformation in the description field).</p>\n<p>However for the hidden test set and for the future targets for the current competition, we will remove targets that have duplicates in the main RNA sequence field.</p>\n<p>It’s a great idea to also curate training data sets to bring together duplicates or near-duplicates as alternative labels for a target sequence.</p>\n<p>If you or others have seen a positive effect on your models from such aggregation of labels, please post and we will considering releasing supplemental files with that rearrangement for the community, or linking to your alternative data set. Please also see the post from <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a>: <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556</a></p>",
      "rawMarkdown": "Thanks for the post! \n\nWe included these duplicate sequences TR1189 and TR1190 in the host provided validation because they were indeed separate targets in CASP15 and have become a widely used benchmark in the field, including for development of approaches that can simulate the protein interactions that are described in `all_sequences` (or ligand information or human readable\nInformation in the description field).\n\nHowever for the hidden test set and for the future targets for the current competition, we will remove targets that have duplicates in the main RNA sequence field.\n\nIt’s a great idea to also curate training data sets to bring together duplicates or near-duplicates as alternative labels for a target sequence.\n\n If you or others have seen a positive effect on your models from such aggregation of labels, please post and we will considering releasing supplemental files with that rearrangement for the community, or linking to your alternative data set. Please also see the post from @ckjoshi9: https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556",
      "votes": null
    },
    {
      "id": "3196898",
      "postDate": "05/07/2025 15:05:02",
      "content": "<p>Thanks for your clarification and the link to the useful post. </p>",
      "rawMarkdown": "Thanks for your clarification and the link to the useful post.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3196838,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "05/07/2025 13:39:33",
      "content": "<p>Thanks for the post! </p>\n<p>We included these duplicate sequences TR1189 and TR1190 in the host provided validation because they were indeed separate targets in CASP15 and have become a widely used benchmark in the field, including for development of approaches that can simulate the protein interactions that are described in <code>all_sequences</code> (or ligand information or human readable<br>\nInformation in the description field).</p>\n<p>However for the hidden test set and for the future targets for the current competition, we will remove targets that have duplicates in the main RNA sequence field.</p>\n<p>It’s a great idea to also curate training data sets to bring together duplicates or near-duplicates as alternative labels for a target sequence.</p>\n<p>If you or others have seen a positive effect on your models from such aggregation of labels, please post and we will considering releasing supplemental files with that rearrangement for the community, or linking to your alternative data set. Please also see the post from <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a>: <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3196898,
          "author_name": "nguyenhoa",
          "author_url": "",
          "post_date": "05/07/2025 15:05:02",
          "content": "<p>Thanks for your clarification and the link to the useful post. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3195371": "Dear @rhijudas das and @shujun717,\n\nThank you again for organizing this exciting challenge. I’d like to revisit [a question I asked earlier](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324) regarding validation, which I believe is still unclear and important for the fairness and consistency of the evaluation.\n\nIn the `validation_sequences.csv` and `validation_labels.csv` files, I noticed that 2 target IDs, R1189 and R1190, have an identical RNA sequence but are treated as separate targets:\n\n* R1189 (7YR7_1) is associated with 6 chains including CsrA protein,\n\n* R1190 (7YR6_1) is associated with 4 chains of CsrA.\n\nMy questions:\n* Why were these considered distinct target IDs, rather than treating them as the same target with multiple possible 3D conformations (i.e., multiple valid ground truths for the same sequence)?\n\n* If the rationale is based on the accompanying `all_sequences` (which lists macromolecule chains), how can we distinguish cases where RNA is co-crystallized with small molecules, ions, or ligands, since these are not included in `all_sequences`?\n\nAs previously raised in [this discussion](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/571324) , small molecules and ions can significantly influence RNA conformation, yet we currently lack the input features needed to account for them.\n\n**Suggestion**:\nWould it be possible to merge all structures with the same RNA sequence into a single target ID, and consider all of their 3D conformations as valid ground truths?\nThis could reduce ambiguity during testing and better reflect the biological reality where a single RNA sequence may adopt multiple conformations depending on context (e.g., protein, ion binding).\n\nThank you for considering this. Looking forward to your thoughts!",
    "3196838": "Thanks for the post! \n\nWe included these duplicate sequences TR1189 and TR1190 in the host provided validation because they were indeed separate targets in CASP15 and have become a widely used benchmark in the field, including for development of approaches that can simulate the protein interactions that are described in `all_sequences` (or ligand information or human readable\nInformation in the description field).\n\nHowever for the hidden test set and for the future targets for the current competition, we will remove targets that have duplicates in the main RNA sequence field.\n\nIt’s a great idea to also curate training data sets to bring together duplicates or near-duplicates as alternative labels for a target sequence.\n\n If you or others have seen a positive effect on your models from such aggregation of labels, please post and we will considering releasing supplemental files with that rearrangement for the community, or linking to your alternative data set. Please also see the post from @ckjoshi9: https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556",
    "3196898": "Thanks for your clarification and the link to the useful post."
  },
  "source": "meta"
}