{
  "id": 569502,
  "title": "Why only 844 sequences and ideas about more data",
  "url": "/competitions/stanford-rna-3d-folding/discussion/569502",
  "author_name": "",
  "post_date": "2025-03-22T08:50:59.459004600Z",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I wondered why there are only 844 sequences in the training data, while there are ~20k PDB entries with RNA chains. Here's what I found in the competition data processing pipeline published <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/569085\" target=\"_blank\">here</a>.<br>\nFrom an initial pool of potentially thousands of RNA chains, the host's pipeline applies stringent filters:</p>\n<ol>\n<li><strong>Initial Pool</strong>: get_pdb_data.py retrieves up to 10,000 RNA-containing entries, but many include proteins or DNA. Splitting into chains might yield ~10,000-20,000 RNA chains if each entry averages 1-2 RNA chains.</li>\n<li><strong>RNA-only Filter</strong>: split_chains.py keeps only pure RNA chains (A, C, G, U), excluding modified nucleotides or hybrids. This could halve the pool (e.g., ~5,000-10,000).</li>\n<li><strong>Structuredness &gt; 0.2</strong>: get_xyz_data.py filters out low-quality structures. If only 20-30% of chains are well-structured, this drops to ~1,000-3,000.</li>\n<li><strong>Chain Breaks &lt; 10 Å</strong>: Excludes discontinuous chains, further reducing to ~1,000-2,000.</li>\n<li><strong>Duplicate Removal and Clustering</strong>: Keeps only unique sequences, collapsing redundant ones (e.g., rRNAs or tRNAs appearing in multiple structures). If 50-70% are redundant, this yields ~500-1,000.</li>\n<li><strong>Length &lt; 1000 nt</strong>: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.</li>\n</ol>\n<p>Now we understand how the data was filtered, we can come up with ideas to lift some of those filters and get more training data. Here is a tradeoff between data quality and quantity. And our model might perform better or worse depending on our choice.</p>",
  "messages": [
    {
      "id": "3156536",
      "postDate": "03/22/2025 08:50:59",
      "content": "<p>I wondered why there are only 844 sequences in the training data, while there are ~20k PDB entries with RNA chains. Here's what I found in the competition data processing pipeline published <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/569085\" target=\"_blank\">here</a>.<br>\nFrom an initial pool of potentially thousands of RNA chains, the host's pipeline applies stringent filters:</p>\n<ol>\n<li><strong>Initial Pool</strong>: get_pdb_data.py retrieves up to 10,000 RNA-containing entries, but many include proteins or DNA. Splitting into chains might yield ~10,000-20,000 RNA chains if each entry averages 1-2 RNA chains.</li>\n<li><strong>RNA-only Filter</strong>: split_chains.py keeps only pure RNA chains (A, C, G, U), excluding modified nucleotides or hybrids. This could halve the pool (e.g., ~5,000-10,000).</li>\n<li><strong>Structuredness &gt; 0.2</strong>: get_xyz_data.py filters out low-quality structures. If only 20-30% of chains are well-structured, this drops to ~1,000-3,000.</li>\n<li><strong>Chain Breaks &lt; 10 Å</strong>: Excludes discontinuous chains, further reducing to ~1,000-2,000.</li>\n<li><strong>Duplicate Removal and Clustering</strong>: Keeps only unique sequences, collapsing redundant ones (e.g., rRNAs or tRNAs appearing in multiple structures). If 50-70% are redundant, this yields ~500-1,000.</li>\n<li><strong>Length &lt; 1000 nt</strong>: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.</li>\n</ol>\n<p>Now we understand how the data was filtered, we can come up with ideas to lift some of those filters and get more training data. Here is a tradeoff between data quality and quantity. And our model might perform better or worse depending on our choice.</p>",
      "rawMarkdown": "I wondered why there are only 844 sequences in the training data, while there are ~20k PDB entries with RNA chains. Here's what I found in the competition data processing pipeline published [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/569085).\nFrom an initial pool of potentially thousands of RNA chains, the host's pipeline applies stringent filters:\n\n1. **Initial Pool**: get_pdb_data.py retrieves up to 10,000 RNA-containing entries, but many include proteins or DNA. Splitting into chains might yield ~10,000-20,000 RNA chains if each entry averages 1-2 RNA chains.\n2. **RNA-only Filter**: split_chains.py keeps only pure RNA chains (A, C, G, U), excluding modified nucleotides or hybrids. This could halve the pool (e.g., ~5,000-10,000).\n3. **Structuredness > 0.2**: get_xyz_data.py filters out low-quality structures. If only 20-30% of chains are well-structured, this drops to ~1,000-3,000.\n4. **Chain Breaks < 10 Å**: Excludes discontinuous chains, further reducing to ~1,000-2,000.\n5. **Duplicate Removal and Clustering**: Keeps only unique sequences, collapsing redundant ones (e.g., rRNAs or tRNAs appearing in multiple structures). If 50-70% are redundant, this yields ~500-1,000.\n6. **Length < 1000 nt**: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.\n\nNow we understand how the data was filtered, we can come up with ideas to lift some of those filters and get more training data. Here is a tradeoff between data quality and quantity. And our model might perform better or worse depending on our choice.",
      "votes": null
    },
    {
      "id": "3156677",
      "postDate": "03/22/2025 11:56:31",
      "content": "<p>Thanks for share. Intereasting. \"Length &lt; 1000 nt: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.\" But at train there is &gt;1000 nt chains.</p>",
      "rawMarkdown": "Thanks for share. Intereasting. \"Length < 1000 nt: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.\" But at train there is >1000 nt chains.",
      "votes": null
    },
    {
      "id": "3157021",
      "postDate": "03/22/2025 21:03:04",
      "content": "<p>I strongly recommend that you avoid very small RNAs bound by proteins. They are most likely unstructured on their own, but in the complex with protein they are in a somewhat fixed conformation.</p>\n<p>The majority of redundant entries will be ribosomal RNAs (rRNAs) in the context of ribosomes. There are excellent sequence-based co-variance models for rRNAs and their folding is well understood. I doubt it will make much of a difference if you add 100-200 structures with rRNAs. Same with tRNAs.</p>\n<p>Including structures with modified nucleotides might help.</p>\n<p>Structures with chain breaks could be useful if one properly accounts for missing nucleotides in terms of numbering.</p>",
      "rawMarkdown": "I strongly recommend that you avoid very small RNAs bound by proteins. They are most likely unstructured on their own, but in the complex with protein they are in a somewhat fixed conformation.\n\nThe majority of redundant entries will be ribosomal RNAs (rRNAs) in the context of ribosomes. There are excellent sequence-based co-variance models for rRNAs and their folding is well understood. I doubt it will make much of a difference if you add 100-200 structures with rRNAs. Same with tRNAs.\n\nIncluding structures with modified nucleotides might help.\n\nStructures with chain breaks could be useful if one properly accounts for missing nucleotides in terms of numbering.",
      "votes": null
    },
    {
      "id": "3157085",
      "postDate": "03/22/2025 23:28:04",
      "content": "<p>my suggestion is first to use dataset that has been used by open source model like af3, drfold, etc. many have their list of pdb downloadable.</p>",
      "rawMarkdown": "my suggestion is first to use dataset that has been used by open source model like af3, drfold, etc. many have their list of pdb downloadable.",
      "votes": null
    },
    {
      "id": "3158415",
      "postDate": "03/24/2025 14:17:54",
      "content": "<p>Can you point to any of the resources you mentioned?</p>",
      "rawMarkdown": "Can you point to any of the resources you mentioned?",
      "votes": null
    },
    {
      "id": "3158895",
      "postDate": "03/25/2025 03:31:31",
      "content": "<p>as an exmple<br>\n<a href=\"https://yanglab.qd.sdu.edu.cn/trRosettaRNA/benchmark/\" target=\"_blank\">https://yanglab.qd.sdu.edu.cn/trRosettaRNA/benchmark/</a></p>",
      "rawMarkdown": "as an exmple\nhttps://yanglab.qd.sdu.edu.cn/trRosettaRNA/benchmark/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3156677,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "03/22/2025 11:56:31",
      "content": "<p>Thanks for share. Intereasting. \"Length &lt; 1000 nt: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.\" But at train there is &gt;1000 nt chains.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3157021,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "03/22/2025 21:03:04",
      "content": "<p>I strongly recommend that you avoid very small RNAs bound by proteins. They are most likely unstructured on their own, but in the complex with protein they are in a somewhat fixed conformation.</p>\n<p>The majority of redundant entries will be ribosomal RNAs (rRNAs) in the context of ribosomes. There are excellent sequence-based co-variance models for rRNAs and their folding is well understood. I doubt it will make much of a difference if you add 100-200 structures with rRNAs. Same with tRNAs.</p>\n<p>Including structures with modified nucleotides might help.</p>\n<p>Structures with chain breaks could be useful if one properly accounts for missing nucleotides in terms of numbering.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3157085,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/22/2025 23:28:04",
      "content": "<p>my suggestion is first to use dataset that has been used by open source model like af3, drfold, etc. many have their list of pdb downloadable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3158415,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "03/24/2025 14:17:54",
          "content": "<p>Can you point to any of the resources you mentioned?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3158895,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/25/2025 03:31:31",
              "content": "<p>as an exmple<br>\n<a href=\"https://yanglab.qd.sdu.edu.cn/trRosettaRNA/benchmark/\" target=\"_blank\">https://yanglab.qd.sdu.edu.cn/trRosettaRNA/benchmark/</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3156536": "I wondered why there are only 844 sequences in the training data, while there are ~20k PDB entries with RNA chains. Here's what I found in the competition data processing pipeline published [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/569085).\nFrom an initial pool of potentially thousands of RNA chains, the host's pipeline applies stringent filters:\n\n1. **Initial Pool**: get_pdb_data.py retrieves up to 10,000 RNA-containing entries, but many include proteins or DNA. Splitting into chains might yield ~10,000-20,000 RNA chains if each entry averages 1-2 RNA chains.\n2. **RNA-only Filter**: split_chains.py keeps only pure RNA chains (A, C, G, U), excluding modified nucleotides or hybrids. This could halve the pool (e.g., ~5,000-10,000).\n3. **Structuredness > 0.2**: get_xyz_data.py filters out low-quality structures. If only 20-30% of chains are well-structured, this drops to ~1,000-3,000.\n4. **Chain Breaks < 10 Å**: Excludes discontinuous chains, further reducing to ~1,000-2,000.\n5. **Duplicate Removal and Clustering**: Keeps only unique sequences, collapsing redundant ones (e.g., rRNAs or tRNAs appearing in multiple structures). If 50-70% are redundant, this yields ~500-1,000.\n6. **Length < 1000 nt**: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.\n\nNow we understand how the data was filtered, we can come up with ideas to lift some of those filters and get more training data. Here is a tradeoff between data quality and quantity. And our model might perform better or worse depending on our choice.",
    "3156677": "Thanks for share. Intereasting. \"Length < 1000 nt: Excludes long RNAs, though most PDB RNAs are short (e.g., tRNAs ~76 nt, riboswitches ~100-200 nt), so this has minimal impact.\" But at train there is >1000 nt chains.",
    "3157021": "I strongly recommend that you avoid very small RNAs bound by proteins. They are most likely unstructured on their own, but in the complex with protein they are in a somewhat fixed conformation.\n\nThe majority of redundant entries will be ribosomal RNAs (rRNAs) in the context of ribosomes. There are excellent sequence-based co-variance models for rRNAs and their folding is well understood. I doubt it will make much of a difference if you add 100-200 structures with rRNAs. Same with tRNAs.\n\nIncluding structures with modified nucleotides might help.\n\nStructures with chain breaks could be useful if one properly accounts for missing nucleotides in terms of numbering.",
    "3157085": "my suggestion is first to use dataset that has been used by open source model like af3, drfold, etc. many have their list of pdb downloadable.",
    "3158415": "Can you point to any of the resources you mentioned?",
    "3158895": "as an exmple\nhttps://yanglab.qd.sdu.edu.cn/trRosettaRNA/benchmark/"
  },
  "source": "meta"
}