{
  "id": 576556,
  "title": "All known RNA 3D structures clustered by sequence and structural similarity",
  "url": "/competitions/stanford-rna-3d-folding/discussion/576556",
  "author_name": "Chaitanya Joshi",
  "post_date": "2025-05-05T18:08:19.427000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey RNA folks!</p>\n<p>If you've been training models with the officially released datasets but want to challenge your models in a more out of distribution setting, I've released <strong>all RNA 3D structures from the PDB</strong> clustered by both sequence and structural similarity. This is part of the <a href=\"https://github.com/chaitjo/geometric-rna-design\" target=\"_blank\">gRNAde</a> project for 3D RNA structure design.</p>\n<p>The data was prepared using <a href=\"https://rnasolo.cs.put.poznan.pl/\" target=\"_blank\">RNASolo</a>, a repository of all known RNA 3D structures from the PDB extracted from solo RNAs as well as biomolecular complexes.</p>\n<p>Processed data in machine learning-ready format (slightly different from this contest) is available here: <a href=\"https://drive.google.com/file/d/1gcUUaRxbGZnGMkLdtVwAILWVerVCbu4Y/view\" target=\"_blank\">https://drive.google.com/file/d/1gcUUaRxbGZnGMkLdtVwAILWVerVCbu4Y/view</a> </p>\n<p>The attached jupyter notebook shows you how to work with the data -- hopefully it <strong>should be very easy to construct a csv file</strong> in the same format as this contest. The notebook links to download processed gRNAde data, access cluster IDs, and handing missing nucleotides and/or missing coordinates.</p>\n<ul>\n<li>There are <strong>over 12 thousand</strong> 3D structures from 4.2K unique RNA sequences.</li>\n<li>The processed coordinates use an Atom27 format where each nucleotide has a subset of 27 possible atoms. The atoms are ordered in a pre-defined manner which is specified in the notebook.</li>\n<li>There are 4223 unique sequences, and 1102 structural clusters determined by USalign with a cutoff of 0.45 TM score.</li>\n<li>And 1002 clusters grouped by sequence identity at a cutoff 0.8 determined by CD-Hit.</li>\n</ul>\n<p>Another good source of external data is <a href=\"https://github.com/marcellszi/rna3db\" target=\"_blank\">RNA3DB</a>, which provides the same functionality with even more rigorous pre-processing and preparation.</p>",
  "messages": [
    {
      "id": 3194342,
      "postDate": "2025-05-05T18:08:19.427Z",
      "content": "<p>Hey RNA folks!</p>\n<p>If you've been training models with the officially released datasets but want to challenge your models in a more out of distribution setting, I've released <strong>all RNA 3D structures from the PDB</strong> clustered by both sequence and structural similarity. This is part of the <a href=\"https://github.com/chaitjo/geometric-rna-design\" target=\"_blank\">gRNAde</a> project for 3D RNA structure design.</p>\n<p>The data was prepared using <a href=\"https://rnasolo.cs.put.poznan.pl/\" target=\"_blank\">RNASolo</a>, a repository of all known RNA 3D structures from the PDB extracted from solo RNAs as well as biomolecular complexes.</p>\n<p>Processed data in machine learning-ready format (slightly different from this contest) is available here: <a href=\"https://drive.google.com/file/d/1gcUUaRxbGZnGMkLdtVwAILWVerVCbu4Y/view\" target=\"_blank\">https://drive.google.com/file/d/1gcUUaRxbGZnGMkLdtVwAILWVerVCbu4Y/view</a> </p>\n<p>The attached jupyter notebook shows you how to work with the data -- hopefully it <strong>should be very easy to construct a csv file</strong> in the same format as this contest. The notebook links to download processed gRNAde data, access cluster IDs, and handing missing nucleotides and/or missing coordinates.</p>\n<ul>\n<li>There are <strong>over 12 thousand</strong> 3D structures from 4.2K unique RNA sequences.</li>\n<li>The processed coordinates use an Atom27 format where each nucleotide has a subset of 27 possible atoms. The atoms are ordered in a pre-defined manner which is specified in the notebook.</li>\n<li>There are 4223 unique sequences, and 1102 structural clusters determined by USalign with a cutoff of 0.45 TM score.</li>\n<li>And 1002 clusters grouped by sequence identity at a cutoff 0.8 determined by CD-Hit.</li>\n</ul>\n<p>Another good source of external data is <a href=\"https://github.com/marcellszi/rna3db\" target=\"_blank\">RNA3DB</a>, which provides the same functionality with even more rigorous pre-processing and preparation.</p>",
      "rawMarkdown": "Hey RNA folks!\n\nIf you've been training models with the officially released datasets but want to challenge your models in a more out of distribution setting, I've released **all RNA 3D structures from the PDB** clustered by both sequence and structural similarity. This is part of the [gRNAde](https://github.com/chaitjo/geometric-rna-design) project for 3D RNA structure design.\n\nThe data was prepared using [RNASolo](https://rnasolo.cs.put.poznan.pl/), a repository of all known RNA 3D structures from the PDB extracted from solo RNAs as well as biomolecular complexes.\n\nProcessed data in machine learning-ready format (slightly different from this contest) is available here: https://drive.google.com/file/d/1gcUUaRxbGZnGMkLdtVwAILWVerVCbu4Y/view \n\nThe attached jupyter notebook shows you how to work with the data -- hopefully it **should be very easy to construct a csv file** in the same format as this contest. The notebook links to download processed gRNAde data, access cluster IDs, and handing missing nucleotides and/or missing coordinates.\n- There are **over 12 thousand** 3D structures from 4.2K unique RNA sequences.\n- The processed coordinates use an Atom27 format where each nucleotide has a subset of 27 possible atoms. The atoms are ordered in a pre-defined manner which is specified in the notebook.\n- There are 4223 unique sequences, and 1102 structural clusters determined by USalign with a cutoff of 0.45 TM score.\n- And 1002 clusters grouped by sequence identity at a cutoff 0.8 determined by CD-Hit.\n\nAnother good source of external data is [RNA3DB](https://github.com/marcellszi/rna3db), which provides the same functionality with even more rigorous pre-processing and preparation.",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3194342": "Hey RNA folks!\n\nIf you've been training models with the officially released datasets but want to challenge your models in a more out of distribution setting, I've released **all RNA 3D structures from the PDB** clustered by both sequence and structural similarity. This is part of the [gRNAde](https://github.com/chaitjo/geometric-rna-design) project for 3D RNA structure design.\n\nThe data was prepared using [RNASolo](https://rnasolo.cs.put.poznan.pl/), a repository of all known RNA 3D structures from the PDB extracted from solo RNAs as well as biomolecular complexes.\n\nProcessed data in machine learning-ready format (slightly different from this contest) is available here: https://drive.google.com/file/d/1gcUUaRxbGZnGMkLdtVwAILWVerVCbu4Y/view \n\nThe attached jupyter notebook shows you how to work with the data -- hopefully it **should be very easy to construct a csv file** in the same format as this contest. The notebook links to download processed gRNAde data, access cluster IDs, and handing missing nucleotides and/or missing coordinates.\n- There are **over 12 thousand** 3D structures from 4.2K unique RNA sequences.\n- The processed coordinates use an Atom27 format where each nucleotide has a subset of 27 possible atoms. The atoms are ordered in a pre-defined manner which is specified in the notebook.\n- There are 4223 unique sequences, and 1102 structural clusters determined by USalign with a cutoff of 0.45 TM score.\n- And 1002 clusters grouped by sequence identity at a cutoff 0.8 determined by CD-Hit.\n\nAnother good source of external data is [RNA3DB](https://github.com/marcellszi/rna3db), which provides the same functionality with even more rigorous pre-processing and preparation."
  }
}