{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":12276181,"sourceType":"competition"},{"sourceId":11644010,"sourceType":"datasetVersion","datasetId":7306643},{"sourceId":11969392,"sourceType":"datasetVersion","datasetId":7526656}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Dataset Enrichment via GraphQL (RCSB PDB API)\n\nThis notebook augments sequence-level training data with structure-aware, evolutionary, and literature-linked metadata retrieved from the **RCSB Protein Data Bank GraphQL API**. Using GraphQL enables precise, schema-driven queries and avoids downloading large, unnecessary metadata tables. The resulting enriched dataset is provided in the notebook’s output directory as a preprocessed CSV, ready for downstream use.\n\n**What metadata is added**\n- **Experimental method** (e.g. **SOLUTION NMR**, **X-ray diffraction**, **Cryo-EM**)\n- **Primary citation information** (**PubMed ID**, **DOI**)\n- **Source organism taxonomy** (NCBI taxonomy ID and scientific name)\n- **Polymer source and evolutionary context** from PDB entity records\n\nAll queried fields are merged into the training dataframe using `rcsb_id` as the key. When metadata is unavailable for a given entry (e.g. missing citations or organism annotations), values are explicitly filled with **NaNs** to preserve dataframe shape and ensure compatibility with downstream filtering, modeling, and batching workflows.\n\n---\n\n### Why these features are useful\n\n**Experimental method (for TBM and template selection)**\n- In **template-based modeling (TBM)**, not all structural templates are equally informative.\n- Experimental method provides a coarse but useful proxy for structural reliability and representational bias:\n  - **X-ray / Cryo-EM** structures often provide well-defined global folds suitable for rigid-template alignment.\n  - **NMR** structures may represent conformational ensembles that are valuable for flexibility-aware modeling but less optimal as single static templates.\n- This metadata enables:\n  - Filtering MMseqs2 hits to prioritize templates solved by specific methods.\n  - Weighting candidate templates differently during TBM rather than treating all hits as equivalent.\n  - Method-aware benchmarking to avoid conflating modeling performance with experimental artifacts.\n\n**Evolutionary and organism-level context (for MMseqs2 filtering and hit prioritization)**\n- Taxonomic annotations provide evolutionary signal that complements sequence similarity scores from **MMseqs2**.\n- Closely related organisms often yield templates with more transferable structural and functional features, even at similar sequence identity.\n- Evolutionary metadata enables:\n  - Filtering or re-ranking MMseqs2 hits based on phylogenetic proximity rather than raw similarity alone.\n  - Emphasizing representative templates from specific clades while down-weighting redundant or evolutionarily distant matches.\n  - More principled template selection in low-identity regimes where sequence similarity alone is ambiguous.\n\n**DOI and PubMed identifiers (for RAG and interpretability)**\n- Citation metadata creates a direct bridge between structured structural data and the primary literature.\n- **PubMed IDs** and **DOIs** can be used in **retrieval-augmented generation (RAG) pipelines** to fetch abstracts or full-text articles associated with selected templates.\n- This enables automated retrieval of experimental details, functional annotations, and known limitations to support literature-grounded interpretation of modeling decisions.\n\n---\n\n### New dataset: train_seqs_extended_pdbdata.csv\n\nFor each entry in `train_seqs[\"target_id\"]`, the corresponding PDB ID is parsed and used to issue a targeted GraphQL query against the PDB `entries` endpoint. Because not all `target_id` values correspond to valid PDB entries (or may be missing entirely), the pipeline explicitly checks for invalid or missing identifiers and fills all unavailable fields with NaNs.\n\nThe resulting columns—`experimental_method`, `pubmed_id`, `doi`, `ncbi_taxonomy_id`, `organism_name`, and `bound_components`—are assigned to the dataframe in a single pass after querying, avoiding per-row overwrites and enabling clean downstream use.","metadata":{}},{"cell_type":"code","source":"import time\n\nimport pandas as pd\nimport numpy as np\n\nimport random\n#from Bio import pairwise2\n#from Bio.Seq import Seq\n%pip install rcsb-api\n\nfrom tqdm import tqdm\nfrom rcsbapi.data import DataQuery as Query\nfrom rcsbapi.data import DataQuery as Query\nimport json  # for easy-to-read output\n\nfrom scipy.spatial.transform import Rotation as R\nfrom sklearn.preprocessing import normalize\nfrom scipy.spatial import distance_matrix\nimport warnings\nwarnings.filterwarnings('ignore')\n\nprint(\"\\nLoading data files...\")\ntrain_seqs = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/train_sequences.csv')\nvalid_seqs = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/validation_sequences.csv')\ntest_seqs = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/test_sequences.csv')\ntrain_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/train_labels.csv')\nvalid_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/validation_labels.csv')\n\nprint(f\"Loaded {len(train_seqs)} training sequences, {len(valid_seqs)} validation sequences, and {len(test_seqs)} test sequences\")","metadata":{"_uuid":"7e2d63ef-4429-4722-95ee-6681af62a6d5","_cell_guid":"0ec49051-b8ac-4f27-b152-0c2fe59d4b5c","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2026-01-12T18:57:16.853770Z","iopub.execute_input":"2026-01-12T18:57:16.854114Z","iopub.status.idle":"2026-01-12T18:57:26.995578Z","shell.execute_reply.started":"2026-01-12T18:57:16.854085Z","shell.execute_reply":"2026-01-12T18:57:26.994294Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"experimental_methods = []\npubmed_ids = []\ndois = []\nncbi_taxonomy_ids = []\norganism_names = []\nbound_components = []\nfor i in train_seqs[\"target_id\"]:\n    if isinstance(i, float):\n        experimental_methods.append(np.nan)\n        pubmed_ids.append(np.nan)\n        dois.append(np.nan)\n        ncbi_taxonomy_ids.append(np.nan)\n        organism_names.append(np.nan)\n        bound_components.append(np.nan)\n        continue\n        \n    pdb_id, chain_id = i.split(\"_\")\n    query = Query(\n    input_type=\"entries\",\n    input_ids=[pdb_id],\n    return_data_list=[\"exptl.method\", \n                          \"rcsb_primary_citation.pdbx_database_id_PubMed\", \n                          \"rcsb_primary_citation.pdbx_database_id_DOI\", \n                          \"rcsb_entity_source_organism.ncbi_taxonomy_id\", \n                          \"rcsb_entity_source_organism.ncbi_scientific_name\",\n                          \"nonpolymer_bound_components\"\n                         ])\n    return_data = query.exec()\n    entry = return_data[\"data\"][\"entries\"][0]\n    \n    experimental_methods.append(\n        entry.get(\"exptl\", [{}])[0].get(\"method\")\n    )\n    \n    pubmed_ids.append(\n        entry.get(\"rcsb_primary_citation\", {}).get(\"pdbx_database_id_PubMed\")\n    )\n    \n    dois.append(\n        entry.get(\"rcsb_primary_citation\", {}).get(\"pdbx_database_id_DOI\")\n    )\n\n    bound_components.append(\n    (entry.get(\"rcsb_entry_info\") or {}).get(\"nonpolymer_bound_components\", np.nan)\n    )\n\n    poly0 = (entry.get(\"polymer_entities\") or [{}])[0]\n    org0  = (poly0.get(\"rcsb_entity_source_organism\") or [{}])[0]\n\n    ncbi_taxonomy_ids.append(org0.get(\"ncbi_taxonomy_id\", np.nan))\n    organism_names.append(org0.get(\"ncbi_scientific_name\", np.nan))\n    \n\n# Assign once (no overwriting per-iteration)\ntrain_seqs[\"experimental_method\"] = experimental_methods\ntrain_seqs[\"pubmed_id\"] = pubmed_ids\ntrain_seqs[\"doi\"] = dois\ntrain_seqs[\"ncbi_taxonomy_id\"] = ncbi_taxonomy_ids\ntrain_seqs[\"organism_name\"] = organism_names\ntrain_seqs[\"bound_components\"] = bound_components","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T18:57:26.996835Z","iopub.execute_input":"2026-01-12T18:57:26.997111Z","iopub.status.idle":"2026-01-12T19:04:00.679247Z","shell.execute_reply.started":"2026-01-12T18:57:26.997088Z","shell.execute_reply":"2026-01-12T19:04:00.678152Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seqs","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T19:04:00.680150Z","iopub.execute_input":"2026-01-12T19:04:00.680385Z","iopub.status.idle":"2026-01-12T19:04:00.712182Z","shell.execute_reply.started":"2026-01-12T19:04:00.680362Z","shell.execute_reply":"2026-01-12T19:04:00.711045Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seqs.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T19:35:32.776396Z","iopub.execute_input":"2026-01-12T19:35:32.776696Z","iopub.status.idle":"2026-01-12T19:35:32.784959Z","shell.execute_reply.started":"2026-01-12T19:35:32.776671Z","shell.execute_reply":"2026-01-12T19:35:32.783938Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#example query using GraphQL\nquery = Query(\n    input_type=\"entries\",\n    input_ids=[\"8Z1F\"],\n    return_data_list=[\"exptl.method\", \"ma_qa_metric_global.type\", \"ma_qa_metric_global.value\", \"rcsb_primary_citation.pdbx_database_id_PubMed\", \"rcsb_primary_citation.pdbx_database_id_DOI\", \"nonpolymer_bound_components\", \"rcsb_entity_source_organism.ncbi_taxonomy_id\", \"rcsb_entity_source_organism.ncbi_scientific_name\"]\n)\nreturn_data = query.exec()\nprint(return_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T19:35:34.049866Z","iopub.execute_input":"2026-01-12T19:35:34.050154Z","iopub.status.idle":"2026-01-12T19:35:34.579143Z","shell.execute_reply.started":"2026-01-12T19:35:34.050136Z","shell.execute_reply":"2026-01-12T19:35:34.578190Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seqs.to_csv('/kaggle/working/trainseqs_extended_pdbdata.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T19:36:57.897223Z","iopub.execute_input":"2026-01-12T19:36:57.897531Z","iopub.status.idle":"2026-01-12T19:36:57.946445Z","shell.execute_reply.started":"2026-01-12T19:36:57.897509Z","shell.execute_reply":"2026-01-12T19:36:57.945340Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}