{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":11553390,"isSourceIdPinned":false,"sourceType":"competition"},{"sourceId":11184470,"sourceType":"datasetVersion","datasetId":6981588},{"sourceId":11185178,"sourceType":"datasetVersion","datasetId":6982132}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Stanford 3D Folding: Minimal EDA\n\nThis notebook provides a basic overview of the dataset, covering:\n\n1. Preprocessing the competition’s sequence and label metadata  \n2. Analyzing sequence length distributions in the training and validation sets  \n3. Visualizing 3D structures of RNA chains using the provided label coordinates (C1' atoms)","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nfrom Bio import SeqIO\nimport polars as pl","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:52:53.525837Z","iopub.execute_input":"2025-03-27T09:52:53.526188Z","iopub.status.idle":"2025-03-27T09:52:53.530732Z","shell.execute_reply.started":"2025-03-27T09:52:53.526152Z","shell.execute_reply":"2025-03-27T09:52:53.529720Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Preprocess Sequence and Label Metadata\n\nIn the code below, we preprocess the training, validation, and test sequence data along with their labels.\n\nThe main purpose of this step is to convert the label data from a **wide format** (one row per sequence with many columns) to a **long format** (one row per residue). This makes it easier to handle specific 3D structures later on.","metadata":{}},{"cell_type":"code","source":"metadata_dir = Path(\"metadata\")\nif not metadata_dir.exists():\n    metadata_dir.mkdir(parents=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:57:23.898412Z","iopub.execute_input":"2025-03-27T09:57:23.898821Z","iopub.status.idle":"2025-03-27T09:57:23.903580Z","shell.execute_reply.started":"2025-03-27T09:57:23.898785Z","shell.execute_reply":"2025-03-27T09:57:23.902585Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import polars as pl\n\n\ndef preprocess_label(df, no_chain=False, invalid_value=-1e18):\n    \"\"\"\n    - Convert wide to long format.\n    - Filter out invalid values.\n    \"\"\"\n\n    # convert to long format\n    df = (\n        df.unpivot(index=[\"ID\", \"resname\", \"resid\"])\n        .with_columns(\n            [\n                pl.col(\"variable\").str.extract(r\"^([xyz])\", 1).alias(\"coord\"),\n                pl.col(\"variable\")\n                .str.extract(r\"_(\\d+)\", 1)\n                .cast(pl.Int32)\n                .alias(\"design_id\"),\n            ]\n        )\n        .pivot(\n            values=\"value\",\n            index=[\"ID\", \"resname\", \"resid\", \"design_id\"],\n            on=\"coord\",\n        )\n        # remove invalid values\n        .filter(pl.col(\"x\").gt(invalid_value))\n    )\n\n    if no_chain:\n        # convert to long format\n        df = (\n            # split ID columns\n            df.with_columns(pl.col(\"ID\").str.splitn(\"_\", 2).alias(\"_a\"))\n            .unnest(\"_a\")\n            .rename({\"field_0\": \"pdb_id\", \"field_1\": \"res_id\"})\n            .with_columns(pl.lit(None).alias(\"chain_id\"))\n        )\n    else:\n        df = (\n            # split ID columns\n            df.with_columns(pl.col(\"ID\").str.splitn(\"_\", 3).alias(\"_a\"))\n            .unnest(\"_a\")\n            .rename({\"field_0\": \"pdb_id\", \"field_1\": \"chain_id\", \"field_2\": \"res_id\"})\n        )\n\n    return (\n        df.with_columns(\n            pl.format(\"{}_{}\", pl.col(\"pdb_id\"), pl.col(\"res_id\")).alias(\"target_id\"),\n        )\n        .sort([\"pdb_id\", \"design_id\", \"chain_id\", \"resid\"])\n        .select(\n            \"ID\",\n            \"target_id\",\n            \"pdb_id\",\n            \"design_id\",\n            \"chain_id\",\n            \"resid\",\n            \"resname\",\n            \"x\",\n            \"y\",\n            \"z\",\n        )\n    )\n\n\ndef preprocess_sequence(df):\n    return (\n        df.with_columns(pl.col(\"target_id\").str.splitn(\"_\", 2).alias(\"_a\"))\n        .unnest(\"_a\")\n        .rename({\"field_0\": \"pdb_id\", \"field_1\": \"chain_id\"})\n        .with_columns(\n            pl.col(\"temporal_cutoff\")\n            .str.strptime(pl.Date, format=\"%Y-%m-%d\")\n            .alias(\"temporal_cutoff\")\n        )\n        .with_columns(\n            pl.col(\"sequence\").str.len_chars().alias(\"sequence_length\"),\n        )\n        .sort([\"pdb_id\", \"chain_id\"])\n        .select(\n            \"target_id\",\n            \"pdb_id\",\n            \"chain_id\",\n            \"temporal_cutoff\",\n            \"sequence\",\n            \"all_sequences\",\n            \"description\",\n            \"sequence_length\",\n        )\n    )\n\n\ntrain_sequence = pl.read_csv(\n    \"/kaggle/input/stanford-rna-3d-folding/train_sequences.csv\"\n)\ntrain_label = pl.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_labels.csv\")\nval_label = pl.read_csv(\"/kaggle/input/stanford-rna-3d-folding/validation_labels.csv\")\nval_sequence = pl.read_csv(\n    \"/kaggle/input/stanford-rna-3d-folding/validation_sequences.csv\"\n)\ntest_sequence = pl.read_csv(\"/kaggle/input/stanford-rna-3d-folding/test_sequences.csv\")\nval_label = preprocess_label(val_label, no_chain=True)\ntrain_label = preprocess_label(train_label)\nval_sequence = preprocess_sequence(val_sequence)\ntrain_sequence = preprocess_sequence(train_sequence)\ntest_sequence = preprocess_sequence(test_sequence)\nsubmission = pl.read_csv(\"/kaggle/input/stanford-rna-3d-folding/sample_submission.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:57:24.083966Z","iopub.execute_input":"2025-03-27T09:57:24.084282Z","iopub.status.idle":"2025-03-27T09:57:24.639297Z","shell.execute_reply.started":"2025-03-27T09:57:24.084257Z","shell.execute_reply":"2025-03-27T09:57:24.638331Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_label.write_parquet(metadata_dir / \"train_label.parquet\")\ntrain_sequence.write_parquet(metadata_dir / \"train_sequence.parquet\")\nval_label.write_parquet(metadata_dir / \"val_label.parquet\")\nval_sequence.write_parquet(metadata_dir / \"val_sequence.parquet\")\ntest_sequence.write_parquet(metadata_dir / \"test_sequence.parquet\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:57:36.820894Z","iopub.execute_input":"2025-03-27T09:57:36.821226Z","iopub.status.idle":"2025-03-27T09:57:36.913194Z","shell.execute_reply.started":"2025-03-27T09:57:36.821198Z","shell.execute_reply":"2025-03-27T09:57:36.912061Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Create MSA Metadata\n\nI processed multiple MSA files (Multiple Sequence Alignments) from `.fasta` format.\n\nMSA is a type of data that contains RNA sequences that are similar to the target sequence.  \nIt helps us find regions that are conserved (similar) across different RNA molecules.\n\nI combined these MSA files into a single metadata table and saved it in Parquet format for efficient use (thought not used in this EDA notebook).","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\n\nmsa_paths = sorted(Path(\"/kaggle/input/stanford-rna-3d-folding/MSA/\").rglob(\"*.fasta\"))\nmsa_names = [path.name for path in msa_paths]\nlen(msa_paths)\ncols = [\n    \"target_id\",\n    \"idx\",\n    \"id\",\n    \"name\",\n    \"description\",\n    \"seq\",\n]\n\nrecords = []\nfor path in tqdm(msa_paths):\n    for i, record in enumerate(SeqIO.parse(path, \"fasta\")):\n        record = vars(record)\n        record[\"target_id\"] = path.name.split(\".\")[0]\n        record[\"idx\"] = i\n        record[\"seq\"] = str(record[\"_seq\"])\n        new_record = {col: record[col] for col in cols if col in record}\n        records.append(new_record)\n\nmsa = pl.DataFrame(records)\ndisplay(msa)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:55:34.088568Z","iopub.execute_input":"2025-03-27T09:55:34.088973Z","iopub.status.idle":"2025-03-27T09:55:55.693181Z","shell.execute_reply.started":"2025-03-27T09:55:34.088938Z","shell.execute_reply":"2025-03-27T09:55:55.692328Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"msa.write_parquet(metadata_dir / \"msa.parquet\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:55:55.694559Z","iopub.execute_input":"2025-03-27T09:55:55.694956Z","iopub.status.idle":"2025-03-27T09:56:00.159945Z","shell.execute_reply.started":"2025-03-27T09:55:55.694918Z","shell.execute_reply":"2025-03-27T09:56:00.158802Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Plot: Histogram of Sequence Length\n\nThe chart below shows histograms of sequence lengths — that is, how many **bases** (A, U, G, C) are present in a single RNA chain.\n\nFrom the graph, we can see that most of the **training sequences** are relatively short, with lengths less than 100.  \nIn contrast, the **validation (and probably test) sequences** tend to be longer, typically ranging from **100 to 700** bases.\n\nNote that we only have **844 training samples** and **12 validation samples**, which is quite a small dataset for machine learning.\nThis makes the task more difficult and requires careful model design.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n_, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))\n\ntrain_sequence[\"sequence_length\"].to_pandas().hist(\n    bins=100, ax=ax1, label=\"sequence length\", color=\"blue\", alpha=0.5\n)\nval_sequence[\"sequence_length\"].to_pandas().hist(\n    bins=10, ax=ax2, label=\"sequence length\", color=\"orange\", alpha=0.5\n)\ntr_mean  = train_sequence[\"sequence_length\"].mean()\nva_mean = val_sequence[\"sequence_length\"].mean()\nax1.axvline(tr_mean, color=\"red\", linestyle=\"--\", label=f\"mean={tr_mean:.0f}\")\nax2.axvline(va_mean, color=\"red\", linestyle=\"--\", label=f\"mean={va_mean:.0f}\")\nax1.set(\n    xlabel=\"sequence length\",\n    ylabel=\"count\",\n    title=f\"train ({train_sequence.shape[0]} samples)\",\n)\nax1.legend()\nax2.set(\n    xlabel=\"sequence length\",\n    ylabel=\"count\",\n    title=f\"val ({val_sequence.shape[0]} samples)\",\n)\nax2.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:54:03.035560Z","iopub.execute_input":"2025-03-27T09:54:03.035932Z","iopub.status.idle":"2025-03-27T09:54:04.218970Z","shell.execute_reply.started":"2025-03-27T09:54:03.035899Z","shell.execute_reply":"2025-03-27T09:54:04.217960Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Save Chains in PDB Format\n\nTo make later steps easier, I split each chain into its own separate PDB file.","metadata":{}},{"cell_type":"code","source":"pdb_dir = Path(\"pdb\")\nif not pdb_dir.exists():\n    pdb_dir.mkdir(parents=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:56:09.514887Z","iopub.execute_input":"2025-03-27T09:56:09.515234Z","iopub.status.idle":"2025-03-27T09:56:09.520262Z","shell.execute_reply.started":"2025-03-27T09:56:09.515198Z","shell.execute_reply":"2025-03-27T09:56:09.519149Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tqdm import tqdm\n\n\ndef make_pdb(df, output_dir):\n    num_groups = df.group_by(\"pdb_id\", \"chain_id\", \"design_id\").len().shape[0]\n    for (pdb_id, chain_id, design_id), df in tqdm(\n        df.group_by(\"pdb_id\", \"chain_id\", \"design_id\", maintain_order=True),\n        total=num_groups,\n    ):\n        chain_id = chain_id if chain_id else \"A\"\n        file_path = output_dir / f\"{pdb_id}_{chain_id}_{design_id}.pdb\"\n        with open(file_path, \"w\") as f:\n            for i, row in enumerate(df.iter_rows(named=True)):\n                design_id = row[\"design_id\"]\n                atom_serial = i + 1\n                atom_name = \"C1\"\n                residue_name = row[\"resname\"]\n                residue_num = int(row[\"resid\"])\n                chain_id = row[\"chain_id\"] if row[\"chain_id\"] else \"A\"\n                x, y, z = row[\"x\"], row[\"y\"], row[\"z\"]\n\n                line = (\n                    f\"ATOM  {atom_serial:5d} {atom_name:<4s} {residue_name:>3s} {chain_id:1s}\"\n                    f\"{residue_num:4d}    {x:8.3f}{y:8.3f}{z:8.3f}  1.00  0.00           C\\n\"\n                )\n                f.write(line)\n            f.write(\"END\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:56:45.262532Z","iopub.execute_input":"2025-03-27T09:56:45.262931Z","iopub.status.idle":"2025-03-27T09:56:45.277622Z","shell.execute_reply.started":"2025-03-27T09:56:45.262896Z","shell.execute_reply":"2025-03-27T09:56:45.276562Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"make_pdb(val_label, pdb_dir)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:56:53.057059Z","iopub.execute_input":"2025-03-27T09:56:53.057434Z","iopub.status.idle":"2025-03-27T09:56:53.147948Z","shell.execute_reply.started":"2025-03-27T09:56:53.057400Z","shell.execute_reply":"2025-03-27T09:56:53.146842Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"make_pdb(train_label, pdb_dir)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T09:56:58.879837Z","iopub.execute_input":"2025-03-27T09:56:58.880215Z","iopub.status.idle":"2025-03-27T09:56:59.923595Z","shell.execute_reply.started":"2025-03-27T09:56:58.880182Z","shell.execute_reply":"2025-03-27T09:56:59.922623Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Visualize 3D Structures","metadata":{}},{"cell_type":"markdown","source":"Since the label data only contains the **positions of the `C1'` atoms**, we visualize them using **spheres**.\n\nThe `C1'` atom is part of the **sugar molecule (ribose)** in RNA. It's the specific carbon that **connects the sugar to a base** (A, U, G, or C).\n\nTo make it easier to understand at a glance, we **color each `C1'` atom according to its base** type.\nThe simplified diagram below shows where `C1'` is located in the sugar ring (ribose):\n\nHere, the base (A/U/G/C) is attached to the `C1'`, and the backbone of RNA is formed through connections between the phosphate (PO₄²⁻), C3', and C5' atoms.","metadata":{}},{"cell_type":"code","source":"import PIL\n\nPIL.Image.open(\"/kaggle/input/s3f-resources/RNA_chemical_structure.gif\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T11:52:21.252595Z","iopub.execute_input":"2025-03-27T11:52:21.252958Z","iopub.status.idle":"2025-03-27T11:52:21.266940Z","shell.execute_reply.started":"2025-03-27T11:52:21.252930Z","shell.execute_reply":"2025-03-27T11:52:21.265931Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\nBy Narayanese, CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=3481560","metadata":{}},{"cell_type":"code","source":"import py3Dmol\n\n\ndef load_pdb(pdb_file, hbondCutoff=4.0, width=800, height=600):\n    \"\"\"PDBファイルを読み込み、py3Dmolのビューを生成する\"\"\"\n    with open(pdb_file, \"r\") as f:\n        pdb_str = f.read()\n    view = py3Dmol.view(js=\"https://3dmol.org/build/3Dmol.js\", width=width, height=height)\n    view.addModel(pdb_str, \"pdb\", {\"hbondCutoff\": hbondCutoff})\n    return view\n\n\ndef apply_cartoon_style(view, filter_option, base_colors, radius):\n    \"\"\"cartoonスタイルの表示設定\"\"\"\n    if filter_option == \"C1\":\n        view.setStyle({\"cartoon\": {\"color\": \"blue\"}})\n        for base, color in base_colors.items():\n            view.setStyle({\"resn\": base}, {\"sphere\": {\"color\": color, \"radius\": radius}})\n    else:\n        view.setStyle({\"cartoon\": {\"color\": \"spectrum\"}})\n\n\ndef apply_chain_style(view, filter_option, chain, radius):\n    \"\"\"チェーン指定の場合の表示設定\"\"\"\n    if filter_option == \"C1\":\n        view.setStyle({\"chain\": chain}, {\"sphere\": {\"color\": \"blue\", \"radius\": radius}})\n    else:\n        view.setStyle({\"chain\": chain}, {\"cartoon\": {\"color\": \"blue\"}})\n\n\ndef apply_stick_style(view):\n    \"\"\"スティック（原子結合）スタイルの表示設定\"\"\"\n    view.setStyle({}, {\"stick\": {}})\n\n\ndef plot_pdb(pdb_file, style=\"cartoon\", filter_option=\"All\", color_mode=\"rainbow\",\n             chain=\"A\", hbondCutoff=4.0, radius=1.0):\n    \"\"\"\n    PDBファイルを表示する関数。\n    \n    Parameters:\n        pdb_file (str): PDBファイルのパス。\n        style (str): 表示スタイル。'cartoon' または 'stick' など。\n        filter_option (str): 表示フィルター（例：\"C1\"）。\n        color_mode (str): カラーモード。'rainbow' または 'chain'。\n        chain (str): チェーン指定。\n        hbondCutoff (float): 水素結合のカットオフ距離。\n        radius (float): 球表示の半径。\n    \"\"\"\n    # 基底ごとの色設定\n    base_colors = {\"A\": \"red\", \"U\": \"blue\", \"G\": \"green\", \"C\": \"orange\"}\n    \n    view = load_pdb(pdb_file, hbondCutoff)\n    \n    if style == \"cartoon\":\n        if color_mode == \"rainbow\":\n            apply_cartoon_style(view, filter_option, base_colors, radius)\n        elif color_mode == \"chain\":\n            apply_chain_style(view, filter_option, chain, radius)\n        else:\n            raise ValueError(\"Invalid color_mode. Use 'rainbow' or 'chain'.\")\n    elif style == \"stick\":\n        apply_stick_style(view)\n    else:\n        raise ValueError(\"Invalid style option. Use 'cartoon' or 'stick'.\")\n    \n    view.zoomTo()\n    return view","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T11:01:42.278893Z","iopub.execute_input":"2025-03-27T11:01:42.279225Z","iopub.status.idle":"2025-03-27T11:01:42.289883Z","shell.execute_reply.started":"2025-03-27T11:01:42.279200Z","shell.execute_reply":"2025-03-27T11:01:42.288585Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"When we try to predict 3D structures of RNA using **only the 1D sequence** (i.e. the base letters A, U, G, C), the result can be **ambiguous** — meaning that **multiple different 3D shapes** can fit the same sequence.\n\nTo handle this, we generate several possible structures for each RNA.　The field `design_id` indicates **which version (or variation) of the structure** it is.\n\nBelow is an example of different structure candidates for the same sequence (`R1156_A`).　At first glance, they may look similar — but if you look closely, you’ll notice subtle differences.","metadata":{}},{"cell_type":"code","source":"plot_pdb(\"pdb/R1156_A_1.pdb\", filter_option=\"C1\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T11:01:57.558606Z","iopub.execute_input":"2025-03-27T11:01:57.558976Z","iopub.status.idle":"2025-03-27T11:01:57.567397Z","shell.execute_reply.started":"2025-03-27T11:01:57.558944Z","shell.execute_reply":"2025-03-27T11:01:57.566416Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_pdb(\"pdb/R1156_A_2.pdb\", filter_option=\"C1\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T11:02:05.287856Z","iopub.execute_input":"2025-03-27T11:02:05.288227Z","iopub.status.idle":"2025-03-27T11:02:05.295644Z","shell.execute_reply.started":"2025-03-27T11:02:05.288194Z","shell.execute_reply":"2025-03-27T11:02:05.294798Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"## Example of a Full 3D Structure (Generated by RFdiffusion)\n\nThe image below shows a structure generated by **RFdiffusion** [1], a method that creates random 3D molecular structures under certain constraints such as length or symmetry.\n\nAlthough this example may represent a protein rather than RNA, it helps illustrate what a **full atom-level structure** might look like — in contrast to our dataset, which only contains the positions of C1' atoms.\n\nNote that this is a **synthetic, cartoon-style rendering** intended for visual understanding, not an actual experimental structure.\n\n- [1] https://github.com/RosettaCommons/RFdiffusion","metadata":{}},{"cell_type":"code","source":"plot_pdb(\"/kaggle/input/s3f-sample-pdb-data-generated-by-rfdiffusion/test_0.pdb\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T11:02:16.894027Z","iopub.execute_input":"2025-03-27T11:02:16.894346Z","iopub.status.idle":"2025-03-27T11:02:16.912517Z","shell.execute_reply.started":"2025-03-27T11:02:16.894319Z","shell.execute_reply":"2025-03-27T11:02:16.911504Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The image below shows the actual chemical bonds between atoms in the molecule. This is the most accurate representation of its structure.","metadata":{}},{"cell_type":"code","source":"plot_pdb(\"/kaggle/input/s3f-sample-pdb-data-generated-by-rfdiffusion/test_0.pdb\", style=\"stick\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-27T11:03:34.914006Z","iopub.execute_input":"2025-03-27T11:03:34.914397Z","iopub.status.idle":"2025-03-27T11:03:34.924671Z","shell.execute_reply.started":"2025-03-27T11:03:34.914362Z","shell.execute_reply":"2025-03-27T11:03:34.923563Z"}},"outputs":[],"execution_count":null}]}