{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"cell-c69b17","cell_type":"markdown","source":"# HMS EEG — EDA Overview (WORKSHOP VERSION, self-contained)\n\nThis notebook covers the data-understanding groundwork for the rest of the workshop:\n\n1. Dataset scale (patients, recordings, labeled windows)\n2. Expert consensus quality — many labels are ambiguous, not clean ground truth\n3. Class imbalance — in the full dataset, and how the workshop subset differs\n4. Why we split by `patient_id`, not by row\n5. Why the labeled window's exact position matters (`eeg_label_offset_seconds`)\n6. Raw signal amplitude — the actual numbers behind the mu-law calibration story\n7. Electrode layout — which electrodes bipolar montage uses, and which it discards\n8. A visual look at one EEG + spectrogram example per class\n\nFully self-contained — no external .py imports.","metadata":{}},{"id":"cell-09ad48","cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.ticker as mticker\n\nIS_KAGGLE = os.path.exists('/kaggle')\nprint(f\"Environment: {'Kaggle' if IS_KAGGLE else 'Local'}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-072702","cell_type":"code","source":"# ============ Config ============\nif IS_KAGGLE:\n    DATA_ROOT       = '/kaggle/input/competitions/hms-harmful-brain-activity-classification'\n    RAW_TRAIN_PATH  = os.path.join(DATA_ROOT, 'train.csv')\n    EEG_DIR         = os.path.join(DATA_ROOT, 'train_eegs')\n    SPEC_DIR        = os.path.join(DATA_ROOT, 'train_spectrograms')\n    SAMPLE_IDS_PATH = '/kaggle/input/datasets/xiaosufrankhu/midas-summer-academy-wk3-eeg/workshop_sample_ids.csv'\nelse:\n    DATA_ROOT       = os.path.abspath('../')\n    RAW_TRAIN_PATH  = os.path.abspath('../data_raw/train.csv')\n    EEG_DIR         = os.path.join(DATA_ROOT, 'train_eegs')\n    SPEC_DIR        = os.path.join(DATA_ROOT, 'train_spectrograms')\n    SAMPLE_IDS_PATH = os.path.abspath('../data_raw/workshop_sample_ids.csv')\n\nVOTE_COLS   = [\"seizure_vote\", \"lpd_vote\", \"gpd_vote\", \"lrda_vote\", \"grda_vote\", \"other_vote\"]\nLABEL_NAMES = [\"Seizure\", \"LPD\", \"GPD\", \"LRDA\", \"GRDA\", \"Other\"]\nCOLORS      = [\"#d62728\", \"#ff7f0e\", \"#bcbd22\", \"#2ca02c\", \"#17becf\", \"#7f7f7f\"]\n\nprint(f\"Competition data path exists: {os.path.exists(DATA_ROOT)}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-43b76e","cell_type":"markdown","source":"## 1 · Dataset scale (full HMS competition data)\n\nThese numbers describe the full training set that participants join the competition for — this is *not* the workshop's 600-row subset.","metadata":{}},{"id":"cell-b07092","cell_type":"code","source":"meta = pd.read_csv(RAW_TRAIN_PATH)\n\nsummary_rows = [\n    (\"Total labeled windows\",         len(meta)),\n    (\"Unique patients\",               meta[\"patient_id\"].nunique()),\n    (\"Unique EEG recordings\",         meta[\"eeg_id\"].nunique()),\n    (\"Unique spectrogram recordings\", meta[\"spectrogram_id\"].nunique()),\n    (\"Avg labeled windows / patient\",  round(len(meta) / meta[\"patient_id\"].nunique(), 1)),\n]\nsummary = pd.DataFrame(summary_rows, columns=[\"Metric\", \"Value\"])\nsummary","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-b7b86d","cell_type":"markdown","source":"## 2 · Expert consensus quality — not all labels are equally trustworthy\n\nEach labeled window was reviewed by up to ~20 experts, who each cast one vote across the 6 categories. A window where the top category gets, say, 18/20 votes is a **confident** label. A window where the top category only gets 3/20 votes is a much **more ambiguous** one — the \"ground truth\" itself is uncertain.","metadata":{}},{"id":"cell-bd02f0","cell_type":"code","source":"HIGH_THRESH = 10\n\nmax_vote   = meta[VOTE_COLS].max(axis=1)\ntotal_vote = meta[VOTE_COLS].sum(axis=1)\n\nhigh_mask = max_vote >= HIGH_THRESH   # matches the n_votes>=10 filter used in the paper pipeline\nlow_mask  = max_vote <  HIGH_THRESH\n\nn_total, n_high, n_low = len(meta), high_mask.sum(), low_mask.sum()\nprint(f\"Total labeled windows          : {n_total:>8,}  (100.0%)\")\nprint(f\"High consensus (max >= {HIGH_THRESH} votes): {n_high:>8,}  ({n_high/n_total:.1%})\")\nprint(f\"Low consensus  (max <  {HIGH_THRESH} votes): {n_low:>8,}  ({n_low/n_total:.1%})\")\n\nfig, ax = plt.subplots(figsize=(9, 4))\nbins = range(0, int(max_vote.max()) + 2)\nax.hist(max_vote, bins=bins, color=\"steelblue\", edgecolor=\"white\", alpha=0.85)\nax.axvline(HIGH_THRESH, color=\"orange\", lw=2, ls=\"--\", label=f\"Threshold = {HIGH_THRESH} (inclusive)\")\nax.set_xlabel(\"Max single-category vote count (per window)\")\nax.set_ylabel(\"Number of windows\")\nax.set_title(\"Expert agreement distribution — many windows are genuinely ambiguous\")\nax.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-1aa4bd","cell_type":"markdown","source":"## 3 · Class balance — the full dataset is heavily imbalanced\n\nThis is the imbalance that motivates 2-step training on the full dataset. The workshop subset (next section) is deliberately built to be balanced instead, so we can run fast, clean comparisons in the time we have.","metadata":{}},{"id":"cell-6fd025","cell_type":"code","source":"consensus_counts = meta[\"expert_consensus\"].value_counts().reindex(LABEL_NAMES)\n\nfig, ax = plt.subplots(figsize=(8, 4.5))\nbars = ax.bar(LABEL_NAMES, consensus_counts.values, color=COLORS, edgecolor=\"white\")\nax.bar_label(bars, padding=3, fontsize=10, fontweight=\"bold\")\nax.set_ylabel(\"Number of labeled windows\")\nax.set_title(f\"Full dataset class distribution — \"\n             f\"{consensus_counts.max()/consensus_counts.min():.1f}x imbalance \"\n             f\"({consensus_counts.idxmax()} vs {consensus_counts.idxmin()})\")\nplt.tight_layout()\nplt.show()\n\nprint(consensus_counts)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"f986a7cb-8cec-4923-b925-3aa3102cb69b","cell_type":"code","source":"# ============ Is the high-quality (>=10 vote) subset also imbalanced? ============\n# Reuses `high_mask` from the consensus-quality section above — no need to recompute.\nhigh_quality_counts = meta.loc[high_mask, \"expert_consensus\"].value_counts().reindex(LABEL_NAMES)\n\nfig, ax = plt.subplots(figsize=(8, 4.5))\nbars = ax.bar(LABEL_NAMES, high_quality_counts.values, color=COLORS, edgecolor=\"white\")\nax.bar_label(bars, padding=3, fontsize=10, fontweight=\"bold\")\nax.set_ylabel(\"Number of labeled windows\")\nax.set_title(f\"High-quality subset (max_vote >= {HIGH_THRESH}) class distribution — \"\n             f\"{high_quality_counts.max()/high_quality_counts.min():.1f}x imbalance \"\n             f\"({high_quality_counts.idxmax()} vs {high_quality_counts.idxmin()})\")\nplt.tight_layout()\nplt.show()\n\nprint(high_quality_counts)\nprint(f\"\\n(for comparison, full dataset imbalance was \"\n      f\"{consensus_counts.max()/consensus_counts.min():.1f}x)\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-ae9f1d","cell_type":"markdown","source":"## 4 · Workshop subset vs full dataset\n\nThe workshop uses `workshop_sample_ids.csv` — a stratified, patient-level sample: 100 rows/class for training, 20 rows/class for validation, with zero patient overlap between the two splits.","metadata":{}},{"id":"cell-6eab4a","cell_type":"code","source":"sample_ids = pd.read_csv(SAMPLE_IDS_PATH)\nworkshop_counts = sample_ids[\"expert_consensus\"].value_counts().reindex(LABEL_NAMES)\n\nfig, axes = plt.subplots(1, 2, figsize=(14, 4.5))\n\nbars0 = axes[0].bar(LABEL_NAMES, consensus_counts.values, color=COLORS, edgecolor=\"white\")\naxes[0].bar_label(bars0, padding=3, fontsize=9, fontweight=\"bold\")\naxes[0].set_title(f\"Full dataset (n={len(meta):,})\\nheavily imbalanced\")\naxes[0].set_ylabel(\"Number of windows\")\n\naxes[1].bar(LABEL_NAMES, workshop_counts.values, color=COLORS, edgecolor=\"white\")\naxes[1].bar_label(axes[1].containers[0], padding=3, fontsize=10, fontweight=\"bold\")\naxes[1].set_title(f\"Workshop subset (n={len(sample_ids):,})\\nbalanced by design\")\n\nplt.tight_layout()\nplt.show()\n\nprint(f\"Unique patients in workshop subset: {sample_ids['patient_id'].nunique()} \"\n      f\"(out of {meta['patient_id'].nunique():,} total)\")\ntrain_p = set(sample_ids[sample_ids['split']=='train']['patient_id'])\nval_p   = set(sample_ids[sample_ids['split']=='val']['patient_id'])\nprint(f\"Train/val patient overlap: {len(train_p & val_p)} (should be 0)\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-e7f482","cell_type":"markdown","source":"## 5 · Why we split by `patient_id`, not by row\n\nSome patients contribute many labeled windows. If we split randomly by row, the same patient's EEG could appear in both train and validation — the model could then \"recognize\" that patient's individual signal characteristics rather than learning to generalize. Splitting by patient closes this leak.","metadata":{}},{"id":"cell-5e6a42","cell_type":"code","source":"windows_per_patient = meta.groupby(\"patient_id\").size()\n\nfig, ax = plt.subplots(figsize=(9, 4))\nax.hist(windows_per_patient, bins=50, color=\"steelblue\", edgecolor=\"white\")\nax.set_xlabel(\"Labeled windows per patient\")\nax.set_ylabel(\"Number of patients\")\nax.set_title(f\"Some patients contribute far more windows than others \"\n             f\"(median={windows_per_patient.median():.0f}, max={windows_per_patient.max()})\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-914dd7","cell_type":"markdown","source":"## 6 · Where in the recording is the labeled window?\n\n`eeg_label_offset_seconds` tells you where the 50-second labeled segment sits within the full EEG recording. It is **not** always near the center of the file — using a fixed center-crop instead of this offset would frequently extract the wrong segment.","metadata":{}},{"id":"cell-f443c1","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(9, 4))\nax.hist(meta[\"eeg_label_offset_seconds\"], bins=60, color=\"steelblue\", edgecolor=\"white\")\nax.set_xlabel(\"eeg_label_offset_seconds\")\nax.set_ylabel(\"Number of labeled windows\")\nax.set_title(\"Labeled window position within the recording — not concentrated at the center\")\nplt.tight_layout()\nplt.show()\n\nprint(meta[\"eeg_label_offset_seconds\"].describe())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-fdf35a","cell_type":"markdown","source":"## 7 · Raw signal amplitude — the numbers behind the mu-law calibration\n\nMu-law encoding assumes its input is roughly within `[-1, 1]` before compression. Here's what the actual raw signal amplitude looks like (after bipolar montage), sampled from the workshop training set — this is exactly the distribution used to calibrate `MULAW_SCALE` in the 1D CNN notebook.","metadata":{}},{"id":"cell-a376d0","cell_type":"code","source":"CHANNELS = ['Fp1','F3','C3','P3','F7','T3','T5','O1','Fz','Cz','Pz',\n            'Fp2','F4','C4','P4','F8','T4','T6','O2']\n\ndef bipolar_montage(eeg_20ch: np.ndarray) -> np.ndarray:\n    # eeg_20ch: (T, 19) in the CHANNELS order above (EKG excluded)\n    idx = {c: i for i, c in enumerate(CHANNELS)}\n    return np.stack([\n        eeg_20ch[:, idx['Fp1']] - eeg_20ch[:, idx['T3']],\n        eeg_20ch[:, idx['T3']]  - eeg_20ch[:, idx['O1']],\n        eeg_20ch[:, idx['Fp1']] - eeg_20ch[:, idx['C3']],\n        eeg_20ch[:, idx['C3']]  - eeg_20ch[:, idx['O1']],\n        eeg_20ch[:, idx['Fp2']] - eeg_20ch[:, idx['C4']],\n        eeg_20ch[:, idx['C4']]  - eeg_20ch[:, idx['O2']],\n        eeg_20ch[:, idx['Fp2']] - eeg_20ch[:, idx['T4']],\n        eeg_20ch[:, idx['T4']]  - eeg_20ch[:, idx['O2']],\n    ], axis=1)  # (T, 8)\n\ntrain_eeg_ids = sample_ids[sample_ids['split'] == 'train']['eeg_id'].unique()[:30]\nabs_vals = []\nfor eeg_id in train_eeg_ids:\n    path = os.path.join(EEG_DIR, f\"{int(eeg_id)}.parquet\")\n    if not os.path.exists(path):\n        continue\n    eeg = pd.read_parquet(path, columns=CHANNELS).to_numpy(dtype=np.float32)\n    eeg = np.nan_to_num(eeg, nan=0.0, posinf=0.0, neginf=0.0)\n    bp  = bipolar_montage(eeg)\n    abs_vals.append(np.abs(bp).ravel())\nabs_vals = np.concatenate(abs_vals)\n\np50, p99 = np.percentile(abs_vals, [50, 99])\nprint(f\"Bipolar signal amplitude — median: {p50:.1f}, 99th percentile: {p99:.1f}, max: {abs_vals.max():.1f}\")\nprint(f\"(units are the raw EEG amplitude units in the parquet files, roughly microvolts)\")\n\nfig, ax = plt.subplots(figsize=(9, 4))\nax.hist(np.clip(abs_vals, 0, 500), bins=80, color=\"steelblue\", edgecolor=\"white\")\nax.axvline(p99, color=\"red\", ls=\"--\", label=f\"99th percentile = {p99:.0f}\")\nax.set_xlabel(\"|bipolar signal amplitude| (clipped at 500 for display)\")\nax.set_ylabel(\"Count\")\nax.set_title(\"This is the real scale mu-law needs calibrating against — not a fixed constant\")\nax.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-442d7e","cell_type":"markdown","source":"## 8 · Electrode layout — what bipolar montage keeps, and what it drops\n\nThe standard double-banana bipolar montage used in the 1D CNN notebook only differences 8 of the 19 EEG electrodes. The midline electrodes (Fz, Cz, Pz) and several others are not used at all — this matters for classes like LRDA vs GRDA, whose main distinguishing feature is *spatial extent* (lateralized vs generalized).","metadata":{}},{"id":"cell-41ccb5","cell_type":"code","source":"# Approximate 10-20 system positions (schematic, top-down view, nose at top)\nPOSITIONS = {\n    'Fp1': (-0.3, 0.9),  'Fp2': (0.3, 0.9),\n    'F7':  (-0.9, 0.5),  'F3':  (-0.5, 0.5), 'Fz': (0.0, 0.5), 'F4': (0.5, 0.5), 'F8': (0.9, 0.5),\n    'T3':  (-1.0, 0.0),  'C3':  (-0.5, 0.0), 'Cz': (0.0, 0.0), 'C4': (0.5, 0.0), 'T4': (1.0, 0.0),\n    'T5':  (-0.9, -0.5), 'P3':  (-0.5, -0.5), 'Pz': (0.0, -0.5), 'P4': (0.5, -0.5), 'T6': (0.9, -0.5),\n    'O1':  (-0.3, -0.9), 'O2':  (0.3, -0.9),\n}\nMONTAGE_ELECTRODES = {'Fp1', 'T3', 'O1', 'C3', 'Fp2', 'C4', 'T4', 'O2'}\n\nfig, ax = plt.subplots(figsize=(6.5, 6.5))\nhead = plt.Circle((0, 0), 1.05, fill=False, color='black', lw=1.5)\nax.add_patch(head)\nax.plot([0], [1.15], marker='^', color='black', markersize=10)  # nose marker\n\nfor ch, (x, y) in POSITIONS.items():\n    used = ch in MONTAGE_ELECTRODES\n    ax.scatter(x, y, s=500, color=('#d62728' if used else '#bbbbbb'),\n               edgecolor='black', zorder=3)\n    ax.text(x, y, ch, ha='center', va='center', fontsize=8, fontweight='bold',\n            color='white' if used else 'black', zorder=4)\n\nax.set_xlim(-1.3, 1.3); ax.set_ylim(-1.3, 1.3)\nax.set_aspect('equal'); ax.axis('off')\nax.set_title(\"Red = used by bipolar montage (8 electrodes)\\nGray = not used (11 electrodes, incl. Fz/Cz/Pz midline)\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-1cda9c","cell_type":"markdown","source":"## 9 · One EEG + spectrogram example per class\n\nA visual sanity check — one representative labeled window per consensus class, from the workshop subset.","metadata":{}},{"id":"cell-833b70","cell_type":"code","source":"# workshop_sample_ids.csv only carries IDs + label + split — merge back to the\n# full train.csv (already loaded as `meta`) to recover eeg_label_offset_seconds,\n# spectrogram_label_offset_seconds, and the vote columns.\nsample_ids_full = sample_ids.merge(\n    meta[[\"eeg_id\", \"eeg_sub_id\", \"eeg_label_offset_seconds\",\n          \"spectrogram_label_offset_seconds\"] + VOTE_COLS],\n    on=[\"eeg_id\", \"eeg_sub_id\"],\n    how=\"left\",\n)\nsample_ids_full[\"max_vote\"] = sample_ids_full[VOTE_COLS].max(axis=1)\n\n# Pick the HIGHEST-confidence (highest max_vote) example per class, not just the first\n# row encountered — section 2 above showed many labels are ambiguous, so an arbitrary\n# pick could easily land on a low-consensus window and undercut this section's own point.\nsample_table = (\n    sample_ids_full[sample_ids_full['split'] == 'train']\n    .sort_values('max_vote', ascending=False)\n    .groupby('expert_consensus')\n    .first()\n    .reindex(LABEL_NAMES)\n    .reset_index()\n)\nsample_table\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-92555a","cell_type":"code","source":"FS = 200\nWINDOW_SEC = 50\n\nfig, axes = plt.subplots(len(sample_table), 1, figsize=(16, 3.2 * len(sample_table)), sharex=False)\n\nfor ax, (_, row) in zip(axes, sample_table.iterrows()):\n    eeg_id = int(row['eeg_id'])\n    offset = float(row['eeg_label_offset_seconds'])\n    eeg = pd.read_parquet(os.path.join(EEG_DIR, f\"{eeg_id}.parquet\"), columns=CHANNELS).to_numpy(dtype=np.float32)\n    eeg = np.nan_to_num(eeg, nan=0.0)\n    time = np.arange(len(eeg)) / FS\n\n    SCALE = 150\n    for i, ch in enumerate(CHANNELS):\n        y_off = (len(CHANNELS) - 1 - i) * SCALE\n        sig = np.clip(eeg[:, i], -500, 500)\n        ax.plot(time, sig + y_off, lw=0.3, color='steelblue')\n\n    ax.axvspan(offset, offset + WINDOW_SEC, color='orange', alpha=0.15, label='Labeled window')\n    ax.set_xlim(max(0, offset - 30), min(time[-1], offset + WINDOW_SEC + 30))\n    ax.set_yticks([])\n    ax.set_title(f\"{row['expert_consensus']}  —  eeg_id={eeg_id}, offset={offset:.0f}s\")\n\naxes[0].legend(loc='upper right', fontsize=8)  # one legend is enough — same label on every subplot\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"cell-5b2c8a","cell_type":"code","source":"fig, axes = plt.subplots(1, len(sample_table), figsize=(4 * len(sample_table), 4))\n\nSPEC_WIDTH = 300  # 300 columns = 600 seconds at 2s/column — matches the 2D CNN notebook's SpectrogramDataset\n\nfor ax, (_, row) in zip(axes, sample_table.iterrows()):\n    spec_id = int(row['spectrogram_id'])\n    offset  = float(row['spectrogram_label_offset_seconds'])\n\n    spec_df = pd.read_parquet(os.path.join(SPEC_DIR, f\"{spec_id}.parquet\"))\n    spec_df = spec_df.drop(columns=['time'], errors='ignore').fillna(0)\n    arr = spec_df.to_numpy(dtype=np.float32).T   # (400, total_time) -> 4 chains x 100 freq bins\n\n    # crop to the labeled 600-second window, same as SpectrogramDataset in the 2D CNN notebook —\n    # otherwise this shows the *entire* recording, which can span multiple labeled windows\n    col_start = int(offset // 2)\n    window = arr[:, col_start:col_start + SPEC_WIDTH]\n    if window.shape[1] < SPEC_WIDTH:\n        window = np.pad(window, ((0, 0), (0, SPEC_WIDTH - window.shape[1])), mode='constant')\n\n    window = np.log1p(np.clip(window, 0, None))\n    ax.imshow(window, aspect='auto', origin='lower', cmap='viridis')\n    ax.set_title(f\"{row['expert_consensus']}\\noffset={offset:.0f}s\", fontsize=10)\n    ax.set_xticks([]); ax.set_yticks([])\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"b9c2ed16","cell_type":"markdown","source":"## Wrap-up — how this connects to the rest of the day, we are not addressing every point with todays hands-on\n\n- **Consensus quality (§2)** is why the paper's benchmark filters to `n_votes >= 10` before\n  training, and why 2-step training exists at all — Step 1 sees the full noisy data, Step 2\n  corrects on the clean subset.\n- **Patient-level split (§5)** is why every workshop/paper notebook splits by `patient_id`,\n  never by row — otherwise validation numbers would be inflated by leakage.\n- **Window position (§6)** is why every notebook uses `eeg_label_offset_seconds` /\n  `spectrogram_label_offset_seconds` to crop, instead of a fixed center-crop.\n- **Signal amplitude (§7)** is exactly what calibrates `MULAW_SCALE` in the 1D CNN notebook —\n  a fixed constant would saturate on the real amplitude range you just saw.\n- **Electrode layout (§8)** is why bipolar montage alone struggles with LRDA vs GRDA — both\n  are largely defined by *spatial extent* (lateralized vs generalized), and the dropped\n  midline electrodes (Fz/Cz/Pz) carry some of that information.\n\nNext up: XGBoost (hand-crafted features, no montage) → 1D CNN (naive vs domain-informed\npreprocessing) → 2D CNN (pretrained EfficientNet on spectrograms).\n","metadata":{}},{"id":"61754bef-b86c-48be-8a90-dce9d7dee7df","cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}