{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Lumbar Vibe 01\n\n## Overview\n\nLumbar Vibe is a research-oriented machine learning project built around the RSNA 2024 Lumbar Spine Degenerative Classification dataset.\n\nThe project's primary goal is to develop, document, and evaluate a complete lumbar MRI machine-learning workflow, beginning with raw DICOM images and ending with model training, inference, validation, and experiment tracking.\n\nThis notebook documents the first complete baseline pipeline developed for the project.\n\n## Objectives\n\n- Explore and understand the RSNA lumbar MRI dataset\n- Build reproducible study-selection and dataset-pruning workflows\n- Construct machine-learning training targets from radiology labels\n- Train and evaluate baseline neural-network models\n- Save and reload trained model artifacts\n- Generate reproducible validation metrics and experiment summaries\n- Establish a foundation for future lumbar MRI classification, localization, and case-study research experiments\n\n## Scope\n\nThis notebook focuses on engineering workflow validation rather than diagnostic performance.\n\nThe initial model is intentionally simple and is used to verify that the complete pipeline functions correctly:\n\n```text\nDICOM images\n→ preprocessing\n→ dataset construction\n→ model training\n→ saved weights\n→ inference\n→ validation metrics\n→ experiment tracking\n```\n\nSubsequent experiments will expand dataset size, training duration, model architecture, evaluation methodology, and target definitions.\n\nFuture notebooks will explore level-specific and side-specific lumbar degeneration targets, including spinal canal stenosis, neural foraminal narrowing, subarticular stenosis, and related lumbar MRI findings.\n\n## Important Disclaimer\n\nThis project is intended for educational, research, and portfolio purposes.\n\nNothing in this notebook should be interpreted as medical advice, medical diagnosis, or a substitute for professional evaluation by a qualified healthcare provider.\n\nAny future experiments involving personal MRI scans should be treated as independent research exercises and not as clinical diagnostic tools.\n\n## Current Milestone\n\nThis notebook represents the completion of the first 200-study baseline experiment:\n\n- Dataset subset construction complete\n- Binary spinal canal stenosis target complete\n- Baseline CNN training complete\n- Model serialization complete\n- Inference pipeline complete\n- Validation metrics complete\n- Experiment summary complete\n\nThis milestone establishes the baseline workflow upon which future experiments will be built.\n\nPlanned follow-on work includes:\n\n- LV 02: Extended spinal canal stenosis baseline development\n- LV 03: L5-S1-focused classification experiments\n- LV 04: Left-sided narrowing and stenosis experiments\n- LV 05: Personal MRI case-study inference workflow (research use only)","metadata":{}},{"cell_type":"markdown","source":"# Runtime & Workflow Instructions\n\n## Notebook Status\n\nThis notebook documents the completed 200-study baseline experiment for the Lumbar Vibe project.\n\nThe notebook supports two workflows:\n\n1. **Recommended Workflow (Run All)**\n2. **Manual Workflow (GPU Quota Conservation)**\n\nMost users should use the Recommended Workflow.\n\n---\n\n## Recommended Workflow (Run All)\n\nFor the simplest and most reliable experience:\n\n1. Open the notebook in Kaggle.\n2. Set **Accelerator = GPU** before running any cells.\n3. Select **Run All**.\n4. Allow the notebook to complete without changing runtime settings.\n5. Download desired outputs from `/kaggle/working`.\n6. Save a Kaggle notebook version.\n\nThis workflow minimizes the risk of runtime resets and missing files.\n\n### Recommended Workflow Summary\n\n```text\nAccelerator = GPU\n↓\nRun All\n↓\nDownload outputs\n↓\nSave notebook version\n```\n\n---\n\n## Manual Workflow (GPU Quota Conservation)\n\nUsers working within Kaggle's free GPU quota may prefer to separate the notebook into two phases.\n\n### Phase A — Dataset Construction\n\nRun:\n\n```text\nSteps 0–13\n```\n\nwith:\n\n```text\nAccelerator = None\n```\n\nPhase A creates the metadata tables, image indexes, labels, training targets, and train/validation split files required for model training.\n\nAfter completing Phase A, save important outputs locally.\n\n### Phase B — Training and Evaluation\n\nBefore beginning Phase B:\n\n1. Change Accelerator to **GPU**.\n2. Kaggle may restart the runtime.\n3. Files previously generated in `/kaggle/working` may disappear.\n4. Run **Step 13B: Rebuild Full Phase A State After Runtime Reset**.\n5. Verify that Step 13B completed successfully.\n6. Continue with:\n\n```text\nSteps 14–18\n```\n\nStep 13B is a required part of the manual workflow.\n\nWithout Step 13B, later training and evaluation steps may fail because required files are no longer present in `/kaggle/working`.\n\nDo not change Accelerator settings again until Phase B is complete.\n\n### Manual Workflow Summary\n\n```text\nPhase A\nAccelerator = None\nSteps 0–13\n↓\nSave important outputs locally\n↓\nAccelerator = GPU\n↓\nRuntime reset may occur\n↓\nRun Step 13B\n↓\nVerify rebuild completed\n↓\nSteps 14–18\n↓\nDownload outputs\n↓\nSave notebook version\n```\n\n---\n\n## Important Kaggle Runtime Behavior\n\nChanging Accelerator settings may restart the Kaggle runtime.\n\nWhen a runtime restart occurs:\n\n- Files stored in `/kaggle/working` may disappear.\n- Later cells may fail with `FileNotFoundError`.\n- Previously generated CSV files may need to be rebuilt.\n- Model weights and intermediate outputs may need to be regenerated.\n\nThis behavior is normal and should be expected.\n\n---\n\n## Step 13B Recovery Cell\n\nStep 13B exists specifically to recover from a Kaggle runtime reset.\n\nIn the recommended **Run All** workflow, Step 13B is skipped and is not needed.\n\nIn the **Manual Workflow**, Step 13B is required after switching from:\n\n```text\nAccelerator = None\n```\n\nto:\n\n```text\nAccelerator = GPU\n```\n\nWhen executed, Step 13B rebuilds the complete Phase A working state, including:\n\n```text\nstudy_metadata_summary.csv\nselected_studies_200.csv\nselected_series_200.csv\nselected_image_index_200.csv\nselected_labels_200.csv\nselected_label_distribution_200.csv\nproject_manifest_200.json\ntraining_index_200_spinal_canal_binary.csv\ntraining_index_200_spinal_canal_binary_split.csv\n```\n\nThis allows Phase B to begin from a clean runtime without manually re-running every earlier notebook step.\n\n### Manual users:\n\nIf you changed Accelerator from None to GPU and /kaggle/working was reset, set:\n\n       RUN_PHASE_A_REBUILD = True\n       \nThen run this cell before Step 14.\n\n---\n\n## About /kaggle/working\n\nMost intermediate outputs generated by this notebook are stored in:\n\n```text\n/kaggle/working\n```\n\nExamples include:\n\n- selected study tables\n- image indexes\n- training targets\n- validation outputs\n- experiment summaries\n- model weights\n\nFiles stored in this directory should be treated as temporary.\n\nDo not assume files stored in `/kaggle/working` will survive runtime changes.\n\n---\n\n## Local Backup Recommendations\n\nSave important outputs locally after major milestones.\n\n### After Phase A\n\nRecommended files:\n\n```text\nselected_studies_200.csv\nselected_series_200.csv\nselected_image_index_200.csv\nselected_labels_200.csv\n\ntraining_index_200_spinal_canal_binary.csv\ntraining_index_200_spinal_canal_binary_split.csv\n```\n\n### After Training\n\nRecommended files:\n\n```text\ntiny_lumbar_cnn_smoke_test.pth\n```\n\n### After Evaluation\n\nRecommended files:\n\n```text\nvalidation_metrics_200_smoke_test.csv\nvalidation_predictions_200_smoke_test.csv\n\nexperiment_summary_200_smoke_test.json\nexperiment_summary_200_smoke_test.csv\n```\n\n### After Milestone Completion\n\nRecommended files:\n\n```text\nmilestone_check_200_smoke_test.csv\n```\n\n---\n\n## Project Philosophy\n\nThis notebook is designed to be:\n\n- reproducible\n- beginner-friendly\n- transparent\n- restartable from a fresh Kaggle environment\n\nWhenever possible, notebook outputs should be reproducible from source data rather than dependent on manually uploaded intermediate files.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 0: SETUP\n# ============================================================\n# Purpose:\n# Define common imports, constants, and paths used throughout\n# the notebook.\n#\n# This cell does NOT load DICOM images.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Shared setup for both Phase A and Phase B.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\nimport os\nimport json\n\nimport numpy as np\nimport pandas as pd\n\n# Reproducibility settings.\nRANDOM_STATE = 42\nSUBSET_SIZE = 200\n\n# Kaggle dataset paths.\nDATASET_DIR = \"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification\"\nTRAIN_IMAGES_DIR = os.path.join(DATASET_DIR, \"train_images\")\nWORKING_DIR = \"/kaggle/working\"\n\n# Core MRI sequence types used in this baseline workflow.\nCORE_SEQUENCES = [\n    \"Axial T2\",\n    \"Sagittal T1\",\n    \"Sagittal T2/STIR\"\n]\n\nprint(\"Step 0 setup complete.\")\nprint(\"DATASET_DIR:\", DATASET_DIR)\nprint(\"TRAIN_IMAGES_DIR:\", TRAIN_IMAGES_DIR)\nprint(\"WORKING_DIR:\", WORKING_DIR)\nprint(\"SUBSET_SIZE:\", SUBSET_SIZE)\nprint(\"RANDOM_STATE:\", RANDOM_STATE)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 1: DATASET INSPECTION\n# ============================================================\n# Purpose:\n# Inspect the attached RSNA lumbar spine dataset and confirm that\n# the expected metadata files and image folders are available.\n#\n# This cell checks:\n#   1. Which datasets are attached to the Kaggle notebook\n#   2. Where the RSNA lumbar dataset is mounted\n#   3. Which top-level files/folders exist\n#   4. What the main CSV metadata files contain\n#   5. How the train_images folder is organized\n#\n# This cell does NOT train a model.\n# This cell does NOT build the selected subset.\n# This cell avoids recursively scanning every DICOM file.\n#\n# Workflow phase:\n# Phase A — Dataset construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Kaggle attaches datasets under the read-only /kaggle/input directory.\nINPUT_DIR = \"/kaggle/input\"\n\nprint(\"=== Attached Kaggle input datasets ===\")\nattached_datasets = os.listdir(INPUT_DIR)\nprint(attached_datasets)\n\n# Attempt to automatically locate the RSNA lumbar dataset.\n# This avoids hard-coding the dataset folder name.\ncandidate_dirs = [\n    name for name in attached_datasets\n    if \"rsna\" in name.lower() and \"lumbar\" in name.lower()\n]\n\nif len(candidate_dirs) == 0:\n    raise FileNotFoundError(\n        \"Could not automatically find an attached RSNA lumbar dataset. \"\n        \"Check the Kaggle notebook's right sidebar and add the dataset as an input.\"\n    )\n\n# Use the first matching dataset folder.\nDATASET_NAME = candidate_dirs[0]\nDATASET_DIR = os.path.join(INPUT_DIR, DATASET_NAME)\nTRAIN_IMAGES_DIR = os.path.join(DATASET_DIR, \"train_images\")\n\nprint(\"\\n=== Selected dataset folder ===\")\nprint(DATASET_DIR)\n\nprint(\"\\n=== Top-level dataset contents ===\")\nprint(os.listdir(DATASET_DIR))\n\n# Verify that the expected image directory exists.\n# Later steps depend on this folder structure.\nif not os.path.exists(TRAIN_IMAGES_DIR):\n    raise FileNotFoundError(\n        f\"Expected image folder not found: {TRAIN_IMAGES_DIR}\"\n    )\n\n# Load the three primary metadata tables supplied by RSNA.\n# These describe:\n#   - study labels\n#   - MRI series information\n#   - coordinate annotations\ntrain_df = pd.read_csv(os.path.join(DATASET_DIR, \"train.csv\"))\nseries_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_series_descriptions.csv\"))\ncoords_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_label_coordinates.csv\"))\n\nprint(\"\\n=== CSV shapes ===\")\nprint(\"train.csv shape:\", train_df.shape)\nprint(\"train_series_descriptions.csv shape:\", series_df.shape)\nprint(\"train_label_coordinates.csv shape:\", coords_df.shape)\n\n# Preview the metadata tables so we understand the structure\n# before constructing derived datasets later.\nprint(\"\\n=== train.csv preview ===\")\ndisplay(train_df.head())\n\nprint(\"\\n=== train_series_descriptions.csv preview ===\")\ndisplay(series_df.head())\n\nprint(\"\\n=== train_label_coordinates.csv preview ===\")\ndisplay(coords_df.head())\n\n# Count the available MRI sequence types.\n# These become important in later pruning and selection steps.\nprint(\"\\n=== Series description counts ===\")\ndisplay(series_df[\"series_description\"].value_counts())\n\n# Inspect a few study folders without scanning the full dataset.\n# This confirms the expected directory structure.\nprint(\"\\n=== First 10 study folders ===\")\nstudy_folders = sorted(os.listdir(TRAIN_IMAGES_DIR))[:10]\nprint(study_folders)\n\nprint(\"\\n=== Example study → series folder structure ===\")\n\nfor study_id in study_folders[:3]:\n    study_path = os.path.join(TRAIN_IMAGES_DIR, study_id)\n\n    if os.path.isdir(study_path):\n\n        # Each study contains one or more MRI series.\n        series_folders = sorted(os.listdir(study_path))\n\n        print(f\"\\nStudy {study_id}:\")\n        print(series_folders[:10])\n\n        if len(series_folders) > 0:\n            first_series = series_folders[0]\n            first_series_path = os.path.join(study_path, first_series)\n\n            if os.path.isdir(first_series_path):\n\n                # Show a few example DICOM filenames.\n                # This verifies the expected file naming convention.\n                example_files = sorted(os.listdir(first_series_path))[:10]\n\n                print(f\"  Example files in series {first_series}:\")\n                print(example_files)\n\n# Summarize how many MRI series exist per study.\n# This provides a quick picture of dataset complexity.\nseries_per_study = series_df.groupby(\"study_id\")[\"series_id\"].count()\n\nprint(\"\\n=== Series per study summary ===\")\ndisplay(series_per_study.describe())\n\nprint(\"\\n=== Done: dataset inspection complete ===\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-06-15T22:13:28.691547Z","iopub.execute_input":"2026-06-15T22:13:28.691770Z","iopub.status.idle":"2026-06-15T22:15:01.910889Z","shell.execute_reply.started":"2026-06-15T22:13:28.691739Z","shell.execute_reply":"2026-06-15T22:15:01.910165Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 2: BUILD STUDY-LEVEL METADATA TABLE\n# ============================================================\n# Purpose:\n# Build a clean study-level summary table from the RSNA metadata.\n#\n# This cell summarizes:\n#   1. How many MRI series each study contains\n#   2. Which core sequence types are present per study\n#   3. Which studies contain all core sequences needed for this baseline\n#\n# This cell does NOT load DICOM images.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load metadata CSVs.\ntrain_df = pd.read_csv(os.path.join(DATASET_DIR, \"train.csv\"))\nseries_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_series_descriptions.csv\"))\ncoords_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_label_coordinates.csv\"))\n\nprint(\"=== Loaded CSVs ===\")\nprint(\"train_df:\", train_df.shape)\nprint(\"series_df:\", series_df.shape)\nprint(\"coords_df:\", coords_df.shape)\n\n# Count how many MRI series each study has.\nseries_count_df = (\n    series_df\n    .groupby(\"study_id\")\n    .agg(\n        num_series=(\"series_id\", \"count\"),\n        series_descriptions=(\"series_description\", lambda x: sorted(list(x)))\n    )\n    .reset_index()\n)\n\n# Build one boolean column per MRI sequence type.\n# A value of True means that study contains that sequence.\nsequence_presence_df = (\n    series_df\n    .assign(value=True)\n    .pivot_table(\n        index=\"study_id\",\n        columns=\"series_description\",\n        values=\"value\",\n        aggfunc=\"any\",\n        fill_value=False\n    )\n    .reset_index()\n)\n\n# Merge sequence counts and sequence presence into one study-level table.\nstudy_meta_df = series_count_df.merge(\n    sequence_presence_df,\n    on=\"study_id\",\n    how=\"left\"\n)\n\n# Identify studies that have all core sequence types used by this baseline.\nstudy_meta_df[\"has_all_core_sequences\"] = study_meta_df[CORE_SEQUENCES].all(axis=1)\n\nprint(\"\\n=== Study metadata preview ===\")\ndisplay(study_meta_df.head())\n\nprint(\"\\n=== Study metadata shape ===\")\nprint(study_meta_df.shape)\n\nprint(\"\\n=== Sequence completeness counts ===\")\nfor sequence_name in CORE_SEQUENCES:\n    if sequence_name in study_meta_df.columns:\n        print(sequence_name, \":\", study_meta_df[sequence_name].sum())\n\nprint(\"\\n=== Studies with all core sequences ===\")\nprint(study_meta_df[\"has_all_core_sequences\"].sum())\n\nprint(\"\\n=== Distribution of number of series per study ===\")\ndisplay(study_meta_df[\"num_series\"].value_counts().sort_index())\n\n# Save this lightweight metadata table to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"study_metadata_summary.csv\")\nstudy_meta_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved study metadata summary to:\")\nprint(output_path)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-15T23:26:43.756670Z","iopub.execute_input":"2026-06-15T23:26:43.756943Z","iopub.status.idle":"2026-06-15T23:26:45.522473Z","shell.execute_reply.started":"2026-06-15T23:26:43.756915Z","shell.execute_reply":"2026-06-15T23:26:45.521469Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 3: CREATE FIRST PRUNED STUDY SUBSET\n# ============================================================\n# Purpose:\n# Select a manageable first subset of complete lumbar MRI studies\n# for baseline model development.\n#\n# This cell:\n#   1. Identifies studies with all core MRI sequences\n#   2. Randomly samples a reproducible 200-study subset\n#   3. Saves the selected study IDs for later steps\n#\n# This cell does NOT copy DICOM files.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load metadata needed to identify complete studies.\nseries_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_series_descriptions.csv\"))\n\n# Build a study-level sequence completeness table.\n# Each sequence column indicates whether that study contains the sequence.\nsequence_presence_df = (\n    series_df\n    .assign(value=True)\n    .pivot_table(\n        index=\"study_id\",\n        columns=\"series_description\",\n        values=\"value\",\n        aggfunc=\"any\",\n        fill_value=False\n    )\n    .reset_index()\n)\n\n# Identify studies that contain all core sequence types used by this baseline.\nsequence_presence_df[\"has_all_core_sequences\"] = (\n    sequence_presence_df[CORE_SEQUENCES].all(axis=1)\n)\n\n# Keep only studies with all required MRI sequence types.\neligible_studies = sequence_presence_df[\n    sequence_presence_df[\"has_all_core_sequences\"] == True\n].copy()\n\nprint(\"Eligible studies with all core sequences:\", len(eligible_studies))\n\n# Randomly sample the first experimental subset.\n# RANDOM_STATE keeps the sample reproducible across reruns.\nselected_studies = eligible_studies.sample(\n    n=SUBSET_SIZE,\n    random_state=RANDOM_STATE\n).copy()\n\n# Keep only columns useful for downstream indexing and documentation.\nselected_studies = selected_studies[\n    [\"study_id\", \"Axial T2\", \"Sagittal T1\", \"Sagittal T2/STIR\", \"has_all_core_sequences\"]\n].sort_values(\"study_id\")\n\n# Save selected study IDs to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"selected_studies_200.csv\")\nselected_studies.to_csv(output_path, index=False)\n\nprint(\"Selected study subset shape:\", selected_studies.shape)\nprint(\"Saved to:\", output_path)\n\ndisplay(selected_studies.head(10))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:43:38.220062Z","iopub.execute_input":"2026-06-16T00:43:38.220235Z","iopub.status.idle":"2026-06-16T00:43:39.747933Z","shell.execute_reply.started":"2026-06-16T00:43:38.220212Z","shell.execute_reply":"2026-06-16T00:43:39.747142Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 4: CREATE SELECTED SERIES TABLE\n# ============================================================\n# Purpose:\n# Link the selected 200 studies to their MRI series IDs.\n#\n# This cell identifies which series_id belongs to each core sequence:\n#   1. Axial T2\n#   2. Sagittal T1\n#   3. Sagittal T2/STIR\n#\n# This cell does NOT load DICOM images.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load selected studies from the previous step.\nselected_studies_path = os.path.join(WORKING_DIR, \"selected_studies_200.csv\")\nselected_studies_df = pd.read_csv(selected_studies_path)\n\n# Load full series metadata from the RSNA dataset.\nseries_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_series_descriptions.csv\"))\n\n# Keep only MRI series belonging to the selected 200 studies.\nselected_series_df = series_df[\n    series_df[\"study_id\"].isin(selected_studies_df[\"study_id\"])\n].copy()\n\n# Sort for readability and stable output ordering.\nselected_series_df = selected_series_df.sort_values(\n    [\"study_id\", \"series_description\", \"series_id\"]\n).reset_index(drop=True)\n\nprint(\"Selected studies:\", selected_studies_df.shape)\nprint(\"Selected series:\", selected_series_df.shape)\n\nprint(\"\\n=== Series description counts in selected subset ===\")\ndisplay(selected_series_df[\"series_description\"].value_counts())\n\nprint(\"\\n=== Preview of selected series table ===\")\ndisplay(selected_series_df.head(20))\n\n# Save selected series table to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"selected_series_200.csv\")\nselected_series_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved selected series table to:\")\nprint(output_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:43:54.445520Z","iopub.execute_input":"2026-06-16T00:43:54.445825Z","iopub.status.idle":"2026-06-16T00:43:54.478486Z","shell.execute_reply.started":"2026-06-16T00:43:54.445805Z","shell.execute_reply":"2026-06-16T00:43:54.477589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 5: CREATE SELECTED DICOM IMAGE INDEX\n# ============================================================\n# Purpose:\n# Build a table of DICOM image file paths for the selected\n# 200-study subset.\n#\n# This cell links:\n#   selected_series_200.csv\n#   → actual DICOM slice paths on disk\n#\n# Output:\n#   selected_image_index_200.csv\n#\n# Each row in the output table represents one DICOM slice/image.\n#\n# This cell does NOT load DICOM pixel data.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load selected series table from the previous step.\nselected_series_path = os.path.join(WORKING_DIR, \"selected_series_200.csv\")\nselected_series_df = pd.read_csv(selected_series_path)\n\nrows = []\n\n# Loop over each selected MRI series and locate its DICOM files.\nfor _, row in selected_series_df.iterrows():\n    study_id = str(row[\"study_id\"])\n    series_id = str(row[\"series_id\"])\n    series_description = row[\"series_description\"]\n\n    # Expected folder pattern:\n    # train_images/<study_id>/<series_id>/\n    series_dir = os.path.join(TRAIN_IMAGES_DIR, study_id, series_id)\n\n    if not os.path.exists(series_dir):\n        print(\"Missing series folder:\", series_dir)\n        continue\n\n    # List DICOM files in this series.\n    dicom_files = [\n        filename for filename in os.listdir(series_dir)\n        if filename.lower().endswith(\".dcm\")\n    ]\n\n    # Sort by numeric instance number:\n    # \"1.dcm\", \"2.dcm\", ..., \"10.dcm\"\n    dicom_files = sorted(\n        dicom_files,\n        key=lambda filename: int(os.path.splitext(filename)[0])\n    )\n\n    # Add one row per DICOM slice.\n    for instance_index, filename in enumerate(dicom_files):\n        rows.append({\n            \"study_id\": int(study_id),\n            \"series_id\": int(series_id),\n            \"series_description\": series_description,\n            \"instance_index\": instance_index,\n            \"filename\": filename,\n            \"dcm_path\": os.path.join(series_dir, filename)\n        })\n\n# Convert collected rows into a DataFrame.\nimage_index_df = pd.DataFrame(rows)\n\nif image_index_df.empty:\n    raise ValueError(\n        \"No DICOM image paths were indexed. \"\n        \"Check selected_series_200.csv and TRAIN_IMAGES_DIR.\"\n    )\n\nprint(\"Selected series:\", selected_series_df.shape)\nprint(\"Selected DICOM image index:\", image_index_df.shape)\n\nprint(\"\\n=== Image count by series description ===\")\ndisplay(image_index_df[\"series_description\"].value_counts())\n\nprint(\"\\n=== Preview ===\")\ndisplay(image_index_df.head(20))\n\n# Save image index to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"selected_image_index_200.csv\")\nimage_index_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved selected image index to:\")\nprint(output_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:44:04.194391Z","iopub.execute_input":"2026-06-16T00:44:04.194684Z","iopub.status.idle":"2026-06-16T00:44:13.651209Z","shell.execute_reply.started":"2026-06-16T00:44:04.194660Z","shell.execute_reply":"2026-06-16T00:44:13.650108Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 6: CREATE SELECTED LABEL TABLE\n# ============================================================\n# Purpose:\n# Create a label table for only the selected 200 studies.\n#\n# This cell links:\n#   selected_studies_200.csv\n#   → RSNA study-level label columns\n#\n# Output:\n#   selected_labels_200.csv\n#\n# This cell does NOT load DICOM images.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load selected studies from Step 3.\nselected_studies_path = os.path.join(WORKING_DIR, \"selected_studies_200.csv\")\nselected_studies_df = pd.read_csv(selected_studies_path)\n\n# Load full RSNA label table.\ntrain_df = pd.read_csv(os.path.join(DATASET_DIR, \"train.csv\"))\n\n# Keep labels only for the selected 200 studies.\nselected_labels_df = train_df[\n    train_df[\"study_id\"].isin(selected_studies_df[\"study_id\"])\n].copy()\n\nif selected_labels_df.empty:\n    raise ValueError(\n        \"No labels found for selected studies. \"\n        \"Check selected_studies_200.csv and train.csv.\"\n    )\n\nprint(\"Full train labels:\", train_df.shape)\nprint(\"Selected labels:\", selected_labels_df.shape)\n\nprint(\"\\n=== Preview of selected labels ===\")\ndisplay(selected_labels_df.head())\n\nprint(\"\\n=== Label columns ===\")\nfor col in selected_labels_df.columns:\n    print(col)\n\n# Save selected labels to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"selected_labels_200.csv\")\nselected_labels_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved selected labels table to:\")\nprint(output_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:44:27.792004Z","iopub.execute_input":"2026-06-16T00:44:27.792295Z","iopub.status.idle":"2026-06-16T00:44:27.830414Z","shell.execute_reply.started":"2026-06-16T00:44:27.792276Z","shell.execute_reply":"2026-06-16T00:44:27.829275Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 7: VISUAL SANITY CHECK OF ONE DICOM IMAGE\n# ============================================================\n# Purpose:\n# Load one DICOM image from the selected 200-study subset and display it.\n#\n# This cell confirms:\n#   1. selected_image_index_200.csv paths are valid\n#   2. pydicom can read the selected DICOM file\n#   3. pixel data displays correctly as a grayscale MRI slice\n#\n# This cell does NOT train a model.\n# This cell does NOT create a model input tensor.\n#\n# Workflow phase:\n# Phase A — Dataset construction and visual quality check.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Specialty imports used for DICOM loading and visualization.\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# Load selected image index from Step 5.\nimage_index_path = os.path.join(WORKING_DIR, \"selected_image_index_200.csv\")\nimage_index_df = pd.read_csv(image_index_path)\n\nprint(\"Selected image index shape:\", image_index_df.shape)\n\n# Prefer a Sagittal T2/STIR image for the first visual check.\n# This sequence is commonly useful for viewing spinal anatomy and disc signal.\nsample_candidates = image_index_df[\n    image_index_df[\"series_description\"] == \"Sagittal T2/STIR\"\n]\n\nif sample_candidates.empty:\n    raise ValueError(\n        \"No Sagittal T2/STIR images found in selected_image_index_200.csv.\"\n    )\n\nsample_row = sample_candidates.iloc[0]\ndcm_path = sample_row[\"dcm_path\"]\n\nprint(\"Sample DICOM path:\")\nprint(dcm_path)\n\nprint(\"\\nSample metadata row:\")\ndisplay(sample_row.to_frame().T)\n\n# Read the selected DICOM file.\nds = pydicom.dcmread(dcm_path)\n\n# Extract the raw pixel array for display.\nimg = ds.pixel_array\n\nprint(\"\\nImage shape:\", img.shape)\nprint(\"Pixel dtype:\", img.dtype)\nprint(\"Pixel min/max:\", img.min(), img.max())\n\n# Display the image as a grayscale MRI slice.\nplt.figure(figsize=(6, 6))\nplt.imshow(img, cmap=\"gray\")\nplt.title(\n    f'{sample_row[\"series_description\"]}\\n'\n    f'Study {sample_row[\"study_id\"]}, Series {sample_row[\"series_id\"]}, File {sample_row[\"filename\"]}'\n)\nplt.axis(\"off\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:15:50.135808Z","iopub.execute_input":"2026-06-16T00:15:50.136146Z","iopub.status.idle":"2026-06-16T00:15:51.383701Z","shell.execute_reply.started":"2026-06-16T00:15:50.136118Z","shell.execute_reply":"2026-06-16T00:15:51.382636Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 8: MULTI-SEQUENCE VISUAL SANITY CHECK\n# ============================================================\n# Purpose:\n# Display one representative slice from each core MRI sequence\n# for the same selected study.\n#\n# This cell checks:\n#   1. The selected study contains all core sequences\n#   2. image paths are valid across multiple series\n#   3. Axial T2, Sagittal T1, and Sagittal T2/STIR appear visually distinct\n#\n# This cell does NOT train a model.\n# This cell does NOT create model inputs.\n#\n# Workflow phase:\n# Phase A — Dataset construction and visual quality check.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load selected image index from Step 5.\nimage_index_path = os.path.join(WORKING_DIR, \"selected_image_index_200.csv\")\nimage_index_df = pd.read_csv(image_index_path)\n\nif image_index_df.empty:\n    raise ValueError(\n        \"selected_image_index_200.csv is empty. \"\n        \"Run Step 5 before this visual sanity check.\"\n    )\n\n# Pick the first selected study for a multi-sequence visual check.\nstudy_id = image_index_df[\"study_id\"].iloc[0]\n\nprint(\"Selected study_id:\", study_id)\n\n# Filter image index to this one study.\nstudy_images_df = image_index_df[\n    image_index_df[\"study_id\"] == study_id\n].copy()\n\nprint(\"\\nSeries available for this study:\")\ndisplay(study_images_df[[\"series_id\", \"series_description\"]].drop_duplicates())\n\nplt.figure(figsize=(15, 5))\n\nfor i, sequence_name in enumerate(CORE_SEQUENCES, start=1):\n\n    # Get all images for this sequence within the selected study.\n    sequence_df = study_images_df[\n        study_images_df[\"series_description\"] == sequence_name\n    ].copy()\n\n    if sequence_df.empty:\n        print(f\"Missing sequence for this study: {sequence_name}\")\n        continue\n\n    # Pick the middle slice from the sequence.\n    # The middle slice is usually more informative than the first or last slice.\n    middle_index = len(sequence_df) // 2\n    sample_row = sequence_df.iloc[middle_index]\n\n    # Read DICOM pixel data for display.\n    ds = pydicom.dcmread(sample_row[\"dcm_path\"])\n    img = ds.pixel_array\n\n    # Plot one representative image for this sequence.\n    plt.subplot(1, len(CORE_SEQUENCES), i)\n    plt.imshow(img, cmap=\"gray\")\n    plt.title(\n        f\"{sequence_name}\\n\"\n        f\"Series {sample_row['series_id']} | File {sample_row['filename']}\"\n    )\n    plt.axis(\"off\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:19:19.335686Z","iopub.execute_input":"2026-06-16T00:19:19.336063Z","iopub.status.idle":"2026-06-16T00:19:20.016250Z","shell.execute_reply.started":"2026-06-16T00:19:19.336034Z","shell.execute_reply":"2026-06-16T00:19:20.013828Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 9: LABEL DISTRIBUTION CHECK\n# ============================================================\n# Purpose:\n# Inspect the label distribution in the selected 200-study subset.\n#\n# This cell summarizes:\n#   1. Overall label frequency across all condition/level columns\n#   2. Overall label percentages\n#   3. Missing label counts\n#   4. Per-label-column distributions\n#\n# This helps determine whether the subset is reasonable for\n# early baseline model experiments.\n#\n# This cell does NOT load DICOM images.\n# This cell does NOT train or evaluate a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction and label review.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load selected labels from Step 6.\nselected_labels_path = os.path.join(WORKING_DIR, \"selected_labels_200.csv\")\nselected_labels_df = pd.read_csv(selected_labels_path)\n\nprint(\"Selected labels shape:\", selected_labels_df.shape)\n\n# All columns except study_id are medical label columns.\nlabel_columns = [col for col in selected_labels_df.columns if col != \"study_id\"]\n\nprint(\"\\nNumber of label columns:\", len(label_columns))\n\n# Count label values across all condition/level columns.\n# This gives a broad view of how common each severity label is.\nall_label_values = selected_labels_df[label_columns].stack()\n\nprint(\"\\n=== Overall label distribution ===\")\ndisplay(all_label_values.value_counts(dropna=False))\n\nprint(\"\\n=== Overall label distribution as percent ===\")\ndisplay((all_label_values.value_counts(normalize=True, dropna=False) * 100).round(2))\n\n# Count missing labels across all selected label columns.\nmissing_count = selected_labels_df[label_columns].isna().sum().sum()\nprint(\"\\nTotal missing label cells:\", missing_count)\n\n# Build a per-column label distribution table.\n# Each row corresponds to one condition/level label column.\nlabel_distribution_rows = []\n\nfor col in label_columns:\n    counts = selected_labels_df[col].value_counts(dropna=False)\n\n    row = {\"label_column\": col}\n\n    for label_value, count in counts.items():\n        row[str(label_value)] = count\n\n    label_distribution_rows.append(row)\n\nlabel_distribution_df = pd.DataFrame(label_distribution_rows).fillna(0)\n\nprint(\"\\n=== Per-label-column distribution preview ===\")\ndisplay(label_distribution_df.head(20))\n\n# Save label distribution summary to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"selected_label_distribution_200.csv\")\nlabel_distribution_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved label distribution summary to:\")\nprint(output_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:20:23.926217Z","iopub.execute_input":"2026-06-16T00:20:23.926561Z","iopub.status.idle":"2026-06-16T00:20:23.979939Z","shell.execute_reply.started":"2026-06-16T00:20:23.926530Z","shell.execute_reply":"2026-06-16T00:20:23.979170Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 10: CREATE PROJECT MANIFEST\n# ============================================================\n# Purpose:\n# Create a project manifest describing the current experiment.\n#\n# This cell records:\n#   1. Dataset information\n#   2. Subset selection details\n#   3. Core MRI sequences used\n#   4. Generated project files\n#   5. Experiment metadata needed for reproducibility\n#\n# This helps future users understand exactly how the\n# current dataset subset was created.\n#\n# This cell does NOT load DICOM pixel data.\n# This cell does NOT train or evaluate a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction and experiment documentation.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load outputs created in earlier Phase A steps.\nselected_studies_df = pd.read_csv(\n    os.path.join(WORKING_DIR, \"selected_studies_200.csv\")\n)\n\nselected_series_df = pd.read_csv(\n    os.path.join(WORKING_DIR, \"selected_series_200.csv\")\n)\n\nselected_image_index_df = pd.read_csv(\n    os.path.join(WORKING_DIR, \"selected_image_index_200.csv\")\n)\n\nselected_labels_df = pd.read_csv(\n    os.path.join(WORKING_DIR, \"selected_labels_200.csv\")\n)\n\n# Build a manifest describing the current experiment state.\nmanifest = {\n    \"project_name\": \"lumbar_vibe\",\n    \"dataset\": \"RSNA 2024 Lumbar Spine Degenerative Classification\",\n    \"dataset_root\": DATASET_DIR,\n    \"subset_name\": f\"selected_{SUBSET_SIZE}_studies\",\n    \"subset_size_studies\": int(len(selected_studies_df)),\n    \"num_selected_series\": int(len(selected_series_df)),\n    \"num_selected_dicom_images\": int(len(selected_image_index_df)),\n    \"num_label_rows\": int(len(selected_labels_df)),\n    \"core_sequences\": CORE_SEQUENCES,\n    \"selection_method\": {\n        \"eligible_rule\": (\n            \"study must contain Axial T2, \"\n            \"Sagittal T1, and Sagittal T2/STIR\"\n        ),\n        \"sampling_method\": \"random sample\",\n        \"random_state\": RANDOM_STATE\n    },\n    \"created_files\": [\n        \"study_metadata_summary.csv\",\n        \"selected_studies_200.csv\",\n        \"selected_series_200.csv\",\n        \"selected_image_index_200.csv\",\n        \"selected_labels_200.csv\",\n        \"selected_label_distribution_200.csv\"\n    ],\n    \"medical_use_note\": (\n        \"This project is for educational and research workflow \"\n        \"development only. It is not a diagnostic medical device \"\n        \"and should not be used for treatment decisions.\"\n    ),\n    \"privacy_note\": (\n        \"Personal MRI DICOMs should be de-identified before being \"\n        \"uploaded, published, or included in any public repository.\"\n    )\n}\n\n# Save manifest JSON to /kaggle/working.\noutput_path = os.path.join(\n    WORKING_DIR,\n    \"project_manifest_200.json\"\n)\n\nwith open(output_path, \"w\") as f:\n    json.dump(manifest, f, indent=2)\n\nprint(\"Saved project manifest to:\")\nprint(output_path)\n\nprint(\"\\nManifest preview:\")\nprint(json.dumps(manifest, indent=2))\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:23:15.965891Z","iopub.execute_input":"2026-06-16T00:23:15.966189Z","iopub.status.idle":"2026-06-16T00:23:16.019655Z","shell.execute_reply.started":"2026-06-16T00:23:15.966165Z","shell.execute_reply":"2026-06-16T00:23:16.018773Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 11: PYTORCH DATASET CLASS SMOKE TEST\n# ============================================================\n# Purpose:\n# Build a basic PyTorch Dataset class that can load selected\n# lumbar MRI DICOM slices into model-ready tensors.\n#\n# This cell checks that the Dataset can:\n#   1. read selected DICOM paths\n#   2. load pixel arrays with pydicom\n#   3. normalize image slices\n#   4. resize images to a fixed shape\n#   5. return image tensors and metadata\n#   6. batch samples with a DataLoader\n#\n# This is a smoke test before model training.\n#\n# This cell does NOT attach diagnosis labels.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Dataset construction and PyTorch input validation.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Specialty imports for DICOM loading and PyTorch data handling.\nimport pydicom\nimport torch\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\n\n# Load selected image index from Step 5.\nimage_index_path = os.path.join(WORKING_DIR, \"selected_image_index_200.csv\")\nimage_index_df = pd.read_csv(image_index_path)\n\nif image_index_df.empty:\n    raise ValueError(\n        \"selected_image_index_200.csv is empty. \"\n        \"Run Step 5 before this Dataset smoke test.\"\n    )\n\nprint(\"Image index shape:\", image_index_df.shape)\n\nclass LumbarSliceDataset(Dataset):\n    \"\"\"\n    Basic PyTorch Dataset for individual lumbar MRI DICOM slices.\n\n    Each item returns:\n      - image tensor with shape (1, image_size, image_size)\n      - metadata describing the source DICOM slice\n\n    This first version does NOT attach diagnosis labels.\n    It confirms that image loading and tensor conversion work correctly.\n    \"\"\"\n\n    def __init__(self, image_index_df, image_size=224):\n        \"\"\"\n        image_index_df:\n            DataFrame containing at least:\n            study_id, series_id, series_description, filename, dcm_path\n\n        image_size:\n            final square image size for model input\n        \"\"\"\n        self.image_index_df = image_index_df.reset_index(drop=True)\n        self.image_size = image_size\n\n    def __len__(self):\n        \"\"\"Return number of DICOM slices available.\"\"\"\n        return len(self.image_index_df)\n\n    def __getitem__(self, idx):\n        \"\"\"Load one DICOM slice and return it as a normalized tensor.\"\"\"\n        row = self.image_index_df.iloc[idx]\n        dcm_path = row[\"dcm_path\"]\n\n        # Read DICOM metadata and pixel data.\n        ds = pydicom.dcmread(dcm_path)\n\n        # Convert pixel array to float32 for model input.\n        img = ds.pixel_array.astype(np.float32)\n\n        # Normalize each slice independently.\n        img = (img - img.mean()) / (img.std() + 1e-8)\n\n        # Convert from NumPy array (H, W) to PyTorch tensor (1, H, W).\n        img_tensor = torch.from_numpy(img).unsqueeze(0)\n\n        # Resize to fixed model input size.\n        # interpolate expects shape (N, C, H, W), so add a batch dimension temporarily.\n        img_tensor = img_tensor.unsqueeze(0)\n        img_tensor = F.interpolate(\n            img_tensor,\n            size=(self.image_size, self.image_size),\n            mode=\"bilinear\",\n            align_corners=False\n        )\n        img_tensor = img_tensor.squeeze(0)\n\n        metadata = {\n            \"study_id\": int(row[\"study_id\"]),\n            \"series_id\": int(row[\"series_id\"]),\n            \"series_description\": row[\"series_description\"],\n            \"filename\": row[\"filename\"],\n            \"dcm_path\": dcm_path\n        }\n\n        return img_tensor, metadata\n\n# Create the Dataset object.\nslice_dataset = LumbarSliceDataset(\n    image_index_df=image_index_df,\n    image_size=224\n)\n\nprint(\"Dataset length:\", len(slice_dataset))\n\n# Load one sample directly.\n# This checks that __getitem__ works for an individual DICOM slice.\nsample_img, sample_meta = slice_dataset[0]\n\nprint(\"\\n=== Single sample test ===\")\nprint(\"Image tensor shape:\", sample_img.shape)\nprint(\"Image tensor dtype:\", sample_img.dtype)\nprint(\"Image tensor min/max:\", sample_img.min().item(), sample_img.max().item())\nprint(\"Metadata:\", sample_meta)\n\n# Test DataLoader batching.\n# This checks that multiple samples can be batched for model training.\nslice_loader = DataLoader(\n    slice_dataset,\n    batch_size=8,\n    shuffle=True,\n    num_workers=0\n)\n\nbatch_imgs, batch_meta = next(iter(slice_loader))\n\nprint(\"\\n=== DataLoader batch test ===\")\nprint(\"Batch image tensor shape:\", batch_imgs.shape)\nprint(\"Batch image tensor dtype:\", batch_imgs.dtype)\n\nprint(\"\\nBatch metadata keys:\")\nprint(batch_meta.keys())\n\nprint(\"\\nDone: Dataset class smoke test passed.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:30:53.487590Z","iopub.execute_input":"2026-06-16T00:30:53.488458Z","iopub.status.idle":"2026-06-16T00:30:57.648164Z","shell.execute_reply.started":"2026-06-16T00:30:53.488419Z","shell.execute_reply":"2026-06-16T00:30:57.647314Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 12: BUILD FIRST TRAINING TARGET\n# ============================================================\n# Purpose:\n# Create the first simple binary classification target for model training.\n#\n# Target:\n#   spinal_canal_stenosis_* columns\n#\n# Binary mapping:\n#   Normal/Mild -> 0\n#   Moderate    -> 1\n#   Severe      -> 1\n#\n# Study-level rule:\n#   If ANY lumbar level has Moderate or Severe spinal canal stenosis,\n#   the study is labeled positive.\n#\n# Output:\n#   training_index_200_spinal_canal_binary.csv\n#\n# This cell does NOT load DICOM pixel data.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Training target construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\n# Load the selected DICOM image index and selected study labels.\nimage_index_df = pd.read_csv(\n    os.path.join(WORKING_DIR, \"selected_image_index_200.csv\")\n)\n\nlabels_df = pd.read_csv(\n    os.path.join(WORKING_DIR, \"selected_labels_200.csv\")\n)\n\n# Identify the spinal canal stenosis label columns.\ntarget_columns = [\n    col for col in labels_df.columns\n    if col.startswith(\"spinal_canal_stenosis\")\n]\n\nif len(target_columns) == 0:\n    raise ValueError(\n        \"No spinal_canal_stenosis columns found in selected_labels_200.csv.\"\n    )\n\nprint(\"Target columns:\")\nfor col in target_columns:\n    print(col)\n\n# Define binary label mapping.\nlabel_map = {\n    \"Normal/Mild\": 0,\n    \"Moderate\": 1,\n    \"Severe\": 1\n}\n\n# Keep study IDs as the anchor for study-level labels.\nbinary_labels = labels_df[[\"study_id\"]].copy()\n\n# Convert each spinal canal stenosis level from text labels to binary labels.\n# This uses column-wise map() to avoid pandas deprecation warnings.\nbinary_target_values = pd.DataFrame({\n    col: labels_df[col].map(label_map)\n    for col in target_columns\n})\n\n# Create one study-level binary target.\n# If any disc level is Moderate/Severe, max() becomes 1.\n# If all disc levels are Normal/Mild, max() remains 0.\nbinary_labels[\"any_moderate_severe_spinal_canal_stenosis\"] = (\n    binary_target_values.max(axis=1)\n)\n\nprint(\"\\n=== Study-level binary label distribution ===\")\ndisplay(binary_labels[\"any_moderate_severe_spinal_canal_stenosis\"].value_counts())\n\n# Join the study-level binary label onto every DICOM slice by study_id.\ntraining_index_df = image_index_df.merge(\n    binary_labels,\n    on=\"study_id\",\n    how=\"left\"\n)\n\n# Drop rows without labels, if any.\ntraining_index_df = training_index_df.dropna(\n    subset=[\"any_moderate_severe_spinal_canal_stenosis\"]\n).copy()\n\nif training_index_df.empty:\n    raise ValueError(\n        \"Training index is empty after merging image paths with labels.\"\n    )\n\n# Create final target column as integer 0/1.\ntraining_index_df[\"target\"] = training_index_df[\n    \"any_moderate_severe_spinal_canal_stenosis\"\n].astype(int)\n\nprint(\"\\n=== Training index shape ===\")\nprint(training_index_df.shape)\n\nprint(\"\\n=== Slice-level target distribution ===\")\ndisplay(training_index_df[\"target\"].value_counts())\n\nprint(\"\\n=== Preview ===\")\ndisplay(training_index_df.head())\n\n# Save training index to /kaggle/working.\noutput_path = os.path.join(\n    WORKING_DIR,\n    \"training_index_200_spinal_canal_binary.csv\"\n)\n\ntraining_index_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved training index to:\")\nprint(output_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:32:16.254608Z","iopub.execute_input":"2026-06-16T00:32:16.255644Z","iopub.status.idle":"2026-06-16T00:32:16.415292Z","shell.execute_reply.started":"2026-06-16T00:32:16.255609Z","shell.execute_reply":"2026-06-16T00:32:16.414381Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 13: STUDY-LEVEL TRAIN/VALIDATION SPLIT\n# ============================================================\n# Purpose:\n# Split the 200-study binary training index into train and validation sets.\n#\n# Important:\n# The split is performed by study_id, NOT by individual DICOM slice.\n# This prevents slices from the same MRI study from appearing in both\n# training and validation.\n#\n# Output:\n#   training_index_200_spinal_canal_binary_split.csv\n#\n# This cell does NOT load DICOM pixel data.\n# This cell does NOT train a model.\n#\n# Workflow phase:\n# Phase A — Training target construction.\n#\n# Manual accelerator note:\n# GPU is NOT needed for this cell.\n# If using the manual quota-saving workflow, Accelerator can be: None.\n# ============================================================\n\nfrom sklearn.model_selection import train_test_split\n\n# Load binary training index from Step 12.\ntraining_index_path = os.path.join(\n    WORKING_DIR,\n    \"training_index_200_spinal_canal_binary.csv\"\n)\n\ntraining_index_df = pd.read_csv(training_index_path)\n\nif training_index_df.empty:\n    raise ValueError(\n        \"training_index_200_spinal_canal_binary.csv is empty. \"\n        \"Run Step 12 before creating the train/validation split.\"\n    )\n\n# Create one row per study with its study-level target.\n# This is the table used for splitting.\nstudy_targets_df = (\n    training_index_df[[\"study_id\", \"target\"]]\n    .drop_duplicates()\n    .reset_index(drop=True)\n)\n\nprint(\"Study-level targets shape:\", study_targets_df.shape)\n\nprint(\"\\n=== Study-level target distribution before split ===\")\ndisplay(study_targets_df[\"target\"].value_counts())\n\n# Stratified split keeps roughly the same positive/negative balance\n# in both train and validation sets.\ntrain_studies_df, val_studies_df = train_test_split(\n    study_targets_df,\n    test_size=0.2,\n    random_state=RANDOM_STATE,\n    stratify=study_targets_df[\"target\"]\n)\n\nprint(\"\\nTrain studies:\", train_studies_df.shape)\nprint(\"Validation studies:\", val_studies_df.shape)\n\nprint(\"\\n=== Train study target distribution ===\")\ndisplay(train_studies_df[\"target\"].value_counts())\n\nprint(\"\\n=== Validation study target distribution ===\")\ndisplay(val_studies_df[\"target\"].value_counts())\n\n# Assign split labels back to every DICOM slice.\n# All slices from a given study inherit that study's split.\ntraining_index_df[\"split\"] = \"unused\"\n\ntraining_index_df.loc[\n    training_index_df[\"study_id\"].isin(train_studies_df[\"study_id\"]),\n    \"split\"\n] = \"train\"\n\ntraining_index_df.loc[\n    training_index_df[\"study_id\"].isin(val_studies_df[\"study_id\"]),\n    \"split\"\n] = \"val\"\n\nprint(\"\\n=== Slice-level split distribution ===\")\ndisplay(training_index_df[\"split\"].value_counts())\n\nprint(\"\\n=== Slice-level target distribution by split ===\")\ndisplay(pd.crosstab(training_index_df[\"split\"], training_index_df[\"target\"]))\n\n# Safety check:\n# No study_id should appear in both train and validation.\ntrain_study_ids = set(train_studies_df[\"study_id\"])\nval_study_ids = set(val_studies_df[\"study_id\"])\noverlap = train_study_ids.intersection(val_study_ids)\n\nprint(\"\\nStudy overlap between train and val:\", len(overlap))\n\nif len(overlap) != 0:\n    raise ValueError(\"Data leakage detected: some studies are in both train and val.\")\n\n# Safety check:\n# Every row should now belong to either train or validation.\nunused_count = (training_index_df[\"split\"] == \"unused\").sum()\n\nif unused_count != 0:\n    raise ValueError(\n        f\"{unused_count} rows were not assigned to train or validation.\"\n    )\n\n# Save split training index to /kaggle/working.\noutput_path = os.path.join(\n    WORKING_DIR,\n    \"training_index_200_spinal_canal_binary_split.csv\"\n)\n\ntraining_index_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved split training index to:\")\nprint(output_path)\n\nprint(\"\\nPreview:\")\ndisplay(training_index_df.head())\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T00:47:24.871418Z","iopub.execute_input":"2026-06-16T00:47:24.871772Z","iopub.status.idle":"2026-06-16T00:47:25.836102Z","shell.execute_reply.started":"2026-06-16T00:47:24.871744Z","shell.execute_reply":"2026-06-16T00:47:25.835041Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 13B: REBUILD FULL PHASE A STATE\n# ============================================================\n# Purpose:\n# Rebuild all Phase A artifacts after a Kaggle runtime reset.\n#\n# Kaggle may wipe /kaggle/working when the runtime restarts,\n# especially after changing Accelerator settings.\n#\n# Run-All workflow:\n#   This step is skipped by default and is not needed.\n#\n# Manual workflow:\n#   This step is required after switching from Accelerator=None\n#   to Accelerator=GPU, before running Step 14.\n#\n# Manual users:\n#   If you changed Accelerator from None to GPU and /kaggle/working\n#   was reset, set:\n#\n#       RUN_PHASE_A_REBUILD = True\n#\n#   Then run this cell before Step 14.\n#\n# This cell rebuilds:\n#   1. study_metadata_summary.csv\n#   2. selected_studies_200.csv\n#   3. selected_series_200.csv\n#   4. selected_image_index_200.csv\n#   5. selected_labels_200.csv\n#   6. selected_label_distribution_200.csv\n#   7. project_manifest_200.json\n#   8. training_index_200_spinal_canal_binary.csv\n#   9. training_index_200_spinal_canal_binary_split.csv\n#\n# Workflow phase:\n# Phase B — Recovery bridge before training.\n#\n# Manual accelerator note:\n# If using the manual quota-saving workflow, run this AFTER\n# switching Accelerator to GPU and BEFORE Step 14.\n# ============================================================\n\nRUN_PHASE_A_REBUILD = False\n\nif not RUN_PHASE_A_REBUILD:\n    print(\n        \"Skipping Step 13B rebuild.\\n\"\n        \"Run-All workflow does not need this step.\\n\"\n        \"For manual workflow after an Accelerator change, set \"\n        \"RUN_PHASE_A_REBUILD = True and rerun this cell.\"\n    )\n\nelse:\n    from sklearn.model_selection import train_test_split\n\n    # ------------------------------------------------------------\n    # Load original RSNA metadata\n    # ------------------------------------------------------------\n    train_df = pd.read_csv(os.path.join(DATASET_DIR, \"train.csv\"))\n    series_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_series_descriptions.csv\"))\n    coords_df = pd.read_csv(os.path.join(DATASET_DIR, \"train_label_coordinates.csv\"))\n\n    # ------------------------------------------------------------\n    # Rebuild Step 2: study_metadata_summary.csv\n    # ------------------------------------------------------------\n    series_count_df = (\n        series_df\n        .groupby(\"study_id\")\n        .agg(\n            num_series=(\"series_id\", \"count\"),\n            series_descriptions=(\"series_description\", lambda x: sorted(list(x)))\n        )\n        .reset_index()\n    )\n\n    sequence_presence_df = (\n        series_df\n        .assign(value=True)\n        .pivot_table(\n            index=\"study_id\",\n            columns=\"series_description\",\n            values=\"value\",\n            aggfunc=\"any\",\n            fill_value=False\n        )\n        .reset_index()\n    )\n\n    study_meta_df = series_count_df.merge(\n        sequence_presence_df,\n        on=\"study_id\",\n        how=\"left\"\n    )\n\n    study_meta_df[\"has_all_core_sequences\"] = (\n        study_meta_df[CORE_SEQUENCES].all(axis=1)\n    )\n\n    study_meta_path = os.path.join(WORKING_DIR, \"study_metadata_summary.csv\")\n    study_meta_df.to_csv(study_meta_path, index=False)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 3: selected_studies_200.csv\n    # ------------------------------------------------------------\n    eligible_studies = study_meta_df[\n        study_meta_df[\"has_all_core_sequences\"] == True\n    ].copy()\n\n    selected_studies_df = eligible_studies.sample(\n        n=SUBSET_SIZE,\n        random_state=RANDOM_STATE\n    ).copy()\n\n    selected_studies_df = selected_studies_df[\n        [\"study_id\", \"Axial T2\", \"Sagittal T1\", \"Sagittal T2/STIR\", \"has_all_core_sequences\"]\n    ].sort_values(\"study_id\")\n\n    selected_studies_path = os.path.join(WORKING_DIR, \"selected_studies_200.csv\")\n    selected_studies_df.to_csv(selected_studies_path, index=False)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 4: selected_series_200.csv\n    # ------------------------------------------------------------\n    selected_series_df = series_df[\n        series_df[\"study_id\"].isin(selected_studies_df[\"study_id\"])\n    ].copy()\n\n    selected_series_df = selected_series_df.sort_values(\n        [\"study_id\", \"series_description\", \"series_id\"]\n    ).reset_index(drop=True)\n\n    selected_series_path = os.path.join(WORKING_DIR, \"selected_series_200.csv\")\n    selected_series_df.to_csv(selected_series_path, index=False)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 5: selected_image_index_200.csv\n    # ------------------------------------------------------------\n    rows = []\n\n    for _, row in selected_series_df.iterrows():\n        study_id = str(row[\"study_id\"])\n        series_id = str(row[\"series_id\"])\n        series_description = row[\"series_description\"]\n\n        series_dir = os.path.join(TRAIN_IMAGES_DIR, study_id, series_id)\n\n        if not os.path.exists(series_dir):\n            print(\"Missing series folder:\", series_dir)\n            continue\n\n        dicom_files = [\n            filename for filename in os.listdir(series_dir)\n            if filename.lower().endswith(\".dcm\")\n        ]\n\n        dicom_files = sorted(\n            dicom_files,\n            key=lambda filename: int(os.path.splitext(filename)[0])\n        )\n\n        for instance_index, filename in enumerate(dicom_files):\n            rows.append({\n                \"study_id\": int(study_id),\n                \"series_id\": int(series_id),\n                \"series_description\": series_description,\n                \"instance_index\": instance_index,\n                \"filename\": filename,\n                \"dcm_path\": os.path.join(series_dir, filename)\n            })\n\n    image_index_df = pd.DataFrame(rows)\n\n    if image_index_df.empty:\n        raise ValueError(\n            \"No DICOM image paths were indexed during Step 13B rebuild.\"\n        )\n\n    image_index_path = os.path.join(WORKING_DIR, \"selected_image_index_200.csv\")\n    image_index_df.to_csv(image_index_path, index=False)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 6: selected_labels_200.csv\n    # ------------------------------------------------------------\n    selected_labels_df = train_df[\n        train_df[\"study_id\"].isin(selected_studies_df[\"study_id\"])\n    ].copy()\n\n    if selected_labels_df.empty:\n        raise ValueError(\n            \"No labels found for selected studies during Step 13B rebuild.\"\n        )\n\n    selected_labels_path = os.path.join(WORKING_DIR, \"selected_labels_200.csv\")\n    selected_labels_df.to_csv(selected_labels_path, index=False)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 9: selected_label_distribution_200.csv\n    # ------------------------------------------------------------\n    label_columns = [\n        col for col in selected_labels_df.columns\n        if col != \"study_id\"\n    ]\n\n    label_distribution_rows = []\n\n    for col in label_columns:\n        counts = selected_labels_df[col].value_counts(dropna=False)\n\n        row = {\"label_column\": col}\n\n        for label_value, count in counts.items():\n            row[str(label_value)] = count\n\n        label_distribution_rows.append(row)\n\n    label_distribution_df = pd.DataFrame(label_distribution_rows).fillna(0)\n\n    label_distribution_path = os.path.join(\n        WORKING_DIR,\n        \"selected_label_distribution_200.csv\"\n    )\n\n    label_distribution_df.to_csv(label_distribution_path, index=False)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 10: project_manifest_200.json\n    # ------------------------------------------------------------\n    manifest = {\n        \"project_name\": \"lumbar_vibe\",\n        \"dataset\": \"RSNA 2024 Lumbar Spine Degenerative Classification\",\n        \"dataset_root\": DATASET_DIR,\n        \"subset_name\": f\"selected_{SUBSET_SIZE}_studies\",\n        \"subset_size_studies\": int(len(selected_studies_df)),\n        \"num_selected_series\": int(len(selected_series_df)),\n        \"num_selected_dicom_images\": int(len(image_index_df)),\n        \"num_label_rows\": int(len(selected_labels_df)),\n        \"core_sequences\": CORE_SEQUENCES,\n        \"selection_method\": {\n            \"eligible_rule\": (\n                \"study must contain Axial T2, \"\n                \"Sagittal T1, and Sagittal T2/STIR\"\n            ),\n            \"sampling_method\": \"random sample\",\n            \"random_state\": RANDOM_STATE\n        },\n        \"created_files\": [\n            \"study_metadata_summary.csv\",\n            \"selected_studies_200.csv\",\n            \"selected_series_200.csv\",\n            \"selected_image_index_200.csv\",\n            \"selected_labels_200.csv\",\n            \"selected_label_distribution_200.csv\"\n        ],\n        \"medical_use_note\": (\n            \"This project is for educational and research workflow \"\n            \"development only. It is not a diagnostic medical device \"\n            \"and should not be used for treatment decisions.\"\n        ),\n        \"privacy_note\": (\n            \"Personal MRI DICOMs should be de-identified before being \"\n            \"uploaded, published, or included in any public repository.\"\n        )\n    }\n\n    manifest_path = os.path.join(WORKING_DIR, \"project_manifest_200.json\")\n\n    with open(manifest_path, \"w\") as f:\n        json.dump(manifest, f, indent=2)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 12: binary training target\n    # ------------------------------------------------------------\n    target_columns = [\n        col for col in selected_labels_df.columns\n        if col.startswith(\"spinal_canal_stenosis\")\n    ]\n\n    if len(target_columns) == 0:\n        raise ValueError(\n            \"No spinal_canal_stenosis columns found during Step 13B rebuild.\"\n        )\n\n    label_map = {\n        \"Normal/Mild\": 0,\n        \"Moderate\": 1,\n        \"Severe\": 1\n    }\n\n    binary_labels = selected_labels_df[[\"study_id\"]].copy()\n\n    binary_target_values = pd.DataFrame({\n        col: selected_labels_df[col].map(label_map)\n        for col in target_columns\n    })\n\n    binary_labels[\"any_moderate_severe_spinal_canal_stenosis\"] = (\n        binary_target_values.max(axis=1)\n    )\n\n    training_index_df = image_index_df.merge(\n        binary_labels,\n        on=\"study_id\",\n        how=\"left\"\n    )\n\n    training_index_df = training_index_df.dropna(\n        subset=[\"any_moderate_severe_spinal_canal_stenosis\"]\n    ).copy()\n\n    if training_index_df.empty:\n        raise ValueError(\n            \"Training index is empty during Step 13B rebuild.\"\n        )\n\n    training_index_df[\"target\"] = training_index_df[\n        \"any_moderate_severe_spinal_canal_stenosis\"\n    ].astype(int)\n\n    training_index_path = os.path.join(\n        WORKING_DIR,\n        \"training_index_200_spinal_canal_binary.csv\"\n    )\n\n    training_index_df.to_csv(training_index_path, index=False)\n\n    # ------------------------------------------------------------\n    # Rebuild Step 13: train/validation split\n    # ------------------------------------------------------------\n    study_targets_df = (\n        training_index_df[[\"study_id\", \"target\"]]\n        .drop_duplicates()\n        .reset_index(drop=True)\n    )\n\n    train_studies_df, val_studies_df = train_test_split(\n        study_targets_df,\n        test_size=0.2,\n        random_state=RANDOM_STATE,\n        stratify=study_targets_df[\"target\"]\n    )\n\n    training_index_df[\"split\"] = \"unused\"\n\n    training_index_df.loc[\n        training_index_df[\"study_id\"].isin(train_studies_df[\"study_id\"]),\n        \"split\"\n    ] = \"train\"\n\n    training_index_df.loc[\n        training_index_df[\"study_id\"].isin(val_studies_df[\"study_id\"]),\n        \"split\"\n    ] = \"val\"\n\n    train_study_ids = set(train_studies_df[\"study_id\"])\n    val_study_ids = set(val_studies_df[\"study_id\"])\n    overlap = train_study_ids.intersection(val_study_ids)\n\n    if len(overlap) != 0:\n        raise ValueError(\n            \"Data leakage detected during Step 13B rebuild.\"\n        )\n\n    unused_count = (training_index_df[\"split\"] == \"unused\").sum()\n\n    if unused_count != 0:\n        raise ValueError(\n            f\"{unused_count} rows were not assigned during Step 13B rebuild.\"\n        )\n\n    split_path = os.path.join(\n        WORKING_DIR,\n        \"training_index_200_spinal_canal_binary_split.csv\"\n    )\n\n    training_index_df.to_csv(split_path, index=False)\n\n    # ------------------------------------------------------------\n    # Final rebuild checks\n    # ------------------------------------------------------------\n    expected_rebuild_files = [\n        \"study_metadata_summary.csv\",\n        \"selected_studies_200.csv\",\n        \"selected_series_200.csv\",\n        \"selected_image_index_200.csv\",\n        \"selected_labels_200.csv\",\n        \"selected_label_distribution_200.csv\",\n        \"project_manifest_200.json\",\n        \"training_index_200_spinal_canal_binary.csv\",\n        \"training_index_200_spinal_canal_binary_split.csv\"\n    ]\n\n    missing_files = [\n        filename for filename in expected_rebuild_files\n        if not os.path.exists(os.path.join(WORKING_DIR, filename))\n    ]\n\n    if len(missing_files) != 0:\n        raise FileNotFoundError(\n            f\"Step 13B rebuild is incomplete. Missing files: {missing_files}\"\n        )\n\n    print(\"Step 13B rebuild complete.\")\n    print(\"\\nFiles rebuilt:\")\n    for filename in expected_rebuild_files:\n        print(\"-\", filename)\n\n    print(\"\\nSelected studies:\", selected_studies_df.shape)\n    print(\"Selected series:\", selected_series_df.shape)\n    print(\"Selected DICOM images:\", image_index_df.shape)\n    print(\"Selected labels:\", selected_labels_df.shape)\n    print(\"Training index:\", training_index_df.shape)\n\n    print(\"\\nStudy-level target distribution:\")\n    display(study_targets_df[\"target\"].value_counts())\n\n    print(\"\\nSlice-level split/target table:\")\n    display(pd.crosstab(training_index_df[\"split\"], training_index_df[\"target\"]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T01:09:35.996365Z","iopub.execute_input":"2026-06-16T01:09:35.997242Z","iopub.status.idle":"2026-06-16T01:09:40.803187Z","shell.execute_reply.started":"2026-06-16T01:09:35.997201Z","shell.execute_reply":"2026-06-16T01:09:40.802502Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 14: FIRST BASELINE TRAINING SMOKE TEST\n# ============================================================\n# Purpose:\n# Train a very small CNN for 1 epoch to confirm the full training pipeline.\n#\n# This cell validates:\n#   1. DICOM paths load through a labeled PyTorch Dataset\n#   2. DataLoader batches reach the model\n#   3. Forward pass, loss calculation, backpropagation, and optimizer steps work\n#   4. Model weights can be saved to /kaggle/working\n#\n# This is NOT the final model.\n# This is a training smoke test.\n#\n# Workflow phase:\n# Phase B — Training and evaluation.\n#\n# Manual accelerator note:\n# GPU is recommended for this cell.\n# If using the manual quota-saving workflow, run Step 13B after\n# switching Accelerator to GPU and before running this cell.\n# ============================================================\n\n# Specialty imports for DICOM loading and PyTorch training.\nimport pydicom\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\n\n# Select compute device.\n# If Kaggle GPU is enabled, this should print \"cuda\".\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(\"Using device:\", device)\n\nif device.type != \"cuda\":\n    print(\n        \"Warning: GPU is not available. \"\n        \"This cell can run on CPU, but training may be slower.\"\n    )\n\n# Load split training index from Step 13 or Step 13B.\nsplit_path = os.path.join(\n    WORKING_DIR,\n    \"training_index_200_spinal_canal_binary_split.csv\"\n)\n\nif not os.path.exists(split_path):\n    raise FileNotFoundError(\n        \"Missing training_index_200_spinal_canal_binary_split.csv. \"\n        \"If you changed Accelerator settings, run Step 13B before Step 14.\"\n    )\n\nsplit_df = pd.read_csv(split_path)\n\n# Separate train and validation slice rows.\ntrain_df = split_df[split_df[\"split\"] == \"train\"].reset_index(drop=True)\nval_df = split_df[split_df[\"split\"] == \"val\"].reset_index(drop=True)\n\nif train_df.empty or val_df.empty:\n    raise ValueError(\n        \"Train or validation split is empty. \"\n        \"Check the split column in training_index_200_spinal_canal_binary_split.csv.\"\n    )\n\nprint(\"Train slices:\", train_df.shape)\nprint(\"Val slices:\", val_df.shape)\n\nclass LumbarSliceBinaryDataset(Dataset):\n    \"\"\"\n    PyTorch Dataset for binary classification using individual DICOM slices.\n\n    Each item returns:\n      - image tensor with shape (1, image_size, image_size)\n      - binary target label\n    \"\"\"\n\n    def __init__(self, df, image_size=224):\n        self.df = df.reset_index(drop=True)\n        self.image_size = image_size\n\n    def __len__(self):\n        \"\"\"Return number of DICOM slices in this split.\"\"\"\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        \"\"\"Load one DICOM slice and its binary target label.\"\"\"\n        row = self.df.iloc[idx]\n\n        # Read DICOM metadata and pixel data.\n        ds = pydicom.dcmread(row[\"dcm_path\"])\n\n        # Convert image pixels to float32 for model input.\n        img = ds.pixel_array.astype(np.float32)\n\n        # Normalize each slice independently.\n        img = (img - img.mean()) / (img.std() + 1e-8)\n\n        # Convert from (H, W) to (1, H, W).\n        img_tensor = torch.from_numpy(img).unsqueeze(0)\n\n        # Resize to fixed model input size.\n        img_tensor = img_tensor.unsqueeze(0)\n        img_tensor = F.interpolate(\n            img_tensor,\n            size=(self.image_size, self.image_size),\n            mode=\"bilinear\",\n            align_corners=False\n        )\n        img_tensor = img_tensor.squeeze(0)\n\n        # Load binary target label.\n        target = torch.tensor(row[\"target\"], dtype=torch.float32)\n\n        return img_tensor, target\n\n# Create train and validation datasets.\ntrain_dataset = LumbarSliceBinaryDataset(train_df, image_size=224)\nval_dataset = LumbarSliceBinaryDataset(val_df, image_size=224)\n\n# Create DataLoaders.\n# num_workers=0 is conservative and stable in notebook environments.\ntrain_loader = DataLoader(\n    train_dataset,\n    batch_size=32,\n    shuffle=True,\n    num_workers=0\n)\n\nval_loader = DataLoader(\n    val_dataset,\n    batch_size=32,\n    shuffle=False,\n    num_workers=0\n)\n\nclass TinyLumbarCNN(nn.Module):\n    \"\"\"\n    Very small baseline CNN.\n\n    This architecture is intentionally simple.\n    Its purpose is pipeline validation, not strong diagnostic performance.\n    \"\"\"\n\n    def __init__(self):\n        super().__init__()\n\n        self.features = nn.Sequential(\n            nn.Conv2d(1, 16, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n\n            nn.Conv2d(16, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n\n            nn.Conv2d(32, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.AdaptiveAvgPool2d((1, 1))\n        )\n\n        # One output logit for binary classification.\n        self.classifier = nn.Linear(64, 1)\n\n    def forward(self, x):\n        \"\"\"Forward pass from image tensor to binary logit.\"\"\"\n        x = self.features(x)\n        x = x.flatten(1)\n        x = self.classifier(x).squeeze(1)\n        return x\n\n# Initialize model, loss function, and optimizer.\nmodel = TinyLumbarCNN().to(device)\n\ncriterion = nn.BCEWithLogitsLoss()\noptimizer = torch.optim.Adam(model.parameters(), lr=1e-3)\n\n# Train for one epoch.\nmodel.train()\nrunning_loss = 0.0\n\nfor batch_idx, (images, targets) in enumerate(train_loader):\n    images = images.to(device)\n    targets = targets.to(device)\n\n    # Reset gradients from the previous batch.\n    optimizer.zero_grad()\n\n    # Forward pass.\n    logits = model(images)\n\n    # Binary classification loss.\n    loss = criterion(logits, targets)\n\n    # Backpropagation and optimizer update.\n    loss.backward()\n    optimizer.step()\n\n    running_loss += loss.item()\n\n    if batch_idx % 25 == 0:\n        print(f\"Batch {batch_idx}/{len(train_loader)} | Loss: {loss.item():.4f}\")\n\navg_train_loss = running_loss / len(train_loader)\n\nprint(\"\\nAverage train loss:\", avg_train_loss)\n\n# Run a validation pass after training.\nmodel.eval()\nval_loss = 0.0\ncorrect = 0\ntotal = 0\n\nwith torch.no_grad():\n    for images, targets in val_loader:\n        images = images.to(device)\n        targets = targets.to(device)\n\n        logits = model(images)\n        loss = criterion(logits, targets)\n        val_loss += loss.item()\n\n        probs = torch.sigmoid(logits)\n        preds = (probs >= 0.5).float()\n\n        correct += (preds == targets).sum().item()\n        total += targets.numel()\n\navg_val_loss = val_loss / len(val_loader)\nval_accuracy = correct / total\n\nprint(\"\\nAverage val loss:\", avg_val_loss)\nprint(\"Validation accuracy:\", val_accuracy)\n\n# Save smoke-test model weights to /kaggle/working.\nmodel_path = os.path.join(WORKING_DIR, \"tiny_lumbar_cnn_smoke_test.pth\")\ntorch.save(model.state_dict(), model_path)\n\nprint(\"\\nSaved smoke-test model weights to:\")\nprint(model_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T01:09:59.765720Z","iopub.execute_input":"2026-06-16T01:09:59.766563Z","iopub.status.idle":"2026-06-16T01:13:10.917743Z","shell.execute_reply.started":"2026-06-16T01:09:59.766531Z","shell.execute_reply":"2026-06-16T01:13:10.916793Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 15: LOAD SAVED MODEL + RUN INFERENCE TEST\n# ============================================================\n# Purpose:\n# Confirm that the saved model weights from Step 14 are reusable.\n#\n# This cell:\n#   1. Rebuilds the same TinyLumbarCNN architecture\n#   2. Loads tiny_lumbar_cnn_smoke_test.pth\n#   3. Runs inference on a small sample of validation slices\n#   4. Saves example probabilities and predictions\n#\n# This is NOT a medical interpretation.\n# This is an engineering check that the saved model artifact works.\n#\n# Workflow phase:\n# Phase B — Training and evaluation.\n#\n# Manual accelerator note:\n# GPU is optional for this cell, but if using the manual workflow,\n# do not change Accelerator settings after Step 14. Keep the same\n# Phase B runtime alive until Steps 15–18 are complete.\n# ============================================================\n\n# Select compute device.\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(\"Using device:\", device)\n\n# Load split index from Step 13 or Step 13B.\nsplit_path = os.path.join(\n    WORKING_DIR,\n    \"training_index_200_spinal_canal_binary_split.csv\"\n)\n\nif not os.path.exists(split_path):\n    raise FileNotFoundError(\n        \"Missing training_index_200_spinal_canal_binary_split.csv. \"\n        \"If you changed Accelerator settings, run Step 13B before Step 15.\"\n    )\n\nsplit_df = pd.read_csv(split_path)\n\n# Use validation split for inference check.\nval_df = split_df[split_df[\"split\"] == \"val\"].reset_index(drop=True)\n\nif val_df.empty:\n    raise ValueError(\n        \"Validation split is empty. \"\n        \"Check training_index_200_spinal_canal_binary_split.csv.\"\n    )\n\nprint(\"Validation slices:\", val_df.shape)\n\nclass LumbarSliceBinaryDataset(Dataset):\n    \"\"\"\n    Dataset for loading individual DICOM slices for binary inference.\n    \"\"\"\n\n    def __init__(self, df, image_size=224):\n        self.df = df.reset_index(drop=True)\n        self.image_size = image_size\n\n    def __len__(self):\n        \"\"\"Return number of validation slices.\"\"\"\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        \"\"\"Load one DICOM slice, target label, and metadata.\"\"\"\n        row = self.df.iloc[idx]\n\n        # Read DICOM metadata and pixel data.\n        ds = pydicom.dcmread(row[\"dcm_path\"])\n\n        # Convert image pixels to float32 for model input.\n        img = ds.pixel_array.astype(np.float32)\n\n        # Normalize each slice independently.\n        img = (img - img.mean()) / (img.std() + 1e-8)\n\n        # Convert from (H, W) to (1, H, W).\n        img_tensor = torch.from_numpy(img).unsqueeze(0)\n\n        # Resize to fixed model input size.\n        img_tensor = img_tensor.unsqueeze(0)\n        img_tensor = F.interpolate(\n            img_tensor,\n            size=(self.image_size, self.image_size),\n            mode=\"bilinear\",\n            align_corners=False\n        )\n        img_tensor = img_tensor.squeeze(0)\n\n        target = torch.tensor(row[\"target\"], dtype=torch.float32)\n\n        metadata = {\n            \"study_id\": int(row[\"study_id\"]),\n            \"series_id\": int(row[\"series_id\"]),\n            \"series_description\": row[\"series_description\"],\n            \"filename\": row[\"filename\"],\n            \"target\": int(row[\"target\"])\n        }\n\n        return img_tensor, target, metadata\n\nclass TinyLumbarCNN(nn.Module):\n    \"\"\"\n    Same architecture used in Step 14.\n\n    The architecture must match exactly before loading saved weights.\n    \"\"\"\n\n    def __init__(self):\n        super().__init__()\n\n        self.features = nn.Sequential(\n            nn.Conv2d(1, 16, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n\n            nn.Conv2d(16, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n\n            nn.Conv2d(32, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.AdaptiveAvgPool2d((1, 1))\n        )\n\n        self.classifier = nn.Linear(64, 1)\n\n    def forward(self, x):\n        \"\"\"Forward pass from image tensor to binary logit.\"\"\"\n        x = self.features(x)\n        x = x.flatten(1)\n        x = self.classifier(x).squeeze(1)\n        return x\n\n# Rebuild model and load saved weights.\nmodel = TinyLumbarCNN().to(device)\n\nweights_path = os.path.join(WORKING_DIR, \"tiny_lumbar_cnn_smoke_test.pth\")\n\nif not os.path.exists(weights_path):\n    raise FileNotFoundError(\n        \"Missing tiny_lumbar_cnn_smoke_test.pth. \"\n        \"Rerun Step 14 before Step 15.\"\n    )\n\nmodel.load_state_dict(torch.load(weights_path, map_location=device))\nmodel.eval()\n\nprint(\"Loaded weights from:\", weights_path)\n\n# Build validation DataLoader.\nval_dataset = LumbarSliceBinaryDataset(val_df, image_size=224)\n\nval_loader = DataLoader(\n    val_dataset,\n    batch_size=32,\n    shuffle=False,\n    num_workers=0\n)\n\nresults = []\n\nwith torch.no_grad():\n    for batch_idx, (images, targets, metadata) in enumerate(val_loader):\n        images = images.to(device)\n\n        # Convert model logits into probabilities and binary predictions.\n        logits = model(images)\n        probs = torch.sigmoid(logits)\n        preds = (probs >= 0.5).float()\n\n        batch_size = images.shape[0]\n\n        for i in range(batch_size):\n            results.append({\n                \"study_id\": int(metadata[\"study_id\"][i]),\n                \"series_id\": int(metadata[\"series_id\"][i]),\n                \"series_description\": metadata[\"series_description\"][i],\n                \"filename\": metadata[\"filename\"][i],\n                \"target\": int(metadata[\"target\"][i]),\n                \"probability_positive\": float(probs[i].cpu().item()),\n                \"prediction\": int(preds[i].cpu().item())\n            })\n\n        # Limit this first inference smoke test to a few batches.\n        if batch_idx >= 4:\n            break\n\nresults_df = pd.DataFrame(results)\n\nif results_df.empty:\n    raise ValueError(\"Inference smoke test produced no results.\")\n\nprint(\"\\nInference results preview:\")\ndisplay(results_df.head(20))\n\nprint(\"\\nPrediction distribution:\")\ndisplay(results_df[\"prediction\"].value_counts())\n\nprint(\"\\nTarget distribution in sampled inference rows:\")\ndisplay(results_df[\"target\"].value_counts())\n\n# Save inference smoke-test output to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"inference_smoke_test_results.csv\")\nresults_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved inference smoke-test results to:\")\nprint(output_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T01:31:45.087961Z","iopub.execute_input":"2026-06-16T01:31:45.088988Z","iopub.status.idle":"2026-06-16T01:31:46.358572Z","shell.execute_reply.started":"2026-06-16T01:31:45.088954Z","shell.execute_reply":"2026-06-16T01:31:46.357892Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 16: FULL VALIDATION METRICS\n# ============================================================\n# Purpose:\n# Evaluate the smoke-test model on the full validation split.\n#\n# This cell calculates:\n#   1. Accuracy\n#   2. Precision\n#   3. Recall\n#   4. F1 score\n#   5. ROC-AUC\n#   6. Confusion matrix\n#\n# Outputs:\n#   validation_metrics_200_smoke_test.csv\n#   validation_predictions_200_smoke_test.csv\n#\n# This is an engineering benchmark, not medical validation.\n#\n# Workflow phase:\n# Phase B — Training and evaluation.\n#\n# Manual accelerator note:\n# GPU is optional for this cell, but if using the manual workflow,\n# do not change Accelerator settings after Step 14. Keep the same\n# Phase B runtime alive until Steps 15–18 are complete.\n# ============================================================\n\nfrom sklearn.metrics import (\n    accuracy_score,\n    precision_score,\n    recall_score,\n    f1_score,\n    roc_auc_score,\n    confusion_matrix\n)\n\n# Select compute device.\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(\"Using device:\", device)\n\n# Required inputs from previous Phase B steps.\nsplit_path = os.path.join(\n    WORKING_DIR,\n    \"training_index_200_spinal_canal_binary_split.csv\"\n)\n\nweights_path = os.path.join(\n    WORKING_DIR,\n    \"tiny_lumbar_cnn_smoke_test.pth\"\n)\n\nif not os.path.exists(split_path):\n    raise FileNotFoundError(\n        \"Missing training_index_200_spinal_canal_binary_split.csv. \"\n        \"If you changed Accelerator settings, run Step 13B before Step 16.\"\n    )\n\nif not os.path.exists(weights_path):\n    raise FileNotFoundError(\n        \"Missing tiny_lumbar_cnn_smoke_test.pth. \"\n        \"Rerun Step 14 before Step 16.\"\n    )\n\n# Load validation rows.\nsplit_df = pd.read_csv(split_path)\nval_df = split_df[split_df[\"split\"] == \"val\"].reset_index(drop=True)\n\nif val_df.empty:\n    raise ValueError(\n        \"Validation split is empty. \"\n        \"Check training_index_200_spinal_canal_binary_split.csv.\"\n    )\n\nprint(\"Validation slices:\", val_df.shape)\n\nclass LumbarSliceBinaryDataset(Dataset):\n    \"\"\"\n    Dataset for loading individual DICOM slices for validation.\n    \"\"\"\n\n    def __init__(self, df, image_size=224):\n        self.df = df.reset_index(drop=True)\n        self.image_size = image_size\n\n    def __len__(self):\n        \"\"\"Return number of validation slices.\"\"\"\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        \"\"\"Load one DICOM slice and target label.\"\"\"\n        row = self.df.iloc[idx]\n\n        # Read DICOM metadata and pixel data.\n        ds = pydicom.dcmread(row[\"dcm_path\"])\n\n        # Convert image pixels to float32 for model input.\n        img = ds.pixel_array.astype(np.float32)\n\n        # Normalize each slice independently.\n        img = (img - img.mean()) / (img.std() + 1e-8)\n\n        # Convert from (H, W) to (1, H, W).\n        img_tensor = torch.from_numpy(img).unsqueeze(0)\n\n        # Resize to fixed model input size.\n        img_tensor = img_tensor.unsqueeze(0)\n        img_tensor = F.interpolate(\n            img_tensor,\n            size=(self.image_size, self.image_size),\n            mode=\"bilinear\",\n            align_corners=False\n        )\n        img_tensor = img_tensor.squeeze(0)\n\n        target = torch.tensor(row[\"target\"], dtype=torch.float32)\n\n        return img_tensor, target\n\nclass TinyLumbarCNN(nn.Module):\n    \"\"\"\n    Same architecture used in Step 14.\n    \"\"\"\n\n    def __init__(self):\n        super().__init__()\n\n        self.features = nn.Sequential(\n            nn.Conv2d(1, 16, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n\n            nn.Conv2d(16, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n\n            nn.Conv2d(32, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.AdaptiveAvgPool2d((1, 1))\n        )\n\n        self.classifier = nn.Linear(64, 1)\n\n    def forward(self, x):\n        \"\"\"Forward pass from image tensor to binary logit.\"\"\"\n        x = self.features(x)\n        x = x.flatten(1)\n        x = self.classifier(x).squeeze(1)\n        return x\n\n# Rebuild model and load saved weights.\nmodel = TinyLumbarCNN().to(device)\nmodel.load_state_dict(torch.load(weights_path, map_location=device))\nmodel.eval()\n\n# Build validation DataLoader.\nval_dataset = LumbarSliceBinaryDataset(val_df, image_size=224)\n\nval_loader = DataLoader(\n    val_dataset,\n    batch_size=32,\n    shuffle=False,\n    num_workers=0\n)\n\nall_targets = []\nall_probs = []\nall_preds = []\n\n# Run inference on the full validation split.\nwith torch.no_grad():\n    for images, targets in val_loader:\n        images = images.to(device)\n\n        logits = model(images)\n        probs = torch.sigmoid(logits).cpu().numpy()\n        preds = (probs >= 0.5).astype(int)\n\n        all_probs.extend(probs.tolist())\n        all_preds.extend(preds.tolist())\n        all_targets.extend(targets.numpy().astype(int).tolist())\n\nif len(all_targets) == 0:\n    raise ValueError(\"Validation evaluation produced no predictions.\")\n\n# Compute validation metrics.\naccuracy = accuracy_score(all_targets, all_preds)\nprecision = precision_score(all_targets, all_preds, zero_division=0)\nrecall = recall_score(all_targets, all_preds, zero_division=0)\nf1 = f1_score(all_targets, all_preds, zero_division=0)\n\n# ROC-AUC can fail if the validation set contains only one target class.\ntry:\n    roc_auc = roc_auc_score(all_targets, all_probs)\nexcept ValueError:\n    roc_auc = None\n\ncm = confusion_matrix(all_targets, all_preds)\n\nmetrics = {\n    \"accuracy\": accuracy,\n    \"precision\": precision,\n    \"recall\": recall,\n    \"f1\": f1,\n    \"roc_auc\": roc_auc,\n    \"num_validation_slices\": len(all_targets),\n    \"positive_rate_targets\": float(np.mean(all_targets)),\n    \"positive_rate_predictions\": float(np.mean(all_preds))\n}\n\nprint(\"\\n=== Validation metrics ===\")\nfor key, value in metrics.items():\n    print(f\"{key}: {value}\")\n\nprint(\"\\n=== Confusion matrix ===\")\nprint(cm)\n\n# Save compact metrics table.\nmetrics_df = pd.DataFrame([metrics])\nmetrics_path = os.path.join(WORKING_DIR, \"validation_metrics_200_smoke_test.csv\")\nmetrics_df.to_csv(metrics_path, index=False)\n\n# Save one prediction row per validation DICOM slice.\npredictions_df = val_df.copy()\npredictions_df[\"probability_positive\"] = all_probs\npredictions_df[\"prediction\"] = all_preds\n\npredictions_path = os.path.join(WORKING_DIR, \"validation_predictions_200_smoke_test.csv\")\npredictions_df.to_csv(predictions_path, index=False)\n\nprint(\"\\nSaved metrics to:\")\nprint(metrics_path)\n\nprint(\"\\nSaved validation predictions to:\")\nprint(predictions_path)\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T01:42:50.709202Z","iopub.execute_input":"2026-06-16T01:42:50.709819Z","iopub.status.idle":"2026-06-16T01:43:20.240351Z","shell.execute_reply.started":"2026-06-16T01:42:50.709785Z","shell.execute_reply":"2026-06-16T01:43:20.239622Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 17: CREATE EXPERIMENT SUMMARY\n# ============================================================\n# Purpose:\n# Create compact summary files for the first 200-study smoke-test model.\n#\n# This cell records:\n#   1. Dataset subset\n#   2. Model artifact\n#   3. Validation metrics\n#   4. Key limitations\n#   5. Generated output files\n#\n# Outputs:\n#   experiment_summary_200_smoke_test.json\n#   experiment_summary_200_smoke_test.csv\n#\n# This cell does NOT train or evaluate a model.\n#\n# Workflow phase:\n# Phase B — Experiment documentation.\n#\n# Manual accelerator note:\n# GPU is not needed for this cell, but if using the manual workflow,\n# do not change Accelerator settings after Step 14. Keep the same\n# Phase B runtime alive until Steps 15–18 are complete.\n# ============================================================\n\n# Input files from previous steps.\nmetrics_path = os.path.join(WORKING_DIR, \"validation_metrics_200_smoke_test.csv\")\npredictions_path = os.path.join(WORKING_DIR, \"validation_predictions_200_smoke_test.csv\")\nweights_path = os.path.join(WORKING_DIR, \"tiny_lumbar_cnn_smoke_test.pth\")\nmanifest_path = os.path.join(WORKING_DIR, \"project_manifest_200.json\")\n\n# These files are required for this experiment summary.\nrequired_files = [\n    metrics_path,\n    predictions_path,\n    weights_path\n]\n\nfor path in required_files:\n    if not os.path.exists(path):\n        raise FileNotFoundError(\n            f\"Missing required file: {path}. \"\n            \"If using the manual workflow, keep the same Phase B runtime alive \"\n            \"or rerun the needed earlier steps.\"\n        )\n\n# Load validation metrics from Step 16.\nmetrics_df = pd.read_csv(metrics_path)\n\nif metrics_df.empty:\n    raise ValueError(\n        \"validation_metrics_200_smoke_test.csv is empty. \"\n        \"Rerun Step 16 before Step 17.\"\n    )\n\nmetrics = metrics_df.iloc[0].to_dict()\n\n# Load manifest if available.\n# This links the model summary back to the dataset-construction summary.\nmanifest = None\n\nif os.path.exists(manifest_path):\n    with open(manifest_path, \"r\") as f:\n        manifest = json.load(f)\nelse:\n    print(\n        \"Warning: project_manifest_200.json not found. \"\n        \"Continuing without manifest metadata.\"\n    )\n\n# Record file sizes for artifact tracking.\nweights_size_mb = os.path.getsize(weights_path) / (1024 * 1024)\npredictions_size_kb = os.path.getsize(predictions_path) / 1024\n\n# Build experiment summary dictionary.\nexperiment_summary = {\n    \"experiment_name\": \"tiny_lumbar_cnn_200_smoke_test\",\n    \"project_name\": \"lumbar_vibe\",\n    \"dataset\": \"RSNA 2024 Lumbar Spine Degenerative Classification\",\n    \"subset\": {\n        \"num_studies\": SUBSET_SIZE,\n        \"target\": \"any moderate/severe spinal canal stenosis\",\n        \"label_mapping\": {\n            \"Normal/Mild\": 0,\n            \"Moderate\": 1,\n            \"Severe\": 1\n        }\n    },\n    \"model\": {\n        \"architecture\": \"TinyLumbarCNN\",\n        \"weights_file\": \"tiny_lumbar_cnn_smoke_test.pth\",\n        \"weights_size_mb\": round(weights_size_mb, 3),\n        \"training_type\": \"1-epoch smoke test\"\n    },\n    \"validation_metrics\": metrics,\n    \"outputs\": {\n        \"metrics_csv\": \"validation_metrics_200_smoke_test.csv\",\n        \"predictions_csv\": \"validation_predictions_200_smoke_test.csv\",\n        \"predictions_size_kb\": round(predictions_size_kb, 3)\n    },\n    \"limitations\": [\n        \"Smoke-test model only\",\n        \"Not clinically meaningful\",\n        \"Study-level labels assigned to individual slices\",\n        \"Tiny CNN architecture\",\n        \"Short training duration\"\n    ],\n    \"conclusion\": (\n        \"End-to-end training and inference pipeline verified. \"\n        \"Ready for improved baseline experiments.\"\n    )\n}\n\n# Include dataset-construction manifest information if available.\nif manifest is not None:\n    experiment_summary[\"dataset_manifest\"] = {\n        \"subset_name\": manifest.get(\"subset_name\"),\n        \"subset_size_studies\": manifest.get(\"subset_size_studies\"),\n        \"num_selected_series\": manifest.get(\"num_selected_series\"),\n        \"num_selected_dicom_images\": manifest.get(\"num_selected_dicom_images\"),\n        \"core_sequences\": manifest.get(\"core_sequences\")\n    }\n\n# Save JSON summary.\njson_output_path = os.path.join(\n    WORKING_DIR,\n    \"experiment_summary_200_smoke_test.json\"\n)\n\nwith open(json_output_path, \"w\") as f:\n    json.dump(experiment_summary, f, indent=2)\n\n# Save flat CSV summary for quick spreadsheet-style review.\ncsv_output_path = os.path.join(\n    WORKING_DIR,\n    \"experiment_summary_200_smoke_test.csv\"\n)\n\npd.DataFrame([{\n    \"experiment_name\": experiment_summary[\"experiment_name\"],\n    \"num_studies\": experiment_summary[\"subset\"][\"num_studies\"],\n    \"model_architecture\": experiment_summary[\"model\"][\"architecture\"],\n    \"weights_file\": experiment_summary[\"model\"][\"weights_file\"],\n    \"accuracy\": metrics.get(\"accuracy\"),\n    \"precision\": metrics.get(\"precision\"),\n    \"recall\": metrics.get(\"recall\"),\n    \"f1\": metrics.get(\"f1\"),\n    \"roc_auc\": metrics.get(\"roc_auc\"),\n    \"positive_rate_targets\": metrics.get(\"positive_rate_targets\"),\n    \"positive_rate_predictions\": metrics.get(\"positive_rate_predictions\"),\n    \"conclusion\": experiment_summary[\"conclusion\"]\n}]).to_csv(csv_output_path, index=False)\n\nprint(\"Saved experiment summary JSON to:\")\nprint(json_output_path)\n\nprint(\"\\nSaved experiment summary CSV to:\")\nprint(csv_output_path)\n\nprint(\"\\nSummary preview:\")\nprint(json.dumps(experiment_summary, indent=2))\n\nprint(\"\\nFiles currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T02:07:03.374621Z","iopub.execute_input":"2026-06-16T02:07:03.375250Z","iopub.status.idle":"2026-06-16T02:07:03.391924Z","shell.execute_reply.started":"2026-06-16T02:07:03.375216Z","shell.execute_reply":"2026-06-16T02:07:03.391211Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# LUMBAR VIBE 01 — STEP 18: MILESTONE CHECK\n# ============================================================\n# Purpose:\n# Verify which expected 200-study smoke-test artifacts currently exist\n# in /kaggle/working.\n#\n# This cell groups artifacts by role:\n#   1. Phase A dataset-construction artifacts\n#   2. Phase B training/evaluation artifacts\n#   3. Documentation and summary artifacts\n#\n# Output:\n#   milestone_check_200_smoke_test.csv\n#\n# Workflow phase:\n# Phase B — Milestone verification.\n#\n# Manual accelerator note:\n# GPU is not needed for this cell, but if using the manual workflow,\n# do not change Accelerator settings before running this check.\n# Run this before stopping or resetting the active Phase B runtime.\n# ============================================================\n\nartifact_groups = {\n    \"phase_a_dataset_construction\": [\n        \"study_metadata_summary.csv\",\n        \"selected_studies_200.csv\",\n        \"selected_series_200.csv\",\n        \"selected_image_index_200.csv\",\n        \"selected_labels_200.csv\",\n        \"selected_label_distribution_200.csv\",\n        \"project_manifest_200.json\",\n        \"training_index_200_spinal_canal_binary.csv\",\n        \"training_index_200_spinal_canal_binary_split.csv\",\n    ],\n    \"phase_b_training_evaluation\": [\n        \"tiny_lumbar_cnn_smoke_test.pth\",\n        \"inference_smoke_test_results.csv\",\n        \"validation_metrics_200_smoke_test.csv\",\n        \"validation_predictions_200_smoke_test.csv\",\n    ],\n    \"experiment_summary\": [\n        \"experiment_summary_200_smoke_test.json\",\n        \"experiment_summary_200_smoke_test.csv\",\n    ]\n}\n\nprint(\"Files currently in /kaggle/working:\")\nprint(os.listdir(WORKING_DIR))\n\nprint(\"\\n=== Milestone artifact check ===\")\n\nrows = []\n\nfor group_name, filenames in artifact_groups.items():\n    for filename in filenames:\n        path = os.path.join(WORKING_DIR, filename)\n        exists = os.path.exists(path)\n        size_kb = os.path.getsize(path) / 1024 if exists else None\n\n        rows.append({\n            \"artifact_group\": group_name,\n            \"filename\": filename,\n            \"exists\": exists,\n            \"size_kb\": round(size_kb, 2) if size_kb is not None else None\n        })\n\ncheck_df = pd.DataFrame(rows)\n\ndisplay(check_df)\n\nprint(\"\\n=== Missing files by artifact group ===\")\n\nmissing_df = check_df[check_df[\"exists\"] == False]\n\nif missing_df.empty:\n    print(\"None. 200-study smoke-test milestone is complete in this session.\")\nelse:\n    display(missing_df)\n\n# Save the milestone check itself to /kaggle/working.\noutput_path = os.path.join(WORKING_DIR, \"milestone_check_200_smoke_test.csv\")\ncheck_df.to_csv(output_path, index=False)\n\nprint(\"\\nSaved milestone check to:\")\nprint(output_path)\n\nprint(\"\\nReminder:\")\nprint(\"Download important outputs before stopping or resetting the runtime.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-16T02:18:53.394486Z","iopub.execute_input":"2026-06-16T02:18:53.394767Z","iopub.status.idle":"2026-06-16T02:18:53.415590Z","shell.execute_reply.started":"2026-06-16T02:18:53.394742Z","shell.execute_reply":"2026-06-16T02:18:53.414808Z"}},"outputs":[],"execution_count":null}]}