{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":118534,"databundleVersionId":14898831,"sourceType":"competition"},{"sourceId":10338,"databundleVersionId":862042,"sourceType":"competition"}],"dockerImageVersionId":31259,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🏥 MedGemma 1.5: Longitudinal X-Ray Analysis\n\n### 👋 Introduction\nMost Medical AI models analyze a single image. But real doctors compare **today's scan** vs. **last month's scan** to see if a patient is getting better or worse.\n\nThis notebook demonstrates **Google's MedGemma 1.5 (VLM)**, which has native \"Longitudinal\" capabilities. We will:\n1. Load `medgemma-1.5-4b-it`.\n2. Simulate a patient timeline (Day 0 vs Day 30) using the RSNA dataset.\n3. Generate an automated **Progress Report** comparing the two images.\n\n---","metadata":{}},{"cell_type":"markdown","source":"## Why this notebook?\n\nMedGemma 1.5 is a powerful medical vision-language model, but running it\noutside Google’s infrastructure can be surprisingly tricky.\n\nThis notebook exists to:\n- Demonstrate **correct multimodal usage** (images + text)\n- Show **longitudinal medical reasoning** (baseline vs follow-up X-rays)\n- Document **real-world deployment issues** (GPU vs CPU stability)\n- Provide a **working, reproducible reference** for Kaggle users","metadata":{}},{"cell_type":"code","source":"# Install dependencies\n!pip install -q -U transformers accelerate\n!pip install -q pydicom pillow","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-15T20:27:48.978004Z","iopub.execute_input":"2026-01-15T20:27:48.978837Z","iopub.status.idle":"2026-01-15T20:28:20.206801Z","shell.execute_reply.started":"2026-01-15T20:27:48.978799Z","shell.execute_reply":"2026-01-15T20:28:20.205154Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport torch\nimport pydicom\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nfrom transformers import AutoProcessor, AutoModelForImageTextToText\nfrom huggingface_hub import login\nfrom kaggle_secrets import UserSecretsClient\n\n# 1. Login (Make sure to add 'HF_TOKEN' in Add-ons -> Secrets)\nuser_secrets = UserSecretsClient()\nhf_token = user_secrets.get_secret(\"HF_TOKEN\")\nlogin(token=hf_token)\n\n# 2. Load MedGemma on CPU (Max Stability)\n# I used CPU to avoid T4 GPU Float16 instability with this specific model architecture (I had lots of problem running this on GPU)\nMODEL_ID = \"google/medgemma-1.5-4b-it\"\n\nprint(f\"⏳ Loading {MODEL_ID}... (This uses System RAM)\")\nprocessor = AutoProcessor.from_pretrained(MODEL_ID)\nmodel = AutoModelForImageTextToText.from_pretrained(\n    MODEL_ID,\n    device_map=\"cpu\",\n    torch_dtype=torch.float32,\n    low_cpu_mem_usage=True\n)\nprint(\"✅ Model Loaded Successfully!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-15T20:28:20.208981Z","iopub.execute_input":"2026-01-15T20:28:20.210060Z","iopub.status.idle":"2026-01-15T20:30:36.893183Z","shell.execute_reply.started":"2026-01-15T20:28:20.209999Z","shell.execute_reply":"2026-01-15T20:30:36.890498Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Configuration ---\nRSNA_DIR = \"/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_train_images\"\nCSV_PATH = \"/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_train_labels.csv\"\n\n# Load Labels\ndf = pd.read_csv(CSV_PATH)\n\n# Select one Normal patient and one Pneumonia patient to simulate progression\n# (You can change these indices to test different cases)\nnormal_pid = df[df['Target'] == 0].iloc[0]['patientId']\npneumonia_pid = df[df['Target'] == 1].iloc[5]['patientId']\n\ndef get_image(pid):\n    dcm = pydicom.dcmread(f\"{RSNA_DIR}/{pid}.dcm\")\n    return dcm.pixel_array\n\n# Load Arrays\nimg1_arr = get_image(normal_pid)\nimg2_arr = get_image(pneumonia_pid)\n\n# Plotting\nfig, axes = plt.subplots(1, 2, figsize=(12, 6))\naxes[0].imshow(img1_arr, cmap='gray')\naxes[0].set_title(\"Visit 1: Baseline (Last Month)\\nStatus: Normal\")\naxes[0].axis('off')\n\naxes[1].imshow(img2_arr, cmap='gray')\naxes[1].set_title(\"Visit 2: Follow-Up (Today)\\nStatus: Opacity Visible\")\naxes[1].axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-15T20:30:36.898724Z","iopub.execute_input":"2026-01-15T20:30:36.901092Z","iopub.status.idle":"2026-01-15T20:30:38.342733Z","shell.execute_reply.started":"2026-01-15T20:30:36.901034Z","shell.execute_reply":"2026-01-15T20:30:38.341468Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⚠️ GPU Inference Notes (Important)\n\nDuring experimentation, MedGemma 1.5 showed **numerical instability**\nwhen running multimodal inference on Kaggle GPUs (T4/L4).\n\nObserved issues included:\n- NaN / Inf probabilities during sampling\n- CUDA device-side assertion failures\n- Immediate EOS termination during greedy decoding\n\nThese issues persisted even with:\n- 4-bit NF4 quantization\n- Different decoding strategies\n\nFor this reason, **CPU inference is used for final results** to ensure\nnumerical stability and reproducibility.","metadata":{}},{"cell_type":"code","source":"# Helper to convert DICOM array to RGB PIL Image\ndef to_pil(img_array):\n    img_array = img_array.astype(float)\n    if img_array.max() == img_array.min():\n        img_array[:] = 0\n    else:\n        img_array = (img_array - img_array.min()) / (img_array.max() - img_array.min()) * 255.0\n    return Image.fromarray(img_array.astype(\"uint8\")).convert(\"RGB\")\n\nimage_1_pil = to_pil(img1_arr)\nimage_2_pil = to_pil(img2_arr)\n\n# The Doctor's Prompt\nmessages = [\n    {\n        \"role\": \"user\",\n        \"content\": [\n            {\"type\": \"image\"},\n            {\"type\": \"image\"}, \n            {\n                \"type\": \"text\",\n                \"text\": (\n                    \"I am providing two Chest X-rays for the same patient. \"\n                    \"The first image is the baseline scan from last month. \"\n                    \"The second image is the new scan from today. \"\n                    \"Compare the second image to the first one. \"\n                    \"Describe if there are any new opacities or signs of pneumonia progression.\"\n                )\n            }\n        ]\n    }\n]\n\n# Process Inputs\nprompt = processor.apply_chat_template(messages, add_generation_prompt=True)\ninputs = processor(text=prompt, images=[image_1_pil, image_2_pil], return_tensors=\"pt\", padding=True)\n\n# Generate\nprint(\"🤖 Dr. MedGemma is analyzing... (Allow ~ 3-5 minutes on CPU)\")\ngenerated_ids = model.generate(\n    **inputs,\n    max_new_tokens=350,\n    min_new_tokens=50,     # Force detailed answer\n    do_sample=True,        # Enable creativity\n    temperature=0.4,\n    repetition_penalty=1.1\n)\n\n# Decode output\ninput_len = inputs[\"input_ids\"].shape[-1]\noutput_tokens = generated_ids[0][input_len:]\nreport = processor.decode(output_tokens, skip_special_tokens=True)\n\n# Markdown Output\nfrom IPython.display import Markdown\ndisplay(Markdown(f\"### 📝 Generated Longitudinal Report\\n\\n{report}\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-15T20:30:38.344790Z","iopub.execute_input":"2026-01-15T20:30:38.345118Z","iopub.status.idle":"2026-01-15T20:37:50.837944Z","shell.execute_reply.started":"2026-01-15T20:30:38.345088Z","shell.execute_reply":"2026-01-15T20:37:50.836783Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## What MedGemma Does Well (and What It Doesn’t)\n\n**Strengths**\n- Produces cautious, clinically aligned language\n- Performs explicit comparison between timepoints\n- Uses appropriate uncertainty and disclaimers\n\n**Limitations**\n- Not a diagnostic tool\n- Sensitive to hardware and numerical precision\n- Conservative outputs may miss subtle findings\n\nMedGemma is best suited for **assistive reporting and educational use**,\nnot for automated diagnosis.\n","metadata":{}},{"cell_type":"markdown","source":"---\n## 🔑 Key Takeaways\n\n- Multimodal medical models can behave very differently depending on\n  hardware and numerical precision.\n- CPU inference, while slower, can be **more reliable** for complex\n  vision–language models like MedGemma.\n- Explicitly framing tasks as **longitudinal comparisons** (baseline vs\n  follow-up) leads to more clinically useful outputs.\n- Always verify that image tokens are actually being consumed — a model\n  running without images can still appear to “work”.\n\nThis notebook focuses on **practical lessons learned**, not just a\nsuccessful demo.\n","metadata":{}},{"cell_type":"markdown","source":"---\n## Closing Thoughts\n\nThis notebook started as an attempt to run MedGemma on chest X-rays and\nended up being a useful lesson in **real-world model deployment**.\n\nIf this saves you time, confusion, or a few failed GPU runs, then it’s\ndone its job 🙂\n\nFeedback and suggestions are always welcome.\n\n---","metadata":{}}]}