{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"cells":[{"cell_type":"markdown","source":"# Competition: physionet-ecg-image-digitization\n\n**Generated by Alexandria Research Assistant**\n\n**Dataset:** physionet-ecg-image-digitization\n\n**Task:** competition - kaggle-competition\n\n---\n\n⚠️ **Note:** This notebook contains Alexandria markers (lines starting with `# ⚠️ ALEXANDRIA MARKER`) at the top of each code cell. These markers enable the 'Sync from Kaggle' feature to track cell outputs. Please do not delete them.","metadata":{}},{"cell_type":"markdown","source":"## Setup & Imports\n\nInstall and import all required libraries for image processing, deep learning, and data manipulation tailored to ECG digitization.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_1_START===\")\n\n# Setup & Imports\n# Install required libraries for image processing, deep learning, and ECG digitization\n\n# Install missing packages (Kaggle environments may already have most, but ensure all are present)\n!pip install albumentations --quiet\n!pip install scikit-image --quiet\n!pip install ecgmentations --quiet\n!pip install ecglib --quiet\n!pip install git+https://github.com/alphanumericslab/ecg-image-kit.git --quiet\n\n# Import libraries\nimport os\nimport random\n\nimport numpy as np\nimport pandas as pd\nimport cv2\nimport torch\nimport torchvision\nimport matplotlib.pyplot as plt\nimport albumentations as A\nimport skimage\nimport scipy\nfrom tqdm import tqdm\n\n# Specialized ECG libraries\ntry:\n    import ecgmentations\nexcept ImportError:\n    print(\"Warning: ecgmentations not found.\")\n\ntry:\n    import ecglib\nexcept ImportError:\n    print(\"Warning: ecglib not found.\")\n\ntry:\n    import ecg_image_kit\nexcept ImportError:\n    print(\"Warning: ECG-Image-Kit not found.\")\n\n# Set random seeds for reproducibility\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\ntorch.cuda.manual_seed_all(SEED)\nos.environ[\"PYTHONHASHSEED\"] = str(SEED)\n\n# Configure device for PyTorch\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"Using device: {device}\")\n\n# Dataset path (use EXACT path as required)\nDATA_PATH = \"/kaggle/input/physionet-ecg-image-digitization\"\n\n# Detect available files in dataset directory\nexpected_files = [\n    \"train.csv\",\n    \"test.csv\",\n    \"sample_submission.csv\",\n    \"train.csv.zip\",\n    \"test.csv.zip\"\n]\nprint(\"Checking dataset files in:\", DATA_PATH)\nfor fname in expected_files:\n    fpath = os.path.join(DATA_PATH, fname)\n    if os.path.exists(fpath):\n        print(f\"Found: {fname}\")\n    else:\n        print(f\"Missing: {fname}\")\n\nprint(\"Setup complete. All libraries imported and dataset files checked.\")","outputs":[],"cell_number":1,"version":1,"status":"generated","created_at":"2025-11-18T22:53:11.710083+00:00","metadata":{},"execution_count":null},{"cell_type":"markdown","source":"## Data Loading\n\nDetect and load the train, test, and sample submission files from the competition dataset directory.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_2_START===\")\n\n# ===ALEXANDRIA_CELL_2_START===\n# Data Loading\n# Detect and load the train, test, and sample submission files from the competition dataset directory.\n# Unzip files if necessary and preview the data.\n\nimport zipfile\n\n# Define dataset path (as required)\nDATA_PATH = \"/kaggle/input/physionet-ecg-image-digitization\"\n\n# Helper function to check file existence\ndef file_exists(filename):\n    return os.path.exists(os.path.join(DATA_PATH, filename))\n\n# Unzip train.csv.zip and test.csv.zip if train.csv or test.csv are missing\nif not file_exists(\"train.csv\") and file_exists(\"train.csv.zip\"):\n    print(\"Extracting train.csv.zip...\")\n    with zipfile.ZipFile(os.path.join(DATA_PATH, \"train.csv.zip\"), \"r\") as z:\n        z.extract(\"train.csv\", DATA_PATH)\n    print(\"train.csv extracted.\")\n\nif not file_exists(\"test.csv\") and file_exists(\"test.csv.zip\"):\n    print(\"Extracting test.csv.zip...\")\n    with zipfile.ZipFile(os.path.join(DATA_PATH, \"test.csv.zip\"), \"r\") as z:\n        z.extract(\"test.csv\", DATA_PATH)\n    print(\"test.csv extracted.\")\n\n# Load train.csv\ntrain_csv_path = os.path.join(DATA_PATH, \"train.csv\")\nif file_exists(\"train.csv\"):\n    print(f\"Loading: {train_csv_path}\")\n    train_df = pd.read_csv(train_csv_path)\n    print(\"train.csv loaded. Shape:\", train_df.shape)\nelse:\n    raise FileNotFoundError(\"train.csv not found in dataset directory.\")\n\n# Load test.csv\ntest_csv_path = os.path.join(DATA_PATH, \"test.csv\")\nif file_exists(\"test.csv\"):\n    print(f\"Loading: {test_csv_path}\")\n    test_df = pd.read_csv(test_csv_path)\n    print(\"test.csv loaded. Shape:\", test_df.shape)\nelse:\n    raise FileNotFoundError(\"test.csv not found in dataset directory.\")\n\n# Preview sample_submission.csv for output format\nsample_sub_path = os.path.join(DATA_PATH, \"sample_submission.csv\")\nif file_exists(\"sample_submission.csv\"):\n    print(f\"Previewing: {sample_sub_path}\")\n    sample_submission_df = pd.read_csv(sample_sub_path)\n    display_cols = sample_submission_df.columns.tolist()\n    print(\"sample_submission.csv columns:\", display_cols)\n    print(sample_submission_df.head())\nelse:\n    print(\"sample_submission.csv not found. Please check dataset directory.\")\n\nprint(\"Data loading complete.\")","outputs":[],"cell_number":2,"version":1,"status":"generated","created_at":"2025-11-18T22:53:18.974746+00:00","metadata":{},"execution_count":null},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)\n\nAnalyze the structure and content of the ECG image dataset and associated metadata.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_3_START===\")\n\n# ===ALEXANDRIA_CELL_3_START===\n# Exploratory Data Analysis (EDA) for ECG Image Dataset and Metadata\n\n# Print cell title and description\nprint(\"=== Exploratory Data Analysis (EDA) ===\")\nprint(\"Analyze the structure and content of the ECG image dataset and associated metadata.\")\n\n# Check train_df structure\nprint(\"\\nTrain DataFrame Info:\")\ntrain_df.info()\n\nprint(\"\\nTrain DataFrame Head:\")\nprint(train_df.head())\n\n# Summarize metadata columns\nprint(\"\\nTrain DataFrame Columns:\")\nprint(train_df.columns.tolist())\n\nprint(\"\\nTrain DataFrame Description (Numerical Columns):\")\nprint(train_df.describe(include=[np.number]))\n\nprint(\"\\nTrain DataFrame Description (All Columns):\")\nprint(train_df.describe(include='all'))\n\n# Check for missing values\nprint(\"\\nMissing Values per Column:\")\nprint(train_df.isnull().sum())\n\n# Check for anomalous values (e.g., negative signal lengths, zero/negative image resolution)\nif 'signal_length' in train_df.columns:\n    print(\"\\nSignal Length Value Counts (including anomalies):\")\n    print(train_df['signal_length'].value_counts(dropna=False))\n    print(\"Signal Length < 0:\", (train_df['signal_length'] < 0).sum())\nif 'image_width' in train_df.columns and 'image_height' in train_df.columns:\n    print(\"\\nImage Resolution Value Counts (including anomalies):\")\n    print(\"Image Width <= 0:\", (train_df['image_width'] <= 0).sum())\n    print(\"Image Height <= 0:\", (train_df['image_height'] <= 0).sum())\n\n# Plot distribution of key variables\nimport seaborn as sns\n\nplt.figure(figsize=(14, 5))\nif 'signal_length' in train_df.columns:\n    plt.subplot(1, 2, 1)\n    sns.histplot(train_df['signal_length'].dropna(), bins=50, kde=True, color='royalblue')\n    plt.title('Distribution of Signal Length')\n    plt.xlabel('Signal Length')\n    plt.ylabel('Count')\n\nif 'image_width' in train_df.columns and 'image_height' in train_df.columns:\n    plt.subplot(1, 2, 2)\n    sns.scatterplot(\n        x=train_df['image_width'],\n        y=train_df['image_height'],\n        alpha=0.3,\n        color='darkorange'\n    )\n    plt.title('Image Resolution Scatterplot')\n    plt.xlabel('Image Width')\n    plt.ylabel('Image Height')\n\nplt.tight_layout()\nplt.show()\n\n# Visualize sample ECG images from train set\n# Detect image file column (commonly 'image_id' or similar)\nimage_col_candidates = [col for col in train_df.columns if 'image' in col and 'id' in col]\nif image_col_candidates:\n    image_col = image_col_candidates\n    print(f\"\\nDetected image column: {image_col}\")\nelse:\n    image_col = None\n    print(\"\\nNo image column detected. Skipping image visualization.\")\n\n# Helper: Find image file extensions in DATA_PATH\ndef find_image_file(image_id, exts=['.png', '.jpg', '.jpeg', '.bmp']):\n    for ext in exts:\n        fpath = os.path.join(DATA_PATH, f\"{image_id}{ext}\")\n        if os.path.exists(fpath):\n            return fpath\n    return None\n\n# Visualize up to 6 random ECG images\nif image_col:\n    sample_images = train_df[image_col].dropna().unique()\n    n_samples = min(6, len(sample_images))\n    sample_ids = np.random.choice(sample_images, n_samples, replace=False)\n    plt.figure(figsize=(16, 8))\n    for i, img_id in enumerate(sample_ids):\n        img_path = find_image_file(img_id)\n        if img_path:\n            img = cv2.imread(img_path, cv2.IMREAD_GRAYSCALE)\n            plt.subplot(2, 3, i+1)\n            plt.imshow(img, cmap='gray')\n            plt.title(f\"{img_id}\")\n            plt.axis('off')\n        else:\n            print(f\"Image file not found for image_id: {img_id}\")\n    plt.suptitle(\"Sample ECG Images from Train Set\")\n    plt.tight_layout(rect=[0, 0, 1, 0.95])\n    plt.show()\nelse:\n    print(\"No image column detected; cannot visualize ECG images.\")\n\n# Summarize patient info columns if present\npatient_cols = [col for col in train_df.columns if 'patient' in col or 'age' in col or 'sex' in col]\nif patient_cols:\n    print(\"\\nPatient Info Columns Summary:\")\n    print(train_df[patient_cols].describe(include='all'))\nelse:\n    print(\"\\nNo patient info columns detected.\")\n\nprint(\"\\nEDA complete.\")","outputs":[],"cell_number":3,"version":1,"status":"generated","created_at":"2025-11-18T22:53:27.907513+00:00","metadata":{},"execution_count":null},{"cell_type":"markdown","source":"## Preprocessing\n\nPrepare images and metadata for model input, including cleaning and normalization.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_4_START===\")\n\nprint(\"===ALEXANDRIA_CELL_4_START===\")\n\n# === Preprocessing ===\n# Prepare images and metadata for model input: cleaning, normalization, denoising, contrast enhancement, ROI extraction, and missing/corrupted data handling.\n\n# 1. Detect image column and available image files\nimage_col_candidates = [col for col in train_df.columns if 'image' in col and 'id' in col]\nif image_col_candidates:\n    image_col = image_col_candidates\n    print(f\"Using image column: {image_col}\")\nelse:\n    raise ValueError(\"No image column found in train_df.\")\n\n# Helper: Find image file path for a given image_id\ndef find_image_file(image_id, exts=['.png', '.jpg', '.jpeg', '.bmp']):\n    for ext in exts:\n        fpath = os.path.join(DATA_PATH, f\"{image_id}{ext}\")\n        if os.path.exists(fpath):\n            return fpath\n    return None\n\n# 2. Preprocessing functions\n\ndef preprocess_ecg_image(img, target_size=(512, 256)):\n    # Convert to grayscale if needed\n    if len(img.shape) == 3:\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n    # Denoising (median filter)\n    img = cv2.medianBlur(img, 3)\n    # Contrast enhancement (CLAHE)\n    clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8))\n    img = clahe.apply(img)\n    # Normalize to [0,1]\n    img = img.astype(np.float32) / 255.0\n    # Resize to target size\n    img = cv2.resize(img, target_size, interpolation=cv2.INTER_AREA)\n    return img\n\ndef extract_ecg_trace(img):\n    # Simple binarization and masking for trace extraction\n    # Assumes img is float32 in [0,1]\n    thresh = 0.7  # Empirically chosen\n    binary = (img < thresh).astype(np.uint8) * 255\n    # Morphological operations to clean up\n    kernel = np.ones((2,2), np.uint8)\n    binary = cv2.morphologyEx(binary, cv2.MORPH_OPEN, kernel)\n    return binary\n\n# 3. Handle missing/corrupted data in metadata\ndef clean_metadata(df):\n    df = df.copy()\n    # Fill missing numerical columns with median\n    for col in df.select_dtypes(include=[np.number]).columns:\n        if df[col].isnull().any():\n            median_val = df[col].median()\n            df[col] = df[col].fillna(median_val)\n    # Fill missing categorical columns with mode\n    for col in df.select_dtypes(include=['object']).columns:\n        if df[col].isnull().any():\n            mode_val = df[col].mode().iloc\n            df[col] = df[col].fillna(mode_val)\n    return df\n\nprint(\"Cleaning metadata...\")\ntrain_df_clean = clean_metadata(train_df)\ntest_df_clean = clean_metadata(test_df)\nprint(\"Metadata cleaning complete.\")\n\n# 4. Preprocess and cache a subset of images for demonstration\npreprocessed_images = {}\ntrace_masks = {}\nmissing_images = []\n\nsample_image_ids = train_df_clean[image_col].dropna().unique()\nn_samples = min(12, len(sample_image_ids))\nsample_ids = np.random.choice(sample_image_ids, n_samples, replace=False)\n\nprint(f\"Preprocessing {n_samples} sample ECG images for demonstration...\")\nfor img_id in tqdm(sample_ids):\n    img_path = find_image_file(img_id)\n    if img_path is None:\n        missing_images.append(img_id)\n        continue\n    img = cv2.imread(img_path, cv2.IMREAD_COLOR)\n    if img is None:\n        missing_images.append(img_id)\n        continue\n    proc_img = preprocess_ecg_image(img)\n    mask = extract_ecg_trace(proc_img)\n    preprocessed_images[img_id] = proc_img\n    trace_masks[img_id] = mask\n\nprint(f\"Preprocessing complete. {len(missing_images)} images missing or unreadable.\")\n\n# 5. Visualize preprocessed images and extracted traces\nplt.figure(figsize=(16, 8))\nfor i, img_id in enumerate(preprocessed_images.keys()):\n    proc_img = preprocessed_images[img_id]\n    mask = trace_masks[img_id]\n    plt.subplot(2, n_samples, i+1)\n    plt.imshow(proc_img, cmap='gray')\n    plt.title(f\"{img_id} (Norm+CLAHE)\")\n    plt.axis('off')\n    plt.subplot(2, n_samples, n_samples+i+1)\n    plt.imshow(mask, cmap='gray')\n    plt.title(f\"{img_id} (Trace Mask)\")\n    plt.axis('off')\nplt.tight_layout()\nplt.show()\n\nif missing_images:\n    print(f\"Warning: {len(missing_images)} sample images could not be found or read.\")\n\nprint(\"Preprocessing pipeline complete. Images and metadata are ready for model input.\")","outputs":[],"cell_number":4,"version":1,"status":"generated","created_at":"2025-11-18T22:53:42.905406+00:00","metadata":{},"execution_count":null},{"cell_type":"markdown","source":"## Feature Engineering\n\nExtract task-specific features from images and metadata to facilitate signal reconstruction.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_5_START===\")\n\nprint(\"===ALEXANDRIA_CELL_5_START===\")\n\n# === Feature Engineering ===\n# Extract task-specific features from ECG images and metadata for signal reconstruction.\n# Operations: edge detection/segmentation, pixel coordinate/time-series extraction, synthetic features, data augmentation.\n\n# 1. Detect image column (from previous cells)\nimage_col_candidates = [col for col in train_df_clean.columns if 'image' in col and 'id' in col]\nif image_col_candidates:\n    image_col = image_col_candidates\n    print(f\"Using image column: {image_col}\")\nelse:\n    raise ValueError(\"No image column found in train_df_clean.\")\n\n# 2. Helper: Find image file path for a given image_id\ndef find_image_file(image_id, exts=['.png', '.jpg', '.jpeg', '.bmp']):\n    for ext in exts:\n        fpath = os.path.join(DATA_PATH, f\"{image_id}{ext}\")\n        if os.path.exists(fpath):\n            return fpath\n    return None\n\n# 3. Edge Detection & Segmentation\nfrom skimage.feature import canny\nfrom skimage import img_as_float\nfrom skimage.morphology import remove_small_objects, binary_opening\n\ndef edge_detect_canny(img, sigma=1.5, low_thresh=0.05, high_thresh=0.2):\n    # img: float32 in [0,1]\n    edges = canny(img, sigma=sigma, low_threshold=low_thresh, high_threshold=high_thresh)\n    # Remove small objects (noise)\n    edges = remove_small_objects(edges, min_size=30)\n    # Morphological opening for clean trace\n    edges = binary_opening(edges)\n    return edges.astype(np.uint8)\n\n# 4. Extract pixel coordinates (trace path) from edge mask\ndef extract_trace_coordinates(edge_mask):\n    # edge_mask: binary mask (0/1)\n    coords = np.column_stack(np.where(edge_mask > 0))\n    # Sort by x (horizontal axis)\n    coords = coords[np.argsort(coords[:,1])]\n    return coords\n\n# 5. Synthetic Features: signal statistics, image quality metrics\ndef compute_signal_stats(trace_coords, img_shape):\n    if len(trace_coords) == 0:\n        return {\n            \"trace_length\": 0,\n            \"trace_height_mean\": np.nan,\n            \"trace_height_std\": np.nan,\n            \"trace_height_range\": np.nan,\n        }\n    heights = trace_coords[:,0]\n    trace_length = len(trace_coords)\n    height_mean = np.mean(heights)\n    height_std = np.std(heights)\n    height_range = np.max(heights) - np.min(heights)\n    return {\n        \"trace_length\": trace_length,\n        \"trace_height_mean\": height_mean,\n        \"trace_height_std\": height_std,\n        \"trace_height_range\": height_range,\n    }\n\ndef compute_image_quality(img):\n    # img: float32 in [0,1]\n    contrast = img.max() - img.min()\n    sharpness = cv2.Laplacian((img*255).astype(np.uint8), cv2.CV_64F).var()\n    return {\n        \"contrast\": contrast,\n        \"sharpness\": sharpness,\n    }\n\n# 6. Data Augmentation: rotation, scaling, translation, brightness/contrast\nAUGMENT = A.Compose([\n    A.Rotate(limit=10, p=0.5, border_mode=cv2.BORDER_REFLECT),\n    A.RandomScale(scale_limit=0.1, p=0.5),\n    A.ShiftScaleRotate(shift_limit=0.05, scale_limit=0.05, rotate_limit=5, p=0.5),\n    A.RandomBrightnessContrast(brightness_limit=0.2, contrast_limit=0.2, p=0.5),\n])\n\ndef augment_image(img):\n    # img: float32 in [0,1], shape HxW\n    img_uint8 = (img*255).astype(np.uint8)\n    aug = AUGMENT(image=img_uint8)\n    img_aug = aug['image'].astype(np.float32) / 255.0\n    return img_aug\n\n# 7. Feature Extraction Pipeline for a single image\ndef extract_features_for_image(img_id):\n    img_path = find_image_file(img_id)\n    if img_path is None:\n        print(f\"Image file not found for image_id: {img_id}\")\n        return None\n    img = cv2.imread(img_path, cv2.IMREAD_COLOR)\n    if img is None:\n        print(f\"Image unreadable for image_id: {img_id}\")\n        return None\n    # Preprocessing (from previous cell)\n    proc_img = preprocess_ecg_image(img)\n    # Edge detection\n    edge_mask = edge_detect_canny(proc_img)\n    # Trace coordinates\n    trace_coords = extract_trace_coordinates(edge_mask)\n    # Signal statistics\n    signal_stats = compute_signal_stats(trace_coords, proc_img.shape)\n    # Image quality metrics\n    img_quality = compute_image_quality(proc_img)\n    # Augmentation (demonstration: single augmented sample)\n    img_aug = augment_image(proc_img)\n    # Augmented edge detection\n    edge_mask_aug = edge_detect_canny(img_aug)\n    trace_coords_aug = extract_trace_coordinates(edge_mask_aug)\n    signal_stats_aug = compute_signal_stats(trace_coords_aug, img_aug.shape)\n    img_quality_aug = compute_image_quality(img_aug)\n    # Combine features\n    features = {\n        \"image_id\": img_id,\n        **signal_stats,\n        **img_quality,\n        \"aug_trace_length\": signal_stats_aug[\"trace_length\"],\n        \"aug_trace_height_mean\": signal_stats_aug[\"trace_height_mean\"],\n        \"aug_contrast\": img_quality_aug[\"contrast\"],\n        \"aug_sharpness\": img_quality_aug[\"sharpness\"],\n    }\n    return features\n\n# 8. Batch Feature Extraction for train set (demo on subset)\nsample_image_ids = train_df_clean[image_col].dropna().unique()\nn_samples = min(24, len(sample_image_ids))\nsample_ids = np.random.choice(sample_image_ids, n_samples, replace=False)\n\nprint(f\"Extracting features for {n_samples} sample ECG images...\")\nfeature_list = []\nfor img_id in tqdm(sample_ids):\n    feats = extract_features_for_image(img_id)\n    if feats is not None:\n        feature_list.append(feats)\n\nfeatures_df = pd.DataFrame(feature_list)\nprint(\"Feature extraction complete. Sample features:\")\nprint(features_df.head())\n\n# 9. Visualize edge detection and trace extraction for a few samples\nplt.figure(figsize=(18, 8))\nfor i, img_id in enumerate(sample_ids[:6]):\n    img_path = find_image_file(img_id)\n    if img_path is None:\n        continue\n    img = cv2.imread(img_path, cv2.IMREAD_COLOR)\n    if img is None:\n        continue\n    proc_img = preprocess_ecg_image(img)\n    edge_mask = edge_detect_canny(proc_img)\n    trace_coords = extract_trace_coordinates(edge_mask)\n    plt.subplot(2, 6, i+1)\n    plt.imshow(proc_img, cmap='gray')\n    if len(trace_coords) > 0:\n        plt.scatter(trace_coords[:,1], trace_coords[:,0], s=1, c='red')\n    plt.title(f\"{img_id} (Trace)\")\n    plt.axis('off')\n    plt.subplot(2, 6, 6+i+1)\n    plt.imshow(edge_mask, cmap='gray')\n    plt.title(f\"{img_id} (Edges)\")\n    plt.axis('off')\nplt.tight_layout()\nplt.show()\n\nprint(\"Feature engineering pipeline complete. Features ready for downstream signal reconstruction/modeling.\")","outputs":[],"cell_number":5,"version":1,"status":"generated","created_at":"2025-11-18T22:53:55.246663+00:00","metadata":{},"execution_count":null},{"cell_type":"markdown","source":"## Model Training\n\nTrain a deep learning model (e.g., CNN, UNet) to digitize ECG signals from images.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_6_START===\")\n\nprint(\"===ALEXANDRIA_CELL_6_START===\")\n\n# === Model Training: ECG Image-to-Signal Digitization ===\n# Train a deep learning model (U-Net) to segment ECG traces from images for digitization.\n\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\nimport albumentations as A\nfrom albumentations.pytorch import ToTensorV2\nimport cv2\nimport numpy as np\nimport os\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\n\n# 1. Detect image column from cleaned train_df\nimage_col_candidates = [col for col in train_df_clean.columns if 'image' in col and 'id' in col]\nif image_col_candidates:\n    image_col = image_col_candidates\n    print(f\"Using image column: {image_col}\")\nelse:\n    raise ValueError(\"No image column found in train_df_clean.\")\n\n# 2. Helper: Find image file path for a given image_id\ndef find_image_file(image_id, exts=['.png', '.jpg', '.jpeg', '.bmp']):\n    for ext in exts:\n        fpath = os.path.join(DATA_PATH, f\"{image_id}{ext}\")\n        if os.path.exists(fpath):\n            return fpath\n    return None\n\n# 3. Dataset class for ECG image segmentation\nclass ECGSegmentationDataset(Dataset):\n    def __init__(self, df, image_col, mask_func, augment=None, target_size=(512,256)):\n        self.df = df.reset_index(drop=True)\n        self.image_col = image_col\n        self.mask_func = mask_func\n        self.augment = augment\n        self.target_size = target_size\n        self.image_ids = self.df[self.image_col].dropna().tolist()\n\n    def __len__(self):\n        return len(self.image_ids)\n\n    def __getitem__(self, idx):\n        img_id = self.image_ids[idx]\n        img_path = find_image_file(img_id)\n        if img_path is None:\n            raise FileNotFoundError(f\"Image file not found for image_id: {img_id}\")\n        img = cv2.imread(img_path, cv2.IMREAD_COLOR)\n        if img is None:\n            raise ValueError(f\"Image unreadable for image_id: {img_id}\")\n        img = preprocess_ecg_image(img, target_size=self.target_size)\n        mask = self.mask_func(img)\n        # Stack image and mask for augmentation\n        augmented = self.augment(image=img, mask=mask) if self.augment else {\"image\": img, \"mask\": mask}\n        image = augmented[\"image\"]\n        mask = augmented[\"mask\"]\n        # Convert to tensor\n        image = torch.tensor(image, dtype=torch.float32).unsqueeze(0)  # (1, H, W)\n        mask = torch.tensor(mask, dtype=torch.float32).unsqueeze(0) / 255.0  # (1, H, W), binary\n        return image, mask\n\n# 4. Data augmentation pipeline\ntrain_aug = A.Compose([\n    A.Rotate(limit=10, p=0.5, border_mode=cv2.BORDER_REFLECT),\n    A.RandomScale(scale_limit=0.1, p=0.5),\n    A.ShiftScaleRotate(shift_limit=0.05, scale_limit=0.05, rotate_limit=5, p=0.5),\n    A.RandomBrightnessContrast(brightness_limit=0.2, contrast_limit=0.2, p=0.5),\n])\n\nval_aug = A.Compose([])  # No augmentation for validation\n\n# 5. Split train/val\nfrom sklearn.model_selection import train_test_split\ntrain_ids, val_ids = train_test_split(\n    train_df_clean[image_col].dropna().unique(),\n    test_size=0.1, random_state=SEED\n)\ntrain_df_seg = train_df_clean[train_df_clean[image_col].isin(train_ids)].reset_index(drop=True)\nval_df_seg = train_df_clean[train_df_clean[image_col].isin(val_ids)].reset_index(drop=True)\n\nprint(f\"Training samples: {len(train_df_seg)}, Validation samples: {len(val_df_seg)}\")\n\n# 6. Create datasets and dataloaders\nBATCH_SIZE = 8\nTARGET_SIZE = (512, 256)\n\ntrain_dataset = ECGSegmentationDataset(\n    train_df_seg, image_col,\n    mask_func=extract_ecg_trace,\n    augment=train_aug,\n    target_size=TARGET_SIZE\n)\nval_dataset = ECGSegmentationDataset(\n    val_df_seg, image_col,\n    mask_func=extract_ecg_trace,\n    augment=val_aug,\n    target_size=TARGET_SIZE\n)\n\ntrain_loader = DataLoader(train_dataset, batch_size=BATCH_SIZE, shuffle=True, num_workers=2, pin_memory=True)\nval_loader = DataLoader(val_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=2, pin_memory=True)\n\n# 7. U-Net Model Definition (simple version)\nclass UNet(nn.Module):\n    def __init__(self, in_channels=1, out_channels=1, features=[32, 64, 128, 256]):\n        super(UNet, self).__init__()\n        self.encoder = nn.ModuleList()\n        self.decoder = nn.ModuleList()\n        # Encoder\n        for feature in features:\n            self.encoder.append(self._block(in_channels, feature))\n            in_channels = feature\n        # Decoder\n        for feature in reversed(features):\n            self.decoder.append(\n                nn.ConvTranspose2d(feature*2, feature, kernel_size=2, stride=2)\n            )\n            self.decoder.append(self._block(feature*2, feature))\n        self.bottleneck = self._block(features[-1], features[-1]*2)\n        self.final_conv = nn.Conv2d(features, out_channels, kernel_size=1)\n\n    def forward(self, x):\n        encs = []\n        for enc in self.encoder:\n            x = enc(x)\n            encs.append(x)\n            x = nn.MaxPool2d(2)(x)\n        x = self.bottleneck(x)\n        for idx in range(0, len(self.decoder), 2):\n            x = self.decoder[idx](x)\n            enc_feat = encs[-(idx//2 + 1)]\n            if x.shape != enc_feat.shape:\n                x = nn.functional.interpolate(x, size=enc_feat.shape[2:])\n            x = torch.cat([enc_feat, x], dim=1)\n            x = self.decoder[idx+1](x)\n        return torch.sigmoid(self.final_conv(x))\n\n    def _block(self, in_channels, out_channels):\n        return nn.Sequential(\n            nn.Conv2d(in_channels, out_channels, 3, padding=1),\n            nn.BatchNorm2d(out_channels),\n            nn.ReLU(inplace=True),\n            nn.Conv2d(out_channels, out_channels, 3, padding=1),\n            nn.BatchNorm2d(out_channels),\n            nn.ReLU(inplace=True),\n        )\n\n# 8. Instantiate model, optimizer, loss\nmodel = UNet(in_channels=1, out_channels=1).to(device)\noptimizer = optim.Adam(model.parameters(), lr=1e-3)\ncriterion = nn.BCELoss()\n\n# 9. Training loop with checkpointing\nNUM_EPOCHS = 20\nbest_val_loss = np.inf\ncheckpoint_path = \"best_unet_ecg.pth\"\n\ntrain_losses = []\nval_losses = []\n\nfor epoch in range(NUM_EPOCHS):\n    model.train()\n    running_loss = 0.0\n    for images, masks in tqdm(train_loader, desc=f\"Epoch {epoch+1}/{NUM_EPOCHS} [Train]\"):\n        images = images.to(device)\n        masks = masks.to(device)\n        optimizer.zero_grad()\n        outputs = model(images)\n        loss = criterion(outputs, masks)\n        loss.backward()\n        optimizer.step()\n        running_loss += loss.item() * images.size(0)\n    avg_train_loss = running_loss / len(train_loader.dataset)\n    train_losses.append(avg_train_loss)\n\n    # Validation\n    model.eval()\n    val_running_loss = 0.0\n    with torch.no_grad():\n        for images, masks in tqdm(val_loader, desc=f\"Epoch {epoch+1}/{NUM_EPOCHS} [Val]\"):\n            images = images.to(device)\n            masks = masks.to(device)\n            outputs = model(images)\n            loss = criterion(outputs, masks)\n            val_running_loss += loss.item() * images.size(0)\n    avg_val_loss = val_running_loss / len(val_loader.dataset)\n    val_losses.append(avg_val_loss)\n\n    print(f\"Epoch {epoch+1}: Train Loss = {avg_train_loss:.4f}, Val Loss = {avg_val_loss:.4f}\")\n\n    # Save best model\n    if avg_val_loss < best_val_loss:\n        best_val_loss = avg_val_loss\n        torch.save(model.state_dict(), checkpoint_path)\n        print(f\"Best model updated. Saved to {checkpoint_path}\")\n\n# 10. Plot training/validation loss curves\nplt.figure(figsize=(8,5))\nplt.plot(range(1, NUM_EPOCHS+1), train_losses, label=\"Train Loss\")\nplt.plot(range(1, NUM_EPOCHS+1), val_losses, label=\"Val Loss\")\nplt.xlabel(\"Epoch\")\nplt.ylabel(\"Loss\")\nplt.title(\"Training & Validation Loss\")\nplt.legend()\nplt.show()\n\nprint(\"Model training complete. Best model saved at:\", checkpoint_path)\n\n# 11. Visualize predictions on validation set\nmodel.load_state_dict(torch.load(checkpoint_path, map_location=device))\nmodel.eval()\nn_vis = min(6, len(val_dataset))\nplt.figure(figsize=(16,8))\nfor i in range(n_vis):\n    image, mask = val_dataset[i]\n    with torch.no_grad():\n        pred = model(image.unsqueeze(0).to(device)).cpu().squeeze().numpy()\n    plt.subplot(3, n_vis, i+1)\n    plt.imshow(image.squeeze().numpy(), cmap='gray')\n    plt.title(\"Input Image\")\n    plt.axis('off')\n    plt.subplot(3, n_vis, n_vis+i+1)\n    plt.imshow(mask.squeeze().numpy(), cmap='gray')\n    plt.title(\"True Mask\")\n    plt.axis('off')\n    plt.subplot(3, n_vis, 2*n_vis+i+1)\n    plt.imshow(pred, cmap='gray')\n    plt.title(\"Predicted Mask\")\n    plt.axis('off')\nplt.tight_layout()\nplt.show()\n\nprint(\"Prediction visualization complete.\")","outputs":[],"cell_number":6,"version":1,"status":"generated","created_at":"2025-11-18T22:54:11.143042+00:00","metadata":{},"execution_count":null},{"cell_type":"markdown","source":"## Evaluation\n\nAssess model performance using competition-specific metrics and visualizations.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_7_START===\")\n\nprint(\"===ALEXANDRIA_CELL_7_START===\")\n\n# === Evaluation ===\n# Assess model performance using competition-specific metrics and visualizations.\n# Predict on validation and test sets, compute evaluation metrics (e.g., RMSE, correlation, custom leaderboard metric),\n# visualize predicted vs. ground truth signals, analyze failure cases and model robustness.\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport torch\nimport os\nfrom tqdm import tqdm\nfrom scipy.stats import pearsonr\n\n# 1. Detect image column and relevant variables from previous cells\nimage_col_candidates = [col for col in train_df_clean.columns if 'image' in col and 'id' in col]\nif image_col_candidates:\n    image_col = image_col_candidates\nelse:\n    raise ValueError(\"No image column found in train_df_clean.\")\n\n# 2. Helper: Find image file path for a given image_id\ndef find_image_file(image_id, exts=['.png', '.jpg', '.jpeg', '.bmp']):\n    for ext in exts:\n        fpath = os.path.join(DATA_PATH, f\"{image_id}{ext}\")\n        if os.path.exists(fpath):\n            return fpath\n    return None\n\n# 3. Helper: Extract predicted ECG trace from image using trained model\ndef predict_ecg_trace_from_image(model, img_path, preprocess_fn, device, target_size=(512,256)):\n    img = cv2.imread(img_path, cv2.IMREAD_COLOR)\n    if img is None:\n        raise ValueError(f\"Image unreadable: {img_path}\")\n    proc_img = preprocess_fn(img, target_size=target_size)\n    tensor_img = torch.tensor(proc_img, dtype=torch.float32).unsqueeze(0).unsqueeze(0).to(device)  # (1,1,H,W)\n    with torch.no_grad():\n        pred_mask = model(tensor_img).cpu().squeeze().numpy()\n    # Post-process: binarize and extract trace coordinates\n    binary_mask = (pred_mask > 0.5).astype(np.uint8)\n    coords = np.column_stack(np.where(binary_mask > 0))\n    if len(coords) == 0:\n        return None\n    # For each column (x), take the median row (y) as the trace\n    x_vals = np.unique(coords[:,1])\n    trace = np.full((target_size[1],), np.nan)\n    for x in x_vals:\n        y_vals = coords[coords[:,1]==x][:,0]\n        trace[x] = np.median(y_vals)\n    # Normalize trace to [0,1]\n    trace_norm = (trace - np.nanmin(trace)) / (np.nanmax(trace) - np.nanmin(trace) + 1e-8)\n    return trace_norm\n\n# 4. Helper: Load ground truth signal for a given image_id (if available)\ndef get_ground_truth_signal(image_id, df, signal_col_candidates=None):\n    # Try to find a column with ground truth signal (e.g., 'signal', 'ecg', etc.)\n    if signal_col_candidates is None:\n        signal_col_candidates = [col for col in df.columns if 'signal' in col or 'ecg' in col]\n    if not signal_col_candidates:\n        return None\n    signal_col = signal_col_candidates\n    row = df[df[image_col]==image_id]\n    if row.empty:\n        return None\n    signal = row.iloc[signal_col]\n    # If signal is a string of comma-separated values, convert to np.array\n    if isinstance(signal, str):\n        signal = np.array([float(x) for x in signal.split(',')])\n    return signal\n\n# 5. Competition metric: SNR (Signal-to-Noise Ratio) in dB\ndef compute_snr_db(y_true, y_pred):\n    # Both y_true and y_pred should be 1D arrays of same length\n    if np.all(np.isnan(y_pred)) or np.all(np.isnan(y_true)):\n        return -384.0  # Minimum score for invalid predictions\n    mask = ~np.isnan(y_true) & ~np.isnan(y_pred)\n    if np.sum(mask) == 0:\n        return -384.0\n    y_true = y_true[mask]\n    y_pred = y_pred[mask]\n    noise = y_true - y_pred\n    signal_power = np.mean(y_true ** 2)\n    noise_power = np.mean(noise ** 2)\n    if noise_power == 0:\n        return 384.0  # Perfect score\n    snr_db = 10 * np.log10(signal_power / noise_power)\n    return np.clip(snr_db, -384.0, 384.0)\n\n# 6. RMSE and Pearson correlation\ndef compute_rmse(y_true, y_pred):\n    mask = ~np.isnan(y_true) & ~np.isnan(y_pred)\n    if np.sum(mask) == 0:\n        return np.nan\n    return np.sqrt(np.mean((y_true[mask] - y_pred[mask]) ** 2))\n\ndef compute_pearson(y_true, y_pred):\n    mask = ~np.isnan(y_true) & ~np.isnan(y_pred)\n    if np.sum(mask) < 2:\n        return np.nan\n    return pearsonr(y_true[mask], y_pred[mask])\n\n# 7. Evaluate on validation set\nprint(\"Evaluating model on validation set...\")\nmodel.load_state_dict(torch.load(\"best_unet_ecg.pth\", map_location=device))\nmodel.eval()\n\nval_image_ids = val_df_seg[image_col].dropna().unique()\nn_eval = min(24, len(val_image_ids))\neval_ids = np.random.choice(val_image_ids, n_eval, replace=False)\n\nresults = []\nfor img_id in tqdm(eval_ids):\n    img_path = find_image_file(img_id)\n    if img_path is None:\n        print(f\"Image file not found: {img_id}\")\n        continue\n    try:\n        pred_trace = predict_ecg_trace_from_image(model, img_path, preprocess_ecg_image, device)\n    except Exception as e:\n        print(f\"Prediction failed for {img_id}: {e}\")\n        continue\n    # Get ground truth signal (if available)\n    gt_signal = get_ground_truth_signal(img_id, val_df_seg)\n    if gt_signal is None or len(gt_signal) != len(pred_trace):\n        # If no ground truth, skip metric computation\n        continue\n    # Resample gt_signal to match pred_trace length if needed\n    if len(gt_signal) != len(pred_trace):\n        gt_signal = np.interp(np.linspace(0, len(gt_signal)-1, len(pred_trace)), np.arange(len(gt_signal)), gt_signal)\n    # Normalize gt_signal to [0,1]\n    gt_signal_norm = (gt_signal - np.nanmin(gt_signal)) / (np.nanmax(gt_signal) - np.nanmin(gt_signal) + 1e-8)\n    snr = compute_snr_db(gt_signal_norm, pred_trace)\n    rmse = compute_rmse(gt_signal_norm, pred_trace)\n    corr = compute_pearson(gt_signal_norm, pred_trace)\n    results.append({\n        \"image_id\": img_id,\n        \"snr_db\": snr,\n        \"rmse\": rmse,\n        \"pearson\": corr\n    })\n\nresults_df = pd.DataFrame(results)\nprint(\"Evaluation complete. Metrics summary:\")\nprint(results_df.describe())\n\n# 8. Visualize predicted vs. ground truth signals for a few samples\nn_vis = min(6, len(results_df))\nplt.figure(figsize=(18, 10))\nfor i, row in enumerate(results_df.head(n_vis).itertuples()):\n    img_id = row.image_id\n    img_path = find_image_file(img_id)\n    pred_trace = predict_ecg_trace_from_image(model, img_path, preprocess_ecg_image, device)\n    gt_signal = get_ground_truth_signal(img_id, val_df_seg)\n    if gt_signal is None or len(gt_signal) != len(pred_trace):\n        continue\n    if len(gt_signal) != len(pred_trace):\n        gt_signal = np.interp(np.linspace(0, len(gt_signal)-1, len(pred_trace)), np.arange(len(gt_signal)), gt_signal)\n    gt_signal_norm = (gt_signal - np.nanmin(gt_signal)) / (np.nanmax(gt_signal) - np.nanmin(gt_signal) + 1e-8)\n    plt.subplot(n_vis, 1, i+1)\n    plt.plot(gt_signal_norm, label=\"Ground Truth\", color='green')\n    plt.plot(pred_trace, label=\"Predicted\", color='red', alpha=0.7)\n    plt.title(f\"Image ID: {img_id} | SNR: {row.snr_db:.2f} dB | RMSE: {row.rmse:.4f} | Pearson: {row.pearson:.3f}\")\n    plt.legend()\n    plt.tight_layout()\nplt.show()\n\n# 9. Analyze failure cases and model robustness\nprint(\"Analyzing failure cases (lowest SNR)...\")\nworst_cases = results_df.nsmallest(3, 'snr_db')\nfor idx, row in worst_cases.iterrows():\n    img_id = row['image_id']\n    img_path = find_image_file(img_id)\n    pred_trace = predict_ecg_trace_from_image(model, img_path, preprocess_ecg_image, device)\n    gt_signal = get_ground_truth_signal(img_id, val_df_seg)\n    if gt_signal is None or len(gt_signal) != len(pred_trace):\n        continue\n    if len(gt_signal) != len(pred_trace):\n        gt_signal = np.interp(np.linspace(0, len(gt_signal)-1, len(pred_trace)), np.arange(len(gt_signal)), gt_signal)\n    gt_signal_norm = (gt_signal - np.nanmin(gt_signal)) / (np.nanmax(gt_signal) - np.nanmin(gt_signal) + 1e-8)\n    plt.figure(figsize=(10,4))\n    plt.plot(gt_signal_norm, label=\"Ground Truth\", color='green')\n    plt.plot(pred_trace, label=\"Predicted\", color='red', alpha=0.7)\n    plt.title(f\"FAILURE CASE | Image ID: {img_id} | SNR: {row['snr_db']:.2f} dB\")\n    plt.legend()\n    plt.tight_layout()\n    plt.show()\n\n# 10. Prepare test set predictions for submission (if required)\nprint(\"Generating predictions for test set...\")\ntest_image_ids = test_df_clean[image_col].dropna().unique()\nsubmission = []\nfor img_id in tqdm(test_image_ids):\n    img_path = find_image_file(img_id)\n    if img_path is None:\n        print(f\"Test image file not found: {img_id}\")\n        continue\n    try:\n        pred_trace = predict_ecg_trace_from_image(model, img_path, preprocess_ecg_image, device)\n    except Exception as e:\n        print(f\"Prediction failed for test image {img_id}: {e}\")\n        continue\n    # Convert prediction to comma-separated string for submission\n    pred_str = \",\".join([f\"{x:.5f}\" if not np.isnan(x) else \"\" for x in pred_trace])\n    submission.append({\n        image_col: img_id,\n        \"ecg_signal\": pred_str\n    })\n\nsubmission_df = pd.DataFrame(submission)\n# If sample_submission.csv exists, align columns\nif 'sample_submission_df' in locals():\n    sub_cols = sample_submission_df.columns.tolist()\n    submission_df = submission_df[sub_cols]\n    print(\"Submission columns aligned to sample_submission.csv.\")\n\n# Save submission file\nsubmission_path = \"submission.csv\"\nsubmission_df.to_csv(submission_path, index=False)\nprint(f\"Test set predictions saved to {submission_path}\")\n\nprint(\"Evaluation pipeline complete. Model performance assessed and predictions generated.\")","outputs":[],"cell_number":7,"version":1,"status":"generated","created_at":"2025-11-18T22:54:37.272637+00:00","metadata":{},"execution_count":null},{"cell_type":"markdown","source":"## Submission & Results\n\nFormat predictions for submission and display final results.","metadata":{}},{"cell_type":"code","source":"# ⚠️ ALEXANDRIA MARKER - DO NOT DELETE (used for syncing outputs from Kaggle)\nprint(\"===ALEXANDRIA_CELL_8_START===\")\n\n# ===ALEXANDRIA_CELL_8_START===\n# Submission & Results\n# Format predictions for submission, save to CSV, display summary, and provide easy re-submission code.\n\nimport os\nimport pandas as pd\nimport numpy as np\n\nprint(\"=== Submission & Results ===\")\nprint(\"Formatting predictions for Kaggle submission and displaying final results.\")\n\n# 1. Detect sample_submission.csv and align columns\nsample_sub_path = os.path.join(DATA_PATH, \"sample_submission.csv\")\nif os.path.exists(sample_sub_path):\n    sample_submission_df = pd.read_csv(sample_sub_path)\n    submission_columns = sample_submission_df.columns.tolist()\n    print(\"Sample submission format detected:\", submission_columns)\nelse:\n    raise FileNotFoundError(\"sample_submission.csv not found in dataset directory.\")\n\n# 2. Prepare submission DataFrame\n# Use 'submission_df' from previous cell if available, else create from test predictions\nif 'submission_df' in locals():\n    sub_df = submission_df.copy()\nelse:\n    raise ValueError(\"submission_df not found. Please run previous prediction cell to generate test predictions.\")\n\n# 3. Align columns to sample_submission.csv\nmissing_cols = [col for col in submission_columns if col not in sub_df.columns]\nif missing_cols:\n    print(\"Adding missing columns to submission:\", missing_cols)\n    for col in missing_cols:\n        sub_df[col] = \"\"\nsub_df = sub_df[submission_columns]\n\n# 4. Save submission file for Kaggle upload\nsubmission_path = \"submission.csv\"\nsub_df.to_csv(submission_path, index=False)\nprint(f\"Submission file saved: {submission_path}\")\n\n# 5. Display summary of submission\nprint(\"\\nSubmission preview:\")\nprint(sub_df.head())\n\nprint(f\"\\nTotal test predictions: {len(sub_df)}\")\nif os.path.exists(submission_path):\n    print(f\"Ready for Kaggle upload: {submission_path}\")\n\n# 6. Provide easy re-submission code\nprint(\"\\nTo submit your predictions to Kaggle, run the following in a new cell:\")\nprint(\"!kaggle competitions submit -c physionet-ecg-image-digitization -f submission.csv -m 'ECG digitization submission'\")\n\n# 7. Display leaderboard score if available (Kaggle will show after upload)\nprint(\"\\nAfter submission, your leaderboard score will appear on the competition page under 'My Submissions'.\")\nprint(\"Submission & Results cell complete.\")","outputs":[],"cell_number":8,"version":1,"status":"generated","created_at":"2025-11-18T22:54:42.935285+00:00","metadata":{},"execution_count":null}],"nbformat":4,"nbformat_minor":4}