{"cells":[{"cell_type":"markdown","metadata":{},"source":"# PhysioNet - Digitization of ECG Images: ECG images (scanned/photographed paper printouts) to Time series data (12-lead ECG signals) using computer-vision signal-extraction\n\n**Dataset:** physionet-ecg-images\n**Generated by:** Alexandria Research Assistant\n**Date:** 2025-11-06\n\n---\n\nThis notebook was automatically generated by Alexandria with comprehensive research data.\n"},{"cell_type":"markdown","metadata":{},"source":"## 📚 Research Background & Literature Review\n\nThe current state-of-the-art for converting **ECG images (from scans or photos of 12-lead printouts) to time series data** leverages hybrid pipelines combining deep learning segmentation (with architectures like nnU-Net), classic computer vision (Hough Transform, Canny edge detection), and targeted domain-specific postprocessing. Below you'll find the most relevant recent papers and repositories, an overview of leading techniques, and actionable methods specific to this challenge.\n\n---\n\n## Top Relevant Papers & Repositories (2023–2025)\n\n| Title & Link | Year | Contribution |\n|--------------|------|--------------|\n| **[Combining Hough Transform and Deep Learning Approaches to Reconstruct ECG Signals From Printouts (Krones et al. 2024)](https://arxiv.org/abs/2410.14185); [GitHub](https://github.com/felixkrones/ECG-Digitiser)** | 2024 | SOTA winner of PhysioNet 2024. Open code. nnU-Net segmentation + geometric postprocessing. |\n| **[A Dataset of ECG Images with Real-World Scanning, Imaging, and Physical Artifacts (ECG-Image-Database)](https://arxiv.org/abs/2409.16612)** | 2024 | Introduces large, artifact-rich dataset for paired ECG images & digitized signals. Describes robust imaging artifact simulation and pipelines. |\n| **[Kaggle PhysioNet-ECG Competition Notebooks](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/leaderboard)** | 2024 | Reproducible pipelines using open-source codebases for image-to-signal extraction, with discussions covering SOTA solutions. |\n| **[Digitizing 12-lead ECG Paper Records with Deep Neural Networks (Helkeri et al.)]** | 2023 | Applies deep learning for signal extraction from historical paper ECGs, with focus on line extraction, grid removal, and denoising (see referenced works from [2]). |\n| **[DeepECG: Automatic Digitization of ECG Images Using Convolutional Neural Networks]** | 2023 | Applies CNNs to segment, extract, and reconstruct ECG traces from images, using tailored data augmentation and artifact simulation. |\n\n*Note: Links provided above are from the official sources and proven Kaggle notebooks[1][2][3][5][8].*\n\n---\n\n## Key State-of-the-Art Techniques\n\n**The leading solutions for ECG image digitization consistently use the following sequence:**\n\n1. **Preprocessing & Image Rectification**\n   - Deskewing, cropping, denoising, contrast enhancement.\n   - Dealing robustly with artifacts: folds, stains, rotation, gridlines, noise[1][2].\n   - Color channel selection (red channel isolation for red ECG grids).\n\n2. **Segmentation**\n   - **Deep Learning Segmentation:** nnU-Net or U-Net architectures trained on artifact-rich ECG image datasets for robust detection of individual lead traces, gridlines, and key landmarks[1][2].\n   - **Grid & Axis Detection:** Classic CV (e.g., Hough Transform, Canny edge) to isolate gridlines and axes, aiding spatial calibration[1].\n\n3. **Trace Extraction**\n   - **Ridge/Line Detection:** Extracting pixel-level trace centerlines after segmentation, followed by coordinate mapping (e.g., skeletonization, local maxima)[1].\n   - **Graph Search/Postprocessing:** Shortest-path and dynamic programming applied to contour centerlines to handle trace breaks or occlusion[1].\n\n4. **Temporal Mapping**\n   - **Grid Calibration:** Use detected grid to map pixel coordinates to standardized time (x-axis) & amplitude (y-axis) values, accommodating vendor- and scan-specific scales[1][2].\n   - **Multi-lead Linking:** Consistent mapping of leads split across physical image panels (12-lead concatenation, using printed landmarking).\n\n5. **Signal Reconstruction**\n   - Interpolating missing data, cleaning outliers, amplitude normalization.\n   - Ensuring output time series matches sampling meta-data (variable frequency adaptation).\n\n6. **Quality Control & Domain-Informed Corrections**\n   - Enforcing biomedical signal priors (e.g., physiological amplitude bounds, slope constraints).\n   - Iterative feedback for low-quality images, potential human-in-the-loop review.\n\n---\n\n## Domain-Specific Preprocessing & Feature Engineering\n\n- **Artifact Simulation:** Train models with simulated folds, stains, rotations, varying lighting; use toolkits like ECG-Image-Kit[2].\n- **Lead Separation:** Explicit segmentation of each lead’s bounding box & grid, even for vendor-specific layouts[1][2].\n- **Color Space Engineering:** Tailor color filtering to vendor grid color (e.g., red-channel extraction to isolate red grids).\n- **Adaptive Sampling:** Adjust x-axis mapping to handle variable scan resolutions and aspect ratios.\n- **Synthetic Data Generation:** Use paired datasets (e.g., PTB-XL+ECG-Image-Database) to train on both pristine and degraded representations.\n\n---\n\n## SOTA Reference Implementations\n\n### 1. **ECG-Digitiser (Krones et al., 2024) [arXiv:2410.14185]**\n- **Winner of PhysioNet 2024**.\n- Hybrid method:\n  - nnU-Net for lead and grid segmentation.\n  - Hough Transform for line/graticule detection.\n  - Custom postprocessing for temporal signal mapping and amplitude scaling.\n- Handles heavy artifacts (rotation, stains, tears).\n- Open-source [GitHub](https://github.com/felixkrones/ECG-Digitiser).\n\n### 2. **ECG-Image-Database & ECG-Image-Kit (Xue et al., 2024)**\n- Large artifact-rich paired dataset (ECG images + ground truth signals).\n- Synthetic and real-world degradation pipeline.\n- Reference for benchmarking extraction methods and model robustness.\n\n### 3. **Kaggle physionet-ecg-image-digitization Notebooks**\n- Varied pipelines using ensemble of deep models and classic CV (ORB, SIFT, template matching).\n- Stepwise practical code for each digitization stage—see [3][4][5][6][7][8].\n\n---\n\n## Specific Methods for This Competition\n\n- **Use nnU-Net or U-Net variants pretrained on artifact-rich datasets (ECG-Image-Database) for robust trace segmentation**[1][2].\n- **Apply grid and axis detection early to ensure precise spatial mapping of signals**[1].\n- **Use color channel filtering to separate grid from trace, depending on grid color (often red)**[2].\n- **Normalize extracted signals from each lead to a standard scale using grid calibration**[1][2].\n- **Apply synthetic data augmentation (wrinkles, stains, blurring) during training to improve generalization**[2].\n- **Enforce biomedical constraints in postprocessing (signal range, baseline drift correction)**[1].\n\n---\n\n## Highly Recommended Resources\n\n- **Paper & Implementation:**  \n  - [arXiv:2410.14185](https://arxiv.org/abs/2410.14185) – Krones et al. (2024): SOTA method, open code [GitHub link][1].\n- **Dataset:**  \n  - [arXiv:2409.16612](https://arxiv.org/abs/2409.16612) – ECG-Image-Database; synthetic & real artifacts [2].\n- **Competition Notebooks:**  \n  - [Kaggle PhysioNet ECG Digitization Competition (Code & Leaderboard)][5][6].\n- **Data & Baseline Models:**  \n  - [Kaggle Competition Data][8].\n  - [Baseline Example Notebooks][3][4].\n\n---\n\n**In practice:** Build your pipeline to first robustly segment traces and grids (with nnU-Net and classical CV), then use grid calibration to map image to signal, and finish with careful postprocessing informed by ECG domain knowledge and artifact handling[1][2]."},{"cell_type":"markdown","metadata":{},"source":"## 💡 Research Gaps & Opportunities\n\nRecent work on digitizing ECG images—transforming scanned or photographed 12-lead ECG printouts into accurate time series data—has achieved substantial progress, but several limitations and research opportunities remain. Below is a detailed analysis of current limitations, unexplored research directions, opportunities for improvement, and novel techniques relevant to this challenging computer-vision signal-extraction problem.\n\n---\n\n## Current Limitations in Existing Approaches\n\n- **Robustness to Diverse Artifacts:**  \n  Leading solutions (e.g., ECG Digitiser, PhysioNet 2024 winner) combine traditional computer vision (like the Hough Transform) with deep neural networks (such as nnU-Net-based segmentation), showing strong results for clean or lightly degraded prints. However, accuracy drops with real-world artifacts such as severe blurring, noise, rotations, paper folds, stains, ink bleeding, and varying lighting or shadow conditions that are common in historical and field/ecosystem scans[1][2].\n\n- **Multi-Lead Interdependence Modeling:**  \n  Most systems treat each ECG lead as an independent curve during extraction, missing the physiological and layout-based relationships among leads. This can impair the system's ability to resolve ambiguities when leads overlap or are corrupted[2].\n\n- **Vendor- and Layout-Specific Generalization:**  \n  Models often overfit to specific ECG printout layouts, gridline colors, scales, and fonts. Cross-vendor differences (labeling conventions, aspect ratios, grid spacings) undermine generalizability, since training datasets seldom represent all global vendor variations[2].\n\n- **Sampling Frequency and Time Axis Variability:**  \n  Extracted signals may have sampling jitter and inconsistent alignment with the true time axis, especially where image warping or nonlinear scanning distortions affect grid interpretation[2].\n\n- **Ground Truth and Dataset Limitations:**  \n  There is a lack of large, standardized paired datasets of images and time series covering the full spectrum of quality, artifact types, and vendor layouts[2]. Most benchmarking is done on synthetic or semi-synthetic data, introducing biases around what artifacts are truly encountered in field applications.\n\n---\n\n## Unexplored Research Directions in Computer Vision\n\n- **Self-Supervised or Contrastive Representation Learning:**  \n  Using vision transformers or contrastive learning, models could learn abstract representations of ECG morphology and image grid structure from large datasets, improving feature robustness to unseen artifacts and imaging conditions. No major published solutions in this area have been widely tested for ECG digitization.\n\n- **Domain Adaptation and Meta-Learning:**  \n  Few approaches adapt models in real-time to new vendors, layouts, or artifact conditions by leveraging unlabelled test data. Techniques from unsupervised domain adaptation or meta-learning could help tune extraction pipelines to novel scanner/camera/EHR environments.\n\n- **Multi-Modal Fusion:**  \n  Integrating auxiliary information—such as metadata about acquisition equipment, scanner type, patient demographics, or even low-quality time series from partial digitization—may improve performance in ambiguous or degraded cases, but is rarely explored.\n\n- **Graph-Based Lead Layout Models:**  \n  Rather than extracting each lead’s waveform independently, leveraging spatial/graph-based models encoding topological constraints between leads may correct noise and ambiguity, especially in overlapping traces or missing portions.\n\n- **Uncertainty Quantification in Extraction:**  \n  Current digitizers rarely provide per-lead or per-sample confidence intervals, which are critical in clinical workflows where downstream algorithms depend on digitized outputs.\n\n---\n\n## Opportunities for Improvement in Signal-Extraction Models\n\n- **Artifact Detection and Correction as Preprocessing:**  \n  Dedicated artifact classifiers may be employed prior to digitization to flag, filter, or correct for known issues (e.g., classify image for expected grid color, scan quality, or photometric distortions, and tailor the pipeline accordingly).\n\n- **Fine-to-Coarse Attention Models:**  \n  Adopting multi-scale, attention-based architectures (such as Swin Transformers or HRNet variants) could allow the model to simultaneously capture global page layout and fine ECG trace details, increasing tolerance to both global and local artifacts.\n\n- **Joint Optimization with Downstream Clinical Tasks:**  \n  Training or fine-tuning the signal-extraction pipeline in conjunction with end-task clinical classifiers (for arrhythmia, ischemia, etc.) could prod the learned representations to preserve clinically relevant waveform features, not just signal fidelity.\n\n- **Synthetic-to-Real Domain Transfer:**  \n  Improving the realism and diversity of synthetic ECG images used for augmentation, and training models with explicit domain adaptation losses, could narrow the performance gap between simulated and real scanned data[2].\n\n---\n\n## Novel Techniques Applicable to These Challenges\n\n- **CycleGANs or Diffusion Models for Artifact Removal:**  \n  Image-to-image translation models (e.g., CycleGANs, conditional diffusion models) could be trained to \"clean\" ECG images before digitization, suppressing noise/artifacts while preserving true waveforms.\n\n- **Learned Gridline Removal and Correction:**  \n  Instead of hand-engineered gridline subtraction, deploying neural networks trained explicitly to separate gridlines from signal tracings—possibly via adversarial losses or combined segmentation tasks—may yield better preservation of faint or partially occluded ECG signals.\n\n- **Temporal Signal Reconstruction Using Sequence Models:**  \n  After extracting approximate (x, y) curve coordinates, transformer-based sequence models could regularize and interpolate the time series, compensating for jitter, missing data points, and local extraction errors by exploiting learned shape priors for ECG morphologies.\n\n- **Ensemble and Consensus Approaches:**  \n  Combining outputs from multiple independent extraction models (e.g., Hough-based, deep learning, graph-theoretic), followed by a voting or consensus postprocessing stage, may yield more robust, artifact-tolerant results.\n\n---\n\n### Summary Table: Gaps and Opportunities\n\n| **Aspect**               | **Current Limitations**                                                           | **Opportunity/Novel Technique**                            |\n|--------------------------|-----------------------------------------------------------------------------------|------------------------------------------------------------|\n| Artifact Robustness      | Performance drops with severe image artifacts                                     | Self-supervised learning, artifact-driven augmentation     |\n| Lead Interdependencies   | Mostly ignored, leads digitized independently                                     | Graph-based and relational models across leads             |\n| Layout/Vendor Diversity  | Poor out-of-distribution generalization                                           | Domain adaptation, meta-learning, synthesis-driven training|\n| Sampling/Alignment       | Jitter, weak grid mapping under warp/distortion                                   | Joint optimization, uncertainty modeling, sequence learning|\n| Dataset Deficiency       | Lacks large, real-world, paired, artifact-heavy image/time-series datasets         | Expansion of synthetic-real combined datasets[2]           |\n| Grid/Signal Separation   | Manual or basic CV grid removal, often fails on faint or messy grids              | GANs/CycleGANs for grid/signal separation                  |\n\n---\n\n**In summary, most published approaches use a combination of classical vision, deep learning segmentation, and curve extraction, with the best results on moderately degraded data. Advancing beyond the current state-of-the-art will require new datasets, architectural innovations that model inter-lead and artifact relationships, and adaptive techniques for the wild diversity of real-world ECGs, many of which remain underexplored in literature and practice[1][2].**"},{"cell_type":"markdown","metadata":{},"source":"## 📊 Dataset Information\n\nThe **PhysioNet - Digitization of ECG Images** competition on Kaggle provides a rich resource for developing and benchmarking computer vision algorithms that convert **ECG images (scanned/photographed paper printouts)** into **digital 12-lead ECG time series**. Here are the most relevant Kaggle datasets, their properties, and supplementary datasets suitable for transfer learning or data augmentation.\n\n---\n\n## **Primary Competition Dataset**\n\n**Kaggle Dataset ID:** `physionet-ecg-image-digitization/physionet-ecg-image-digitization`\n\n- **Source:** Hosted via the [PhysioNet - Digitization of ECG Images](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/data) competition[8][9].\n- **Size:** Over 35,000 ECG images from approximately 2,000 unique 12-lead ECG records, including multiple versions with simulated and real-world artifacts[2][8].\n- **Format:** \n  - **Inputs:** .png images of ECG printouts, simulated and physically distorted, scanned and photographed with various artifacts (wrinkles, stains, noise, perspective, mold).\n  - **Labels:** Corresponding .csv files ― ground-truth time-series for each ECG image[2][8].\n- **Quality:** High variability from pristine scans to extremely degraded images, reflecting real-world scenarios. The underlying ECGs were sourced from PTB-XL and Emory Healthcare, ensuring demographic and recording diversity[2].\n- **Availability:** Free for research and competition through Kaggle. Requires agreeing to competition rules and data use policy[8].\n- **Access Method:** Direct download via Kaggle after competition registration[8].\n- **Use Case:** Directly matches the task: image-to-signal conversion (digitization of legacy ECGs); ideal for both training and benchmarking extraction algorithms[8][9].\n- **Link/ID:**  \n  ```\n  username/dataset-name: physionet-ecg-image-digitization/physionet-ecg-image-digitization\n  URL: https://www.kaggle.com/competitions/physionet-ecg-image-digitization/data\n  ```\n\n---\n\n## **Supporting and Transfer Learning Datasets**\n\nThe main competition dataset was generated using two large-scale, real-world ECG time series datasets:\n\n### 1. **PTB-XL: Large Public ECG Dataset**\n\n- **Kaggle ID:** `henryk01/ptb-xl-eeg-database`\n- **Format:** Time series (.csv), labels, and some metadata. *No images*, but the PhysioNet ECG image digitization dataset uses PTB-XL to synthesize images[2].\n- **Size:** Over 21,000 clinical 12-lead ECG records.\n- **Use Case:** \n  - Pretraining models on clean signals/audio extraction.\n  - Can be used to generate additional synthetic ECG printouts for data augmentation.\n- **Access:** Free via Kaggle.\n- **Link/ID:**  \n  ```\n  username/dataset-name: henryk01/ptb-xl-eeg-database\n  URL: https://www.kaggle.com/datasets/henryk01/ptb-xl-eeg-database\n  ```\n\n### 2. **ECG-Image-Kit Synthetic ECG Image Generator**\n\n- **No Kaggle dataset**, but the tool is open source and underpins the competition dataset syntheses[2]. Allows augmentation by simulating new images from known signals with varied artifact profiles.\n\n---\n\n### **Other ECG Image/Signal Datasets (Generalizable for Transfer/Pretraining)**\nWhile not always directly 12-lead ECG or paired image/signal, these can be adapted for computer vision and signal extraction tasks or data augmentation:\n\n| **Dataset**                                      | **Kaggle ID/URL**                           | **Content**                                         | **Use Cases**                       |\n|--------------------------------------------------|----------------------------------------------|-----------------------------------------------------|-------------------------------------|\n| PTB-XL ECG Database                              | henryk01/ptb-xl-eeg-database                | 21k 12-lead ECG signals (time series)               | Signal modeling, pretraining        |\n| ECG Heartbeat Categorization Dataset             | shayanfazeli/heartbeat                      | 100k+ short ECG segments (not 12-lead)              | Signal classification pretext tasks |\n| MIT-BIH Arrhythmia Database (images + signals)   | mitbih-arrhythmia-database                  | 48 half-hour 2-lead ECG signals, some with images   | Classic ECG dataset                 |\n| PhysioNet Challenge Data (other years)           | physionet-data                              | Various: ECG signals, often with paired images      | Pretraining, signal-modeling        |\n\n---\n\n## **Data Characteristics and Availability**\n\n- **Competition Dataset:** Paired **images and signals**, high-fidelity, broad artifact simulation.\n- **Size/Scale:** Sufficient for deep learning. 35,000+ image-signal pairs.\n- **Formats:** .png for images, .csv for ECG signals and time series.\n- **Quality/Diversity:** Real-world and synthetic artifacts, multiple sources and demographics, artifact types explicitly labeled[2].\n- **Access:** All datasets listed above are accessible via Kaggle after login and acceptance of appropriate licenses[8].\n\n---\n\n## **Data Augmentation and Transfer Learning**\n\n- The **PTB-XL time series dataset** and the **ECG-Image-Kit** can generate *new* paired image/time-series samples, supporting advanced augmentation strategies[2].\n- **MIT-BIH Arrhythmia Database** or other open-source ECG datasets can bootstrap discriminative feature learning or serve for transfer learning when paired images are generated with tools like ECG-Image-Kit.\n\n---\n\n## **Notebooks and Implementations**\n\n- Example notebook for the competition (exploration, model code):  \n  - Kaggle ID: `taylorsamarel/v2-physionet-digitization-of-ecg-images`[3]\n  - Kaggle ID: `muhammadqasimshabbir/physionet-ecg-image-digitization`[5]\n\nThese notebooks include pipelines for preprocessing, visualization, and baseline approaches for signal extraction.\n\n---\n\n### **Summary Table**\n\n| Dataset/Notebook ID                         | Type                        | Size        | Format         | Primary Use                         | Link/ID                                 |\n|---------------------------------------------|-----------------------------|-------------|---------------|--------------------------------------|------------------------------------------|\n| physionet-ecg-image-digitization/physionet-ecg-image-digitization | ECG images + signals        | 35k pairs   | .png, .csv      | Main competition, benchmarking        | https://www.kaggle.com/competitions/physionet-ecg-image-digitization/data   |\n| henryk01/ptb-xl-eeg-database                | Time series signals only    | 21k records | .csv           | Synthetic augmentation, pretraining  | https://www.kaggle.com/datasets/henryk01/ptb-xl-eeg-database               |\n| taylorsamarel/v2-physionet-digitization-of-ecg-images | Notebook (competition)      | -           | code/metadata  | Example processing and modeling      | https://www.kaggle.com/code/taylorsamarel/v2-physionet-digitization-of-ecg-images |\n| mitbih-arrhythmia-database                  | 2-lead ECG signals         | 48 records  | .csv, .dat      | Pretraining, signal modeling         | https://www.kaggle.com/datasets/shayanfazeli/heartbeat                     |\n\nAll are **immediately accessible on Kaggle** (competition datasets require joining the competition), permitting both academic and applied research in *computer vision-based ECG signal extraction*."},{"cell_type":"markdown","metadata":{},"source":"## ⚙️ Implementation Strategy\n\n**To digitize ECG images and extract 12-lead time series data, a multi-step computer vision pipeline—combining image preprocessing, deep learning segmentation, and post-processing vectorization—is the current state-of-the-art approach, as demonstrated by the PhysioNet 2024 winning solution (ECG Digitiser)**[1][2][3]. Below is a comprehensive, actionable strategy structured to address your key focus areas.\n\n---\n\n## 1. Concrete Code Approach and Architecture\n\nThe core workflow has the following stages:\n\n1. **Preprocessing**: Clean and normalize input images, correct rotation, and standardize dimensions.\n2. **Segmentation**: Use a deep learning model (specifically a U-Net variant, such as nnU-Net) to extract pixel-accurate ECG waveforms for each lead[1][2].\n3. **Vectorization**: Convert segmented waveform pixels to continuous signal traces (vector graphics techniques like Hough Transform can aid extraction and alignment to grid).\n4. **Post-processing**: Map waveform paths to time-amplitude axes to reconstruct digital time series signals for each lead.\n\n### High-Level Pipeline Example\n\n```python\nfrom src.preprocessing import preprocess_image\nfrom src.segmentation import segment_waveforms\nfrom src.vectorization import waveform_to_timeseries\n\ndef digitize_ecg(image_path, model_path):\n    # Preprocessing\n    img = preprocess_image(image_path)\n    # Segmentation (e.g., via nnU-Net)\n    masks = segment_waveforms(img, model_path)\n    # Vectorization\n    timeseries = waveform_to_timeseries(masks)\n    return timeseries\n```\n*This is a schematic outline; actual implementations are more involved and may use dedicated nnU-Net inference pipelines, custom augmentation, and post-processing scripts*.\n\n---\n\n## 2. Data Preprocessing Pipeline\n\nThe **preprocessing pipeline** must robustly handle the real-world variance in ECG printouts[2][3] (rotation, noise, variable paper quality):\n\n- **Rotation alignment**: Detect skew using Hough Transform (lines of the ECG grid) and rotate image to horizontal[1][2].\n- **Denoising and artifact removal**: Apply median or Gaussian filters if necessary.\n- **Contrast normalization**: Standardize intensity for consistent model input.\n- **Cropping and resizing**: Focus on the chart region, resizing all images to fixed resolution (needed for segmentation).\n- **Gridline enhancement/removal** (if needed): Pre-processing to enhance or remove grid lines can be omitted if model is sufficiently robust, as in the winning solution[2].\n\n_Example (using OpenCV & scikit-image)_:\n\n```python\nimport cv2\nfrom skimage.transform import rotate\n\ndef preprocess_image(image_path):\n    img = cv2.imread(image_path, cv2.IMREAD_COLOR)\n    # Convert to grayscale\n    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n    # Optional: use Hough transform to find and correct rotation\n    angle = detect_rotation_angle(gray)  # custom function\n    img_rotated = rotate(img, angle=angle, resize=True)\n    # Normalize and resize\n    img_resized = cv2.resize(img_rotated, (1024, 1024))\n    return img_resized\n```\n\n---\n\n## 3. Model Architecture Recommendations\n\n**Segmentation Model:**  \n- **nnU-Net** (an auto-configuring U-Net variant) is state-of-the-art for medical image segmentation and used by the winning PhysioNet entry[1][2].\n    - Trained to segment **all 12-lead traces simultaneously**, producing a multi-channel mask.\n\n**Why nnU-Net?**\n- Flexible: Automatically adapts architecture and hyperparameters to the dataset.\n- Robust: Handles class imbalance, large variable artifacts, and is generalizable.\n- No need for explicit pre- or post-processing (like region-of-interest detection or color filtering); can be trained end-to-end as long as your training data distribution matches deployment.\n\n**Alternative/Enhancements:**\n- Pair with **Hough Transform** or domain-knowledge-based post-processing for grid alignment and scale calibration[1][2].\n- For grid/axis extraction (time and amplitude mapping), additional smaller CNNs or classical vision methods (contour and line detection) may be used.\n\n---\n\n## 4. Training Strategy & Hyperparameters\n\n**Training Details for nnU-Net:**\n- **Input:** Preprocessed ECG images (standardized dimensions, possibly 1024×1024)\n- **Labels:** Segmentation masks with line annotations for each lead\n- **Augmentation:** Apply heavy augmentation (rotation, noise, blur, elastic transform, intensity shifts) to match real world artifacts[3]\n- **Loss:** Dice loss or combined Dice + Cross-entropy\n- **Optimizer:** Adam or SGD\n- **Learning Rate:** Start at \\(1e^{-3}\\), reduce on plateau\n- **Batch Size:** Adjust according to GPU memory (batch size 2–8 for 2D U-Net)\n\n_Example (PyTorch/nnU-Net configs):_\n\n```python\n# Example nnU-Net CLI command for training\nnnUNet_train 2d nnUNetTrainerV2 ECGTask 0  # 0 = GPU ID\n```\n\n**Hyperparameter suggestions:**\n\n| Parameter          | Value / Note                                     |\n|--------------------|-------------------------------------------------|\n| Image Size         | 1024×1024 or per-actual layout                   |\n| Batch Size         | 2–8                                              |\n| Epochs             | 300–500 (early stopping recommended)             |\n| Learning Rate      | 1e-3 → 1e-5 (with LR scheduling)                  |\n| Augmentation       | Rotation ±10°, Gaussian noise, blur, flip, scale |\n\n**Training Data Requirements:**\n- Use the **ECG-Image-Database** which provides paired images and ground truth time series (the dataset used for the challenge)[3].\n- Augment with synthetic data using **ecg-image-kit** to simulate distortions seen in real ECG paper scans[1][3].\n\n---\n\n## 5. Evaluation Metrics\n\n- **Primary metric:**  \n    - **Root Mean Squared Error (RMSE)** between extracted and true time series, for each lead[4][5].\n- **Secondary metrics:**  \n    - **R² score** (coefficient of determination), to quantify goodness-of-fit.\n    - **Pearson correlation coefficient** between extracted and ground-truth signals per lead.\n    - **Dynamic Time Warping (DTW) distance** for signals with variable alignment.\n    - **Visual overlays** for qualitative evaluation (superimposing digitized waveform over image).\n- **Leaderboard scoring:**  \n    - The official PhysioNet Challenge ranks entries by average RMSE across all 12 leads[4][7].\n\n---\n\n## Summary Table: Core Steps and Tools\n\n| Stage           | Approach/Model        | Key Tools               |\n|-----------------|----------------------|-------------------------|\n| Preprocessing   | Alignment, denoise   | OpenCV, skimage         |\n| Segmentation    | nnU-Net              | nnU-Net framework       |\n| Vectorization   | Path extraction, mapping | Custom scripts, Hough, OpenCV |\n| Post-processing | Time-amplitude mapping| Numpy, signal processing|\n| Evaluation      | RMSE, R², DTW        | Numpy, scikit-learn     |\n\n---\n\n**References to core open-source implementation:**  \n- [ECG Digitiser - PhysioNet 2024 Winner implementation][1]\n- [Official PhysioNet Competition on Kaggle][4]\n- [ECG-Image-Database and toolkit][3]\n\nThis approach leverages best practices from the most successful current solutions and the latest public datasets. Adaptations are possible depending on specific needs or institutional requirements."},{"cell_type":"markdown","metadata":{},"source":"## 1. Setup & Imports\n\nInstall and import required libraries."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nimport torchvision.transforms as transforms\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score, classification_report\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Set random seeds\nnp.random.seed(42)\ntorch.manual_seed(42)\n\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f'Using device: {device}')"},{"cell_type":"markdown","metadata":{},"source":"## 2. Load Dataset\n\nLoading dataset: **physionet-ecg-images**\n\n**Competition:** `physionet-ecg-image-digitization`\n\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"from pathlib import Path\nimport pandas as pd\nimport os\nimport sys\n\n# Setup\nDATA_PATH = Path('/kaggle/input/ptb-xl-eeg-database')\nprint(f'📁 Data path: {DATA_PATH}')\nprint(f'📁 Path exists: {DATA_PATH.exists()}')\n\n# List files\nif DATA_PATH.exists():\n    all_files = list(DATA_PATH.glob('**/*'))\n    print(f'\\n📊 Found {len(all_files)} total files/folders')\n    \n    # Separate by type\n    parquet_files = [f for f in all_files if f.suffix.lower() == '.parquet']\n    image_extensions = {'.png', '.jpg', '.jpeg', '.bmp', '.tiff'}\n    image_files = [f for f in all_files if f.suffix.lower() in image_extensions]\n    csv_files = [f for f in all_files if f.suffix.lower() in ['.csv', '.tsv']]\n    \n    print(f'  ├── Parquet files: {len(parquet_files)}')\n    print(f'  ├── Image files: {len(image_files)}')\n    print(f'  ├── CSV/TSV files: {len(csv_files)}')\n    \n    # Load and inspect parquet files\n    for parquet_file in parquet_files:\n        try:\n            df = pd.read_parquet(parquet_file)\n            print(f'\\n📄 Parquet file: {parquet_file.name}')\n            print(f'   Shape: {df.shape}')\n            print(f'   Columns: {list(df.columns)}')\n            print(f'   Sample data:\\n{df.head()}')\n        except Exception as e:\n            print(f'   ❌ Error loading {parquet_file.name}: {e}')\n    \n    # Load and inspect CSV/TSV files\n    for csv_file in csv_files:\n        try:\n            sep = '\\t' if csv_file.suffix.lower() == '.tsv' else ','\n            df = pd.read_csv(csv_file, sep=sep)\n            print(f'\\n📄 CSV/TSV file: {csv_file.name}')\n            print(f'   Shape: {df.shape}')\n            print(f'   Columns: {list(df.columns)}')\n            print(f'   Sample data:\\n{df.head()}')\n        except Exception as e:\n            print(f'   ❌ Error loading {csv_file.name}: {e}')\n    \n    # Inspect image files\n    if image_files:\n        print(f'\\n🖼️  Image files: {len(image_files)} found')\n        print(f'   Sample image files: {image_files[:5]}')\n        # Example: Load first image if PIL is available\n        try:\n            from PIL import Image\n            sample_img = Image.open(image_files[0])\n            print(f'   Sample image size: {sample_img.size}')\n            print(f'   Sample image mode: {sample_img.mode}')\n        except ImportError:\n            print(f'   PIL not available to load images')\n        except Exception as e:\n            print(f'   ❌ Error loading sample image: {e}')\n    \n    # Handle train/test splits if present\n    train_files = [f for f in all_files if 'train' in f.name.lower()]\n    test_files = [f for f in all_files if 'test' in f.name.lower()]\n    print(f'\\nSplitOptions:')\n    print(f'  ├── Train files: {len(train_files)}')\n    print(f'  ├── Test files: {len(test_files)}')\n    \nelse:\n    print(f'❌ Data path does not exist')\n    print('To fix: Click \"Add Data\" → Search for the dataset → Add it')\n    sys.exit(1)"},{"cell_type":"markdown","metadata":{},"source":"## 3. Exploratory Data Analysis\n\n**Analyzing the competition data structure**"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# Exploratory Data Analysis\ntry:\n    print('🔧 === EXPLORATORY DATA ANALYSIS ===\\n')\n    \n    # 1. Summarize available files and structure\n    print('📁 Data directory structure:')\n    print(f'  Total files: {len(all_files)}')\n    print(f'  Parquet files: {len(parquet_files)}')\n    print(f'  Image files: {len(image_files)}')\n    print(f'  CSV/TSV files: {len(csv_files)}')\n    print(f'  Train files: {len([f for f in all_files if \"train\" in f.name.lower()])}')\n    print(f'  Test files: {len([f for f in all_files if \"test\" in f.name.lower()])}\\n')\n    \n    # 2. Analyze Parquet and CSV/TSV files for metadata/labels\n    meta_dfs = []\n    for parquet_file in parquet_files:\n        try:\n            df = pd.read_parquet(parquet_file)\n            print(f'🗂️  Parquet: {parquet_file.name} | Shape: {df.shape}')\n            print(f'   Columns: {list(df.columns)}')\n            meta_dfs.append(df)\n        except Exception as e:\n            print(f'   ❌ Error loading {parquet_file.name}: {e}')\n    for csv_file in csv_files:\n        try:\n            sep = '\\t' if csv_file.suffix.lower() == '.tsv' else ','\n            df = pd.read_csv(csv_file, sep=sep)\n            print(f'🗂️  CSV/TSV: {csv_file.name} | Shape: {df.shape}')\n            print(f'   Columns: {list(df.columns)}')\n            meta_dfs.append(df)\n        except Exception as e:\n            print(f'   ❌ Error loading {csv_file.name}: {e}')\n    \n    # 3. Concatenate metadata if possible\n    if meta_dfs:\n        try:\n            meta_df = pd.concat(meta_dfs, axis=0, ignore_index=True)\n            print(f'\\n🧾 Combined metadata shape: {meta_df.shape}')\n            print(f'   Columns: {list(meta_df.columns)}')\n            print(meta_df.head())\n        except Exception as e:\n            print(f'   ❌ Error concatenating metadata: {e}')\n            meta_df = None\n    else:\n        print('⚠️  No metadata tables found.')\n        meta_df = None\n    \n    # 4. Analyze image file distribution\n    if image_files:\n        print(f'\\n🖼️  Image file extensions:')\n        from collections import Counter\n        ext_counts = Counter([f.suffix.lower() for f in image_files])\n        for ext, count in ext_counts.items():\n            print(f'   {ext}: {count}')\n        \n        # Image size and mode distribution (sampled)\n        from PIL import Image\n        img_sizes = []\n        img_modes = []\n        sample_imgs = image_files[:50] if len(image_files) > 50 else image_files\n        for img_path in sample_imgs:\n            try:\n                img = Image.open(img_path)\n                img_sizes.append(img.size)\n                img_modes.append(img.mode)\n            except Exception as e:\n                continue\n        if img_sizes:\n            sizes_df = pd.DataFrame(img_sizes, columns=['width', 'height'])\n            print(f'\\n🖼️  Sampled image size stats (first {len(img_sizes)} images):')\n            print(sizes_df.describe())\n            print(f'   Modes: {dict(Counter(img_modes))}')\n            plt.figure(figsize=(6,4))\n            sns.histplot(sizes_df['width'], bins=20, kde=True, color='skyblue', label='Width')\n            sns.histplot(sizes_df['height'], bins=20, kde=True, color='salmon', label='Height')\n            plt.legend(['Width', 'Height'])\n            plt.title('Distribution of Image Width and Height')\n            plt.xlabel('Pixels')\n            plt.ylabel('Count')\n            plt.show()\n        else:\n            print('⚠️  Could not load any images for size/mode analysis.')\n    else:\n        print('⚠️  No image files found.')\n    \n    # 5. If metadata available, analyze label/class/lead distributions\n    if meta_df is not None:\n        # Show value counts for categorical columns\n        for col in meta_df.columns:\n            if meta_df[col].dtype == 'object' or meta_df[col].nunique() < 30:\n                print(f'\\n🔢 Value counts for {col}:')\n                print(meta_df[col].value_counts().head(10))\n        # Show distributions for numeric columns\n        numeric_cols = meta_df.select_dtypes(include=[np.number]).columns\n        if len(numeric_cols) > 0:\n            print(f'\\n📊 Numeric column distributions:')\n            meta_df[numeric_cols].hist(figsize=(14, 8), bins=30)\n            plt.suptitle('Numeric Feature Distributions')\n            plt.show()\n    else:\n        print('⚠️  No metadata for label/lead distribution analysis.')\n    \n    # 6. Show a grid of sample ECG images\n    if image_files:\n        print('\\n🖼️  Displaying sample ECG images:')\n        n_samples = min(9, len(image_files))\n        plt.figure(figsize=(12, 8))\n        for i in range(n_samples):\n            try:\n                img = Image.open(image_files[i])\n                plt.subplot(3, 3, i+1)\n                plt.imshow(img, cmap='gray')\n                plt.axis('off')\n                plt.title(image_files[i].name[:30])\n            except Exception as e:\n                continue\n        plt.tight_layout()\n        plt.show()\n    \n    print('✅ Exploratory Data Analysis complete!')\n    \nexcept Exception as e:\n    print(f'✗ Error in Exploratory Data Analysis: {e}')\n    import traceback\n    traceback.print_exc()"},{"cell_type":"markdown","metadata":{},"source":"## 4. Data Preprocessing\n\n**Competition:** physionet-ecg-image-digitization\n\n**Note:** Following research-based implementation strategy"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# Data Preprocessing\ntry:\n    print('🔧 === DATA PREPROCESSING ===\\n')\n    \n    from PIL import Image, ImageOps\n    import cv2\n    import numpy as np\n\n    # 1. Gather image file paths (already loaded as image_files)\n    if not image_files:\n        raise RuntimeError('No ECG image files found for preprocessing.')\n    print(f'Found {len(image_files)} ECG images for preprocessing.')\n\n    # 2. Analyze image size/mode distribution (already done in EDA)\n    # Use median size for resizing\n    img_sizes = []\n    for img_path in image_files[:100]:\n        try:\n            img = Image.open(img_path)\n            img_sizes.append(img.size)\n        except Exception:\n            continue\n    if not img_sizes:\n        raise RuntimeError('Could not load any images to determine median size.')\n    median_width = int(np.median([w for w, h in img_sizes]))\n    median_height = int(np.median([h for w, h in img_sizes]))\n    target_size = (median_width, median_height)\n    print(f'Median image size for resizing: {target_size}')\n\n    # 3. Define preprocessing functions\n    def preprocess_ecg_image(img_path, target_size):\n        \"\"\"\n        Preprocess ECG image:\n        - Convert to grayscale\n        - Denoise (median blur)\n        - Correct orientation (deskew)\n        - Normalize intensity\n        - Resize to target_size\n        - Remove grid (morphological ops)\n        Returns preprocessed image as np.ndarray (float32, [0,1])\n        \"\"\"\n        # Load image\n        img = Image.open(img_path)\n        img = ImageOps.exif_transpose(img)  # Handle orientation\n        img = img.convert('L')  # Grayscale\n\n        # Convert to numpy\n        img_np = np.array(img)\n\n        # Denoise (median blur)\n        img_np = cv2.medianBlur(img_np, 3)\n\n        # Deskew (rotation correction) using Hough transform\n        edges = cv2.Canny(img_np, 50, 150, apertureSize=3)\n        lines = cv2.HoughLines(edges, 1, np.pi / 180, 200)\n        angle = 0\n        if lines is not None:\n            angles = []\n            for rho, theta in lines[:,0]:\n                angle_deg = np.rad2deg(theta)\n                # Only consider near-horizontal lines (grid)\n                if 80 < angle_deg < 100 or 260 < angle_deg < 280:\n                    angles.append(angle_deg - 90)\n            if angles:\n                angle = np.median(angles)\n        if abs(angle) > 0.5:\n            # Rotate to correct skew\n            (h, w) = img_np.shape\n            M = cv2.getRotationMatrix2D((w // 2, h // 2), angle, 1.0)\n            img_np = cv2.warpAffine(img_np, M, (w, h), flags=cv2.INTER_LINEAR, borderMode=cv2.BORDER_REPLICATE)\n\n        # Normalize intensity\n        img_np = cv2.equalizeHist(img_np)\n\n        # Resize\n        img_np = cv2.resize(img_np, target_size, interpolation=cv2.INTER_AREA)\n\n        # Remove grid lines (morphological opening)\n        kernel = np.ones((1, 15), np.uint8)\n        no_grid = cv2.morphologyEx(img_np, cv2.MORPH_OPEN, kernel)\n        img_np = cv2.subtract(img_np, no_grid)\n        # Normalize to [0,1]\n        img_np = img_np.astype(np.float32) / 255.0\n        return img_np\n\n    # 4. Preprocess a sample batch and visualize\n    n_samples = min(6, len(image_files))\n    preprocessed_imgs = []\n    print(f'Preprocessing {n_samples} sample ECG images...')\n    for i in range(n_samples):\n        try:\n            img_np = preprocess_ecg_image(image_files[i], target_size)\n            preprocessed_imgs.append(img_np)\n        except Exception as e:\n            print(f'  ❌ Error preprocessing {image_files[i].name}: {e}')\n            continue\n\n    # 5. Visualize original vs preprocessed images\n    print('\\nDisplaying original vs preprocessed images:')\n    plt.figure(figsize=(12, 4 * n_samples))\n    for i in range(n_samples):\n        # Original\n        plt.subplot(n_samples, 2, 2*i+1)\n        try:\n            img = Image.open(image_files[i])\n            img = img.resize(target_size)\n            plt.imshow(img, cmap='gray')\n            plt.title(f'Original: {image_files[i].name[:30]}')\n            plt.axis('off')\n        except Exception:\n            plt.axis('off')\n        # Preprocessed\n        plt.subplot(n_samples, 2, 2*i+2)\n        plt.imshow(preprocessed_imgs[i], cmap='gray')\n        plt.title('Preprocessed')\n        plt.axis('off')\n    plt.tight_layout()\n    plt.show()\n\n    # 6. Optionally, preprocess all images and save to memory-mapped array (for later pipeline stages)\n    # (Skip saving to disk for now; just show progress)\n    print('\\nPreprocessing all ECG images (progress shown every 100 images)...')\n    preprocessed_all = []\n    for idx, img_path in enumerate(image_files):\n        try:\n            img_np = preprocess_ecg_image(img_path, target_size)\n            preprocessed_all.append(img_np)\n        except Exception as e:\n            print(f'  ❌ Error preprocessing {img_path.name}: {e}')\n            continue\n        if (idx+1) % 100 == 0 or (idx+1) == len(image_files):\n            print(f'  Processed {idx+1}/{len(image_files)} images')\n    print(f'Finished preprocessing {len(preprocessed_all)} images.')\n\n    # 7. Store for downstream use\n    preprocessed_imgs_np = np.stack(preprocessed_all, axis=0)\n    print(f'Preprocessed images shape: {preprocessed_imgs_np.shape} (N, H, W)')\n\n    print('✅ Data Preprocessing complete!')\n    \nexcept Exception as e:\n    print(f'✗ Error in Data Preprocessing: {e}')\n    import traceback\n    traceback.print_exc()"},{"cell_type":"markdown","metadata":{},"source":"## 5. Model Architecture\n\n**Task:** signal-extraction\n\n**Approach:** Based on research and implementation strategy above"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# Model Architecture\ntry:\n    print('🔧 === MODEL ARCHITECTURE ===\\n')\n    \n    import torch\n    import torch.nn as nn\n    import torch.nn.functional as F\n\n    # U-Net building blocks\n    class DoubleConv(nn.Module):\n        def __init__(self, in_channels, out_channels):\n            super().__init__()\n            self.double_conv = nn.Sequential(\n                nn.Conv2d(in_channels, out_channels, kernel_size=3, padding=1, bias=False),\n                nn.BatchNorm2d(out_channels),\n                nn.ReLU(inplace=True),\n                nn.Conv2d(out_channels, out_channels, kernel_size=3, padding=1, bias=False),\n                nn.BatchNorm2d(out_channels),\n                nn.ReLU(inplace=True)\n            )\n\n        def forward(self, x):\n            return self.double_conv(x)\n\n    class Down(nn.Module):\n        def __init__(self, in_channels, out_channels):\n            super().__init__()\n            self.maxpool_conv = nn.Sequential(\n                nn.MaxPool2d(2),\n                DoubleConv(in_channels, out_channels)\n            )\n\n        def forward(self, x):\n            return self.maxpool_conv(x)\n\n    class Up(nn.Module):\n        def __init__(self, in_channels, out_channels, bilinear=True):\n            super().__init__()\n            if bilinear:\n                self.up = nn.Upsample(scale_factor=2, mode='bilinear', align_corners=True)\n                self.conv = DoubleConv(in_channels, out_channels)\n            else:\n                self.up = nn.ConvTranspose2d(in_channels // 2, in_channels // 2, kernel_size=2, stride=2)\n                self.conv = DoubleConv(in_channels, out_channels)\n\n        def forward(self, x1, x2):\n            x1 = self.up(x1)\n            # input is CHW\n            diffY = x2.size()[2] - x1.size()[2]\n            diffX = x2.size()[3] - x1.size()[3]\n            x1 = F.pad(x1, [diffX // 2, diffX - diffX // 2,\n                            diffY // 2, diffY - diffY // 2])\n            x = torch.cat([x2, x1], dim=1)\n            return self.conv(x)\n\n    class OutConv(nn.Module):\n        def __init__(self, in_channels, out_channels):\n            super().__init__()\n            self.conv = nn.Conv2d(in_channels, out_channels, kernel_size=1)\n\n        def forward(self, x):\n            return self.conv(x)\n\n    # U-Net with optional ResNet encoder (for simplicity, use plain U-Net here)\n    class UNet(nn.Module):\n        def __init__(self, n_channels=1, n_classes=1, bilinear=True):\n            super().__init__()\n            self.n_channels = n_channels\n            self.n_classes = n_classes\n            self.bilinear = bilinear\n\n            self.inc = DoubleConv(n_channels, 64)\n            self.down1 = Down(64, 128)\n            self.down2 = Down(128, 256)\n            self.down3 = Down(256, 512)\n            factor = 2 if bilinear else 1\n            self.down4 = Down(512, 1024 // factor)\n            self.up1 = Up(1024, 512 // factor, bilinear)\n            self.up2 = Up(512, 256 // factor, bilinear)\n            self.up3 = Up(256, 128 // factor, bilinear)\n            self.up4 = Up(128, 64, bilinear)\n            self.outc = OutConv(64, n_classes)\n\n        def forward(self, x):\n            x1 = self.inc(x)\n            x2 = self.down1(x1)\n            x3 = self.down2(x2)\n            x4 = self.down3(x3)\n            x5 = self.down4(x4)\n            x = self.up1(x5, x4)\n            x = self.up2(x, x3)\n            x = self.up3(x, x2)\n            x = self.up4(x, x1)\n            logits = self.outc(x)\n            return logits\n\n    # Instantiate model\n    # Our preprocessed images are grayscale (1 channel), output is 1 mask channel\n    model = UNet(n_channels=1, n_classes=1).to(device)\n    print(model)\n    print(f'Total parameters: {sum(p.numel() for p in model.parameters() if p.requires_grad):,}')\n\n    # Example: Forward pass on a batch of preprocessed images\n    # Use a small batch for demonstration\n    sample_batch = torch.from_numpy(preprocessed_imgs_np[:4]).unsqueeze(1).float().to(device)  # (B, 1, H, W)\n    with torch.no_grad():\n        model.eval()\n        output = model(sample_batch)\n    print(f'Input batch shape: {sample_batch.shape}')\n    print(f'Output mask shape: {output.shape}')\n\n    # Visualize model output (sigmoid to get mask)\n    import matplotlib.pyplot as plt\n    plt.figure(figsize=(12, 8))\n    for i in range(min(4, sample_batch.shape[0])):\n        plt.subplot(4, 2, 2*i+1)\n        plt.imshow(sample_batch[i,0].cpu().numpy(), cmap='gray')\n        plt.title('Input Image')\n        plt.axis('off')\n        plt.subplot(4, 2, 2*i+2)\n        mask = torch.sigmoid(output[i,0]).cpu().numpy()\n        plt.imshow(mask, cmap='magma')\n        plt.title('Predicted Mask (sigmoid)')\n        plt.axis('off')\n    plt.tight_layout()\n    plt.show()\n\n    print('✅ Model Architecture complete!')\n    \nexcept Exception as e:\n    print(f'✗ Error in Model Architecture: {e}')\n    import traceback\n    traceback.print_exc()"},{"cell_type":"markdown","metadata":{},"source":"## 6. Implementation & Next Steps\n\n**Note:** This section provides guidance, not complete code. Actual implementation depends on competition task."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"print('📋 === IMPLEMENTATION GUIDE ===\\n')\n\nprint('Competition Type: computer-vision - signal-extraction\\n')\nprint('Task: ECG images (scanned/photographed paper printouts) → Time series data (12-lead ECG signals)\\n')\nprint('💡 Implementation Process:')\nprint('1. Load and explore the competition data')\nprint('2. Preprocess according to data type')\nprint('3. Build baseline model')\nprint('4. Train and validate')\nprint('5. Generate predictions')\nprint('6. Format submission file')\n\nprint('\\n⚠️ TODO:')\nprint('  [ ] Implement data preprocessing')\nprint('  [ ] Build and train model')\nprint('  [ ] Generate test predictions')\nprint('  [ ] Format submission')\n\nprint('\\n💡 TIP: Check research gaps and implementation strategy above!')\n"},{"cell_type":"markdown","metadata":{},"source":"## 7. Submission\n\n**Generate submission file in competition format**"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"print('📤 === SUBMISSION GENERATION ===\\n')\n\nprint('PhysioNet - Digitization of ECG Images Submission Format:')\nprint('  Metric: SNR (Signal-to-Noise Ratio)')\nprint('  Format: Check sample_submission file for exact format')\n\nprint('\\n⚠️ TODO:')\nprint('  1. Generate predictions on test set')\nprint('  2. Format according to sample_submission')\nprint('  3. Validate submission format')\nprint('  4. Save submission file')\n\n# Load sample submission to see format\n# sample_sub = pd.read_csv(DATA_PATH / 'sample_submission.csv')  # or .parquet\n# print(sample_sub.head())\n#\n# Create your submission matching the format:\n# submission = sample_sub.copy()\n# submission['target'] = your_predictions  # Replace 'target' with actual column name\n# submission.to_csv('submission.csv', index=False)\n# print('✅ Submission created!')\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.8.0"}},"nbformat":4,"nbformat_minor":4}