{
  "id": 669558,
  "title": "14th Place Solution",
  "url": "/competitions/physionet-ecg-image-digitization/writeups/14th-place-solution",
  "author_name": "",
  "post_date": "2026-01-23T03:37:54.760Z",
  "votes": 22,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>PhysioNet ECG Digitization Challenge - Solution Writeup</h1>\n<h2>Overview</h2>\n<p>This solution addresses the PhysioNet challenge of digitizing ECG images—converting printed/photographed ECG recordings back into numerical waveform data. The approach uses a <strong>3-stage deep learning pipeline</strong> that progressively transforms raw ECG images into high-fidelity digital waveforms.</p>\n<p><strong>Final Public Leaderboard Score: 21.56 SNR</strong>\n<strong>Final Private Leaderboard Score: 21.37 SNR</strong></p>\n<hr>\n<h2>Problem Statement</h2>\n<p>ECG images from various sources (printed records, mobile phone photos, scans of damaged documents) need to be converted back to digital waveforms. The challenge involves:</p>\n<ol>\n<li><strong>Diverse image sources</strong>: High-quality prints, mobile photos, screen captures, damaged/moldy documents</li>\n<li><strong>Geometric distortions</strong>: Perspective warps, rotations, varying aspect ratios</li>\n<li><strong>Visual degradations</strong>: Water stains, fold creases, mold, low resolution, compression artifacts</li>\n<li><strong>Layout parsing</strong>: Standard 4-strip ECG layout with 12 leads arranged in specific positions</li>\n</ol>\n<hr>\n<h2>Solution Architecture</h2>\n<h3>Three-Stage Pipeline</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2Fea72f9f6a77a90a276213a31965d4755%2Fdiagram1.jpg?generation=1769138841225117&amp;alt=media\" alt=\"Pipeline Diagram\"></p>\n<h3>Stage 0: Normalization</h3>\n<p><strong>Purpose</strong>: Detect lead markers and correct image orientation</p>\n<ul>\n<li><strong>Architecture</strong>: ResNet18 encoder + U-Net decoder</li>\n<li><strong>Outputs</strong>: <ul>\n<li>Lead marker segmentation (13 leads + background)</li>\n<li>Image orientation classification (8 orientations)</li></ul></li>\n<li><strong>Processing</strong>: Applies homography to normalize image to canonical size (3024×4032)</li>\n</ul>\n<blockquote>\n  <p><strong>Note</strong>: For this stage, I utilized the pre-trained Stage 0 checkpoint provided in the <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">original baseline</a>.</p>\n</blockquote>\n<h3>Stage 1: Rectification</h3>\n<p><strong>Purpose</strong>: Detect and correct grid distortions</p>\n<ul>\n<li><strong>Architecture</strong>: ResNet34 encoder + U-Net decoder  </li>\n<li><strong>Outputs</strong>:<ul>\n<li>Grid point detection (keypoints for perspective correction)</li>\n<li>Horizontal line classification (44 classes)</li>\n<li>Vertical line classification (57 classes)</li></ul></li>\n<li><strong>Processing</strong>: Uses detected grid points to apply perspective transformation, producing a perfectly aligned rectified image</li>\n</ul>\n<blockquote>\n  <p><strong>Note</strong>: Similar to Stage 0, I used the pre-trained Stage 1 checkpoint from the <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">original baseline</a>.</p>\n</blockquote>\n<h3>Stage 2: Waveform Prediction</h3>\n<p><strong>Purpose</strong>: Extract ECG waveforms from rectified images</p>\n<p>This is the main prediction stage where most of the innovation lies.</p>\n<h4>Model Architecture</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2F28e82aa96fa43e3593089860871597e3%2Fdiagram2.jpg?generation=1769138929087253&amp;alt=media\" alt=\"Stage 2 Architecture\"></p>\n<h4>Key Design Decisions</h4>\n<ol>\n<li><p><strong>HRNet-W48 Backbone</strong>: Maintains high-resolution representations throughout the network via parallel branches, crucial for precise pixel-level waveform prediction</p></li>\n<li><p><strong>CoordConv Decoder</strong>: Injects normalized 2D coordinates at each decoder block, enabling position-aware predictions essential for ECG layout understanding</p></li>\n<li><p><strong>RGB Skip Connection</strong>: Direct connection from input image to prediction head preserves fine spatial details that may be lost in encoder compression</p></li>\n<li><p><strong>PixelShuffle Upsampling</strong>: Memory-efficient learnable upsampling using sub-pixel convolution instead of transposed convolution</p></li>\n<li><p><strong>GroupNorm</strong>: Enables stable training with batch size 1 (necessary for high-resolution 1696×4352 images on limited GPU memory)</p></li>\n</ol>\n<hr>\n<h2>Training Strategy</h2>\n<h3>Loss Function</h3>\n<p>The model uses <strong>Binary Cross-Entropy (BCE) with positive class weighting</strong> for pixel-level supervision:</p>\n<pre><code>loss = F.binary_cross_entropy_with_logits(\n    pixel_logit,\n    target_mask,\n    pos_weight=10.0  # Handle sparse ECG line pixels\n)\n</code></pre>\n<p>This was found to outperform regression-based losses (MSE, SNR) for this task.</p>\n<h3>Data Augmentation</h3>\n<p>A comprehensive <strong>segment-aware augmentation pipeline</strong> was developed to match the diverse training data distribution:</p>\n<h4>Universal Augmentations</h4>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Probability</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Perlin Noise</td>\n<td>85%</td>\n<td>Simulate scanner noise, paper texture</td>\n</tr>\n<tr>\n<td>Texture Noise</td>\n<td>70%</td>\n<td>Grid-like artifacts, scan lines</td>\n</tr>\n<tr>\n<td>Color Jitter</td>\n<td>80%</td>\n<td>Brightness, contrast, saturation, hue variation</td>\n</tr>\n<tr>\n<td>Gaussian Noise</td>\n<td>50%</td>\n<td>Scanner ISO noise</td>\n</tr>\n<tr>\n<td>Gaussian Blur</td>\n<td>35%</td>\n<td>Low-quality scans, out-of-focus photos</td>\n</tr>\n<tr>\n<td>Shadow Overlay</td>\n<td>35%</td>\n<td>Uneven lighting</td>\n</tr>\n<tr>\n<td>Paper Texture</td>\n<td>30%</td>\n<td>Physical paper grain simulation</td>\n</tr>\n<tr>\n<td>JPEG Artifacts</td>\n<td>30%</td>\n<td>Compression artifacts</td>\n</tr>\n<tr>\n<td>Horizontal Flip</td>\n<td>50%</td>\n<td>Data augmentation (with label flip)</td>\n</tr>\n</tbody>\n</table>\n<h4>Segment-Specific Augmentations</h4>\n<p>The training data contains 12 segment types with distinct characteristics. The augmentation pipeline applies targeted degradations based on segment type:</p>\n<table>\n<thead>\n<tr>\n<th>Segment</th>\n<th>Characteristics</th>\n<th>Applied Augmentations</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0003</td>\n<td>Standard color scans</td>\n<td>Baseline augmentations only</td>\n</tr>\n<tr>\n<td>0004</td>\n<td>Black &amp; white scans</td>\n<td>Grayscale conversion + paper tint</td>\n</tr>\n<tr>\n<td>0005</td>\n<td>Mobile phone photos</td>\n<td>Vignetting + shadow gradients</td>\n</tr>\n<tr>\n<td>0006</td>\n<td>Screen photos</td>\n<td>Moiré patterns + screen glare</td>\n</tr>\n<tr>\n<td>0009</td>\n<td>Water-damaged</td>\n<td>Water stains + color bleeding</td>\n</tr>\n<tr>\n<td>0010</td>\n<td>Extensively damaged</td>\n<td>Fold creases + heavy noise</td>\n</tr>\n<tr>\n<td>0011, 0012</td>\n<td>Moldy scans</td>\n<td>Mold patches + grayscale</td>\n</tr>\n</tbody>\n</table>\n<h4>Augmentation Caching</h4>\n<p>To avoid expensive on-the-fly augmentation operations (especially Perlin noise generation), a <strong>memory-mapped caching system</strong> pre-computes augmentation masks:</p>\n<ul>\n<li>Perlin noise cache: 512 pre-computed noise maps</li>\n<li>Texture cache: 256 patterns</li>\n<li>Shadow/vignette cache: 256 gradient patterns</li>\n</ul>\n<h3>Training Configuration</h3>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Batch Size</td>\n<td>1 (4352×1696 full resolution)</td>\n</tr>\n<tr>\n<td>Optimizer</td>\n<td>AdamW</td>\n</tr>\n<tr>\n<td>Learning Rate</td>\n<td>1e-4</td>\n</tr>\n<tr>\n<td>Weight Decay</td>\n<td>0.05</td>\n</tr>\n<tr>\n<td>Epochs</td>\n<td>80</td>\n</tr>\n<tr>\n<td>EMA Decay</td>\n<td>0.999</td>\n</tr>\n<tr>\n<td>Train/Val Split</td>\n<td>100/0 (full training set)</td>\n</tr>\n<tr>\n<td>Mixed Precision</td>\n<td>FP16 (AMP)</td>\n</tr>\n<tr>\n<td>Gradient Checkpointing</td>\n<td>Enabled for decoder</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Inference Pipeline</h2>\n<h3>Test-Time Augmentation (TTA)</h3>\n<p>Horizontal flip TTA improves robustness:</p>\n<ol>\n<li>Predict on original image</li>\n<li>Horizontally flip input image</li>\n<li>Predict on flipped image</li>\n<li>Flip prediction back to original orientation</li>\n<li>Average both predictions</li>\n</ol>\n<h3>Einthoven's Law Correction</h3>\n<p>A physics-based post-processing step enforces the electrocardiogram constraint:</p>\n<p><strong>Einthoven's Law</strong>: <code>Lead II = Lead I + Lead III</code></p>\n<p>The correction redistributes violation errors across all three limb leads:</p>\n<pre><code>error = II - (I + III)\nI_corrected = I + α × error\nIII_corrected = III + α × error\nII_corrected = II - α × error  # α = 0.33\n</code></pre>\n<h3>Waveform Extraction</h3>\n<p>The 4-strip pixel predictions are converted to waveforms via:</p>\n<ol>\n<li>Apply softmax on y-dimension to get probability distributions</li>\n<li>Compute weighted centroid (argmax or soft-argmax) for each x-column</li>\n<li>Convert pixel coordinates to mV using calibration parameters</li>\n<li>Split each strip into constituent leads based on temporal layout</li>\n<li>Resample to target sample lengths specified in test.csv</li>\n</ol>\n<hr>\n<h2>Ablation Studies</h2>\n<blockquote>\n  <p><strong>Note</strong>: These scores are approximate. Changes were not isolated individually to facilitate faster iteration and save compute costs.</p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>Public LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline (BCE loss + comprehensive augmentations, 90/10 split)</td>\n<td>~15.31</td>\n</tr>\n<tr>\n<td>+ High Resolution (1696×4352)</td>\n<td>~18.22</td>\n</tr>\n<tr>\n<td>+ HRNet-W48 backbone (replacing ResNet)</td>\n<td>~20.90</td>\n</tr>\n<tr>\n<td>+ 100/0 train/val split (use all training data)</td>\n<td>~21.01</td>\n</tr>\n<tr>\n<td>+ Horizontal Flip TTA</td>\n<td>~21.28</td>\n</tr>\n<tr>\n<td>+ Einthoven's Law correction</td>\n<td>~21.32</td>\n</tr>\n<tr>\n<td>+ RGB Skip + PixelShuffle + Fusion Head</td>\n<td>~21.56</td>\n</tr>\n</tbody>\n</table>\n<h3>Key Insights</h3>\n<ol>\n<li><p><strong>Resolution matters</strong>: Increasing from standard resolution to 1696×4352 provided the biggest single improvement (~3 SNR)</p></li>\n<li><p><strong>HRNet superiority</strong>: The multi-resolution parallel branches preserve spatial precision better than traditional encoder-decoder architectures</p></li>\n<li><p><strong>Full training data</strong>: With comprehensive augmentation, using 100% of data for training outperformed keeping a validation set</p></li>\n<li><p><strong>Physics constraints help</strong>: Einthoven correction provides consistent small improvements by enforcing known ECG relationships</p></li>\n</ol>\n<hr>\n<h2>Technical Implementation Details</h2>\n<h3>Memory Optimization</h3>\n<p>Training on high-resolution images (1696×4352) with batch size 1 required careful memory management:</p>\n<ul>\n<li><strong>Gradient Checkpointing</strong>: Enabled for decoder blocks to trade compute for memory</li>\n<li><strong>GroupNorm</strong>: Used instead of BatchNorm for stable single-sample training</li>\n<li><strong>PixelShuffle</strong>: Used instead of ConvTranspose2d for 2× upsampling (more memory efficient)</li>\n<li><strong>Mixed Precision (FP16)</strong>: Enabled via PyTorch AMP</li>\n</ul>\n<h3>Distributed Training</h3>\n<ul>\n<li>DDP (Distributed Data Parallel) support for multi-GPU training</li>\n<li>Per-worker augmentation cache to avoid contention</li>\n<li>Memory-mapped cache files shared across workers</li>\n</ul>\n<hr>\n<h2>Repository Structure</h2>\n<pre><code>├── stage0_model.py          # Stage 0: Marker detection &amp; normalization\n├── stage1_model.py          # Stage 1: Grid detection &amp; rectification\n├── stage2_model.py          # Stage 2: Waveform prediction (main model)\n├── predictor_dataset.py     # Training data loading &amp; preprocessing\n├── predictor_trainer.py     # Training loop with EMA, metrics\n├── augmentations.py         # Comprehensive augmentation pipeline\n├── augmentation_cache.py    # Memory-mapped augmentation caching\n├── demo-submission.py       # End-to-end inference &amp; submission generation\n├── config_predictor.yaml    # Training configuration\n└── archive/\n    ├── stage0-last.checkpoint.pth  # Stage 0 pretrained weights\n    └── stage1-last.checkpoint.pth  # Stage 1 pretrained weights\n</code></pre>\n<hr>\n<h2>Code</h2>\n<ul>\n<li><strong>GitHub Repository</strong>: <a href=\"https://github.com/starrynites/PhysioNet_Digitization_of_ECG_Images\" target=\"_blank\">starrynites/PhysioNet_Digitization_of_ECG_Images</a></li>\n<li><strong>Kaggle Inference Notebook</strong>: <a href=\"https://www.kaggle.com/code/sshiyu/physionet-final-submission\" target=\"_blank\">sshiyu/physionet-final-submission</a></li>\n</ul>\n<h2>Conclusion</h2>\n<p>This solution achieves strong performance on ECG digitization through:</p>\n<ol>\n<li><strong>Robust preprocessing</strong>: Stage 0+1 pipeline handles diverse image sources and geometric distortions</li>\n<li><strong>High-resolution prediction</strong>: Full-resolution HRNet enables precise pixel-level waveform detection</li>\n<li><strong>Comprehensive augmentation</strong>: Segment-aware degradation simulation improves generalization</li>\n<li><strong>Physics-informed post-processing</strong>: Einthoven correction enforces known ECG constraints</li>\n<li><strong>Efficient architecture</strong>: RGB skip connection + PixelShuffle fusion preserves fine details while maintaining training efficiency</li>\n</ol>\n<p>Hope you enjoyed my writeup. Thanks for reading!</p>\n<hr>\n<h2>Acknowledgements</h2>\n<p>This solution builds upon the excellent work shared by the Kaggle community:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">hengck23/demo-submission</a> — Original low resolution baseline</li>\n<li><a href=\"https://www.kaggle.com/code/wasupandceacar/physio-v2-3-public\" target=\"_blank\">wasupandceacar/physio-v2-3-public</a> — High resolution baseline reference</li>\n<li><a href=\"https://www.kaggle.com/code/tonylica/physionet-ecg-streamlined-inference\" target=\"_blank\">tonylica/physionet-ecg-streamlined-inference</a> — Einthoven's law correction idea</li>\n</ul>\n<h3>References</h3>\n<ul>\n<li>Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954</li>\n<li>Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.</li>\n<li>Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Seyedi S, Elola A, Bahrami Rad A, Shah AJ, Bhatia NK, Clifford GD, Sameni R. Digitization and Classification of ECG Images: The George B. Moody PhysioNet Challenge 2024; Computing in Cardiology 2024; 51: 1-4.</li>\n<li>Matthew A. Reyna, Deepanshi, James Weigle, Zuzana Koscova, Kiersten Campbell, Salman Seyedi, Andoni Elola, Ali Bahrami Rad, Amit J Shah, Neal K. Bhatia, Yao Yan, Sohier Dane, Addison Howard, Gari D. Clifford, and Reza Sameni. PhysioNet - Digitization of ECG Images. <a href=\"https://kaggle.com/competitions/physionet-ecg-image-digitization\" target=\"_blank\">https://kaggle.com/competitions/physionet-ecg-image-digitization</a>, 2025. Kaggle.</li>\n</ul>\n<hr>",
  "messages": [
    {
      "id": "3395469",
      "postDate": "01/23/2026 03:36:15",
      "content": "<h1>PhysioNet ECG Digitization Challenge - Solution Writeup</h1>\n<h2>Overview</h2>\n<p>This solution addresses the PhysioNet challenge of digitizing ECG images—converting printed/photographed ECG recordings back into numerical waveform data. The approach uses a <strong>3-stage deep learning pipeline</strong> that progressively transforms raw ECG images into high-fidelity digital waveforms.</p>\n<p><strong>Final Public Leaderboard Score: 21.56 SNR</strong>\n<strong>Final Private Leaderboard Score: 21.37 SNR</strong></p>\n<hr>\n<h2>Problem Statement</h2>\n<p>ECG images from various sources (printed records, mobile phone photos, scans of damaged documents) need to be converted back to digital waveforms. The challenge involves:</p>\n<ol>\n<li><strong>Diverse image sources</strong>: High-quality prints, mobile photos, screen captures, damaged/moldy documents</li>\n<li><strong>Geometric distortions</strong>: Perspective warps, rotations, varying aspect ratios</li>\n<li><strong>Visual degradations</strong>: Water stains, fold creases, mold, low resolution, compression artifacts</li>\n<li><strong>Layout parsing</strong>: Standard 4-strip ECG layout with 12 leads arranged in specific positions</li>\n</ol>\n<hr>\n<h2>Solution Architecture</h2>\n<h3>Three-Stage Pipeline</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2Fea72f9f6a77a90a276213a31965d4755%2Fdiagram1.jpg?generation=1769138841225117&amp;alt=media\" alt=\"Pipeline Diagram\"></p>\n<h3>Stage 0: Normalization</h3>\n<p><strong>Purpose</strong>: Detect lead markers and correct image orientation</p>\n<ul>\n<li><strong>Architecture</strong>: ResNet18 encoder + U-Net decoder</li>\n<li><strong>Outputs</strong>: <ul>\n<li>Lead marker segmentation (13 leads + background)</li>\n<li>Image orientation classification (8 orientations)</li></ul></li>\n<li><strong>Processing</strong>: Applies homography to normalize image to canonical size (3024×4032)</li>\n</ul>\n<blockquote>\n  <p><strong>Note</strong>: For this stage, I utilized the pre-trained Stage 0 checkpoint provided in the <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">original baseline</a>.</p>\n</blockquote>\n<h3>Stage 1: Rectification</h3>\n<p><strong>Purpose</strong>: Detect and correct grid distortions</p>\n<ul>\n<li><strong>Architecture</strong>: ResNet34 encoder + U-Net decoder  </li>\n<li><strong>Outputs</strong>:<ul>\n<li>Grid point detection (keypoints for perspective correction)</li>\n<li>Horizontal line classification (44 classes)</li>\n<li>Vertical line classification (57 classes)</li></ul></li>\n<li><strong>Processing</strong>: Uses detected grid points to apply perspective transformation, producing a perfectly aligned rectified image</li>\n</ul>\n<blockquote>\n  <p><strong>Note</strong>: Similar to Stage 0, I used the pre-trained Stage 1 checkpoint from the <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">original baseline</a>.</p>\n</blockquote>\n<h3>Stage 2: Waveform Prediction</h3>\n<p><strong>Purpose</strong>: Extract ECG waveforms from rectified images</p>\n<p>This is the main prediction stage where most of the innovation lies.</p>\n<h4>Model Architecture</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2F28e82aa96fa43e3593089860871597e3%2Fdiagram2.jpg?generation=1769138929087253&amp;alt=media\" alt=\"Stage 2 Architecture\"></p>\n<h4>Key Design Decisions</h4>\n<ol>\n<li><p><strong>HRNet-W48 Backbone</strong>: Maintains high-resolution representations throughout the network via parallel branches, crucial for precise pixel-level waveform prediction</p></li>\n<li><p><strong>CoordConv Decoder</strong>: Injects normalized 2D coordinates at each decoder block, enabling position-aware predictions essential for ECG layout understanding</p></li>\n<li><p><strong>RGB Skip Connection</strong>: Direct connection from input image to prediction head preserves fine spatial details that may be lost in encoder compression</p></li>\n<li><p><strong>PixelShuffle Upsampling</strong>: Memory-efficient learnable upsampling using sub-pixel convolution instead of transposed convolution</p></li>\n<li><p><strong>GroupNorm</strong>: Enables stable training with batch size 1 (necessary for high-resolution 1696×4352 images on limited GPU memory)</p></li>\n</ol>\n<hr>\n<h2>Training Strategy</h2>\n<h3>Loss Function</h3>\n<p>The model uses <strong>Binary Cross-Entropy (BCE) with positive class weighting</strong> for pixel-level supervision:</p>\n<pre><code>loss = F.binary_cross_entropy_with_logits(\n    pixel_logit,\n    target_mask,\n    pos_weight=10.0  # Handle sparse ECG line pixels\n)\n</code></pre>\n<p>This was found to outperform regression-based losses (MSE, SNR) for this task.</p>\n<h3>Data Augmentation</h3>\n<p>A comprehensive <strong>segment-aware augmentation pipeline</strong> was developed to match the diverse training data distribution:</p>\n<h4>Universal Augmentations</h4>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Probability</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Perlin Noise</td>\n<td>85%</td>\n<td>Simulate scanner noise, paper texture</td>\n</tr>\n<tr>\n<td>Texture Noise</td>\n<td>70%</td>\n<td>Grid-like artifacts, scan lines</td>\n</tr>\n<tr>\n<td>Color Jitter</td>\n<td>80%</td>\n<td>Brightness, contrast, saturation, hue variation</td>\n</tr>\n<tr>\n<td>Gaussian Noise</td>\n<td>50%</td>\n<td>Scanner ISO noise</td>\n</tr>\n<tr>\n<td>Gaussian Blur</td>\n<td>35%</td>\n<td>Low-quality scans, out-of-focus photos</td>\n</tr>\n<tr>\n<td>Shadow Overlay</td>\n<td>35%</td>\n<td>Uneven lighting</td>\n</tr>\n<tr>\n<td>Paper Texture</td>\n<td>30%</td>\n<td>Physical paper grain simulation</td>\n</tr>\n<tr>\n<td>JPEG Artifacts</td>\n<td>30%</td>\n<td>Compression artifacts</td>\n</tr>\n<tr>\n<td>Horizontal Flip</td>\n<td>50%</td>\n<td>Data augmentation (with label flip)</td>\n</tr>\n</tbody>\n</table>\n<h4>Segment-Specific Augmentations</h4>\n<p>The training data contains 12 segment types with distinct characteristics. The augmentation pipeline applies targeted degradations based on segment type:</p>\n<table>\n<thead>\n<tr>\n<th>Segment</th>\n<th>Characteristics</th>\n<th>Applied Augmentations</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0003</td>\n<td>Standard color scans</td>\n<td>Baseline augmentations only</td>\n</tr>\n<tr>\n<td>0004</td>\n<td>Black &amp; white scans</td>\n<td>Grayscale conversion + paper tint</td>\n</tr>\n<tr>\n<td>0005</td>\n<td>Mobile phone photos</td>\n<td>Vignetting + shadow gradients</td>\n</tr>\n<tr>\n<td>0006</td>\n<td>Screen photos</td>\n<td>Moiré patterns + screen glare</td>\n</tr>\n<tr>\n<td>0009</td>\n<td>Water-damaged</td>\n<td>Water stains + color bleeding</td>\n</tr>\n<tr>\n<td>0010</td>\n<td>Extensively damaged</td>\n<td>Fold creases + heavy noise</td>\n</tr>\n<tr>\n<td>0011, 0012</td>\n<td>Moldy scans</td>\n<td>Mold patches + grayscale</td>\n</tr>\n</tbody>\n</table>\n<h4>Augmentation Caching</h4>\n<p>To avoid expensive on-the-fly augmentation operations (especially Perlin noise generation), a <strong>memory-mapped caching system</strong> pre-computes augmentation masks:</p>\n<ul>\n<li>Perlin noise cache: 512 pre-computed noise maps</li>\n<li>Texture cache: 256 patterns</li>\n<li>Shadow/vignette cache: 256 gradient patterns</li>\n</ul>\n<h3>Training Configuration</h3>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Batch Size</td>\n<td>1 (4352×1696 full resolution)</td>\n</tr>\n<tr>\n<td>Optimizer</td>\n<td>AdamW</td>\n</tr>\n<tr>\n<td>Learning Rate</td>\n<td>1e-4</td>\n</tr>\n<tr>\n<td>Weight Decay</td>\n<td>0.05</td>\n</tr>\n<tr>\n<td>Epochs</td>\n<td>80</td>\n</tr>\n<tr>\n<td>EMA Decay</td>\n<td>0.999</td>\n</tr>\n<tr>\n<td>Train/Val Split</td>\n<td>100/0 (full training set)</td>\n</tr>\n<tr>\n<td>Mixed Precision</td>\n<td>FP16 (AMP)</td>\n</tr>\n<tr>\n<td>Gradient Checkpointing</td>\n<td>Enabled for decoder</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Inference Pipeline</h2>\n<h3>Test-Time Augmentation (TTA)</h3>\n<p>Horizontal flip TTA improves robustness:</p>\n<ol>\n<li>Predict on original image</li>\n<li>Horizontally flip input image</li>\n<li>Predict on flipped image</li>\n<li>Flip prediction back to original orientation</li>\n<li>Average both predictions</li>\n</ol>\n<h3>Einthoven's Law Correction</h3>\n<p>A physics-based post-processing step enforces the electrocardiogram constraint:</p>\n<p><strong>Einthoven's Law</strong>: <code>Lead II = Lead I + Lead III</code></p>\n<p>The correction redistributes violation errors across all three limb leads:</p>\n<pre><code>error = II - (I + III)\nI_corrected = I + α × error\nIII_corrected = III + α × error\nII_corrected = II - α × error  # α = 0.33\n</code></pre>\n<h3>Waveform Extraction</h3>\n<p>The 4-strip pixel predictions are converted to waveforms via:</p>\n<ol>\n<li>Apply softmax on y-dimension to get probability distributions</li>\n<li>Compute weighted centroid (argmax or soft-argmax) for each x-column</li>\n<li>Convert pixel coordinates to mV using calibration parameters</li>\n<li>Split each strip into constituent leads based on temporal layout</li>\n<li>Resample to target sample lengths specified in test.csv</li>\n</ol>\n<hr>\n<h2>Ablation Studies</h2>\n<blockquote>\n  <p><strong>Note</strong>: These scores are approximate. Changes were not isolated individually to facilitate faster iteration and save compute costs.</p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>Public LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline (BCE loss + comprehensive augmentations, 90/10 split)</td>\n<td>~15.31</td>\n</tr>\n<tr>\n<td>+ High Resolution (1696×4352)</td>\n<td>~18.22</td>\n</tr>\n<tr>\n<td>+ HRNet-W48 backbone (replacing ResNet)</td>\n<td>~20.90</td>\n</tr>\n<tr>\n<td>+ 100/0 train/val split (use all training data)</td>\n<td>~21.01</td>\n</tr>\n<tr>\n<td>+ Horizontal Flip TTA</td>\n<td>~21.28</td>\n</tr>\n<tr>\n<td>+ Einthoven's Law correction</td>\n<td>~21.32</td>\n</tr>\n<tr>\n<td>+ RGB Skip + PixelShuffle + Fusion Head</td>\n<td>~21.56</td>\n</tr>\n</tbody>\n</table>\n<h3>Key Insights</h3>\n<ol>\n<li><p><strong>Resolution matters</strong>: Increasing from standard resolution to 1696×4352 provided the biggest single improvement (~3 SNR)</p></li>\n<li><p><strong>HRNet superiority</strong>: The multi-resolution parallel branches preserve spatial precision better than traditional encoder-decoder architectures</p></li>\n<li><p><strong>Full training data</strong>: With comprehensive augmentation, using 100% of data for training outperformed keeping a validation set</p></li>\n<li><p><strong>Physics constraints help</strong>: Einthoven correction provides consistent small improvements by enforcing known ECG relationships</p></li>\n</ol>\n<hr>\n<h2>Technical Implementation Details</h2>\n<h3>Memory Optimization</h3>\n<p>Training on high-resolution images (1696×4352) with batch size 1 required careful memory management:</p>\n<ul>\n<li><strong>Gradient Checkpointing</strong>: Enabled for decoder blocks to trade compute for memory</li>\n<li><strong>GroupNorm</strong>: Used instead of BatchNorm for stable single-sample training</li>\n<li><strong>PixelShuffle</strong>: Used instead of ConvTranspose2d for 2× upsampling (more memory efficient)</li>\n<li><strong>Mixed Precision (FP16)</strong>: Enabled via PyTorch AMP</li>\n</ul>\n<h3>Distributed Training</h3>\n<ul>\n<li>DDP (Distributed Data Parallel) support for multi-GPU training</li>\n<li>Per-worker augmentation cache to avoid contention</li>\n<li>Memory-mapped cache files shared across workers</li>\n</ul>\n<hr>\n<h2>Repository Structure</h2>\n<pre><code>├── stage0_model.py          # Stage 0: Marker detection &amp; normalization\n├── stage1_model.py          # Stage 1: Grid detection &amp; rectification\n├── stage2_model.py          # Stage 2: Waveform prediction (main model)\n├── predictor_dataset.py     # Training data loading &amp; preprocessing\n├── predictor_trainer.py     # Training loop with EMA, metrics\n├── augmentations.py         # Comprehensive augmentation pipeline\n├── augmentation_cache.py    # Memory-mapped augmentation caching\n├── demo-submission.py       # End-to-end inference &amp; submission generation\n├── config_predictor.yaml    # Training configuration\n└── archive/\n    ├── stage0-last.checkpoint.pth  # Stage 0 pretrained weights\n    └── stage1-last.checkpoint.pth  # Stage 1 pretrained weights\n</code></pre>\n<hr>\n<h2>Code</h2>\n<ul>\n<li><strong>GitHub Repository</strong>: <a href=\"https://github.com/starrynites/PhysioNet_Digitization_of_ECG_Images\" target=\"_blank\">starrynites/PhysioNet_Digitization_of_ECG_Images</a></li>\n<li><strong>Kaggle Inference Notebook</strong>: <a href=\"https://www.kaggle.com/code/sshiyu/physionet-final-submission\" target=\"_blank\">sshiyu/physionet-final-submission</a></li>\n</ul>\n<h2>Conclusion</h2>\n<p>This solution achieves strong performance on ECG digitization through:</p>\n<ol>\n<li><strong>Robust preprocessing</strong>: Stage 0+1 pipeline handles diverse image sources and geometric distortions</li>\n<li><strong>High-resolution prediction</strong>: Full-resolution HRNet enables precise pixel-level waveform detection</li>\n<li><strong>Comprehensive augmentation</strong>: Segment-aware degradation simulation improves generalization</li>\n<li><strong>Physics-informed post-processing</strong>: Einthoven correction enforces known ECG constraints</li>\n<li><strong>Efficient architecture</strong>: RGB skip connection + PixelShuffle fusion preserves fine details while maintaining training efficiency</li>\n</ol>\n<p>Hope you enjoyed my writeup. Thanks for reading!</p>\n<hr>\n<h2>Acknowledgements</h2>\n<p>This solution builds upon the excellent work shared by the Kaggle community:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">hengck23/demo-submission</a> — Original low resolution baseline</li>\n<li><a href=\"https://www.kaggle.com/code/wasupandceacar/physio-v2-3-public\" target=\"_blank\">wasupandceacar/physio-v2-3-public</a> — High resolution baseline reference</li>\n<li><a href=\"https://www.kaggle.com/code/tonylica/physionet-ecg-streamlined-inference\" target=\"_blank\">tonylica/physionet-ecg-streamlined-inference</a> — Einthoven's law correction idea</li>\n</ul>\n<h3>References</h3>\n<ul>\n<li>Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954</li>\n<li>Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.</li>\n<li>Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Seyedi S, Elola A, Bahrami Rad A, Shah AJ, Bhatia NK, Clifford GD, Sameni R. Digitization and Classification of ECG Images: The George B. Moody PhysioNet Challenge 2024; Computing in Cardiology 2024; 51: 1-4.</li>\n<li>Matthew A. Reyna, Deepanshi, James Weigle, Zuzana Koscova, Kiersten Campbell, Salman Seyedi, Andoni Elola, Ali Bahrami Rad, Amit J Shah, Neal K. Bhatia, Yao Yan, Sohier Dane, Addison Howard, Gari D. Clifford, and Reza Sameni. PhysioNet - Digitization of ECG Images. <a href=\"https://kaggle.com/competitions/physionet-ecg-image-digitization\" target=\"_blank\">https://kaggle.com/competitions/physionet-ecg-image-digitization</a>, 2025. Kaggle.</li>\n</ul>\n<hr>",
      "rawMarkdown": "# PhysioNet ECG Digitization Challenge - Solution Writeup\n\n## Overview\nThis solution addresses the PhysioNet challenge of digitizing ECG images—converting printed/photographed ECG recordings back into numerical waveform data. The approach uses a **3-stage deep learning pipeline** that progressively transforms raw ECG images into high-fidelity digital waveforms.\n\n**Final Public Leaderboard Score: 21.56 SNR**\n**Final Private Leaderboard Score: 21.37 SNR**\n\n---\n\n## Problem Statement\n\nECG images from various sources (printed records, mobile phone photos, scans of damaged documents) need to be converted back to digital waveforms. The challenge involves:\n\n1. **Diverse image sources**: High-quality prints, mobile photos, screen captures, damaged/moldy documents\n2. **Geometric distortions**: Perspective warps, rotations, varying aspect ratios\n3. **Visual degradations**: Water stains, fold creases, mold, low resolution, compression artifacts\n4. **Layout parsing**: Standard 4-strip ECG layout with 12 leads arranged in specific positions\n\n---\n## Solution Architecture\n\n### Three-Stage Pipeline\n\n![Pipeline Diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2Fea72f9f6a77a90a276213a31965d4755%2Fdiagram1.jpg?generation=1769138841225117&alt=media)\n\n### Stage 0: Normalization\n\n**Purpose**: Detect lead markers and correct image orientation\n\n- **Architecture**: ResNet18 encoder + U-Net decoder\n- **Outputs**: \n  - Lead marker segmentation (13 leads + background)\n  - Image orientation classification (8 orientations)\n- **Processing**: Applies homography to normalize image to canonical size (3024×4032)\n\n> **Note**: For this stage, I utilized the pre-trained Stage 0 checkpoint provided in the [original baseline](https://www.kaggle.com/code/hengck23/demo-submission).\n\n### Stage 1: Rectification\n\n**Purpose**: Detect and correct grid distortions\n\n- **Architecture**: ResNet34 encoder + U-Net decoder  \n- **Outputs**:\n  - Grid point detection (keypoints for perspective correction)\n  - Horizontal line classification (44 classes)\n  - Vertical line classification (57 classes)\n- **Processing**: Uses detected grid points to apply perspective transformation, producing a perfectly aligned rectified image\n\n> **Note**: Similar to Stage 0, I used the pre-trained Stage 1 checkpoint from the [original baseline](https://www.kaggle.com/code/hengck23/demo-submission).\n\n### Stage 2: Waveform Prediction\n\n**Purpose**: Extract ECG waveforms from rectified images\n\nThis is the main prediction stage where most of the innovation lies.\n\n#### Model Architecture\n\n![Stage 2 Architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2F28e82aa96fa43e3593089860871597e3%2Fdiagram2.jpg?generation=1769138929087253&alt=media)\n\n#### Key Design Decisions\n\n1. **HRNet-W48 Backbone**: Maintains high-resolution representations throughout the network via parallel branches, crucial for precise pixel-level waveform prediction\n\n2. **CoordConv Decoder**: Injects normalized 2D coordinates at each decoder block, enabling position-aware predictions essential for ECG layout understanding\n\n3. **RGB Skip Connection**: Direct connection from input image to prediction head preserves fine spatial details that may be lost in encoder compression\n\n4. **PixelShuffle Upsampling**: Memory-efficient learnable upsampling using sub-pixel convolution instead of transposed convolution\n\n5. **GroupNorm**: Enables stable training with batch size 1 (necessary for high-resolution 1696×4352 images on limited GPU memory)\n\n---\n\n## Training Strategy\n\n### Loss Function\n\nThe model uses **Binary Cross-Entropy (BCE) with positive class weighting** for pixel-level supervision:\n\n```python\nloss = F.binary_cross_entropy_with_logits(\n    pixel_logit,\n    target_mask,\n    pos_weight=10.0  # Handle sparse ECG line pixels\n)\n```\n\nThis was found to outperform regression-based losses (MSE, SNR) for this task.\n\n### Data Augmentation\n\nA comprehensive **segment-aware augmentation pipeline** was developed to match the diverse training data distribution:\n\n#### Universal Augmentations\n| Augmentation | Probability | Purpose |\n|-------------|-------------|---------|\n| Perlin Noise | 85% | Simulate scanner noise, paper texture |\n| Texture Noise | 70% | Grid-like artifacts, scan lines |\n| Color Jitter | 80% | Brightness, contrast, saturation, hue variation |\n| Gaussian Noise | 50% | Scanner ISO noise |\n| Gaussian Blur | 35% | Low-quality scans, out-of-focus photos |\n| Shadow Overlay | 35% | Uneven lighting |\n| Paper Texture | 30% | Physical paper grain simulation |\n| JPEG Artifacts | 30% | Compression artifacts |\n| Horizontal Flip | 50% | Data augmentation (with label flip) |\n\n#### Segment-Specific Augmentations\n\nThe training data contains 12 segment types with distinct characteristics. The augmentation pipeline applies targeted degradations based on segment type:\n\n| Segment | Characteristics | Applied Augmentations |\n|---------|-----------------|----------------------|\n| 0003 | Standard color scans | Baseline augmentations only |\n| 0004 | Black & white scans | Grayscale conversion + paper tint |\n| 0005 | Mobile phone photos | Vignetting + shadow gradients |\n| 0006 | Screen photos | Moiré patterns + screen glare |\n| 0009 | Water-damaged | Water stains + color bleeding |\n| 0010 | Extensively damaged | Fold creases + heavy noise |\n| 0011, 0012 | Moldy scans | Mold patches + grayscale |\n\n#### Augmentation Caching\n\nTo avoid expensive on-the-fly augmentation operations (especially Perlin noise generation), a **memory-mapped caching system** pre-computes augmentation masks:\n\n- Perlin noise cache: 512 pre-computed noise maps\n- Texture cache: 256 patterns\n- Shadow/vignette cache: 256 gradient patterns\n\n### Training Configuration\n\n| Parameter | Value |\n|-----------|-------|\n| Batch Size | 1 (4352×1696 full resolution) |\n| Optimizer | AdamW |\n| Learning Rate | 1e-4 |\n| Weight Decay | 0.05 |\n| Epochs | 80 |\n| EMA Decay | 0.999 |\n| Train/Val Split | 100/0 (full training set) |\n| Mixed Precision | FP16 (AMP) |\n| Gradient Checkpointing | Enabled for decoder |\n\n---\n\n## Inference Pipeline\n\n### Test-Time Augmentation (TTA)\n\nHorizontal flip TTA improves robustness:\n\n1. Predict on original image\n2. Horizontally flip input image\n3. Predict on flipped image\n4. Flip prediction back to original orientation\n5. Average both predictions\n\n### Einthoven's Law Correction\n\nA physics-based post-processing step enforces the electrocardiogram constraint:\n\n**Einthoven's Law**: `Lead II = Lead I + Lead III`\n\nThe correction redistributes violation errors across all three limb leads:\n\n```python\nerror = II - (I + III)\nI_corrected = I + α × error\nIII_corrected = III + α × error\nII_corrected = II - α × error  # α = 0.33\n```\n\n### Waveform Extraction\n\nThe 4-strip pixel predictions are converted to waveforms via:\n\n1. Apply softmax on y-dimension to get probability distributions\n2. Compute weighted centroid (argmax or soft-argmax) for each x-column\n3. Convert pixel coordinates to mV using calibration parameters\n4. Split each strip into constituent leads based on temporal layout\n5. Resample to target sample lengths specified in test.csv\n\n---\n\n## Ablation Studies\n\n> **Note**: These scores are approximate. Changes were not isolated individually to facilitate faster iteration and save compute costs.\n\n| Configuration | Public LB Score |\n|---------------|-----------------|\n| Baseline (BCE loss + comprehensive augmentations, 90/10 split) | ~15.31 |\n| + High Resolution (1696×4352) | ~18.22 |\n| + HRNet-W48 backbone (replacing ResNet) | ~20.90 |\n| + 100/0 train/val split (use all training data) | ~21.01 |\n| + Horizontal Flip TTA | ~21.28 |\n| + Einthoven's Law correction | ~21.32 |\n| + RGB Skip + PixelShuffle + Fusion Head | ~21.56 |\n\n### Key Insights\n\n1. **Resolution matters**: Increasing from standard resolution to 1696×4352 provided the biggest single improvement (~3 SNR)\n\n2. **HRNet superiority**: The multi-resolution parallel branches preserve spatial precision better than traditional encoder-decoder architectures\n\n3. **Full training data**: With comprehensive augmentation, using 100% of data for training outperformed keeping a validation set\n\n4. **Physics constraints help**: Einthoven correction provides consistent small improvements by enforcing known ECG relationships\n\n---\n\n## Technical Implementation Details\n\n### Memory Optimization\n\nTraining on high-resolution images (1696×4352) with batch size 1 required careful memory management:\n\n- **Gradient Checkpointing**: Enabled for decoder blocks to trade compute for memory\n- **GroupNorm**: Used instead of BatchNorm for stable single-sample training\n- **PixelShuffle**: Used instead of ConvTranspose2d for 2× upsampling (more memory efficient)\n- **Mixed Precision (FP16)**: Enabled via PyTorch AMP\n\n### Distributed Training\n\n- DDP (Distributed Data Parallel) support for multi-GPU training\n- Per-worker augmentation cache to avoid contention\n- Memory-mapped cache files shared across workers\n\n---\n\n## Repository Structure\n\n```\n├── stage0_model.py          # Stage 0: Marker detection & normalization\n├── stage1_model.py          # Stage 1: Grid detection & rectification\n├── stage2_model.py          # Stage 2: Waveform prediction (main model)\n├── predictor_dataset.py     # Training data loading & preprocessing\n├── predictor_trainer.py     # Training loop with EMA, metrics\n├── augmentations.py         # Comprehensive augmentation pipeline\n├── augmentation_cache.py    # Memory-mapped augmentation caching\n├── demo-submission.py       # End-to-end inference & submission generation\n├── config_predictor.yaml    # Training configuration\n└── archive/\n    ├── stage0-last.checkpoint.pth  # Stage 0 pretrained weights\n    └── stage1-last.checkpoint.pth  # Stage 1 pretrained weights\n```\n\n---\n\n## Code\n\n- **GitHub Repository**: [starrynites/PhysioNet_Digitization_of_ECG_Images](https://github.com/starrynites/PhysioNet_Digitization_of_ECG_Images)\n- **Kaggle Inference Notebook**: [sshiyu/physionet-final-submission](https://www.kaggle.com/code/sshiyu/physionet-final-submission)\n\n## Conclusion\n\nThis solution achieves strong performance on ECG digitization through:\n\n1. **Robust preprocessing**: Stage 0+1 pipeline handles diverse image sources and geometric distortions\n2. **High-resolution prediction**: Full-resolution HRNet enables precise pixel-level waveform detection\n3. **Comprehensive augmentation**: Segment-aware degradation simulation improves generalization\n4. **Physics-informed post-processing**: Einthoven correction enforces known ECG constraints\n5. **Efficient architecture**: RGB skip connection + PixelShuffle fusion preserves fine details while maintaining training efficiency\n\nHope you enjoyed my writeup. Thanks for reading!\n\n---\n\n## Acknowledgements\n\nThis solution builds upon the excellent work shared by the Kaggle community:\n\n- [hengck23/demo-submission](https://www.kaggle.com/code/hengck23/demo-submission) — Original low resolution baseline\n- [wasupandceacar/physio-v2-3-public](https://www.kaggle.com/code/wasupandceacar/physio-v2-3-public) — High resolution baseline reference\n- [tonylica/physionet-ecg-streamlined-inference](https://www.kaggle.com/code/tonylica/physionet-ecg-streamlined-inference) — Einthoven's law correction idea\n\n### References\n\n- Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954\n- Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.\n- Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Seyedi S, Elola A, Bahrami Rad A, Shah AJ, Bhatia NK, Clifford GD, Sameni R. Digitization and Classification of ECG Images: The George B. Moody PhysioNet Challenge 2024; Computing in Cardiology 2024; 51: 1-4.\n- Matthew A. Reyna, Deepanshi, James Weigle, Zuzana Koscova, Kiersten Campbell, Salman Seyedi, Andoni Elola, Ali Bahrami Rad, Amit J Shah, Neal K. Bhatia, Yao Yan, Sohier Dane, Addison Howard, Gari D. Clifford, and Reza Sameni. PhysioNet - Digitization of ECG Images. https://kaggle.com/competitions/physionet-ecg-image-digitization, 2025. Kaggle.\n\n---",
      "votes": null
    },
    {
      "id": "3395554",
      "postDate": "01/23/2026 07:45:49",
      "content": "<p>Thank you so much for sharing such a detailed and thorough solution! When I was going through your code and solution, I noticed that you use all types of training images after rectification as the input for Stage 2 during training. However, in my tests, the rectified non-0001 type images are not perfectly aligned with the GT signals. I wonder how you addressed this issue, or if I have any misunderstanding of your solution.</p>",
      "rawMarkdown": "Thank you so much for sharing such a detailed and thorough solution! When I was going through your code and solution, I noticed that you use all types of training images after rectification as the input for Stage 2 during training. However, in my tests, the rectified non-0001 type images are not perfectly aligned with the GT signals. I wonder how you addressed this issue, or if I have any misunderstanding of your solution.",
      "votes": null
    },
    {
      "id": "3395576",
      "postDate": "01/23/2026 08:55:46",
      "content": "<p>That is a great observation! I also noticed that due to rectification errors in stage 0 and 1, the GT signals are not always perfectly aligned with what we see visually (mostly due to horizontal misalignment).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14897182%2F671aa7d0472f2321356fc7045ed06e58%2FScreenshot%202026-01-01%20at%2011.42.46PM.png?generation=1769156126338577&amp;alt=media\" alt=\"Visualisation of ground truth horizontal misalignment\"></p>\n<p>I do agree that this creates some training confusion. However, I believe that this actually does not significantly hurt performance for the following reasons:</p>\n<ol>\n<li><strong>Most samples are reasonably well-aligned</strong> - Stage 1 rectification works well for the majority of images</li>\n<li><strong>The spatial prior is \"soft\" rather than \"hard\"</strong> - the model can adapt to misalignments</li>\n</ol>\n<h2>Soft vs Hard Spatial Prior</h2>\n<p>The key distinction is <strong>where and how</strong> the timespan constraint is applied:</p>\n<h3>Hard Spatial Prior (NOT what I do)</h3>\n<pre><code># Hypothetical: Model architecture prevents predictions outside timespan\nclass HardConstrainedModel:\n    def forward(self, batch):\n        x = batch['image']\n        features = self.encoder(x)\n        output = torch.zeros(B, 4, H, W)\n        # Only allow decoder to predict in specific columns\n        output[:, :, :, 235:4161] = self.decoder(features)[:, :, :, 235:4161]\n        return {'pixel_logit': output}  # Forced zeros outside [235, 4161]\n</code></pre>\n<p>This would completely prevent the model from learning waveform features outside [235-4161].</p>\n<h3>Soft Spatial Prior (My Implementation)</h3>\n<p><strong>Model architecture</strong> (<code>stage2_model.py</code>):</p>\n<pre><code>def forward(self, batch, L=None):\n    image = batch['image']  # (B, 3, H, W)\n    # ... encoder + decoder ...\n    pixel_logit = self.pixel_out(last)  # (B, 4, H, W) - FULL width prediction\n    return {'pixel_logit': pixel_logit}  # No spatial masking\n</code></pre>\n<p><strong>Training</strong> (<code>predictor_trainer.py</code>):</p>\n<pre><code># Loss is computed on FULL image\npixel_loss = F.binary_cross_entropy_with_logits(\n    pixel_logit,     # (B, 4, H, W=4352) - all columns\n    pixel_target,    # (B, 4, H, W=4352) - GT drawn at [235, 4161]\n    pos_weight=pos_weight,\n)\n# Model learns to activate [235-4161] due to GT distribution,\n# but architecture doesn't prevent activations elsewhere\n</code></pre>\n<p><strong>Inference</strong> (<code>predictor_trainer.py</code>):</p>\n<pre><code>def _predict_waveform(self, image_u8, target_len):\n    out = self.model(batch)\n    pixel_logit = out['pixel_logit']  # (B, 4, H, W) - full width\n\n    # Convert ALL columns to y-positions\n    y_hat = pixel_logits_to_y(pixel_logit, ...)  # (B, 4, W)\n\n    # Crop to timespan AFTER conversion (post-processing, not architectural)\n    y_hat = y_hat[..., self.t0:self.t1]  # (B, 4, 3926)\n\n    mv_hat = y_to_mv(y_hat, ...)\n    return mv_hat\n</code></pre>\n<h2>Why This Matters for Misalignment</h2>\n<p>For a misaligned sample where the waveform appears at columns [200-4126] instead of [235-4161]:</p>\n<p><strong>With Hard Prior:</strong></p>\n<ul>\n<li>Model cannot predict at [200-234] (architecture constraint)</li>\n<li>Training fails - visual features are blocked</li>\n</ul>\n<p><strong>With Soft Prior:</strong></p>\n<ul>\n<li>Model processes visual features at [200-234] through convolutional layers</li>\n<li>Training encourages predictions at [235-4161] via GT loss</li>\n<li>Model learns a compromise weighted by data distribution</li>\n<li>Inference crops to [235-4161], losing edge [200-234] but keeping [235-4126]</li>\n</ul>\n<p>The GT is drawn into [235-4161], but the model sees and processes the entire image during training. Visual features from misaligned regions still contribute to learning through the convolutional receptive field.</p>\n<p>The system degrades gracefully rather than failing catastrophically.</p>\n<p>This design choice accepts some label noise from imperfect rectifications in exchange for:</p>\n<ul>\n<li>More diverse training data (all segment types)</li>\n<li>Improved robustness to real-world rectification errors</li>\n</ul>",
      "rawMarkdown": "That is a great observation! I also noticed that due to rectification errors in stage 0 and 1, the GT signals are not always perfectly aligned with what we see visually (mostly due to horizontal misalignment).\n\n![Visualisation of ground truth horizontal misalignment](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14897182%2F671aa7d0472f2321356fc7045ed06e58%2FScreenshot%202026-01-01%20at%2011.42.46PM.png?generation=1769156126338577&alt=media)\n\nI do agree that this creates some training confusion. However, I believe that this actually does not significantly hurt performance for the following reasons:\n\n1. **Most samples are reasonably well-aligned** - Stage 1 rectification works well for the majority of images\n2. **The spatial prior is \"soft\" rather than \"hard\"** - the model can adapt to misalignments\n\n## Soft vs Hard Spatial Prior\n\nThe key distinction is **where and how** the timespan constraint is applied:\n\n### Hard Spatial Prior (NOT what I do)\n\n```python\n# Hypothetical: Model architecture prevents predictions outside timespan\nclass HardConstrainedModel:\n    def forward(self, batch):\n        x = batch['image']\n        features = self.encoder(x)\n        output = torch.zeros(B, 4, H, W)\n        # Only allow decoder to predict in specific columns\n        output[:, :, :, 235:4161] = self.decoder(features)[:, :, :, 235:4161]\n        return {'pixel_logit': output}  # Forced zeros outside [235, 4161]\n```\n\nThis would completely prevent the model from learning waveform features outside [235-4161].\n\n### Soft Spatial Prior (My Implementation)\n\n**Model architecture** (`stage2_model.py`):\n\n```python\ndef forward(self, batch, L=None):\n    image = batch['image']  # (B, 3, H, W)\n    # ... encoder + decoder ...\n    pixel_logit = self.pixel_out(last)  # (B, 4, H, W) - FULL width prediction\n    return {'pixel_logit': pixel_logit}  # No spatial masking\n```\n\n**Training** (`predictor_trainer.py`):\n\n```python\n# Loss is computed on FULL image\npixel_loss = F.binary_cross_entropy_with_logits(\n    pixel_logit,     # (B, 4, H, W=4352) - all columns\n    pixel_target,    # (B, 4, H, W=4352) - GT drawn at [235, 4161]\n    pos_weight=pos_weight,\n)\n# Model learns to activate [235-4161] due to GT distribution,\n# but architecture doesn't prevent activations elsewhere\n```\n\n**Inference** (`predictor_trainer.py`):\n\n```python\ndef _predict_waveform(self, image_u8, target_len):\n    out = self.model(batch)\n    pixel_logit = out['pixel_logit']  # (B, 4, H, W) - full width\n    \n    # Convert ALL columns to y-positions\n    y_hat = pixel_logits_to_y(pixel_logit, ...)  # (B, 4, W)\n    \n    # Crop to timespan AFTER conversion (post-processing, not architectural)\n    y_hat = y_hat[..., self.t0:self.t1]  # (B, 4, 3926)\n    \n    mv_hat = y_to_mv(y_hat, ...)\n    return mv_hat\n```\n\n## Why This Matters for Misalignment\n\nFor a misaligned sample where the waveform appears at columns [200-4126] instead of [235-4161]:\n\n**With Hard Prior:**\n- Model cannot predict at [200-234] (architecture constraint)\n- Training fails - visual features are blocked\n\n**With Soft Prior:**\n- Model processes visual features at [200-234] through convolutional layers\n- Training encourages predictions at [235-4161] via GT loss\n- Model learns a compromise weighted by data distribution\n- Inference crops to [235-4161], losing edge [200-234] but keeping [235-4126]\n\nThe GT is drawn into [235-4161], but the model sees and processes the entire image during training. Visual features from misaligned regions still contribute to learning through the convolutional receptive field.\n\nThe system degrades gracefully rather than failing catastrophically.\n\nThis design choice accepts some label noise from imperfect rectifications in exchange for:\n- More diverse training data (all segment types)\n- Improved robustness to real-world rectification errors",
      "votes": null
    },
    {
      "id": "3395721",
      "postDate": "01/23/2026 14:44:42",
      "content": "<p>Thank you so much for your patient explanation. Your answer has cleared up my confusion. I have one more question: when constructing the training data, did you manually check and exclude some severely misaligned samples, or did you use all the rectified images directly?</p>",
      "rawMarkdown": "Thank you so much for your patient explanation. Your answer has cleared up my confusion. I have one more question: when constructing the training data, did you manually check and exclude some severely misaligned samples, or did you use all the rectified images directly?",
      "votes": null
    },
    {
      "id": "3395753",
      "postDate": "01/23/2026 15:57:51",
      "content": "<p>I used all rectified images directly without manual filtering or exclusion.</p>\n<p>I generally avoid removing data due to fear of information loss. In my experience, severely misaligned samples are quite rare - the vast majority of rectifications from Stage 0+1 are reasonably accurate. The few badly misaligned samples act as noise in the training set, but with enough well-aligned data, the model naturally learns to filter out these outliers and converge to the correct visual features.</p>\n<p>The benefits of keeping all data (more diverse training samples, better coverage of edge cases) outweighed the cost of a small amount of label noise from misalignments.</p>",
      "rawMarkdown": "I used all rectified images directly without manual filtering or exclusion.\n\nI generally avoid removing data due to fear of information loss. In my experience, severely misaligned samples are quite rare - the vast majority of rectifications from Stage 0+1 are reasonably accurate. The few badly misaligned samples act as noise in the training set, but with enough well-aligned data, the model naturally learns to filter out these outliers and converge to the correct visual features.\n\nThe benefits of keeping all data (more diverse training samples, better coverage of edge cases) outweighed the cost of a small amount of label noise from misalignments.",
      "votes": null
    },
    {
      "id": "3395806",
      "postDate": "01/23/2026 17:16:58",
      "content": "<p>Regarding the alignment issue:\nI think teams that use direct regression (i.e., mapping from images to values) will not have this issue. This is because crooked lines provide cues for the network to adjust misaligned pixel locations and still predict the correct values.</p>",
      "rawMarkdown": "Regarding the alignment issue:\nI think teams that use direct regression (i.e., mapping from images to values) will not have this issue. This is because crooked lines provide cues for the network to adjust misaligned pixel locations and still predict the correct values.",
      "votes": null
    },
    {
      "id": "3395917",
      "postDate": "01/23/2026 23:44:16",
      "content": "<p>I think so too. And thanks again for your great work, it helps us a lot.😃</p>",
      "rawMarkdown": "I think so too. And thanks again for your great work, it helps us a lot.😃",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3395554,
      "author_name": "xie233",
      "author_url": "",
      "post_date": "01/23/2026 07:45:49",
      "content": "<p>Thank you so much for sharing such a detailed and thorough solution! When I was going through your code and solution, I noticed that you use all types of training images after rectification as the input for Stage 2 during training. However, in my tests, the rectified non-0001 type images are not perfectly aligned with the GT signals. I wonder how you addressed this issue, or if I have any misunderstanding of your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3395576,
          "author_name": "sshiyu",
          "author_url": "",
          "post_date": "01/23/2026 08:55:46",
          "content": "<p>That is a great observation! I also noticed that due to rectification errors in stage 0 and 1, the GT signals are not always perfectly aligned with what we see visually (mostly due to horizontal misalignment).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14897182%2F671aa7d0472f2321356fc7045ed06e58%2FScreenshot%202026-01-01%20at%2011.42.46PM.png?generation=1769156126338577&amp;alt=media\" alt=\"Visualisation of ground truth horizontal misalignment\"></p>\n<p>I do agree that this creates some training confusion. However, I believe that this actually does not significantly hurt performance for the following reasons:</p>\n<ol>\n<li><strong>Most samples are reasonably well-aligned</strong> - Stage 1 rectification works well for the majority of images</li>\n<li><strong>The spatial prior is \"soft\" rather than \"hard\"</strong> - the model can adapt to misalignments</li>\n</ol>\n<h2>Soft vs Hard Spatial Prior</h2>\n<p>The key distinction is <strong>where and how</strong> the timespan constraint is applied:</p>\n<h3>Hard Spatial Prior (NOT what I do)</h3>\n<pre><code># Hypothetical: Model architecture prevents predictions outside timespan\nclass HardConstrainedModel:\n    def forward(self, batch):\n        x = batch['image']\n        features = self.encoder(x)\n        output = torch.zeros(B, 4, H, W)\n        # Only allow decoder to predict in specific columns\n        output[:, :, :, 235:4161] = self.decoder(features)[:, :, :, 235:4161]\n        return {'pixel_logit': output}  # Forced zeros outside [235, 4161]\n</code></pre>\n<p>This would completely prevent the model from learning waveform features outside [235-4161].</p>\n<h3>Soft Spatial Prior (My Implementation)</h3>\n<p><strong>Model architecture</strong> (<code>stage2_model.py</code>):</p>\n<pre><code>def forward(self, batch, L=None):\n    image = batch['image']  # (B, 3, H, W)\n    # ... encoder + decoder ...\n    pixel_logit = self.pixel_out(last)  # (B, 4, H, W) - FULL width prediction\n    return {'pixel_logit': pixel_logit}  # No spatial masking\n</code></pre>\n<p><strong>Training</strong> (<code>predictor_trainer.py</code>):</p>\n<pre><code># Loss is computed on FULL image\npixel_loss = F.binary_cross_entropy_with_logits(\n    pixel_logit,     # (B, 4, H, W=4352) - all columns\n    pixel_target,    # (B, 4, H, W=4352) - GT drawn at [235, 4161]\n    pos_weight=pos_weight,\n)\n# Model learns to activate [235-4161] due to GT distribution,\n# but architecture doesn't prevent activations elsewhere\n</code></pre>\n<p><strong>Inference</strong> (<code>predictor_trainer.py</code>):</p>\n<pre><code>def _predict_waveform(self, image_u8, target_len):\n    out = self.model(batch)\n    pixel_logit = out['pixel_logit']  # (B, 4, H, W) - full width\n\n    # Convert ALL columns to y-positions\n    y_hat = pixel_logits_to_y(pixel_logit, ...)  # (B, 4, W)\n\n    # Crop to timespan AFTER conversion (post-processing, not architectural)\n    y_hat = y_hat[..., self.t0:self.t1]  # (B, 4, 3926)\n\n    mv_hat = y_to_mv(y_hat, ...)\n    return mv_hat\n</code></pre>\n<h2>Why This Matters for Misalignment</h2>\n<p>For a misaligned sample where the waveform appears at columns [200-4126] instead of [235-4161]:</p>\n<p><strong>With Hard Prior:</strong></p>\n<ul>\n<li>Model cannot predict at [200-234] (architecture constraint)</li>\n<li>Training fails - visual features are blocked</li>\n</ul>\n<p><strong>With Soft Prior:</strong></p>\n<ul>\n<li>Model processes visual features at [200-234] through convolutional layers</li>\n<li>Training encourages predictions at [235-4161] via GT loss</li>\n<li>Model learns a compromise weighted by data distribution</li>\n<li>Inference crops to [235-4161], losing edge [200-234] but keeping [235-4126]</li>\n</ul>\n<p>The GT is drawn into [235-4161], but the model sees and processes the entire image during training. Visual features from misaligned regions still contribute to learning through the convolutional receptive field.</p>\n<p>The system degrades gracefully rather than failing catastrophically.</p>\n<p>This design choice accepts some label noise from imperfect rectifications in exchange for:</p>\n<ul>\n<li>More diverse training data (all segment types)</li>\n<li>Improved robustness to real-world rectification errors</li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 3395721,
              "author_name": "xie233",
              "author_url": "",
              "post_date": "01/23/2026 14:44:42",
              "content": "<p>Thank you so much for your patient explanation. Your answer has cleared up my confusion. I have one more question: when constructing the training data, did you manually check and exclude some severely misaligned samples, or did you use all the rectified images directly?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3395753,
                  "author_name": "sshiyu",
                  "author_url": "",
                  "post_date": "01/23/2026 15:57:51",
                  "content": "<p>I used all rectified images directly without manual filtering or exclusion.</p>\n<p>I generally avoid removing data due to fear of information loss. In my experience, severely misaligned samples are quite rare - the vast majority of rectifications from Stage 0+1 are reasonably accurate. The few badly misaligned samples act as noise in the training set, but with enough well-aligned data, the model naturally learns to filter out these outliers and converge to the correct visual features.</p>\n<p>The benefits of keeping all data (more diverse training samples, better coverage of edge cases) outweighed the cost of a small amount of label noise from misalignments.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3395806,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/23/2026 17:16:58",
      "content": "<p>Regarding the alignment issue:\nI think teams that use direct regression (i.e., mapping from images to values) will not have this issue. This is because crooked lines provide cues for the network to adjust misaligned pixel locations and still predict the correct values.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3395917,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "01/23/2026 23:44:16",
          "content": "<p>I think so too. And thanks again for your great work, it helps us a lot.😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3395469": "# PhysioNet ECG Digitization Challenge - Solution Writeup\n\n## Overview\nThis solution addresses the PhysioNet challenge of digitizing ECG images—converting printed/photographed ECG recordings back into numerical waveform data. The approach uses a **3-stage deep learning pipeline** that progressively transforms raw ECG images into high-fidelity digital waveforms.\n\n**Final Public Leaderboard Score: 21.56 SNR**\n**Final Private Leaderboard Score: 21.37 SNR**\n\n---\n\n## Problem Statement\n\nECG images from various sources (printed records, mobile phone photos, scans of damaged documents) need to be converted back to digital waveforms. The challenge involves:\n\n1. **Diverse image sources**: High-quality prints, mobile photos, screen captures, damaged/moldy documents\n2. **Geometric distortions**: Perspective warps, rotations, varying aspect ratios\n3. **Visual degradations**: Water stains, fold creases, mold, low resolution, compression artifacts\n4. **Layout parsing**: Standard 4-strip ECG layout with 12 leads arranged in specific positions\n\n---\n## Solution Architecture\n\n### Three-Stage Pipeline\n\n![Pipeline Diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2Fea72f9f6a77a90a276213a31965d4755%2Fdiagram1.jpg?generation=1769138841225117&alt=media)\n\n### Stage 0: Normalization\n\n**Purpose**: Detect lead markers and correct image orientation\n\n- **Architecture**: ResNet18 encoder + U-Net decoder\n- **Outputs**: \n  - Lead marker segmentation (13 leads + background)\n  - Image orientation classification (8 orientations)\n- **Processing**: Applies homography to normalize image to canonical size (3024×4032)\n\n> **Note**: For this stage, I utilized the pre-trained Stage 0 checkpoint provided in the [original baseline](https://www.kaggle.com/code/hengck23/demo-submission).\n\n### Stage 1: Rectification\n\n**Purpose**: Detect and correct grid distortions\n\n- **Architecture**: ResNet34 encoder + U-Net decoder  \n- **Outputs**:\n  - Grid point detection (keypoints for perspective correction)\n  - Horizontal line classification (44 classes)\n  - Vertical line classification (57 classes)\n- **Processing**: Uses detected grid points to apply perspective transformation, producing a perfectly aligned rectified image\n\n> **Note**: Similar to Stage 0, I used the pre-trained Stage 1 checkpoint from the [original baseline](https://www.kaggle.com/code/hengck23/demo-submission).\n\n### Stage 2: Waveform Prediction\n\n**Purpose**: Extract ECG waveforms from rectified images\n\nThis is the main prediction stage where most of the innovation lies.\n\n#### Model Architecture\n\n![Stage 2 Architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14897182%2F28e82aa96fa43e3593089860871597e3%2Fdiagram2.jpg?generation=1769138929087253&alt=media)\n\n#### Key Design Decisions\n\n1. **HRNet-W48 Backbone**: Maintains high-resolution representations throughout the network via parallel branches, crucial for precise pixel-level waveform prediction\n\n2. **CoordConv Decoder**: Injects normalized 2D coordinates at each decoder block, enabling position-aware predictions essential for ECG layout understanding\n\n3. **RGB Skip Connection**: Direct connection from input image to prediction head preserves fine spatial details that may be lost in encoder compression\n\n4. **PixelShuffle Upsampling**: Memory-efficient learnable upsampling using sub-pixel convolution instead of transposed convolution\n\n5. **GroupNorm**: Enables stable training with batch size 1 (necessary for high-resolution 1696×4352 images on limited GPU memory)\n\n---\n\n## Training Strategy\n\n### Loss Function\n\nThe model uses **Binary Cross-Entropy (BCE) with positive class weighting** for pixel-level supervision:\n\n```python\nloss = F.binary_cross_entropy_with_logits(\n    pixel_logit,\n    target_mask,\n    pos_weight=10.0  # Handle sparse ECG line pixels\n)\n```\n\nThis was found to outperform regression-based losses (MSE, SNR) for this task.\n\n### Data Augmentation\n\nA comprehensive **segment-aware augmentation pipeline** was developed to match the diverse training data distribution:\n\n#### Universal Augmentations\n| Augmentation | Probability | Purpose |\n|-------------|-------------|---------|\n| Perlin Noise | 85% | Simulate scanner noise, paper texture |\n| Texture Noise | 70% | Grid-like artifacts, scan lines |\n| Color Jitter | 80% | Brightness, contrast, saturation, hue variation |\n| Gaussian Noise | 50% | Scanner ISO noise |\n| Gaussian Blur | 35% | Low-quality scans, out-of-focus photos |\n| Shadow Overlay | 35% | Uneven lighting |\n| Paper Texture | 30% | Physical paper grain simulation |\n| JPEG Artifacts | 30% | Compression artifacts |\n| Horizontal Flip | 50% | Data augmentation (with label flip) |\n\n#### Segment-Specific Augmentations\n\nThe training data contains 12 segment types with distinct characteristics. The augmentation pipeline applies targeted degradations based on segment type:\n\n| Segment | Characteristics | Applied Augmentations |\n|---------|-----------------|----------------------|\n| 0003 | Standard color scans | Baseline augmentations only |\n| 0004 | Black & white scans | Grayscale conversion + paper tint |\n| 0005 | Mobile phone photos | Vignetting + shadow gradients |\n| 0006 | Screen photos | Moiré patterns + screen glare |\n| 0009 | Water-damaged | Water stains + color bleeding |\n| 0010 | Extensively damaged | Fold creases + heavy noise |\n| 0011, 0012 | Moldy scans | Mold patches + grayscale |\n\n#### Augmentation Caching\n\nTo avoid expensive on-the-fly augmentation operations (especially Perlin noise generation), a **memory-mapped caching system** pre-computes augmentation masks:\n\n- Perlin noise cache: 512 pre-computed noise maps\n- Texture cache: 256 patterns\n- Shadow/vignette cache: 256 gradient patterns\n\n### Training Configuration\n\n| Parameter | Value |\n|-----------|-------|\n| Batch Size | 1 (4352×1696 full resolution) |\n| Optimizer | AdamW |\n| Learning Rate | 1e-4 |\n| Weight Decay | 0.05 |\n| Epochs | 80 |\n| EMA Decay | 0.999 |\n| Train/Val Split | 100/0 (full training set) |\n| Mixed Precision | FP16 (AMP) |\n| Gradient Checkpointing | Enabled for decoder |\n\n---\n\n## Inference Pipeline\n\n### Test-Time Augmentation (TTA)\n\nHorizontal flip TTA improves robustness:\n\n1. Predict on original image\n2. Horizontally flip input image\n3. Predict on flipped image\n4. Flip prediction back to original orientation\n5. Average both predictions\n\n### Einthoven's Law Correction\n\nA physics-based post-processing step enforces the electrocardiogram constraint:\n\n**Einthoven's Law**: `Lead II = Lead I + Lead III`\n\nThe correction redistributes violation errors across all three limb leads:\n\n```python\nerror = II - (I + III)\nI_corrected = I + α × error\nIII_corrected = III + α × error\nII_corrected = II - α × error  # α = 0.33\n```\n\n### Waveform Extraction\n\nThe 4-strip pixel predictions are converted to waveforms via:\n\n1. Apply softmax on y-dimension to get probability distributions\n2. Compute weighted centroid (argmax or soft-argmax) for each x-column\n3. Convert pixel coordinates to mV using calibration parameters\n4. Split each strip into constituent leads based on temporal layout\n5. Resample to target sample lengths specified in test.csv\n\n---\n\n## Ablation Studies\n\n> **Note**: These scores are approximate. Changes were not isolated individually to facilitate faster iteration and save compute costs.\n\n| Configuration | Public LB Score |\n|---------------|-----------------|\n| Baseline (BCE loss + comprehensive augmentations, 90/10 split) | ~15.31 |\n| + High Resolution (1696×4352) | ~18.22 |\n| + HRNet-W48 backbone (replacing ResNet) | ~20.90 |\n| + 100/0 train/val split (use all training data) | ~21.01 |\n| + Horizontal Flip TTA | ~21.28 |\n| + Einthoven's Law correction | ~21.32 |\n| + RGB Skip + PixelShuffle + Fusion Head | ~21.56 |\n\n### Key Insights\n\n1. **Resolution matters**: Increasing from standard resolution to 1696×4352 provided the biggest single improvement (~3 SNR)\n\n2. **HRNet superiority**: The multi-resolution parallel branches preserve spatial precision better than traditional encoder-decoder architectures\n\n3. **Full training data**: With comprehensive augmentation, using 100% of data for training outperformed keeping a validation set\n\n4. **Physics constraints help**: Einthoven correction provides consistent small improvements by enforcing known ECG relationships\n\n---\n\n## Technical Implementation Details\n\n### Memory Optimization\n\nTraining on high-resolution images (1696×4352) with batch size 1 required careful memory management:\n\n- **Gradient Checkpointing**: Enabled for decoder blocks to trade compute for memory\n- **GroupNorm**: Used instead of BatchNorm for stable single-sample training\n- **PixelShuffle**: Used instead of ConvTranspose2d for 2× upsampling (more memory efficient)\n- **Mixed Precision (FP16)**: Enabled via PyTorch AMP\n\n### Distributed Training\n\n- DDP (Distributed Data Parallel) support for multi-GPU training\n- Per-worker augmentation cache to avoid contention\n- Memory-mapped cache files shared across workers\n\n---\n\n## Repository Structure\n\n```\n├── stage0_model.py          # Stage 0: Marker detection & normalization\n├── stage1_model.py          # Stage 1: Grid detection & rectification\n├── stage2_model.py          # Stage 2: Waveform prediction (main model)\n├── predictor_dataset.py     # Training data loading & preprocessing\n├── predictor_trainer.py     # Training loop with EMA, metrics\n├── augmentations.py         # Comprehensive augmentation pipeline\n├── augmentation_cache.py    # Memory-mapped augmentation caching\n├── demo-submission.py       # End-to-end inference & submission generation\n├── config_predictor.yaml    # Training configuration\n└── archive/\n    ├── stage0-last.checkpoint.pth  # Stage 0 pretrained weights\n    └── stage1-last.checkpoint.pth  # Stage 1 pretrained weights\n```\n\n---\n\n## Code\n\n- **GitHub Repository**: [starrynites/PhysioNet_Digitization_of_ECG_Images](https://github.com/starrynites/PhysioNet_Digitization_of_ECG_Images)\n- **Kaggle Inference Notebook**: [sshiyu/physionet-final-submission](https://www.kaggle.com/code/sshiyu/physionet-final-submission)\n\n## Conclusion\n\nThis solution achieves strong performance on ECG digitization through:\n\n1. **Robust preprocessing**: Stage 0+1 pipeline handles diverse image sources and geometric distortions\n2. **High-resolution prediction**: Full-resolution HRNet enables precise pixel-level waveform detection\n3. **Comprehensive augmentation**: Segment-aware degradation simulation improves generalization\n4. **Physics-informed post-processing**: Einthoven correction enforces known ECG constraints\n5. **Efficient architecture**: RGB skip connection + PixelShuffle fusion preserves fine details while maintaining training efficiency\n\nHope you enjoyed my writeup. Thanks for reading!\n\n---\n\n## Acknowledgements\n\nThis solution builds upon the excellent work shared by the Kaggle community:\n\n- [hengck23/demo-submission](https://www.kaggle.com/code/hengck23/demo-submission) — Original low resolution baseline\n- [wasupandceacar/physio-v2-3-public](https://www.kaggle.com/code/wasupandceacar/physio-v2-3-public) — High resolution baseline reference\n- [tonylica/physionet-ecg-streamlined-inference](https://www.kaggle.com/code/tonylica/physionet-ecg-streamlined-inference) — Einthoven's law correction idea\n\n### References\n\n- Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954\n- Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.\n- Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Seyedi S, Elola A, Bahrami Rad A, Shah AJ, Bhatia NK, Clifford GD, Sameni R. Digitization and Classification of ECG Images: The George B. Moody PhysioNet Challenge 2024; Computing in Cardiology 2024; 51: 1-4.\n- Matthew A. Reyna, Deepanshi, James Weigle, Zuzana Koscova, Kiersten Campbell, Salman Seyedi, Andoni Elola, Ali Bahrami Rad, Amit J Shah, Neal K. Bhatia, Yao Yan, Sohier Dane, Addison Howard, Gari D. Clifford, and Reza Sameni. PhysioNet - Digitization of ECG Images. https://kaggle.com/competitions/physionet-ecg-image-digitization, 2025. Kaggle.\n\n---",
    "3395554": "Thank you so much for sharing such a detailed and thorough solution! When I was going through your code and solution, I noticed that you use all types of training images after rectification as the input for Stage 2 during training. However, in my tests, the rectified non-0001 type images are not perfectly aligned with the GT signals. I wonder how you addressed this issue, or if I have any misunderstanding of your solution.",
    "3395576": "That is a great observation! I also noticed that due to rectification errors in stage 0 and 1, the GT signals are not always perfectly aligned with what we see visually (mostly due to horizontal misalignment).\n\n![Visualisation of ground truth horizontal misalignment](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14897182%2F671aa7d0472f2321356fc7045ed06e58%2FScreenshot%202026-01-01%20at%2011.42.46PM.png?generation=1769156126338577&alt=media)\n\nI do agree that this creates some training confusion. However, I believe that this actually does not significantly hurt performance for the following reasons:\n\n1. **Most samples are reasonably well-aligned** - Stage 1 rectification works well for the majority of images\n2. **The spatial prior is \"soft\" rather than \"hard\"** - the model can adapt to misalignments\n\n## Soft vs Hard Spatial Prior\n\nThe key distinction is **where and how** the timespan constraint is applied:\n\n### Hard Spatial Prior (NOT what I do)\n\n```python\n# Hypothetical: Model architecture prevents predictions outside timespan\nclass HardConstrainedModel:\n    def forward(self, batch):\n        x = batch['image']\n        features = self.encoder(x)\n        output = torch.zeros(B, 4, H, W)\n        # Only allow decoder to predict in specific columns\n        output[:, :, :, 235:4161] = self.decoder(features)[:, :, :, 235:4161]\n        return {'pixel_logit': output}  # Forced zeros outside [235, 4161]\n```\n\nThis would completely prevent the model from learning waveform features outside [235-4161].\n\n### Soft Spatial Prior (My Implementation)\n\n**Model architecture** (`stage2_model.py`):\n\n```python\ndef forward(self, batch, L=None):\n    image = batch['image']  # (B, 3, H, W)\n    # ... encoder + decoder ...\n    pixel_logit = self.pixel_out(last)  # (B, 4, H, W) - FULL width prediction\n    return {'pixel_logit': pixel_logit}  # No spatial masking\n```\n\n**Training** (`predictor_trainer.py`):\n\n```python\n# Loss is computed on FULL image\npixel_loss = F.binary_cross_entropy_with_logits(\n    pixel_logit,     # (B, 4, H, W=4352) - all columns\n    pixel_target,    # (B, 4, H, W=4352) - GT drawn at [235, 4161]\n    pos_weight=pos_weight,\n)\n# Model learns to activate [235-4161] due to GT distribution,\n# but architecture doesn't prevent activations elsewhere\n```\n\n**Inference** (`predictor_trainer.py`):\n\n```python\ndef _predict_waveform(self, image_u8, target_len):\n    out = self.model(batch)\n    pixel_logit = out['pixel_logit']  # (B, 4, H, W) - full width\n    \n    # Convert ALL columns to y-positions\n    y_hat = pixel_logits_to_y(pixel_logit, ...)  # (B, 4, W)\n    \n    # Crop to timespan AFTER conversion (post-processing, not architectural)\n    y_hat = y_hat[..., self.t0:self.t1]  # (B, 4, 3926)\n    \n    mv_hat = y_to_mv(y_hat, ...)\n    return mv_hat\n```\n\n## Why This Matters for Misalignment\n\nFor a misaligned sample where the waveform appears at columns [200-4126] instead of [235-4161]:\n\n**With Hard Prior:**\n- Model cannot predict at [200-234] (architecture constraint)\n- Training fails - visual features are blocked\n\n**With Soft Prior:**\n- Model processes visual features at [200-234] through convolutional layers\n- Training encourages predictions at [235-4161] via GT loss\n- Model learns a compromise weighted by data distribution\n- Inference crops to [235-4161], losing edge [200-234] but keeping [235-4126]\n\nThe GT is drawn into [235-4161], but the model sees and processes the entire image during training. Visual features from misaligned regions still contribute to learning through the convolutional receptive field.\n\nThe system degrades gracefully rather than failing catastrophically.\n\nThis design choice accepts some label noise from imperfect rectifications in exchange for:\n- More diverse training data (all segment types)\n- Improved robustness to real-world rectification errors",
    "3395721": "Thank you so much for your patient explanation. Your answer has cleared up my confusion. I have one more question: when constructing the training data, did you manually check and exclude some severely misaligned samples, or did you use all the rectified images directly?",
    "3395753": "I used all rectified images directly without manual filtering or exclusion.\n\nI generally avoid removing data due to fear of information loss. In my experience, severely misaligned samples are quite rare - the vast majority of rectifications from Stage 0+1 are reasonably accurate. The few badly misaligned samples act as noise in the training set, but with enough well-aligned data, the model naturally learns to filter out these outliers and converge to the correct visual features.\n\nThe benefits of keeping all data (more diverse training samples, better coverage of edge cases) outweighed the cost of a small amount of label noise from misalignments.",
    "3395806": "Regarding the alignment issue:\nI think teams that use direct regression (i.e., mapping from images to values) will not have this issue. This is because crooked lines provide cues for the network to adjust misaligned pixel locations and still predict the correct values.",
    "3395917": "I think so too. And thanks again for your great work, it helps us a lot.😃"
  },
  "source": "meta"
}