{
  "id": 669697,
  "title": "15th place solution",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/669697",
  "author_name": "Arunodhayan",
  "post_date": "2026-01-23T19:07:12.090000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Team mate: <a href=\"https://www.kaggle.com/fatliuyun\" target=\"_blank\">@fatliuyun</a> </p>\n<h1>Overview</h1>\n<p><strong>ECG image → signal mask prediction → post-processed waveform extraction</strong></p>\n<p>The core idea is to separate signal extraction from signal reconstruction. By learning to segment ECG traces at the pixel level, the model becomes significantly more robust to real-world artifacts such as grid distortion, scanning noise, paper folds, ink fading, and occlusions.</p>\n<p>To capture both global ECG layout and fine-grained waveform details, two complementary training strategies are used:</p>\n<ul>\n<li>Full-crop training, which preserves the complete ECG context and long-range signal continuity.</li>\n<li>Half-crop training, which focuses on localized waveform structure and improves sensitivity to subtle signal variations.</li>\n</ul>\n<p>Both strategies employ lightweight ResNet-based encoders (ResNet-18 and ResNet-34d) and share a common decoding head. Their predictions are combined during inference to improve robustness and generalization</p>\n<h1>Rectified image and Mask generation</h1>\n<h2>Rectified image</h2>\n<p>The early rectification version from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> suffered from a subtle boundary issue where the predicted grid did not fully span the image extent, causing the last 1–2 pixel columns to be mis-mapped or border-filled during warping. This was fixed by enforcing a fixed output canvas, normalizing grid points using (W−1, H−1), and enabling corner-aligned interpolation (align_corners=True) when upsampling the deformation field. As a result, the dense grid now covers the full image width, eliminating edge pixel loss and significantly improving pixel-level alignment between rectified images and supervision masks, especially at waveform endpoints.</p>\n<p>` </p>\n<pre><code># Trim unstable grid columns\ngridpoint_xy = gridpoint_xy[:, :56]\n\n # Empirical offsets to correct scanner margins\noffset_y = 7\noffset_x = 34\n\n# Remove noisy right-side region\nimage = image[:, :2166]\n\n# Reference output size\nH, W = 1700 - offset_y, 2200 - offset_x\n\n# Allocate padded output canvas\nempty_image = np.zeros((H + offset_y, W, 3), np.uint8)\n\n# Normalize grid points to [-1, 1]\nH1, W1 = image.shape[:2]\nsparse_map = gridpoint_xy / [[[W1 - 1, H1 - 1]]] * 2 - 1\nsparse_map = torch.from_numpy(\n    np.ascontiguousarray(sparse_map.transpose(2, 0, 1))\n).unsqueeze(0).float()\n\n# Interpolate sparse grid to dense deformation field\ndense_map = F.interpolate(\n    sparse_map,\n    size=(H, W),\n    mode=\"bilinear\",\n    align_corners=True\n)\n\n# Apply grid-based warping\ndistort = torch.from_numpy(\n    np.ascontiguousarray(image.transpose(2, 0, 1))\n).unsqueeze(0).float()\n\nrectified = F.grid_sample(\n    distort,\n    dense_map.permute(0, 2, 3, 1),\n    mode=\"bilinear\",\n    padding_mode=\"border\",\n    align_corners=True\n)\n\n# Convert back to image format\nrectified = rectified.data.cpu().numpy()\nrectified = rectified[0].transpose(1, 2, 0).astype(np.uint8)\n\n# Restore vertical offset for alignment\nempty_image[offset_y:, :, :] = rectified\n\nreturn empty_image\n</code></pre>\n<p>`</p>\n<h2>Mask generation</h2>\n<p>Number of series: 4</p>\n<p>Output mask shape: (4, 1696, 4352)\n(height × width corresponds to the rectified ECG image)</p>\n<table>\n<thead>\n<tr>\n<th>Series</th>\n<th>Zero-MV</th>\n<th>mV --&gt; Pixel</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>707</td>\n<td>78.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>991</td>\n<td>79.5</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1273</td>\n<td>78.5</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1533</td>\n<td>78.0</td>\n</tr>\n</tbody>\n</table>\n<h1>Training</h1>\n<p>To balance global ECG layout understanding with local waveform precision, two complementary training strategies were employed: Full-Crop training and Half-Crop training. Each addresses a different failure mode observed during early experiments.</p>\n<h2>Full-Crop Training (Global Context)</h2>\n<p>In full-crop training, the model is trained on the entire rectified ECG image, preserving the complete temporal span and vertical alignment of all leads.</p>\n<ul>\n<li><p>Captures long-range temporal continuity of ECG signals</p></li>\n<li><p>Preserves relative lead positioning and baseline consistency</p></li>\n<li><p>Improves stability near signal boundaries (start/end regions)</p></li>\n<li><p>Helps the model learn overall ECG structure and layout</p></li>\n</ul>\n<h2>Half-Crop Training (Local Detail)</h2>\n<p>Half-crop training uses random horizontal crops covering approximately half the image width, while preserving full vertical resolution. Cropped regions are then resized back to the original dimensions before training.</p>\n<p>Why half-crop matters:</p>\n<ul>\n<li><p>Forces the model to focus on local waveform morphology</p></li>\n<li><p>Increases effective resolution per signal segment</p></li>\n<li><p>Improves robustness to local noise, breaks, and occlusions</p></li>\n<li><p>Acts as a strong form of data augmentation</p></li>\n</ul>\n<h1>DATA AUGMENTATION STRATEGY</h1>\n<h2>GEOMETRIC AUGMENTATIONS</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Parameters</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Affine transform</td>\n<td>Scale 0.90–1.10, Rotate ±8°, Translate ≤4%</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>Perspective transform</td>\n<td>Scale 0.02–0.07</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>Optical distortion</td>\n<td>Distort ≤0.03, Shift ≤0.02</td>\n<td>0.15</td>\n</tr>\n<tr>\n<td>Elastic transform</td>\n<td>Alpha = 10, Sigma = 3</td>\n<td>0.03</td>\n</tr>\n<tr>\n<td>Grid distortion</td>\n<td>Steps = 3, Distort limit = 0.03</td>\n<td>0.05</td>\n</tr>\n<tr>\n<td>Affine (shear)</td>\n<td>Shear ±3°</td>\n<td>0.15</td>\n</tr>\n</tbody>\n</table>\n<h2>PHOTOMETRIC AUGMENTATIONS</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Parameters</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Brightness / Contrast</td>\n<td>±0.15</td>\n<td>0.50 (OneOf)</td>\n</tr>\n<tr>\n<td>CLAHE</td>\n<td>Clip = 2.0, Grid = 8×8</td>\n<td>0.30 (OneOf)</td>\n</tr>\n<tr>\n<td>Random Gamma</td>\n<td>0.8–1.2</td>\n<td>0.20 (OneOf)</td>\n</tr>\n<tr>\n<td>Gaussian Noise</td>\n<td>Variance 5–20</td>\n<td>0.50 (OneOf)</td>\n</tr>\n<tr>\n<td>ISO Noise</td>\n<td>—</td>\n<td>0.30 (OneOf)</td>\n</tr>\n<tr>\n<td>Multiplicative Noise</td>\n<td>0.9–1.1</td>\n<td>0.20 (OneOf)</td>\n</tr>\n<tr>\n<td>Hue / Saturation / Value</td>\n<td>H±2, S±5, V±5</td>\n<td>0.15</td>\n</tr>\n</tbody>\n</table>\n<h2>BLUR &amp; SHARPENING</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Parameters</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Sharpen</td>\n<td>Alpha 0.05–0.15</td>\n<td>0.20</td>\n</tr>\n<tr>\n<td>Gaussian Blur</td>\n<td>Kernel size 3–5</td>\n<td>0.60 (OneOf)</td>\n</tr>\n<tr>\n<td>Motion Blur</td>\n<td>Kernel size 3–5</td>\n<td>0.40 (OneOf)</td>\n</tr>\n</tbody>\n</table>\n<h2>ECG-SPECIFIC OCCLUSIONS</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Purpose</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CoarseDropout (small)</td>\n<td>Ink breaks / stains</td>\n<td>0.25</td>\n</tr>\n<tr>\n<td>CoarseDropout (vertical)</td>\n<td>Printer / scan vertical lines</td>\n<td>0.20</td>\n</tr>\n<tr>\n<td>Perlin noise patches</td>\n<td>Paper dirt / aging artifacts</td>\n<td>~0.50 (custom)</td>\n</tr>\n</tbody>\n</table>\n<h1># POST-PROCESSING &amp; SIGNAL REFINEMENT (STAGE-2)</h1>\n<p>After segmentation-based signal extraction, additional temporal refinement and gap filling are applied to recover physiologically consistent ECG waveforms. This post-processing focuses on handling missing segments, enforcing periodic consistency, and stabilizing QRS morphology.</p>\n<h2>1) MISSING-SEGMENT DETECTION &amp; FILLING</h2>\n<ul>\n<li>A binary mask indicates missing points (mask = 1).</li>\n<li>Consecutive missing runs longer than a minimum length are detected per lead.</li>\n<li>These runs are filled with the mean of nearby observed values in the same lead.</li>\n<li>Prevents long flat gaps while preserving baseline consistency.</li>\n</ul>\n<h2>2) BEAT-PHASE–AWARE IMPUTATION (KEY STEP)</h2>\n<ul>\n<li>ECG is divided into fixed 1250-sample segments.</li>\n<li>Global R-peaks are detected across all leads.</li>\n<li>For a missing point t:<ul>\n<li>Find nearest R-peak in the same segment.</li>\n<li>Compute phase offset k = t − R_peak.</li>\n<li>Collect corresponding points (R_j + k) from other cycles in the same segment.</li>\n<li>If multiple valid values exist, average them to fill t.</li></ul></li>\n<li>Enforces periodic consistency and preserves waveform morphology.</li>\n</ul>\n<h2>3) ROBUST R-PEAK DETECTION</h2>\n<ul>\n<li>Band-pass filtering (5–18 Hz) emphasizes QRS complexes.</li>\n<li>Pan-Tompkins–style moving window integration (MWI) is applied.</li>\n<li>Energy envelopes are fused across leads.</li>\n<li>Adaptive thresholding with search-back logic recovers missed beats.</li>\n<li>Peaks are refined to the true positive QRS maximum.</li>\n<li>A reference lead with strongest QRS energy is selected automatically.</li>\n</ul>\n<h2>4) MODEL-LEVEL TEMPORAL REFINEMENT</h2>\n<ul>\n<li>Multi-scale temporal convolution (TCN) captures local and long-range dependencies.</li>\n<li>Fourier harmonic features model dominant cardiac periodicity.</li>\n<li>Gated fusion blends time-domain and frequency-domain information.</li>\n<li>Residual bidirectional LSTM enforces temporal smoothness.</li>\n<li>The model learns corrections only for missing regions; observed points are preserved.</li>\n</ul>\n<h2>5) OUTPUT GUARANTEES</h2>\n<ul>\n<li>Observed (non-missing) points are never modified.</li>\n<li>Filled values respect local continuity and global cardiac rhythm.</li>\n<li>No aggressive low-pass filtering is applied; sharp QRS morphology is preserved.</li>\n</ul>",
  "messages": [
    {
      "id": 3395851,
      "postDate": "2026-01-23T19:07:12.090Z",
      "content": "<p>Team mate: <a href=\"https://www.kaggle.com/fatliuyun\" target=\"_blank\">@fatliuyun</a> </p>\n<h1>Overview</h1>\n<p><strong>ECG image → signal mask prediction → post-processed waveform extraction</strong></p>\n<p>The core idea is to separate signal extraction from signal reconstruction. By learning to segment ECG traces at the pixel level, the model becomes significantly more robust to real-world artifacts such as grid distortion, scanning noise, paper folds, ink fading, and occlusions.</p>\n<p>To capture both global ECG layout and fine-grained waveform details, two complementary training strategies are used:</p>\n<ul>\n<li>Full-crop training, which preserves the complete ECG context and long-range signal continuity.</li>\n<li>Half-crop training, which focuses on localized waveform structure and improves sensitivity to subtle signal variations.</li>\n</ul>\n<p>Both strategies employ lightweight ResNet-based encoders (ResNet-18 and ResNet-34d) and share a common decoding head. Their predictions are combined during inference to improve robustness and generalization</p>\n<h1>Rectified image and Mask generation</h1>\n<h2>Rectified image</h2>\n<p>The early rectification version from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> suffered from a subtle boundary issue where the predicted grid did not fully span the image extent, causing the last 1–2 pixel columns to be mis-mapped or border-filled during warping. This was fixed by enforcing a fixed output canvas, normalizing grid points using (W−1, H−1), and enabling corner-aligned interpolation (align_corners=True) when upsampling the deformation field. As a result, the dense grid now covers the full image width, eliminating edge pixel loss and significantly improving pixel-level alignment between rectified images and supervision masks, especially at waveform endpoints.</p>\n<p>` </p>\n<pre><code># Trim unstable grid columns\ngridpoint_xy = gridpoint_xy[:, :56]\n\n # Empirical offsets to correct scanner margins\noffset_y = 7\noffset_x = 34\n\n# Remove noisy right-side region\nimage = image[:, :2166]\n\n# Reference output size\nH, W = 1700 - offset_y, 2200 - offset_x\n\n# Allocate padded output canvas\nempty_image = np.zeros((H + offset_y, W, 3), np.uint8)\n\n# Normalize grid points to [-1, 1]\nH1, W1 = image.shape[:2]\nsparse_map = gridpoint_xy / [[[W1 - 1, H1 - 1]]] * 2 - 1\nsparse_map = torch.from_numpy(\n    np.ascontiguousarray(sparse_map.transpose(2, 0, 1))\n).unsqueeze(0).float()\n\n# Interpolate sparse grid to dense deformation field\ndense_map = F.interpolate(\n    sparse_map,\n    size=(H, W),\n    mode=\"bilinear\",\n    align_corners=True\n)\n\n# Apply grid-based warping\ndistort = torch.from_numpy(\n    np.ascontiguousarray(image.transpose(2, 0, 1))\n).unsqueeze(0).float()\n\nrectified = F.grid_sample(\n    distort,\n    dense_map.permute(0, 2, 3, 1),\n    mode=\"bilinear\",\n    padding_mode=\"border\",\n    align_corners=True\n)\n\n# Convert back to image format\nrectified = rectified.data.cpu().numpy()\nrectified = rectified[0].transpose(1, 2, 0).astype(np.uint8)\n\n# Restore vertical offset for alignment\nempty_image[offset_y:, :, :] = rectified\n\nreturn empty_image\n</code></pre>\n<p>`</p>\n<h2>Mask generation</h2>\n<p>Number of series: 4</p>\n<p>Output mask shape: (4, 1696, 4352)\n(height × width corresponds to the rectified ECG image)</p>\n<table>\n<thead>\n<tr>\n<th>Series</th>\n<th>Zero-MV</th>\n<th>mV --&gt; Pixel</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>707</td>\n<td>78.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>991</td>\n<td>79.5</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1273</td>\n<td>78.5</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1533</td>\n<td>78.0</td>\n</tr>\n</tbody>\n</table>\n<h1>Training</h1>\n<p>To balance global ECG layout understanding with local waveform precision, two complementary training strategies were employed: Full-Crop training and Half-Crop training. Each addresses a different failure mode observed during early experiments.</p>\n<h2>Full-Crop Training (Global Context)</h2>\n<p>In full-crop training, the model is trained on the entire rectified ECG image, preserving the complete temporal span and vertical alignment of all leads.</p>\n<ul>\n<li><p>Captures long-range temporal continuity of ECG signals</p></li>\n<li><p>Preserves relative lead positioning and baseline consistency</p></li>\n<li><p>Improves stability near signal boundaries (start/end regions)</p></li>\n<li><p>Helps the model learn overall ECG structure and layout</p></li>\n</ul>\n<h2>Half-Crop Training (Local Detail)</h2>\n<p>Half-crop training uses random horizontal crops covering approximately half the image width, while preserving full vertical resolution. Cropped regions are then resized back to the original dimensions before training.</p>\n<p>Why half-crop matters:</p>\n<ul>\n<li><p>Forces the model to focus on local waveform morphology</p></li>\n<li><p>Increases effective resolution per signal segment</p></li>\n<li><p>Improves robustness to local noise, breaks, and occlusions</p></li>\n<li><p>Acts as a strong form of data augmentation</p></li>\n</ul>\n<h1>DATA AUGMENTATION STRATEGY</h1>\n<h2>GEOMETRIC AUGMENTATIONS</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Parameters</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Affine transform</td>\n<td>Scale 0.90–1.10, Rotate ±8°, Translate ≤4%</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>Perspective transform</td>\n<td>Scale 0.02–0.07</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>Optical distortion</td>\n<td>Distort ≤0.03, Shift ≤0.02</td>\n<td>0.15</td>\n</tr>\n<tr>\n<td>Elastic transform</td>\n<td>Alpha = 10, Sigma = 3</td>\n<td>0.03</td>\n</tr>\n<tr>\n<td>Grid distortion</td>\n<td>Steps = 3, Distort limit = 0.03</td>\n<td>0.05</td>\n</tr>\n<tr>\n<td>Affine (shear)</td>\n<td>Shear ±3°</td>\n<td>0.15</td>\n</tr>\n</tbody>\n</table>\n<h2>PHOTOMETRIC AUGMENTATIONS</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Parameters</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Brightness / Contrast</td>\n<td>±0.15</td>\n<td>0.50 (OneOf)</td>\n</tr>\n<tr>\n<td>CLAHE</td>\n<td>Clip = 2.0, Grid = 8×8</td>\n<td>0.30 (OneOf)</td>\n</tr>\n<tr>\n<td>Random Gamma</td>\n<td>0.8–1.2</td>\n<td>0.20 (OneOf)</td>\n</tr>\n<tr>\n<td>Gaussian Noise</td>\n<td>Variance 5–20</td>\n<td>0.50 (OneOf)</td>\n</tr>\n<tr>\n<td>ISO Noise</td>\n<td>—</td>\n<td>0.30 (OneOf)</td>\n</tr>\n<tr>\n<td>Multiplicative Noise</td>\n<td>0.9–1.1</td>\n<td>0.20 (OneOf)</td>\n</tr>\n<tr>\n<td>Hue / Saturation / Value</td>\n<td>H±2, S±5, V±5</td>\n<td>0.15</td>\n</tr>\n</tbody>\n</table>\n<h2>BLUR &amp; SHARPENING</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Parameters</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Sharpen</td>\n<td>Alpha 0.05–0.15</td>\n<td>0.20</td>\n</tr>\n<tr>\n<td>Gaussian Blur</td>\n<td>Kernel size 3–5</td>\n<td>0.60 (OneOf)</td>\n</tr>\n<tr>\n<td>Motion Blur</td>\n<td>Kernel size 3–5</td>\n<td>0.40 (OneOf)</td>\n</tr>\n</tbody>\n</table>\n<h2>ECG-SPECIFIC OCCLUSIONS</h2>\n<table>\n<thead>\n<tr>\n<th>Augmentation</th>\n<th>Purpose</th>\n<th>Probability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CoarseDropout (small)</td>\n<td>Ink breaks / stains</td>\n<td>0.25</td>\n</tr>\n<tr>\n<td>CoarseDropout (vertical)</td>\n<td>Printer / scan vertical lines</td>\n<td>0.20</td>\n</tr>\n<tr>\n<td>Perlin noise patches</td>\n<td>Paper dirt / aging artifacts</td>\n<td>~0.50 (custom)</td>\n</tr>\n</tbody>\n</table>\n<h1># POST-PROCESSING &amp; SIGNAL REFINEMENT (STAGE-2)</h1>\n<p>After segmentation-based signal extraction, additional temporal refinement and gap filling are applied to recover physiologically consistent ECG waveforms. This post-processing focuses on handling missing segments, enforcing periodic consistency, and stabilizing QRS morphology.</p>\n<h2>1) MISSING-SEGMENT DETECTION &amp; FILLING</h2>\n<ul>\n<li>A binary mask indicates missing points (mask = 1).</li>\n<li>Consecutive missing runs longer than a minimum length are detected per lead.</li>\n<li>These runs are filled with the mean of nearby observed values in the same lead.</li>\n<li>Prevents long flat gaps while preserving baseline consistency.</li>\n</ul>\n<h2>2) BEAT-PHASE–AWARE IMPUTATION (KEY STEP)</h2>\n<ul>\n<li>ECG is divided into fixed 1250-sample segments.</li>\n<li>Global R-peaks are detected across all leads.</li>\n<li>For a missing point t:<ul>\n<li>Find nearest R-peak in the same segment.</li>\n<li>Compute phase offset k = t − R_peak.</li>\n<li>Collect corresponding points (R_j + k) from other cycles in the same segment.</li>\n<li>If multiple valid values exist, average them to fill t.</li></ul></li>\n<li>Enforces periodic consistency and preserves waveform morphology.</li>\n</ul>\n<h2>3) ROBUST R-PEAK DETECTION</h2>\n<ul>\n<li>Band-pass filtering (5–18 Hz) emphasizes QRS complexes.</li>\n<li>Pan-Tompkins–style moving window integration (MWI) is applied.</li>\n<li>Energy envelopes are fused across leads.</li>\n<li>Adaptive thresholding with search-back logic recovers missed beats.</li>\n<li>Peaks are refined to the true positive QRS maximum.</li>\n<li>A reference lead with strongest QRS energy is selected automatically.</li>\n</ul>\n<h2>4) MODEL-LEVEL TEMPORAL REFINEMENT</h2>\n<ul>\n<li>Multi-scale temporal convolution (TCN) captures local and long-range dependencies.</li>\n<li>Fourier harmonic features model dominant cardiac periodicity.</li>\n<li>Gated fusion blends time-domain and frequency-domain information.</li>\n<li>Residual bidirectional LSTM enforces temporal smoothness.</li>\n<li>The model learns corrections only for missing regions; observed points are preserved.</li>\n</ul>\n<h2>5) OUTPUT GUARANTEES</h2>\n<ul>\n<li>Observed (non-missing) points are never modified.</li>\n<li>Filled values respect local continuity and global cardiac rhythm.</li>\n<li>No aggressive low-pass filtering is applied; sharp QRS morphology is preserved.</li>\n</ul>",
      "rawMarkdown": "Team mate: @fatliuyun \n# Overview\n**ECG image → signal mask prediction → post-processed waveform extraction**\n\nThe core idea is to separate signal extraction from signal reconstruction. By learning to segment ECG traces at the pixel level, the model becomes significantly more robust to real-world artifacts such as grid distortion, scanning noise, paper folds, ink fading, and occlusions.\n\nTo capture both global ECG layout and fine-grained waveform details, two complementary training strategies are used:\n\n- Full-crop training, which preserves the complete ECG context and long-range signal continuity.\n- Half-crop training, which focuses on localized waveform structure and improves sensitivity to subtle signal variations.\n\nBoth strategies employ lightweight ResNet-based encoders (ResNet-18 and ResNet-34d) and share a common decoding head. Their predictions are combined during inference to improve robustness and generalization\n\n\n# Rectified image and Mask generation\n## Rectified image \nThe early rectification version from @hengck23 suffered from a subtle boundary issue where the predicted grid did not fully span the image extent, causing the last 1–2 pixel columns to be mis-mapped or border-filled during warping. This was fixed by enforcing a fixed output canvas, normalizing grid points using (W−1, H−1), and enabling corner-aligned interpolation (align_corners=True) when upsampling the deformation field. As a result, the dense grid now covers the full image width, eliminating edge pixel loss and significantly improving pixel-level alignment between rectified images and supervision masks, especially at waveform endpoints.\n\n\n` \n\n\n\n    # Trim unstable grid columns\n    gridpoint_xy = gridpoint_xy[:, :56]\n\n     # Empirical offsets to correct scanner margins\n    offset_y = 7\n    offset_x = 34\n\n    # Remove noisy right-side region\n    image = image[:, :2166]\n\n    # Reference output size\n    H, W = 1700 - offset_y, 2200 - offset_x\n\n    # Allocate padded output canvas\n    empty_image = np.zeros((H + offset_y, W, 3), np.uint8)\n\n    # Normalize grid points to [-1, 1]\n    H1, W1 = image.shape[:2]\n    sparse_map = gridpoint_xy / [[[W1 - 1, H1 - 1]]] * 2 - 1\n    sparse_map = torch.from_numpy(\n        np.ascontiguousarray(sparse_map.transpose(2, 0, 1))\n    ).unsqueeze(0).float()\n\n    # Interpolate sparse grid to dense deformation field\n    dense_map = F.interpolate(\n        sparse_map,\n        size=(H, W),\n        mode=\"bilinear\",\n        align_corners=True\n    )\n\n    # Apply grid-based warping\n    distort = torch.from_numpy(\n        np.ascontiguousarray(image.transpose(2, 0, 1))\n    ).unsqueeze(0).float()\n\n    rectified = F.grid_sample(\n        distort,\n        dense_map.permute(0, 2, 3, 1),\n        mode=\"bilinear\",\n        padding_mode=\"border\",\n        align_corners=True\n    )\n\n    # Convert back to image format\n    rectified = rectified.data.cpu().numpy()\n    rectified = rectified[0].transpose(1, 2, 0).astype(np.uint8)\n\n    # Restore vertical offset for alignment\n    empty_image[offset_y:, :, :] = rectified\n\n    return empty_image\n`\n\n## Mask generation\nNumber of series: 4\n\nOutput mask shape: (4, 1696, 4352)\n(height × width corresponds to the rectified ECG image)\n|  Series|Zero-MV  | mV --> Pixel\n| --- | --- |\n| 0 | 707  | 78.0\n| 1 | 991  | 79.5\n| 2 | 1273  | 78.5\n| 3 | 1533  | 78.0\n\n# Training\nTo balance global ECG layout understanding with local waveform precision, two complementary training strategies were employed: Full-Crop training and Half-Crop training. Each addresses a different failure mode observed during early experiments.\n\n## Full-Crop Training (Global Context)\n\nIn full-crop training, the model is trained on the entire rectified ECG image, preserving the complete temporal span and vertical alignment of all leads.\n\n\n- Captures long-range temporal continuity of ECG signals\n\n- Preserves relative lead positioning and baseline consistency\n\n- Improves stability near signal boundaries (start/end regions)\n\n- Helps the model learn overall ECG structure and layout\n\n## Half-Crop Training (Local Detail)\n\nHalf-crop training uses random horizontal crops covering approximately half the image width, while preserving full vertical resolution. Cropped regions are then resized back to the original dimensions before training.\n\nWhy half-crop matters:\n\n- Forces the model to focus on local waveform morphology\n\n- Increases effective resolution per signal segment\n\n- Improves robustness to local noise, breaks, and occlusions\n\n- Acts as a strong form of data augmentation\n\nDATA AUGMENTATION STRATEGY\n==========================\n\nGEOMETRIC AUGMENTATIONS\n-----------------------\nAugmentation              | Parameters                                              | Probability\n--------------------------|---------------------------------------------------------|------------\nAffine transform          | Scale 0.90–1.10, Rotate ±8°, Translate ≤4%              | 0.60\nPerspective transform     | Scale 0.02–0.07                                         | 0.40\nOptical distortion        | Distort ≤0.03, Shift ≤0.02                              | 0.15\nElastic transform         | Alpha = 10, Sigma = 3                                   | 0.03\nGrid distortion           | Steps = 3, Distort limit = 0.03                         | 0.05\nAffine (shear)            | Shear ±3°                                               | 0.15\n\nPHOTOMETRIC AUGMENTATIONS\n-------------------------\nAugmentation              | Parameters                                | Probability\n--------------------------|-------------------------------------------|------------\nBrightness / Contrast     | ±0.15                                     | 0.50 (OneOf)\nCLAHE                     | Clip = 2.0, Grid = 8×8                    | 0.30 (OneOf)\nRandom Gamma              | 0.8–1.2                                   | 0.20 (OneOf)\nGaussian Noise            | Variance 5–20                             | 0.50 (OneOf)\nISO Noise                 | —                                         | 0.30 (OneOf)\nMultiplicative Noise      | 0.9–1.1                                   | 0.20 (OneOf)\nHue / Saturation / Value  | H±2, S±5, V±5                             | 0.15\n\nBLUR & SHARPENING\n-----------------\nAugmentation              | Parameters              | Probability\n--------------------------|-------------------------|------------\nSharpen                   | Alpha 0.05–0.15         | 0.20\nGaussian Blur             | Kernel size 3–5         | 0.60 (OneOf)\nMotion Blur               | Kernel size 3–5         | 0.40 (OneOf)\n\nECG-SPECIFIC OCCLUSIONS\n----------------------\nAugmentation              | Purpose                                   | Probability\n--------------------------|-------------------------------------------|------------\nCoarseDropout (small)     | Ink breaks / stains                       | 0.25\nCoarseDropout (vertical)  | Printer / scan vertical lines             | 0.20\nPerlin noise patches      | Paper dirt / aging artifacts              | ~0.50 (custom)\n\n\n\n# POST-PROCESSING & SIGNAL REFINEMENT (STAGE-2)\n======================================\n\nAfter segmentation-based signal extraction, additional temporal refinement and gap filling are applied to recover physiologically consistent ECG waveforms. This post-processing focuses on handling missing segments, enforcing periodic consistency, and stabilizing QRS morphology.\n\n1) MISSING-SEGMENT DETECTION & FILLING\n-------------------------------------\n- A binary mask indicates missing points (mask = 1).\n- Consecutive missing runs longer than a minimum length are detected per lead.\n- These runs are filled with the mean of nearby observed values in the same lead.\n- Prevents long flat gaps while preserving baseline consistency.\n\n2) BEAT-PHASE–AWARE IMPUTATION (KEY STEP)\n-----------------------------------------\n- ECG is divided into fixed 1250-sample segments.\n- Global R-peaks are detected across all leads.\n- For a missing point t:\n  * Find nearest R-peak in the same segment.\n  * Compute phase offset k = t − R_peak.\n  * Collect corresponding points (R_j + k) from other cycles in the same segment.\n  * If multiple valid values exist, average them to fill t.\n- Enforces periodic consistency and preserves waveform morphology.\n\n3) ROBUST R-PEAK DETECTION\n-------------------------\n- Band-pass filtering (5–18 Hz) emphasizes QRS complexes.\n- Pan-Tompkins–style moving window integration (MWI) is applied.\n- Energy envelopes are fused across leads.\n- Adaptive thresholding with search-back logic recovers missed beats.\n- Peaks are refined to the true positive QRS maximum.\n- A reference lead with strongest QRS energy is selected automatically.\n\n4) MODEL-LEVEL TEMPORAL REFINEMENT\n---------------------------------\n- Multi-scale temporal convolution (TCN) captures local and long-range dependencies.\n- Fourier harmonic features model dominant cardiac periodicity.\n- Gated fusion blends time-domain and frequency-domain information.\n- Residual bidirectional LSTM enforces temporal smoothness.\n- The model learns corrections only for missing regions; observed points are preserved.\n\n5) OUTPUT GUARANTEES\n-------------------\n- Observed (non-missing) points are never modified.\n- Filled values respect local continuity and global cardiac rhythm.\n- No aggressive low-pass filtering is applied; sharp QRS morphology is preserved.\n\n\n\n\n",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3395851": "Team mate: @fatliuyun \n# Overview\n**ECG image → signal mask prediction → post-processed waveform extraction**\n\nThe core idea is to separate signal extraction from signal reconstruction. By learning to segment ECG traces at the pixel level, the model becomes significantly more robust to real-world artifacts such as grid distortion, scanning noise, paper folds, ink fading, and occlusions.\n\nTo capture both global ECG layout and fine-grained waveform details, two complementary training strategies are used:\n\n- Full-crop training, which preserves the complete ECG context and long-range signal continuity.\n- Half-crop training, which focuses on localized waveform structure and improves sensitivity to subtle signal variations.\n\nBoth strategies employ lightweight ResNet-based encoders (ResNet-18 and ResNet-34d) and share a common decoding head. Their predictions are combined during inference to improve robustness and generalization\n\n\n# Rectified image and Mask generation\n## Rectified image \nThe early rectification version from @hengck23 suffered from a subtle boundary issue where the predicted grid did not fully span the image extent, causing the last 1–2 pixel columns to be mis-mapped or border-filled during warping. This was fixed by enforcing a fixed output canvas, normalizing grid points using (W−1, H−1), and enabling corner-aligned interpolation (align_corners=True) when upsampling the deformation field. As a result, the dense grid now covers the full image width, eliminating edge pixel loss and significantly improving pixel-level alignment between rectified images and supervision masks, especially at waveform endpoints.\n\n\n` \n\n\n\n    # Trim unstable grid columns\n    gridpoint_xy = gridpoint_xy[:, :56]\n\n     # Empirical offsets to correct scanner margins\n    offset_y = 7\n    offset_x = 34\n\n    # Remove noisy right-side region\n    image = image[:, :2166]\n\n    # Reference output size\n    H, W = 1700 - offset_y, 2200 - offset_x\n\n    # Allocate padded output canvas\n    empty_image = np.zeros((H + offset_y, W, 3), np.uint8)\n\n    # Normalize grid points to [-1, 1]\n    H1, W1 = image.shape[:2]\n    sparse_map = gridpoint_xy / [[[W1 - 1, H1 - 1]]] * 2 - 1\n    sparse_map = torch.from_numpy(\n        np.ascontiguousarray(sparse_map.transpose(2, 0, 1))\n    ).unsqueeze(0).float()\n\n    # Interpolate sparse grid to dense deformation field\n    dense_map = F.interpolate(\n        sparse_map,\n        size=(H, W),\n        mode=\"bilinear\",\n        align_corners=True\n    )\n\n    # Apply grid-based warping\n    distort = torch.from_numpy(\n        np.ascontiguousarray(image.transpose(2, 0, 1))\n    ).unsqueeze(0).float()\n\n    rectified = F.grid_sample(\n        distort,\n        dense_map.permute(0, 2, 3, 1),\n        mode=\"bilinear\",\n        padding_mode=\"border\",\n        align_corners=True\n    )\n\n    # Convert back to image format\n    rectified = rectified.data.cpu().numpy()\n    rectified = rectified[0].transpose(1, 2, 0).astype(np.uint8)\n\n    # Restore vertical offset for alignment\n    empty_image[offset_y:, :, :] = rectified\n\n    return empty_image\n`\n\n## Mask generation\nNumber of series: 4\n\nOutput mask shape: (4, 1696, 4352)\n(height × width corresponds to the rectified ECG image)\n|  Series|Zero-MV  | mV --> Pixel\n| --- | --- |\n| 0 | 707  | 78.0\n| 1 | 991  | 79.5\n| 2 | 1273  | 78.5\n| 3 | 1533  | 78.0\n\n# Training\nTo balance global ECG layout understanding with local waveform precision, two complementary training strategies were employed: Full-Crop training and Half-Crop training. Each addresses a different failure mode observed during early experiments.\n\n## Full-Crop Training (Global Context)\n\nIn full-crop training, the model is trained on the entire rectified ECG image, preserving the complete temporal span and vertical alignment of all leads.\n\n\n- Captures long-range temporal continuity of ECG signals\n\n- Preserves relative lead positioning and baseline consistency\n\n- Improves stability near signal boundaries (start/end regions)\n\n- Helps the model learn overall ECG structure and layout\n\n## Half-Crop Training (Local Detail)\n\nHalf-crop training uses random horizontal crops covering approximately half the image width, while preserving full vertical resolution. Cropped regions are then resized back to the original dimensions before training.\n\nWhy half-crop matters:\n\n- Forces the model to focus on local waveform morphology\n\n- Increases effective resolution per signal segment\n\n- Improves robustness to local noise, breaks, and occlusions\n\n- Acts as a strong form of data augmentation\n\nDATA AUGMENTATION STRATEGY\n==========================\n\nGEOMETRIC AUGMENTATIONS\n-----------------------\nAugmentation              | Parameters                                              | Probability\n--------------------------|---------------------------------------------------------|------------\nAffine transform          | Scale 0.90–1.10, Rotate ±8°, Translate ≤4%              | 0.60\nPerspective transform     | Scale 0.02–0.07                                         | 0.40\nOptical distortion        | Distort ≤0.03, Shift ≤0.02                              | 0.15\nElastic transform         | Alpha = 10, Sigma = 3                                   | 0.03\nGrid distortion           | Steps = 3, Distort limit = 0.03                         | 0.05\nAffine (shear)            | Shear ±3°                                               | 0.15\n\nPHOTOMETRIC AUGMENTATIONS\n-------------------------\nAugmentation              | Parameters                                | Probability\n--------------------------|-------------------------------------------|------------\nBrightness / Contrast     | ±0.15                                     | 0.50 (OneOf)\nCLAHE                     | Clip = 2.0, Grid = 8×8                    | 0.30 (OneOf)\nRandom Gamma              | 0.8–1.2                                   | 0.20 (OneOf)\nGaussian Noise            | Variance 5–20                             | 0.50 (OneOf)\nISO Noise                 | —                                         | 0.30 (OneOf)\nMultiplicative Noise      | 0.9–1.1                                   | 0.20 (OneOf)\nHue / Saturation / Value  | H±2, S±5, V±5                             | 0.15\n\nBLUR & SHARPENING\n-----------------\nAugmentation              | Parameters              | Probability\n--------------------------|-------------------------|------------\nSharpen                   | Alpha 0.05–0.15         | 0.20\nGaussian Blur             | Kernel size 3–5         | 0.60 (OneOf)\nMotion Blur               | Kernel size 3–5         | 0.40 (OneOf)\n\nECG-SPECIFIC OCCLUSIONS\n----------------------\nAugmentation              | Purpose                                   | Probability\n--------------------------|-------------------------------------------|------------\nCoarseDropout (small)     | Ink breaks / stains                       | 0.25\nCoarseDropout (vertical)  | Printer / scan vertical lines             | 0.20\nPerlin noise patches      | Paper dirt / aging artifacts              | ~0.50 (custom)\n\n\n\n# POST-PROCESSING & SIGNAL REFINEMENT (STAGE-2)\n======================================\n\nAfter segmentation-based signal extraction, additional temporal refinement and gap filling are applied to recover physiologically consistent ECG waveforms. This post-processing focuses on handling missing segments, enforcing periodic consistency, and stabilizing QRS morphology.\n\n1) MISSING-SEGMENT DETECTION & FILLING\n-------------------------------------\n- A binary mask indicates missing points (mask = 1).\n- Consecutive missing runs longer than a minimum length are detected per lead.\n- These runs are filled with the mean of nearby observed values in the same lead.\n- Prevents long flat gaps while preserving baseline consistency.\n\n2) BEAT-PHASE–AWARE IMPUTATION (KEY STEP)\n-----------------------------------------\n- ECG is divided into fixed 1250-sample segments.\n- Global R-peaks are detected across all leads.\n- For a missing point t:\n  * Find nearest R-peak in the same segment.\n  * Compute phase offset k = t − R_peak.\n  * Collect corresponding points (R_j + k) from other cycles in the same segment.\n  * If multiple valid values exist, average them to fill t.\n- Enforces periodic consistency and preserves waveform morphology.\n\n3) ROBUST R-PEAK DETECTION\n-------------------------\n- Band-pass filtering (5–18 Hz) emphasizes QRS complexes.\n- Pan-Tompkins–style moving window integration (MWI) is applied.\n- Energy envelopes are fused across leads.\n- Adaptive thresholding with search-back logic recovers missed beats.\n- Peaks are refined to the true positive QRS maximum.\n- A reference lead with strongest QRS energy is selected automatically.\n\n4) MODEL-LEVEL TEMPORAL REFINEMENT\n---------------------------------\n- Multi-scale temporal convolution (TCN) captures local and long-range dependencies.\n- Fourier harmonic features model dominant cardiac periodicity.\n- Gated fusion blends time-domain and frequency-domain information.\n- Residual bidirectional LSTM enforces temporal smoothness.\n- The model learns corrections only for missing regions; observed points are preserved.\n\n5) OUTPUT GUARANTEES\n-------------------\n- Observed (non-missing) points are never modified.\n- Filled values respect local continuity and global cardiac rhythm.\n- No aggressive low-pass filtering is applied; sharp QRS morphology is preserved.\n\n\n\n\n"
  }
}