{
  "id": 670511,
  "title": "27th place solution - Only Pre&Postprocessing ",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/670511",
  "author_name": "BanhMiMatOng",
  "post_date": "2026-01-28T10:36:06.693000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Acknowledgements</h2>\n<p>First of all, <strong>huge thanks to PhysioNet and Kaggle</strong> for organizing this challenging and meaningful competition.\nWe would also like to express our sincere gratitude to the community members who shared high-quality public baselines and insights. In particular, <strong>huge thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/wasupandceacar\" target=\"_blank\">@wasupandceacar</a></strong>, and others for their generous public releases, which provided extremely solid foundations to build upon.\nFinally, <strong>congratulations to the winning teams</strong> — their solutions were inspiring and very enlightening :).<br>\nP/s: we will public the notebook soon after some cleaning.</p>\n<h2>Overview</h2>\n<p>Our final solution achieved <strong>19.59 dB SNR on the Private Leaderboard</strong>, ranking <strong>27th</strong>.</p>\n<p>Overall, our pipeline follows a standard and well-proven structure:</p>\n<p><strong>Normalization / Rectification → Segmentation → Post-processing</strong></p>\n<ul>\n<li><p><strong>Stage 0 &amp; Stage 1 (Normalization / Rectification)</strong><br>\nTaken directly from <strong>hengck23’s pipeline</strong>, which remains one of the strongest and most robust approaches for ECG image alignment and grid rectification.</p></li>\n<li><p><strong>Stage 2 (Segmentation)</strong><br>\nBased on the U-Net style segmentation model released by <strong>wasupandceacar</strong>, serving as a very strong baseline for pixel-level ECG trace extraction.</p></li>\n</ul>\n<p>On top of these baselines, our main contributions are:</p>\n<ul>\n<li>Improved preprocessing before Stage 0   </li>\n<li>A new pixel-to-series postprocessing function   </li>\n<li>Additional safeguards against outliers and noise   </li>\n<li>A learned signal refinement model applied after postprocessing   </li>\n</ul>\n<h2>Improved preprocessing before stage 0</h2>\n<p>While Stage 0 and Stage 1 from hengck23 are already strong, we observed that <strong>input image quality</strong> (contrast, illumination, stains, shadows) significantly affects downstream rectification quality.</p>\n<p>Inspired by the notebook  <em>Visual QA for all stages of ECG digitization</em> by <a href=\"https://www.kaggle.com/sanpier\" target=\"_blank\">@sanpier</a>,  we adopted a <strong>dual-path strategy</strong>:</p>\n<ul>\n<li>Run Stage 0 → Stage 1 <strong>without preprocessing</strong>   </li>\n<li>Run Stage 0 → Stage 1 <strong>with preprocessing</strong>   </li>\n<li>Compare Stage-1 quality metrics   </li>\n<li>Select the better output for Stage 2<br>\nThis alone improved our public leaderboard score by <strong>~0.3 dB</strong>.   </li>\n</ul>\n<h2>Image-type-aware preprocessing</h2>\n<p>We further observed that:   </p>\n<ul>\n<li>The public image-type classifier is not always reliable   </li>\n<li>Different ECG image sources degrade in very different ways   </li>\n</ul>\n<p>Therefore, we trained <strong>our own image-type classifier</strong> and designed a <strong>safe, gated preprocessing function</strong> that:   </p>\n<ul>\n<li>Never applies destructive operations unconditionally   </li>\n<li>Uses illumination and contrast as safety gates   </li>\n<li>Applies only mild, source-specific adjustments   </li>\n</ul>\n<p>This refinement provided another <strong>~0.3 dB improvement</strong> (so in total 0.6dB).<br>\nExample code:</p>\n<pre><code>def preprocess_by_source(img_bgr, pred_src):\n   gray = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2GRAY)\n   std0 = float(gray.std())\n   illum = illumination_strength(img_bgr, sigma=35)\n   # universal safe fixes\n   if illum &gt; 0.14:\n       img_bgr = bg_correct_lab_l(img_bgr)\n   if std0 &lt; 30:\n       img_bgr = clahe_luminance_bgr(img_bgr, clip=1.15, tile=8)\n   if s in [\"0003\", \"0011\"]:  # color scans / mold color scans\n       x = grayworld_white_balance(x)\n       if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) &lt; 35:\n           x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n   elif s in [\"0006\"]:  # screen photos\n       x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n       if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) &lt; 35:\n           x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n   elif s in [\"0005\", \"0009\", \"0010\"]:  # mobile printed / stained / damaged\n       if illum &gt; 0.14:\n           x = denoise_median(x, k=3)\n       elif std0 &lt; 35:\n           x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n       # DO NOT apply strong CLAHE for 0009 (stained) unless really low contrast\n       if s == \"0009\":\n           if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) &lt; 25:\n               x = clahe_luminance_bgr(x, clip=1.1, tile=8)\n   elif s in [\"0004\", \"0012\", \"0001\"]:\n       pass\n   return img_bgr\n</code></pre>\n<p>More details and full code are available in our public notebook.   </p>\n<h2>Postprocessing: Pixel → Series</h2>\n<p>Our <strong>largest single gain (+ ~1.3 dB compared to baseline)</strong> came from redesigning the pixel-to-series conversion.</p>\n<p>Baseline approaches typically:   </p>\n<ul>\n<li>Take argmax per column   </li>\n<li>Replace low-confidence columns with a fixed baseline<br>\nWe found this often introduces quantization noise, sudden jumps and artificial flat regions. Instead, we built a more stable conversion function with three key ideas:<br>\n1) <strong>Quadratic sub-pixel refinement</strong><br>\nFor each column, we first take the hard argmax y₀, then refine it using a 3-point quadratic (parabolic) fit around (y₀−1, y₀, y₀+1). This provides sub-pixel accuracy and reduces staircase artifacts in the recovered waveform.<br>\n2) <strong>Controlled gap handling via NaN + interpolation</strong><br>\nColumns with no foreground (or very low confidence) are marked as missing (NaN), rather than being replaced with a constant baseline. We then interpolate both the extracted y-path and its confidence across these gaps using neighboring columns. This preserves continuity and avoids introducing artificial plateaus.<br>\n3) <strong>Confidence-aware despiking in pixel space</strong><br>\nAfter interpolation, we apply a small median-filter based “despike” step that only removes sharp outliers when the model confidence is low. This helps remove random jumps caused by noise/grid artifacts without suppressing real ECG peaks (e.g. QRS complexes).<br>\nFinally, the cleaned sub-pixel trace is converted into mV using the known <code>zero_mv</code> offsets and <code>mv_to_pixel</code> scale, and lightly smoothed with Savitzky–Golay filtering.   </li>\n</ul>\n<pre><code>def pixel_to_series_v2_47_quadratic_interpnan(pixel, zero_mv, length, mv_to_pixel, miss_thr=0.1):\n   \"\"\"\n   Convert stage-2 pixel probabilities (4,H,W) into 4-channel voltage series (4,W),\n   then optionally resample to `length`.\n   Key ideas:\n     - Gaussian blur for argmax stability\n     - Quadratic sub-pixel refinement around the argmax\n     - Mark low-confidence / empty columns as NaN, then interpolate (instead of baseline fill)\n     - Confidence-aware despiking in pixel space\n     - Convert pixels -&gt; mV and apply light SavGol smoothing\n   \"\"\"\n   _, H, W = pixel.shape\n   series = []\n   for j in range(4):\n       p = pixel[j]  # (H, W)\n       # 1) Smooth only to stabilize argmax (thin traces are noisy)\n       p_smooth = cv2.GaussianBlur(p, (3, 1), 0)\n       # 2) Hard argmax per column (integer y)\n       idx = p_smooth.argmax(axis=0)\n       s = idx.astype(np.float32)\n       conf = p.max(axis=0).astype(np.float32)\n       # 3) Quadratic sub-pixel refinement using (y-1, y, y+1)\n       for x in range(W):\n           y0 = idx[x]\n           if 1 &lt; y0 &lt; H - 2 and p[y0, x] &gt; miss_thr:\n               vL = float(p[y0 - 1, x])\n               vC = float(p[y0,     x])\n               vR = float(p[y0 + 1, x])\n               denom = 2.0 * (vL - 2.0 * vC + vR)\n               if abs(denom) &gt; 1e-6:\n                   delta = (vL - vR) / denom\n                   if -0.7 &lt;= delta &lt;= 0.7:\n                       s[x] = float(y0) + float(delta)\n       # 4) Missing/low-confidence columns -&gt; NaN, then interpolate\n       miss = ((p &gt; miss_thr).sum(axis=0) == 0) | (conf &lt;= miss_thr)\n       s = s.astype(float)\n       s[miss] = np.nan\n       conf = conf.astype(float)\n       conf[miss] = np.nan\n       if np.isnan(s).all():\n           s[:] = float(zero_mv[j])\n           conf[:] = 0.0\n       else:\n           s = interpolate_nans(s)\n           conf = interpolate_nans(conf)\n       series.append((s.astype(np.float32), conf.astype(np.float32)))\n   # 5) Postprocessing per channel: despike (pixel space) -&gt; convert to mV -&gt; SavGol\n   final_series = []\n   for k, (s, conf) in enumerate(series):\n       s_clean = conf_aware_despike(s, conf, kernel_size=3, threshold=2.0, conf_guard=0.3)\n       s_mv = (zero_mv[k] - s_clean) / mv_to_pixel\n       s_mv = savgol_filter(s_mv, window_length=7, polyorder=5)\n       final_series.append(s_mv)\n   series_final = np.stack(final_series).astype(np.float32)\n   # 6) Optional resample along time axis\n   if length is not None and length != W:\n       series_final = torch.from_numpy(series_final).unsqueeze(1)  # (4,1,W)\n       series_final = F.interpolate(series_final, size=length, mode=\"linear\", align_corners=False)\n       series_final = series_final.squeeze(1).cpu().numpy()\n   return series_final\n</code></pre>\n<h2>Other minor improvements</h2>\n<p>We added several lightweight postprocessing safeguards to improve robustness on difficult images. Each provided small but consistent gains (<strong>≈ +0.01 to +0.05 dB individually</strong>).   </p>\n<ul>\n<li><strong>Boundary spike cleaning:</strong> fixes short artificial jumps at lead boundaries after lead splitting.   </li>\n<li><strong>Physiological consistency:</strong> applies a soft Einthoven correction (II ≈ I + III) on short leads.   </li>\n<li><strong>Image-type–dependent calibration:</strong> adjusts pixel-to-mV scaling based on image source:   </li>\n</ul>\n<pre><code>if src in ['0005', '0006', '0009', '0010']:\n   mv_to_pixel = 78.3\nelse:\n   mv_to_pixel = 79.5\n</code></pre>\n<h2>Learned signal refinement (Black-Box)</h2>\n<p>Finally, we introduced a <strong>learned refinement stage</strong>:<br>\nThis black-box refinement consistently improved results by +0.4 to +0.6 dB, depending on configuration, and turned out to be one of the most effective final steps.<br>\nSpecifically, after extracting the 4-row voltage series, we applied a <strong>black-box 1D refinement network</strong> to denoise and correct systematic extraction errors. The refiner is an enhanced <strong>1D U-Net++</strong> (and other variants for ensembling).<br>\nPractical setup:   </p>\n<ul>\n<li>Train 1D neural networks directly on:<br><ul>\n<li>Extracted ECG signals (after postprocessing)   </li>\n<li>Ground-truth ECG waveforms   </li></ul></li>\n<li>We trained <strong>separate refiners per target sampling length / frequency</strong> (e.g., 2500/2560/5000/5120/10000/10250),   </li>\n<li>Trained multiple folds / variants, then <strong>ensembled</strong> them at inference (simple averaging).<br>\nThis final refinement step was one of our biggest gains, improving roughly <strong>+0.4 to +0.6 dB</strong> depending on configuration.   </li>\n</ul>\n<h2>Some visuals explaining some of our ideas</h2>\n<p>In these image the extracted series are in red, the blue series are ground truth ones.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F48290088cba5ded83fde0047306c341e%2Fimage_1%20(1).png?generation=1769595407668225&amp;alt=media\" alt=\"\">\nHere the segmentation is less confident and all values turn to baseline.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff62964573ee53089bc4c1e8189ce336a%2Fimage_2%20(1).png?generation=1769595434457066&amp;alt=media\" alt=\"\">\nHere the segmentation create weird confident pixel (sometimes due to grid, sometimes due to noise in B/W scans).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1e407fdabd30b0f1bfd52ebd82b2e51c%2Fimage_3%20(1).png?generation=1769595455111012&amp;alt=media\" alt=\"\">\nHere outlier peak\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1377ed8c77cfe6d756c9c7a1746b5aef%2Fimage_4%20(1).png?generation=1769595473117005&amp;alt=media\" alt=\"\">\nHere outlier due to miss-predict grid as signal\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff673b473d8f8fb3e1bce2bb7ea75e670%2Fimage_5%20(1).png?generation=1769595489612472&amp;alt=media\" alt=\"\">\nHere outlier peak\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Fb1dbe4ea26a133e677807ad0b88d55ba%2Fimage_6%20(1).png?generation=1769595505436361&amp;alt=media\" alt=\"\">\nHere the case where we haven't yet added some specific tailoring :)</p>\n<h2>Discussion on other things we have tried but not too successful :)</h2>\n<p>We also invested significant effort into rebuilding <strong>Stage 2 (segmentation)</strong> from scratch, but did not achieve improvements over the public baselines.</p>\n<ul>\n<li><p>For ground truth generation, we followed a slightly different path than most write-ups:<br>\nwe converted the provided CSV signals into <code>.dat</code> / <code>.hea</code> files and reused <strong>ecg_image_kit</strong> to regenerate ECG images, allowing us to extract point-level annotations from the resulting JSON files.</p></li>\n<li><p>From these points, we experimented with many ways of rendering segmentation masks:   </p>\n<ul>\n<li>Different drawing backends (OpenCV, Pillow, Matplotlib),   </li>\n<li>Different line modes (anti-aliasing, <code>cv2.LINE_8</code>, etc.),   </li>\n<li>Different line thicknesses.<br>\nWe observed that <strong>thinner lines consistently gave better extraction metrics</strong>.   </li></ul></li>\n<li><p>We also tried:   </p>\n<ul>\n<li>Upscaling point coordinates (×2, ×4) before drawing,   </li>\n<li>Resampling points so that, e.g., 250 Hz signals mapped exactly to 2500 columns.   </li></ul></li>\n<li><p>However, even on clean image types (e.g. <code>0001</code>), these generated masks only achieved <strong>~24–28 dB upper-bound SNR</strong> when re-extracted, and performance degraded quickly across configurations.   </p></li>\n<li><p>We trained segmentation models:   </p>\n<ul>\n<li>On <strong>non-rectified images</strong> with heavy augmentation (Public LB ≈ 13.5 dB),   </li>\n<li>On <strong>Stage-1-rectified images</strong>, where masks also had to be rectified.<br>\nWe tried two rectification strategies:<br>\n1) Draw mask first, then warp using the homography,<br>\n2) Warp point coordinates first, then draw the mask.<br>\nDespite trying multiple losses and regularizations, these models plateaued at <strong>~14–15 dB Public LB</strong>.   </li></ul></li>\n</ul>\n<p>Overall, despite extensive experiments, we were unable to outperform the public Stage-2 segmentation models.<br>\nWe suspect the issue may stem from subtle differences in GT construction (our ecg_image_kit-based pipeline vs. others), but we are still curious why rebuilding Stage 2 did not yield gains.<br>\nFinally, despite what did not work, we learned a huge amount from this competition.<br>\nThe required precision—both in image geometry and signal reconstruction—was truly impressive, and it gave us a much deeper appreciation of how sensitive ECG digitization is to even sub-pixel errors.<br>\nThanks a lot for reading through to the end 🙂<br>\nWe’re very happy to discuss ideas, failures, or details further in the comments.   </p>",
  "messages": [
    {
      "id": 3397960,
      "postDate": "2026-01-28T10:36:06.693Z",
      "content": "<h2>Acknowledgements</h2>\n<p>First of all, <strong>huge thanks to PhysioNet and Kaggle</strong> for organizing this challenging and meaningful competition.\nWe would also like to express our sincere gratitude to the community members who shared high-quality public baselines and insights. In particular, <strong>huge thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/wasupandceacar\" target=\"_blank\">@wasupandceacar</a></strong>, and others for their generous public releases, which provided extremely solid foundations to build upon.\nFinally, <strong>congratulations to the winning teams</strong> — their solutions were inspiring and very enlightening :).<br>\nP/s: we will public the notebook soon after some cleaning.</p>\n<h2>Overview</h2>\n<p>Our final solution achieved <strong>19.59 dB SNR on the Private Leaderboard</strong>, ranking <strong>27th</strong>.</p>\n<p>Overall, our pipeline follows a standard and well-proven structure:</p>\n<p><strong>Normalization / Rectification → Segmentation → Post-processing</strong></p>\n<ul>\n<li><p><strong>Stage 0 &amp; Stage 1 (Normalization / Rectification)</strong><br>\nTaken directly from <strong>hengck23’s pipeline</strong>, which remains one of the strongest and most robust approaches for ECG image alignment and grid rectification.</p></li>\n<li><p><strong>Stage 2 (Segmentation)</strong><br>\nBased on the U-Net style segmentation model released by <strong>wasupandceacar</strong>, serving as a very strong baseline for pixel-level ECG trace extraction.</p></li>\n</ul>\n<p>On top of these baselines, our main contributions are:</p>\n<ul>\n<li>Improved preprocessing before Stage 0   </li>\n<li>A new pixel-to-series postprocessing function   </li>\n<li>Additional safeguards against outliers and noise   </li>\n<li>A learned signal refinement model applied after postprocessing   </li>\n</ul>\n<h2>Improved preprocessing before stage 0</h2>\n<p>While Stage 0 and Stage 1 from hengck23 are already strong, we observed that <strong>input image quality</strong> (contrast, illumination, stains, shadows) significantly affects downstream rectification quality.</p>\n<p>Inspired by the notebook  <em>Visual QA for all stages of ECG digitization</em> by <a href=\"https://www.kaggle.com/sanpier\" target=\"_blank\">@sanpier</a>,  we adopted a <strong>dual-path strategy</strong>:</p>\n<ul>\n<li>Run Stage 0 → Stage 1 <strong>without preprocessing</strong>   </li>\n<li>Run Stage 0 → Stage 1 <strong>with preprocessing</strong>   </li>\n<li>Compare Stage-1 quality metrics   </li>\n<li>Select the better output for Stage 2<br>\nThis alone improved our public leaderboard score by <strong>~0.3 dB</strong>.   </li>\n</ul>\n<h2>Image-type-aware preprocessing</h2>\n<p>We further observed that:   </p>\n<ul>\n<li>The public image-type classifier is not always reliable   </li>\n<li>Different ECG image sources degrade in very different ways   </li>\n</ul>\n<p>Therefore, we trained <strong>our own image-type classifier</strong> and designed a <strong>safe, gated preprocessing function</strong> that:   </p>\n<ul>\n<li>Never applies destructive operations unconditionally   </li>\n<li>Uses illumination and contrast as safety gates   </li>\n<li>Applies only mild, source-specific adjustments   </li>\n</ul>\n<p>This refinement provided another <strong>~0.3 dB improvement</strong> (so in total 0.6dB).<br>\nExample code:</p>\n<pre><code>def preprocess_by_source(img_bgr, pred_src):\n   gray = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2GRAY)\n   std0 = float(gray.std())\n   illum = illumination_strength(img_bgr, sigma=35)\n   # universal safe fixes\n   if illum &gt; 0.14:\n       img_bgr = bg_correct_lab_l(img_bgr)\n   if std0 &lt; 30:\n       img_bgr = clahe_luminance_bgr(img_bgr, clip=1.15, tile=8)\n   if s in [\"0003\", \"0011\"]:  # color scans / mold color scans\n       x = grayworld_white_balance(x)\n       if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) &lt; 35:\n           x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n   elif s in [\"0006\"]:  # screen photos\n       x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n       if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) &lt; 35:\n           x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n   elif s in [\"0005\", \"0009\", \"0010\"]:  # mobile printed / stained / damaged\n       if illum &gt; 0.14:\n           x = denoise_median(x, k=3)\n       elif std0 &lt; 35:\n           x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n       # DO NOT apply strong CLAHE for 0009 (stained) unless really low contrast\n       if s == \"0009\":\n           if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) &lt; 25:\n               x = clahe_luminance_bgr(x, clip=1.1, tile=8)\n   elif s in [\"0004\", \"0012\", \"0001\"]:\n       pass\n   return img_bgr\n</code></pre>\n<p>More details and full code are available in our public notebook.   </p>\n<h2>Postprocessing: Pixel → Series</h2>\n<p>Our <strong>largest single gain (+ ~1.3 dB compared to baseline)</strong> came from redesigning the pixel-to-series conversion.</p>\n<p>Baseline approaches typically:   </p>\n<ul>\n<li>Take argmax per column   </li>\n<li>Replace low-confidence columns with a fixed baseline<br>\nWe found this often introduces quantization noise, sudden jumps and artificial flat regions. Instead, we built a more stable conversion function with three key ideas:<br>\n1) <strong>Quadratic sub-pixel refinement</strong><br>\nFor each column, we first take the hard argmax y₀, then refine it using a 3-point quadratic (parabolic) fit around (y₀−1, y₀, y₀+1). This provides sub-pixel accuracy and reduces staircase artifacts in the recovered waveform.<br>\n2) <strong>Controlled gap handling via NaN + interpolation</strong><br>\nColumns with no foreground (or very low confidence) are marked as missing (NaN), rather than being replaced with a constant baseline. We then interpolate both the extracted y-path and its confidence across these gaps using neighboring columns. This preserves continuity and avoids introducing artificial plateaus.<br>\n3) <strong>Confidence-aware despiking in pixel space</strong><br>\nAfter interpolation, we apply a small median-filter based “despike” step that only removes sharp outliers when the model confidence is low. This helps remove random jumps caused by noise/grid artifacts without suppressing real ECG peaks (e.g. QRS complexes).<br>\nFinally, the cleaned sub-pixel trace is converted into mV using the known <code>zero_mv</code> offsets and <code>mv_to_pixel</code> scale, and lightly smoothed with Savitzky–Golay filtering.   </li>\n</ul>\n<pre><code>def pixel_to_series_v2_47_quadratic_interpnan(pixel, zero_mv, length, mv_to_pixel, miss_thr=0.1):\n   \"\"\"\n   Convert stage-2 pixel probabilities (4,H,W) into 4-channel voltage series (4,W),\n   then optionally resample to `length`.\n   Key ideas:\n     - Gaussian blur for argmax stability\n     - Quadratic sub-pixel refinement around the argmax\n     - Mark low-confidence / empty columns as NaN, then interpolate (instead of baseline fill)\n     - Confidence-aware despiking in pixel space\n     - Convert pixels -&gt; mV and apply light SavGol smoothing\n   \"\"\"\n   _, H, W = pixel.shape\n   series = []\n   for j in range(4):\n       p = pixel[j]  # (H, W)\n       # 1) Smooth only to stabilize argmax (thin traces are noisy)\n       p_smooth = cv2.GaussianBlur(p, (3, 1), 0)\n       # 2) Hard argmax per column (integer y)\n       idx = p_smooth.argmax(axis=0)\n       s = idx.astype(np.float32)\n       conf = p.max(axis=0).astype(np.float32)\n       # 3) Quadratic sub-pixel refinement using (y-1, y, y+1)\n       for x in range(W):\n           y0 = idx[x]\n           if 1 &lt; y0 &lt; H - 2 and p[y0, x] &gt; miss_thr:\n               vL = float(p[y0 - 1, x])\n               vC = float(p[y0,     x])\n               vR = float(p[y0 + 1, x])\n               denom = 2.0 * (vL - 2.0 * vC + vR)\n               if abs(denom) &gt; 1e-6:\n                   delta = (vL - vR) / denom\n                   if -0.7 &lt;= delta &lt;= 0.7:\n                       s[x] = float(y0) + float(delta)\n       # 4) Missing/low-confidence columns -&gt; NaN, then interpolate\n       miss = ((p &gt; miss_thr).sum(axis=0) == 0) | (conf &lt;= miss_thr)\n       s = s.astype(float)\n       s[miss] = np.nan\n       conf = conf.astype(float)\n       conf[miss] = np.nan\n       if np.isnan(s).all():\n           s[:] = float(zero_mv[j])\n           conf[:] = 0.0\n       else:\n           s = interpolate_nans(s)\n           conf = interpolate_nans(conf)\n       series.append((s.astype(np.float32), conf.astype(np.float32)))\n   # 5) Postprocessing per channel: despike (pixel space) -&gt; convert to mV -&gt; SavGol\n   final_series = []\n   for k, (s, conf) in enumerate(series):\n       s_clean = conf_aware_despike(s, conf, kernel_size=3, threshold=2.0, conf_guard=0.3)\n       s_mv = (zero_mv[k] - s_clean) / mv_to_pixel\n       s_mv = savgol_filter(s_mv, window_length=7, polyorder=5)\n       final_series.append(s_mv)\n   series_final = np.stack(final_series).astype(np.float32)\n   # 6) Optional resample along time axis\n   if length is not None and length != W:\n       series_final = torch.from_numpy(series_final).unsqueeze(1)  # (4,1,W)\n       series_final = F.interpolate(series_final, size=length, mode=\"linear\", align_corners=False)\n       series_final = series_final.squeeze(1).cpu().numpy()\n   return series_final\n</code></pre>\n<h2>Other minor improvements</h2>\n<p>We added several lightweight postprocessing safeguards to improve robustness on difficult images. Each provided small but consistent gains (<strong>≈ +0.01 to +0.05 dB individually</strong>).   </p>\n<ul>\n<li><strong>Boundary spike cleaning:</strong> fixes short artificial jumps at lead boundaries after lead splitting.   </li>\n<li><strong>Physiological consistency:</strong> applies a soft Einthoven correction (II ≈ I + III) on short leads.   </li>\n<li><strong>Image-type–dependent calibration:</strong> adjusts pixel-to-mV scaling based on image source:   </li>\n</ul>\n<pre><code>if src in ['0005', '0006', '0009', '0010']:\n   mv_to_pixel = 78.3\nelse:\n   mv_to_pixel = 79.5\n</code></pre>\n<h2>Learned signal refinement (Black-Box)</h2>\n<p>Finally, we introduced a <strong>learned refinement stage</strong>:<br>\nThis black-box refinement consistently improved results by +0.4 to +0.6 dB, depending on configuration, and turned out to be one of the most effective final steps.<br>\nSpecifically, after extracting the 4-row voltage series, we applied a <strong>black-box 1D refinement network</strong> to denoise and correct systematic extraction errors. The refiner is an enhanced <strong>1D U-Net++</strong> (and other variants for ensembling).<br>\nPractical setup:   </p>\n<ul>\n<li>Train 1D neural networks directly on:<br><ul>\n<li>Extracted ECG signals (after postprocessing)   </li>\n<li>Ground-truth ECG waveforms   </li></ul></li>\n<li>We trained <strong>separate refiners per target sampling length / frequency</strong> (e.g., 2500/2560/5000/5120/10000/10250),   </li>\n<li>Trained multiple folds / variants, then <strong>ensembled</strong> them at inference (simple averaging).<br>\nThis final refinement step was one of our biggest gains, improving roughly <strong>+0.4 to +0.6 dB</strong> depending on configuration.   </li>\n</ul>\n<h2>Some visuals explaining some of our ideas</h2>\n<p>In these image the extracted series are in red, the blue series are ground truth ones.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F48290088cba5ded83fde0047306c341e%2Fimage_1%20(1).png?generation=1769595407668225&amp;alt=media\" alt=\"\">\nHere the segmentation is less confident and all values turn to baseline.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff62964573ee53089bc4c1e8189ce336a%2Fimage_2%20(1).png?generation=1769595434457066&amp;alt=media\" alt=\"\">\nHere the segmentation create weird confident pixel (sometimes due to grid, sometimes due to noise in B/W scans).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1e407fdabd30b0f1bfd52ebd82b2e51c%2Fimage_3%20(1).png?generation=1769595455111012&amp;alt=media\" alt=\"\">\nHere outlier peak\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1377ed8c77cfe6d756c9c7a1746b5aef%2Fimage_4%20(1).png?generation=1769595473117005&amp;alt=media\" alt=\"\">\nHere outlier due to miss-predict grid as signal\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff673b473d8f8fb3e1bce2bb7ea75e670%2Fimage_5%20(1).png?generation=1769595489612472&amp;alt=media\" alt=\"\">\nHere outlier peak\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Fb1dbe4ea26a133e677807ad0b88d55ba%2Fimage_6%20(1).png?generation=1769595505436361&amp;alt=media\" alt=\"\">\nHere the case where we haven't yet added some specific tailoring :)</p>\n<h2>Discussion on other things we have tried but not too successful :)</h2>\n<p>We also invested significant effort into rebuilding <strong>Stage 2 (segmentation)</strong> from scratch, but did not achieve improvements over the public baselines.</p>\n<ul>\n<li><p>For ground truth generation, we followed a slightly different path than most write-ups:<br>\nwe converted the provided CSV signals into <code>.dat</code> / <code>.hea</code> files and reused <strong>ecg_image_kit</strong> to regenerate ECG images, allowing us to extract point-level annotations from the resulting JSON files.</p></li>\n<li><p>From these points, we experimented with many ways of rendering segmentation masks:   </p>\n<ul>\n<li>Different drawing backends (OpenCV, Pillow, Matplotlib),   </li>\n<li>Different line modes (anti-aliasing, <code>cv2.LINE_8</code>, etc.),   </li>\n<li>Different line thicknesses.<br>\nWe observed that <strong>thinner lines consistently gave better extraction metrics</strong>.   </li></ul></li>\n<li><p>We also tried:   </p>\n<ul>\n<li>Upscaling point coordinates (×2, ×4) before drawing,   </li>\n<li>Resampling points so that, e.g., 250 Hz signals mapped exactly to 2500 columns.   </li></ul></li>\n<li><p>However, even on clean image types (e.g. <code>0001</code>), these generated masks only achieved <strong>~24–28 dB upper-bound SNR</strong> when re-extracted, and performance degraded quickly across configurations.   </p></li>\n<li><p>We trained segmentation models:   </p>\n<ul>\n<li>On <strong>non-rectified images</strong> with heavy augmentation (Public LB ≈ 13.5 dB),   </li>\n<li>On <strong>Stage-1-rectified images</strong>, where masks also had to be rectified.<br>\nWe tried two rectification strategies:<br>\n1) Draw mask first, then warp using the homography,<br>\n2) Warp point coordinates first, then draw the mask.<br>\nDespite trying multiple losses and regularizations, these models plateaued at <strong>~14–15 dB Public LB</strong>.   </li></ul></li>\n</ul>\n<p>Overall, despite extensive experiments, we were unable to outperform the public Stage-2 segmentation models.<br>\nWe suspect the issue may stem from subtle differences in GT construction (our ecg_image_kit-based pipeline vs. others), but we are still curious why rebuilding Stage 2 did not yield gains.<br>\nFinally, despite what did not work, we learned a huge amount from this competition.<br>\nThe required precision—both in image geometry and signal reconstruction—was truly impressive, and it gave us a much deeper appreciation of how sensitive ECG digitization is to even sub-pixel errors.<br>\nThanks a lot for reading through to the end 🙂<br>\nWe’re very happy to discuss ideas, failures, or details further in the comments.   </p>",
      "rawMarkdown": "\n## Acknowledgements\n\nFirst of all, **huge thanks to PhysioNet and Kaggle** for organizing this challenging and meaningful competition.\n\nWe would also like to express our sincere gratitude to the community members who shared high-quality public baselines and insights. In particular, **huge thanks to [@hengck23](https://www.kaggle.com/hengck23), [@wasupandceacar](https://www.kaggle.com/wasupandceacar)**, and others for their generous public releases, which provided extremely solid foundations to build upon.\n\nFinally, **congratulations to the winning teams** — their solutions were inspiring and very enlightening :).  \n\nP/s: we will public the notebook soon after some cleaning.\n\n## Overview\n\nOur final solution achieved **19.59 dB SNR on the Private Leaderboard**, ranking **27th**.\n  \nOverall, our pipeline follows a standard and well-proven structure:\n  \n**Normalization / Rectification → Segmentation → Post-processing**\n  \n- **Stage 0 & Stage 1 (Normalization / Rectification)**  \n  Taken directly from **hengck23’s pipeline**, which remains one of the strongest and most robust approaches for ECG image alignment and grid rectification.\n  \n- **Stage 2 (Segmentation)**  \n  Based on the U-Net style segmentation model released by **wasupandceacar**, serving as a very strong baseline for pixel-level ECG trace extraction.\n  \nOn top of these baselines, our main contributions are:\n- Improved preprocessing before Stage 0   \n- A new pixel-to-series postprocessing function   \n- Additional safeguards against outliers and noise   \n- A learned signal refinement model applied after postprocessing   \n\n## Improved preprocessing before stage 0\n   \nWhile Stage 0 and Stage 1 from hengck23 are already strong, we observed that **input image quality** (contrast, illumination, stains, shadows) significantly affects downstream rectification quality.\n   \nInspired by the notebook  *Visual QA for all stages of ECG digitization* by [@sanpier](https://www.kaggle.com/sanpier),  we adopted a **dual-path strategy**:\n   \n- Run Stage 0 → Stage 1 **without preprocessing**   \n- Run Stage 0 → Stage 1 **with preprocessing**   \n- Compare Stage-1 quality metrics   \n- Select the better output for Stage 2   \n\nThis alone improved our public leaderboard score by **~0.3 dB**.   \n\n## Image-type-aware preprocessing\n\nWe further observed that:   \n- The public image-type classifier is not always reliable   \n- Different ECG image sources degrade in very different ways   \n  \nTherefore, we trained **our own image-type classifier** and designed a **safe, gated preprocessing function** that:   \n- Never applies destructive operations unconditionally   \n- Uses illumination and contrast as safety gates   \n- Applies only mild, source-specific adjustments   \n    \nThis refinement provided another **~0.3 dB improvement** (so in total 0.6dB).   \n\nExample code:\n\n```python\ndef preprocess_by_source(img_bgr, pred_src):\n    gray = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2GRAY)\n    std0 = float(gray.std())\n    illum = illumination_strength(img_bgr, sigma=35)\n\n    # universal safe fixes\n    if illum > 0.14:\n        img_bgr = bg_correct_lab_l(img_bgr)\n\n    if std0 < 30:\n        img_bgr = clahe_luminance_bgr(img_bgr, clip=1.15, tile=8)\n\n    if s in [\"0003\", \"0011\"]:  # color scans / mold color scans\n        x = grayworld_white_balance(x)\n        if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) < 35:\n            x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n\n    elif s in [\"0006\"]:  # screen photos\n        x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n        if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) < 35:\n            x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n\n    elif s in [\"0005\", \"0009\", \"0010\"]:  # mobile printed / stained / damaged\n        if illum > 0.14:\n            x = denoise_median(x, k=3)\n        elif std0 < 35:\n            x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n\n        # DO NOT apply strong CLAHE for 0009 (stained) unless really low contrast\n        if s == \"0009\":\n            if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) < 25:\n                x = clahe_luminance_bgr(x, clip=1.1, tile=8)\n\n    elif s in [\"0004\", \"0012\", \"0001\"]:\n        pass\n\n    return img_bgr\n```\n\nMore details and full code are available in our public notebook.   \n\n## Postprocessing: Pixel → Series\n\nOur **largest single gain (+ ~1.3 dB compared to baseline)** came from redesigning the pixel-to-series conversion.\n   \nBaseline approaches typically:   \n- Take argmax per column   \n- Replace low-confidence columns with a fixed baseline   \n\nWe found this often introduces quantization noise, sudden jumps and artificial flat regions. Instead, we built a more stable conversion function with three key ideas:   \n1) **Quadratic sub-pixel refinement**     \n   For each column, we first take the hard argmax y₀, then refine it using a 3-point quadratic (parabolic) fit around (y₀−1, y₀, y₀+1). This provides sub-pixel accuracy and reduces staircase artifacts in the recovered waveform.   \n\n2) **Controlled gap handling via NaN + interpolation**     \n   Columns with no foreground (or very low confidence) are marked as missing (NaN), rather than being replaced with a constant baseline. We then interpolate both the extracted y-path and its confidence across these gaps using neighboring columns. This preserves continuity and avoids introducing artificial plateaus.   \n\n3) **Confidence-aware despiking in pixel space**     \n   After interpolation, we apply a small median-filter based “despike” step that only removes sharp outliers when the model confidence is low. This helps remove random jumps caused by noise/grid artifacts without suppressing real ECG peaks (e.g. QRS complexes).   \n\nFinally, the cleaned sub-pixel trace is converted into mV using the known `zero_mv` offsets and `mv_to_pixel` scale, and lightly smoothed with Savitzky–Golay filtering.   \n\n```python\ndef pixel_to_series_v2_47_quadratic_interpnan(pixel, zero_mv, length, mv_to_pixel, miss_thr=0.1):\n    \"\"\"\n    Convert stage-2 pixel probabilities (4,H,W) into 4-channel voltage series (4,W),\n    then optionally resample to `length`.\n\n    Key ideas:\n      - Gaussian blur for argmax stability\n      - Quadratic sub-pixel refinement around the argmax\n      - Mark low-confidence / empty columns as NaN, then interpolate (instead of baseline fill)\n      - Confidence-aware despiking in pixel space\n      - Convert pixels -> mV and apply light SavGol smoothing\n    \"\"\"\n    _, H, W = pixel.shape\n    series = []\n\n    for j in range(4):\n        p = pixel[j]  # (H, W)\n\n        # 1) Smooth only to stabilize argmax (thin traces are noisy)\n        p_smooth = cv2.GaussianBlur(p, (3, 1), 0)\n\n        # 2) Hard argmax per column (integer y)\n        idx = p_smooth.argmax(axis=0)\n        s = idx.astype(np.float32)\n        conf = p.max(axis=0).astype(np.float32)\n\n        # 3) Quadratic sub-pixel refinement using (y-1, y, y+1)\n        for x in range(W):\n            y0 = idx[x]\n            if 1 < y0 < H - 2 and p[y0, x] > miss_thr:\n                vL = float(p[y0 - 1, x])\n                vC = float(p[y0,     x])\n                vR = float(p[y0 + 1, x])\n                denom = 2.0 * (vL - 2.0 * vC + vR)\n                if abs(denom) > 1e-6:\n                    delta = (vL - vR) / denom\n                    if -0.7 <= delta <= 0.7:\n                        s[x] = float(y0) + float(delta)\n\n        # 4) Missing/low-confidence columns -> NaN, then interpolate\n        miss = ((p > miss_thr).sum(axis=0) == 0) | (conf <= miss_thr)\n        s = s.astype(float)\n        s[miss] = np.nan\n        conf = conf.astype(float)\n        conf[miss] = np.nan\n\n        if np.isnan(s).all():\n            s[:] = float(zero_mv[j])\n            conf[:] = 0.0\n        else:\n            s = interpolate_nans(s)\n            conf = interpolate_nans(conf)\n\n        series.append((s.astype(np.float32), conf.astype(np.float32)))\n\n    # 5) Postprocessing per channel: despike (pixel space) -> convert to mV -> SavGol\n    final_series = []\n    for k, (s, conf) in enumerate(series):\n        s_clean = conf_aware_despike(s, conf, kernel_size=3, threshold=2.0, conf_guard=0.3)\n        s_mv = (zero_mv[k] - s_clean) / mv_to_pixel\n        s_mv = savgol_filter(s_mv, window_length=7, polyorder=5)\n        final_series.append(s_mv)\n\n    series_final = np.stack(final_series).astype(np.float32)\n\n    # 6) Optional resample along time axis\n    if length is not None and length != W:\n        series_final = torch.from_numpy(series_final).unsqueeze(1)  # (4,1,W)\n        series_final = F.interpolate(series_final, size=length, mode=\"linear\", align_corners=False)\n        series_final = series_final.squeeze(1).cpu().numpy()\n\n    return series_final\n\n```\n\n## Other minor improvements   \n\nWe added several lightweight postprocessing safeguards to improve robustness on difficult images. Each provided small but consistent gains (**≈ +0.01 to +0.05 dB individually**).   \n\n- **Boundary spike cleaning:** fixes short artificial jumps at lead boundaries after lead splitting.   \n- **Physiological consistency:** applies a soft Einthoven correction (II ≈ I + III) on short leads.   \n- **Image-type–dependent calibration:** adjusts pixel-to-mV scaling based on image source:   \n```python\nif src in ['0005', '0006', '0009', '0010']:\n    mv_to_pixel = 78.3\nelse:\n    mv_to_pixel = 79.5\n```\n\n## Learned signal refinement (Black-Box)\n\nFinally, we introduced a **learned refinement stage**:   \n\nThis black-box refinement consistently improved results by +0.4 to +0.6 dB, depending on configuration, and turned out to be one of the most effective final steps.   \n\nSpecifically, after extracting the 4-row voltage series, we applied a **black-box 1D refinement network** to denoise and correct systematic extraction errors. The refiner is an enhanced **1D U-Net++** (and other variants for ensembling).   \n\nPractical setup:   \n- Train 1D neural networks directly on:   \n  - Extracted ECG signals (after postprocessing)   \n  - Ground-truth ECG waveforms   \n\n- We trained **separate refiners per target sampling length / frequency** (e.g., 2500/2560/5000/5120/10000/10250),   \n- Trained multiple folds / variants, then **ensembled** them at inference (simple averaging).   \n\nThis final refinement step was one of our biggest gains, improving roughly **+0.4 to +0.6 dB** depending on configuration.   \n\n## Some visuals explaining some of our ideas\nIn these image the extracted series are in red, the blue series are ground truth ones.   \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F48290088cba5ded83fde0047306c341e%2Fimage_1%20(1).png?generation=1769595407668225&alt=media)\n\nHere the segmentation is less confident and all values turn to baseline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff62964573ee53089bc4c1e8189ce336a%2Fimage_2%20(1).png?generation=1769595434457066&alt=media)\n\nHere the segmentation create weird confident pixel (sometimes due to grid, sometimes due to noise in B/W scans).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1e407fdabd30b0f1bfd52ebd82b2e51c%2Fimage_3%20(1).png?generation=1769595455111012&alt=media)\n\nHere outlier peak\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1377ed8c77cfe6d756c9c7a1746b5aef%2Fimage_4%20(1).png?generation=1769595473117005&alt=media)\n\nHere outlier due to miss-predict grid as signal\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff673b473d8f8fb3e1bce2bb7ea75e670%2Fimage_5%20(1).png?generation=1769595489612472&alt=media)\n\nHere outlier peak\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Fb1dbe4ea26a133e677807ad0b88d55ba%2Fimage_6%20(1).png?generation=1769595505436361&alt=media)\n\nHere the case where we haven't yet added some specific tailoring :)\n\n\n## Discussion on other things we have tried but not too successful :)\n\nWe also invested significant effort into rebuilding **Stage 2 (segmentation)** from scratch, but did not achieve improvements over the public baselines.\n   \n- For ground truth generation, we followed a slightly different path than most write-ups:  \n  we converted the provided CSV signals into `.dat` / `.hea` files and reused **ecg_image_kit** to regenerate ECG images, allowing us to extract point-level annotations from the resulting JSON files.\n   \n- From these points, we experimented with many ways of rendering segmentation masks:   \n  - Different drawing backends (OpenCV, Pillow, Matplotlib),   \n  - Different line modes (anti-aliasing, `cv2.LINE_8`, etc.),   \n  - Different line thicknesses.     \n  We observed that **thinner lines consistently gave better extraction metrics**.   \n   \n- We also tried:   \n  - Upscaling point coordinates (×2, ×4) before drawing,   \n  - Resampling points so that, e.g., 250 Hz signals mapped exactly to 2500 columns.   \n\n- However, even on clean image types (e.g. `0001`), these generated masks only achieved **~24–28 dB upper-bound SNR** when re-extracted, and performance degraded quickly across configurations.   \n\n- We trained segmentation models:   \n  - On **non-rectified images** with heavy augmentation (Public LB ≈ 13.5 dB),   \n  - On **Stage-1-rectified images**, where masks also had to be rectified.   \n    We tried two rectification strategies:   \n    1) Draw mask first, then warp using the homography,   \n    2) Warp point coordinates first, then draw the mask.     \n    Despite trying multiple losses and regularizations, these models plateaued at **~14–15 dB Public LB**.   \n  \nOverall, despite extensive experiments, we were unable to outperform the public Stage-2 segmentation models.     \nWe suspect the issue may stem from subtle differences in GT construction (our ecg_image_kit-based pipeline vs. others), but we are still curious why rebuilding Stage 2 did not yield gains.   \n\nFinally, despite what did not work, we learned a huge amount from this competition.     \nThe required precision—both in image geometry and signal reconstruction—was truly impressive, and it gave us a much deeper appreciation of how sensitive ECG digitization is to even sub-pixel errors.   \n\nThanks a lot for reading through to the end 🙂     \nWe’re very happy to discuss ideas, failures, or details further in the comments.   \n",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3397960": "\n## Acknowledgements\n\nFirst of all, **huge thanks to PhysioNet and Kaggle** for organizing this challenging and meaningful competition.\n\nWe would also like to express our sincere gratitude to the community members who shared high-quality public baselines and insights. In particular, **huge thanks to [@hengck23](https://www.kaggle.com/hengck23), [@wasupandceacar](https://www.kaggle.com/wasupandceacar)**, and others for their generous public releases, which provided extremely solid foundations to build upon.\n\nFinally, **congratulations to the winning teams** — their solutions were inspiring and very enlightening :).  \n\nP/s: we will public the notebook soon after some cleaning.\n\n## Overview\n\nOur final solution achieved **19.59 dB SNR on the Private Leaderboard**, ranking **27th**.\n  \nOverall, our pipeline follows a standard and well-proven structure:\n  \n**Normalization / Rectification → Segmentation → Post-processing**\n  \n- **Stage 0 & Stage 1 (Normalization / Rectification)**  \n  Taken directly from **hengck23’s pipeline**, which remains one of the strongest and most robust approaches for ECG image alignment and grid rectification.\n  \n- **Stage 2 (Segmentation)**  \n  Based on the U-Net style segmentation model released by **wasupandceacar**, serving as a very strong baseline for pixel-level ECG trace extraction.\n  \nOn top of these baselines, our main contributions are:\n- Improved preprocessing before Stage 0   \n- A new pixel-to-series postprocessing function   \n- Additional safeguards against outliers and noise   \n- A learned signal refinement model applied after postprocessing   \n\n## Improved preprocessing before stage 0\n   \nWhile Stage 0 and Stage 1 from hengck23 are already strong, we observed that **input image quality** (contrast, illumination, stains, shadows) significantly affects downstream rectification quality.\n   \nInspired by the notebook  *Visual QA for all stages of ECG digitization* by [@sanpier](https://www.kaggle.com/sanpier),  we adopted a **dual-path strategy**:\n   \n- Run Stage 0 → Stage 1 **without preprocessing**   \n- Run Stage 0 → Stage 1 **with preprocessing**   \n- Compare Stage-1 quality metrics   \n- Select the better output for Stage 2   \n\nThis alone improved our public leaderboard score by **~0.3 dB**.   \n\n## Image-type-aware preprocessing\n\nWe further observed that:   \n- The public image-type classifier is not always reliable   \n- Different ECG image sources degrade in very different ways   \n  \nTherefore, we trained **our own image-type classifier** and designed a **safe, gated preprocessing function** that:   \n- Never applies destructive operations unconditionally   \n- Uses illumination and contrast as safety gates   \n- Applies only mild, source-specific adjustments   \n    \nThis refinement provided another **~0.3 dB improvement** (so in total 0.6dB).   \n\nExample code:\n\n```python\ndef preprocess_by_source(img_bgr, pred_src):\n    gray = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2GRAY)\n    std0 = float(gray.std())\n    illum = illumination_strength(img_bgr, sigma=35)\n\n    # universal safe fixes\n    if illum > 0.14:\n        img_bgr = bg_correct_lab_l(img_bgr)\n\n    if std0 < 30:\n        img_bgr = clahe_luminance_bgr(img_bgr, clip=1.15, tile=8)\n\n    if s in [\"0003\", \"0011\"]:  # color scans / mold color scans\n        x = grayworld_white_balance(x)\n        if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) < 35:\n            x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n\n    elif s in [\"0006\"]:  # screen photos\n        x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n        if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) < 35:\n            x = clahe_luminance_bgr(x, clip=1.2, tile=8)\n\n    elif s in [\"0005\", \"0009\", \"0010\"]:  # mobile printed / stained / damaged\n        if illum > 0.14:\n            x = denoise_median(x, k=3)\n        elif std0 < 35:\n            x = denoise_bilateral(x, d=5, sigmaColor=25, sigmaSpace=25)\n\n        # DO NOT apply strong CLAHE for 0009 (stained) unless really low contrast\n        if s == \"0009\":\n            if float(cv2.cvtColor(x, cv2.COLOR_BGR2GRAY).std()) < 25:\n                x = clahe_luminance_bgr(x, clip=1.1, tile=8)\n\n    elif s in [\"0004\", \"0012\", \"0001\"]:\n        pass\n\n    return img_bgr\n```\n\nMore details and full code are available in our public notebook.   \n\n## Postprocessing: Pixel → Series\n\nOur **largest single gain (+ ~1.3 dB compared to baseline)** came from redesigning the pixel-to-series conversion.\n   \nBaseline approaches typically:   \n- Take argmax per column   \n- Replace low-confidence columns with a fixed baseline   \n\nWe found this often introduces quantization noise, sudden jumps and artificial flat regions. Instead, we built a more stable conversion function with three key ideas:   \n1) **Quadratic sub-pixel refinement**     \n   For each column, we first take the hard argmax y₀, then refine it using a 3-point quadratic (parabolic) fit around (y₀−1, y₀, y₀+1). This provides sub-pixel accuracy and reduces staircase artifacts in the recovered waveform.   \n\n2) **Controlled gap handling via NaN + interpolation**     \n   Columns with no foreground (or very low confidence) are marked as missing (NaN), rather than being replaced with a constant baseline. We then interpolate both the extracted y-path and its confidence across these gaps using neighboring columns. This preserves continuity and avoids introducing artificial plateaus.   \n\n3) **Confidence-aware despiking in pixel space**     \n   After interpolation, we apply a small median-filter based “despike” step that only removes sharp outliers when the model confidence is low. This helps remove random jumps caused by noise/grid artifacts without suppressing real ECG peaks (e.g. QRS complexes).   \n\nFinally, the cleaned sub-pixel trace is converted into mV using the known `zero_mv` offsets and `mv_to_pixel` scale, and lightly smoothed with Savitzky–Golay filtering.   \n\n```python\ndef pixel_to_series_v2_47_quadratic_interpnan(pixel, zero_mv, length, mv_to_pixel, miss_thr=0.1):\n    \"\"\"\n    Convert stage-2 pixel probabilities (4,H,W) into 4-channel voltage series (4,W),\n    then optionally resample to `length`.\n\n    Key ideas:\n      - Gaussian blur for argmax stability\n      - Quadratic sub-pixel refinement around the argmax\n      - Mark low-confidence / empty columns as NaN, then interpolate (instead of baseline fill)\n      - Confidence-aware despiking in pixel space\n      - Convert pixels -> mV and apply light SavGol smoothing\n    \"\"\"\n    _, H, W = pixel.shape\n    series = []\n\n    for j in range(4):\n        p = pixel[j]  # (H, W)\n\n        # 1) Smooth only to stabilize argmax (thin traces are noisy)\n        p_smooth = cv2.GaussianBlur(p, (3, 1), 0)\n\n        # 2) Hard argmax per column (integer y)\n        idx = p_smooth.argmax(axis=0)\n        s = idx.astype(np.float32)\n        conf = p.max(axis=0).astype(np.float32)\n\n        # 3) Quadratic sub-pixel refinement using (y-1, y, y+1)\n        for x in range(W):\n            y0 = idx[x]\n            if 1 < y0 < H - 2 and p[y0, x] > miss_thr:\n                vL = float(p[y0 - 1, x])\n                vC = float(p[y0,     x])\n                vR = float(p[y0 + 1, x])\n                denom = 2.0 * (vL - 2.0 * vC + vR)\n                if abs(denom) > 1e-6:\n                    delta = (vL - vR) / denom\n                    if -0.7 <= delta <= 0.7:\n                        s[x] = float(y0) + float(delta)\n\n        # 4) Missing/low-confidence columns -> NaN, then interpolate\n        miss = ((p > miss_thr).sum(axis=0) == 0) | (conf <= miss_thr)\n        s = s.astype(float)\n        s[miss] = np.nan\n        conf = conf.astype(float)\n        conf[miss] = np.nan\n\n        if np.isnan(s).all():\n            s[:] = float(zero_mv[j])\n            conf[:] = 0.0\n        else:\n            s = interpolate_nans(s)\n            conf = interpolate_nans(conf)\n\n        series.append((s.astype(np.float32), conf.astype(np.float32)))\n\n    # 5) Postprocessing per channel: despike (pixel space) -> convert to mV -> SavGol\n    final_series = []\n    for k, (s, conf) in enumerate(series):\n        s_clean = conf_aware_despike(s, conf, kernel_size=3, threshold=2.0, conf_guard=0.3)\n        s_mv = (zero_mv[k] - s_clean) / mv_to_pixel\n        s_mv = savgol_filter(s_mv, window_length=7, polyorder=5)\n        final_series.append(s_mv)\n\n    series_final = np.stack(final_series).astype(np.float32)\n\n    # 6) Optional resample along time axis\n    if length is not None and length != W:\n        series_final = torch.from_numpy(series_final).unsqueeze(1)  # (4,1,W)\n        series_final = F.interpolate(series_final, size=length, mode=\"linear\", align_corners=False)\n        series_final = series_final.squeeze(1).cpu().numpy()\n\n    return series_final\n\n```\n\n## Other minor improvements   \n\nWe added several lightweight postprocessing safeguards to improve robustness on difficult images. Each provided small but consistent gains (**≈ +0.01 to +0.05 dB individually**).   \n\n- **Boundary spike cleaning:** fixes short artificial jumps at lead boundaries after lead splitting.   \n- **Physiological consistency:** applies a soft Einthoven correction (II ≈ I + III) on short leads.   \n- **Image-type–dependent calibration:** adjusts pixel-to-mV scaling based on image source:   \n```python\nif src in ['0005', '0006', '0009', '0010']:\n    mv_to_pixel = 78.3\nelse:\n    mv_to_pixel = 79.5\n```\n\n## Learned signal refinement (Black-Box)\n\nFinally, we introduced a **learned refinement stage**:   \n\nThis black-box refinement consistently improved results by +0.4 to +0.6 dB, depending on configuration, and turned out to be one of the most effective final steps.   \n\nSpecifically, after extracting the 4-row voltage series, we applied a **black-box 1D refinement network** to denoise and correct systematic extraction errors. The refiner is an enhanced **1D U-Net++** (and other variants for ensembling).   \n\nPractical setup:   \n- Train 1D neural networks directly on:   \n  - Extracted ECG signals (after postprocessing)   \n  - Ground-truth ECG waveforms   \n\n- We trained **separate refiners per target sampling length / frequency** (e.g., 2500/2560/5000/5120/10000/10250),   \n- Trained multiple folds / variants, then **ensembled** them at inference (simple averaging).   \n\nThis final refinement step was one of our biggest gains, improving roughly **+0.4 to +0.6 dB** depending on configuration.   \n\n## Some visuals explaining some of our ideas\nIn these image the extracted series are in red, the blue series are ground truth ones.   \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F48290088cba5ded83fde0047306c341e%2Fimage_1%20(1).png?generation=1769595407668225&alt=media)\n\nHere the segmentation is less confident and all values turn to baseline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff62964573ee53089bc4c1e8189ce336a%2Fimage_2%20(1).png?generation=1769595434457066&alt=media)\n\nHere the segmentation create weird confident pixel (sometimes due to grid, sometimes due to noise in B/W scans).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1e407fdabd30b0f1bfd52ebd82b2e51c%2Fimage_3%20(1).png?generation=1769595455111012&alt=media)\n\nHere outlier peak\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2F1377ed8c77cfe6d756c9c7a1746b5aef%2Fimage_4%20(1).png?generation=1769595473117005&alt=media)\n\nHere outlier due to miss-predict grid as signal\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Ff673b473d8f8fb3e1bce2bb7ea75e670%2Fimage_5%20(1).png?generation=1769595489612472&alt=media)\n\nHere outlier peak\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21537961%2Fb1dbe4ea26a133e677807ad0b88d55ba%2Fimage_6%20(1).png?generation=1769595505436361&alt=media)\n\nHere the case where we haven't yet added some specific tailoring :)\n\n\n## Discussion on other things we have tried but not too successful :)\n\nWe also invested significant effort into rebuilding **Stage 2 (segmentation)** from scratch, but did not achieve improvements over the public baselines.\n   \n- For ground truth generation, we followed a slightly different path than most write-ups:  \n  we converted the provided CSV signals into `.dat` / `.hea` files and reused **ecg_image_kit** to regenerate ECG images, allowing us to extract point-level annotations from the resulting JSON files.\n   \n- From these points, we experimented with many ways of rendering segmentation masks:   \n  - Different drawing backends (OpenCV, Pillow, Matplotlib),   \n  - Different line modes (anti-aliasing, `cv2.LINE_8`, etc.),   \n  - Different line thicknesses.     \n  We observed that **thinner lines consistently gave better extraction metrics**.   \n   \n- We also tried:   \n  - Upscaling point coordinates (×2, ×4) before drawing,   \n  - Resampling points so that, e.g., 250 Hz signals mapped exactly to 2500 columns.   \n\n- However, even on clean image types (e.g. `0001`), these generated masks only achieved **~24–28 dB upper-bound SNR** when re-extracted, and performance degraded quickly across configurations.   \n\n- We trained segmentation models:   \n  - On **non-rectified images** with heavy augmentation (Public LB ≈ 13.5 dB),   \n  - On **Stage-1-rectified images**, where masks also had to be rectified.   \n    We tried two rectification strategies:   \n    1) Draw mask first, then warp using the homography,   \n    2) Warp point coordinates first, then draw the mask.     \n    Despite trying multiple losses and regularizations, these models plateaued at **~14–15 dB Public LB**.   \n  \nOverall, despite extensive experiments, we were unable to outperform the public Stage-2 segmentation models.     \nWe suspect the issue may stem from subtle differences in GT construction (our ecg_image_kit-based pipeline vs. others), but we are still curious why rebuilding Stage 2 did not yield gains.   \n\nFinally, despite what did not work, we learned a huge amount from this competition.     \nThe required precision—both in image geometry and signal reconstruction—was truly impressive, and it gave us a much deeper appreciation of how sensitive ECG digitization is to even sub-pixel errors.   \n\nThanks a lot for reading through to the end 🙂     \nWe’re very happy to discuss ideas, failures, or details further in the comments.   \n"
  }
}