{
  "id": 672944,
  "title": "51st place solution",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/672944",
  "author_name": "COLA KE LIU",
  "post_date": "2026-02-11T11:15:08.544000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>3-Stage Pipeline with Robust Signal Extraction</h1>\n<p>First of all, thanks to the organizers for hosting this rerun of the ECG digitization challenge. Here is an overview of our team's approach.</p>\n<h2>Overview</h2>\n<p>Our solution breaks the problem down into three distinct sub-tasks:</p>\n<ol>\n<li><strong>Stage 0: ROI Extraction &amp; Orientation:</strong> Locating the ECG chart and correcting perspective.</li>\n<li><strong>Stage 1: Rectification:</strong> Unwarping the image to correct non-linear paper deformations.</li>\n<li><strong>Stage 2: Segmentation:</strong> Predicting the probability map of the signal trace.</li>\n</ol>\n<p><strong>Key Improvement:</strong> The critical difference in our final submission lies in our <strong>Signal Extraction (Post-processing)</strong> strategy (<code>fill_pro.py</code>), where we replaced simple peak detection with Gaussian smoothing and advanced interpolation to handle broken signal lines and noise.</p>\n<hr>\n<h2>1. Preprocessing &amp; Source Classification</h2>\n<p>We noticed that the test set contains images from various sources (scanners, photos, different resolutions). Applying the same preprocessing to all images hurts performance.</p>\n<ul>\n<li><strong>Source Classifier:</strong> We trained an <code>EfficientNet-B2</code> to classify the image \"source\" (12 classes).</li>\n<li><strong>Adaptive Preprocessing:</strong> Based on the predicted source ID, we apply specific OpenCV operations:<ul>\n<li><strong>CLAHE:</strong> For low-contrast images.</li>\n<li><strong>Gray-world White Balance:</strong> For photos with color casts.</li>\n<li><strong>Bilateral/Median Filtering:</strong> For noisy scans.</li>\n<li><strong>Background Removal:</strong> Using morphological opening (Lab color space) to remove uneven lighting.</li></ul></li>\n</ul>\n<h2>2. The Pipeline Architecture</h2>\n<h3>Stage 0: Global Alignment (Homography)</h3>\n<ul>\n<li><strong>Model:</strong> ResNet34 Encoder.</li>\n<li><strong>Task:</strong> Predicts keypoints (corners/anchors) and orientation.</li>\n<li><strong>Action:</strong> We calculate a Homography matrix to warp the input image into a standard coordinate system. This handles rotation and perspective distortion.</li>\n</ul>\n<h3>Stage 1: Fine-Grained Rectification</h3>\n<ul>\n<li><strong>Model:</strong> U-Net style architecture.</li>\n<li><strong>Task:</strong> Detects the grid lines (horizontal and vertical) and grid intersections.</li>\n<li><strong>Action:</strong> We use the predicted grid points to create a dense flow map. Using <code>torch.nn.functional.grid_sample</code>, we \"unwarp\" the image. This is crucial for photos of paper that is curved or folded, ensuring the time-axis is linear.</li>\n</ul>\n<h3>Stage 2: 1D Signal Segmentation</h3>\n<ul>\n<li><strong>Model:</strong> ResNet34 Encoder + <code>CoordUnetDecoder</code>.</li>\n<li><strong>Feature:</strong> We inject Coordinate Convolutions (<code>CoordConv</code>) in the decoder to help the model understand the spatial position of the signals (since ECG leads have fixed positions).</li>\n<li><strong>Output:</strong> A probability map (Heatmap) of the signal trace.</li>\n</ul>\n<hr>\n<h2>3. The \"Secret Sauce\": Robust Signal Extraction (<code>fill_pro</code>)</h2>\n<p>In the baseline approach, extracting the time series from the Stage 2 heatmap was done via a simple <code>argmax</code>. However, this fails when:</p>\n<ol>\n<li>The ink is faded (missing signal).</li>\n<li>There is noise (salt-and-pepper artifacts).</li>\n<li>The line is jagged.</li>\n</ol>\n<p>We implemented a robust extraction algorithm (<code>fill_pro.py</code>) that significantly improved our SNR score.</p>\n<h3>A. Gaussian Probability Smoothing</h3>\n<p>Before extracting the coordinates, we apply a Gaussian Blur to the probability map output by the model.</p>\n<ul>\n<li><strong>Why:</strong> The raw pixel predictions can be noisy. Blurring aggregates local evidence, making the <code>argmax</code> (peak detection) much more stable and continuous.</li>\n<li><strong>Parameters:</strong> We tuned <code>sigma=0.29</code> and <code>threshold=0.11</code> for the best leaderboard performance.</li>\n</ul>\n<h3>B. Handling Signal Gaps (Interpolation vs. Zero-Filling)</h3>\n<p>The baseline approach filled missing data (where model confidence &lt; threshold) with the <code>zero_mv</code> (isoelectric line). This causes massive penalties in the SNR metric because a missing peak is treated as a flat line (high error).</p>\n<p>We replaced this with <strong>Linear/Spline Interpolation</strong>.</p>\n<ul>\n<li><strong>Impact:</strong> If a QRS complex has a gap in the middle due to poor scanning, interpolation connects the dots, preserving the shape much better than dropping to zero.</li>\n</ul>\n<h3>C. Einthoven's Law Correction</h3>\n<p>We apply a physics-based constraint for the limb leads. According to Einthoven's Law, Lead II should equal Lead I + Lead III. We calculate the error between the measured leads and distribute this error to correct the signals.</p>\n<p>This post-processing step ensures the extracted signals are electrically consistent, acting as a powerful regularizer against noise in any single lead.</p>\n<hr>\n<h2>4. Summary of Parameters</h2>\n<p>For the final submission, we used the following configuration in <code>fill_pro</code>:</p>\n<ul>\n<li><strong>Smoothing:</strong> Gaussian (<code>sigma=0.29</code>).</li>\n<li><strong>Interpolation:</strong> Linear (<code>scipy.interpolate.interp1d</code>).</li>\n<li><strong>Threshold:</strong> <code>0.11</code> (Pixels with probability below this are treated as missing and interpolated).</li>\n<li><strong>Post-Smoothing:</strong> Savitzky-Golay filter (<code>window=7</code>, <code>poly=2</code>) applied to the final 1D series to remove high-frequency digitization noise.</li>\n</ul>",
  "messages": [
    {
      "id": 3404855,
      "postDate": "2026-02-11T11:15:08.543Z",
      "content": "<h1>3-Stage Pipeline with Robust Signal Extraction</h1>\n<p>First of all, thanks to the organizers for hosting this rerun of the ECG digitization challenge. Here is an overview of our team's approach.</p>\n<h2>Overview</h2>\n<p>Our solution breaks the problem down into three distinct sub-tasks:</p>\n<ol>\n<li><strong>Stage 0: ROI Extraction &amp; Orientation:</strong> Locating the ECG chart and correcting perspective.</li>\n<li><strong>Stage 1: Rectification:</strong> Unwarping the image to correct non-linear paper deformations.</li>\n<li><strong>Stage 2: Segmentation:</strong> Predicting the probability map of the signal trace.</li>\n</ol>\n<p><strong>Key Improvement:</strong> The critical difference in our final submission lies in our <strong>Signal Extraction (Post-processing)</strong> strategy (<code>fill_pro.py</code>), where we replaced simple peak detection with Gaussian smoothing and advanced interpolation to handle broken signal lines and noise.</p>\n<hr>\n<h2>1. Preprocessing &amp; Source Classification</h2>\n<p>We noticed that the test set contains images from various sources (scanners, photos, different resolutions). Applying the same preprocessing to all images hurts performance.</p>\n<ul>\n<li><strong>Source Classifier:</strong> We trained an <code>EfficientNet-B2</code> to classify the image \"source\" (12 classes).</li>\n<li><strong>Adaptive Preprocessing:</strong> Based on the predicted source ID, we apply specific OpenCV operations:<ul>\n<li><strong>CLAHE:</strong> For low-contrast images.</li>\n<li><strong>Gray-world White Balance:</strong> For photos with color casts.</li>\n<li><strong>Bilateral/Median Filtering:</strong> For noisy scans.</li>\n<li><strong>Background Removal:</strong> Using morphological opening (Lab color space) to remove uneven lighting.</li></ul></li>\n</ul>\n<h2>2. The Pipeline Architecture</h2>\n<h3>Stage 0: Global Alignment (Homography)</h3>\n<ul>\n<li><strong>Model:</strong> ResNet34 Encoder.</li>\n<li><strong>Task:</strong> Predicts keypoints (corners/anchors) and orientation.</li>\n<li><strong>Action:</strong> We calculate a Homography matrix to warp the input image into a standard coordinate system. This handles rotation and perspective distortion.</li>\n</ul>\n<h3>Stage 1: Fine-Grained Rectification</h3>\n<ul>\n<li><strong>Model:</strong> U-Net style architecture.</li>\n<li><strong>Task:</strong> Detects the grid lines (horizontal and vertical) and grid intersections.</li>\n<li><strong>Action:</strong> We use the predicted grid points to create a dense flow map. Using <code>torch.nn.functional.grid_sample</code>, we \"unwarp\" the image. This is crucial for photos of paper that is curved or folded, ensuring the time-axis is linear.</li>\n</ul>\n<h3>Stage 2: 1D Signal Segmentation</h3>\n<ul>\n<li><strong>Model:</strong> ResNet34 Encoder + <code>CoordUnetDecoder</code>.</li>\n<li><strong>Feature:</strong> We inject Coordinate Convolutions (<code>CoordConv</code>) in the decoder to help the model understand the spatial position of the signals (since ECG leads have fixed positions).</li>\n<li><strong>Output:</strong> A probability map (Heatmap) of the signal trace.</li>\n</ul>\n<hr>\n<h2>3. The \"Secret Sauce\": Robust Signal Extraction (<code>fill_pro</code>)</h2>\n<p>In the baseline approach, extracting the time series from the Stage 2 heatmap was done via a simple <code>argmax</code>. However, this fails when:</p>\n<ol>\n<li>The ink is faded (missing signal).</li>\n<li>There is noise (salt-and-pepper artifacts).</li>\n<li>The line is jagged.</li>\n</ol>\n<p>We implemented a robust extraction algorithm (<code>fill_pro.py</code>) that significantly improved our SNR score.</p>\n<h3>A. Gaussian Probability Smoothing</h3>\n<p>Before extracting the coordinates, we apply a Gaussian Blur to the probability map output by the model.</p>\n<ul>\n<li><strong>Why:</strong> The raw pixel predictions can be noisy. Blurring aggregates local evidence, making the <code>argmax</code> (peak detection) much more stable and continuous.</li>\n<li><strong>Parameters:</strong> We tuned <code>sigma=0.29</code> and <code>threshold=0.11</code> for the best leaderboard performance.</li>\n</ul>\n<h3>B. Handling Signal Gaps (Interpolation vs. Zero-Filling)</h3>\n<p>The baseline approach filled missing data (where model confidence &lt; threshold) with the <code>zero_mv</code> (isoelectric line). This causes massive penalties in the SNR metric because a missing peak is treated as a flat line (high error).</p>\n<p>We replaced this with <strong>Linear/Spline Interpolation</strong>.</p>\n<ul>\n<li><strong>Impact:</strong> If a QRS complex has a gap in the middle due to poor scanning, interpolation connects the dots, preserving the shape much better than dropping to zero.</li>\n</ul>\n<h3>C. Einthoven's Law Correction</h3>\n<p>We apply a physics-based constraint for the limb leads. According to Einthoven's Law, Lead II should equal Lead I + Lead III. We calculate the error between the measured leads and distribute this error to correct the signals.</p>\n<p>This post-processing step ensures the extracted signals are electrically consistent, acting as a powerful regularizer against noise in any single lead.</p>\n<hr>\n<h2>4. Summary of Parameters</h2>\n<p>For the final submission, we used the following configuration in <code>fill_pro</code>:</p>\n<ul>\n<li><strong>Smoothing:</strong> Gaussian (<code>sigma=0.29</code>).</li>\n<li><strong>Interpolation:</strong> Linear (<code>scipy.interpolate.interp1d</code>).</li>\n<li><strong>Threshold:</strong> <code>0.11</code> (Pixels with probability below this are treated as missing and interpolated).</li>\n<li><strong>Post-Smoothing:</strong> Savitzky-Golay filter (<code>window=7</code>, <code>poly=2</code>) applied to the final 1D series to remove high-frequency digitization noise.</li>\n</ul>",
      "rawMarkdown": "# 3-Stage Pipeline with Robust Signal Extraction\n\nFirst of all, thanks to the organizers for hosting this rerun of the ECG digitization challenge. Here is an overview of our team's approach.\n\n## Overview\nOur solution breaks the problem down into three distinct sub-tasks:\n1.  **Stage 0: ROI Extraction & Orientation:** Locating the ECG chart and correcting perspective.\n2.  **Stage 1: Rectification:** Unwarping the image to correct non-linear paper deformations.\n3.  **Stage 2: Segmentation:** Predicting the probability map of the signal trace.\n\n**Key Improvement:** The critical difference in our final submission lies in our **Signal Extraction (Post-processing)** strategy (`fill_pro.py`), where we replaced simple peak detection with Gaussian smoothing and advanced interpolation to handle broken signal lines and noise.\n\n---\n\n## 1. Preprocessing & Source Classification\nWe noticed that the test set contains images from various sources (scanners, photos, different resolutions). Applying the same preprocessing to all images hurts performance.\n\n*   **Source Classifier:** We trained an `EfficientNet-B2` to classify the image \"source\" (12 classes).\n*   **Adaptive Preprocessing:** Based on the predicted source ID, we apply specific OpenCV operations:\n    *   **CLAHE:** For low-contrast images.\n    *   **Gray-world White Balance:** For photos with color casts.\n    *   **Bilateral/Median Filtering:** For noisy scans.\n    *   **Background Removal:** Using morphological opening (Lab color space) to remove uneven lighting.\n\n## 2. The Pipeline Architecture\n\n### Stage 0: Global Alignment (Homography)\n*   **Model:** ResNet34 Encoder.\n*   **Task:** Predicts keypoints (corners/anchors) and orientation.\n*   **Action:** We calculate a Homography matrix to warp the input image into a standard coordinate system. This handles rotation and perspective distortion.\n\n### Stage 1: Fine-Grained Rectification\n*   **Model:** U-Net style architecture.\n*   **Task:** Detects the grid lines (horizontal and vertical) and grid intersections.\n*   **Action:** We use the predicted grid points to create a dense flow map. Using `torch.nn.functional.grid_sample`, we \"unwarp\" the image. This is crucial for photos of paper that is curved or folded, ensuring the time-axis is linear.\n\n### Stage 2: 1D Signal Segmentation\n*   **Model:** ResNet34 Encoder + `CoordUnetDecoder`.\n*   **Feature:** We inject Coordinate Convolutions (`CoordConv`) in the decoder to help the model understand the spatial position of the signals (since ECG leads have fixed positions).\n*   **Output:** A probability map (Heatmap) of the signal trace.\n\n---\n\n## 3. The \"Secret Sauce\": Robust Signal Extraction (`fill_pro`)\n\nIn the baseline approach, extracting the time series from the Stage 2 heatmap was done via a simple `argmax`. However, this fails when:\n1.  The ink is faded (missing signal).\n2.  There is noise (salt-and-pepper artifacts).\n3.  The line is jagged.\n\nWe implemented a robust extraction algorithm (`fill_pro.py`) that significantly improved our SNR score.\n\n### A. Gaussian Probability Smoothing\nBefore extracting the coordinates, we apply a Gaussian Blur to the probability map output by the model.\n\n*   **Why:** The raw pixel predictions can be noisy. Blurring aggregates local evidence, making the `argmax` (peak detection) much more stable and continuous.\n*   **Parameters:** We tuned `sigma=0.29` and `threshold=0.11` for the best leaderboard performance.\n\n### B. Handling Signal Gaps (Interpolation vs. Zero-Filling)\nThe baseline approach filled missing data (where model confidence < threshold) with the `zero_mv` (isoelectric line). This causes massive penalties in the SNR metric because a missing peak is treated as a flat line (high error).\n\nWe replaced this with **Linear/Spline Interpolation**.\n\n*   **Impact:** If a QRS complex has a gap in the middle due to poor scanning, interpolation connects the dots, preserving the shape much better than dropping to zero.\n\n### C. Einthoven's Law Correction\nWe apply a physics-based constraint for the limb leads. According to Einthoven's Law, Lead II should equal Lead I + Lead III. We calculate the error between the measured leads and distribute this error to correct the signals.\n\nThis post-processing step ensures the extracted signals are electrically consistent, acting as a powerful regularizer against noise in any single lead.\n\n---\n\n## 4. Summary of Parameters\nFor the final submission, we used the following configuration in `fill_pro`:\n*   **Smoothing:** Gaussian (`sigma=0.29`).\n*   **Interpolation:** Linear (`scipy.interpolate.interp1d`).\n*   **Threshold:** `0.11` (Pixels with probability below this are treated as missing and interpolated).\n*   **Post-Smoothing:** Savitzky-Golay filter (`window=7`, `poly=2`) applied to the final 1D series to remove high-frequency digitization noise.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3404855": "# 3-Stage Pipeline with Robust Signal Extraction\n\nFirst of all, thanks to the organizers for hosting this rerun of the ECG digitization challenge. Here is an overview of our team's approach.\n\n## Overview\nOur solution breaks the problem down into three distinct sub-tasks:\n1.  **Stage 0: ROI Extraction & Orientation:** Locating the ECG chart and correcting perspective.\n2.  **Stage 1: Rectification:** Unwarping the image to correct non-linear paper deformations.\n3.  **Stage 2: Segmentation:** Predicting the probability map of the signal trace.\n\n**Key Improvement:** The critical difference in our final submission lies in our **Signal Extraction (Post-processing)** strategy (`fill_pro.py`), where we replaced simple peak detection with Gaussian smoothing and advanced interpolation to handle broken signal lines and noise.\n\n---\n\n## 1. Preprocessing & Source Classification\nWe noticed that the test set contains images from various sources (scanners, photos, different resolutions). Applying the same preprocessing to all images hurts performance.\n\n*   **Source Classifier:** We trained an `EfficientNet-B2` to classify the image \"source\" (12 classes).\n*   **Adaptive Preprocessing:** Based on the predicted source ID, we apply specific OpenCV operations:\n    *   **CLAHE:** For low-contrast images.\n    *   **Gray-world White Balance:** For photos with color casts.\n    *   **Bilateral/Median Filtering:** For noisy scans.\n    *   **Background Removal:** Using morphological opening (Lab color space) to remove uneven lighting.\n\n## 2. The Pipeline Architecture\n\n### Stage 0: Global Alignment (Homography)\n*   **Model:** ResNet34 Encoder.\n*   **Task:** Predicts keypoints (corners/anchors) and orientation.\n*   **Action:** We calculate a Homography matrix to warp the input image into a standard coordinate system. This handles rotation and perspective distortion.\n\n### Stage 1: Fine-Grained Rectification\n*   **Model:** U-Net style architecture.\n*   **Task:** Detects the grid lines (horizontal and vertical) and grid intersections.\n*   **Action:** We use the predicted grid points to create a dense flow map. Using `torch.nn.functional.grid_sample`, we \"unwarp\" the image. This is crucial for photos of paper that is curved or folded, ensuring the time-axis is linear.\n\n### Stage 2: 1D Signal Segmentation\n*   **Model:** ResNet34 Encoder + `CoordUnetDecoder`.\n*   **Feature:** We inject Coordinate Convolutions (`CoordConv`) in the decoder to help the model understand the spatial position of the signals (since ECG leads have fixed positions).\n*   **Output:** A probability map (Heatmap) of the signal trace.\n\n---\n\n## 3. The \"Secret Sauce\": Robust Signal Extraction (`fill_pro`)\n\nIn the baseline approach, extracting the time series from the Stage 2 heatmap was done via a simple `argmax`. However, this fails when:\n1.  The ink is faded (missing signal).\n2.  There is noise (salt-and-pepper artifacts).\n3.  The line is jagged.\n\nWe implemented a robust extraction algorithm (`fill_pro.py`) that significantly improved our SNR score.\n\n### A. Gaussian Probability Smoothing\nBefore extracting the coordinates, we apply a Gaussian Blur to the probability map output by the model.\n\n*   **Why:** The raw pixel predictions can be noisy. Blurring aggregates local evidence, making the `argmax` (peak detection) much more stable and continuous.\n*   **Parameters:** We tuned `sigma=0.29` and `threshold=0.11` for the best leaderboard performance.\n\n### B. Handling Signal Gaps (Interpolation vs. Zero-Filling)\nThe baseline approach filled missing data (where model confidence < threshold) with the `zero_mv` (isoelectric line). This causes massive penalties in the SNR metric because a missing peak is treated as a flat line (high error).\n\nWe replaced this with **Linear/Spline Interpolation**.\n\n*   **Impact:** If a QRS complex has a gap in the middle due to poor scanning, interpolation connects the dots, preserving the shape much better than dropping to zero.\n\n### C. Einthoven's Law Correction\nWe apply a physics-based constraint for the limb leads. According to Einthoven's Law, Lead II should equal Lead I + Lead III. We calculate the error between the measured leads and distribute this error to correct the signals.\n\nThis post-processing step ensures the extracted signals are electrically consistent, acting as a powerful regularizer against noise in any single lead.\n\n---\n\n## 4. Summary of Parameters\nFor the final submission, we used the following configuration in `fill_pro`:\n*   **Smoothing:** Gaussian (`sigma=0.29`).\n*   **Interpolation:** Linear (`scipy.interpolate.interp1d`).\n*   **Threshold:** `0.11` (Pixels with probability below this are treated as missing and interpolated).\n*   **Post-Smoothing:** Savitzky-Golay filter (`window=7`, `poly=2`) applied to the final 1D series to remove high-frequency digitization noise."
  }
}