{
  "id": 669553,
  "title": "97th Place Solution",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/669553",
  "author_name": "Evan",
  "post_date": "2026-01-23T02:49:48.502000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to the organizers for this fun competition. This is our writeup for <strong>PhysioNet - Digitization of ECG Images</strong>. We did not train new models. We focused on a clean inference pipeline and a small ensemble that is stable on many kinds of ECG scans.</p>\n<h2>Overview</h2>\n<p>The competition gives an ECG image and asks us to output the ECG signal values in a CSV file. Our approach follows a simple idea: first make the image “clean and aligned”, then extract wave pixels, then convert pixels into signal values. After that, we run the same pipeline three times with small variations and average the results.</p>\n<h2>Pipeline Architecture</h2>\n<p>Our solution uses a 3-stage pipeline and runs it with 3 different input views.</p>\n<pre><code>ECG image\n  -&gt; Stage 0: alignment (rotation / position)\n  -&gt; Stage 1: grid rectification (make the paper grid straight)\n  -&gt; Stage 2: waveform digitization (predict waveform pixels)\n  -&gt; pixel-to-signal conversion (mv_to_pixel, zero_mv)\n  -&gt; optional signal smoothing + simple ECG constraint\n  -&gt; build submission rows (id, value)\n</code></pre>\n<h2>Stage 0 / Stage 1 (Preprocessing)</h2>\n<p>For Stage 0 and Stage 1 we reuse public weights. Stage 0 normalizes the image geometry, and Stage 1 rectifies the ECG grid. This helps Stage 2 see a consistent layout.</p>\n<p>In this notebook we keep all models in eval mode and use AMP (<code>autocast</code>) for faster inference.</p>\n<h2>Stage 2 (Digitization)</h2>\n<p>Stage 2 predicts waveform pixels for the 4 ECG rows. In our code this model is <code>Net3</code>, which uses a <code>resnet34</code> encoder and a UNet-style decoder. We apply <code>sigmoid</code> to the model output and convert it to a waveform using <code>pixel_to_series</code>.</p>\n<p>To convert pixels into mV values, we use the <code>zero_mv</code> reference lines and a pixel-to-mV scale. We use two close scales (78.5 and 78.8) as a small diversity trick.</p>\n<h2>Three Branches (Input diversity)</h2>\n<p>We run the full pipeline three times per image. The goal is to make three reasonable predictions that are not identical.</p>\n<p>Branch A uses a grayscale-based enhancement. We convert the image to grayscale, denoise it, and apply CLAHE to increase contrast. This can make faint traces easier to see.</p>\n<p>Branch B uses the original image without extra preprocessing. This baseline is usually the most stable.</p>\n<p>Branch C uses an HSV-based enhancement. We denoise the V channel and apply CLAHE with an adaptive clip limit. This often helps on scans with uneven lighting.</p>\n<h2>Signal Refinement (Post-processing)</h2>\n<p>We apply light post-processing on some branches.</p>\n<p>First, we sometimes use Savitzky–Golay smoothing (<code>savgol_filter</code>, window 7, poly 2). This reduces small noise but keeps most wave shape.</p>\n<p>Second, we sometimes apply a simple ECG rule (Einthoven’s law). Lead II is close to Lead I plus Lead III. We compute the error and distribute it with <code>alpha = 0.33</code>. The notebook handles length mismatch safely, so this step does not crash.</p>\n<p>We do not force the same post-processing on every branch. Keeping differences between branches is important for the ensemble.</p>\n<h2>Building the Submission</h2>\n<p>The pipeline produces a 4-row signal layout. Each row contains multiple leads placed next to each other, so we split each row into segments and map them back to the standard 12-lead names. We also use the long rhythm strip (Lead II) as the full-length Lead II signal.</p>\n<p>For each lead in <code>test.csv</code>, we make sure the signal length matches <code>number_of_rows</code>. If lengths differ, we resample with <code>np.interp</code>. We then write three intermediate files (<code>submission_a.csv</code>, <code>submission_b.csv</code>, <code>submission_c.csv</code>).</p>\n<h2>Ensemble Strategy</h2>\n<p>We combine the three intermediate submissions with a weighted average (after aligning by <code>id</code>). These are the weights used in the notebook:</p>\n<table>\n<thead>\n<tr>\n<th>Branch</th>\n<th>Preprocess</th>\n<th>Savgol</th>\n<th>Einthoven (dw)</th>\n<th>mv_to_pixel</th>\n<th>Weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>grayscale guidance</td>\n<td>yes</td>\n<td>yes</td>\n<td>78.5</td>\n<td>0.325</td>\n</tr>\n<tr>\n<td>B</td>\n<td>none (baseline)</td>\n<td>no</td>\n<td>no</td>\n<td>78.5</td>\n<td>0.350</td>\n</tr>\n<tr>\n<td>C</td>\n<td>change color (HSV)</td>\n<td>yes</td>\n<td>yes</td>\n<td>78.8</td>\n<td>0.325</td>\n</tr>\n</tbody>\n</table>\n<h2>What Worked</h2>\n<ol>\n<li>Using a clean and stable 3-stage pipeline (align -&gt; rectify -&gt; digitize) was a strong baseline.</li>\n<li>Keeping one branch as “no preprocess” (Branch B) helped a lot on hard scans where heavy enhancement can hurt.</li>\n<li>Using two different preprocessing styles (grayscale guidance and HSV change) gave useful diversity.</li>\n<li>Light refinement (Savgol + Einthoven) helped on some cases, but we only applied it to some branches to keep diversity.</li>\n<li>A simple weighted average was enough to improve stability without making the notebook complex.</li>\n</ol>\n<h2>What Didn’t Work (for us)</h2>\n<p>We did not train models end-to-end in this solution. We stayed close to the public weights and focused on inference.</p>\n<p>We also found that using too much post-processing everywhere can reduce diversity. For example, if all branches are smoothed in the same way, the final average can become less robust on sharp peaks.</p>\n<p>Finally, we did not build a large ensemble of many different models. We aimed for a fast notebook that is easy to run and share.</p>\n<h2>Reproducibility / Environment</h2>\n<p>This solution was designed to run as a Kaggle notebook with GPU (we used a single GPU setup). The notebook removes TensorFlow to avoid package conflicts and uses AMP (<code>autocast</code>) for faster inference. Running the notebook top to bottom will produce <code>submission.csv</code>.</p>\n<h2>Acknowledgements</h2>\n<p>Thanks to the organizers and to the Kaggle community. Our pipeline is built on top of public notebooks and public weights, which made a strong baseline possible.</p>",
  "messages": [
    {
      "id": 3395458,
      "postDate": "2026-01-23T02:49:48.503Z",
      "content": "<p>Thanks to the organizers for this fun competition. This is our writeup for <strong>PhysioNet - Digitization of ECG Images</strong>. We did not train new models. We focused on a clean inference pipeline and a small ensemble that is stable on many kinds of ECG scans.</p>\n<h2>Overview</h2>\n<p>The competition gives an ECG image and asks us to output the ECG signal values in a CSV file. Our approach follows a simple idea: first make the image “clean and aligned”, then extract wave pixels, then convert pixels into signal values. After that, we run the same pipeline three times with small variations and average the results.</p>\n<h2>Pipeline Architecture</h2>\n<p>Our solution uses a 3-stage pipeline and runs it with 3 different input views.</p>\n<pre><code>ECG image\n  -&gt; Stage 0: alignment (rotation / position)\n  -&gt; Stage 1: grid rectification (make the paper grid straight)\n  -&gt; Stage 2: waveform digitization (predict waveform pixels)\n  -&gt; pixel-to-signal conversion (mv_to_pixel, zero_mv)\n  -&gt; optional signal smoothing + simple ECG constraint\n  -&gt; build submission rows (id, value)\n</code></pre>\n<h2>Stage 0 / Stage 1 (Preprocessing)</h2>\n<p>For Stage 0 and Stage 1 we reuse public weights. Stage 0 normalizes the image geometry, and Stage 1 rectifies the ECG grid. This helps Stage 2 see a consistent layout.</p>\n<p>In this notebook we keep all models in eval mode and use AMP (<code>autocast</code>) for faster inference.</p>\n<h2>Stage 2 (Digitization)</h2>\n<p>Stage 2 predicts waveform pixels for the 4 ECG rows. In our code this model is <code>Net3</code>, which uses a <code>resnet34</code> encoder and a UNet-style decoder. We apply <code>sigmoid</code> to the model output and convert it to a waveform using <code>pixel_to_series</code>.</p>\n<p>To convert pixels into mV values, we use the <code>zero_mv</code> reference lines and a pixel-to-mV scale. We use two close scales (78.5 and 78.8) as a small diversity trick.</p>\n<h2>Three Branches (Input diversity)</h2>\n<p>We run the full pipeline three times per image. The goal is to make three reasonable predictions that are not identical.</p>\n<p>Branch A uses a grayscale-based enhancement. We convert the image to grayscale, denoise it, and apply CLAHE to increase contrast. This can make faint traces easier to see.</p>\n<p>Branch B uses the original image without extra preprocessing. This baseline is usually the most stable.</p>\n<p>Branch C uses an HSV-based enhancement. We denoise the V channel and apply CLAHE with an adaptive clip limit. This often helps on scans with uneven lighting.</p>\n<h2>Signal Refinement (Post-processing)</h2>\n<p>We apply light post-processing on some branches.</p>\n<p>First, we sometimes use Savitzky–Golay smoothing (<code>savgol_filter</code>, window 7, poly 2). This reduces small noise but keeps most wave shape.</p>\n<p>Second, we sometimes apply a simple ECG rule (Einthoven’s law). Lead II is close to Lead I plus Lead III. We compute the error and distribute it with <code>alpha = 0.33</code>. The notebook handles length mismatch safely, so this step does not crash.</p>\n<p>We do not force the same post-processing on every branch. Keeping differences between branches is important for the ensemble.</p>\n<h2>Building the Submission</h2>\n<p>The pipeline produces a 4-row signal layout. Each row contains multiple leads placed next to each other, so we split each row into segments and map them back to the standard 12-lead names. We also use the long rhythm strip (Lead II) as the full-length Lead II signal.</p>\n<p>For each lead in <code>test.csv</code>, we make sure the signal length matches <code>number_of_rows</code>. If lengths differ, we resample with <code>np.interp</code>. We then write three intermediate files (<code>submission_a.csv</code>, <code>submission_b.csv</code>, <code>submission_c.csv</code>).</p>\n<h2>Ensemble Strategy</h2>\n<p>We combine the three intermediate submissions with a weighted average (after aligning by <code>id</code>). These are the weights used in the notebook:</p>\n<table>\n<thead>\n<tr>\n<th>Branch</th>\n<th>Preprocess</th>\n<th>Savgol</th>\n<th>Einthoven (dw)</th>\n<th>mv_to_pixel</th>\n<th>Weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>grayscale guidance</td>\n<td>yes</td>\n<td>yes</td>\n<td>78.5</td>\n<td>0.325</td>\n</tr>\n<tr>\n<td>B</td>\n<td>none (baseline)</td>\n<td>no</td>\n<td>no</td>\n<td>78.5</td>\n<td>0.350</td>\n</tr>\n<tr>\n<td>C</td>\n<td>change color (HSV)</td>\n<td>yes</td>\n<td>yes</td>\n<td>78.8</td>\n<td>0.325</td>\n</tr>\n</tbody>\n</table>\n<h2>What Worked</h2>\n<ol>\n<li>Using a clean and stable 3-stage pipeline (align -&gt; rectify -&gt; digitize) was a strong baseline.</li>\n<li>Keeping one branch as “no preprocess” (Branch B) helped a lot on hard scans where heavy enhancement can hurt.</li>\n<li>Using two different preprocessing styles (grayscale guidance and HSV change) gave useful diversity.</li>\n<li>Light refinement (Savgol + Einthoven) helped on some cases, but we only applied it to some branches to keep diversity.</li>\n<li>A simple weighted average was enough to improve stability without making the notebook complex.</li>\n</ol>\n<h2>What Didn’t Work (for us)</h2>\n<p>We did not train models end-to-end in this solution. We stayed close to the public weights and focused on inference.</p>\n<p>We also found that using too much post-processing everywhere can reduce diversity. For example, if all branches are smoothed in the same way, the final average can become less robust on sharp peaks.</p>\n<p>Finally, we did not build a large ensemble of many different models. We aimed for a fast notebook that is easy to run and share.</p>\n<h2>Reproducibility / Environment</h2>\n<p>This solution was designed to run as a Kaggle notebook with GPU (we used a single GPU setup). The notebook removes TensorFlow to avoid package conflicts and uses AMP (<code>autocast</code>) for faster inference. Running the notebook top to bottom will produce <code>submission.csv</code>.</p>\n<h2>Acknowledgements</h2>\n<p>Thanks to the organizers and to the Kaggle community. Our pipeline is built on top of public notebooks and public weights, which made a strong baseline possible.</p>",
      "rawMarkdown": "Thanks to the organizers for this fun competition. This is our writeup for **PhysioNet - Digitization of ECG Images**. We did not train new models. We focused on a clean inference pipeline and a small ensemble that is stable on many kinds of ECG scans.\n\n## Overview\n\nThe competition gives an ECG image and asks us to output the ECG signal values in a CSV file. Our approach follows a simple idea: first make the image “clean and aligned”, then extract wave pixels, then convert pixels into signal values. After that, we run the same pipeline three times with small variations and average the results.\n\n## Pipeline Architecture\n\nOur solution uses a 3-stage pipeline and runs it with 3 different input views.\n\n```text\nECG image\n  -> Stage 0: alignment (rotation / position)\n  -> Stage 1: grid rectification (make the paper grid straight)\n  -> Stage 2: waveform digitization (predict waveform pixels)\n  -> pixel-to-signal conversion (mv_to_pixel, zero_mv)\n  -> optional signal smoothing + simple ECG constraint\n  -> build submission rows (id, value)\n```\n\n## Stage 0 / Stage 1 (Preprocessing)\n\nFor Stage 0 and Stage 1 we reuse public weights. Stage 0 normalizes the image geometry, and Stage 1 rectifies the ECG grid. This helps Stage 2 see a consistent layout.\n\nIn this notebook we keep all models in eval mode and use AMP (`autocast`) for faster inference.\n\n## Stage 2 (Digitization)\n\nStage 2 predicts waveform pixels for the 4 ECG rows. In our code this model is `Net3`, which uses a `resnet34` encoder and a UNet-style decoder. We apply `sigmoid` to the model output and convert it to a waveform using `pixel_to_series`.\n\nTo convert pixels into mV values, we use the `zero_mv` reference lines and a pixel-to-mV scale. We use two close scales (78.5 and 78.8) as a small diversity trick.\n\n## Three Branches (Input diversity)\n\nWe run the full pipeline three times per image. The goal is to make three reasonable predictions that are not identical.\n\nBranch A uses a grayscale-based enhancement. We convert the image to grayscale, denoise it, and apply CLAHE to increase contrast. This can make faint traces easier to see.\n\nBranch B uses the original image without extra preprocessing. This baseline is usually the most stable.\n\nBranch C uses an HSV-based enhancement. We denoise the V channel and apply CLAHE with an adaptive clip limit. This often helps on scans with uneven lighting.\n\n## Signal Refinement (Post-processing)\n\nWe apply light post-processing on some branches.\n\nFirst, we sometimes use Savitzky–Golay smoothing (`savgol_filter`, window 7, poly 2). This reduces small noise but keeps most wave shape.\n\nSecond, we sometimes apply a simple ECG rule (Einthoven’s law). Lead II is close to Lead I plus Lead III. We compute the error and distribute it with `alpha = 0.33`. The notebook handles length mismatch safely, so this step does not crash.\n\nWe do not force the same post-processing on every branch. Keeping differences between branches is important for the ensemble.\n\n## Building the Submission\n\nThe pipeline produces a 4-row signal layout. Each row contains multiple leads placed next to each other, so we split each row into segments and map them back to the standard 12-lead names. We also use the long rhythm strip (Lead II) as the full-length Lead II signal.\n\nFor each lead in `test.csv`, we make sure the signal length matches `number_of_rows`. If lengths differ, we resample with `np.interp`. We then write three intermediate files (`submission_a.csv`, `submission_b.csv`, `submission_c.csv`).\n\n## Ensemble Strategy\n\nWe combine the three intermediate submissions with a weighted average (after aligning by `id`). These are the weights used in the notebook:\n\n| Branch | Preprocess         | Savgol | Einthoven (dw) | mv_to_pixel | Weight |\n| ------ | ------------------ | ------ | -------------- | ----------: | -----: |\n| A      | grayscale guidance | yes    | yes            |        78.5 |  0.325 |\n| B      | none (baseline)    | no     | no             |        78.5 |  0.350 |\n| C      | change color (HSV) | yes    | yes            |        78.8 |  0.325 |\n\n## What Worked\n\n1. Using a clean and stable 3-stage pipeline (align -> rectify -> digitize) was a strong baseline.\n2. Keeping one branch as “no preprocess” (Branch B) helped a lot on hard scans where heavy enhancement can hurt.\n3. Using two different preprocessing styles (grayscale guidance and HSV change) gave useful diversity.\n4. Light refinement (Savgol + Einthoven) helped on some cases, but we only applied it to some branches to keep diversity.\n5. A simple weighted average was enough to improve stability without making the notebook complex.\n\n## What Didn’t Work (for us)\n\nWe did not train models end-to-end in this solution. We stayed close to the public weights and focused on inference.\n\nWe also found that using too much post-processing everywhere can reduce diversity. For example, if all branches are smoothed in the same way, the final average can become less robust on sharp peaks.\n\nFinally, we did not build a large ensemble of many different models. We aimed for a fast notebook that is easy to run and share.\n\n## Reproducibility / Environment\n\nThis solution was designed to run as a Kaggle notebook with GPU (we used a single GPU setup). The notebook removes TensorFlow to avoid package conflicts and uses AMP (`autocast`) for faster inference. Running the notebook top to bottom will produce `submission.csv`.\n\n## Acknowledgements\n\nThanks to the organizers and to the Kaggle community. Our pipeline is built on top of public notebooks and public weights, which made a strong baseline possible.\n",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3395458": "Thanks to the organizers for this fun competition. This is our writeup for **PhysioNet - Digitization of ECG Images**. We did not train new models. We focused on a clean inference pipeline and a small ensemble that is stable on many kinds of ECG scans.\n\n## Overview\n\nThe competition gives an ECG image and asks us to output the ECG signal values in a CSV file. Our approach follows a simple idea: first make the image “clean and aligned”, then extract wave pixels, then convert pixels into signal values. After that, we run the same pipeline three times with small variations and average the results.\n\n## Pipeline Architecture\n\nOur solution uses a 3-stage pipeline and runs it with 3 different input views.\n\n```text\nECG image\n  -> Stage 0: alignment (rotation / position)\n  -> Stage 1: grid rectification (make the paper grid straight)\n  -> Stage 2: waveform digitization (predict waveform pixels)\n  -> pixel-to-signal conversion (mv_to_pixel, zero_mv)\n  -> optional signal smoothing + simple ECG constraint\n  -> build submission rows (id, value)\n```\n\n## Stage 0 / Stage 1 (Preprocessing)\n\nFor Stage 0 and Stage 1 we reuse public weights. Stage 0 normalizes the image geometry, and Stage 1 rectifies the ECG grid. This helps Stage 2 see a consistent layout.\n\nIn this notebook we keep all models in eval mode and use AMP (`autocast`) for faster inference.\n\n## Stage 2 (Digitization)\n\nStage 2 predicts waveform pixels for the 4 ECG rows. In our code this model is `Net3`, which uses a `resnet34` encoder and a UNet-style decoder. We apply `sigmoid` to the model output and convert it to a waveform using `pixel_to_series`.\n\nTo convert pixels into mV values, we use the `zero_mv` reference lines and a pixel-to-mV scale. We use two close scales (78.5 and 78.8) as a small diversity trick.\n\n## Three Branches (Input diversity)\n\nWe run the full pipeline three times per image. The goal is to make three reasonable predictions that are not identical.\n\nBranch A uses a grayscale-based enhancement. We convert the image to grayscale, denoise it, and apply CLAHE to increase contrast. This can make faint traces easier to see.\n\nBranch B uses the original image without extra preprocessing. This baseline is usually the most stable.\n\nBranch C uses an HSV-based enhancement. We denoise the V channel and apply CLAHE with an adaptive clip limit. This often helps on scans with uneven lighting.\n\n## Signal Refinement (Post-processing)\n\nWe apply light post-processing on some branches.\n\nFirst, we sometimes use Savitzky–Golay smoothing (`savgol_filter`, window 7, poly 2). This reduces small noise but keeps most wave shape.\n\nSecond, we sometimes apply a simple ECG rule (Einthoven’s law). Lead II is close to Lead I plus Lead III. We compute the error and distribute it with `alpha = 0.33`. The notebook handles length mismatch safely, so this step does not crash.\n\nWe do not force the same post-processing on every branch. Keeping differences between branches is important for the ensemble.\n\n## Building the Submission\n\nThe pipeline produces a 4-row signal layout. Each row contains multiple leads placed next to each other, so we split each row into segments and map them back to the standard 12-lead names. We also use the long rhythm strip (Lead II) as the full-length Lead II signal.\n\nFor each lead in `test.csv`, we make sure the signal length matches `number_of_rows`. If lengths differ, we resample with `np.interp`. We then write three intermediate files (`submission_a.csv`, `submission_b.csv`, `submission_c.csv`).\n\n## Ensemble Strategy\n\nWe combine the three intermediate submissions with a weighted average (after aligning by `id`). These are the weights used in the notebook:\n\n| Branch | Preprocess         | Savgol | Einthoven (dw) | mv_to_pixel | Weight |\n| ------ | ------------------ | ------ | -------------- | ----------: | -----: |\n| A      | grayscale guidance | yes    | yes            |        78.5 |  0.325 |\n| B      | none (baseline)    | no     | no             |        78.5 |  0.350 |\n| C      | change color (HSV) | yes    | yes            |        78.8 |  0.325 |\n\n## What Worked\n\n1. Using a clean and stable 3-stage pipeline (align -> rectify -> digitize) was a strong baseline.\n2. Keeping one branch as “no preprocess” (Branch B) helped a lot on hard scans where heavy enhancement can hurt.\n3. Using two different preprocessing styles (grayscale guidance and HSV change) gave useful diversity.\n4. Light refinement (Savgol + Einthoven) helped on some cases, but we only applied it to some branches to keep diversity.\n5. A simple weighted average was enough to improve stability without making the notebook complex.\n\n## What Didn’t Work (for us)\n\nWe did not train models end-to-end in this solution. We stayed close to the public weights and focused on inference.\n\nWe also found that using too much post-processing everywhere can reduce diversity. For example, if all branches are smoothed in the same way, the final average can become less robust on sharp peaks.\n\nFinally, we did not build a large ensemble of many different models. We aimed for a fast notebook that is easy to run and share.\n\n## Reproducibility / Environment\n\nThis solution was designed to run as a Kaggle notebook with GPU (we used a single GPU setup). The notebook removes TensorFlow to avoid package conflicts and uses AMP (`autocast`) for faster inference. Running the notebook top to bottom will produce `submission.csv`.\n\n## Acknowledgements\n\nThanks to the organizers and to the Kaggle community. Our pipeline is built on top of public notebooks and public weights, which made a strong baseline possible.\n"
  }
}