{
  "id": 669740,
  "title": "8th Place Solution",
  "url": "/competitions/physionet-ecg-image-digitization/writeups/8th-place-solution",
  "author_name": "",
  "post_date": "2026-01-24T04:32:11.433Z",
  "votes": 18,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>Overview</h1>\n<p>Our best submission works by running three separate solutions, then computing a weighted average of their predictions. We formed our team late in the competition, so our pipelines were developed with relatively few shared assumptions. This independence increased model diversity, thereby allowing our ensemble to significantly outperform the individual scores.</p>\n<table>\n<thead>\n<tr>\n<th>Solution</th>\n<th>Ensemble Weight</th>\n<th>Public LB Score</th>\n<th>Private LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Imanishi's pipeline</td>\n<td>45%</td>\n<td>21.58</td>\n<td>21.43</td>\n</tr>\n<tr>\n<td>James's pipeline</td>\n<td>32%</td>\n<td>21.17</td>\n<td>20.98</td>\n</tr>\n<tr>\n<td>Liu's pipeline</td>\n<td>23%</td>\n<td>20.84</td>\n<td>20.74</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>N/A</td>\n<td>22.16</td>\n<td>22.03</td>\n</tr>\n</tbody>\n</table>\n<p>At a high level, all of our solutions work by rectifying the input images to correct for distortions, then using segmentation models and softmax operations to extract the voltage signals. However, there were some substantial differences in our preprocessing methodology (1-stage vs. 2-stage), model architectures (CNNs vs. Transformers), segmentation loss functions (binary cross entropy vs. custom \"coord\" loss), and signal post-processing methodology (anomaly suppression techniques &amp; signal sharpening transforms). As a result, the inaccuracies in our predictions are largely uncorrelated and weighted averaging improves the signal to noise ratio substantially.</p>\n<p>Our individual solutions are described in greater detail below.</p>\n<h1>Imanishi's Pipeline</h1>\n<h3>Overview of Imanishi's Part</h3>\n<p>Using <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">hengck23’s excellent notebook</a> as a baseline, I introduced the following improvements:</p>\n<ul>\n<li><p>Replaced the stage0 model with my own trained model.\nThe original motivation was to completely exclude pretrained data when computing CV scores, but since it also improved the LB score, I adopted it. (Stage1 ultimately remains hengck23’s model.)</p></li>\n<li><p>When converting to rectified input images for stage2, hengck23 performed the transformation in two steps. To reduce image quality degradation, I instead computed a grid from the outputs of stage0 and stage1 and transformed the original image in a single step using <code>F.grid_sample()</code>.</p></li>\n<li><p>Increased the stage2 input resolution, especially in the x-direction (input size: <strong>1632 × 4480</strong>).</p></li>\n<li><p>Introduced the <strong>Coord loss</strong> idea from my teammate liuzhangzhen into stage2.</p></li>\n<li><p>Added y-direction clustering–based masking in stage2.</p></li>\n<li><p>Instead of filling low-confidence predictions in stage2 with 0 mV, I switched to <strong>linear interpolation</strong>.</p></li>\n<li><p>Adopted <strong>EfficientNetV2B1-UNet</strong> for the stage0 and stage2 models (it is unclear how much this contributed to accuracy).</p></li>\n<li><p>Added the following post-processing:</p>\n<ul>\n<li>Correction of Leads I, II, and III using <strong>Einthoven’s law</strong>.</li>\n<li>Since the first quarter of Lead II overlaps with two waveforms, apply average ensembling.</li>\n<li>Apply <code>savgol_filter</code> with a dynamically adjusted <code>window_length</code> depending on <code>sig_len</code>.</li></ul></li>\n</ul>\n<h3>Details of Stage2 Model</h3>\n<p>In Liu’s implementation of Coord loss, both the segmentation loss and the Coord loss are applied to the same logits. In my case, this did not work well, so I split the model head into two heads (segmentation head, y-coordinate head) and applied <strong>segmentation loss (dice loss + BCE loss)</strong> and <strong>Coord loss</strong> separately.</p>\n<p>Both heads have the same output shape: [height, width, 4ch].</p>\n<p>The segmentation loss acts like an auxiliary loss during training, but the segmentation output is also used during inference for masking low-confidence regions in the x-direction and for the y-clustering described later.</p>\n<h3>Coord Loss</h3>\n<p>Although Liu’s section already explains Coord Loss, I will describe what I consider important.</p>\n<p>With a segmentation-based approach like hengck23’s baseline, where the curve mask is predicted and then <code>argmax</code> is used to obtain y-coordinates, it is difficult to estimate optimal y positions. This becomes clear when looking at the ground truth for the two heads (following figure).</p>\n<p>With Coord Loss, the logits output (<strong>height × width × 4 channels</strong>) are the same as in segmentation head up to a point, but then a <strong>softmax over the y-dimension</strong> is applied to directly predict the y-coordinate.</p>\n<p>If we use a one-hot label with only one pixel as ground truth, the model becomes very sensitive to small coordinate shifts. Instead, the ground truth is represented as a <strong>Gaussian-shaped probability distribution</strong> along the y-axis. The Gaussian parameters (sigma and radius) were kept the same as Liu’s because I did not have time to tune them.</p>\n<p>The loss function is standard <strong>cross-entropy loss</strong>, treating the problem as multi-class classification over y positions.</p>\n<p>By applying <code>argmax</code> along the y-direction on the Coord Loss head output, we can obtain much more accurate y-coordinates. Adding Coord Loss improved my public LB score by about <strong>+1.0 dB</strong>, which is a large gain.</p>\n<p><a href=\"https://postimg.cc/p59nH9h1\" target=\"_blank\"><img src=\"https://i.postimg.cc/Fzpb039m/imanishi-figure1.png\" alt=\"imanishi-figure1.png\"></a></p>\n<h3>Y-Clustering</h3>\n<p>For some reason, my stage2 model occasionally misclassified lead labels, which could significantly degrade the SNR. I wanted to prevent this during training, but in the end I could not, so I handled it with a <strong>post-processing step using y-direction clustering</strong>.</p>\n<p>For y-clustering, I used the output from the segmentation head, and the generated masks were then applied to the y-coordinate head.</p>\n<p><a href=\"https://postimg.cc/JtXg1pNv\" target=\"_blank\"><img src=\"https://i.postimg.cc/L4QM37n9/imanishi-figure2.png\" alt=\"imanishi-figure2.png\"></a></p>\n<h3>Other Tricks</h3>\n<p>In the submission notebook, I split the entire test set into two parts and assigned one T4 GPU to each process, as shown in the code below, in order to speed up inference.</p>\n<p>In practice, my part of the pipeline runs in about <strong>70 minutes</strong>, which was the fastest among the team members.</p>\n<pre><code>env1 = os.environ.copy()\nenv2 = os.environ.copy()\nenv1['CUDA_VISIBLE_DEVICES'] = '0'\nenv2['CUDA_VISIBLE_DEVICES'] = '1'\n\ncmd1 = f'python run.py --half 0'\nproc1 = subprocess.Popen(cmd1.split(' '), env=env1)\n\ncmd2 = f'python run.py --half 1'\nproc2 = subprocess.Popen(cmd2.split(' '), env=env2)\n\n_ = proc1.communicate()\n_ = proc2.communicate()\n</code></pre>\n<h1>James's Pipeline</h1>\n<h3>Overview</h3>\n<p>My inference pipeline has 3 main steps:</p>\n<ol>\n<li><strong>Preprocessing:</strong> Detect landmarks in the images and use them to correct for affine warping (such as camera perspective variations &amp; rotation, but not wrinkles in the paper).</li>\n<li><strong>Lead segmentation:</strong> Predict a heatmap describing where the leads hypothetically would be located if the input image was perfectly \"clean\" and undistorted. This is better at handling \"natural\" distortions than unnatural ones caused by imperfect non-affine warp correction, so attempting to correct for non-affine warping with preprocessing logic similar to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s baseline is harmful in my pipeline.</li>\n<li><strong>Heatmap-to-signal conversion:</strong> I use top-k softmax operations to predict estimated voltages &amp; confidence intervals for each lead at each timestep, replace extremely low confidence predictions with ones interpolated from context, use a top-hat transform to \"sharpen\" the remaining voltage spikes, and use a little linear algebra to exploit redundancy between leads for denoising purposes.</li>\n</ol>\n<p>Steps 1 and 2 both use finetuned versions of <a href=\"https://huggingface.co/timm/vit_base_patch14_reg4_dinov2.lvd142m\" target=\"_blank\">DINOv2-base</a>. I used an ensemble of 6 models in total, 2 in step 1, 4 in step 2. Half of the models process the images in an intentionally flipped orientation as a form of test time data augmentation.</p>\n<h3>Data generation</h3>\n<p>I generated ~175K training examples based on ~22K unique ECG records from the PTB-XL dataset.</p>\n<p>This was done by generating \"clean\" ECG images with <a href=\"https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator\" target=\"_blank\">ECG-image-kit</a>, then intentionally corrupting them with some custom data augmentation code which tries to roughly imitate they types of images that appear in the host's data (cell phone pictures of paper with heavy damage, cell phone pictures of computer screens, black and white scans, etc). </p>\n<p>Each image type was roughly simulated using a combination of the following (in no particular order):</p>\n<ul>\n<li><strong>Affine warping</strong></li>\n<li><strong>Background image insertion:</strong> For simulated cellphone pics, the warped ECG images were overlaid on top of images from <a href=\"https://www.kaggle.com/datasets/ntsv648/unsplash-25k\" target=\"_blank\">unsplash-25k</a>. For scanner pics, I used a white background.</li>\n<li><strong>Stain insertion:</strong> This works by overlaying translucent mold and stain images on top of the ECG images. The stains were randomly drawn from a pool of 26 that I generated semi-manually by prompting a diffusion model. In hindsight, I think the stain intensity distribution I used was a bit unrealistic (typically too transparent) and wonder if maybe pwelin noise would work better than the stain images I generated, but never got around to tinkering with a second version of this.</li>\n<li><strong>Wrinkle shading:</strong> This simulates the shadows from hypothetical wrinkles without introducing any non-affine warping. It is very similar to functionality from ECG-image-kit, I just had ChatGPT port it into my script.</li>\n<li><strong>Black-and-white scanner simulation:</strong> This ain't an off the shelf greyscale conversion. I tried to simulate the behavior of black and white scanners by randomizing brightness, contrast, gamma, and the way the input color channels are weighted. It also injects a little gaussian noise to mimic sensor noise &amp; quantization artifacts.</li>\n</ul>\n<p>Additional augmentations were performed on-the-fly in my training scripts, those are just the ones I applied before training.</p>\n<p>Of the ~175K extra training examples, only ~50K were used to train my final models. Primarily because (1) the accuracy of the keypoint detectors didn't seem to meaningfully improve beyond ~17K training examples and (2) the gains from pretraining the lead segmentation models on synthetic data before finetuning on the host images was pretty small relative to the amount of GPU time it was consuming (a little over two RTX 5090 days of extra compute for a +0.17 dB gain was a bit disappointing… I wanted to try other things instead of having my hardware tied up pushing further in that direction).</p>\n<h3>Camera perspective correction</h3>\n<p>This works by detecting keypoints in the images, then using those keypoint locations to compute a homography matrix that can be used to correct for differences in camera angle, zoom, and rotation. It is fairly similar to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s stage 0, with the main differences being that I used a vision transformer instead of a CNN and training data that I generated myself.</p>\n<p><img src=\"https://i.postimg.cc/xdh0Y7yV/raw-rectified-comparison.png\" alt=\"Figure 1: Sample ECG before and after perspective correction\"></p>\n<p>Keypoints were detected at a resolution of 1036x1036, then used to directly rectify images from their native resolutions --&gt; 1694x2198.</p>\n<p>Training details:</p>\n<ul>\n<li>Model architecture: DINOv2-base backbone with linear prediction head that produces 30 channel outputs (1 per target landmark, all lead label text + the start and end of each signal's x axis).</li>\n<li>Trained to imitate ground truth heatmaps with 1 \"gaussian blob\" per keypoint. These blobs have a value of 1 in the center and decay towards zero as distance from the center increases.</li>\n<li>BCE loss</li>\n<li>AdamW optimizer</li>\n<li>One-cycle learning rate schedule</li>\n<li>Data augmentations (applied using <code>albumentations</code>):<ul>\n<li><code>RandomRotate90</code></li>\n<li><code>GridDistortion</code></li>\n<li><code>ElasticTransform</code></li>\n<li><code>OpticalDistortion</code></li>\n<li><code>GaussNoise</code></li>\n<li><code>GaussianBlur</code></li>\n<li><code>RandomBrightnessContrast</code></li>\n<li><code>ColorJitter</code></li>\n<li><code>CoarseDropout</code></li></ul></li>\n</ul>\n<p>Both of the keypoint detection models in my final ensemble were trained on 17.4K of my synthetic training examples (and none of the host data). They primarily differed in the data agmentation settings &amp; training epoch count. One was trained with moderately heavy agumentation &amp; 24 epochs, the other was trained with heavier augmentation &amp; 48 epochs.</p>\n<p>Tripling the training data to ~50K examples improved cross validation when testing against other synthetic images, but did not improve the scores of my full pipeline when testing against the ECG images provided by the host, so only ~10% of the available synthetic data was used to train the keypoint detectors in my final ensemble.</p>\n<h3>Lead segmentation</h3>\n<p>I used the ground-truth signal data to generate \"perfect\" heatmaps describing where the leads ought to be located in the images if they were completely undistored, then finetuned DINOv2-base to predict those heatmaps based on images with realistic distortions. This teaches it to automatically correct for any warping which makes it past my preprocessing.</p>\n<p><img src=\"https://i.postimg.cc/Yqzqgdks/heatmap-pred.png\" alt=\"Predicted heatmap overlaid on original image\"></p>\n<p>I found it very beneficial to use horizontal resolutions higher than the native resolution of the ECG plots, so this uses an input resolution of 1694x4396 (~2x wider than native) and I inserted a <code>ConvTranspose2d</code> layer between the transformer backbone and the prediction head, which increased resolution another 2x shortly before producing the outputs. As a result, the heatmaps have a resolution of <strong>1694x8792</strong>. This provided <strong>MASSIVE score improvements in comparison to just using the naive resolution, roughly a 4.1 dB gain</strong> in early experiments. Using an input resolution 2x higher than native and output resolution 4x higher than native appeared to be ~optimal, adjusting either of those figures by a factor of 2 makes the score worse.</p>\n<p>Training took place in 2 stages:</p>\n<ol>\n<li>Pretraining on my synthetic images</li>\n<li>Finetuning on the host images</li>\n</ol>\n<p>Training details:</p>\n<ul>\n<li><strong>Data:</strong> Stage 1 used 17.4K synthetic images per model with different images used for each model in the ensemble. Stage 2 used 80% of the host data, with 20% held in reserve for cross-validation.</li>\n<li><strong>Mask generation:</strong> Unlike the keypoint detection models, I found it beneficial for the ground-truth segmentation masks used to train these models to be very sharp. The ground-truth lead lines are only a single pixel thick with no blur.</li>\n<li><strong>Activation checkpointing:</strong> Training vision transformers at high resolutions uses a lot of memory. I used <a href=\"https://pytorch.org/blog/activation-checkpointing-techniques/\" target=\"_blank\">activation checkpointing</a> to mitigate this. It allows for a configurable tradeoff between speed and memory usage. I found the speed drawbacks to be extremely minor. It can cut memory usage in half with almost zero slowdown.</li>\n<li><strong>Data augmentation:</strong> Aggressive data augmentation seems to do more harm than good for these models, so I used relatively light configs. For most models in the ensemble, I just used <code>A.GridDistortion(num_steps=5, distort_limit=0.2, p=0.15)</code> during pretraining and didn't have any <em>explicit</em> data augmentation during finetuning. However, finetuning intentionally used rectified images from an older, less accurate, version of my preprocessing pipeline, so models are exposed to more rectification errors during training than they are at test time. Using more accurate rectification for the training data makes my scores worse, so I believe rectification inaccuracies act as a form of sneaky data augmentation during finetuning.</li>\n<li><strong>Loss functions:</strong> 3 of the 4 lead segmentation models in my ensemble were just trained to minimize binary cross entropy loss. One of them used a hybrid loss during finetuning in which it also tries to minimize the mean squared error of the predicted y pixel coordinates at each timestep. That seemed to be <em>slightly</em> beneficial (0.05 dB in cross validation, even less on the leaderboard), but I didn't have time to propagate it to all models in the ensemble. Applying L1 or L2 losses like that is something which consistently did more harm than good to me earlier in the competition, so I initially abandoned it, but it seemed somewhat beneficial after the pretraining was added; I think pretraining purely with BCE before adding MSE or MAE helps to prevent much of the overfitting &amp; instability I encountered earlier.</li>\n<li><strong>Misc:</strong> AdamW optimizer &amp; one cycle learning rate scheduler, similar to the keypoint detector.</li>\n</ul>\n<h3>Test time augmentation &amp; ensembling</h3>\n<p>My pipeline rectifies each image twice using separate keypoint detection models, then flips one of the resulting images and feeds them to a collection of 4 lead segmentation models, half of which were trained to process images in the flipped orientation. The resulting heatmaps were then blended by averaging the pixel logits.</p>\n<p><img src=\"https://i.postimg.cc/yNbNxWKp/ECG-ensembling.png\" alt=\"TTA approach\"></p>\n<p>The approach above is based on the following observations:</p>\n<ol>\n<li>Averaging predictions from models in the standard &amp; flipped orientations provides a gain of roughly 0.28 dB.</li>\n<li>Using 2 lead segmentation models per orientation provides a gain of roughly 0.18 dB.</li>\n<li>Using 2 models for rectification (instead of 1) provides a gain of roughly 0.13 dB.</li>\n<li>Averaging the pixel logits scores ~0.02 dB better than averaging after signal extraction.</li>\n<li>Flipping or rotating the images before rectification does not help.</li>\n</ol>\n<p>The score could likely be improved further by processing the images in more orientations, applying some of the test time augmentations unrelated to rotation &amp; flipping from other top solutions &amp; public notebooks, and figuring out why applying TTA before rectification didn't help (maybe I had a bug?), but I didn't have time to experiment with this super extensively.</p>\n<h3>Signal extraction &amp; post processing</h3>\n<p>I convert the raw predicted lead location heatmaps to signals via the following steps:</p>\n<ol>\n<li><strong>Heatmap --&gt; raw signal:</strong> Within each column of the image where a lead is expected to be located (based on the canonical layout), I use a top-k softmax to compute the estimated probability of the lead being located at each vertical pixel coordinate, then use those probability estimates to compute the expected signal value. Using a top-k softmax with k=10 scored ~0.28 dB better than using a hard argmax. I tried k ∈ {3, 5, 10, 20, unlimited} and found 5 &amp; 10 to be roughly tied for \"best\". The optimal choice varies by model depending on minor differences in other hyperparameters.</li>\n<li><strong>Unconfident prediction replacement:</strong> The top-k softmax operations are also used to compute upper and lower confidence bounds for the signal values at each timestep. If those bounds are more than 0.4 mV apart, the sample is dropped and the gaps are filled in by linearly interpolating from the surrounding context samples. The range covered by the confidence interval corresponds to either the top-10 most likely vertical pixel coordinates or the min and max pixel coordinates with associated probability estimates above 0.1%, whichever is narrower for each timestep. This provided a gain of roughly 0.13 dB in comparison to a baseline without confidence filtering.</li>\n<li><strong>Edge artifact suppression:</strong> I remove edge artifacts by replacing the rightmost 2 voltage samples for each lead with the one located 3rd from the right. The rightmost samples tend to be inaccurate because the \"ground truth\" training signals contain some voltage spikes that are not visible in the printed images, which causes the model to be prone to hallucinating at the end. Partially suppressing those errors gave me a ~0.05 dB gain (some of them leak through this filtering, a 2px safety margin is not wide enough to fully eliminate them).</li>\n<li><strong>Signal sharpening:</strong> The peaks of the voltage spikes tend to be a bit \"rounded\" due to the input images having lower resolution than the raw signals they're based on, so a <a href=\"https://en.wikipedia.org/wiki/Top-hat_transform\" target=\"_blank\">top-hat transform</a> is used to sharpen them. This provided a gain of roughly 0.04 dB. It worked better for me than sharpening with a shock filter (+0.01 dB gain in CV, did not test on LB) or 1D unsharp masking (harmful in CV, did not test on LB).</li>\n<li><strong>I/II/III redundancy exploitation (einthoven's law):</strong> This takes place in several stages. First, the I and III leads are aligned with II to correct for small timing discrepancies. Then the II = I + III relation (einthoven's law) is used to compute \"expected\" signal values for each of those 3 leads based on the other two. Finally, the raw extracted signals are blended with their expected values via weighted averaging with two thirds of the weight assigned to the raw values. This provides a gain of roughly 0.05 dB. It was critically important to align the signals first, otherwise this is harmful for me. I also tried exploiting the aVR + aVL + aVF = 0 relation in a similar manner, and observed a similar gain in local cross validation from doing so, but unlike I/II/III the aV* post processing did not work well on the leaderboard, so my final pipeline only does einthoven post-processing for I/II/III.</li>\n<li><strong>II &amp; rhythm strip redundancy exploitation:</strong> The II signal appears twice in each image, once as a 2.5 second segment, then again as a 10 second segment (rhythm strip) whose prefix should match the shorter segment. My II predictions are generated by aligning the shorter signal with the longer one (correcting for small timing discrepancies), then computing a weighted average with ~56% of the weight given to the long signal, ~44% given to the shorter one. This provides a gain of roughly 0.01 dB… barely measurable, but was consistent for both CV &amp; LB.</li>\n<li><strong>Sample rate adjustment:</strong> Signals are re-sampled to match the host's desired sample rates via linear interpolation. I also tried cubic, akima spline, and PCHIP interpolation, but linear worked best.</li>\n</ol>\n<h3>Things that didn't work well for me</h3>\n<ul>\n<li>Correcting for non-affine image warping during preprocessing</li>\n<li>DeepLabV3 segmentation models</li>\n<li>Larger DINO v2 models</li>\n<li>DINO v3</li>\n<li>Post-processing with 1D CNNs</li>\n<li>Using differentiable warping operations so that the preprocessing model(s) can be trained end-to-end with the final lead segmentation model</li>\n</ul>\n<p>… as usual, many of the things that don't work well for me could potentially work well with additional effort and GPU time for tuning. I frequently move on to testing other ideas when early results don't seem promising. Many of the \"bad\" ideas above wound up working well for other top competitors 😅</p>\n<h1>Liu's Pipeline</h1>\n<h3>Acknowledgments</h3>\n<p>I would like to express my sincere gratitude to Kaggle and the competition organizers for providing this invaluable opportunity. Special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing his strong baseline, which served as a crucial foundation for my work.</p>\n<h3>Summary</h3>\n<p>My solution optimizes Stage 2 of <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>’s baseline. The key insight is to treat waveform extraction as <strong>per-column coordinate regression</strong>, rather than relying purely on a “segmentation → post-processing” pipeline.</p>\n<p>I keep the U-Net–style heatmap prediction for stability, but add <strong>CoordLoss</strong>: a GT-centered <strong>local Gaussian cross-entropy</strong> that directly supervises the centerline y-coordinate per column. On my local validation split, this single change improved SNR by <strong>more than +1 dB</strong>, and it was the dominant contributor to overall quality.</p>\n<p>To further reduce train–inference mismatch, validation/inference uses the same <strong>old-subpixel-compatible y extractor</strong> (NaN + interpolation behavior). I also add lightweight “safe” regularizers enforcing physically plausible lead relationships.</p>\n<h3>Overall Pipeline</h3>\n<p>Stage 0/1 (baseline): detect and rectify ECG sheets into canonical coordinates (rectified strips).</p>\n<p>Stage 2 (this work):</p>\n<ul>\n<li><p>Model: ResNet34 encoder + U-Net decoder → <strong>4 heatmaps</strong> (3 short rows + long Lead II).</p></li>\n<li><p>Losses:</p>\n<ul>\n<li>Masked BCE (heatmap supervision)</li>\n<li><strong>CoordLoss (main gain)</strong></li>\n<li>Small auxiliary losses (consistency + lead relationships)</li></ul></li>\n<li><p>Inference:</p>\n<ul>\n<li>Predict full-width logits</li>\n<li>Logits → y(px) using the same old-subpixel extraction as validation (NaN → interpolation)</li>\n<li>Convert y(px) → mV</li>\n<li>Apply Einthoven-based patching for long Lead II near the boundary</li></ul></li>\n</ul>\n<h3>Stage 2 Model</h3>\n<p>Architecture</p>\n<ul>\n<li>Encoder: <code>timm resnet34.a3_in1k</code></li>\n<li>Decoder: U-Net style decoder (MyCoordUnetDecoder) with multi-scale skip connections</li>\n<li>Head: 1×1 conv → 4 heatmaps</li>\n</ul>\n<p>BatchNorm freezing\nAll BN layers are forced into eval mode during training, which stabilizes optimization under batch size 1 with gradient accumulation.</p>\n<p>Losses</p>\n<ol>\n<li><p>Masked Pixel BCE (baseline supervision)\nI rasterize GT polylines into thin heatmap masks and apply BCEWithLogitsLoss, masked by valid columns. Since the final metric is strongly influenced by the rhythm strip, I upweight <strong>Long Lead II</strong> (via channel weighting and/or multiplicity depending on the run).</p></li>\n<li><p>CoordLoss (main gain, +1 dB)\nMotivation: segmentation-style BCE does not directly optimize the quantity we ultimately need—<strong>the centerline y-coordinate per column</strong>. Even small vertical errors can degrade SNR after alignment and interpolation.</p></li>\n</ol>\n<p>Method: for each channel c and valid column x</p>\n<ul>\n<li><p>Treat logits over height as a categorical distribution:</p>\n<ul>\n<li>( p(y\\mid x,c)=\\mathrm{softmax}(z_c[:,x]/T) )</li></ul></li>\n<li><p>Build a GT-centered local Gaussian target around the continuous ground-truth (y^*(x,c)):</p>\n<ul>\n<li>( w_k \\propto \\exp(-(k-y^*)^2/(2\\sigma^2)) ) within a window ±R</li></ul></li>\n<li><p>Compute cross-entropy using log-softmax values in that local window (local Gaussian CE)</p></li>\n</ul>\n<p>This provides dense, well-shaped gradients for precise localization, even when the line is thin or partially missing. In ablations, CoordLoss produced the largest improvement and drove the overall gain.</p>\n<ol>\n<li>Auxiliary “safe” regularizers (small weights)</li>\n</ol>\n<ul>\n<li>Short II vs Long II dy consistency: applied only on segment0 (Short II time), using SmoothL1 on centered dy to reduce drift/slope mismatch.</li>\n<li>Einthoven bridge (mV domain): encourage Long II near the boundary to match an Einthoven-corrected (II_{\\text{corr}}=(1-w),II+w,(I+III)), with warmup/ramp for stability.</li>\n<li>Lead relationship penalty (hinge): enforce (II \\approx I + III) on segment0 using a tolerance-based squared hinge.</li>\n</ul>\n<p>These are intentionally lightweight; they mainly help avoid implausible outputs and reduce boundary artifacts, while CoordLoss provides the primary accuracy boost.</p>\n<h1>Ensembling Approach</h1>\n<p>We ensemble our predictions by:</p>\n<ol>\n<li>Aligning the signals to correct for minor time discrepancies that would otherwise act as a blur filter if the signals were naively averaged together. This is done using code copied from the evaluation metric, very similar to how it aligns signals with the ground truth.</li>\n<li>Computing a weighted average of the predictions. We only did 2 submissions to try tuning the weights, so they're primarily based on rough guesswork, i.e. the assumption that the optimal weight for each pipeline's predictions would probably be correlated with its public LB score. Using weights which lean more heavily in that direction scored better than using relatively even weights (on both the public &amp; private leaderboard), so that assumption appears to have been correct.</li>\n</ol>\n<p>This provided a gain of 0.6 dB over our best individual submission 😁.</p>",
  "messages": [
    {
      "id": "3395990",
      "postDate": "01/24/2026 04:31:00",
      "content": "<h1>Overview</h1>\n<p>Our best submission works by running three separate solutions, then computing a weighted average of their predictions. We formed our team late in the competition, so our pipelines were developed with relatively few shared assumptions. This independence increased model diversity, thereby allowing our ensemble to significantly outperform the individual scores.</p>\n<table>\n<thead>\n<tr>\n<th>Solution</th>\n<th>Ensemble Weight</th>\n<th>Public LB Score</th>\n<th>Private LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Imanishi's pipeline</td>\n<td>45%</td>\n<td>21.58</td>\n<td>21.43</td>\n</tr>\n<tr>\n<td>James's pipeline</td>\n<td>32%</td>\n<td>21.17</td>\n<td>20.98</td>\n</tr>\n<tr>\n<td>Liu's pipeline</td>\n<td>23%</td>\n<td>20.84</td>\n<td>20.74</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>N/A</td>\n<td>22.16</td>\n<td>22.03</td>\n</tr>\n</tbody>\n</table>\n<p>At a high level, all of our solutions work by rectifying the input images to correct for distortions, then using segmentation models and softmax operations to extract the voltage signals. However, there were some substantial differences in our preprocessing methodology (1-stage vs. 2-stage), model architectures (CNNs vs. Transformers), segmentation loss functions (binary cross entropy vs. custom \"coord\" loss), and signal post-processing methodology (anomaly suppression techniques &amp; signal sharpening transforms). As a result, the inaccuracies in our predictions are largely uncorrelated and weighted averaging improves the signal to noise ratio substantially.</p>\n<p>Our individual solutions are described in greater detail below.</p>\n<h1>Imanishi's Pipeline</h1>\n<h3>Overview of Imanishi's Part</h3>\n<p>Using <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">hengck23’s excellent notebook</a> as a baseline, I introduced the following improvements:</p>\n<ul>\n<li><p>Replaced the stage0 model with my own trained model.\nThe original motivation was to completely exclude pretrained data when computing CV scores, but since it also improved the LB score, I adopted it. (Stage1 ultimately remains hengck23’s model.)</p></li>\n<li><p>When converting to rectified input images for stage2, hengck23 performed the transformation in two steps. To reduce image quality degradation, I instead computed a grid from the outputs of stage0 and stage1 and transformed the original image in a single step using <code>F.grid_sample()</code>.</p></li>\n<li><p>Increased the stage2 input resolution, especially in the x-direction (input size: <strong>1632 × 4480</strong>).</p></li>\n<li><p>Introduced the <strong>Coord loss</strong> idea from my teammate liuzhangzhen into stage2.</p></li>\n<li><p>Added y-direction clustering–based masking in stage2.</p></li>\n<li><p>Instead of filling low-confidence predictions in stage2 with 0 mV, I switched to <strong>linear interpolation</strong>.</p></li>\n<li><p>Adopted <strong>EfficientNetV2B1-UNet</strong> for the stage0 and stage2 models (it is unclear how much this contributed to accuracy).</p></li>\n<li><p>Added the following post-processing:</p>\n<ul>\n<li>Correction of Leads I, II, and III using <strong>Einthoven’s law</strong>.</li>\n<li>Since the first quarter of Lead II overlaps with two waveforms, apply average ensembling.</li>\n<li>Apply <code>savgol_filter</code> with a dynamically adjusted <code>window_length</code> depending on <code>sig_len</code>.</li></ul></li>\n</ul>\n<h3>Details of Stage2 Model</h3>\n<p>In Liu’s implementation of Coord loss, both the segmentation loss and the Coord loss are applied to the same logits. In my case, this did not work well, so I split the model head into two heads (segmentation head, y-coordinate head) and applied <strong>segmentation loss (dice loss + BCE loss)</strong> and <strong>Coord loss</strong> separately.</p>\n<p>Both heads have the same output shape: [height, width, 4ch].</p>\n<p>The segmentation loss acts like an auxiliary loss during training, but the segmentation output is also used during inference for masking low-confidence regions in the x-direction and for the y-clustering described later.</p>\n<h3>Coord Loss</h3>\n<p>Although Liu’s section already explains Coord Loss, I will describe what I consider important.</p>\n<p>With a segmentation-based approach like hengck23’s baseline, where the curve mask is predicted and then <code>argmax</code> is used to obtain y-coordinates, it is difficult to estimate optimal y positions. This becomes clear when looking at the ground truth for the two heads (following figure).</p>\n<p>With Coord Loss, the logits output (<strong>height × width × 4 channels</strong>) are the same as in segmentation head up to a point, but then a <strong>softmax over the y-dimension</strong> is applied to directly predict the y-coordinate.</p>\n<p>If we use a one-hot label with only one pixel as ground truth, the model becomes very sensitive to small coordinate shifts. Instead, the ground truth is represented as a <strong>Gaussian-shaped probability distribution</strong> along the y-axis. The Gaussian parameters (sigma and radius) were kept the same as Liu’s because I did not have time to tune them.</p>\n<p>The loss function is standard <strong>cross-entropy loss</strong>, treating the problem as multi-class classification over y positions.</p>\n<p>By applying <code>argmax</code> along the y-direction on the Coord Loss head output, we can obtain much more accurate y-coordinates. Adding Coord Loss improved my public LB score by about <strong>+1.0 dB</strong>, which is a large gain.</p>\n<p><a href=\"https://postimg.cc/p59nH9h1\" target=\"_blank\"><img src=\"https://i.postimg.cc/Fzpb039m/imanishi-figure1.png\" alt=\"imanishi-figure1.png\"></a></p>\n<h3>Y-Clustering</h3>\n<p>For some reason, my stage2 model occasionally misclassified lead labels, which could significantly degrade the SNR. I wanted to prevent this during training, but in the end I could not, so I handled it with a <strong>post-processing step using y-direction clustering</strong>.</p>\n<p>For y-clustering, I used the output from the segmentation head, and the generated masks were then applied to the y-coordinate head.</p>\n<p><a href=\"https://postimg.cc/JtXg1pNv\" target=\"_blank\"><img src=\"https://i.postimg.cc/L4QM37n9/imanishi-figure2.png\" alt=\"imanishi-figure2.png\"></a></p>\n<h3>Other Tricks</h3>\n<p>In the submission notebook, I split the entire test set into two parts and assigned one T4 GPU to each process, as shown in the code below, in order to speed up inference.</p>\n<p>In practice, my part of the pipeline runs in about <strong>70 minutes</strong>, which was the fastest among the team members.</p>\n<pre><code>env1 = os.environ.copy()\nenv2 = os.environ.copy()\nenv1['CUDA_VISIBLE_DEVICES'] = '0'\nenv2['CUDA_VISIBLE_DEVICES'] = '1'\n\ncmd1 = f'python run.py --half 0'\nproc1 = subprocess.Popen(cmd1.split(' '), env=env1)\n\ncmd2 = f'python run.py --half 1'\nproc2 = subprocess.Popen(cmd2.split(' '), env=env2)\n\n_ = proc1.communicate()\n_ = proc2.communicate()\n</code></pre>\n<h1>James's Pipeline</h1>\n<h3>Overview</h3>\n<p>My inference pipeline has 3 main steps:</p>\n<ol>\n<li><strong>Preprocessing:</strong> Detect landmarks in the images and use them to correct for affine warping (such as camera perspective variations &amp; rotation, but not wrinkles in the paper).</li>\n<li><strong>Lead segmentation:</strong> Predict a heatmap describing where the leads hypothetically would be located if the input image was perfectly \"clean\" and undistorted. This is better at handling \"natural\" distortions than unnatural ones caused by imperfect non-affine warp correction, so attempting to correct for non-affine warping with preprocessing logic similar to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s baseline is harmful in my pipeline.</li>\n<li><strong>Heatmap-to-signal conversion:</strong> I use top-k softmax operations to predict estimated voltages &amp; confidence intervals for each lead at each timestep, replace extremely low confidence predictions with ones interpolated from context, use a top-hat transform to \"sharpen\" the remaining voltage spikes, and use a little linear algebra to exploit redundancy between leads for denoising purposes.</li>\n</ol>\n<p>Steps 1 and 2 both use finetuned versions of <a href=\"https://huggingface.co/timm/vit_base_patch14_reg4_dinov2.lvd142m\" target=\"_blank\">DINOv2-base</a>. I used an ensemble of 6 models in total, 2 in step 1, 4 in step 2. Half of the models process the images in an intentionally flipped orientation as a form of test time data augmentation.</p>\n<h3>Data generation</h3>\n<p>I generated ~175K training examples based on ~22K unique ECG records from the PTB-XL dataset.</p>\n<p>This was done by generating \"clean\" ECG images with <a href=\"https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator\" target=\"_blank\">ECG-image-kit</a>, then intentionally corrupting them with some custom data augmentation code which tries to roughly imitate they types of images that appear in the host's data (cell phone pictures of paper with heavy damage, cell phone pictures of computer screens, black and white scans, etc). </p>\n<p>Each image type was roughly simulated using a combination of the following (in no particular order):</p>\n<ul>\n<li><strong>Affine warping</strong></li>\n<li><strong>Background image insertion:</strong> For simulated cellphone pics, the warped ECG images were overlaid on top of images from <a href=\"https://www.kaggle.com/datasets/ntsv648/unsplash-25k\" target=\"_blank\">unsplash-25k</a>. For scanner pics, I used a white background.</li>\n<li><strong>Stain insertion:</strong> This works by overlaying translucent mold and stain images on top of the ECG images. The stains were randomly drawn from a pool of 26 that I generated semi-manually by prompting a diffusion model. In hindsight, I think the stain intensity distribution I used was a bit unrealistic (typically too transparent) and wonder if maybe pwelin noise would work better than the stain images I generated, but never got around to tinkering with a second version of this.</li>\n<li><strong>Wrinkle shading:</strong> This simulates the shadows from hypothetical wrinkles without introducing any non-affine warping. It is very similar to functionality from ECG-image-kit, I just had ChatGPT port it into my script.</li>\n<li><strong>Black-and-white scanner simulation:</strong> This ain't an off the shelf greyscale conversion. I tried to simulate the behavior of black and white scanners by randomizing brightness, contrast, gamma, and the way the input color channels are weighted. It also injects a little gaussian noise to mimic sensor noise &amp; quantization artifacts.</li>\n</ul>\n<p>Additional augmentations were performed on-the-fly in my training scripts, those are just the ones I applied before training.</p>\n<p>Of the ~175K extra training examples, only ~50K were used to train my final models. Primarily because (1) the accuracy of the keypoint detectors didn't seem to meaningfully improve beyond ~17K training examples and (2) the gains from pretraining the lead segmentation models on synthetic data before finetuning on the host images was pretty small relative to the amount of GPU time it was consuming (a little over two RTX 5090 days of extra compute for a +0.17 dB gain was a bit disappointing… I wanted to try other things instead of having my hardware tied up pushing further in that direction).</p>\n<h3>Camera perspective correction</h3>\n<p>This works by detecting keypoints in the images, then using those keypoint locations to compute a homography matrix that can be used to correct for differences in camera angle, zoom, and rotation. It is fairly similar to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s stage 0, with the main differences being that I used a vision transformer instead of a CNN and training data that I generated myself.</p>\n<p><img src=\"https://i.postimg.cc/xdh0Y7yV/raw-rectified-comparison.png\" alt=\"Figure 1: Sample ECG before and after perspective correction\"></p>\n<p>Keypoints were detected at a resolution of 1036x1036, then used to directly rectify images from their native resolutions --&gt; 1694x2198.</p>\n<p>Training details:</p>\n<ul>\n<li>Model architecture: DINOv2-base backbone with linear prediction head that produces 30 channel outputs (1 per target landmark, all lead label text + the start and end of each signal's x axis).</li>\n<li>Trained to imitate ground truth heatmaps with 1 \"gaussian blob\" per keypoint. These blobs have a value of 1 in the center and decay towards zero as distance from the center increases.</li>\n<li>BCE loss</li>\n<li>AdamW optimizer</li>\n<li>One-cycle learning rate schedule</li>\n<li>Data augmentations (applied using <code>albumentations</code>):<ul>\n<li><code>RandomRotate90</code></li>\n<li><code>GridDistortion</code></li>\n<li><code>ElasticTransform</code></li>\n<li><code>OpticalDistortion</code></li>\n<li><code>GaussNoise</code></li>\n<li><code>GaussianBlur</code></li>\n<li><code>RandomBrightnessContrast</code></li>\n<li><code>ColorJitter</code></li>\n<li><code>CoarseDropout</code></li></ul></li>\n</ul>\n<p>Both of the keypoint detection models in my final ensemble were trained on 17.4K of my synthetic training examples (and none of the host data). They primarily differed in the data agmentation settings &amp; training epoch count. One was trained with moderately heavy agumentation &amp; 24 epochs, the other was trained with heavier augmentation &amp; 48 epochs.</p>\n<p>Tripling the training data to ~50K examples improved cross validation when testing against other synthetic images, but did not improve the scores of my full pipeline when testing against the ECG images provided by the host, so only ~10% of the available synthetic data was used to train the keypoint detectors in my final ensemble.</p>\n<h3>Lead segmentation</h3>\n<p>I used the ground-truth signal data to generate \"perfect\" heatmaps describing where the leads ought to be located in the images if they were completely undistored, then finetuned DINOv2-base to predict those heatmaps based on images with realistic distortions. This teaches it to automatically correct for any warping which makes it past my preprocessing.</p>\n<p><img src=\"https://i.postimg.cc/Yqzqgdks/heatmap-pred.png\" alt=\"Predicted heatmap overlaid on original image\"></p>\n<p>I found it very beneficial to use horizontal resolutions higher than the native resolution of the ECG plots, so this uses an input resolution of 1694x4396 (~2x wider than native) and I inserted a <code>ConvTranspose2d</code> layer between the transformer backbone and the prediction head, which increased resolution another 2x shortly before producing the outputs. As a result, the heatmaps have a resolution of <strong>1694x8792</strong>. This provided <strong>MASSIVE score improvements in comparison to just using the naive resolution, roughly a 4.1 dB gain</strong> in early experiments. Using an input resolution 2x higher than native and output resolution 4x higher than native appeared to be ~optimal, adjusting either of those figures by a factor of 2 makes the score worse.</p>\n<p>Training took place in 2 stages:</p>\n<ol>\n<li>Pretraining on my synthetic images</li>\n<li>Finetuning on the host images</li>\n</ol>\n<p>Training details:</p>\n<ul>\n<li><strong>Data:</strong> Stage 1 used 17.4K synthetic images per model with different images used for each model in the ensemble. Stage 2 used 80% of the host data, with 20% held in reserve for cross-validation.</li>\n<li><strong>Mask generation:</strong> Unlike the keypoint detection models, I found it beneficial for the ground-truth segmentation masks used to train these models to be very sharp. The ground-truth lead lines are only a single pixel thick with no blur.</li>\n<li><strong>Activation checkpointing:</strong> Training vision transformers at high resolutions uses a lot of memory. I used <a href=\"https://pytorch.org/blog/activation-checkpointing-techniques/\" target=\"_blank\">activation checkpointing</a> to mitigate this. It allows for a configurable tradeoff between speed and memory usage. I found the speed drawbacks to be extremely minor. It can cut memory usage in half with almost zero slowdown.</li>\n<li><strong>Data augmentation:</strong> Aggressive data augmentation seems to do more harm than good for these models, so I used relatively light configs. For most models in the ensemble, I just used <code>A.GridDistortion(num_steps=5, distort_limit=0.2, p=0.15)</code> during pretraining and didn't have any <em>explicit</em> data augmentation during finetuning. However, finetuning intentionally used rectified images from an older, less accurate, version of my preprocessing pipeline, so models are exposed to more rectification errors during training than they are at test time. Using more accurate rectification for the training data makes my scores worse, so I believe rectification inaccuracies act as a form of sneaky data augmentation during finetuning.</li>\n<li><strong>Loss functions:</strong> 3 of the 4 lead segmentation models in my ensemble were just trained to minimize binary cross entropy loss. One of them used a hybrid loss during finetuning in which it also tries to minimize the mean squared error of the predicted y pixel coordinates at each timestep. That seemed to be <em>slightly</em> beneficial (0.05 dB in cross validation, even less on the leaderboard), but I didn't have time to propagate it to all models in the ensemble. Applying L1 or L2 losses like that is something which consistently did more harm than good to me earlier in the competition, so I initially abandoned it, but it seemed somewhat beneficial after the pretraining was added; I think pretraining purely with BCE before adding MSE or MAE helps to prevent much of the overfitting &amp; instability I encountered earlier.</li>\n<li><strong>Misc:</strong> AdamW optimizer &amp; one cycle learning rate scheduler, similar to the keypoint detector.</li>\n</ul>\n<h3>Test time augmentation &amp; ensembling</h3>\n<p>My pipeline rectifies each image twice using separate keypoint detection models, then flips one of the resulting images and feeds them to a collection of 4 lead segmentation models, half of which were trained to process images in the flipped orientation. The resulting heatmaps were then blended by averaging the pixel logits.</p>\n<p><img src=\"https://i.postimg.cc/yNbNxWKp/ECG-ensembling.png\" alt=\"TTA approach\"></p>\n<p>The approach above is based on the following observations:</p>\n<ol>\n<li>Averaging predictions from models in the standard &amp; flipped orientations provides a gain of roughly 0.28 dB.</li>\n<li>Using 2 lead segmentation models per orientation provides a gain of roughly 0.18 dB.</li>\n<li>Using 2 models for rectification (instead of 1) provides a gain of roughly 0.13 dB.</li>\n<li>Averaging the pixel logits scores ~0.02 dB better than averaging after signal extraction.</li>\n<li>Flipping or rotating the images before rectification does not help.</li>\n</ol>\n<p>The score could likely be improved further by processing the images in more orientations, applying some of the test time augmentations unrelated to rotation &amp; flipping from other top solutions &amp; public notebooks, and figuring out why applying TTA before rectification didn't help (maybe I had a bug?), but I didn't have time to experiment with this super extensively.</p>\n<h3>Signal extraction &amp; post processing</h3>\n<p>I convert the raw predicted lead location heatmaps to signals via the following steps:</p>\n<ol>\n<li><strong>Heatmap --&gt; raw signal:</strong> Within each column of the image where a lead is expected to be located (based on the canonical layout), I use a top-k softmax to compute the estimated probability of the lead being located at each vertical pixel coordinate, then use those probability estimates to compute the expected signal value. Using a top-k softmax with k=10 scored ~0.28 dB better than using a hard argmax. I tried k ∈ {3, 5, 10, 20, unlimited} and found 5 &amp; 10 to be roughly tied for \"best\". The optimal choice varies by model depending on minor differences in other hyperparameters.</li>\n<li><strong>Unconfident prediction replacement:</strong> The top-k softmax operations are also used to compute upper and lower confidence bounds for the signal values at each timestep. If those bounds are more than 0.4 mV apart, the sample is dropped and the gaps are filled in by linearly interpolating from the surrounding context samples. The range covered by the confidence interval corresponds to either the top-10 most likely vertical pixel coordinates or the min and max pixel coordinates with associated probability estimates above 0.1%, whichever is narrower for each timestep. This provided a gain of roughly 0.13 dB in comparison to a baseline without confidence filtering.</li>\n<li><strong>Edge artifact suppression:</strong> I remove edge artifacts by replacing the rightmost 2 voltage samples for each lead with the one located 3rd from the right. The rightmost samples tend to be inaccurate because the \"ground truth\" training signals contain some voltage spikes that are not visible in the printed images, which causes the model to be prone to hallucinating at the end. Partially suppressing those errors gave me a ~0.05 dB gain (some of them leak through this filtering, a 2px safety margin is not wide enough to fully eliminate them).</li>\n<li><strong>Signal sharpening:</strong> The peaks of the voltage spikes tend to be a bit \"rounded\" due to the input images having lower resolution than the raw signals they're based on, so a <a href=\"https://en.wikipedia.org/wiki/Top-hat_transform\" target=\"_blank\">top-hat transform</a> is used to sharpen them. This provided a gain of roughly 0.04 dB. It worked better for me than sharpening with a shock filter (+0.01 dB gain in CV, did not test on LB) or 1D unsharp masking (harmful in CV, did not test on LB).</li>\n<li><strong>I/II/III redundancy exploitation (einthoven's law):</strong> This takes place in several stages. First, the I and III leads are aligned with II to correct for small timing discrepancies. Then the II = I + III relation (einthoven's law) is used to compute \"expected\" signal values for each of those 3 leads based on the other two. Finally, the raw extracted signals are blended with their expected values via weighted averaging with two thirds of the weight assigned to the raw values. This provides a gain of roughly 0.05 dB. It was critically important to align the signals first, otherwise this is harmful for me. I also tried exploiting the aVR + aVL + aVF = 0 relation in a similar manner, and observed a similar gain in local cross validation from doing so, but unlike I/II/III the aV* post processing did not work well on the leaderboard, so my final pipeline only does einthoven post-processing for I/II/III.</li>\n<li><strong>II &amp; rhythm strip redundancy exploitation:</strong> The II signal appears twice in each image, once as a 2.5 second segment, then again as a 10 second segment (rhythm strip) whose prefix should match the shorter segment. My II predictions are generated by aligning the shorter signal with the longer one (correcting for small timing discrepancies), then computing a weighted average with ~56% of the weight given to the long signal, ~44% given to the shorter one. This provides a gain of roughly 0.01 dB… barely measurable, but was consistent for both CV &amp; LB.</li>\n<li><strong>Sample rate adjustment:</strong> Signals are re-sampled to match the host's desired sample rates via linear interpolation. I also tried cubic, akima spline, and PCHIP interpolation, but linear worked best.</li>\n</ol>\n<h3>Things that didn't work well for me</h3>\n<ul>\n<li>Correcting for non-affine image warping during preprocessing</li>\n<li>DeepLabV3 segmentation models</li>\n<li>Larger DINO v2 models</li>\n<li>DINO v3</li>\n<li>Post-processing with 1D CNNs</li>\n<li>Using differentiable warping operations so that the preprocessing model(s) can be trained end-to-end with the final lead segmentation model</li>\n</ul>\n<p>… as usual, many of the things that don't work well for me could potentially work well with additional effort and GPU time for tuning. I frequently move on to testing other ideas when early results don't seem promising. Many of the \"bad\" ideas above wound up working well for other top competitors 😅</p>\n<h1>Liu's Pipeline</h1>\n<h3>Acknowledgments</h3>\n<p>I would like to express my sincere gratitude to Kaggle and the competition organizers for providing this invaluable opportunity. Special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing his strong baseline, which served as a crucial foundation for my work.</p>\n<h3>Summary</h3>\n<p>My solution optimizes Stage 2 of <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>’s baseline. The key insight is to treat waveform extraction as <strong>per-column coordinate regression</strong>, rather than relying purely on a “segmentation → post-processing” pipeline.</p>\n<p>I keep the U-Net–style heatmap prediction for stability, but add <strong>CoordLoss</strong>: a GT-centered <strong>local Gaussian cross-entropy</strong> that directly supervises the centerline y-coordinate per column. On my local validation split, this single change improved SNR by <strong>more than +1 dB</strong>, and it was the dominant contributor to overall quality.</p>\n<p>To further reduce train–inference mismatch, validation/inference uses the same <strong>old-subpixel-compatible y extractor</strong> (NaN + interpolation behavior). I also add lightweight “safe” regularizers enforcing physically plausible lead relationships.</p>\n<h3>Overall Pipeline</h3>\n<p>Stage 0/1 (baseline): detect and rectify ECG sheets into canonical coordinates (rectified strips).</p>\n<p>Stage 2 (this work):</p>\n<ul>\n<li><p>Model: ResNet34 encoder + U-Net decoder → <strong>4 heatmaps</strong> (3 short rows + long Lead II).</p></li>\n<li><p>Losses:</p>\n<ul>\n<li>Masked BCE (heatmap supervision)</li>\n<li><strong>CoordLoss (main gain)</strong></li>\n<li>Small auxiliary losses (consistency + lead relationships)</li></ul></li>\n<li><p>Inference:</p>\n<ul>\n<li>Predict full-width logits</li>\n<li>Logits → y(px) using the same old-subpixel extraction as validation (NaN → interpolation)</li>\n<li>Convert y(px) → mV</li>\n<li>Apply Einthoven-based patching for long Lead II near the boundary</li></ul></li>\n</ul>\n<h3>Stage 2 Model</h3>\n<p>Architecture</p>\n<ul>\n<li>Encoder: <code>timm resnet34.a3_in1k</code></li>\n<li>Decoder: U-Net style decoder (MyCoordUnetDecoder) with multi-scale skip connections</li>\n<li>Head: 1×1 conv → 4 heatmaps</li>\n</ul>\n<p>BatchNorm freezing\nAll BN layers are forced into eval mode during training, which stabilizes optimization under batch size 1 with gradient accumulation.</p>\n<p>Losses</p>\n<ol>\n<li><p>Masked Pixel BCE (baseline supervision)\nI rasterize GT polylines into thin heatmap masks and apply BCEWithLogitsLoss, masked by valid columns. Since the final metric is strongly influenced by the rhythm strip, I upweight <strong>Long Lead II</strong> (via channel weighting and/or multiplicity depending on the run).</p></li>\n<li><p>CoordLoss (main gain, +1 dB)\nMotivation: segmentation-style BCE does not directly optimize the quantity we ultimately need—<strong>the centerline y-coordinate per column</strong>. Even small vertical errors can degrade SNR after alignment and interpolation.</p></li>\n</ol>\n<p>Method: for each channel c and valid column x</p>\n<ul>\n<li><p>Treat logits over height as a categorical distribution:</p>\n<ul>\n<li>( p(y\\mid x,c)=\\mathrm{softmax}(z_c[:,x]/T) )</li></ul></li>\n<li><p>Build a GT-centered local Gaussian target around the continuous ground-truth (y^*(x,c)):</p>\n<ul>\n<li>( w_k \\propto \\exp(-(k-y^*)^2/(2\\sigma^2)) ) within a window ±R</li></ul></li>\n<li><p>Compute cross-entropy using log-softmax values in that local window (local Gaussian CE)</p></li>\n</ul>\n<p>This provides dense, well-shaped gradients for precise localization, even when the line is thin or partially missing. In ablations, CoordLoss produced the largest improvement and drove the overall gain.</p>\n<ol>\n<li>Auxiliary “safe” regularizers (small weights)</li>\n</ol>\n<ul>\n<li>Short II vs Long II dy consistency: applied only on segment0 (Short II time), using SmoothL1 on centered dy to reduce drift/slope mismatch.</li>\n<li>Einthoven bridge (mV domain): encourage Long II near the boundary to match an Einthoven-corrected (II_{\\text{corr}}=(1-w),II+w,(I+III)), with warmup/ramp for stability.</li>\n<li>Lead relationship penalty (hinge): enforce (II \\approx I + III) on segment0 using a tolerance-based squared hinge.</li>\n</ul>\n<p>These are intentionally lightweight; they mainly help avoid implausible outputs and reduce boundary artifacts, while CoordLoss provides the primary accuracy boost.</p>\n<h1>Ensembling Approach</h1>\n<p>We ensemble our predictions by:</p>\n<ol>\n<li>Aligning the signals to correct for minor time discrepancies that would otherwise act as a blur filter if the signals were naively averaged together. This is done using code copied from the evaluation metric, very similar to how it aligns signals with the ground truth.</li>\n<li>Computing a weighted average of the predictions. We only did 2 submissions to try tuning the weights, so they're primarily based on rough guesswork, i.e. the assumption that the optimal weight for each pipeline's predictions would probably be correlated with its public LB score. Using weights which lean more heavily in that direction scored better than using relatively even weights (on both the public &amp; private leaderboard), so that assumption appears to have been correct.</li>\n</ol>\n<p>This provided a gain of 0.6 dB over our best individual submission 😁.</p>",
      "rawMarkdown": "# Overview\n\nOur best submission works by running three separate solutions, then computing a weighted average of their predictions. We formed our team late in the competition, so our pipelines were developed with relatively few shared assumptions. This independence increased model diversity, thereby allowing our ensemble to significantly outperform the individual scores.\n\n| Solution | Ensemble Weight |  Public LB Score | Private LB Score |\n| -------- | --------------- | --------------- | ---------------- |\n| Imanishi's pipeline | 45% | 21.58 | 21.43 |\n| James's pipeline | 32% | 21.17 | 20.98 |\n| Liu's pipeline | 23% | 20.84 | 20.74 |\n| Ensemble | N/A | 22.16 | 22.03 |\n\nAt a high level, all of our solutions work by rectifying the input images to correct for distortions, then using segmentation models and softmax operations to extract the voltage signals. However, there were some substantial differences in our preprocessing methodology (1-stage vs. 2-stage), model architectures (CNNs vs. Transformers), segmentation loss functions (binary cross entropy vs. custom \"coord\" loss), and signal post-processing methodology (anomaly suppression techniques & signal sharpening transforms). As a result, the inaccuracies in our predictions are largely uncorrelated and weighted averaging improves the signal to noise ratio substantially.\n\nOur individual solutions are described in greater detail below.\n\n# Imanishi's Pipeline\n\n### Overview of Imanishi's Part\n\nUsing [hengck23’s excellent notebook](https://www.kaggle.com/code/hengck23/demo-submission) as a baseline, I introduced the following improvements:\n\n* Replaced the stage0 model with my own trained model.\n  The original motivation was to completely exclude pretrained data when computing CV scores, but since it also improved the LB score, I adopted it. (Stage1 ultimately remains hengck23’s model.)\n* When converting to rectified input images for stage2, hengck23 performed the transformation in two steps. To reduce image quality degradation, I instead computed a grid from the outputs of stage0 and stage1 and transformed the original image in a single step using `F.grid_sample()`.\n* Increased the stage2 input resolution, especially in the x-direction (input size: **1632 × 4480**).\n* Introduced the **Coord loss** idea from my teammate liuzhangzhen into stage2.\n* Added y-direction clustering–based masking in stage2.\n* Instead of filling low-confidence predictions in stage2 with 0 mV, I switched to **linear interpolation**.\n* Adopted **EfficientNetV2B1-UNet** for the stage0 and stage2 models (it is unclear how much this contributed to accuracy).\n* Added the following post-processing:\n\n  * Correction of Leads I, II, and III using **Einthoven’s law**.\n  * Since the first quarter of Lead II overlaps with two waveforms, apply average ensembling.\n  * Apply `savgol_filter` with a dynamically adjusted `window_length` depending on `sig_len`.\n\n\n\n### Details of Stage2 Model\n\nIn Liu’s implementation of Coord loss, both the segmentation loss and the Coord loss are applied to the same logits. In my case, this did not work well, so I split the model head into two heads (segmentation head, y-coordinate head) and applied **segmentation loss (dice loss + BCE loss)** and **Coord loss** separately.\n\nBoth heads have the same output shape: [height, width, 4ch].\n\nThe segmentation loss acts like an auxiliary loss during training, but the segmentation output is also used during inference for masking low-confidence regions in the x-direction and for the y-clustering described later.\n\n\n\n### Coord Loss\n\nAlthough Liu’s section already explains Coord Loss, I will describe what I consider important.\n\nWith a segmentation-based approach like hengck23’s baseline, where the curve mask is predicted and then `argmax` is used to obtain y-coordinates, it is difficult to estimate optimal y positions. This becomes clear when looking at the ground truth for the two heads (following figure).\n\nWith Coord Loss, the logits output (**height × width × 4 channels**) are the same as in segmentation head up to a point, but then a **softmax over the y-dimension** is applied to directly predict the y-coordinate.\n\nIf we use a one-hot label with only one pixel as ground truth, the model becomes very sensitive to small coordinate shifts. Instead, the ground truth is represented as a **Gaussian-shaped probability distribution** along the y-axis. The Gaussian parameters (sigma and radius) were kept the same as Liu’s because I did not have time to tune them.\n\nThe loss function is standard **cross-entropy loss**, treating the problem as multi-class classification over y positions.\n\nBy applying `argmax` along the y-direction on the Coord Loss head output, we can obtain much more accurate y-coordinates. Adding Coord Loss improved my public LB score by about **+1.0 dB**, which is a large gain.\n\n[![imanishi-figure1.png](https://i.postimg.cc/Fzpb039m/imanishi-figure1.png)](https://postimg.cc/p59nH9h1)\n\n\n### Y-Clustering\n\nFor some reason, my stage2 model occasionally misclassified lead labels, which could significantly degrade the SNR. I wanted to prevent this during training, but in the end I could not, so I handled it with a **post-processing step using y-direction clustering**.\n\nFor y-clustering, I used the output from the segmentation head, and the generated masks were then applied to the y-coordinate head.\n\n[![imanishi-figure2.png](https://i.postimg.cc/L4QM37n9/imanishi-figure2.png)](https://postimg.cc/JtXg1pNv)\n\n\n### Other Tricks\n\nIn the submission notebook, I split the entire test set into two parts and assigned one T4 GPU to each process, as shown in the code below, in order to speed up inference.\n\nIn practice, my part of the pipeline runs in about **70 minutes**, which was the fastest among the team members.\n\n```python\nenv1 = os.environ.copy()\nenv2 = os.environ.copy()\nenv1['CUDA_VISIBLE_DEVICES'] = '0'\nenv2['CUDA_VISIBLE_DEVICES'] = '1'\n\ncmd1 = f'python run.py --half 0'\nproc1 = subprocess.Popen(cmd1.split(' '), env=env1)\n\ncmd2 = f'python run.py --half 1'\nproc2 = subprocess.Popen(cmd2.split(' '), env=env2)\n\n_ = proc1.communicate()\n_ = proc2.communicate()\n```\n\n# James's Pipeline\n### Overview\nMy inference pipeline has 3 main steps:\n1. **Preprocessing:** Detect landmarks in the images and use them to correct for affine warping (such as camera perspective variations & rotation, but not wrinkles in the paper).\n2. **Lead segmentation:** Predict a heatmap describing where the leads hypothetically would be located if the input image was perfectly \"clean\" and undistorted. This is better at handling \"natural\" distortions than unnatural ones caused by imperfect non-affine warp correction, so attempting to correct for non-affine warping with preprocessing logic similar to [@hengck23](https://www.kaggle.com/hengck23)'s baseline is harmful in my pipeline.\n3. **Heatmap-to-signal conversion:** I use top-k softmax operations to predict estimated voltages & confidence intervals for each lead at each timestep, replace extremely low confidence predictions with ones interpolated from context, use a top-hat transform to \"sharpen\" the remaining voltage spikes, and use a little linear algebra to exploit redundancy between leads for denoising purposes.\n\nSteps 1 and 2 both use finetuned versions of [DINOv2-base](https://huggingface.co/timm/vit_base_patch14_reg4_dinov2.lvd142m). I used an ensemble of 6 models in total, 2 in step 1, 4 in step 2. Half of the models process the images in an intentionally flipped orientation as a form of test time data augmentation.\n\n### Data generation\n\nI generated ~175K training examples based on ~22K unique ECG records from the PTB-XL dataset.\n\nThis was done by generating \"clean\" ECG images with [ECG-image-kit](https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator), then intentionally corrupting them with some custom data augmentation code which tries to roughly imitate they types of images that appear in the host's data (cell phone pictures of paper with heavy damage, cell phone pictures of computer screens, black and white scans, etc). \n\nEach image type was roughly simulated using a combination of the following (in no particular order):\n* **Affine warping**\n* **Background image insertion:** For simulated cellphone pics, the warped ECG images were overlaid on top of images from [unsplash-25k](https://www.kaggle.com/datasets/ntsv648/unsplash-25k). For scanner pics, I used a white background.\n* **Stain insertion:** This works by overlaying translucent mold and stain images on top of the ECG images. The stains were randomly drawn from a pool of 26 that I generated semi-manually by prompting a diffusion model. In hindsight, I think the stain intensity distribution I used was a bit unrealistic (typically too transparent) and wonder if maybe pwelin noise would work better than the stain images I generated, but never got around to tinkering with a second version of this.\n* **Wrinkle shading:** This simulates the shadows from hypothetical wrinkles without introducing any non-affine warping. It is very similar to functionality from ECG-image-kit, I just had ChatGPT port it into my script.\n* **Black-and-white scanner simulation:** This ain't an off the shelf greyscale conversion. I tried to simulate the behavior of black and white scanners by randomizing brightness, contrast, gamma, and the way the input color channels are weighted. It also injects a little gaussian noise to mimic sensor noise & quantization artifacts.\n\nAdditional augmentations were performed on-the-fly in my training scripts, those are just the ones I applied before training.\n\nOf the ~175K extra training examples, only ~50K were used to train my final models. Primarily because (1) the accuracy of the keypoint detectors didn't seem to meaningfully improve beyond ~17K training examples and (2) the gains from pretraining the lead segmentation models on synthetic data before finetuning on the host images was pretty small relative to the amount of GPU time it was consuming (a little over two RTX 5090 days of extra compute for a +0.17 dB gain was a bit disappointing... I wanted to try other things instead of having my hardware tied up pushing further in that direction).\n\n### Camera perspective correction\n\nThis works by detecting keypoints in the images, then using those keypoint locations to compute a homography matrix that can be used to correct for differences in camera angle, zoom, and rotation. It is fairly similar to @hengck23's stage 0, with the main differences being that I used a vision transformer instead of a CNN and training data that I generated myself.\n\n![Figure 1: Sample ECG before and after perspective correction](https://i.postimg.cc/xdh0Y7yV/raw-rectified-comparison.png)\n\nKeypoints were detected at a resolution of 1036x1036, then used to directly rectify images from their native resolutions --> 1694x2198.\n\nTraining details:\n* Model architecture: DINOv2-base backbone with linear prediction head that produces 30 channel outputs (1 per target landmark, all lead label text + the start and end of each signal's x axis).\n* Trained to imitate ground truth heatmaps with 1 \"gaussian blob\" per keypoint. These blobs have a value of 1 in the center and decay towards zero as distance from the center increases.\n* BCE loss\n* AdamW optimizer\n* One-cycle learning rate schedule\n* Data augmentations (applied using `albumentations`):\n    * `RandomRotate90`\n    * `GridDistortion`\n    * `ElasticTransform`\n    * `OpticalDistortion`\n    * `GaussNoise`\n    * `GaussianBlur`\n    * `RandomBrightnessContrast`\n    * `ColorJitter`\n    * `CoarseDropout`\n\nBoth of the keypoint detection models in my final ensemble were trained on 17.4K of my synthetic training examples (and none of the host data). They primarily differed in the data agmentation settings & training epoch count. One was trained with moderately heavy agumentation & 24 epochs, the other was trained with heavier augmentation & 48 epochs.\n\nTripling the training data to ~50K examples improved cross validation when testing against other synthetic images, but did not improve the scores of my full pipeline when testing against the ECG images provided by the host, so only ~10% of the available synthetic data was used to train the keypoint detectors in my final ensemble.\n\n### Lead segmentation\n\nI used the ground-truth signal data to generate \"perfect\" heatmaps describing where the leads ought to be located in the images if they were completely undistored, then finetuned DINOv2-base to predict those heatmaps based on images with realistic distortions. This teaches it to automatically correct for any warping which makes it past my preprocessing.\n\n![Predicted heatmap overlaid on original image](https://i.postimg.cc/Yqzqgdks/heatmap-pred.png)\n\nI found it very beneficial to use horizontal resolutions higher than the native resolution of the ECG plots, so this uses an input resolution of 1694x4396 (~2x wider than native) and I inserted a `ConvTranspose2d` layer between the transformer backbone and the prediction head, which increased resolution another 2x shortly before producing the outputs. As a result, the heatmaps have a resolution of **1694x8792**. This provided **MASSIVE score improvements in comparison to just using the naive resolution, roughly a 4.1 dB gain** in early experiments. Using an input resolution 2x higher than native and output resolution 4x higher than native appeared to be ~optimal, adjusting either of those figures by a factor of 2 makes the score worse.\n\nTraining took place in 2 stages:\n1. Pretraining on my synthetic images\n2. Finetuning on the host images\n\nTraining details:\n* **Data:** Stage 1 used 17.4K synthetic images per model with different images used for each model in the ensemble. Stage 2 used 80% of the host data, with 20% held in reserve for cross-validation.\n* **Mask generation:** Unlike the keypoint detection models, I found it beneficial for the ground-truth segmentation masks used to train these models to be very sharp. The ground-truth lead lines are only a single pixel thick with no blur.\n* **Activation checkpointing:** Training vision transformers at high resolutions uses a lot of memory. I used [activation checkpointing](https://pytorch.org/blog/activation-checkpointing-techniques/) to mitigate this. It allows for a configurable tradeoff between speed and memory usage. I found the speed drawbacks to be extremely minor. It can cut memory usage in half with almost zero slowdown.\n* **Data augmentation:** Aggressive data augmentation seems to do more harm than good for these models, so I used relatively light configs. For most models in the ensemble, I just used `A.GridDistortion(num_steps=5, distort_limit=0.2, p=0.15)` during pretraining and didn't have any *explicit* data augmentation during finetuning. However, finetuning intentionally used rectified images from an older, less accurate, version of my preprocessing pipeline, so models are exposed to more rectification errors during training than they are at test time. Using more accurate rectification for the training data makes my scores worse, so I believe rectification inaccuracies act as a form of sneaky data augmentation during finetuning.\n* **Loss functions:** 3 of the 4 lead segmentation models in my ensemble were just trained to minimize binary cross entropy loss. One of them used a hybrid loss during finetuning in which it also tries to minimize the mean squared error of the predicted y pixel coordinates at each timestep. That seemed to be *slightly* beneficial (0.05 dB in cross validation, even less on the leaderboard), but I didn't have time to propagate it to all models in the ensemble. Applying L1 or L2 losses like that is something which consistently did more harm than good to me earlier in the competition, so I initially abandoned it, but it seemed somewhat beneficial after the pretraining was added; I think pretraining purely with BCE before adding MSE or MAE helps to prevent much of the overfitting & instability I encountered earlier.\n* **Misc:** AdamW optimizer & one cycle learning rate scheduler, similar to the keypoint detector.\n\n### Test time augmentation & ensembling\n\nMy pipeline rectifies each image twice using separate keypoint detection models, then flips one of the resulting images and feeds them to a collection of 4 lead segmentation models, half of which were trained to process images in the flipped orientation. The resulting heatmaps were then blended by averaging the pixel logits.\n\n![TTA approach](https://i.postimg.cc/yNbNxWKp/ECG-ensembling.png)\n\nThe approach above is based on the following observations:\n1. Averaging predictions from models in the standard & flipped orientations provides a gain of roughly 0.28 dB.\n2. Using 2 lead segmentation models per orientation provides a gain of roughly 0.18 dB.\n3. Using 2 models for rectification (instead of 1) provides a gain of roughly 0.13 dB.\n4. Averaging the pixel logits scores ~0.02 dB better than averaging after signal extraction.\n5. Flipping or rotating the images before rectification does not help.\n\nThe score could likely be improved further by processing the images in more orientations, applying some of the test time augmentations unrelated to rotation & flipping from other top solutions & public notebooks, and figuring out why applying TTA before rectification didn't help (maybe I had a bug?), but I didn't have time to experiment with this super extensively.\n\n### Signal extraction & post processing\nI convert the raw predicted lead location heatmaps to signals via the following steps:\n1. **Heatmap --> raw signal:** Within each column of the image where a lead is expected to be located (based on the canonical layout), I use a top-k softmax to compute the estimated probability of the lead being located at each vertical pixel coordinate, then use those probability estimates to compute the expected signal value. Using a top-k softmax with k=10 scored ~0.28 dB better than using a hard argmax. I tried k ∈ {3, 5, 10, 20, unlimited} and found 5 & 10 to be roughly tied for \"best\". The optimal choice varies by model depending on minor differences in other hyperparameters.\n2. **Unconfident prediction replacement:** The top-k softmax operations are also used to compute upper and lower confidence bounds for the signal values at each timestep. If those bounds are more than 0.4 mV apart, the sample is dropped and the gaps are filled in by linearly interpolating from the surrounding context samples. The range covered by the confidence interval corresponds to either the top-10 most likely vertical pixel coordinates or the min and max pixel coordinates with associated probability estimates above 0.1%, whichever is narrower for each timestep. This provided a gain of roughly 0.13 dB in comparison to a baseline without confidence filtering.\n3. **Edge artifact suppression:** I remove edge artifacts by replacing the rightmost 2 voltage samples for each lead with the one located 3rd from the right. The rightmost samples tend to be inaccurate because the \"ground truth\" training signals contain some voltage spikes that are not visible in the printed images, which causes the model to be prone to hallucinating at the end. Partially suppressing those errors gave me a ~0.05 dB gain (some of them leak through this filtering, a 2px safety margin is not wide enough to fully eliminate them).\n4. **Signal sharpening:** The peaks of the voltage spikes tend to be a bit \"rounded\" due to the input images having lower resolution than the raw signals they're based on, so a [top-hat transform](https://en.wikipedia.org/wiki/Top-hat_transform) is used to sharpen them. This provided a gain of roughly 0.04 dB. It worked better for me than sharpening with a shock filter (+0.01 dB gain in CV, did not test on LB) or 1D unsharp masking (harmful in CV, did not test on LB).\n5. **I/II/III redundancy exploitation (einthoven's law):** This takes place in several stages. First, the I and III leads are aligned with II to correct for small timing discrepancies. Then the II = I + III relation (einthoven's law) is used to compute \"expected\" signal values for each of those 3 leads based on the other two. Finally, the raw extracted signals are blended with their expected values via weighted averaging with two thirds of the weight assigned to the raw values. This provides a gain of roughly 0.05 dB. It was critically important to align the signals first, otherwise this is harmful for me. I also tried exploiting the aVR + aVL + aVF = 0 relation in a similar manner, and observed a similar gain in local cross validation from doing so, but unlike I/II/III the aV* post processing did not work well on the leaderboard, so my final pipeline only does einthoven post-processing for I/II/III.\n6. **II & rhythm strip redundancy exploitation:** The II signal appears twice in each image, once as a 2.5 second segment, then again as a 10 second segment (rhythm strip) whose prefix should match the shorter segment. My II predictions are generated by aligning the shorter signal with the longer one (correcting for small timing discrepancies), then computing a weighted average with ~56% of the weight given to the long signal, ~44% given to the shorter one. This provides a gain of roughly 0.01 dB... barely measurable, but was consistent for both CV & LB.\n7. **Sample rate adjustment:** Signals are re-sampled to match the host's desired sample rates via linear interpolation. I also tried cubic, akima spline, and PCHIP interpolation, but linear worked best.\n\n### Things that didn't work well for me\n* Correcting for non-affine image warping during preprocessing\n* DeepLabV3 segmentation models\n* Larger DINO v2 models\n* DINO v3\n* Post-processing with 1D CNNs\n* Using differentiable warping operations so that the preprocessing model(s) can be trained end-to-end with the final lead segmentation model\n\n... as usual, many of the things that don't work well for me could potentially work well with additional effort and GPU time for tuning. I frequently move on to testing other ideas when early results don't seem promising. Many of the \"bad\" ideas above wound up working well for other top competitors 😅\n\n# Liu's Pipeline\n\n### Acknowledgments\n\nI would like to express my sincere gratitude to Kaggle and the competition organizers for providing this invaluable opportunity. Special thanks to @hengck23 for sharing his strong baseline, which served as a crucial foundation for my work.\n\n### Summary\n\nMy solution optimizes Stage 2 of @hengck23’s baseline. The key insight is to treat waveform extraction as **per-column coordinate regression**, rather than relying purely on a “segmentation → post-processing” pipeline.\n\nI keep the U-Net–style heatmap prediction for stability, but add **CoordLoss**: a GT-centered **local Gaussian cross-entropy** that directly supervises the centerline y-coordinate per column. On my local validation split, this single change improved SNR by **more than +1 dB**, and it was the dominant contributor to overall quality.\n\nTo further reduce train–inference mismatch, validation/inference uses the same **old-subpixel-compatible y extractor** (NaN + interpolation behavior). I also add lightweight “safe” regularizers enforcing physically plausible lead relationships.\n\n### Overall Pipeline\n\nStage 0/1 (baseline): detect and rectify ECG sheets into canonical coordinates (rectified strips).\n\nStage 2 (this work):\n\n* Model: ResNet34 encoder + U-Net decoder → **4 heatmaps** (3 short rows + long Lead II).\n* Losses:\n\n  * Masked BCE (heatmap supervision)\n  * **CoordLoss (main gain)**\n  * Small auxiliary losses (consistency + lead relationships)\n* Inference:\n\n  * Predict full-width logits\n  * Logits → y(px) using the same old-subpixel extraction as validation (NaN → interpolation)\n  * Convert y(px) → mV\n  * Apply Einthoven-based patching for long Lead II near the boundary\n\n### Stage 2 Model\n\nArchitecture\n\n* Encoder: `timm resnet34.a3_in1k`\n* Decoder: U-Net style decoder (MyCoordUnetDecoder) with multi-scale skip connections\n* Head: 1×1 conv → 4 heatmaps\n\nBatchNorm freezing\nAll BN layers are forced into eval mode during training, which stabilizes optimization under batch size 1 with gradient accumulation.\n\nLosses\n\n1. Masked Pixel BCE (baseline supervision)\n   I rasterize GT polylines into thin heatmap masks and apply BCEWithLogitsLoss, masked by valid columns. Since the final metric is strongly influenced by the rhythm strip, I upweight **Long Lead II** (via channel weighting and/or multiplicity depending on the run).\n\n2. CoordLoss (main gain, +1 dB)\n   Motivation: segmentation-style BCE does not directly optimize the quantity we ultimately need—**the centerline y-coordinate per column**. Even small vertical errors can degrade SNR after alignment and interpolation.\n\nMethod: for each channel c and valid column x\n\n* Treat logits over height as a categorical distribution:\n\n  * ( p(y\\mid x,c)=\\mathrm{softmax}(z_c[:,x]/T) )\n* Build a GT-centered local Gaussian target around the continuous ground-truth (y^*(x,c)):\n\n  * ( w_k \\propto \\exp(-(k-y^*)^2/(2\\sigma^2)) ) within a window ±R\n* Compute cross-entropy using log-softmax values in that local window (local Gaussian CE)\n\nThis provides dense, well-shaped gradients for precise localization, even when the line is thin or partially missing. In ablations, CoordLoss produced the largest improvement and drove the overall gain.\n\n3. Auxiliary “safe” regularizers (small weights)\n\n* Short II vs Long II dy consistency: applied only on segment0 (Short II time), using SmoothL1 on centered dy to reduce drift/slope mismatch.\n* Einthoven bridge (mV domain): encourage Long II near the boundary to match an Einthoven-corrected (II_{\\text{corr}}=(1-w),II+w,(I+III)), with warmup/ramp for stability.\n* Lead relationship penalty (hinge): enforce (II \\approx I + III) on segment0 using a tolerance-based squared hinge.\n\nThese are intentionally lightweight; they mainly help avoid implausible outputs and reduce boundary artifacts, while CoordLoss provides the primary accuracy boost.\n\n# Ensembling Approach\n\nWe ensemble our predictions by:\n1. Aligning the signals to correct for minor time discrepancies that would otherwise act as a blur filter if the signals were naively averaged together. This is done using code copied from the evaluation metric, very similar to how it aligns signals with the ground truth.\n2. Computing a weighted average of the predictions. We only did 2 submissions to try tuning the weights, so they're primarily based on rough guesswork, i.e. the assumption that the optimal weight for each pipeline's predictions would probably be correlated with its public LB score. Using weights which lean more heavily in that direction scored better than using relatively even weights (on both the public & private leaderboard), so that assumption appears to have been correct.\n\nThis provided a gain of 0.6 dB over our best individual submission 😁.",
      "votes": null
    },
    {
      "id": "3396049",
      "postDate": "01/24/2026 07:19:47",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a> <a href=\"https://www.kaggle.com/liuzhangzhen\" target=\"_blank\">@liuzhangzhen</a> </p>",
      "rawMarkdown": "Congrats @jsday96 @dimanishi @liuzhangzhen",
      "votes": null
    },
    {
      "id": "3396063",
      "postDate": "01/24/2026 07:54:07",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/navneetbende\" target=\"_blank\">@navneetbende</a> </p>",
      "rawMarkdown": "Thanks @navneetbende",
      "votes": null
    },
    {
      "id": "3396137",
      "postDate": "01/24/2026 12:11:01",
      "content": "<p>Congratulations on the work. Could you share how you generated the ground-truth (GT) masks from the GT signals for the signal segmentation task?\nIn our case, we used plotted pixel coordinates from a JSON file generated by ECG-image-kit while regenerating image 0001 from the GT signal, and drew the signal mask with a line thickness of 1 pixel. However, even though the segmentation results looked visually good, the reconstructed signal achieved only modest SNR (around 14–15 dB). Our GT masks with resultion 1700-2200 give only 20-22 db</p>",
      "rawMarkdown": "Congratulations on the work. Could you share how you generated the ground-truth (GT) masks from the GT signals for the signal segmentation task?\nIn our case, we used plotted pixel coordinates from a JSON file generated by ECG-image-kit while regenerating image 0001 from the GT signal, and drew the signal mask with a line thickness of 1 pixel. However, even though the segmentation results looked visually good, the reconstructed signal achieved only modest SNR (around 14–15 dB). Our GT masks with resultion 1700-2200 give only 20-22 db",
      "votes": null
    },
    {
      "id": "3396215",
      "postDate": "01/24/2026 14:45:28",
      "content": "<p>Thank you! Each of us three used a different method, but I used matplotlib.</p>",
      "rawMarkdown": "Thank you! Each of us three used a different method, but I used matplotlib.",
      "votes": null
    },
    {
      "id": "3396238",
      "postDate": "01/24/2026 15:19:29",
      "content": "<p>I used <code>cv2.polylines</code> to generate one binary mask per lead, then stacked together those masks to form the full 13 channel target masks.</p>",
      "rawMarkdown": "I used `cv2.polylines` to generate one binary mask per lead, then stacked together those masks to form the full 13 channel target masks.",
      "votes": null
    },
    {
      "id": "3396382",
      "postDate": "01/24/2026 21:28:25",
      "content": "<p><a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> congratulations with GM 🎉</p>",
      "rawMarkdown": "jsday96 congratulations with GM 🎉",
      "votes": null
    },
    {
      "id": "3396404",
      "postDate": "01/25/2026 00:25:50",
      "content": "<p>Thanks! 🎉</p>",
      "rawMarkdown": "Thanks! 🎉",
      "votes": null
    },
    {
      "id": "3396405",
      "postDate": "01/25/2026 00:33:28",
      "content": "<p>I rasterize each valid centerline into a thin 4-channel binary mask by drawing the polyline on a blank canvas per channel using cv2.polylines</p>",
      "rawMarkdown": "I rasterize each valid centerline into a thin 4-channel binary mask by drawing the polyline on a blank canvas per channel using cv2.polylines",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3396049,
      "author_name": "navneetbende",
      "author_url": "",
      "post_date": "01/24/2026 07:19:47",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a> <a href=\"https://www.kaggle.com/liuzhangzhen\" target=\"_blank\">@liuzhangzhen</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3396063,
          "author_name": "liuzhangzhen",
          "author_url": "",
          "post_date": "01/24/2026 07:54:07",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/navneetbende\" target=\"_blank\">@navneetbende</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3396137,
      "author_name": "vandongtran",
      "author_url": "",
      "post_date": "01/24/2026 12:11:01",
      "content": "<p>Congratulations on the work. Could you share how you generated the ground-truth (GT) masks from the GT signals for the signal segmentation task?\nIn our case, we used plotted pixel coordinates from a JSON file generated by ECG-image-kit while regenerating image 0001 from the GT signal, and drew the signal mask with a line thickness of 1 pixel. However, even though the segmentation results looked visually good, the reconstructed signal achieved only modest SNR (around 14–15 dB). Our GT masks with resultion 1700-2200 give only 20-22 db</p>",
      "votes": null,
      "replies": [
        {
          "id": 3396215,
          "author_name": "dimanishi",
          "author_url": "",
          "post_date": "01/24/2026 14:45:28",
          "content": "<p>Thank you! Each of us three used a different method, but I used matplotlib.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3396238,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "01/24/2026 15:19:29",
          "content": "<p>I used <code>cv2.polylines</code> to generate one binary mask per lead, then stacked together those masks to form the full 13 channel target masks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3396405,
          "author_name": "liuzhangzhen",
          "author_url": "",
          "post_date": "01/25/2026 00:33:28",
          "content": "<p>I rasterize each valid centerline into a thin 4-channel binary mask by drawing the polyline on a blank canvas per channel using cv2.polylines</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3396382,
      "author_name": "bluepill",
      "author_url": "",
      "post_date": "01/24/2026 21:28:25",
      "content": "<p><a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> congratulations with GM 🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 3396404,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "01/25/2026 00:25:50",
          "content": "<p>Thanks! 🎉</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3395990": "# Overview\n\nOur best submission works by running three separate solutions, then computing a weighted average of their predictions. We formed our team late in the competition, so our pipelines were developed with relatively few shared assumptions. This independence increased model diversity, thereby allowing our ensemble to significantly outperform the individual scores.\n\n| Solution | Ensemble Weight |  Public LB Score | Private LB Score |\n| -------- | --------------- | --------------- | ---------------- |\n| Imanishi's pipeline | 45% | 21.58 | 21.43 |\n| James's pipeline | 32% | 21.17 | 20.98 |\n| Liu's pipeline | 23% | 20.84 | 20.74 |\n| Ensemble | N/A | 22.16 | 22.03 |\n\nAt a high level, all of our solutions work by rectifying the input images to correct for distortions, then using segmentation models and softmax operations to extract the voltage signals. However, there were some substantial differences in our preprocessing methodology (1-stage vs. 2-stage), model architectures (CNNs vs. Transformers), segmentation loss functions (binary cross entropy vs. custom \"coord\" loss), and signal post-processing methodology (anomaly suppression techniques & signal sharpening transforms). As a result, the inaccuracies in our predictions are largely uncorrelated and weighted averaging improves the signal to noise ratio substantially.\n\nOur individual solutions are described in greater detail below.\n\n# Imanishi's Pipeline\n\n### Overview of Imanishi's Part\n\nUsing [hengck23’s excellent notebook](https://www.kaggle.com/code/hengck23/demo-submission) as a baseline, I introduced the following improvements:\n\n* Replaced the stage0 model with my own trained model.\n  The original motivation was to completely exclude pretrained data when computing CV scores, but since it also improved the LB score, I adopted it. (Stage1 ultimately remains hengck23’s model.)\n* When converting to rectified input images for stage2, hengck23 performed the transformation in two steps. To reduce image quality degradation, I instead computed a grid from the outputs of stage0 and stage1 and transformed the original image in a single step using `F.grid_sample()`.\n* Increased the stage2 input resolution, especially in the x-direction (input size: **1632 × 4480**).\n* Introduced the **Coord loss** idea from my teammate liuzhangzhen into stage2.\n* Added y-direction clustering–based masking in stage2.\n* Instead of filling low-confidence predictions in stage2 with 0 mV, I switched to **linear interpolation**.\n* Adopted **EfficientNetV2B1-UNet** for the stage0 and stage2 models (it is unclear how much this contributed to accuracy).\n* Added the following post-processing:\n\n  * Correction of Leads I, II, and III using **Einthoven’s law**.\n  * Since the first quarter of Lead II overlaps with two waveforms, apply average ensembling.\n  * Apply `savgol_filter` with a dynamically adjusted `window_length` depending on `sig_len`.\n\n\n\n### Details of Stage2 Model\n\nIn Liu’s implementation of Coord loss, both the segmentation loss and the Coord loss are applied to the same logits. In my case, this did not work well, so I split the model head into two heads (segmentation head, y-coordinate head) and applied **segmentation loss (dice loss + BCE loss)** and **Coord loss** separately.\n\nBoth heads have the same output shape: [height, width, 4ch].\n\nThe segmentation loss acts like an auxiliary loss during training, but the segmentation output is also used during inference for masking low-confidence regions in the x-direction and for the y-clustering described later.\n\n\n\n### Coord Loss\n\nAlthough Liu’s section already explains Coord Loss, I will describe what I consider important.\n\nWith a segmentation-based approach like hengck23’s baseline, where the curve mask is predicted and then `argmax` is used to obtain y-coordinates, it is difficult to estimate optimal y positions. This becomes clear when looking at the ground truth for the two heads (following figure).\n\nWith Coord Loss, the logits output (**height × width × 4 channels**) are the same as in segmentation head up to a point, but then a **softmax over the y-dimension** is applied to directly predict the y-coordinate.\n\nIf we use a one-hot label with only one pixel as ground truth, the model becomes very sensitive to small coordinate shifts. Instead, the ground truth is represented as a **Gaussian-shaped probability distribution** along the y-axis. The Gaussian parameters (sigma and radius) were kept the same as Liu’s because I did not have time to tune them.\n\nThe loss function is standard **cross-entropy loss**, treating the problem as multi-class classification over y positions.\n\nBy applying `argmax` along the y-direction on the Coord Loss head output, we can obtain much more accurate y-coordinates. Adding Coord Loss improved my public LB score by about **+1.0 dB**, which is a large gain.\n\n[![imanishi-figure1.png](https://i.postimg.cc/Fzpb039m/imanishi-figure1.png)](https://postimg.cc/p59nH9h1)\n\n\n### Y-Clustering\n\nFor some reason, my stage2 model occasionally misclassified lead labels, which could significantly degrade the SNR. I wanted to prevent this during training, but in the end I could not, so I handled it with a **post-processing step using y-direction clustering**.\n\nFor y-clustering, I used the output from the segmentation head, and the generated masks were then applied to the y-coordinate head.\n\n[![imanishi-figure2.png](https://i.postimg.cc/L4QM37n9/imanishi-figure2.png)](https://postimg.cc/JtXg1pNv)\n\n\n### Other Tricks\n\nIn the submission notebook, I split the entire test set into two parts and assigned one T4 GPU to each process, as shown in the code below, in order to speed up inference.\n\nIn practice, my part of the pipeline runs in about **70 minutes**, which was the fastest among the team members.\n\n```python\nenv1 = os.environ.copy()\nenv2 = os.environ.copy()\nenv1['CUDA_VISIBLE_DEVICES'] = '0'\nenv2['CUDA_VISIBLE_DEVICES'] = '1'\n\ncmd1 = f'python run.py --half 0'\nproc1 = subprocess.Popen(cmd1.split(' '), env=env1)\n\ncmd2 = f'python run.py --half 1'\nproc2 = subprocess.Popen(cmd2.split(' '), env=env2)\n\n_ = proc1.communicate()\n_ = proc2.communicate()\n```\n\n# James's Pipeline\n### Overview\nMy inference pipeline has 3 main steps:\n1. **Preprocessing:** Detect landmarks in the images and use them to correct for affine warping (such as camera perspective variations & rotation, but not wrinkles in the paper).\n2. **Lead segmentation:** Predict a heatmap describing where the leads hypothetically would be located if the input image was perfectly \"clean\" and undistorted. This is better at handling \"natural\" distortions than unnatural ones caused by imperfect non-affine warp correction, so attempting to correct for non-affine warping with preprocessing logic similar to [@hengck23](https://www.kaggle.com/hengck23)'s baseline is harmful in my pipeline.\n3. **Heatmap-to-signal conversion:** I use top-k softmax operations to predict estimated voltages & confidence intervals for each lead at each timestep, replace extremely low confidence predictions with ones interpolated from context, use a top-hat transform to \"sharpen\" the remaining voltage spikes, and use a little linear algebra to exploit redundancy between leads for denoising purposes.\n\nSteps 1 and 2 both use finetuned versions of [DINOv2-base](https://huggingface.co/timm/vit_base_patch14_reg4_dinov2.lvd142m). I used an ensemble of 6 models in total, 2 in step 1, 4 in step 2. Half of the models process the images in an intentionally flipped orientation as a form of test time data augmentation.\n\n### Data generation\n\nI generated ~175K training examples based on ~22K unique ECG records from the PTB-XL dataset.\n\nThis was done by generating \"clean\" ECG images with [ECG-image-kit](https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator), then intentionally corrupting them with some custom data augmentation code which tries to roughly imitate they types of images that appear in the host's data (cell phone pictures of paper with heavy damage, cell phone pictures of computer screens, black and white scans, etc). \n\nEach image type was roughly simulated using a combination of the following (in no particular order):\n* **Affine warping**\n* **Background image insertion:** For simulated cellphone pics, the warped ECG images were overlaid on top of images from [unsplash-25k](https://www.kaggle.com/datasets/ntsv648/unsplash-25k). For scanner pics, I used a white background.\n* **Stain insertion:** This works by overlaying translucent mold and stain images on top of the ECG images. The stains were randomly drawn from a pool of 26 that I generated semi-manually by prompting a diffusion model. In hindsight, I think the stain intensity distribution I used was a bit unrealistic (typically too transparent) and wonder if maybe pwelin noise would work better than the stain images I generated, but never got around to tinkering with a second version of this.\n* **Wrinkle shading:** This simulates the shadows from hypothetical wrinkles without introducing any non-affine warping. It is very similar to functionality from ECG-image-kit, I just had ChatGPT port it into my script.\n* **Black-and-white scanner simulation:** This ain't an off the shelf greyscale conversion. I tried to simulate the behavior of black and white scanners by randomizing brightness, contrast, gamma, and the way the input color channels are weighted. It also injects a little gaussian noise to mimic sensor noise & quantization artifacts.\n\nAdditional augmentations were performed on-the-fly in my training scripts, those are just the ones I applied before training.\n\nOf the ~175K extra training examples, only ~50K were used to train my final models. Primarily because (1) the accuracy of the keypoint detectors didn't seem to meaningfully improve beyond ~17K training examples and (2) the gains from pretraining the lead segmentation models on synthetic data before finetuning on the host images was pretty small relative to the amount of GPU time it was consuming (a little over two RTX 5090 days of extra compute for a +0.17 dB gain was a bit disappointing... I wanted to try other things instead of having my hardware tied up pushing further in that direction).\n\n### Camera perspective correction\n\nThis works by detecting keypoints in the images, then using those keypoint locations to compute a homography matrix that can be used to correct for differences in camera angle, zoom, and rotation. It is fairly similar to @hengck23's stage 0, with the main differences being that I used a vision transformer instead of a CNN and training data that I generated myself.\n\n![Figure 1: Sample ECG before and after perspective correction](https://i.postimg.cc/xdh0Y7yV/raw-rectified-comparison.png)\n\nKeypoints were detected at a resolution of 1036x1036, then used to directly rectify images from their native resolutions --> 1694x2198.\n\nTraining details:\n* Model architecture: DINOv2-base backbone with linear prediction head that produces 30 channel outputs (1 per target landmark, all lead label text + the start and end of each signal's x axis).\n* Trained to imitate ground truth heatmaps with 1 \"gaussian blob\" per keypoint. These blobs have a value of 1 in the center and decay towards zero as distance from the center increases.\n* BCE loss\n* AdamW optimizer\n* One-cycle learning rate schedule\n* Data augmentations (applied using `albumentations`):\n    * `RandomRotate90`\n    * `GridDistortion`\n    * `ElasticTransform`\n    * `OpticalDistortion`\n    * `GaussNoise`\n    * `GaussianBlur`\n    * `RandomBrightnessContrast`\n    * `ColorJitter`\n    * `CoarseDropout`\n\nBoth of the keypoint detection models in my final ensemble were trained on 17.4K of my synthetic training examples (and none of the host data). They primarily differed in the data agmentation settings & training epoch count. One was trained with moderately heavy agumentation & 24 epochs, the other was trained with heavier augmentation & 48 epochs.\n\nTripling the training data to ~50K examples improved cross validation when testing against other synthetic images, but did not improve the scores of my full pipeline when testing against the ECG images provided by the host, so only ~10% of the available synthetic data was used to train the keypoint detectors in my final ensemble.\n\n### Lead segmentation\n\nI used the ground-truth signal data to generate \"perfect\" heatmaps describing where the leads ought to be located in the images if they were completely undistored, then finetuned DINOv2-base to predict those heatmaps based on images with realistic distortions. This teaches it to automatically correct for any warping which makes it past my preprocessing.\n\n![Predicted heatmap overlaid on original image](https://i.postimg.cc/Yqzqgdks/heatmap-pred.png)\n\nI found it very beneficial to use horizontal resolutions higher than the native resolution of the ECG plots, so this uses an input resolution of 1694x4396 (~2x wider than native) and I inserted a `ConvTranspose2d` layer between the transformer backbone and the prediction head, which increased resolution another 2x shortly before producing the outputs. As a result, the heatmaps have a resolution of **1694x8792**. This provided **MASSIVE score improvements in comparison to just using the naive resolution, roughly a 4.1 dB gain** in early experiments. Using an input resolution 2x higher than native and output resolution 4x higher than native appeared to be ~optimal, adjusting either of those figures by a factor of 2 makes the score worse.\n\nTraining took place in 2 stages:\n1. Pretraining on my synthetic images\n2. Finetuning on the host images\n\nTraining details:\n* **Data:** Stage 1 used 17.4K synthetic images per model with different images used for each model in the ensemble. Stage 2 used 80% of the host data, with 20% held in reserve for cross-validation.\n* **Mask generation:** Unlike the keypoint detection models, I found it beneficial for the ground-truth segmentation masks used to train these models to be very sharp. The ground-truth lead lines are only a single pixel thick with no blur.\n* **Activation checkpointing:** Training vision transformers at high resolutions uses a lot of memory. I used [activation checkpointing](https://pytorch.org/blog/activation-checkpointing-techniques/) to mitigate this. It allows for a configurable tradeoff between speed and memory usage. I found the speed drawbacks to be extremely minor. It can cut memory usage in half with almost zero slowdown.\n* **Data augmentation:** Aggressive data augmentation seems to do more harm than good for these models, so I used relatively light configs. For most models in the ensemble, I just used `A.GridDistortion(num_steps=5, distort_limit=0.2, p=0.15)` during pretraining and didn't have any *explicit* data augmentation during finetuning. However, finetuning intentionally used rectified images from an older, less accurate, version of my preprocessing pipeline, so models are exposed to more rectification errors during training than they are at test time. Using more accurate rectification for the training data makes my scores worse, so I believe rectification inaccuracies act as a form of sneaky data augmentation during finetuning.\n* **Loss functions:** 3 of the 4 lead segmentation models in my ensemble were just trained to minimize binary cross entropy loss. One of them used a hybrid loss during finetuning in which it also tries to minimize the mean squared error of the predicted y pixel coordinates at each timestep. That seemed to be *slightly* beneficial (0.05 dB in cross validation, even less on the leaderboard), but I didn't have time to propagate it to all models in the ensemble. Applying L1 or L2 losses like that is something which consistently did more harm than good to me earlier in the competition, so I initially abandoned it, but it seemed somewhat beneficial after the pretraining was added; I think pretraining purely with BCE before adding MSE or MAE helps to prevent much of the overfitting & instability I encountered earlier.\n* **Misc:** AdamW optimizer & one cycle learning rate scheduler, similar to the keypoint detector.\n\n### Test time augmentation & ensembling\n\nMy pipeline rectifies each image twice using separate keypoint detection models, then flips one of the resulting images and feeds them to a collection of 4 lead segmentation models, half of which were trained to process images in the flipped orientation. The resulting heatmaps were then blended by averaging the pixel logits.\n\n![TTA approach](https://i.postimg.cc/yNbNxWKp/ECG-ensembling.png)\n\nThe approach above is based on the following observations:\n1. Averaging predictions from models in the standard & flipped orientations provides a gain of roughly 0.28 dB.\n2. Using 2 lead segmentation models per orientation provides a gain of roughly 0.18 dB.\n3. Using 2 models for rectification (instead of 1) provides a gain of roughly 0.13 dB.\n4. Averaging the pixel logits scores ~0.02 dB better than averaging after signal extraction.\n5. Flipping or rotating the images before rectification does not help.\n\nThe score could likely be improved further by processing the images in more orientations, applying some of the test time augmentations unrelated to rotation & flipping from other top solutions & public notebooks, and figuring out why applying TTA before rectification didn't help (maybe I had a bug?), but I didn't have time to experiment with this super extensively.\n\n### Signal extraction & post processing\nI convert the raw predicted lead location heatmaps to signals via the following steps:\n1. **Heatmap --> raw signal:** Within each column of the image where a lead is expected to be located (based on the canonical layout), I use a top-k softmax to compute the estimated probability of the lead being located at each vertical pixel coordinate, then use those probability estimates to compute the expected signal value. Using a top-k softmax with k=10 scored ~0.28 dB better than using a hard argmax. I tried k ∈ {3, 5, 10, 20, unlimited} and found 5 & 10 to be roughly tied for \"best\". The optimal choice varies by model depending on minor differences in other hyperparameters.\n2. **Unconfident prediction replacement:** The top-k softmax operations are also used to compute upper and lower confidence bounds for the signal values at each timestep. If those bounds are more than 0.4 mV apart, the sample is dropped and the gaps are filled in by linearly interpolating from the surrounding context samples. The range covered by the confidence interval corresponds to either the top-10 most likely vertical pixel coordinates or the min and max pixel coordinates with associated probability estimates above 0.1%, whichever is narrower for each timestep. This provided a gain of roughly 0.13 dB in comparison to a baseline without confidence filtering.\n3. **Edge artifact suppression:** I remove edge artifacts by replacing the rightmost 2 voltage samples for each lead with the one located 3rd from the right. The rightmost samples tend to be inaccurate because the \"ground truth\" training signals contain some voltage spikes that are not visible in the printed images, which causes the model to be prone to hallucinating at the end. Partially suppressing those errors gave me a ~0.05 dB gain (some of them leak through this filtering, a 2px safety margin is not wide enough to fully eliminate them).\n4. **Signal sharpening:** The peaks of the voltage spikes tend to be a bit \"rounded\" due to the input images having lower resolution than the raw signals they're based on, so a [top-hat transform](https://en.wikipedia.org/wiki/Top-hat_transform) is used to sharpen them. This provided a gain of roughly 0.04 dB. It worked better for me than sharpening with a shock filter (+0.01 dB gain in CV, did not test on LB) or 1D unsharp masking (harmful in CV, did not test on LB).\n5. **I/II/III redundancy exploitation (einthoven's law):** This takes place in several stages. First, the I and III leads are aligned with II to correct for small timing discrepancies. Then the II = I + III relation (einthoven's law) is used to compute \"expected\" signal values for each of those 3 leads based on the other two. Finally, the raw extracted signals are blended with their expected values via weighted averaging with two thirds of the weight assigned to the raw values. This provides a gain of roughly 0.05 dB. It was critically important to align the signals first, otherwise this is harmful for me. I also tried exploiting the aVR + aVL + aVF = 0 relation in a similar manner, and observed a similar gain in local cross validation from doing so, but unlike I/II/III the aV* post processing did not work well on the leaderboard, so my final pipeline only does einthoven post-processing for I/II/III.\n6. **II & rhythm strip redundancy exploitation:** The II signal appears twice in each image, once as a 2.5 second segment, then again as a 10 second segment (rhythm strip) whose prefix should match the shorter segment. My II predictions are generated by aligning the shorter signal with the longer one (correcting for small timing discrepancies), then computing a weighted average with ~56% of the weight given to the long signal, ~44% given to the shorter one. This provides a gain of roughly 0.01 dB... barely measurable, but was consistent for both CV & LB.\n7. **Sample rate adjustment:** Signals are re-sampled to match the host's desired sample rates via linear interpolation. I also tried cubic, akima spline, and PCHIP interpolation, but linear worked best.\n\n### Things that didn't work well for me\n* Correcting for non-affine image warping during preprocessing\n* DeepLabV3 segmentation models\n* Larger DINO v2 models\n* DINO v3\n* Post-processing with 1D CNNs\n* Using differentiable warping operations so that the preprocessing model(s) can be trained end-to-end with the final lead segmentation model\n\n... as usual, many of the things that don't work well for me could potentially work well with additional effort and GPU time for tuning. I frequently move on to testing other ideas when early results don't seem promising. Many of the \"bad\" ideas above wound up working well for other top competitors 😅\n\n# Liu's Pipeline\n\n### Acknowledgments\n\nI would like to express my sincere gratitude to Kaggle and the competition organizers for providing this invaluable opportunity. Special thanks to @hengck23 for sharing his strong baseline, which served as a crucial foundation for my work.\n\n### Summary\n\nMy solution optimizes Stage 2 of @hengck23’s baseline. The key insight is to treat waveform extraction as **per-column coordinate regression**, rather than relying purely on a “segmentation → post-processing” pipeline.\n\nI keep the U-Net–style heatmap prediction for stability, but add **CoordLoss**: a GT-centered **local Gaussian cross-entropy** that directly supervises the centerline y-coordinate per column. On my local validation split, this single change improved SNR by **more than +1 dB**, and it was the dominant contributor to overall quality.\n\nTo further reduce train–inference mismatch, validation/inference uses the same **old-subpixel-compatible y extractor** (NaN + interpolation behavior). I also add lightweight “safe” regularizers enforcing physically plausible lead relationships.\n\n### Overall Pipeline\n\nStage 0/1 (baseline): detect and rectify ECG sheets into canonical coordinates (rectified strips).\n\nStage 2 (this work):\n\n* Model: ResNet34 encoder + U-Net decoder → **4 heatmaps** (3 short rows + long Lead II).\n* Losses:\n\n  * Masked BCE (heatmap supervision)\n  * **CoordLoss (main gain)**\n  * Small auxiliary losses (consistency + lead relationships)\n* Inference:\n\n  * Predict full-width logits\n  * Logits → y(px) using the same old-subpixel extraction as validation (NaN → interpolation)\n  * Convert y(px) → mV\n  * Apply Einthoven-based patching for long Lead II near the boundary\n\n### Stage 2 Model\n\nArchitecture\n\n* Encoder: `timm resnet34.a3_in1k`\n* Decoder: U-Net style decoder (MyCoordUnetDecoder) with multi-scale skip connections\n* Head: 1×1 conv → 4 heatmaps\n\nBatchNorm freezing\nAll BN layers are forced into eval mode during training, which stabilizes optimization under batch size 1 with gradient accumulation.\n\nLosses\n\n1. Masked Pixel BCE (baseline supervision)\n   I rasterize GT polylines into thin heatmap masks and apply BCEWithLogitsLoss, masked by valid columns. Since the final metric is strongly influenced by the rhythm strip, I upweight **Long Lead II** (via channel weighting and/or multiplicity depending on the run).\n\n2. CoordLoss (main gain, +1 dB)\n   Motivation: segmentation-style BCE does not directly optimize the quantity we ultimately need—**the centerline y-coordinate per column**. Even small vertical errors can degrade SNR after alignment and interpolation.\n\nMethod: for each channel c and valid column x\n\n* Treat logits over height as a categorical distribution:\n\n  * ( p(y\\mid x,c)=\\mathrm{softmax}(z_c[:,x]/T) )\n* Build a GT-centered local Gaussian target around the continuous ground-truth (y^*(x,c)):\n\n  * ( w_k \\propto \\exp(-(k-y^*)^2/(2\\sigma^2)) ) within a window ±R\n* Compute cross-entropy using log-softmax values in that local window (local Gaussian CE)\n\nThis provides dense, well-shaped gradients for precise localization, even when the line is thin or partially missing. In ablations, CoordLoss produced the largest improvement and drove the overall gain.\n\n3. Auxiliary “safe” regularizers (small weights)\n\n* Short II vs Long II dy consistency: applied only on segment0 (Short II time), using SmoothL1 on centered dy to reduce drift/slope mismatch.\n* Einthoven bridge (mV domain): encourage Long II near the boundary to match an Einthoven-corrected (II_{\\text{corr}}=(1-w),II+w,(I+III)), with warmup/ramp for stability.\n* Lead relationship penalty (hinge): enforce (II \\approx I + III) on segment0 using a tolerance-based squared hinge.\n\nThese are intentionally lightweight; they mainly help avoid implausible outputs and reduce boundary artifacts, while CoordLoss provides the primary accuracy boost.\n\n# Ensembling Approach\n\nWe ensemble our predictions by:\n1. Aligning the signals to correct for minor time discrepancies that would otherwise act as a blur filter if the signals were naively averaged together. This is done using code copied from the evaluation metric, very similar to how it aligns signals with the ground truth.\n2. Computing a weighted average of the predictions. We only did 2 submissions to try tuning the weights, so they're primarily based on rough guesswork, i.e. the assumption that the optimal weight for each pipeline's predictions would probably be correlated with its public LB score. Using weights which lean more heavily in that direction scored better than using relatively even weights (on both the public & private leaderboard), so that assumption appears to have been correct.\n\nThis provided a gain of 0.6 dB over our best individual submission 😁.",
    "3396049": "Congrats @jsday96 @dimanishi @liuzhangzhen",
    "3396063": "Thanks @navneetbende",
    "3396137": "Congratulations on the work. Could you share how you generated the ground-truth (GT) masks from the GT signals for the signal segmentation task?\nIn our case, we used plotted pixel coordinates from a JSON file generated by ECG-image-kit while regenerating image 0001 from the GT signal, and drew the signal mask with a line thickness of 1 pixel. However, even though the segmentation results looked visually good, the reconstructed signal achieved only modest SNR (around 14–15 dB). Our GT masks with resultion 1700-2200 give only 20-22 db",
    "3396215": "Thank you! Each of us three used a different method, but I used matplotlib.",
    "3396238": "I used `cv2.polylines` to generate one binary mask per lead, then stacked together those masks to form the full 13 channel target masks.",
    "3396382": "jsday96 congratulations with GM 🎉",
    "3396404": "Thanks! 🎉",
    "3396405": "I rasterize each valid centerline into a thin 4-channel binary mask by drawing the polyline on a blank canvas per channel using cv2.polylines"
  },
  "source": "meta"
}