{
  "id": 670060,
  "title": "24th Place Solution",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/670060",
  "author_name": "Uday Bhatia",
  "post_date": "2026-01-25T20:04:55.401000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>24th Place Solution: V19 + V16V18 Ensemble (19.59 SNR)</h1>\n<h2>Approach: Regression, Not Segmentation</h2>\n<p>We initially tried segmentation (predicting binary trace masks), but it failed—thick traces, overlapping leads, and the need for sub-pixel accuracy made it impractical. Instead, we frame this as <strong>per-pixel y-coordinate regression</strong>: for each x-position, predict the continuous y-coordinate of the ECG trace.</p>\n<hr>\n<h2>Data Strategy</h2>\n<p><strong>Problem</strong>: Kaggle GT is in millivolts, not pixels. Converting back introduces alignment errors.</p>\n<p><strong>Solution</strong>: Train primarily on <strong>synthetic ECGs</strong> with pixel-perfect ground truth, then mix in Kaggle data for domain adaptation.</p>\n<p>We extract <strong>individual rows</strong> as separate samples (4× more data), cropping each row centered on its baseline (±250px) and using only the signal region (3926px wide, no margins).</p>\n<hr>\n<h2>Architecture</h2>\n<h3>V16: Per-Lead Regression Baseline</h3>\n<pre><code>Input: [B, 3, 500, 3926] (single row crop)\n  │\n  ▼\nConvNeXt-Base Encoder (ImageNet-22k pretrained)\n  │\n  ▼\nU-Net Decoder with Skip Connections\n  │\n  ▼\nHeight Attention (softmax over y-axis)\n  │\n  ▼\n1D Conv Regression Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926] normalized y-coordinates\n</code></pre>\n<h3>V18: Cross-Row Refiner (Stacked on V16)</h3>\n<p>Takes V16's predictions and refines them by looking at <strong>all 4 rows simultaneously</strong>:</p>\n<ul>\n<li>Creates Gaussian \"guide channel\" from V16 prediction (σ=15px)</li>\n<li>EfficientNet-B0 encoder (4 channels: RGB + guide)</li>\n<li><strong>3 Cross-Row Transformer blocks</strong> with self-attention across rows</li>\n<li>Outputs small residual corrections (Tanh × learnable scale ~0.1)</li>\n</ul>\n<p><strong>Key insight</strong>: Cross-row attention implicitly learns Einthoven's Law (II = I + III) and catches inter-row inconsistencies.</p>\n<h3>V19: BiLSTM + Deformable Conv</h3>\n<pre><code>Input: [B, 3, 500, 3926]\n  │\n  ▼\nConvNeXt-Base Encoder\n  │\n  ▼\nU-Net Decoder with Deformable Conv (last 2 stages)\n  │\n  ▼\nHeight Attention\n  │\n  ▼\nBidirectional LSTM (2 layers, hidden=128)\n  │\n  ▼\nLinear Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926]\n</code></pre>\n<p><strong>Deformable Conv</strong>: Learns adaptive receptive fields for curved/warped traces.\n<strong>BiLSTM</strong>: Captures long-range temporal dependencies across the full 10-second ECG.</p>\n<hr>\n<h2>Why Ensemble Works: Complementary Failure Modes</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>V16+V18</th>\n<th>V19</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Strength</strong></td>\n<td>Cross-row consistency</td>\n<td>Long-range temporal</td>\n</tr>\n<tr>\n<td><strong>Weakness</strong></td>\n<td>Local receptive field</td>\n<td>No cross-row awareness</td>\n</tr>\n<tr>\n<td><strong>Fails on</strong></td>\n<td>Baseline drift, rhythm consistency</td>\n<td>Cross-row artifacts, lead relationships</td>\n</tr>\n</tbody>\n</table>\n<p>The architectures have <strong>uncorrelated errors</strong> (~0.3-0.4 correlation). When averaged, mistakes cancel out.</p>\n<pre><code>ensemble = 0.4 * V19 + 0.6 * V16V18  # Empirically best ratio\n</code></pre>\n<hr>\n<h2>Training</h2>\n<p><strong>Loss Function</strong>:</p>\n<pre><code>Loss = 1.0 × MaskedL1 + 0.2 × SNRLoss + 0.1 × GradientLoss\n</code></pre>\n<ul>\n<li><strong>MaskedL1</strong>: Pixel accuracy on valid regions</li>\n<li><strong>SNRLoss</strong>: Direct optimization of competition metric (<code>-10*log10(signal²/noise²)</code>)</li>\n<li><strong>GradientLoss</strong>: Smoothness constraint on first derivative</li>\n</ul>\n<p><strong>Setup</strong>:</p>\n<ul>\n<li>AdamW (lr=1e-4, weight_decay=0.01)</li>\n<li>Cosine annealing schedule</li>\n<li>Mixed precision (FP16) with gradient scaling</li>\n<li>Heavy augmentation (noise, blur, color jitter, coarse dropout)</li>\n<li>Multi-GPU via DistributedDataParallel</li>\n</ul>\n<p><strong>V18 Training</strong>: Freeze V16, train refiner to predict small residual corrections.</p>\n<hr>\n<h2>Post-Processing</h2>\n<ul>\n<li>Savitzky-Golay smoothing (window=7, polyorder=2)</li>\n<li>Einthoven's Law correction (α=0.25 error distribution)</li>\n<li>Amplitude clamping (±10 mV)</li>\n</ul>\n<hr>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Val SNR</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>V16 only</td>\n<td>~19.5 dB</td>\n<td>~19.1</td>\n</tr>\n<tr>\n<td>V16 + V18</td>\n<td>~19.7 dB</td>\n<td>~19.5</td>\n</tr>\n<tr>\n<td>V19 only</td>\n<td>~20 dB</td>\n<td>~19.5</td>\n</tr>\n<tr>\n<td><strong>Ensemble</strong></td>\n<td>—</td>\n<td><strong>~19.75</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Key Takeaways</h2>\n<ol>\n<li><strong>Regression &gt; Segmentation</strong> for sub-pixel trace extraction</li>\n<li><strong>Synthetic data</strong> with pixel-perfect GT is essential</li>\n<li><strong>Architectural diversity</strong> → uncorrelated errors → effective ensemble</li>\n<li><strong>Cross-row attention</strong> (V18) and <strong>BiLSTM</strong> (V19) complement each other perfectly</li>\n<li><strong>Direct SNR loss</strong> optimization helps</li>\n</ol>\n<p>Thanks for reading!</p>\n<h3>Special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the stage 0/1 pipeline</h3>",
  "messages": [
    {
      "id": 3396784,
      "postDate": "2026-01-25T20:04:55.403Z",
      "content": "<h1>24th Place Solution: V19 + V16V18 Ensemble (19.59 SNR)</h1>\n<h2>Approach: Regression, Not Segmentation</h2>\n<p>We initially tried segmentation (predicting binary trace masks), but it failed—thick traces, overlapping leads, and the need for sub-pixel accuracy made it impractical. Instead, we frame this as <strong>per-pixel y-coordinate regression</strong>: for each x-position, predict the continuous y-coordinate of the ECG trace.</p>\n<hr>\n<h2>Data Strategy</h2>\n<p><strong>Problem</strong>: Kaggle GT is in millivolts, not pixels. Converting back introduces alignment errors.</p>\n<p><strong>Solution</strong>: Train primarily on <strong>synthetic ECGs</strong> with pixel-perfect ground truth, then mix in Kaggle data for domain adaptation.</p>\n<p>We extract <strong>individual rows</strong> as separate samples (4× more data), cropping each row centered on its baseline (±250px) and using only the signal region (3926px wide, no margins).</p>\n<hr>\n<h2>Architecture</h2>\n<h3>V16: Per-Lead Regression Baseline</h3>\n<pre><code>Input: [B, 3, 500, 3926] (single row crop)\n  │\n  ▼\nConvNeXt-Base Encoder (ImageNet-22k pretrained)\n  │\n  ▼\nU-Net Decoder with Skip Connections\n  │\n  ▼\nHeight Attention (softmax over y-axis)\n  │\n  ▼\n1D Conv Regression Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926] normalized y-coordinates\n</code></pre>\n<h3>V18: Cross-Row Refiner (Stacked on V16)</h3>\n<p>Takes V16's predictions and refines them by looking at <strong>all 4 rows simultaneously</strong>:</p>\n<ul>\n<li>Creates Gaussian \"guide channel\" from V16 prediction (σ=15px)</li>\n<li>EfficientNet-B0 encoder (4 channels: RGB + guide)</li>\n<li><strong>3 Cross-Row Transformer blocks</strong> with self-attention across rows</li>\n<li>Outputs small residual corrections (Tanh × learnable scale ~0.1)</li>\n</ul>\n<p><strong>Key insight</strong>: Cross-row attention implicitly learns Einthoven's Law (II = I + III) and catches inter-row inconsistencies.</p>\n<h3>V19: BiLSTM + Deformable Conv</h3>\n<pre><code>Input: [B, 3, 500, 3926]\n  │\n  ▼\nConvNeXt-Base Encoder\n  │\n  ▼\nU-Net Decoder with Deformable Conv (last 2 stages)\n  │\n  ▼\nHeight Attention\n  │\n  ▼\nBidirectional LSTM (2 layers, hidden=128)\n  │\n  ▼\nLinear Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926]\n</code></pre>\n<p><strong>Deformable Conv</strong>: Learns adaptive receptive fields for curved/warped traces.\n<strong>BiLSTM</strong>: Captures long-range temporal dependencies across the full 10-second ECG.</p>\n<hr>\n<h2>Why Ensemble Works: Complementary Failure Modes</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>V16+V18</th>\n<th>V19</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Strength</strong></td>\n<td>Cross-row consistency</td>\n<td>Long-range temporal</td>\n</tr>\n<tr>\n<td><strong>Weakness</strong></td>\n<td>Local receptive field</td>\n<td>No cross-row awareness</td>\n</tr>\n<tr>\n<td><strong>Fails on</strong></td>\n<td>Baseline drift, rhythm consistency</td>\n<td>Cross-row artifacts, lead relationships</td>\n</tr>\n</tbody>\n</table>\n<p>The architectures have <strong>uncorrelated errors</strong> (~0.3-0.4 correlation). When averaged, mistakes cancel out.</p>\n<pre><code>ensemble = 0.4 * V19 + 0.6 * V16V18  # Empirically best ratio\n</code></pre>\n<hr>\n<h2>Training</h2>\n<p><strong>Loss Function</strong>:</p>\n<pre><code>Loss = 1.0 × MaskedL1 + 0.2 × SNRLoss + 0.1 × GradientLoss\n</code></pre>\n<ul>\n<li><strong>MaskedL1</strong>: Pixel accuracy on valid regions</li>\n<li><strong>SNRLoss</strong>: Direct optimization of competition metric (<code>-10*log10(signal²/noise²)</code>)</li>\n<li><strong>GradientLoss</strong>: Smoothness constraint on first derivative</li>\n</ul>\n<p><strong>Setup</strong>:</p>\n<ul>\n<li>AdamW (lr=1e-4, weight_decay=0.01)</li>\n<li>Cosine annealing schedule</li>\n<li>Mixed precision (FP16) with gradient scaling</li>\n<li>Heavy augmentation (noise, blur, color jitter, coarse dropout)</li>\n<li>Multi-GPU via DistributedDataParallel</li>\n</ul>\n<p><strong>V18 Training</strong>: Freeze V16, train refiner to predict small residual corrections.</p>\n<hr>\n<h2>Post-Processing</h2>\n<ul>\n<li>Savitzky-Golay smoothing (window=7, polyorder=2)</li>\n<li>Einthoven's Law correction (α=0.25 error distribution)</li>\n<li>Amplitude clamping (±10 mV)</li>\n</ul>\n<hr>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Val SNR</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>V16 only</td>\n<td>~19.5 dB</td>\n<td>~19.1</td>\n</tr>\n<tr>\n<td>V16 + V18</td>\n<td>~19.7 dB</td>\n<td>~19.5</td>\n</tr>\n<tr>\n<td>V19 only</td>\n<td>~20 dB</td>\n<td>~19.5</td>\n</tr>\n<tr>\n<td><strong>Ensemble</strong></td>\n<td>—</td>\n<td><strong>~19.75</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Key Takeaways</h2>\n<ol>\n<li><strong>Regression &gt; Segmentation</strong> for sub-pixel trace extraction</li>\n<li><strong>Synthetic data</strong> with pixel-perfect GT is essential</li>\n<li><strong>Architectural diversity</strong> → uncorrelated errors → effective ensemble</li>\n<li><strong>Cross-row attention</strong> (V18) and <strong>BiLSTM</strong> (V19) complement each other perfectly</li>\n<li><strong>Direct SNR loss</strong> optimization helps</li>\n</ol>\n<p>Thanks for reading!</p>\n<h3>Special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the stage 0/1 pipeline</h3>",
      "rawMarkdown": "# 24th Place Solution: V19 + V16V18 Ensemble (19.59 SNR)\n\n## Approach: Regression, Not Segmentation\n\nWe initially tried segmentation (predicting binary trace masks), but it failed—thick traces, overlapping leads, and the need for sub-pixel accuracy made it impractical. Instead, we frame this as **per-pixel y-coordinate regression**: for each x-position, predict the continuous y-coordinate of the ECG trace.\n\n---\n\n## Data Strategy\n\n**Problem**: Kaggle GT is in millivolts, not pixels. Converting back introduces alignment errors.\n\n**Solution**: Train primarily on **synthetic ECGs** with pixel-perfect ground truth, then mix in Kaggle data for domain adaptation.\n\nWe extract **individual rows** as separate samples (4× more data), cropping each row centered on its baseline (±250px) and using only the signal region (3926px wide, no margins).\n\n---\n\n## Architecture\n\n### V16: Per-Lead Regression Baseline\n\n```\nInput: [B, 3, 500, 3926] (single row crop)\n  │\n  ▼\nConvNeXt-Base Encoder (ImageNet-22k pretrained)\n  │\n  ▼\nU-Net Decoder with Skip Connections\n  │\n  ▼\nHeight Attention (softmax over y-axis)\n  │\n  ▼\n1D Conv Regression Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926] normalized y-coordinates\n```\n\n### V18: Cross-Row Refiner (Stacked on V16)\n\nTakes V16's predictions and refines them by looking at **all 4 rows simultaneously**:\n\n- Creates Gaussian \"guide channel\" from V16 prediction (σ=15px)\n- EfficientNet-B0 encoder (4 channels: RGB + guide)\n- **3 Cross-Row Transformer blocks** with self-attention across rows\n- Outputs small residual corrections (Tanh × learnable scale ~0.1)\n\n**Key insight**: Cross-row attention implicitly learns Einthoven's Law (II = I + III) and catches inter-row inconsistencies.\n\n### V19: BiLSTM + Deformable Conv\n\n```\nInput: [B, 3, 500, 3926]\n  │\n  ▼\nConvNeXt-Base Encoder\n  │\n  ▼\nU-Net Decoder with Deformable Conv (last 2 stages)\n  │\n  ▼\nHeight Attention\n  │\n  ▼\nBidirectional LSTM (2 layers, hidden=128)\n  │\n  ▼\nLinear Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926]\n```\n\n**Deformable Conv**: Learns adaptive receptive fields for curved/warped traces.\n**BiLSTM**: Captures long-range temporal dependencies across the full 10-second ECG.\n\n---\n\n## Why Ensemble Works: Complementary Failure Modes\n\n| | V16+V18 | V19 |\n|---|---------|-----|\n| **Strength** | Cross-row consistency | Long-range temporal |\n| **Weakness** | Local receptive field | No cross-row awareness |\n| **Fails on** | Baseline drift, rhythm consistency | Cross-row artifacts, lead relationships |\n\nThe architectures have **uncorrelated errors** (~0.3-0.4 correlation). When averaged, mistakes cancel out.\n\n```python\nensemble = 0.4 * V19 + 0.6 * V16V18  # Empirically best ratio\n```\n\n---\n\n## Training\n\n**Loss Function**:\n```python\nLoss = 1.0 × MaskedL1 + 0.2 × SNRLoss + 0.1 × GradientLoss\n```\n\n- **MaskedL1**: Pixel accuracy on valid regions\n- **SNRLoss**: Direct optimization of competition metric (`-10*log10(signal²/noise²)`)\n- **GradientLoss**: Smoothness constraint on first derivative\n\n**Setup**:\n- AdamW (lr=1e-4, weight_decay=0.01)\n- Cosine annealing schedule\n- Mixed precision (FP16) with gradient scaling\n- Heavy augmentation (noise, blur, color jitter, coarse dropout)\n- Multi-GPU via DistributedDataParallel\n\n**V18 Training**: Freeze V16, train refiner to predict small residual corrections.\n\n---\n\n## Post-Processing\n\n- Savitzky-Golay smoothing (window=7, polyorder=2)\n- Einthoven's Law correction (α=0.25 error distribution)\n- Amplitude clamping (±10 mV)\n\n---\n\n## Results\n\n| Model | Val SNR | Public LB |\n|-------|---------|-----------|\n| V16 only | ~19.5 dB | ~19.1 |\n| V16 + V18 | ~19.7 dB | ~19.5 |\n| V19 only | ~20 dB | ~19.5 |\n| **Ensemble** | — | **~19.75** |\n\n---\n\n## Key Takeaways\n\n1. **Regression > Segmentation** for sub-pixel trace extraction\n2. **Synthetic data** with pixel-perfect GT is essential\n3. **Architectural diversity** → uncorrelated errors → effective ensemble\n4. **Cross-row attention** (V18) and **BiLSTM** (V19) complement each other perfectly\n5. **Direct SNR loss** optimization helps\n\n\nThanks for reading!\n\n### Special thanks to @hengck23 for the stage 0/1 pipeline",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3396784": "# 24th Place Solution: V19 + V16V18 Ensemble (19.59 SNR)\n\n## Approach: Regression, Not Segmentation\n\nWe initially tried segmentation (predicting binary trace masks), but it failed—thick traces, overlapping leads, and the need for sub-pixel accuracy made it impractical. Instead, we frame this as **per-pixel y-coordinate regression**: for each x-position, predict the continuous y-coordinate of the ECG trace.\n\n---\n\n## Data Strategy\n\n**Problem**: Kaggle GT is in millivolts, not pixels. Converting back introduces alignment errors.\n\n**Solution**: Train primarily on **synthetic ECGs** with pixel-perfect ground truth, then mix in Kaggle data for domain adaptation.\n\nWe extract **individual rows** as separate samples (4× more data), cropping each row centered on its baseline (±250px) and using only the signal region (3926px wide, no margins).\n\n---\n\n## Architecture\n\n### V16: Per-Lead Regression Baseline\n\n```\nInput: [B, 3, 500, 3926] (single row crop)\n  │\n  ▼\nConvNeXt-Base Encoder (ImageNet-22k pretrained)\n  │\n  ▼\nU-Net Decoder with Skip Connections\n  │\n  ▼\nHeight Attention (softmax over y-axis)\n  │\n  ▼\n1D Conv Regression Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926] normalized y-coordinates\n```\n\n### V18: Cross-Row Refiner (Stacked on V16)\n\nTakes V16's predictions and refines them by looking at **all 4 rows simultaneously**:\n\n- Creates Gaussian \"guide channel\" from V16 prediction (σ=15px)\n- EfficientNet-B0 encoder (4 channels: RGB + guide)\n- **3 Cross-Row Transformer blocks** with self-attention across rows\n- Outputs small residual corrections (Tanh × learnable scale ~0.1)\n\n**Key insight**: Cross-row attention implicitly learns Einthoven's Law (II = I + III) and catches inter-row inconsistencies.\n\n### V19: BiLSTM + Deformable Conv\n\n```\nInput: [B, 3, 500, 3926]\n  │\n  ▼\nConvNeXt-Base Encoder\n  │\n  ▼\nU-Net Decoder with Deformable Conv (last 2 stages)\n  │\n  ▼\nHeight Attention\n  │\n  ▼\nBidirectional LSTM (2 layers, hidden=128)\n  │\n  ▼\nLinear Head → Sigmoid\n  │\n  ▼\nOutput: [B, 3926]\n```\n\n**Deformable Conv**: Learns adaptive receptive fields for curved/warped traces.\n**BiLSTM**: Captures long-range temporal dependencies across the full 10-second ECG.\n\n---\n\n## Why Ensemble Works: Complementary Failure Modes\n\n| | V16+V18 | V19 |\n|---|---------|-----|\n| **Strength** | Cross-row consistency | Long-range temporal |\n| **Weakness** | Local receptive field | No cross-row awareness |\n| **Fails on** | Baseline drift, rhythm consistency | Cross-row artifacts, lead relationships |\n\nThe architectures have **uncorrelated errors** (~0.3-0.4 correlation). When averaged, mistakes cancel out.\n\n```python\nensemble = 0.4 * V19 + 0.6 * V16V18  # Empirically best ratio\n```\n\n---\n\n## Training\n\n**Loss Function**:\n```python\nLoss = 1.0 × MaskedL1 + 0.2 × SNRLoss + 0.1 × GradientLoss\n```\n\n- **MaskedL1**: Pixel accuracy on valid regions\n- **SNRLoss**: Direct optimization of competition metric (`-10*log10(signal²/noise²)`)\n- **GradientLoss**: Smoothness constraint on first derivative\n\n**Setup**:\n- AdamW (lr=1e-4, weight_decay=0.01)\n- Cosine annealing schedule\n- Mixed precision (FP16) with gradient scaling\n- Heavy augmentation (noise, blur, color jitter, coarse dropout)\n- Multi-GPU via DistributedDataParallel\n\n**V18 Training**: Freeze V16, train refiner to predict small residual corrections.\n\n---\n\n## Post-Processing\n\n- Savitzky-Golay smoothing (window=7, polyorder=2)\n- Einthoven's Law correction (α=0.25 error distribution)\n- Amplitude clamping (±10 mV)\n\n---\n\n## Results\n\n| Model | Val SNR | Public LB |\n|-------|---------|-----------|\n| V16 only | ~19.5 dB | ~19.1 |\n| V16 + V18 | ~19.7 dB | ~19.5 |\n| V19 only | ~20 dB | ~19.5 |\n| **Ensemble** | — | **~19.75** |\n\n---\n\n## Key Takeaways\n\n1. **Regression > Segmentation** for sub-pixel trace extraction\n2. **Synthetic data** with pixel-perfect GT is essential\n3. **Architectural diversity** → uncorrelated errors → effective ensemble\n4. **Cross-row attention** (V18) and **BiLSTM** (V19) complement each other perfectly\n5. **Direct SNR loss** optimization helps\n\n\nThanks for reading!\n\n### Special thanks to @hengck23 for the stage 0/1 pipeline"
  }
}