{
  "id": 669556,
  "title": "26th Place Solution",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/669556",
  "author_name": "Boredom",
  "post_date": "2026-01-23T03:15:46.649000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to PhysioNet and Kaggle for hosting such a fantastic competition. Huge thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">hengck23</a>, <a href=\"https://www.kaggle.com/wasupandceacar\" target=\"_blank\">wasupandceacar</a>, and others for their selfless open-source contributions.  We only started putting serious effort into this during the final week, so we focused on optimizing the top-scoring open-source solutions.  We are thrilled to have achieved this rank and are eagerly looking forward to seeing the winning solutions!</p>\n<hr>\n<h3><strong>Overview</strong>：</h3>\n<p>Our solution is primarily based on hengck23's open-source code.     We focused our optimization on Stage 1 and Stage 2 (we left Stage 0 as is, because we ran out of time).     The biggest improvement came from the two-stage training for Stage 2, which combines segmentation and signal extraction.  This allowed the raw signal data to be directly involved in the training process—an idea that had already been suggested in hengck23's discussion thread (though I didn't quite grasp it from the diagrams at first…).</p>\n<hr>\n<h1><strong>Stage 1</strong></h1>\n<h3><strong>Data Preparation:</strong></h3>\n<p>We utilized a modified version of ecg-image-kit based on several external datasets (e.g., PTB, PTB-XL, Georgia, CPSC-2018, etc.) to randomly generate approximately 10k raw ECG images.  Combined with the official dataset, we processed these through our Stage 0 and Stage 1 inference pipelines.  The resulting outputs served as pre-annotations, which were then manually refined using our custom-built annotation software.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F1de6cec56fd2ddb7f54f24a216081038%2F20260123-102527.jpg?generation=1769135930758823&amp;alt=media\" alt=\"\"></p>\n<h3><strong>Model:</strong></h3>\n<p>Standard UNet. Backbone: EfficientNet-V2-S, ResNet34. (AdamW (lr: 1e-4, wd: 2e-5).)</p>\n<h3><strong>Augmentation:</strong></h3>\n<p>Various Albumentations techniques including Noise, Blur, ToGray, and CLAHE.</p>\n<h3><strong>Simple Weight Averaging (SWA):</strong></h3>\n<p>During training, we maintained the top three models based on validation loss and scores, then performed offline weight averaging.</p>\n<h3><strong>Inference Post-processing:</strong></h3>\n<ol>\n<li>Optimized output_to_predict: When assigning values within the 44/57 range for each point, the logic was enhanced to consider not just the current line position but also a voting value from a surrounding 3x3 window.  </li>\n<li>Grid Completion: Missing grid points from the segmentation were interpolated and filled based on directional vectors.</li>\n</ol>\n<hr>\n<h1><strong>Stage 2</strong></h1>\n<h3><strong>Data Preparation:</strong></h3>\n<p>Generated 20k ECG images (with lead masks) using a modified ecg-image-kit, combined with the official dataset.       All training data were pre-processed through our optimized Stage 0 + Stage 1 inference pipeline to obtain rectified, undistorted images.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F646407128f6e358008d8477e681ba63b%2F_20260122122451_7615_6.png?generation=1769137987162403&amp;alt=media\" alt=\"\"></p>\n<h3><strong>Cross-Validation:</strong></h3>\n<p>Performed 5-fold CV on the official dataset using patient IDs as groups.       Synthetic data were included only in the training set of each fold and excluded from the validation set.</p>\n<h3><strong>Two-Stage Training:</strong></h3>\n<ul>\n<li>Phase 1: Fine-tuned wasupandceacar’s best open-source segmentation model using synthetic data and masks.     </li>\n<li>Phase 2: Used Phase 1 weights as the pre-trained backbone.       We implemented a Differentiable Soft-Argmax layer after the segmentation head to extract signals from the probability maps (logits).       Using fixed cropping parameters and zero-baseline coordinates, the predicted pixel-level signals were converted into physical signals to calculate L1 + L2 loss against the ground truth (GT) signals.</li>\n</ul>\n<h3><strong>Model:</strong></h3>\n<p>wasupandceacar’s ResNet34-UNet. imagesize:1696 x 4352</p>\n<h3><strong>Segmentation Phase Augmentation:</strong></h3>\n<p>Various Albumentations techniques including Noise, Blur, ToGray, and CLAHE.</p>\n<h3><strong>Signal Phase Augmentation:</strong></h3>\n<p>Combined segmentation-phase augmentations with custom-designed effects: simulated black/yellow stains, lens distortion, paper creases, partial shadows, and scanner noise.</p>\n<h3><strong>Training Configuration:</strong></h3>\n<p>4x RTX 4090 (DDP), EMA (0.992), AdamW, Warmup + Cosine Scheduler, 2e-5 learning rate.       Data sampling for each epoch: 80% official data and 20% synthetic data.</p>\n<h3><strong>Inference Configuration:</strong></h3>\n<ol>\n<li>Ensemble Strategy: Integrated three best models (varying in augmentations, learning rates, and folds). We performed simple averaging at the probability level (segmentation logits) before passing it through the same differentiable post-processing used in training.       </li>\n<li>Legacy Logic: Retained the open-source logic for ECG type classification and its corresponding pre-processing.       Despite imperfect classification accuracy, it provided a 0.2 LB boost.</li>\n</ol>",
  "messages": [
    {
      "id": 3395467,
      "postDate": "2026-01-23T03:15:46.650Z",
      "content": "<p>Thanks to PhysioNet and Kaggle for hosting such a fantastic competition. Huge thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">hengck23</a>, <a href=\"https://www.kaggle.com/wasupandceacar\" target=\"_blank\">wasupandceacar</a>, and others for their selfless open-source contributions.  We only started putting serious effort into this during the final week, so we focused on optimizing the top-scoring open-source solutions.  We are thrilled to have achieved this rank and are eagerly looking forward to seeing the winning solutions!</p>\n<hr>\n<h3><strong>Overview</strong>：</h3>\n<p>Our solution is primarily based on hengck23's open-source code.     We focused our optimization on Stage 1 and Stage 2 (we left Stage 0 as is, because we ran out of time).     The biggest improvement came from the two-stage training for Stage 2, which combines segmentation and signal extraction.  This allowed the raw signal data to be directly involved in the training process—an idea that had already been suggested in hengck23's discussion thread (though I didn't quite grasp it from the diagrams at first…).</p>\n<hr>\n<h1><strong>Stage 1</strong></h1>\n<h3><strong>Data Preparation:</strong></h3>\n<p>We utilized a modified version of ecg-image-kit based on several external datasets (e.g., PTB, PTB-XL, Georgia, CPSC-2018, etc.) to randomly generate approximately 10k raw ECG images.  Combined with the official dataset, we processed these through our Stage 0 and Stage 1 inference pipelines.  The resulting outputs served as pre-annotations, which were then manually refined using our custom-built annotation software.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F1de6cec56fd2ddb7f54f24a216081038%2F20260123-102527.jpg?generation=1769135930758823&amp;alt=media\" alt=\"\"></p>\n<h3><strong>Model:</strong></h3>\n<p>Standard UNet. Backbone: EfficientNet-V2-S, ResNet34. (AdamW (lr: 1e-4, wd: 2e-5).)</p>\n<h3><strong>Augmentation:</strong></h3>\n<p>Various Albumentations techniques including Noise, Blur, ToGray, and CLAHE.</p>\n<h3><strong>Simple Weight Averaging (SWA):</strong></h3>\n<p>During training, we maintained the top three models based on validation loss and scores, then performed offline weight averaging.</p>\n<h3><strong>Inference Post-processing:</strong></h3>\n<ol>\n<li>Optimized output_to_predict: When assigning values within the 44/57 range for each point, the logic was enhanced to consider not just the current line position but also a voting value from a surrounding 3x3 window.  </li>\n<li>Grid Completion: Missing grid points from the segmentation were interpolated and filled based on directional vectors.</li>\n</ol>\n<hr>\n<h1><strong>Stage 2</strong></h1>\n<h3><strong>Data Preparation:</strong></h3>\n<p>Generated 20k ECG images (with lead masks) using a modified ecg-image-kit, combined with the official dataset.       All training data were pre-processed through our optimized Stage 0 + Stage 1 inference pipeline to obtain rectified, undistorted images.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F646407128f6e358008d8477e681ba63b%2F_20260122122451_7615_6.png?generation=1769137987162403&amp;alt=media\" alt=\"\"></p>\n<h3><strong>Cross-Validation:</strong></h3>\n<p>Performed 5-fold CV on the official dataset using patient IDs as groups.       Synthetic data were included only in the training set of each fold and excluded from the validation set.</p>\n<h3><strong>Two-Stage Training:</strong></h3>\n<ul>\n<li>Phase 1: Fine-tuned wasupandceacar’s best open-source segmentation model using synthetic data and masks.     </li>\n<li>Phase 2: Used Phase 1 weights as the pre-trained backbone.       We implemented a Differentiable Soft-Argmax layer after the segmentation head to extract signals from the probability maps (logits).       Using fixed cropping parameters and zero-baseline coordinates, the predicted pixel-level signals were converted into physical signals to calculate L1 + L2 loss against the ground truth (GT) signals.</li>\n</ul>\n<h3><strong>Model:</strong></h3>\n<p>wasupandceacar’s ResNet34-UNet. imagesize:1696 x 4352</p>\n<h3><strong>Segmentation Phase Augmentation:</strong></h3>\n<p>Various Albumentations techniques including Noise, Blur, ToGray, and CLAHE.</p>\n<h3><strong>Signal Phase Augmentation:</strong></h3>\n<p>Combined segmentation-phase augmentations with custom-designed effects: simulated black/yellow stains, lens distortion, paper creases, partial shadows, and scanner noise.</p>\n<h3><strong>Training Configuration:</strong></h3>\n<p>4x RTX 4090 (DDP), EMA (0.992), AdamW, Warmup + Cosine Scheduler, 2e-5 learning rate.       Data sampling for each epoch: 80% official data and 20% synthetic data.</p>\n<h3><strong>Inference Configuration:</strong></h3>\n<ol>\n<li>Ensemble Strategy: Integrated three best models (varying in augmentations, learning rates, and folds). We performed simple averaging at the probability level (segmentation logits) before passing it through the same differentiable post-processing used in training.       </li>\n<li>Legacy Logic: Retained the open-source logic for ECG type classification and its corresponding pre-processing.       Despite imperfect classification accuracy, it provided a 0.2 LB boost.</li>\n</ol>",
      "rawMarkdown": "Thanks to PhysioNet and Kaggle for hosting such a fantastic competition. Huge thanks to [hengck23](https://www.kaggle.com/hengck23), [wasupandceacar](https://www.kaggle.com/wasupandceacar), and others for their selfless open-source contributions.  We only started putting serious effort into this during the final week, so we focused on optimizing the top-scoring open-source solutions.  We are thrilled to have achieved this rank and are eagerly looking forward to seeing the winning solutions!\n- --\n### **Overview**：\nOur solution is primarily based on hengck23's open-source code.     We focused our optimization on Stage 1 and Stage 2 (we left Stage 0 as is, because we ran out of time).     The biggest improvement came from the two-stage training for Stage 2, which combines segmentation and signal extraction.  This allowed the raw signal data to be directly involved in the training process—an idea that had already been suggested in hengck23's discussion thread (though I didn't quite grasp it from the diagrams at first...).\n- --\n# **Stage 1**\n\n### **Data Preparation:**\nWe utilized a modified version of ecg-image-kit based on several external datasets (e.g., PTB, PTB-XL, Georgia, CPSC-2018, etc.) to randomly generate approximately 10k raw ECG images.  Combined with the official dataset, we processed these through our Stage 0 and Stage 1 inference pipelines.  The resulting outputs served as pre-annotations, which were then manually refined using our custom-built annotation software.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F1de6cec56fd2ddb7f54f24a216081038%2F20260123-102527.jpg?generation=1769135930758823&alt=media)\n### **Model:**   \nStandard UNet. Backbone: EfficientNet-V2-S, ResNet34. (AdamW (lr: 1e-4, wd: 2e-5).)\n\n### **Augmentation:**   \nVarious Albumentations techniques including Noise, Blur, ToGray, and CLAHE.\n\n### **Simple Weight Averaging (SWA):**  \nDuring training, we maintained the top three models based on validation loss and scores, then performed offline weight averaging.\n\n### **Inference Post-processing:**  \n1.  Optimized output_to_predict: When assigning values within the 44/57 range for each point, the logic was enhanced to consider not just the current line position but also a voting value from a surrounding 3x3 window.  \n2.  Grid Completion: Missing grid points from the segmentation were interpolated and filled based on directional vectors.\n\n- --\n# **Stage 2**\n\n### **Data Preparation:**  \nGenerated 20k ECG images (with lead masks) using a modified ecg-image-kit, combined with the official dataset.       All training data were pre-processed through our optimized Stage 0 + Stage 1 inference pipeline to obtain rectified, undistorted images.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F646407128f6e358008d8477e681ba63b%2F_20260122122451_7615_6.png?generation=1769137987162403&alt=media)\n### **Cross-Validation:** \nPerformed 5-fold CV on the official dataset using patient IDs as groups.       Synthetic data were included only in the training set of each fold and excluded from the validation set.\n\n### **Two-Stage Training:** \n* Phase 1: Fine-tuned wasupandceacar’s best open-source segmentation model using synthetic data and masks.     \n* Phase 2: Used Phase 1 weights as the pre-trained backbone.       We implemented a Differentiable Soft-Argmax layer after the segmentation head to extract signals from the probability maps (logits).       Using fixed cropping parameters and zero-baseline coordinates, the predicted pixel-level signals were converted into physical signals to calculate L1 + L2 loss against the ground truth (GT) signals.\n\n### **Model:**  \nwasupandceacar’s ResNet34-UNet. imagesize:1696 x 4352\n\n### **Segmentation Phase Augmentation:**   \nVarious Albumentations techniques including Noise, Blur, ToGray, and CLAHE.\n\n### **Signal Phase Augmentation:**  \nCombined segmentation-phase augmentations with custom-designed effects: simulated black/yellow stains, lens distortion, paper creases, partial shadows, and scanner noise.\n\n### **Training Configuration:**  \n4x RTX 4090 (DDP), EMA (0.992), AdamW, Warmup + Cosine Scheduler, 2e-5 learning rate.       Data sampling for each epoch: 80% official data and 20% synthetic data.\n\n### **Inference Configuration:**\n1. Ensemble Strategy: Integrated three best models (varying in augmentations, learning rates, and folds). We performed simple averaging at the probability level (segmentation logits) before passing it through the same differentiable post-processing used in training.       \n2. Legacy Logic: Retained the open-source logic for ECG type classification and its corresponding pre-processing.       Despite imperfect classification accuracy, it provided a 0.2 LB boost.",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3395467": "Thanks to PhysioNet and Kaggle for hosting such a fantastic competition. Huge thanks to [hengck23](https://www.kaggle.com/hengck23), [wasupandceacar](https://www.kaggle.com/wasupandceacar), and others for their selfless open-source contributions.  We only started putting serious effort into this during the final week, so we focused on optimizing the top-scoring open-source solutions.  We are thrilled to have achieved this rank and are eagerly looking forward to seeing the winning solutions!\n- --\n### **Overview**：\nOur solution is primarily based on hengck23's open-source code.     We focused our optimization on Stage 1 and Stage 2 (we left Stage 0 as is, because we ran out of time).     The biggest improvement came from the two-stage training for Stage 2, which combines segmentation and signal extraction.  This allowed the raw signal data to be directly involved in the training process—an idea that had already been suggested in hengck23's discussion thread (though I didn't quite grasp it from the diagrams at first...).\n- --\n# **Stage 1**\n\n### **Data Preparation:**\nWe utilized a modified version of ecg-image-kit based on several external datasets (e.g., PTB, PTB-XL, Georgia, CPSC-2018, etc.) to randomly generate approximately 10k raw ECG images.  Combined with the official dataset, we processed these through our Stage 0 and Stage 1 inference pipelines.  The resulting outputs served as pre-annotations, which were then manually refined using our custom-built annotation software.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F1de6cec56fd2ddb7f54f24a216081038%2F20260123-102527.jpg?generation=1769135930758823&alt=media)\n### **Model:**   \nStandard UNet. Backbone: EfficientNet-V2-S, ResNet34. (AdamW (lr: 1e-4, wd: 2e-5).)\n\n### **Augmentation:**   \nVarious Albumentations techniques including Noise, Blur, ToGray, and CLAHE.\n\n### **Simple Weight Averaging (SWA):**  \nDuring training, we maintained the top three models based on validation loss and scores, then performed offline weight averaging.\n\n### **Inference Post-processing:**  \n1.  Optimized output_to_predict: When assigning values within the 44/57 range for each point, the logic was enhanced to consider not just the current line position but also a voting value from a surrounding 3x3 window.  \n2.  Grid Completion: Missing grid points from the segmentation were interpolated and filled based on directional vectors.\n\n- --\n# **Stage 2**\n\n### **Data Preparation:**  \nGenerated 20k ECG images (with lead masks) using a modified ecg-image-kit, combined with the official dataset.       All training data were pre-processed through our optimized Stage 0 + Stage 1 inference pipeline to obtain rectified, undistorted images.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F646407128f6e358008d8477e681ba63b%2F_20260122122451_7615_6.png?generation=1769137987162403&alt=media)\n### **Cross-Validation:** \nPerformed 5-fold CV on the official dataset using patient IDs as groups.       Synthetic data were included only in the training set of each fold and excluded from the validation set.\n\n### **Two-Stage Training:** \n* Phase 1: Fine-tuned wasupandceacar’s best open-source segmentation model using synthetic data and masks.     \n* Phase 2: Used Phase 1 weights as the pre-trained backbone.       We implemented a Differentiable Soft-Argmax layer after the segmentation head to extract signals from the probability maps (logits).       Using fixed cropping parameters and zero-baseline coordinates, the predicted pixel-level signals were converted into physical signals to calculate L1 + L2 loss against the ground truth (GT) signals.\n\n### **Model:**  \nwasupandceacar’s ResNet34-UNet. imagesize:1696 x 4352\n\n### **Segmentation Phase Augmentation:**   \nVarious Albumentations techniques including Noise, Blur, ToGray, and CLAHE.\n\n### **Signal Phase Augmentation:**  \nCombined segmentation-phase augmentations with custom-designed effects: simulated black/yellow stains, lens distortion, paper creases, partial shadows, and scanner noise.\n\n### **Training Configuration:**  \n4x RTX 4090 (DDP), EMA (0.992), AdamW, Warmup + Cosine Scheduler, 2e-5 learning rate.       Data sampling for each epoch: 80% official data and 20% synthetic data.\n\n### **Inference Configuration:**\n1. Ensemble Strategy: Integrated three best models (varying in augmentations, learning rates, and folds). We performed simple averaging at the probability level (segmentation logits) before passing it through the same differentiable post-processing used in training.       \n2. Legacy Logic: Retained the open-source logic for ECG type classification and its corresponding pre-processing.       Despite imperfect classification accuracy, it provided a 0.2 LB boost."
  }
}