{
  "id": 670169,
  "title": "9th Place Solution",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/670169",
  "author_name": "YumeNeko",
  "post_date": "2026-01-26T13:59:12.103000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to pay tribute to all the participants who worked on this competition.\nI would also like to thank the hosts for organizing this interesting task competition.</p>\n<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F389d63e6a6f3b3a7103b645812f652b0%2Foverview.jpg?generation=1769435859268811&amp;alt=media\" alt=\"\">\nMy solution consists of a two-stage pipeline: image rectify and signal reconstruction.</p>\n<p>For image rectify, I first apply coarse alignment using a homography transformation based on image matching, then detect grid point coordinates and perform precise alignment using Piecewise Affine transformation.</p>\n<p>For signal reconstruction, I extend the decoder of a UNet with a signal reconstruction module that directly outputs signals from images, and train the model to optimize SNR.</p>\n<h1>Pipeline</h1>\n<h2>1. Image rectify</h2>\n<h3>1.a Image Type Classification</h3>\n<ul>\n<li>Since performance is more stable when some image types (006) are handled separately, images are first classified by type.</li>\n<li>Only competition data is used for training, and a simple classification model with convnext_large_384_in22ft1k as the backbone is trained.</li>\n<li>There is nothing particularly special here, but the model achieves about 99.9% accuracy in CV.</li>\n</ul>\n<h2>1.b Image Matching</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F329c1c81d9b9cf8db9c62434e5b763b7%2Fimage_matching.jpg?generation=1769435872559085&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Feature point matching with a template image is performed using ALIKED + LightGlue, followed by coarse alignment via a homography transformation estimated with RANSAC.<ul>\n<li>The template image is created by averaging competition images of image type 001.</li>\n<li>Publicly available pretrained weights are used as-is for both ALIKED and LightGlue.</li></ul></li>\n<li>To simultaneously correct rotation, the input image is rotated by 0°, 90°, 180°, 270°, matching is performed for each rotation, and the rotation with the largest number of matches is selected.</li>\n<li>Although some distortion remains at this stage, almost all images are aligned to the same orientation and composition as the template (001), which stabilizes downstream training and inference.</li>\n</ul>\n<h2>1.c Grid Detection</h2>\n<ul>\n<li>A model is built to detect grid point coordinates from the projectively transformed images.<ul>\n<li>A UNet with ResNeSt-14d as the encoder is used.</li></ul></li>\n<li>Only about 22,000 synthetic images are used for training.<ul>\n<li>Since image matching already roughly normalizes the input images, augmentation on synthetic data alone is sufficient to generalize to real data.</li>\n<li>No annotation on real data is performed, significantly reducing annotation cost.</li></ul></li>\n<li>For image type 006, false detections sometimes occur due to moiré patterns, so grid detection is applied after moiré removal using <a href=\"https://github.com/CVMI-Lab/UHDM\" target=\"_blank\">UHDM</a>.</li>\n</ul>\n<h2>1.d Precise Alignment</h2>\n<ul>\n<li>Finally, precise alignment is performed using Piecewise Affine transformation based on the outputs of image matching and grid detection.</li>\n</ul>\n<ol>\n<li>From the correspondence points obtained by image matching, regions with small reprojection error and minimal distortion are selected as initial regions. Detected grid points in these regions are matched to ideal grid points using the Hungarian algorithm.</li>\n<li>Starting from the initial correspondences, assuming a grid spacing of approximately 40 px, correspondences are incrementally expanded to neighboring grid points using breadth-first search in the up/down/left/right directions.</li>\n<li>The final set of correspondences is used as control points to apply a Piecewise Affine transformation to correct distortion over the entire image.</li>\n</ol>\n<h2>2. Signal Reconstruction</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F18cf7c0c16852c73f0e4dc6f2f6b87fb%2Fnetwork.jpg?generation=1769435905967605&amp;alt=media\" alt=\"\"></p>\n<h3>Architecture</h3>\n<ul>\n<li>Based on a UNet with a ResNeSt-14d backbone, an architecture is designed by adding a dedicated module to directly estimate signals from the decoder output.<ul>\n<li>Feature maps from the UNet decoder are expanded 6× along the temporal axis (W direction), then concatenated with sampling frequency (fs) and lead type as features.</li>\n<li>After concatenation, convolution along the W direction is applied to output logits representing “signal-likeness” along the y direction for each x column.</li>\n<li>From the y-direction logits, two soft expectations are computed: one biased toward the top edge and one toward the bottom edge, and a gate function predicts (0–1) which to use for each x.</li>\n<li>The final estimated y (pixel coordinates) is converted to a signal value in mV based on ECG drawing geometry (paper size / resolution / row layout).</li></ul></li>\n</ul>\n<h3>Inputs, Outputs, and Loss Functions</h3>\n<ul>\n<li><p>Inputs</p>\n<ul>\n<li>Cropped images for each lead<ul>\n<li>Fixed crop size: H×W = 600×491</li>\n<li>RGB images with positional encoding added for x and y directions, resulting in 5-channel input</li>\n<li>The long lead (II) is split into 4 parts to match the temporal length of other leads</li></ul></li>\n<li>Target sampling frequency (fs)</li>\n<li>Lead type<ul>\n<li>Adding target fs and lead type as input features improved CV performance by about +1.8 dB, but unfortunately did not yield a clear improvement on the LB.</li></ul></li></ul></li>\n<li><p>Outputs</p>\n<ul>\n<li>Segmentation maps<ul>\n<li>Three classes: background, target signal to be reconstructed, and non-target signals</li></ul></li>\n<li>Reconstructed signal waveform</li></ul></li>\n<li><p>Loss</p>\n<ul>\n<li>Segmentation loss: Dice + CE, weighted 0.5 : 0.5</li>\n<li>Waveform loss: after interpolating the predicted waveform to match the GT length, an SNR-based loss is used as the main loss, with L1 loss added as an auxiliary term</li>\n<li>The final loss is a weighted sum, with emphasis on waveform reconstruction (SNR loss)</li></ul></li>\n</ul>\n<h3>Multi-stage Training</h3>\n<p>Training is performed in three stages, gradually switching datasets and augmentation strategies.</p>\n<p>This multi-stage training yields approximately +1.0 dB from 1st-stage pretraining and an additional +0.2 dB from image-type-specific tuning in the 3rd stage, consistently on both Public and Private sets.</p>\n<h4>1st Stage Training (Pretraining with Synthetic Data)</h4>\n<ul>\n<li>Trained using about 21,000 synthetic images generated from PTB-XL (500 Hz), excluding overlaps with competition data.</li>\n<li>Image-only augmentations include blur, brightness changes, grid coordinate shifts up to 1 pixel via Piecewise Affine, and custom augmentations such as DirtPatch / CreaseWrinkles.</li>\n<li>Image-and-signal joint augmentations include random segment dropout and vertical/horizontal flips.</li>\n<li>GT signals are randomly resampled using scipy.resample_poly to one of\n[250, 256, 512, 500, 1000, 1025].</li>\n<li>Trained for 15 epochs with lr = 4e-4.</li>\n</ul>\n<h4>2nd Stage Training (Main Training)</h4>\n<ul>\n<li>Initialized with weights from the 1st stage and trained using competition data only.</li>\n<li>No image-only augmentation is applied; only image+signal augmentations from the 1st stage are used.</li>\n<li>Trained for 30 epochs with lr = 1e-4.</li>\n</ul>\n<h4>3rd Stage Training (Image-Type-Specific Fine-tuning)</h4>\n<ul>\n<li>Initialized with weights from the 2nd stage and further trained using only image type 006 data.</li>\n<li>Augmentation settings are the same as in the 2nd stage.</li>\n<li>Trained for 50 epochs with lr = 1e-4.</li>\n<li>Although CV performance improved for other image types as well, LB improvements were unstable, so the final submission uses the 3rd-stage model only for image type 006, and the 2nd-stage model for others.</li>\n</ul>\n<h3>TTA</h3>\n<ul>\n<li>During inference, predictions are made on the original image plus left-right flip, top-bottom flip, and both flips, totaling 4 patterns, and the results are averaged.</li>\n<li>This TTA consistently provides about +0.15–0.2 dB improvement on both Public and Private sets.</li>\n</ul>\n<h1>Score</h1>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>synthetic_data pretrain (1st stage train)</th>\n<th>006 fine tuning (3rd stage train)</th>\n<th>TTA</th>\n<th>full data train</th>\n<th>use target_fs</th>\n<th>use_lead_id</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td>20.17</td>\n<td>20.11</td>\n</tr>\n<tr>\n<td>2</td>\n<td></td>\n<td></td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td></td>\n<td>20.32</td>\n<td>20.26</td>\n</tr>\n<tr>\n<td>3</td>\n<td>✅</td>\n<td></td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td></td>\n<td>21.43</td>\n<td>21.30</td>\n</tr>\n<tr>\n<td>4</td>\n<td>✅</td>\n<td></td>\n<td>✅</td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td>21.54</td>\n<td>21.38</td>\n</tr>\n<tr>\n<td>5</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td>21.80</td>\n<td>21.60</td>\n</tr>\n<tr>\n<td>6</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td></td>\n<td>21.66</td>\n<td>21.56</td>\n</tr>\n<tr>\n<td>7</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>21.81</td>\n<td>21.71</td>\n</tr>\n</tbody>\n</table>\n<p><strong>final submission</strong><br>\nEnsemble #5 + #6 + #7</p>\n<ul>\n<li>Public: 22.04  </li>\n<li>Private: 21.92</li>\n</ul>",
  "messages": [
    {
      "id": 3397112,
      "postDate": "2026-01-26T13:59:12.103Z",
      "content": "<p>First of all, I would like to pay tribute to all the participants who worked on this competition.\nI would also like to thank the hosts for organizing this interesting task competition.</p>\n<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F389d63e6a6f3b3a7103b645812f652b0%2Foverview.jpg?generation=1769435859268811&amp;alt=media\" alt=\"\">\nMy solution consists of a two-stage pipeline: image rectify and signal reconstruction.</p>\n<p>For image rectify, I first apply coarse alignment using a homography transformation based on image matching, then detect grid point coordinates and perform precise alignment using Piecewise Affine transformation.</p>\n<p>For signal reconstruction, I extend the decoder of a UNet with a signal reconstruction module that directly outputs signals from images, and train the model to optimize SNR.</p>\n<h1>Pipeline</h1>\n<h2>1. Image rectify</h2>\n<h3>1.a Image Type Classification</h3>\n<ul>\n<li>Since performance is more stable when some image types (006) are handled separately, images are first classified by type.</li>\n<li>Only competition data is used for training, and a simple classification model with convnext_large_384_in22ft1k as the backbone is trained.</li>\n<li>There is nothing particularly special here, but the model achieves about 99.9% accuracy in CV.</li>\n</ul>\n<h2>1.b Image Matching</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F329c1c81d9b9cf8db9c62434e5b763b7%2Fimage_matching.jpg?generation=1769435872559085&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Feature point matching with a template image is performed using ALIKED + LightGlue, followed by coarse alignment via a homography transformation estimated with RANSAC.<ul>\n<li>The template image is created by averaging competition images of image type 001.</li>\n<li>Publicly available pretrained weights are used as-is for both ALIKED and LightGlue.</li></ul></li>\n<li>To simultaneously correct rotation, the input image is rotated by 0°, 90°, 180°, 270°, matching is performed for each rotation, and the rotation with the largest number of matches is selected.</li>\n<li>Although some distortion remains at this stage, almost all images are aligned to the same orientation and composition as the template (001), which stabilizes downstream training and inference.</li>\n</ul>\n<h2>1.c Grid Detection</h2>\n<ul>\n<li>A model is built to detect grid point coordinates from the projectively transformed images.<ul>\n<li>A UNet with ResNeSt-14d as the encoder is used.</li></ul></li>\n<li>Only about 22,000 synthetic images are used for training.<ul>\n<li>Since image matching already roughly normalizes the input images, augmentation on synthetic data alone is sufficient to generalize to real data.</li>\n<li>No annotation on real data is performed, significantly reducing annotation cost.</li></ul></li>\n<li>For image type 006, false detections sometimes occur due to moiré patterns, so grid detection is applied after moiré removal using <a href=\"https://github.com/CVMI-Lab/UHDM\" target=\"_blank\">UHDM</a>.</li>\n</ul>\n<h2>1.d Precise Alignment</h2>\n<ul>\n<li>Finally, precise alignment is performed using Piecewise Affine transformation based on the outputs of image matching and grid detection.</li>\n</ul>\n<ol>\n<li>From the correspondence points obtained by image matching, regions with small reprojection error and minimal distortion are selected as initial regions. Detected grid points in these regions are matched to ideal grid points using the Hungarian algorithm.</li>\n<li>Starting from the initial correspondences, assuming a grid spacing of approximately 40 px, correspondences are incrementally expanded to neighboring grid points using breadth-first search in the up/down/left/right directions.</li>\n<li>The final set of correspondences is used as control points to apply a Piecewise Affine transformation to correct distortion over the entire image.</li>\n</ol>\n<h2>2. Signal Reconstruction</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F18cf7c0c16852c73f0e4dc6f2f6b87fb%2Fnetwork.jpg?generation=1769435905967605&amp;alt=media\" alt=\"\"></p>\n<h3>Architecture</h3>\n<ul>\n<li>Based on a UNet with a ResNeSt-14d backbone, an architecture is designed by adding a dedicated module to directly estimate signals from the decoder output.<ul>\n<li>Feature maps from the UNet decoder are expanded 6× along the temporal axis (W direction), then concatenated with sampling frequency (fs) and lead type as features.</li>\n<li>After concatenation, convolution along the W direction is applied to output logits representing “signal-likeness” along the y direction for each x column.</li>\n<li>From the y-direction logits, two soft expectations are computed: one biased toward the top edge and one toward the bottom edge, and a gate function predicts (0–1) which to use for each x.</li>\n<li>The final estimated y (pixel coordinates) is converted to a signal value in mV based on ECG drawing geometry (paper size / resolution / row layout).</li></ul></li>\n</ul>\n<h3>Inputs, Outputs, and Loss Functions</h3>\n<ul>\n<li><p>Inputs</p>\n<ul>\n<li>Cropped images for each lead<ul>\n<li>Fixed crop size: H×W = 600×491</li>\n<li>RGB images with positional encoding added for x and y directions, resulting in 5-channel input</li>\n<li>The long lead (II) is split into 4 parts to match the temporal length of other leads</li></ul></li>\n<li>Target sampling frequency (fs)</li>\n<li>Lead type<ul>\n<li>Adding target fs and lead type as input features improved CV performance by about +1.8 dB, but unfortunately did not yield a clear improvement on the LB.</li></ul></li></ul></li>\n<li><p>Outputs</p>\n<ul>\n<li>Segmentation maps<ul>\n<li>Three classes: background, target signal to be reconstructed, and non-target signals</li></ul></li>\n<li>Reconstructed signal waveform</li></ul></li>\n<li><p>Loss</p>\n<ul>\n<li>Segmentation loss: Dice + CE, weighted 0.5 : 0.5</li>\n<li>Waveform loss: after interpolating the predicted waveform to match the GT length, an SNR-based loss is used as the main loss, with L1 loss added as an auxiliary term</li>\n<li>The final loss is a weighted sum, with emphasis on waveform reconstruction (SNR loss)</li></ul></li>\n</ul>\n<h3>Multi-stage Training</h3>\n<p>Training is performed in three stages, gradually switching datasets and augmentation strategies.</p>\n<p>This multi-stage training yields approximately +1.0 dB from 1st-stage pretraining and an additional +0.2 dB from image-type-specific tuning in the 3rd stage, consistently on both Public and Private sets.</p>\n<h4>1st Stage Training (Pretraining with Synthetic Data)</h4>\n<ul>\n<li>Trained using about 21,000 synthetic images generated from PTB-XL (500 Hz), excluding overlaps with competition data.</li>\n<li>Image-only augmentations include blur, brightness changes, grid coordinate shifts up to 1 pixel via Piecewise Affine, and custom augmentations such as DirtPatch / CreaseWrinkles.</li>\n<li>Image-and-signal joint augmentations include random segment dropout and vertical/horizontal flips.</li>\n<li>GT signals are randomly resampled using scipy.resample_poly to one of\n[250, 256, 512, 500, 1000, 1025].</li>\n<li>Trained for 15 epochs with lr = 4e-4.</li>\n</ul>\n<h4>2nd Stage Training (Main Training)</h4>\n<ul>\n<li>Initialized with weights from the 1st stage and trained using competition data only.</li>\n<li>No image-only augmentation is applied; only image+signal augmentations from the 1st stage are used.</li>\n<li>Trained for 30 epochs with lr = 1e-4.</li>\n</ul>\n<h4>3rd Stage Training (Image-Type-Specific Fine-tuning)</h4>\n<ul>\n<li>Initialized with weights from the 2nd stage and further trained using only image type 006 data.</li>\n<li>Augmentation settings are the same as in the 2nd stage.</li>\n<li>Trained for 50 epochs with lr = 1e-4.</li>\n<li>Although CV performance improved for other image types as well, LB improvements were unstable, so the final submission uses the 3rd-stage model only for image type 006, and the 2nd-stage model for others.</li>\n</ul>\n<h3>TTA</h3>\n<ul>\n<li>During inference, predictions are made on the original image plus left-right flip, top-bottom flip, and both flips, totaling 4 patterns, and the results are averaged.</li>\n<li>This TTA consistently provides about +0.15–0.2 dB improvement on both Public and Private sets.</li>\n</ul>\n<h1>Score</h1>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>synthetic_data pretrain (1st stage train)</th>\n<th>006 fine tuning (3rd stage train)</th>\n<th>TTA</th>\n<th>full data train</th>\n<th>use target_fs</th>\n<th>use_lead_id</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td>20.17</td>\n<td>20.11</td>\n</tr>\n<tr>\n<td>2</td>\n<td></td>\n<td></td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td></td>\n<td>20.32</td>\n<td>20.26</td>\n</tr>\n<tr>\n<td>3</td>\n<td>✅</td>\n<td></td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td></td>\n<td>21.43</td>\n<td>21.30</td>\n</tr>\n<tr>\n<td>4</td>\n<td>✅</td>\n<td></td>\n<td>✅</td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td>21.54</td>\n<td>21.38</td>\n</tr>\n<tr>\n<td>5</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td></td>\n<td></td>\n<td>21.80</td>\n<td>21.60</td>\n</tr>\n<tr>\n<td>6</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td></td>\n<td>21.66</td>\n<td>21.56</td>\n</tr>\n<tr>\n<td>7</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>✅</td>\n<td>21.81</td>\n<td>21.71</td>\n</tr>\n</tbody>\n</table>\n<p><strong>final submission</strong><br>\nEnsemble #5 + #6 + #7</p>\n<ul>\n<li>Public: 22.04  </li>\n<li>Private: 21.92</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to pay tribute to all the participants who worked on this competition.\nI would also like to thank the hosts for organizing this interesting task competition.\n\n\n# Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F389d63e6a6f3b3a7103b645812f652b0%2Foverview.jpg?generation=1769435859268811&alt=media)\nMy solution consists of a two-stage pipeline: image rectify and signal reconstruction.\n\nFor image rectify, I first apply coarse alignment using a homography transformation based on image matching, then detect grid point coordinates and perform precise alignment using Piecewise Affine transformation.\n\nFor signal reconstruction, I extend the decoder of a UNet with a signal reconstruction module that directly outputs signals from images, and train the model to optimize SNR.\n\n# Pipeline\n## 1. Image rectify\n### 1.a Image Type Classification\n- Since performance is more stable when some image types (006) are handled separately, images are first classified by type.\n- Only competition data is used for training, and a simple classification model with convnext_large_384_in22ft1k as the backbone is trained.\n- There is nothing particularly special here, but the model achieves about 99.9% accuracy in CV.\n\n## 1.b Image Matching\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F329c1c81d9b9cf8db9c62434e5b763b7%2Fimage_matching.jpg?generation=1769435872559085&alt=media)\n- Feature point matching with a template image is performed using ALIKED + LightGlue, followed by coarse alignment via a homography transformation estimated with RANSAC.\n  - The template image is created by averaging competition images of image type 001.\n  - Publicly available pretrained weights are used as-is for both ALIKED and LightGlue.\n- To simultaneously correct rotation, the input image is rotated by 0°, 90°, 180°, 270°, matching is performed for each rotation, and the rotation with the largest number of matches is selected.\n- Although some distortion remains at this stage, almost all images are aligned to the same orientation and composition as the template (001), which stabilizes downstream training and inference.\n\n## 1.c Grid Detection\n- A model is built to detect grid point coordinates from the projectively transformed images.\n    - A UNet with ResNeSt-14d as the encoder is used.\n- Only about 22,000 synthetic images are used for training.\n    - Since image matching already roughly normalizes the input images, augmentation on synthetic data alone is sufficient to generalize to real data.\n    - No annotation on real data is performed, significantly reducing annotation cost.\n- For image type 006, false detections sometimes occur due to moiré patterns, so grid detection is applied after moiré removal using [UHDM](https://github.com/CVMI-Lab/UHDM).\n\n## 1.d Precise Alignment\n- Finally, precise alignment is performed using Piecewise Affine transformation based on the outputs of image matching and grid detection.\n1. From the correspondence points obtained by image matching, regions with small reprojection error and minimal distortion are selected as initial regions. Detected grid points in these regions are matched to ideal grid points using the Hungarian algorithm.\n2. Starting from the initial correspondences, assuming a grid spacing of approximately 40 px, correspondences are incrementally expanded to neighboring grid points using breadth-first search in the up/down/left/right directions.\n3. The final set of correspondences is used as control points to apply a Piecewise Affine transformation to correct distortion over the entire image.\n\n\n## 2. Signal Reconstruction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F18cf7c0c16852c73f0e4dc6f2f6b87fb%2Fnetwork.jpg?generation=1769435905967605&alt=media)\n### Architecture\n- Based on a UNet with a ResNeSt-14d backbone, an architecture is designed by adding a dedicated module to directly estimate signals from the decoder output.\n    - Feature maps from the UNet decoder are expanded 6× along the temporal axis (W direction), then concatenated with sampling frequency (fs) and lead type as features.\n    - After concatenation, convolution along the W direction is applied to output logits representing “signal-likeness” along the y direction for each x column.\n    - From the y-direction logits, two soft expectations are computed: one biased toward the top edge and one toward the bottom edge, and a gate function predicts (0–1) which to use for each x.\n    - The final estimated y (pixel coordinates) is converted to a signal value in mV based on ECG drawing geometry (paper size / resolution / row layout).\n\n### Inputs, Outputs, and Loss Functions\n- Inputs\n  - Cropped images for each lead\n      - Fixed crop size: H×W = 600×491\n      - RGB images with positional encoding added for x and y directions, resulting in 5-channel input\n      - The long lead (II) is split into 4 parts to match the temporal length of other leads\n  - Target sampling frequency (fs)\n  - Lead type\n      - Adding target fs and lead type as input features improved CV performance by about +1.8 dB, but unfortunately did not yield a clear improvement on the LB.\n\n- Outputs\n  - Segmentation maps\n      - Three classes: background, target signal to be reconstructed, and non-target signals\n  - Reconstructed signal waveform\n\n- Loss\n  - Segmentation loss: Dice + CE, weighted 0.5 : 0.5\n  - Waveform loss: after interpolating the predicted waveform to match the GT length, an SNR-based loss is used as the main loss, with L1 loss added as an auxiliary term\n  - The final loss is a weighted sum, with emphasis on waveform reconstruction (SNR loss)\n\n\n### Multi-stage Training\nTraining is performed in three stages, gradually switching datasets and augmentation strategies.\n\nThis multi-stage training yields approximately +1.0 dB from 1st-stage pretraining and an additional +0.2 dB from image-type-specific tuning in the 3rd stage, consistently on both Public and Private sets.\n\n#### 1st Stage Training (Pretraining with Synthetic Data)\n- Trained using about 21,000 synthetic images generated from PTB-XL (500 Hz), excluding overlaps with competition data.\n- Image-only augmentations include blur, brightness changes, grid coordinate shifts up to 1 pixel via Piecewise Affine, and custom augmentations such as DirtPatch / CreaseWrinkles.\n- Image-and-signal joint augmentations include random segment dropout and vertical/horizontal flips.\n- GT signals are randomly resampled using scipy.resample_poly to one of\n[250, 256, 512, 500, 1000, 1025].\n- Trained for 15 epochs with lr = 4e-4.\n\n#### 2nd Stage Training (Main Training)\n- Initialized with weights from the 1st stage and trained using competition data only.\n- No image-only augmentation is applied; only image+signal augmentations from the 1st stage are used.\n- Trained for 30 epochs with lr = 1e-4.\n\n#### 3rd Stage Training (Image-Type-Specific Fine-tuning)\n- Initialized with weights from the 2nd stage and further trained using only image type 006 data.\n- Augmentation settings are the same as in the 2nd stage.\n- Trained for 50 epochs with lr = 1e-4.\n- Although CV performance improved for other image types as well, LB improvements were unstable, so the final submission uses the 3rd-stage model only for image type 006, and the 2nd-stage model for others.\n\n### TTA\n- During inference, predictions are made on the original image plus left-right flip, top-bottom flip, and both flips, totaling 4 patterns, and the results are averaged.\n- This TTA consistently provides about +0.15–0.2 dB improvement on both Public and Private sets.\n\n\n# Score\n| # | synthetic_data pretrain (1st stage train) | 006 fine tuning (3rd stage train) | TTA | full data train | use target_fs | use_lead_id | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- |\n| 1 |  |  |  |  |  |  | 20.17 | 20.11 |\n| 2 |  |  | ✅ |  |  |  | 20.32 | 20.26 |\n| 3 | ✅ |  | ✅ |  |  |  | 21.43 | 21.30 |\n| 4 | ✅ |  | ✅ | ✅ |  |  | 21.54 | 21.38 |\n| 5 | ✅ | ✅ | ✅ | ✅ |  |  | 21.80 | 21.60 |\n| 6 | ✅ | ✅ | ✅ | ✅ | ✅ |  | 21.66 | 21.56 |\n| 7 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 21.81 | 21.71 |\n\n**final submission**  \nEnsemble #5 + #6 + #7\n- Public: 22.04  \n- Private: 21.92",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3397112": "First of all, I would like to pay tribute to all the participants who worked on this competition.\nI would also like to thank the hosts for organizing this interesting task competition.\n\n\n# Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F389d63e6a6f3b3a7103b645812f652b0%2Foverview.jpg?generation=1769435859268811&alt=media)\nMy solution consists of a two-stage pipeline: image rectify and signal reconstruction.\n\nFor image rectify, I first apply coarse alignment using a homography transformation based on image matching, then detect grid point coordinates and perform precise alignment using Piecewise Affine transformation.\n\nFor signal reconstruction, I extend the decoder of a UNet with a signal reconstruction module that directly outputs signals from images, and train the model to optimize SNR.\n\n# Pipeline\n## 1. Image rectify\n### 1.a Image Type Classification\n- Since performance is more stable when some image types (006) are handled separately, images are first classified by type.\n- Only competition data is used for training, and a simple classification model with convnext_large_384_in22ft1k as the backbone is trained.\n- There is nothing particularly special here, but the model achieves about 99.9% accuracy in CV.\n\n## 1.b Image Matching\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F329c1c81d9b9cf8db9c62434e5b763b7%2Fimage_matching.jpg?generation=1769435872559085&alt=media)\n- Feature point matching with a template image is performed using ALIKED + LightGlue, followed by coarse alignment via a homography transformation estimated with RANSAC.\n  - The template image is created by averaging competition images of image type 001.\n  - Publicly available pretrained weights are used as-is for both ALIKED and LightGlue.\n- To simultaneously correct rotation, the input image is rotated by 0°, 90°, 180°, 270°, matching is performed for each rotation, and the rotation with the largest number of matches is selected.\n- Although some distortion remains at this stage, almost all images are aligned to the same orientation and composition as the template (001), which stabilizes downstream training and inference.\n\n## 1.c Grid Detection\n- A model is built to detect grid point coordinates from the projectively transformed images.\n    - A UNet with ResNeSt-14d as the encoder is used.\n- Only about 22,000 synthetic images are used for training.\n    - Since image matching already roughly normalizes the input images, augmentation on synthetic data alone is sufficient to generalize to real data.\n    - No annotation on real data is performed, significantly reducing annotation cost.\n- For image type 006, false detections sometimes occur due to moiré patterns, so grid detection is applied after moiré removal using [UHDM](https://github.com/CVMI-Lab/UHDM).\n\n## 1.d Precise Alignment\n- Finally, precise alignment is performed using Piecewise Affine transformation based on the outputs of image matching and grid detection.\n1. From the correspondence points obtained by image matching, regions with small reprojection error and minimal distortion are selected as initial regions. Detected grid points in these regions are matched to ideal grid points using the Hungarian algorithm.\n2. Starting from the initial correspondences, assuming a grid spacing of approximately 40 px, correspondences are incrementally expanded to neighboring grid points using breadth-first search in the up/down/left/right directions.\n3. The final set of correspondences is used as control points to apply a Piecewise Affine transformation to correct distortion over the entire image.\n\n\n## 2. Signal Reconstruction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3823496%2F18cf7c0c16852c73f0e4dc6f2f6b87fb%2Fnetwork.jpg?generation=1769435905967605&alt=media)\n### Architecture\n- Based on a UNet with a ResNeSt-14d backbone, an architecture is designed by adding a dedicated module to directly estimate signals from the decoder output.\n    - Feature maps from the UNet decoder are expanded 6× along the temporal axis (W direction), then concatenated with sampling frequency (fs) and lead type as features.\n    - After concatenation, convolution along the W direction is applied to output logits representing “signal-likeness” along the y direction for each x column.\n    - From the y-direction logits, two soft expectations are computed: one biased toward the top edge and one toward the bottom edge, and a gate function predicts (0–1) which to use for each x.\n    - The final estimated y (pixel coordinates) is converted to a signal value in mV based on ECG drawing geometry (paper size / resolution / row layout).\n\n### Inputs, Outputs, and Loss Functions\n- Inputs\n  - Cropped images for each lead\n      - Fixed crop size: H×W = 600×491\n      - RGB images with positional encoding added for x and y directions, resulting in 5-channel input\n      - The long lead (II) is split into 4 parts to match the temporal length of other leads\n  - Target sampling frequency (fs)\n  - Lead type\n      - Adding target fs and lead type as input features improved CV performance by about +1.8 dB, but unfortunately did not yield a clear improvement on the LB.\n\n- Outputs\n  - Segmentation maps\n      - Three classes: background, target signal to be reconstructed, and non-target signals\n  - Reconstructed signal waveform\n\n- Loss\n  - Segmentation loss: Dice + CE, weighted 0.5 : 0.5\n  - Waveform loss: after interpolating the predicted waveform to match the GT length, an SNR-based loss is used as the main loss, with L1 loss added as an auxiliary term\n  - The final loss is a weighted sum, with emphasis on waveform reconstruction (SNR loss)\n\n\n### Multi-stage Training\nTraining is performed in three stages, gradually switching datasets and augmentation strategies.\n\nThis multi-stage training yields approximately +1.0 dB from 1st-stage pretraining and an additional +0.2 dB from image-type-specific tuning in the 3rd stage, consistently on both Public and Private sets.\n\n#### 1st Stage Training (Pretraining with Synthetic Data)\n- Trained using about 21,000 synthetic images generated from PTB-XL (500 Hz), excluding overlaps with competition data.\n- Image-only augmentations include blur, brightness changes, grid coordinate shifts up to 1 pixel via Piecewise Affine, and custom augmentations such as DirtPatch / CreaseWrinkles.\n- Image-and-signal joint augmentations include random segment dropout and vertical/horizontal flips.\n- GT signals are randomly resampled using scipy.resample_poly to one of\n[250, 256, 512, 500, 1000, 1025].\n- Trained for 15 epochs with lr = 4e-4.\n\n#### 2nd Stage Training (Main Training)\n- Initialized with weights from the 1st stage and trained using competition data only.\n- No image-only augmentation is applied; only image+signal augmentations from the 1st stage are used.\n- Trained for 30 epochs with lr = 1e-4.\n\n#### 3rd Stage Training (Image-Type-Specific Fine-tuning)\n- Initialized with weights from the 2nd stage and further trained using only image type 006 data.\n- Augmentation settings are the same as in the 2nd stage.\n- Trained for 50 epochs with lr = 1e-4.\n- Although CV performance improved for other image types as well, LB improvements were unstable, so the final submission uses the 3rd-stage model only for image type 006, and the 2nd-stage model for others.\n\n### TTA\n- During inference, predictions are made on the original image plus left-right flip, top-bottom flip, and both flips, totaling 4 patterns, and the results are averaged.\n- This TTA consistently provides about +0.15–0.2 dB improvement on both Public and Private sets.\n\n\n# Score\n| # | synthetic_data pretrain (1st stage train) | 006 fine tuning (3rd stage train) | TTA | full data train | use target_fs | use_lead_id | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- |\n| 1 |  |  |  |  |  |  | 20.17 | 20.11 |\n| 2 |  |  | ✅ |  |  |  | 20.32 | 20.26 |\n| 3 | ✅ |  | ✅ |  |  |  | 21.43 | 21.30 |\n| 4 | ✅ |  | ✅ | ✅ |  |  | 21.54 | 21.38 |\n| 5 | ✅ | ✅ | ✅ | ✅ |  |  | 21.80 | 21.60 |\n| 6 | ✅ | ✅ | ✅ | ✅ | ✅ |  | 21.66 | 21.56 |\n| 7 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 21.81 | 21.71 |\n\n**final submission**  \nEnsemble #5 + #6 + #7\n- Public: 22.04  \n- Private: 21.92"
  }
}