{
  "id": 669897,
  "title": "4th Place Solution",
  "url": "/competitions/physionet-ecg-image-digitization/writeups/4th-place-solution",
  "author_name": "",
  "post_date": "2026-01-25T00:32:19.667Z",
  "votes": 27,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First, I would like to thank the hosts and Kaggle for organizing this competition.\nI will explain my 4th place solution, focusing on the processing pipeline and model/training strategy.</p>\n<h2>Solution Overview</h2>\n<p>My approach focuses on rectification to make it easier for the model to recognize the signals, and then estimating the waveform using a regression model. Signal segmentation is treated only as an auxiliary task.</p>\n<p>The main points are:</p>\n<ul>\n<li><strong>Robust Rectification</strong>: Estimate grid intersections and layout, then map from an indexed grid to a normalized coordinate system (based on <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s Discussion).</li>\n<li><strong>Vertical Splitting</strong>: Split the rectified image vertically by leads. This reduces the complexity of the prediction task.</li>\n<li><strong>SNR-based Training</strong>: Use a loss function based on SNR, which is close to the evaluation metric.</li>\n</ul>\n<hr>\n<h2>Processing Pipeline</h2>\n<p>The process from input image to signal estimation is divided into 4 stages.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F902ffc495a55d8b24e4bb4a14f411c44%2Foverview.jpg?generation=1769251811275527&amp;alt=media\" alt=\"\"></p>\n<h3>Stage 1: Grid Intersection &amp; Orientation</h3>\n<h4>Grid Intersection &amp; Layout Detection</h4>\n<p>For one high-resolution image, I use a sliding window (704x704 size, 0.5 overlap) to estimate the following using a multi-task model:</p>\n<ol>\n<li><strong>Grid Intersection Heatmap</strong>: All grid intersections.</li>\n<li><strong>Layout Segmentation</strong>: Positions of lead separator lines and calibration signals.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Fcad3ae422c01202661fdd7c80355967f%2F02_grid_overlay.jpg?generation=1769251996864767&amp;alt=media\" alt=\"\"></p>\n<h4>Orientation Correction</h4>\n<p>Using the detected layout information, I estimate the image orientation and rotate it (0/90/180/270 degrees) to a unified upright orientation. The detected intersection coordinates are also rotated.</p>\n<h3>Stage 2: Grid Line Indexing</h3>\n<p>Using the corrected image, the model estimates an index for each pixel indicating which vertical or horizontal line it belongs to.</p>\n<ul>\n<li><strong>Input</strong>: Resized and padded to 1536x1536.</li>\n<li><strong>Output</strong>: Vertical line index (55 classes + background), Horizontal line index (43 classes + background).</li>\n<li><strong>Assignment</strong>: For each intersection detected in Stage 1, I aggregate the predicted labels in its neighborhood (radius 6px) and determine the row/col index by majority vote.</li>\n</ul>\n<p>This determines where each intersection corresponds to in the grid <code>(row, col)</code>.</p>\n<h3>Stage 3: Rectification &amp; Splitting</h3>\n<h4>Rectification (Grid Unwarping)</h4>\n<p>I create a sampling coordinate field from the indexed intersections and warp the image to a fixed-size coordinate system. This method is the same as the <code>grid_sample</code> approach shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.\nThe output size is fixed at 1700x2200.</p>\n<h4>Crop &amp; Vertical Splitting</h4>\n<p>The rectified image often contains headers or margins, so I crop the top 25% to keep only the signal area.\nThen, I split the image width into <strong>4 segments</strong>. The segment boundaries are determined by converting fixed column indices to pixels in the rectified image.</p>\n<ul>\n<li>Each segment contains 3 short leads and 2.5 seconds of the long Lead II stacked vertically.</li>\n<li><strong>Purpose</strong>:<ul>\n<li><strong>Reduce Complexity</strong>: The model does not need to search spatially for where each lead is in the whole image.</li>\n<li><strong>Information Interpolation</strong>: Since the time axis is aligned within the same segment, taking 4 signals as input allows the model to use information from signals above and below to fill in gaps if grid lines are missing or noisy.</li></ul></li>\n</ul>\n<p>Each segment image is resized to 800x800.</p>\n<h3>Stage 4: Signal Estimation (Regression)</h3>\n<h4>Waveform Estimation Model</h4>\n<p>The model takes the 4 segmented images as input and regresses the time-series values for 12 leads.</p>\n<ul>\n<li><strong>Input</strong>: 800x800 x 3ch (x 4 segments)</li>\n<li><strong>Output</strong>:<ul>\n<li><strong>Short Leads</strong>: Outputs waveforms for the corresponding 3 leads from each of the 4 segments.</li>\n<li><strong>Long Lead II</strong>: Concatenates features from the 4 segments and outputs the full-length waveform.</li></ul></li>\n<li><strong>Resampling</strong>: The fixed-length output sequence is resampled to the required number of samples based on the test data's <code>fs</code> (sampling frequency) using linear interpolation.</li>\n</ul>\n<h4>Ensemble</h4>\n<p>I trained multiple models with different backbones and ensembled them using weighted median. The median was more robust against spike-like outlier predictions than the mean.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Weight</th>\n<th>CV (SNR)</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnetaa101d.sw_in12k_ft_in1k</td>\n<td>1</td>\n<td>25.62</td>\n<td>21.99</td>\n<td>22.03</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m.in21k_ft_in1k</td>\n<td>3</td>\n<td>26.05</td>\n<td>22.39</td>\n<td>22.42</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b6.ns_jft_in1k</td>\n<td>4</td>\n<td>26.05</td>\n<td>22.50</td>\n<td>22.44</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_l.in21k_ft_in1k</td>\n<td>2</td>\n<td>26.08</td>\n<td>22.50</td>\n<td>22.53</td>\n</tr>\n<tr>\n<td>Ensemble (Mean)</td>\n<td></td>\n<td>26.58</td>\n<td>22.30</td>\n<td>22.37</td>\n</tr>\n<tr>\n<td>Ensemble (Median)</td>\n<td></td>\n<td>26.45</td>\n<td><strong>22.71</strong></td>\n<td><strong>22.69</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Models &amp; Training Strategy</h2>\n<h3>Stage 1 &amp; 2 Models (Grid &amp; Index)</h3>\n<h4>Dataset &amp; Annotation</h4>\n<p>I created annotations (intersections and line indices) myself. I used a cycle of \"Initial labeling by rule-based processing\" -&gt; \"Model training\" -&gt; \"Manual correction of inference results\" to create about 1000 annotated images. Note that \"lines\" are stored as sequences of intersections, not as pixel-level continuous curves.</p>\n<h4>Backbone &amp; Pretraining</h4>\n<p>I used ConvNeXt Small (<code>convnext_small.dinov3_lvd1689m</code>) for the backbones of Stage 1 and 2.\nNotably, I performed pretraining in the ECG image domain using FCMAE (Fully Convolutional Masked Autoencoder). This helped the model converge faster and improved detection accuracy even with limited labels.</p>\n<h3>Stage 4 Model (Digitization)</h3>\n<h4>Network Architecture</h4>\n<p>I used a U-Net based architecture with specific modifications for ECG Digitization.</p>\n<ul>\n<li><strong>Asymmetric Decoder</strong>:\nThe decoder performs normal 2D Upsampling initially. However, after the feature map size becomes large enough (1/8 scale), it performs Upsampling only in the horizontal (time) direction.<ul>\n<li>This reduces the computational cost of unnecessary vertical resolution while ensuring high temporal resolution.</li></ul></li>\n<li><strong>Multi-head Regression</strong>:<ul>\n<li><strong>Short Leads Head</strong>: Performs Global Average Pooling vertically on the decoder output, then regresses using 1D Conv.</li>\n<li><strong>Long Lead II Head</strong>: Concatenates features between segments along the time axis, then regresses using another 1D Conv.</li></ul></li>\n<li><strong>Auxiliary Head</strong>:\nI added a segmentation head to estimate the signal area for each lead. This was an auxiliary task to improve signal discrimination.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9a24b2d1314f48bbda103756ecd8481e%2Fmodel.png?generation=1769301568006031&amp;alt=media\" alt=\"\"></p>\n<h4>Loss Function</h4>\n<p>To prevent discrepancy with the evaluation metric, I designed a loss function based on SNR. I optimized the SNR term and the Mean term separately.</p>\n<h4>External Datasets</h4>\n<p>In addition to the training data, I used the following external datasets converted into images using <code>ecg-image-kit</code>:</p>\n<ul>\n<li>PTB-XL</li>\n<li>CODE-15%</li>\n</ul>\n<p>Since the generated data has accurate ground truth masks for which pixel corresponds to which lead, I was able to use them for training the auxiliary segmentation task.</p>\n<hr>\n<h2>References</h2>\n<ol>\n<li><p>Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., &amp; Xie, S. (2023). Convnext v2: Co-designing and scaling convnets with masked autoencoders. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 16133-16142).</p></li>\n<li><p>Kshama Kodthalu Shivashankara, Deepanshi, Afagh Mehri Shervedani, Matthew A. Reyna, Gari D. Clifford, Reza Sameni (2024). ECG-image-kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. In Physiological Measurement. IOP Publishing. doi: 10.1088/1361-6579/ad4954</p></li>\n<li><p>ECG-Image-Kit: A Toolkit for Synthesis, Analysis, and Digitization of Electrocardiogram Images, (2024). URL: <a href=\"https://github.com/alphanumericslab/ecg-image-kit\" target=\"_blank\">https://github.com/alphanumericslab/ecg-image-kit</a></p></li>\n<li><p>Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., &amp; Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. <a href=\"https://doi.org/10.13026/kfzx-aw45\" target=\"_blank\">https://doi.org/10.13026/kfzx-aw45</a></p></li>\n<li><p>Ribeiro, A. H., Paixao, G. M. M., Lima, E. M., Horta Ribeiro, M., Pinto Filho, M. M., Gomes, P. R., Oliveira, D. M., Meira Jr, W., Schon, T. B., &amp; Ribeiro, A. L. P. (2021). CODE-15%: a large scale annotated dataset of 12-lead ECGs (1.0.0) [Data set]. Zenodo. <a href=\"https://doi.org/10.5281/zenodo.4916206\" target=\"_blank\">https://doi.org/10.5281/zenodo.4916206</a></p></li>\n</ol>\n<h2>Code Release</h2>\n<p>Code: <a href=\"https://github.com/uchiyama33/physionet_4th_place\" target=\"_blank\">https://github.com/uchiyama33/physionet_4th_place</a></p>\n<p>Submission notebook: <a href=\"https://www.kaggle.com/code/tomoon33/physionet-submission-4th-place\" target=\"_blank\">https://www.kaggle.com/code/tomoon33/physionet-submission-4th-place</a></p>",
  "messages": [
    {
      "id": "3396403",
      "postDate": "01/25/2026 00:23:13",
      "content": "<p>First, I would like to thank the hosts and Kaggle for organizing this competition.\nI will explain my 4th place solution, focusing on the processing pipeline and model/training strategy.</p>\n<h2>Solution Overview</h2>\n<p>My approach focuses on rectification to make it easier for the model to recognize the signals, and then estimating the waveform using a regression model. Signal segmentation is treated only as an auxiliary task.</p>\n<p>The main points are:</p>\n<ul>\n<li><strong>Robust Rectification</strong>: Estimate grid intersections and layout, then map from an indexed grid to a normalized coordinate system (based on <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s Discussion).</li>\n<li><strong>Vertical Splitting</strong>: Split the rectified image vertically by leads. This reduces the complexity of the prediction task.</li>\n<li><strong>SNR-based Training</strong>: Use a loss function based on SNR, which is close to the evaluation metric.</li>\n</ul>\n<hr>\n<h2>Processing Pipeline</h2>\n<p>The process from input image to signal estimation is divided into 4 stages.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F902ffc495a55d8b24e4bb4a14f411c44%2Foverview.jpg?generation=1769251811275527&amp;alt=media\" alt=\"\"></p>\n<h3>Stage 1: Grid Intersection &amp; Orientation</h3>\n<h4>Grid Intersection &amp; Layout Detection</h4>\n<p>For one high-resolution image, I use a sliding window (704x704 size, 0.5 overlap) to estimate the following using a multi-task model:</p>\n<ol>\n<li><strong>Grid Intersection Heatmap</strong>: All grid intersections.</li>\n<li><strong>Layout Segmentation</strong>: Positions of lead separator lines and calibration signals.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Fcad3ae422c01202661fdd7c80355967f%2F02_grid_overlay.jpg?generation=1769251996864767&amp;alt=media\" alt=\"\"></p>\n<h4>Orientation Correction</h4>\n<p>Using the detected layout information, I estimate the image orientation and rotate it (0/90/180/270 degrees) to a unified upright orientation. The detected intersection coordinates are also rotated.</p>\n<h3>Stage 2: Grid Line Indexing</h3>\n<p>Using the corrected image, the model estimates an index for each pixel indicating which vertical or horizontal line it belongs to.</p>\n<ul>\n<li><strong>Input</strong>: Resized and padded to 1536x1536.</li>\n<li><strong>Output</strong>: Vertical line index (55 classes + background), Horizontal line index (43 classes + background).</li>\n<li><strong>Assignment</strong>: For each intersection detected in Stage 1, I aggregate the predicted labels in its neighborhood (radius 6px) and determine the row/col index by majority vote.</li>\n</ul>\n<p>This determines where each intersection corresponds to in the grid <code>(row, col)</code>.</p>\n<h3>Stage 3: Rectification &amp; Splitting</h3>\n<h4>Rectification (Grid Unwarping)</h4>\n<p>I create a sampling coordinate field from the indexed intersections and warp the image to a fixed-size coordinate system. This method is the same as the <code>grid_sample</code> approach shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.\nThe output size is fixed at 1700x2200.</p>\n<h4>Crop &amp; Vertical Splitting</h4>\n<p>The rectified image often contains headers or margins, so I crop the top 25% to keep only the signal area.\nThen, I split the image width into <strong>4 segments</strong>. The segment boundaries are determined by converting fixed column indices to pixels in the rectified image.</p>\n<ul>\n<li>Each segment contains 3 short leads and 2.5 seconds of the long Lead II stacked vertically.</li>\n<li><strong>Purpose</strong>:<ul>\n<li><strong>Reduce Complexity</strong>: The model does not need to search spatially for where each lead is in the whole image.</li>\n<li><strong>Information Interpolation</strong>: Since the time axis is aligned within the same segment, taking 4 signals as input allows the model to use information from signals above and below to fill in gaps if grid lines are missing or noisy.</li></ul></li>\n</ul>\n<p>Each segment image is resized to 800x800.</p>\n<h3>Stage 4: Signal Estimation (Regression)</h3>\n<h4>Waveform Estimation Model</h4>\n<p>The model takes the 4 segmented images as input and regresses the time-series values for 12 leads.</p>\n<ul>\n<li><strong>Input</strong>: 800x800 x 3ch (x 4 segments)</li>\n<li><strong>Output</strong>:<ul>\n<li><strong>Short Leads</strong>: Outputs waveforms for the corresponding 3 leads from each of the 4 segments.</li>\n<li><strong>Long Lead II</strong>: Concatenates features from the 4 segments and outputs the full-length waveform.</li></ul></li>\n<li><strong>Resampling</strong>: The fixed-length output sequence is resampled to the required number of samples based on the test data's <code>fs</code> (sampling frequency) using linear interpolation.</li>\n</ul>\n<h4>Ensemble</h4>\n<p>I trained multiple models with different backbones and ensembled them using weighted median. The median was more robust against spike-like outlier predictions than the mean.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Weight</th>\n<th>CV (SNR)</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnetaa101d.sw_in12k_ft_in1k</td>\n<td>1</td>\n<td>25.62</td>\n<td>21.99</td>\n<td>22.03</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m.in21k_ft_in1k</td>\n<td>3</td>\n<td>26.05</td>\n<td>22.39</td>\n<td>22.42</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b6.ns_jft_in1k</td>\n<td>4</td>\n<td>26.05</td>\n<td>22.50</td>\n<td>22.44</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_l.in21k_ft_in1k</td>\n<td>2</td>\n<td>26.08</td>\n<td>22.50</td>\n<td>22.53</td>\n</tr>\n<tr>\n<td>Ensemble (Mean)</td>\n<td></td>\n<td>26.58</td>\n<td>22.30</td>\n<td>22.37</td>\n</tr>\n<tr>\n<td>Ensemble (Median)</td>\n<td></td>\n<td>26.45</td>\n<td><strong>22.71</strong></td>\n<td><strong>22.69</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Models &amp; Training Strategy</h2>\n<h3>Stage 1 &amp; 2 Models (Grid &amp; Index)</h3>\n<h4>Dataset &amp; Annotation</h4>\n<p>I created annotations (intersections and line indices) myself. I used a cycle of \"Initial labeling by rule-based processing\" -&gt; \"Model training\" -&gt; \"Manual correction of inference results\" to create about 1000 annotated images. Note that \"lines\" are stored as sequences of intersections, not as pixel-level continuous curves.</p>\n<h4>Backbone &amp; Pretraining</h4>\n<p>I used ConvNeXt Small (<code>convnext_small.dinov3_lvd1689m</code>) for the backbones of Stage 1 and 2.\nNotably, I performed pretraining in the ECG image domain using FCMAE (Fully Convolutional Masked Autoencoder). This helped the model converge faster and improved detection accuracy even with limited labels.</p>\n<h3>Stage 4 Model (Digitization)</h3>\n<h4>Network Architecture</h4>\n<p>I used a U-Net based architecture with specific modifications for ECG Digitization.</p>\n<ul>\n<li><strong>Asymmetric Decoder</strong>:\nThe decoder performs normal 2D Upsampling initially. However, after the feature map size becomes large enough (1/8 scale), it performs Upsampling only in the horizontal (time) direction.<ul>\n<li>This reduces the computational cost of unnecessary vertical resolution while ensuring high temporal resolution.</li></ul></li>\n<li><strong>Multi-head Regression</strong>:<ul>\n<li><strong>Short Leads Head</strong>: Performs Global Average Pooling vertically on the decoder output, then regresses using 1D Conv.</li>\n<li><strong>Long Lead II Head</strong>: Concatenates features between segments along the time axis, then regresses using another 1D Conv.</li></ul></li>\n<li><strong>Auxiliary Head</strong>:\nI added a segmentation head to estimate the signal area for each lead. This was an auxiliary task to improve signal discrimination.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9a24b2d1314f48bbda103756ecd8481e%2Fmodel.png?generation=1769301568006031&amp;alt=media\" alt=\"\"></p>\n<h4>Loss Function</h4>\n<p>To prevent discrepancy with the evaluation metric, I designed a loss function based on SNR. I optimized the SNR term and the Mean term separately.</p>\n<h4>External Datasets</h4>\n<p>In addition to the training data, I used the following external datasets converted into images using <code>ecg-image-kit</code>:</p>\n<ul>\n<li>PTB-XL</li>\n<li>CODE-15%</li>\n</ul>\n<p>Since the generated data has accurate ground truth masks for which pixel corresponds to which lead, I was able to use them for training the auxiliary segmentation task.</p>\n<hr>\n<h2>References</h2>\n<ol>\n<li><p>Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., &amp; Xie, S. (2023). Convnext v2: Co-designing and scaling convnets with masked autoencoders. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 16133-16142).</p></li>\n<li><p>Kshama Kodthalu Shivashankara, Deepanshi, Afagh Mehri Shervedani, Matthew A. Reyna, Gari D. Clifford, Reza Sameni (2024). ECG-image-kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. In Physiological Measurement. IOP Publishing. doi: 10.1088/1361-6579/ad4954</p></li>\n<li><p>ECG-Image-Kit: A Toolkit for Synthesis, Analysis, and Digitization of Electrocardiogram Images, (2024). URL: <a href=\"https://github.com/alphanumericslab/ecg-image-kit\" target=\"_blank\">https://github.com/alphanumericslab/ecg-image-kit</a></p></li>\n<li><p>Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., &amp; Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. <a href=\"https://doi.org/10.13026/kfzx-aw45\" target=\"_blank\">https://doi.org/10.13026/kfzx-aw45</a></p></li>\n<li><p>Ribeiro, A. H., Paixao, G. M. M., Lima, E. M., Horta Ribeiro, M., Pinto Filho, M. M., Gomes, P. R., Oliveira, D. M., Meira Jr, W., Schon, T. B., &amp; Ribeiro, A. L. P. (2021). CODE-15%: a large scale annotated dataset of 12-lead ECGs (1.0.0) [Data set]. Zenodo. <a href=\"https://doi.org/10.5281/zenodo.4916206\" target=\"_blank\">https://doi.org/10.5281/zenodo.4916206</a></p></li>\n</ol>\n<h2>Code Release</h2>\n<p>Code: <a href=\"https://github.com/uchiyama33/physionet_4th_place\" target=\"_blank\">https://github.com/uchiyama33/physionet_4th_place</a></p>\n<p>Submission notebook: <a href=\"https://www.kaggle.com/code/tomoon33/physionet-submission-4th-place\" target=\"_blank\">https://www.kaggle.com/code/tomoon33/physionet-submission-4th-place</a></p>",
      "rawMarkdown": "First, I would like to thank the hosts and Kaggle for organizing this competition.\nI will explain my 4th place solution, focusing on the processing pipeline and model/training strategy.\n\n## Solution Overview\n\nMy approach focuses on rectification to make it easier for the model to recognize the signals, and then estimating the waveform using a regression model. Signal segmentation is treated only as an auxiliary task.\n\nThe main points are:\n\n- **Robust Rectification**: Estimate grid intersections and layout, then map from an indexed grid to a normalized coordinate system (based on @hengck23's Discussion).\n- **Vertical Splitting**: Split the rectified image vertically by leads. This reduces the complexity of the prediction task.\n- **SNR-based Training**: Use a loss function based on SNR, which is close to the evaluation metric.\n\n---\n\n## Processing Pipeline\n\nThe process from input image to signal estimation is divided into 4 stages.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F902ffc495a55d8b24e4bb4a14f411c44%2Foverview.jpg?generation=1769251811275527&alt=media)\n\n### Stage 1: Grid Intersection & Orientation\n\n#### Grid Intersection & Layout Detection\nFor one high-resolution image, I use a sliding window (704x704 size, 0.5 overlap) to estimate the following using a multi-task model:\n\n1.  **Grid Intersection Heatmap**: All grid intersections.\n2.  **Layout Segmentation**: Positions of lead separator lines and calibration signals.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Fcad3ae422c01202661fdd7c80355967f%2F02_grid_overlay.jpg?generation=1769251996864767&alt=media)\n\n#### Orientation Correction\nUsing the detected layout information, I estimate the image orientation and rotate it (0/90/180/270 degrees) to a unified upright orientation. The detected intersection coordinates are also rotated.\n\n### Stage 2: Grid Line Indexing\n\nUsing the corrected image, the model estimates an index for each pixel indicating which vertical or horizontal line it belongs to.\n\n- **Input**: Resized and padded to 1536x1536.\n- **Output**: Vertical line index (55 classes + background), Horizontal line index (43 classes + background).\n- **Assignment**: For each intersection detected in Stage 1, I aggregate the predicted labels in its neighborhood (radius 6px) and determine the row/col index by majority vote.\n\nThis determines where each intersection corresponds to in the grid `(row, col)`.\n\n### Stage 3: Rectification & Splitting\n\n#### Rectification (Grid Unwarping)\nI create a sampling coordinate field from the indexed intersections and warp the image to a fixed-size coordinate system. This method is the same as the `grid_sample` approach shared by @hengck23.\nThe output size is fixed at 1700x2200.\n\n#### Crop & Vertical Splitting\nThe rectified image often contains headers or margins, so I crop the top 25% to keep only the signal area.\nThen, I split the image width into **4 segments**. The segment boundaries are determined by converting fixed column indices to pixels in the rectified image.\n\n- Each segment contains 3 short leads and 2.5 seconds of the long Lead II stacked vertically.\n- **Purpose**:\n    - **Reduce Complexity**: The model does not need to search spatially for where each lead is in the whole image.\n    - **Information Interpolation**: Since the time axis is aligned within the same segment, taking 4 signals as input allows the model to use information from signals above and below to fill in gaps if grid lines are missing or noisy.\n\nEach segment image is resized to 800x800.\n\n### Stage 4: Signal Estimation (Regression)\n\n#### Waveform Estimation Model\nThe model takes the 4 segmented images as input and regresses the time-series values for 12 leads.\n\n- **Input**: 800x800 x 3ch (x 4 segments)\n- **Output**:\n    - **Short Leads**: Outputs waveforms for the corresponding 3 leads from each of the 4 segments.\n    - **Long Lead II**: Concatenates features from the 4 segments and outputs the full-length waveform.\n- **Resampling**: The fixed-length output sequence is resampled to the required number of samples based on the test data's `fs` (sampling frequency) using linear interpolation.\n\n#### Ensemble\nI trained multiple models with different backbones and ensembled them using weighted median. The median was more robust against spike-like outlier predictions than the mean.\n\n| Backbone | Weight | CV (SNR) | Public LB | Private LB |\n|---|---:|---:|---:|---:|\n| resnetaa101d.sw_in12k_ft_in1k | 1 | 25.62 | 21.99 | 22.03 |\n| tf_efficientnetv2_m.in21k_ft_in1k | 3 | 26.05 | 22.39 | 22.42 |\n| tf_efficientnet_b6.ns_jft_in1k | 4 | 26.05 | 22.50 | 22.44 |\n| tf_efficientnetv2_l.in21k_ft_in1k | 2 | 26.08 | 22.50 | 22.53 |\n| Ensemble (Mean) |  | 26.58 | 22.30 | 22.37 |\n| Ensemble (Median) |  | 26.45 | **22.71** | **22.69** |\n\n---\n\n## Models & Training Strategy\n\n### Stage 1 & 2 Models (Grid & Index)\n\n#### Dataset & Annotation\nI created annotations (intersections and line indices) myself. I used a cycle of \"Initial labeling by rule-based processing\" -> \"Model training\" -> \"Manual correction of inference results\" to create about 1000 annotated images. Note that \"lines\" are stored as sequences of intersections, not as pixel-level continuous curves.\n\n#### Backbone & Pretraining\nI used ConvNeXt Small (`convnext_small.dinov3_lvd1689m`) for the backbones of Stage 1 and 2.\nNotably, I performed pretraining in the ECG image domain using FCMAE (Fully Convolutional Masked Autoencoder). This helped the model converge faster and improved detection accuracy even with limited labels.\n\n### Stage 4 Model (Digitization)\n\n#### Network Architecture\nI used a U-Net based architecture with specific modifications for ECG Digitization.\n\n- **Asymmetric Decoder**:\n    The decoder performs normal 2D Upsampling initially. However, after the feature map size becomes large enough (1/8 scale), it performs Upsampling only in the horizontal (time) direction.\n    - This reduces the computational cost of unnecessary vertical resolution while ensuring high temporal resolution.\n- **Multi-head Regression**:\n    - **Short Leads Head**: Performs Global Average Pooling vertically on the decoder output, then regresses using 1D Conv.\n    - **Long Lead II Head**: Concatenates features between segments along the time axis, then regresses using another 1D Conv.\n- **Auxiliary Head**:\n    I added a segmentation head to estimate the signal area for each lead. This was an auxiliary task to improve signal discrimination.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9a24b2d1314f48bbda103756ecd8481e%2Fmodel.png?generation=1769301568006031&alt=media)\n\n#### Loss Function\nTo prevent discrepancy with the evaluation metric, I designed a loss function based on SNR. I optimized the SNR term and the Mean term separately.\n\n#### External Datasets\nIn addition to the training data, I used the following external datasets converted into images using `ecg-image-kit`:\n- PTB-XL\n- CODE-15%\n\nSince the generated data has accurate ground truth masks for which pixel corresponds to which lead, I was able to use them for training the auxiliary segmentation task.\n\n---\n\n## References\n\n1. Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., & Xie, S. (2023). Convnext v2: Co-designing and scaling convnets with masked autoencoders. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 16133-16142).\n\n1. Kshama Kodthalu Shivashankara, Deepanshi, Afagh Mehri Shervedani, Matthew A. Reyna, Gari D. Clifford, Reza Sameni (2024). ECG-image-kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. In Physiological Measurement. IOP Publishing. doi: 10.1088/1361-6579/ad4954\n\n1. ECG-Image-Kit: A Toolkit for Synthesis, Analysis, and Digitization of Electrocardiogram Images, (2024). URL: https://github.com/alphanumericslab/ecg-image-kit\n\n1. Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., & Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/kfzx-aw45\n\n1. Ribeiro, A. H., Paixao, G. M. M., Lima, E. M., Horta Ribeiro, M., Pinto Filho, M. M., Gomes, P. R., Oliveira, D. M., Meira Jr, W., Schon, T. B., & Ribeiro, A. L. P. (2021). CODE-15%: a large scale annotated dataset of 12-lead ECGs (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.4916206\n\n## Code Release\n\nCode: https://github.com/uchiyama33/physionet_4th_place\n\nSubmission notebook: https://www.kaggle.com/code/tomoon33/physionet-submission-4th-place",
      "votes": null
    },
    {
      "id": "3398690",
      "postDate": "01/29/2026 14:48:56",
      "content": "<p>Congrats on the win I was going through your git hub code trying to reproduce the work, but got stuck with a missing json file error. data/annotations/multitask_lines_v2/. could you update that or provide resources to reproduce that file</p>",
      "rawMarkdown": "Congrats on the win I was going through your git hub code trying to reproduce the work, but got stuck with a missing json file error. data/annotations/multitask_lines_v2/. could you update that or provide resources to reproduce that file",
      "votes": null
    },
    {
      "id": "3398714",
      "postDate": "01/29/2026 15:36:12",
      "content": "<p>Thanks for checking out the code. I’ve uploaded the missing annotation files to GitHub under data/annotations/. Please pull the latest commit and try again.</p>",
      "rawMarkdown": "Thanks for checking out the code. I’ve uploaded the missing annotation files to GitHub under data/annotations/. Please pull the latest commit and try again.",
      "votes": null
    },
    {
      "id": "3490238",
      "postDate": "07/06/2026 06:55:55",
      "content": "<p>hi, <a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a> . what type of GPU did u use to train it and how long did it take?\nCongrats on this amazing work btw.</p>",
      "rawMarkdown": "hi, @tomoon33 . what type of GPU did u use to train it and how long did it take?\nCongrats on this amazing work btw.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3398690,
      "author_name": "cartoonistbeard",
      "author_url": "",
      "post_date": "01/29/2026 14:48:56",
      "content": "<p>Congrats on the win I was going through your git hub code trying to reproduce the work, but got stuck with a missing json file error. data/annotations/multitask_lines_v2/. could you update that or provide resources to reproduce that file</p>",
      "votes": null,
      "replies": [
        {
          "id": 3398714,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "01/29/2026 15:36:12",
          "content": "<p>Thanks for checking out the code. I’ve uploaded the missing annotation files to GitHub under data/annotations/. Please pull the latest commit and try again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3490238,
      "author_name": "beckpro",
      "author_url": "",
      "post_date": "07/06/2026 06:55:55",
      "content": "<p>hi, <a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a> . what type of GPU did u use to train it and how long did it take?\nCongrats on this amazing work btw.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3396403": "First, I would like to thank the hosts and Kaggle for organizing this competition.\nI will explain my 4th place solution, focusing on the processing pipeline and model/training strategy.\n\n## Solution Overview\n\nMy approach focuses on rectification to make it easier for the model to recognize the signals, and then estimating the waveform using a regression model. Signal segmentation is treated only as an auxiliary task.\n\nThe main points are:\n\n- **Robust Rectification**: Estimate grid intersections and layout, then map from an indexed grid to a normalized coordinate system (based on @hengck23's Discussion).\n- **Vertical Splitting**: Split the rectified image vertically by leads. This reduces the complexity of the prediction task.\n- **SNR-based Training**: Use a loss function based on SNR, which is close to the evaluation metric.\n\n---\n\n## Processing Pipeline\n\nThe process from input image to signal estimation is divided into 4 stages.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F902ffc495a55d8b24e4bb4a14f411c44%2Foverview.jpg?generation=1769251811275527&alt=media)\n\n### Stage 1: Grid Intersection & Orientation\n\n#### Grid Intersection & Layout Detection\nFor one high-resolution image, I use a sliding window (704x704 size, 0.5 overlap) to estimate the following using a multi-task model:\n\n1.  **Grid Intersection Heatmap**: All grid intersections.\n2.  **Layout Segmentation**: Positions of lead separator lines and calibration signals.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Fcad3ae422c01202661fdd7c80355967f%2F02_grid_overlay.jpg?generation=1769251996864767&alt=media)\n\n#### Orientation Correction\nUsing the detected layout information, I estimate the image orientation and rotate it (0/90/180/270 degrees) to a unified upright orientation. The detected intersection coordinates are also rotated.\n\n### Stage 2: Grid Line Indexing\n\nUsing the corrected image, the model estimates an index for each pixel indicating which vertical or horizontal line it belongs to.\n\n- **Input**: Resized and padded to 1536x1536.\n- **Output**: Vertical line index (55 classes + background), Horizontal line index (43 classes + background).\n- **Assignment**: For each intersection detected in Stage 1, I aggregate the predicted labels in its neighborhood (radius 6px) and determine the row/col index by majority vote.\n\nThis determines where each intersection corresponds to in the grid `(row, col)`.\n\n### Stage 3: Rectification & Splitting\n\n#### Rectification (Grid Unwarping)\nI create a sampling coordinate field from the indexed intersections and warp the image to a fixed-size coordinate system. This method is the same as the `grid_sample` approach shared by @hengck23.\nThe output size is fixed at 1700x2200.\n\n#### Crop & Vertical Splitting\nThe rectified image often contains headers or margins, so I crop the top 25% to keep only the signal area.\nThen, I split the image width into **4 segments**. The segment boundaries are determined by converting fixed column indices to pixels in the rectified image.\n\n- Each segment contains 3 short leads and 2.5 seconds of the long Lead II stacked vertically.\n- **Purpose**:\n    - **Reduce Complexity**: The model does not need to search spatially for where each lead is in the whole image.\n    - **Information Interpolation**: Since the time axis is aligned within the same segment, taking 4 signals as input allows the model to use information from signals above and below to fill in gaps if grid lines are missing or noisy.\n\nEach segment image is resized to 800x800.\n\n### Stage 4: Signal Estimation (Regression)\n\n#### Waveform Estimation Model\nThe model takes the 4 segmented images as input and regresses the time-series values for 12 leads.\n\n- **Input**: 800x800 x 3ch (x 4 segments)\n- **Output**:\n    - **Short Leads**: Outputs waveforms for the corresponding 3 leads from each of the 4 segments.\n    - **Long Lead II**: Concatenates features from the 4 segments and outputs the full-length waveform.\n- **Resampling**: The fixed-length output sequence is resampled to the required number of samples based on the test data's `fs` (sampling frequency) using linear interpolation.\n\n#### Ensemble\nI trained multiple models with different backbones and ensembled them using weighted median. The median was more robust against spike-like outlier predictions than the mean.\n\n| Backbone | Weight | CV (SNR) | Public LB | Private LB |\n|---|---:|---:|---:|---:|\n| resnetaa101d.sw_in12k_ft_in1k | 1 | 25.62 | 21.99 | 22.03 |\n| tf_efficientnetv2_m.in21k_ft_in1k | 3 | 26.05 | 22.39 | 22.42 |\n| tf_efficientnet_b6.ns_jft_in1k | 4 | 26.05 | 22.50 | 22.44 |\n| tf_efficientnetv2_l.in21k_ft_in1k | 2 | 26.08 | 22.50 | 22.53 |\n| Ensemble (Mean) |  | 26.58 | 22.30 | 22.37 |\n| Ensemble (Median) |  | 26.45 | **22.71** | **22.69** |\n\n---\n\n## Models & Training Strategy\n\n### Stage 1 & 2 Models (Grid & Index)\n\n#### Dataset & Annotation\nI created annotations (intersections and line indices) myself. I used a cycle of \"Initial labeling by rule-based processing\" -> \"Model training\" -> \"Manual correction of inference results\" to create about 1000 annotated images. Note that \"lines\" are stored as sequences of intersections, not as pixel-level continuous curves.\n\n#### Backbone & Pretraining\nI used ConvNeXt Small (`convnext_small.dinov3_lvd1689m`) for the backbones of Stage 1 and 2.\nNotably, I performed pretraining in the ECG image domain using FCMAE (Fully Convolutional Masked Autoencoder). This helped the model converge faster and improved detection accuracy even with limited labels.\n\n### Stage 4 Model (Digitization)\n\n#### Network Architecture\nI used a U-Net based architecture with specific modifications for ECG Digitization.\n\n- **Asymmetric Decoder**:\n    The decoder performs normal 2D Upsampling initially. However, after the feature map size becomes large enough (1/8 scale), it performs Upsampling only in the horizontal (time) direction.\n    - This reduces the computational cost of unnecessary vertical resolution while ensuring high temporal resolution.\n- **Multi-head Regression**:\n    - **Short Leads Head**: Performs Global Average Pooling vertically on the decoder output, then regresses using 1D Conv.\n    - **Long Lead II Head**: Concatenates features between segments along the time axis, then regresses using another 1D Conv.\n- **Auxiliary Head**:\n    I added a segmentation head to estimate the signal area for each lead. This was an auxiliary task to improve signal discrimination.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9a24b2d1314f48bbda103756ecd8481e%2Fmodel.png?generation=1769301568006031&alt=media)\n\n#### Loss Function\nTo prevent discrepancy with the evaluation metric, I designed a loss function based on SNR. I optimized the SNR term and the Mean term separately.\n\n#### External Datasets\nIn addition to the training data, I used the following external datasets converted into images using `ecg-image-kit`:\n- PTB-XL\n- CODE-15%\n\nSince the generated data has accurate ground truth masks for which pixel corresponds to which lead, I was able to use them for training the auxiliary segmentation task.\n\n---\n\n## References\n\n1. Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., & Xie, S. (2023). Convnext v2: Co-designing and scaling convnets with masked autoencoders. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 16133-16142).\n\n1. Kshama Kodthalu Shivashankara, Deepanshi, Afagh Mehri Shervedani, Matthew A. Reyna, Gari D. Clifford, Reza Sameni (2024). ECG-image-kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. In Physiological Measurement. IOP Publishing. doi: 10.1088/1361-6579/ad4954\n\n1. ECG-Image-Kit: A Toolkit for Synthesis, Analysis, and Digitization of Electrocardiogram Images, (2024). URL: https://github.com/alphanumericslab/ecg-image-kit\n\n1. Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., & Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/kfzx-aw45\n\n1. Ribeiro, A. H., Paixao, G. M. M., Lima, E. M., Horta Ribeiro, M., Pinto Filho, M. M., Gomes, P. R., Oliveira, D. M., Meira Jr, W., Schon, T. B., & Ribeiro, A. L. P. (2021). CODE-15%: a large scale annotated dataset of 12-lead ECGs (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.4916206\n\n## Code Release\n\nCode: https://github.com/uchiyama33/physionet_4th_place\n\nSubmission notebook: https://www.kaggle.com/code/tomoon33/physionet-submission-4th-place",
    "3398690": "Congrats on the win I was going through your git hub code trying to reproduce the work, but got stuck with a missing json file error. data/annotations/multitask_lines_v2/. could you update that or provide resources to reproduce that file",
    "3398714": "Thanks for checking out the code. I’ve uploaded the missing annotation files to GitHub under data/annotations/. Please pull the latest commit and try again.",
    "3490238": "hi, @tomoon33 . what type of GPU did u use to train it and how long did it take?\nCongrats on this amazing work btw."
  },
  "source": "meta"
}