{
  "id": 669871,
  "title": "2nd place solution",
  "url": "/competitions/physionet-ecg-image-digitization/writeups/2nd-place-solution",
  "author_name": "",
  "post_date": "2026-01-24T19:23:05.327Z",
  "votes": 36,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>Acknowledgements</h1>\n<p>I would like to thank the organizers and Kaggle staff for hosting and running this excellent competition. I also extend my sincere gratitude to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for publishing outstanding baseline notebooks and discussions.</p>\n<h1>Overview</h1>\n<p>My pipeline uses <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">public implementation</a> for stage 0 and stage 1 without modifications. Therefore, this solution primarily focuses on stage 2 segmentation and post-processing techniques.</p>\n<p>Key innovations:</p>\n<ul>\n<li>Replace competition time-series data with original PTB-XL Dataset signals (500Hz)</li>\n<li>Predict signal sampling positions directly using sparse masks</li>\n<li>Build a 2.5D segmentation model (series model) that fuses phase and amplitude information across leads</li>\n<li>Develop a whole-image segmentation model (whole model) combining timm encoders with <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <code>MyCoordUnetDecoder</code></li>\n</ul>\n<p><strong>Note</strong>: In this solution writeup, I refer to each row of the standard 12-lead ECG as \"series (0-3)\".</p>\n<h1>Strategy</h1>\n<p>The competition metric computes SNR for each image, averages it in the linear SNR domain, and then converts it to SNR(dB)(<a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/663901\" target=\"_blank\">discussion</a>).</p>\n<p>As a result, pushing medium-to-high SNR images higher tends to improve the final score more than spending effort on the hardest low-SNR cases. After analyzing out-of-fold (OOF) predictions, I therefore focused on improving medium-difficulty and easy images, which were also more numerous and easier to improve.</p>\n<p>To minimize prediction errors by staying close to sampling points, I refined my approach in three key areas:</p>\n<ul>\n<li>Segmentation mask creation</li>\n<li>Modeling architecture</li>\n<li>Post-processing methods</li>\n</ul>\n<h1>Data</h1>\n<h2>Competition Data</h2>\n<p>The competition data is a subset of the <a href=\"https://physionet.org/content/ptb-xl/1.0.3/\" target=\"_blank\">PTB-XL Dataset</a>.</p>\n<p>While the PTB-XL dataset provides all time-series data at a 500Hz sampling frequency, the competition data uses multiple frequencies: 250, 256, 500, 512, 1000, and 1025Hz. This indicates that the competition organizers resampled the original 500Hz signals to these different frequencies.</p>\n<p>I matched the competition data back to the original PTB-XL time-series by computing correlation coefficients, and replaced all signals with their 500Hz originals.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F650cdd44cf06ff5bbd90b934718d14a6%2F500hz.png?generation=1769278608892672&amp;alt=media\" alt=\"500hz\"></p>\n<p>This replacement improved consistency in sampling positions, which led to noticeable CV improvements during training. A single fold model's LB score boosted from 21.67dB to 22.49dB.</p>\n<h2>Synthetic Data</h2>\n<p>Since I achieved satisfactory scores with the competition data alone, and further improvements from synthetic data were marginal, I did not invest much time in additional data synthesis (using PTB-XL dataset and ECG-Image-Kit).</p>\n<p>Considering that the private dataset might contain difficult <code>0015</code> type wrinkled data, I added a small amount of randomly selected <code>0015</code> data to the training set each epoch.</p>\n<h2>Segmentation Mask</h2>\n<p>The method for creating segmentation masks is crucial. This is because the SNR achieved when reconstructing signals through the pipeline (time-series → segmentation mask → time-series) directly correlates with the model's upper performance limit.</p>\n<p>Dense masks (covering the entire signal line) did not yield high reconstruction SNR. Instead, I created sparse masks that annotate at most 2 pixels per column.</p>\n<p>To enable reconstruction at sub-pixel precision in the post-processing stage, I distribute labels to two pixels (the integer part and integer part + 1) according to the fractional part of the y-coordinate. During reconstruction, I convert these to time-series data by computing a weighted average of y-coordinates using label values as weights.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F57bb4cef0794abd7b91aa65ee4c82235%2Fmask_recostruction.png?generation=1769278666595454&amp;alt=media\" alt=\"mask_recostruction\"></p>\n<p>To map from 500Hz (sig_len=5000) to masks without resampling, I set the mask width to 5600 pixels and drew masks in the range [301:5301], covering 5000 pixels.</p>\n<p>Below are the created masks and OOF prediction results from a model trained on these masks (yellow: GT, green: prediction).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F0c336a5c1dfe45c26386e58cefed6551%2Fmask_prediction_rev3.png?generation=1769341874859694&amp;alt=media\" alt=\"mask_prediction_rev3\"></p>\n<h1>Models</h1>\n<p>I used two types of models:</p>\n<ul>\n<li>whole model : A U-Net model combining timm encoders with <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <code>MyCoordUnetDecoder</code>.</li>\n<li>series model : A 2.5D segmentation model that fuses series images. I use this to share phase and amplitude information between leads.</li>\n</ul>\n<h2>Whole Model Architecture</h2>\n<p>I based the whole model architecture on <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s model. I changed the encoder to timm encoders from segmentation_models_pytorch. I cropped the top portion of the image and resized it to (1280, 5600).</p>\n<h2>Series Model Architecture</h2>\n<p>Specific ECG leads have strong correlations (e.g., Einthoven's Law). To incorporate these relationships into the model, I built a 2.5D model that takes four series as input.</p>\n<p>I crop each series within a ±3mV (240 pixel) range centered on the zero mV y-coordinate position. A shared U-Net encoder processes each series image, then I fuse features across series in the connection paths to each U-Net decoder layer. The decoder extracts features that the segmentation head converts to mask predictions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Febe06e7dd533692b750ead6625354935%2Fmodel_v5.png?generation=1769278736457754&amp;alt=media\" alt=\"model_v5\"></p>\n<h3>Fusion Module</h3>\n<p>I tested three fusion module variants:</p>\n<ul>\n<li>conv2d</li>\n<li>shared conv2d</li>\n<li>conv3d</li>\n</ul>\n<p>Typically, 2.5D models use conv3d, LSTM, or transformers to extract depth-wise information. LSTM did not work well (possibly due to poor parameter settings). I did not try transformers due to limited experience.</p>\n<p>Since I wanted to mix all depth (series stacking order) information together, I also experimented with feature fusion using conv2d. While differences between variants were small, conv2d showed the best CV results.</p>\n<h4>conv2d</h4>\n<p>I apply conv2d blocks (Conv2d→BatchNorm2d→ReLU) for feature reduction and feature fusion.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Fcae941cf33bcb2dc4e67179089b2be6d%2Fconv2d_fusion_v2.png?generation=1769280668345882&amp;alt=media\" alt=\"conv2d_fusion\"></p>\n<h4>shared conv2d</h4>\n<p>To save parameters, shared conv2d shares the reduce conv2d block across all series.</p>\n<h4>conv3d</h4>\n<p>I reshape features to (C, 4, H, W) and apply conv3d blocks (Conv3d→BatchNorm3d→LeakyReLU) multiple times.</p>\n<h2>Model Comparison</h2>\n<p>To verify the fusion module's effectiveness, I compared prediction results.</p>\n<p>The top row shows images with unmasked signals and GT masks. The middle and bottom rows show images with masked regions (simulating image artifacts) and predictions from the whole model and series model, respectively.</p>\n<p>While the whole model fails to predict the masked regions, the series model shares information across leads, enabling it to reasonably predict peak phase and amplitude.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Ffc67adf0cc00216fb93436f91d234fd6%2Fcombined_result.png?generation=1769278842918294&amp;alt=media\" alt=\"masked_prediction\"></p>\n<h1>Training Strategy</h1>\n<h2>Parameters</h2>\n<ul>\n<li>Input image: Stage 1 rectified image<ul>\n<li>whole model: input shape = (3, 1280, 5600)</li>\n<li>series model: input shape = (4, 3, 480, 5600)</li></ul></li>\n<li>Loss: BCEWithLogitsLoss(pos_weight=20)</li>\n<li>Epoch: 50</li>\n<li>Batch size: 4</li>\n<li>Optimizer: AdamW<ul>\n<li>Learning rate: 5e-4 ~ 1e-3</li>\n<li>Weight decay: 0.01</li></ul></li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<p><strong>Tip: Gradient checkpointing is effective for reducing activation memory consumption with high-resolution images</strong></p>\n<h2>Augmentation</h2>\n<pre><code>image_only_aug = A.Compose([\n    A.RandomBrightnessContrast(brightness_limit=(-0.1,0.1), contrast_limit=(-0.1, 0.1), p=0.2),\n    A.RandomShadow(p=0.2),\n    A.GaussianBlur(p=0.2),\n    A.CoarseDropout(num_holes_range=(1,8), hole_height_range=(0.01, 0.1), hole_width_range=(0.01, 0.05), fill=0, p=0.1),\n    A.ToGray(p=0.25),\n])\n</code></pre>\n<pre><code>image_and_mask_aug = A.HorizontalFlip(p=0.5)\n</code></pre>\n<h1>Post-Processing</h1>\n<h2>Mask Post-Processing</h2>\n<p>I compute a weighted average for each column of the mask using prediction results to obtain sub-pixel level predictions.</p>\n<h2>Resampling Method</h2>\n<p>I tested:</p>\n<ul>\n<li>scipy.signal.resample</li>\n<li>scipy.signal.resample_poly(padtype='line')</li>\n<li>torch.nn.functional.interpolate(mode=\"linear\")</li>\n</ul>\n<p>scipy.signal.resample gave the best results.</p>\n<h2>Pixel to Series</h2>\n<p>Code for the complete mask→time-series conversion pipeline, including mask post-processing and resampling:</p>\n<pre><code>def pixel_to_series(pixel, length):\n    _, H, W = pixel.shape\n    eps=1e-8\n    y_idx = np.arange(H, dtype=np.float32)[:, None]\n\n    series = []\n    for j in [0, 1, 2, 3]:\n        p = pixel[j]\n        denom = p.sum(axis=0)\n        y_exp = (p * y_idx).sum(axis=0) / (denom + eps)\n        series.append(y_exp)\n    series = np.stack(series).astype(np.float32)\n\n    if length!=W:\n        resampled_series = []\n        for s in series:\n            rs = signal.resample(s, length).astype(np.float32)\n            resampled_series.append(rs)\n        series = np.stack(resampled_series)\n    return series\n</code></pre>\n<h1>Inference</h1>\n<h2>TTA</h2>\n<ul>\n<li>Horizontal flip</li>\n</ul>\n<h2>Results</h2>\n<h3>Submission A (6 models)</h3>\n<p><strong>Public LB:</strong> 23.37, <strong>Private LB:</strong> 23.27</p>\n<h3>Submission B (11 models)</h3>\n<p><strong>Public LB:</strong> 23.34, <strong>Private LB:</strong> 23.23</p>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>Model (fusion module)</th>\n<th>Synthetic Data</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Submission A</th>\n<th>Submission B</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet B7</td>\n<td>whole</td>\n<td></td>\n<td>22.93</td>\n<td>22.65</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 L</td>\n<td>whole</td>\n<td>✅</td>\n<td>22.60</td>\n<td>22.55</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>whole</td>\n<td></td>\n<td>22.52</td>\n<td>22.38</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>whole</td>\n<td>✅</td>\n<td>22.58</td>\n<td>22.39</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>series (shared conv2d)</td>\n<td>✅</td>\n<td>23.10</td>\n<td>22.92</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>series (shared conv2d)</td>\n<td></td>\n<td>23.00</td>\n<td>22.81</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNet B4</td>\n<td>series (shared conv2d)</td>\n<td>✅</td>\n<td>22.93</td>\n<td>22.73</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 L</td>\n<td>series (conv3d)</td>\n<td></td>\n<td>22.92</td>\n<td>22.77</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 L</td>\n<td>series (conv2d)</td>\n<td></td>\n<td>22.85</td>\n<td>22.70</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>series (conv3d)</td>\n<td>✅</td>\n<td>22.73</td>\n<td>22.49</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>series (conv2d)</td>\n<td>✅</td>\n<td>22.76</td>\n<td>22.55</td>\n<td></td>\n<td>☑️</td>\n</tr>\n</tbody>\n</table>\n<h1>References</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">https://www.kaggle.com/code/hengck23/demo-submission</a></li>\n<li>Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954</li>\n<li>Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.</li>\n<li>Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., &amp; Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. <a href=\"https://doi.org/10.13026/kfzx-aw45\" target=\"_blank\">https://doi.org/10.13026/kfzx-aw45</a></li>\n<li>Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F.I., Samek, W., Schaeffter, T. (2020), PTB-XL: A Large Publicly Available ECG Dataset. Scientific Data. <a href=\"https://doi.org/10.1038/s41597-020-0495-6\" target=\"_blank\">https://doi.org/10.1038/s41597-020-0495-6</a></li>\n<li>Goldberger, A., Amaral, L., Glass, L., Hausdorff, J., Ivanov, P. C., Mark, R., … &amp; Stanley, H. E. (2000). PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation [Online]. 101 (23), pp. e215–e220. RRID:SCR_007345.</li>\n</ul>\n<h1>Code Availability</h1>\n<ul>\n<li>Training code: <a href=\"https://github.com/someya-takashi/physionet-image-digitization/tree/main\" target=\"_blank\">https://github.com/someya-takashi/physionet-image-digitization/tree/main</a></li>\n<li>Inference code: <a href=\"https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission\" target=\"_blank\">https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission</a></li>\n</ul>",
  "messages": [
    {
      "id": "3396350",
      "postDate": "01/24/2026 19:17:27",
      "content": "<h1>Acknowledgements</h1>\n<p>I would like to thank the organizers and Kaggle staff for hosting and running this excellent competition. I also extend my sincere gratitude to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for publishing outstanding baseline notebooks and discussions.</p>\n<h1>Overview</h1>\n<p>My pipeline uses <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">public implementation</a> for stage 0 and stage 1 without modifications. Therefore, this solution primarily focuses on stage 2 segmentation and post-processing techniques.</p>\n<p>Key innovations:</p>\n<ul>\n<li>Replace competition time-series data with original PTB-XL Dataset signals (500Hz)</li>\n<li>Predict signal sampling positions directly using sparse masks</li>\n<li>Build a 2.5D segmentation model (series model) that fuses phase and amplitude information across leads</li>\n<li>Develop a whole-image segmentation model (whole model) combining timm encoders with <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <code>MyCoordUnetDecoder</code></li>\n</ul>\n<p><strong>Note</strong>: In this solution writeup, I refer to each row of the standard 12-lead ECG as \"series (0-3)\".</p>\n<h1>Strategy</h1>\n<p>The competition metric computes SNR for each image, averages it in the linear SNR domain, and then converts it to SNR(dB)(<a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/663901\" target=\"_blank\">discussion</a>).</p>\n<p>As a result, pushing medium-to-high SNR images higher tends to improve the final score more than spending effort on the hardest low-SNR cases. After analyzing out-of-fold (OOF) predictions, I therefore focused on improving medium-difficulty and easy images, which were also more numerous and easier to improve.</p>\n<p>To minimize prediction errors by staying close to sampling points, I refined my approach in three key areas:</p>\n<ul>\n<li>Segmentation mask creation</li>\n<li>Modeling architecture</li>\n<li>Post-processing methods</li>\n</ul>\n<h1>Data</h1>\n<h2>Competition Data</h2>\n<p>The competition data is a subset of the <a href=\"https://physionet.org/content/ptb-xl/1.0.3/\" target=\"_blank\">PTB-XL Dataset</a>.</p>\n<p>While the PTB-XL dataset provides all time-series data at a 500Hz sampling frequency, the competition data uses multiple frequencies: 250, 256, 500, 512, 1000, and 1025Hz. This indicates that the competition organizers resampled the original 500Hz signals to these different frequencies.</p>\n<p>I matched the competition data back to the original PTB-XL time-series by computing correlation coefficients, and replaced all signals with their 500Hz originals.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F650cdd44cf06ff5bbd90b934718d14a6%2F500hz.png?generation=1769278608892672&amp;alt=media\" alt=\"500hz\"></p>\n<p>This replacement improved consistency in sampling positions, which led to noticeable CV improvements during training. A single fold model's LB score boosted from 21.67dB to 22.49dB.</p>\n<h2>Synthetic Data</h2>\n<p>Since I achieved satisfactory scores with the competition data alone, and further improvements from synthetic data were marginal, I did not invest much time in additional data synthesis (using PTB-XL dataset and ECG-Image-Kit).</p>\n<p>Considering that the private dataset might contain difficult <code>0015</code> type wrinkled data, I added a small amount of randomly selected <code>0015</code> data to the training set each epoch.</p>\n<h2>Segmentation Mask</h2>\n<p>The method for creating segmentation masks is crucial. This is because the SNR achieved when reconstructing signals through the pipeline (time-series → segmentation mask → time-series) directly correlates with the model's upper performance limit.</p>\n<p>Dense masks (covering the entire signal line) did not yield high reconstruction SNR. Instead, I created sparse masks that annotate at most 2 pixels per column.</p>\n<p>To enable reconstruction at sub-pixel precision in the post-processing stage, I distribute labels to two pixels (the integer part and integer part + 1) according to the fractional part of the y-coordinate. During reconstruction, I convert these to time-series data by computing a weighted average of y-coordinates using label values as weights.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F57bb4cef0794abd7b91aa65ee4c82235%2Fmask_recostruction.png?generation=1769278666595454&amp;alt=media\" alt=\"mask_recostruction\"></p>\n<p>To map from 500Hz (sig_len=5000) to masks without resampling, I set the mask width to 5600 pixels and drew masks in the range [301:5301], covering 5000 pixels.</p>\n<p>Below are the created masks and OOF prediction results from a model trained on these masks (yellow: GT, green: prediction).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F0c336a5c1dfe45c26386e58cefed6551%2Fmask_prediction_rev3.png?generation=1769341874859694&amp;alt=media\" alt=\"mask_prediction_rev3\"></p>\n<h1>Models</h1>\n<p>I used two types of models:</p>\n<ul>\n<li>whole model : A U-Net model combining timm encoders with <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <code>MyCoordUnetDecoder</code>.</li>\n<li>series model : A 2.5D segmentation model that fuses series images. I use this to share phase and amplitude information between leads.</li>\n</ul>\n<h2>Whole Model Architecture</h2>\n<p>I based the whole model architecture on <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s model. I changed the encoder to timm encoders from segmentation_models_pytorch. I cropped the top portion of the image and resized it to (1280, 5600).</p>\n<h2>Series Model Architecture</h2>\n<p>Specific ECG leads have strong correlations (e.g., Einthoven's Law). To incorporate these relationships into the model, I built a 2.5D model that takes four series as input.</p>\n<p>I crop each series within a ±3mV (240 pixel) range centered on the zero mV y-coordinate position. A shared U-Net encoder processes each series image, then I fuse features across series in the connection paths to each U-Net decoder layer. The decoder extracts features that the segmentation head converts to mask predictions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Febe06e7dd533692b750ead6625354935%2Fmodel_v5.png?generation=1769278736457754&amp;alt=media\" alt=\"model_v5\"></p>\n<h3>Fusion Module</h3>\n<p>I tested three fusion module variants:</p>\n<ul>\n<li>conv2d</li>\n<li>shared conv2d</li>\n<li>conv3d</li>\n</ul>\n<p>Typically, 2.5D models use conv3d, LSTM, or transformers to extract depth-wise information. LSTM did not work well (possibly due to poor parameter settings). I did not try transformers due to limited experience.</p>\n<p>Since I wanted to mix all depth (series stacking order) information together, I also experimented with feature fusion using conv2d. While differences between variants were small, conv2d showed the best CV results.</p>\n<h4>conv2d</h4>\n<p>I apply conv2d blocks (Conv2d→BatchNorm2d→ReLU) for feature reduction and feature fusion.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Fcae941cf33bcb2dc4e67179089b2be6d%2Fconv2d_fusion_v2.png?generation=1769280668345882&amp;alt=media\" alt=\"conv2d_fusion\"></p>\n<h4>shared conv2d</h4>\n<p>To save parameters, shared conv2d shares the reduce conv2d block across all series.</p>\n<h4>conv3d</h4>\n<p>I reshape features to (C, 4, H, W) and apply conv3d blocks (Conv3d→BatchNorm3d→LeakyReLU) multiple times.</p>\n<h2>Model Comparison</h2>\n<p>To verify the fusion module's effectiveness, I compared prediction results.</p>\n<p>The top row shows images with unmasked signals and GT masks. The middle and bottom rows show images with masked regions (simulating image artifacts) and predictions from the whole model and series model, respectively.</p>\n<p>While the whole model fails to predict the masked regions, the series model shares information across leads, enabling it to reasonably predict peak phase and amplitude.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Ffc67adf0cc00216fb93436f91d234fd6%2Fcombined_result.png?generation=1769278842918294&amp;alt=media\" alt=\"masked_prediction\"></p>\n<h1>Training Strategy</h1>\n<h2>Parameters</h2>\n<ul>\n<li>Input image: Stage 1 rectified image<ul>\n<li>whole model: input shape = (3, 1280, 5600)</li>\n<li>series model: input shape = (4, 3, 480, 5600)</li></ul></li>\n<li>Loss: BCEWithLogitsLoss(pos_weight=20)</li>\n<li>Epoch: 50</li>\n<li>Batch size: 4</li>\n<li>Optimizer: AdamW<ul>\n<li>Learning rate: 5e-4 ~ 1e-3</li>\n<li>Weight decay: 0.01</li></ul></li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<p><strong>Tip: Gradient checkpointing is effective for reducing activation memory consumption with high-resolution images</strong></p>\n<h2>Augmentation</h2>\n<pre><code>image_only_aug = A.Compose([\n    A.RandomBrightnessContrast(brightness_limit=(-0.1,0.1), contrast_limit=(-0.1, 0.1), p=0.2),\n    A.RandomShadow(p=0.2),\n    A.GaussianBlur(p=0.2),\n    A.CoarseDropout(num_holes_range=(1,8), hole_height_range=(0.01, 0.1), hole_width_range=(0.01, 0.05), fill=0, p=0.1),\n    A.ToGray(p=0.25),\n])\n</code></pre>\n<pre><code>image_and_mask_aug = A.HorizontalFlip(p=0.5)\n</code></pre>\n<h1>Post-Processing</h1>\n<h2>Mask Post-Processing</h2>\n<p>I compute a weighted average for each column of the mask using prediction results to obtain sub-pixel level predictions.</p>\n<h2>Resampling Method</h2>\n<p>I tested:</p>\n<ul>\n<li>scipy.signal.resample</li>\n<li>scipy.signal.resample_poly(padtype='line')</li>\n<li>torch.nn.functional.interpolate(mode=\"linear\")</li>\n</ul>\n<p>scipy.signal.resample gave the best results.</p>\n<h2>Pixel to Series</h2>\n<p>Code for the complete mask→time-series conversion pipeline, including mask post-processing and resampling:</p>\n<pre><code>def pixel_to_series(pixel, length):\n    _, H, W = pixel.shape\n    eps=1e-8\n    y_idx = np.arange(H, dtype=np.float32)[:, None]\n\n    series = []\n    for j in [0, 1, 2, 3]:\n        p = pixel[j]\n        denom = p.sum(axis=0)\n        y_exp = (p * y_idx).sum(axis=0) / (denom + eps)\n        series.append(y_exp)\n    series = np.stack(series).astype(np.float32)\n\n    if length!=W:\n        resampled_series = []\n        for s in series:\n            rs = signal.resample(s, length).astype(np.float32)\n            resampled_series.append(rs)\n        series = np.stack(resampled_series)\n    return series\n</code></pre>\n<h1>Inference</h1>\n<h2>TTA</h2>\n<ul>\n<li>Horizontal flip</li>\n</ul>\n<h2>Results</h2>\n<h3>Submission A (6 models)</h3>\n<p><strong>Public LB:</strong> 23.37, <strong>Private LB:</strong> 23.27</p>\n<h3>Submission B (11 models)</h3>\n<p><strong>Public LB:</strong> 23.34, <strong>Private LB:</strong> 23.23</p>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>Model (fusion module)</th>\n<th>Synthetic Data</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Submission A</th>\n<th>Submission B</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet B7</td>\n<td>whole</td>\n<td></td>\n<td>22.93</td>\n<td>22.65</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 L</td>\n<td>whole</td>\n<td>✅</td>\n<td>22.60</td>\n<td>22.55</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>whole</td>\n<td></td>\n<td>22.52</td>\n<td>22.38</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>whole</td>\n<td>✅</td>\n<td>22.58</td>\n<td>22.39</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>series (shared conv2d)</td>\n<td>✅</td>\n<td>23.10</td>\n<td>22.92</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNet B6</td>\n<td>series (shared conv2d)</td>\n<td></td>\n<td>23.00</td>\n<td>22.81</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNet B4</td>\n<td>series (shared conv2d)</td>\n<td>✅</td>\n<td>22.93</td>\n<td>22.73</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 L</td>\n<td>series (conv3d)</td>\n<td></td>\n<td>22.92</td>\n<td>22.77</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 L</td>\n<td>series (conv2d)</td>\n<td></td>\n<td>22.85</td>\n<td>22.70</td>\n<td>☑️</td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>series (conv3d)</td>\n<td>✅</td>\n<td>22.73</td>\n<td>22.49</td>\n<td></td>\n<td>☑️</td>\n</tr>\n<tr>\n<td>EfficientNetV2 M</td>\n<td>series (conv2d)</td>\n<td>✅</td>\n<td>22.76</td>\n<td>22.55</td>\n<td></td>\n<td>☑️</td>\n</tr>\n</tbody>\n</table>\n<h1>References</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/demo-submission\" target=\"_blank\">https://www.kaggle.com/code/hengck23/demo-submission</a></li>\n<li>Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954</li>\n<li>Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.</li>\n<li>Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., &amp; Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. <a href=\"https://doi.org/10.13026/kfzx-aw45\" target=\"_blank\">https://doi.org/10.13026/kfzx-aw45</a></li>\n<li>Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F.I., Samek, W., Schaeffter, T. (2020), PTB-XL: A Large Publicly Available ECG Dataset. Scientific Data. <a href=\"https://doi.org/10.1038/s41597-020-0495-6\" target=\"_blank\">https://doi.org/10.1038/s41597-020-0495-6</a></li>\n<li>Goldberger, A., Amaral, L., Glass, L., Hausdorff, J., Ivanov, P. C., Mark, R., … &amp; Stanley, H. E. (2000). PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation [Online]. 101 (23), pp. e215–e220. RRID:SCR_007345.</li>\n</ul>\n<h1>Code Availability</h1>\n<ul>\n<li>Training code: <a href=\"https://github.com/someya-takashi/physionet-image-digitization/tree/main\" target=\"_blank\">https://github.com/someya-takashi/physionet-image-digitization/tree/main</a></li>\n<li>Inference code: <a href=\"https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission\" target=\"_blank\">https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission</a></li>\n</ul>",
      "rawMarkdown": "# Acknowledgements\n\nI would like to thank the organizers and Kaggle staff for hosting and running this excellent competition. I also extend my sincere gratitude to @hengck23 for publishing outstanding baseline notebooks and discussions.\n\n# Overview\n\nMy pipeline uses @hengck23's [public implementation](https://www.kaggle.com/code/hengck23/demo-submission) for stage 0 and stage 1 without modifications. Therefore, this solution primarily focuses on stage 2 segmentation and post-processing techniques.\n\nKey innovations:\n- Replace competition time-series data with original PTB-XL Dataset signals (500Hz)\n- Predict signal sampling positions directly using sparse masks\n- Build a 2.5D segmentation model (series model) that fuses phase and amplitude information across leads\n- Develop a whole-image segmentation model (whole model) combining timm encoders with @hengck23's `MyCoordUnetDecoder`\n\n**Note**: In this solution writeup, I refer to each row of the standard 12-lead ECG as \"series (0-3)\".\n\n# Strategy\n\nThe competition metric computes SNR for each image, averages it in the linear SNR domain, and then converts it to SNR(dB)([discussion](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/663901)).\n\nAs a result, pushing medium-to-high SNR images higher tends to improve the final score more than spending effort on the hardest low-SNR cases. After analyzing out-of-fold (OOF) predictions, I therefore focused on improving medium-difficulty and easy images, which were also more numerous and easier to improve.\n\nTo minimize prediction errors by staying close to sampling points, I refined my approach in three key areas:\n- Segmentation mask creation\n- Modeling architecture\n- Post-processing methods\n\n# Data\n\n## Competition Data\n\nThe competition data is a subset of the [PTB-XL Dataset](https://physionet.org/content/ptb-xl/1.0.3/).\n\nWhile the PTB-XL dataset provides all time-series data at a 500Hz sampling frequency, the competition data uses multiple frequencies: 250, 256, 500, 512, 1000, and 1025Hz. This indicates that the competition organizers resampled the original 500Hz signals to these different frequencies.\n\nI matched the competition data back to the original PTB-XL time-series by computing correlation coefficients, and replaced all signals with their 500Hz originals.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F650cdd44cf06ff5bbd90b934718d14a6%2F500hz.png?generation=1769278608892672&alt=media\" alt=\"500hz\" width=\"526\" height=\"220\">\n\nThis replacement improved consistency in sampling positions, which led to noticeable CV improvements during training. A single fold model's LB score boosted from 21.67dB to 22.49dB.\n\n## Synthetic Data\n\nSince I achieved satisfactory scores with the competition data alone, and further improvements from synthetic data were marginal, I did not invest much time in additional data synthesis (using PTB-XL dataset and ECG-Image-Kit).\n\nConsidering that the private dataset might contain difficult `0015` type wrinkled data, I added a small amount of randomly selected `0015` data to the training set each epoch.\n\n## Segmentation Mask\n\nThe method for creating segmentation masks is crucial. This is because the SNR achieved when reconstructing signals through the pipeline (time-series → segmentation mask → time-series) directly correlates with the model's upper performance limit.\n\nDense masks (covering the entire signal line) did not yield high reconstruction SNR. Instead, I created sparse masks that annotate at most 2 pixels per column.\n\nTo enable reconstruction at sub-pixel precision in the post-processing stage, I distribute labels to two pixels (the integer part and integer part + 1) according to the fractional part of the y-coordinate. During reconstruction, I convert these to time-series data by computing a weighted average of y-coordinates using label values as weights.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F57bb4cef0794abd7b91aa65ee4c82235%2Fmask_recostruction.png?generation=1769278666595454&alt=media\" alt=\"mask_recostruction\" width=\"730\" height=\"331\">\n\nTo map from 500Hz (sig_len=5000) to masks without resampling, I set the mask width to 5600 pixels and drew masks in the range [301:5301], covering 5000 pixels.\n\nBelow are the created masks and OOF prediction results from a model trained on these masks (yellow: GT, green: prediction).\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F0c336a5c1dfe45c26386e58cefed6551%2Fmask_prediction_rev3.png?generation=1769341874859694&alt=media\" alt=\"mask_prediction_rev3\" width=\"910\" height=\"310\">\n\n# Models\n\nI used two types of models:\n\n- whole model : A U-Net model combining timm encoders with @hengck23's `MyCoordUnetDecoder`.\n- series model : A 2.5D segmentation model that fuses series images. I use this to share phase and amplitude information between leads.\n\n## Whole Model Architecture\n\nI based the whole model architecture on @hengck23's model. I changed the encoder to timm encoders from segmentation_models_pytorch. I cropped the top portion of the image and resized it to (1280, 5600).\n\n## Series Model Architecture\n\nSpecific ECG leads have strong correlations (e.g., Einthoven's Law). To incorporate these relationships into the model, I built a 2.5D model that takes four series as input.\n\nI crop each series within a ±3mV (240 pixel) range centered on the zero mV y-coordinate position. A shared U-Net encoder processes each series image, then I fuse features across series in the connection paths to each U-Net decoder layer. The decoder extracts features that the segmentation head converts to mask predictions.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Febe06e7dd533692b750ead6625354935%2Fmodel_v5.png?generation=1769278736457754&alt=media\" alt=\"model_v5\" width=\"863\" height=\"472\">\n\n### Fusion Module\n\nI tested three fusion module variants:\n- conv2d\n- shared conv2d\n- conv3d\n\nTypically, 2.5D models use conv3d, LSTM, or transformers to extract depth-wise information. LSTM did not work well (possibly due to poor parameter settings). I did not try transformers due to limited experience.\n\nSince I wanted to mix all depth (series stacking order) information together, I also experimented with feature fusion using conv2d. While differences between variants were small, conv2d showed the best CV results.\n\n#### conv2d\nI apply conv2d blocks (Conv2d→BatchNorm2d→ReLU) for feature reduction and feature fusion.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Fcae941cf33bcb2dc4e67179089b2be6d%2Fconv2d_fusion_v2.png?generation=1769280668345882&alt=media\" alt=\"conv2d_fusion\" width=\"899\" height=\"382\">\n\n#### shared conv2d\nTo save parameters, shared conv2d shares the reduce conv2d block across all series.\n\n#### conv3d\nI reshape features to (C, 4, H, W) and apply conv3d blocks (Conv3d→BatchNorm3d→LeakyReLU) multiple times.\n\n## Model Comparison\n\nTo verify the fusion module's effectiveness, I compared prediction results.\n\nThe top row shows images with unmasked signals and GT masks. The middle and bottom rows show images with masked regions (simulating image artifacts) and predictions from the whole model and series model, respectively.\n\nWhile the whole model fails to predict the masked regions, the series model shares information across leads, enabling it to reasonably predict peak phase and amplitude.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Ffc67adf0cc00216fb93436f91d234fd6%2Fcombined_result.png?generation=1769278842918294&alt=media\" alt=\"masked_prediction\" width=\"800\" height=\"540\">\n\n# Training Strategy\n\n## Parameters\n- Input image: Stage 1 rectified image\n  - whole model: input shape = (3, 1280, 5600)\n  - series model: input shape = (4, 3, 480, 5600)\n- Loss: BCEWithLogitsLoss(pos_weight=20)\n- Epoch: 50\n- Batch size: 4\n- Optimizer: AdamW\n  - Learning rate: 5e-4 ~ 1e-3\n  - Weight decay: 0.01\n- Scheduler: CosineAnnealingLR\n\n**Tip: Gradient checkpointing is effective for reducing activation memory consumption with high-resolution images**\n\n## Augmentation\n```python\nimage_only_aug = A.Compose([\n    A.RandomBrightnessContrast(brightness_limit=(-0.1,0.1), contrast_limit=(-0.1, 0.1), p=0.2),\n    A.RandomShadow(p=0.2),\n    A.GaussianBlur(p=0.2),\n    A.CoarseDropout(num_holes_range=(1,8), hole_height_range=(0.01, 0.1), hole_width_range=(0.01, 0.05), fill=0, p=0.1),\n    A.ToGray(p=0.25),\n])\n```\n```python\nimage_and_mask_aug = A.HorizontalFlip(p=0.5)\n```\n\n# Post-Processing\n\n## Mask Post-Processing\nI compute a weighted average for each column of the mask using prediction results to obtain sub-pixel level predictions.\n\n## Resampling Method\nI tested:\n- scipy.signal.resample\n- scipy.signal.resample_poly(padtype='line')\n- torch.nn.functional.interpolate(mode=\"linear\")\n\nscipy.signal.resample gave the best results.\n\n## Pixel to Series\nCode for the complete mask→time-series conversion pipeline, including mask post-processing and resampling:\n\n```python\ndef pixel_to_series(pixel, length):\n    _, H, W = pixel.shape\n    eps=1e-8\n    y_idx = np.arange(H, dtype=np.float32)[:, None]\n\n    series = []\n    for j in [0, 1, 2, 3]:\n        p = pixel[j]\n        denom = p.sum(axis=0)\n        y_exp = (p * y_idx).sum(axis=0) / (denom + eps)\n        series.append(y_exp)\n    series = np.stack(series).astype(np.float32)\n\n    if length!=W:\n        resampled_series = []\n        for s in series:\n            rs = signal.resample(s, length).astype(np.float32)\n            resampled_series.append(rs)\n        series = np.stack(resampled_series)\n    return series\n```\n\n# Inference\n\n## TTA\n- Horizontal flip\n\n## Results\n\n### Submission A (6 models)\n**Public LB:** 23.37, **Private LB:** 23.27\n\n### Submission B (11 models)\n**Public LB:** 23.34, **Private LB:** 23.23\n\n| Encoder                       | Model (fusion module)    | Synthetic Data | Public LB | Private LB | Submission A | Submission B |\n|-------------------------------|------------------------|----------------|-----------|------------|-------|-------|\n| EfficientNet B7               | whole                  |                | 22.93     | 22.65      | ☑️    | ☑️    |\n| EfficientNetV2 L             | whole                  | ✅             | 22.60     | 22.55      | ☑️    | ☑️    |\n| EfficientNetV2 M             | whole                  |                | 22.52     | 22.38      |       | ☑️    |\n| EfficientNetV2 M             | whole                  | ✅             | 22.58     | 22.39      |       | ☑️    |\n| EfficientNet B6               | series (shared conv2d) | ✅             | 23.10     | 22.92      | ☑️    | ☑️    |\n| EfficientNet B6               | series (shared conv2d) |                | 23.00     | 22.81      | ☑️    | ☑️    |\n| EfficientNet B4               | series (shared conv2d) | ✅             | 22.93     | 22.73      |       | ☑️    |\n| EfficientNetV2 L             | series (conv3d)        |                | 22.92     | 22.77      | ☑️    | ☑️    |\n| EfficientNetV2 L             | series (conv2d)        |                | 22.85     | 22.70      | ☑️    | ☑️    |\n| EfficientNetV2 M             | series (conv3d)        | ✅             | 22.73     | 22.49      |       | ☑️    |\n| EfficientNetV2 M             | series (conv2d)        | ✅             | 22.76     | 22.55      |       | ☑️    |\n\n# References\n\n- https://www.kaggle.com/code/hengck23/demo-submission\n- Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954\n- Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.\n- Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., & Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/kfzx-aw45\n- Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F.I., Samek, W., Schaeffter, T. (2020), PTB-XL: A Large Publicly Available ECG Dataset. Scientific Data. https://doi.org/10.1038/s41597-020-0495-6\n- Goldberger, A., Amaral, L., Glass, L., Hausdorff, J., Ivanov, P. C., Mark, R., ... & Stanley, H. E. (2000). PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation [Online]. 101 (23), pp. e215–e220. RRID:SCR_007345.\n\n# Code Availability\n\n- Training code: https://github.com/someya-takashi/physionet-image-digitization/tree/main\n- Inference code: https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission",
      "votes": null
    },
    {
      "id": "3396393",
      "postDate": "01/24/2026 22:55:11",
      "content": "<p>Congrats! Using original sampling data is smart, we noticed competition data is re-sampled and this will impact the result but do nothing 😅.</p>",
      "rawMarkdown": "Congrats! Using original sampling data is smart, we noticed competition data is re-sampled and this will impact the result but do nothing 😅.",
      "votes": null
    },
    {
      "id": "3396579",
      "postDate": "01/25/2026 10:52:54",
      "content": "<p>Thanks! And congrats on 1st place!\nYour team’s approach of predicting the signal directly is really insightful—I learned a lot from it! 😄</p>",
      "rawMarkdown": "Thanks! And congrats on 1st place!\nYour team’s approach of predicting the signal directly is really insightful—I learned a lot from it! 😄",
      "votes": null
    },
    {
      "id": "3401307",
      "postDate": "02/03/2026 12:27:13",
      "content": "<p>Finally published my training and inference code.</p>\n<ul>\n<li>Training code: <a href=\"https://github.com/someya-takashi/physionet-image-digitization/tree/main\" target=\"_blank\">https://github.com/someya-takashi/physionet-image-digitization/tree/main</a></li>\n<li>Inference code: <a href=\"https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission\" target=\"_blank\">https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission</a></li>\n</ul>\n<p>Hope it helps!</p>",
      "rawMarkdown": "Finally published my training and inference code.\n\n- Training code: https://github.com/someya-takashi/physionet-image-digitization/tree/main\n- Inference code: https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission\n\nHope it helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3396393,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "01/24/2026 22:55:11",
      "content": "<p>Congrats! Using original sampling data is smart, we noticed competition data is re-sampled and this will impact the result but do nothing 😅.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3396579,
          "author_name": "takashisomeya",
          "author_url": "",
          "post_date": "01/25/2026 10:52:54",
          "content": "<p>Thanks! And congrats on 1st place!\nYour team’s approach of predicting the signal directly is really insightful—I learned a lot from it! 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3401307,
      "author_name": "takashisomeya",
      "author_url": "",
      "post_date": "02/03/2026 12:27:13",
      "content": "<p>Finally published my training and inference code.</p>\n<ul>\n<li>Training code: <a href=\"https://github.com/someya-takashi/physionet-image-digitization/tree/main\" target=\"_blank\">https://github.com/someya-takashi/physionet-image-digitization/tree/main</a></li>\n<li>Inference code: <a href=\"https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission\" target=\"_blank\">https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission</a></li>\n</ul>\n<p>Hope it helps!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3396350": "# Acknowledgements\n\nI would like to thank the organizers and Kaggle staff for hosting and running this excellent competition. I also extend my sincere gratitude to @hengck23 for publishing outstanding baseline notebooks and discussions.\n\n# Overview\n\nMy pipeline uses @hengck23's [public implementation](https://www.kaggle.com/code/hengck23/demo-submission) for stage 0 and stage 1 without modifications. Therefore, this solution primarily focuses on stage 2 segmentation and post-processing techniques.\n\nKey innovations:\n- Replace competition time-series data with original PTB-XL Dataset signals (500Hz)\n- Predict signal sampling positions directly using sparse masks\n- Build a 2.5D segmentation model (series model) that fuses phase and amplitude information across leads\n- Develop a whole-image segmentation model (whole model) combining timm encoders with @hengck23's `MyCoordUnetDecoder`\n\n**Note**: In this solution writeup, I refer to each row of the standard 12-lead ECG as \"series (0-3)\".\n\n# Strategy\n\nThe competition metric computes SNR for each image, averages it in the linear SNR domain, and then converts it to SNR(dB)([discussion](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/663901)).\n\nAs a result, pushing medium-to-high SNR images higher tends to improve the final score more than spending effort on the hardest low-SNR cases. After analyzing out-of-fold (OOF) predictions, I therefore focused on improving medium-difficulty and easy images, which were also more numerous and easier to improve.\n\nTo minimize prediction errors by staying close to sampling points, I refined my approach in three key areas:\n- Segmentation mask creation\n- Modeling architecture\n- Post-processing methods\n\n# Data\n\n## Competition Data\n\nThe competition data is a subset of the [PTB-XL Dataset](https://physionet.org/content/ptb-xl/1.0.3/).\n\nWhile the PTB-XL dataset provides all time-series data at a 500Hz sampling frequency, the competition data uses multiple frequencies: 250, 256, 500, 512, 1000, and 1025Hz. This indicates that the competition organizers resampled the original 500Hz signals to these different frequencies.\n\nI matched the competition data back to the original PTB-XL time-series by computing correlation coefficients, and replaced all signals with their 500Hz originals.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F650cdd44cf06ff5bbd90b934718d14a6%2F500hz.png?generation=1769278608892672&alt=media\" alt=\"500hz\" width=\"526\" height=\"220\">\n\nThis replacement improved consistency in sampling positions, which led to noticeable CV improvements during training. A single fold model's LB score boosted from 21.67dB to 22.49dB.\n\n## Synthetic Data\n\nSince I achieved satisfactory scores with the competition data alone, and further improvements from synthetic data were marginal, I did not invest much time in additional data synthesis (using PTB-XL dataset and ECG-Image-Kit).\n\nConsidering that the private dataset might contain difficult `0015` type wrinkled data, I added a small amount of randomly selected `0015` data to the training set each epoch.\n\n## Segmentation Mask\n\nThe method for creating segmentation masks is crucial. This is because the SNR achieved when reconstructing signals through the pipeline (time-series → segmentation mask → time-series) directly correlates with the model's upper performance limit.\n\nDense masks (covering the entire signal line) did not yield high reconstruction SNR. Instead, I created sparse masks that annotate at most 2 pixels per column.\n\nTo enable reconstruction at sub-pixel precision in the post-processing stage, I distribute labels to two pixels (the integer part and integer part + 1) according to the fractional part of the y-coordinate. During reconstruction, I convert these to time-series data by computing a weighted average of y-coordinates using label values as weights.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F57bb4cef0794abd7b91aa65ee4c82235%2Fmask_recostruction.png?generation=1769278666595454&alt=media\" alt=\"mask_recostruction\" width=\"730\" height=\"331\">\n\nTo map from 500Hz (sig_len=5000) to masks without resampling, I set the mask width to 5600 pixels and drew masks in the range [301:5301], covering 5000 pixels.\n\nBelow are the created masks and OOF prediction results from a model trained on these masks (yellow: GT, green: prediction).\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2F0c336a5c1dfe45c26386e58cefed6551%2Fmask_prediction_rev3.png?generation=1769341874859694&alt=media\" alt=\"mask_prediction_rev3\" width=\"910\" height=\"310\">\n\n# Models\n\nI used two types of models:\n\n- whole model : A U-Net model combining timm encoders with @hengck23's `MyCoordUnetDecoder`.\n- series model : A 2.5D segmentation model that fuses series images. I use this to share phase and amplitude information between leads.\n\n## Whole Model Architecture\n\nI based the whole model architecture on @hengck23's model. I changed the encoder to timm encoders from segmentation_models_pytorch. I cropped the top portion of the image and resized it to (1280, 5600).\n\n## Series Model Architecture\n\nSpecific ECG leads have strong correlations (e.g., Einthoven's Law). To incorporate these relationships into the model, I built a 2.5D model that takes four series as input.\n\nI crop each series within a ±3mV (240 pixel) range centered on the zero mV y-coordinate position. A shared U-Net encoder processes each series image, then I fuse features across series in the connection paths to each U-Net decoder layer. The decoder extracts features that the segmentation head converts to mask predictions.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Febe06e7dd533692b750ead6625354935%2Fmodel_v5.png?generation=1769278736457754&alt=media\" alt=\"model_v5\" width=\"863\" height=\"472\">\n\n### Fusion Module\n\nI tested three fusion module variants:\n- conv2d\n- shared conv2d\n- conv3d\n\nTypically, 2.5D models use conv3d, LSTM, or transformers to extract depth-wise information. LSTM did not work well (possibly due to poor parameter settings). I did not try transformers due to limited experience.\n\nSince I wanted to mix all depth (series stacking order) information together, I also experimented with feature fusion using conv2d. While differences between variants were small, conv2d showed the best CV results.\n\n#### conv2d\nI apply conv2d blocks (Conv2d→BatchNorm2d→ReLU) for feature reduction and feature fusion.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Fcae941cf33bcb2dc4e67179089b2be6d%2Fconv2d_fusion_v2.png?generation=1769280668345882&alt=media\" alt=\"conv2d_fusion\" width=\"899\" height=\"382\">\n\n#### shared conv2d\nTo save parameters, shared conv2d shares the reduce conv2d block across all series.\n\n#### conv3d\nI reshape features to (C, 4, H, W) and apply conv3d blocks (Conv3d→BatchNorm3d→LeakyReLU) multiple times.\n\n## Model Comparison\n\nTo verify the fusion module's effectiveness, I compared prediction results.\n\nThe top row shows images with unmasked signals and GT masks. The middle and bottom rows show images with masked regions (simulating image artifacts) and predictions from the whole model and series model, respectively.\n\nWhile the whole model fails to predict the masked regions, the series model shares information across leads, enabling it to reasonably predict peak phase and amplitude.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F6102861%2Ffc67adf0cc00216fb93436f91d234fd6%2Fcombined_result.png?generation=1769278842918294&alt=media\" alt=\"masked_prediction\" width=\"800\" height=\"540\">\n\n# Training Strategy\n\n## Parameters\n- Input image: Stage 1 rectified image\n  - whole model: input shape = (3, 1280, 5600)\n  - series model: input shape = (4, 3, 480, 5600)\n- Loss: BCEWithLogitsLoss(pos_weight=20)\n- Epoch: 50\n- Batch size: 4\n- Optimizer: AdamW\n  - Learning rate: 5e-4 ~ 1e-3\n  - Weight decay: 0.01\n- Scheduler: CosineAnnealingLR\n\n**Tip: Gradient checkpointing is effective for reducing activation memory consumption with high-resolution images**\n\n## Augmentation\n```python\nimage_only_aug = A.Compose([\n    A.RandomBrightnessContrast(brightness_limit=(-0.1,0.1), contrast_limit=(-0.1, 0.1), p=0.2),\n    A.RandomShadow(p=0.2),\n    A.GaussianBlur(p=0.2),\n    A.CoarseDropout(num_holes_range=(1,8), hole_height_range=(0.01, 0.1), hole_width_range=(0.01, 0.05), fill=0, p=0.1),\n    A.ToGray(p=0.25),\n])\n```\n```python\nimage_and_mask_aug = A.HorizontalFlip(p=0.5)\n```\n\n# Post-Processing\n\n## Mask Post-Processing\nI compute a weighted average for each column of the mask using prediction results to obtain sub-pixel level predictions.\n\n## Resampling Method\nI tested:\n- scipy.signal.resample\n- scipy.signal.resample_poly(padtype='line')\n- torch.nn.functional.interpolate(mode=\"linear\")\n\nscipy.signal.resample gave the best results.\n\n## Pixel to Series\nCode for the complete mask→time-series conversion pipeline, including mask post-processing and resampling:\n\n```python\ndef pixel_to_series(pixel, length):\n    _, H, W = pixel.shape\n    eps=1e-8\n    y_idx = np.arange(H, dtype=np.float32)[:, None]\n\n    series = []\n    for j in [0, 1, 2, 3]:\n        p = pixel[j]\n        denom = p.sum(axis=0)\n        y_exp = (p * y_idx).sum(axis=0) / (denom + eps)\n        series.append(y_exp)\n    series = np.stack(series).astype(np.float32)\n\n    if length!=W:\n        resampled_series = []\n        for s in series:\n            rs = signal.resample(s, length).astype(np.float32)\n            resampled_series.append(rs)\n        series = np.stack(resampled_series)\n    return series\n```\n\n# Inference\n\n## TTA\n- Horizontal flip\n\n## Results\n\n### Submission A (6 models)\n**Public LB:** 23.37, **Private LB:** 23.27\n\n### Submission B (11 models)\n**Public LB:** 23.34, **Private LB:** 23.23\n\n| Encoder                       | Model (fusion module)    | Synthetic Data | Public LB | Private LB | Submission A | Submission B |\n|-------------------------------|------------------------|----------------|-----------|------------|-------|-------|\n| EfficientNet B7               | whole                  |                | 22.93     | 22.65      | ☑️    | ☑️    |\n| EfficientNetV2 L             | whole                  | ✅             | 22.60     | 22.55      | ☑️    | ☑️    |\n| EfficientNetV2 M             | whole                  |                | 22.52     | 22.38      |       | ☑️    |\n| EfficientNetV2 M             | whole                  | ✅             | 22.58     | 22.39      |       | ☑️    |\n| EfficientNet B6               | series (shared conv2d) | ✅             | 23.10     | 22.92      | ☑️    | ☑️    |\n| EfficientNet B6               | series (shared conv2d) |                | 23.00     | 22.81      | ☑️    | ☑️    |\n| EfficientNet B4               | series (shared conv2d) | ✅             | 22.93     | 22.73      |       | ☑️    |\n| EfficientNetV2 L             | series (conv3d)        |                | 22.92     | 22.77      | ☑️    | ☑️    |\n| EfficientNetV2 L             | series (conv2d)        |                | 22.85     | 22.70      | ☑️    | ☑️    |\n| EfficientNetV2 M             | series (conv3d)        | ✅             | 22.73     | 22.49      |       | ☑️    |\n| EfficientNetV2 M             | series (conv2d)        | ✅             | 22.76     | 22.55      |       | ☑️    |\n\n# References\n\n- https://www.kaggle.com/code/hengck23/demo-submission\n- Shivashankara KK, Deepanshi, Shervedani AM, Reyna MA, Clifford GD, Sameni R. ECG-Image-Kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiological Measurement 2024; 45:055019. DOI: 10.1088/1361-6579/ad4954\n- Reyna MA, Deepanshi, Weigle J, Koscova Z, Campbell K, Shivashankara KK, Saghafi S, Nikookar S, Motie-Shirazi M, Kiarashi Y, Seyedi S, Hassannia M, Bjørnstad AM, Stenhede E, Ranjbar A, Clifford GD, and Sameni R. ECG-Image-Database: A dataset of ECG images with real-world imaging and scanning artifacts; a foundation for computerized ECG image digitization and analysis, 2024. DOI: 10.48550/arXiv.2409.16612.\n- Wagner, P., Strodthoff, N., Bousseljot, R., Samek, W., & Schaeffter, T. (2022). PTB-XL, a large publicly available electrocardiography dataset (version 1.0.3). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/kfzx-aw45\n- Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F.I., Samek, W., Schaeffter, T. (2020), PTB-XL: A Large Publicly Available ECG Dataset. Scientific Data. https://doi.org/10.1038/s41597-020-0495-6\n- Goldberger, A., Amaral, L., Glass, L., Hausdorff, J., Ivanov, P. C., Mark, R., ... & Stanley, H. E. (2000). PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation [Online]. 101 (23), pp. e215–e220. RRID:SCR_007345.\n\n# Code Availability\n\n- Training code: https://github.com/someya-takashi/physionet-image-digitization/tree/main\n- Inference code: https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission",
    "3396393": "Congrats! Using original sampling data is smart, we noticed competition data is re-sampled and this will impact the result but do nothing 😅.",
    "3396579": "Thanks! And congrats on 1st place!\nYour team’s approach of predicting the signal directly is really insightful—I learned a lot from it! 😄",
    "3401307": "Finally published my training and inference code.\n\n- Training code: https://github.com/someya-takashi/physionet-image-digitization/tree/main\n- Inference code: https://www.kaggle.com/code/takashisomeya/physionet-2nd-place-submission\n\nHope it helps!"
  },
  "source": "meta"
}