{
  "id": 669548,
  "title": "7th place solution",
  "url": "/competitions/physionet-ecg-image-digitization/writeups/7th-place-solution",
  "author_name": "",
  "post_date": "2026-01-23T01:19:03.987Z",
  "votes": 48,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Thanks to the PhysioNet hosts and Kaggle for this fun competition. There were many stages to optimize, and we took the opportunity to learn as much as we could from all stages. Congrats to everyone who competed, Harshit and I are excited to read through all the solution write-ups!</p>\n<h2>TLDR</h2>\n<p>Our pipeline uses rotation, lead detection, lead segmentation, digitization, and out-of-distribution detection models. We modified the ECG-image-kit repository to create lead detection training data, relied on the competition data for digitization, and used out-of-distribution detection models to optimize our ensemble. For more details, keep reading!</p>\n<h2>Cross Validation</h2>\n<p>For validating experiments, we used a k-fold cross-validation scheme across all samples. We used lightweight models on all folds to detect edge cases; for more computationally expensive models, we only validated on 100 samples. We found that 100 samples were enough for a strong CV/LB correlation and increased the speed of experiments.</p>\n<h2>Data Generation</h2>\n<p>Next, due to a lack of lead annotations in the competition data, we modified ECG-image-kit to create data for the rotation, lead detection, and lead segmentation models. We updated the codebase to insert ECG plots into backgrounds, simulate shadows, and use more variable colors and textures in the ECG plots. All modifications we made to the codebase were done in an attempt to make the artifacts more realistic.</p>\n<p>We used the <a href=\"https://physionet.org/content/ptb-xl/1.0.3/\" target=\"_blank\">PTBXL dataset</a> for the raw ECG values, and the <a href=\"https://www.robots.ox.ac.uk/~vgg/data/dtd/\" target=\"_blank\">Describable Textures Dataset (DTD)</a> for backgrounds and shadows. Here are a few samples of what we generated with bounding boxes and segmentation labels overlaid.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fbeeed8c1741a9b4311dbd1c3b9a2eac3%2FIMG_4.jpg?generation=1769131261528664&amp;alt=media\" alt=\"IMG_0\"></p>\n<h2>Rotation</h2>\n<p>The first model in the pipeline was a simple classification model to predict when an image needed rotation. We used the B4 and B5 variants from the HGNet-V2 model family to predict 4 classes (0,90,180 or 270 degree rotation). We detected 69 images that required rotation in the training dataset, and applied this model first during inference.</p>\n<h2>Lead Detection / Segmentation</h2>\n<p>Next, we trained a set of hybrid lead detection/segmentation models. This was a great learning curve for me (Bartley), as I have always wanted detection models that come without a confusing license. To do this, we designed a model to predict objectiveness, class scores, and offsets. It was important to add sufficient capacity in the bounding box detection head (&gt;=128 channels) for the model to be able to learn the signal. We found that the ConvNeXt model family worked best, though the architecture supports any backbone from the timm library. </p>\n<p>We first ran the detection model to locate the AOI (area of interest). We then cropped the image and re-ran the model for a more precise result. We also added a minimum crop height of 16 pixels and a width of 64 pixels to account for small lead crops. This fixed all catastrophic detections we observed in the training set, and boosted LB by ~0.2-3. Here are some sample predictions on the training set.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F3e3547c29e7459fd238e239e2262d341%2FIMG_5.jpg?generation=1769130824182840&amp;alt=media\" alt=\"IMG_1\"></p>\n<p>We also added a segmentation branch to predict the pixels corresponding to the 13 different lead classes. We used this predicted segmentation as an input to our digitization models in the next stage.</p>\n<h2>Digitization (Bartley)</h2>\n<p>The first digitization model we used was a <code>maxxvitv2_nano_rw_256.sw_in1k</code> model with a 1D unet decoder. We pooled the encoder features before passing them into the decoder. The model input was a 5-channel image. Three RGB channels, one for the target lead probability, and one for the maximum probability of other leads. The 5-channel input helped generate more robust predictions on crops with overlapping leads.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fd1931b0d7ab3ca66ddd2668d8b13393b%2FIMG_6.jpg?generation=1769130842998048&amp;alt=media\" alt=\"IMG_2\"></p>\n<p>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.</p>\n<p>We used a couple of variations of SNRloss. During the low-resolution stage, we used a variation of SNRloss that forces the model to learn the optimal vertical shift. In the later stages (once vertical shift was learned), we used a variation that used the optimal vertical shift to more closely align with the competition metric. The latter approach was able to achieve higher final scores.</p>\n<p>For the lead II full crops, we used a VIT model with a linear head to go from patch embeddings to pixel-level predictions. This architecture was identical to Harshit's, and we will go into more details in the next section.</p>\n<h2>Digitization (Harshit)</h2>\n<p>For our next set of digitization models, we rely on the same preprocessing steps from Bartley's pipeline. All our models here used a <code>vit_small_patch16_dinov3.lvd1689m</code> encoder with slight differences in the training setup. The architecture was heavily inspired by Harshit’s 1st Place Solution in the Yale Competition <a href=\"https://www.kaggle.com/competitions/waveform-inversion/writeups/harshit-sheoran-1st-place-solution\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F7d1d8f99365f40f5662bb4a367daa825%2FIMG_7.png?generation=1769130914922529&amp;alt=media\" alt=\"IMG_4\"></p>\n<p>We developed three variations of this model to improve diversity. While the architecture and training method remained consistent, we varied the input data sources (crops) and loss functions. We used random padding augmentation during training on crops derived from the sources below.</p>\n<table>\n<thead>\n<tr>\n<th>Model Version</th>\n<th>Input Data Sources (Crops)</th>\n<th>Loss Function</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>V27</strong></td>\n<td><code>train_crops5</code>, <code>train_bartley_crops3</code></td>\n<td>MAE</td>\n</tr>\n<tr>\n<td><strong>V27-SNR</strong></td>\n<td><code>train_crops5</code>, <code>train_bartley_crops3</code></td>\n<td>SNR</td>\n</tr>\n<tr>\n<td><strong>V6</strong></td>\n<td><code>train_gen_crops1</code>, <code>train_crops4</code>, <code>train_bartley_crops2</code>, <code>train_bartley_crops4</code></td>\n<td>MAE</td>\n</tr>\n</tbody>\n</table>\n<p>In addition, we employed a multi-stage training approach with progressive image size scaling to ensure stable convergence. Each stage consisted of <strong>20 epochs</strong>, scaling up the resolution as follows:</p>\n<pre><code>`224x896` → `224x1782` → `224x2688` → `224x3584` → `336x3584`\n</code></pre>\n<p>We intentionally kept augmentations minimal, only using horizontal flips during training. Despite the light augmentation pipeline, the models converged effectively. To test the robustness of this pipeline, we took pictures of ECG plots on different monitors to simulate a distribution shift. Scores were consistent with those on the competition set indicating that the models were robust.</p>\n<h2>Ensemble</h2>\n<p>We expected a large boost when combining our approaches as we used a diverse set of architectures, training pipelines, and loss functions. Harshit’s models excelled on in-distribution samples and when there was a significant drift in the signal. Bartley’s models excelled on out-of-distribution samples and on crops with overlapping leads. A simple mean ensemble of our pipeline scored <strong>22.54/22.10</strong> on the Public/Private LB.</p>\n<h3>Out-of-Distribution (OOD) Detection</h3>\n<p>To further improve ensemble performance, we implemented an out-of-distribution (OOD) detection model. Since Harshit’s models were highly specialized for in-distribution samples, we wanted to detect when to mask his predictions during inference.</p>\n<p>To do this, we trained a feature extractor using <code>tf_efficientnetv2_s</code> and ArcFace Loss. During inference, we calculated the embedding of each test image and compared it to the average embedding of each image type in the training set. If the Cosine Similarity between the test image and any of the average embeddings was &lt;0.5, we masked Harshit's predictions. This strategy boosted our score further to <strong>22.80/22.48</strong>.</p>\n<p>All the models we trained in this competition (Harshit and Bartley) used the Muon optimizer from timm. We found that this significantly outperformed all others and thought it was worth a mention.</p>\n<h2>Final Note</h2>\n<p>Last thing, a quick shout-out to <a href=\"https://www.kaggle.com/TheoViel\" target=\"_blank\">@TheoViel</a>. I recently modified my training pipeline to follow a similar structure to his <a href=\"https://github.com/TheoViel/kaggle_rsna_abdominal_trauma\" target=\"_blank\">RSNA 2023 Solution</a>. It’s an excellent repository that I would recommend checking out.</p>\n<p>Thanks for reading, and as always, Happy Kaggling!</p>\n<p>Code: <a href=\"https://github.com/brendanartley/PhysioNet-Competition\" target=\"_blank\">here</a></p>\n<p>Datasets: <a href=\"https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission---other-data\" target=\"_blank\">here</a></p>\n<p>Inference: <a href=\"https://www.kaggle.com/code/harshitsheoran/physionet-infer-v-final\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "3395437",
      "postDate": "01/23/2026 01:18:34",
      "content": "<p>Thanks to the PhysioNet hosts and Kaggle for this fun competition. There were many stages to optimize, and we took the opportunity to learn as much as we could from all stages. Congrats to everyone who competed, Harshit and I are excited to read through all the solution write-ups!</p>\n<h2>TLDR</h2>\n<p>Our pipeline uses rotation, lead detection, lead segmentation, digitization, and out-of-distribution detection models. We modified the ECG-image-kit repository to create lead detection training data, relied on the competition data for digitization, and used out-of-distribution detection models to optimize our ensemble. For more details, keep reading!</p>\n<h2>Cross Validation</h2>\n<p>For validating experiments, we used a k-fold cross-validation scheme across all samples. We used lightweight models on all folds to detect edge cases; for more computationally expensive models, we only validated on 100 samples. We found that 100 samples were enough for a strong CV/LB correlation and increased the speed of experiments.</p>\n<h2>Data Generation</h2>\n<p>Next, due to a lack of lead annotations in the competition data, we modified ECG-image-kit to create data for the rotation, lead detection, and lead segmentation models. We updated the codebase to insert ECG plots into backgrounds, simulate shadows, and use more variable colors and textures in the ECG plots. All modifications we made to the codebase were done in an attempt to make the artifacts more realistic.</p>\n<p>We used the <a href=\"https://physionet.org/content/ptb-xl/1.0.3/\" target=\"_blank\">PTBXL dataset</a> for the raw ECG values, and the <a href=\"https://www.robots.ox.ac.uk/~vgg/data/dtd/\" target=\"_blank\">Describable Textures Dataset (DTD)</a> for backgrounds and shadows. Here are a few samples of what we generated with bounding boxes and segmentation labels overlaid.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fbeeed8c1741a9b4311dbd1c3b9a2eac3%2FIMG_4.jpg?generation=1769131261528664&amp;alt=media\" alt=\"IMG_0\"></p>\n<h2>Rotation</h2>\n<p>The first model in the pipeline was a simple classification model to predict when an image needed rotation. We used the B4 and B5 variants from the HGNet-V2 model family to predict 4 classes (0,90,180 or 270 degree rotation). We detected 69 images that required rotation in the training dataset, and applied this model first during inference.</p>\n<h2>Lead Detection / Segmentation</h2>\n<p>Next, we trained a set of hybrid lead detection/segmentation models. This was a great learning curve for me (Bartley), as I have always wanted detection models that come without a confusing license. To do this, we designed a model to predict objectiveness, class scores, and offsets. It was important to add sufficient capacity in the bounding box detection head (&gt;=128 channels) for the model to be able to learn the signal. We found that the ConvNeXt model family worked best, though the architecture supports any backbone from the timm library. </p>\n<p>We first ran the detection model to locate the AOI (area of interest). We then cropped the image and re-ran the model for a more precise result. We also added a minimum crop height of 16 pixels and a width of 64 pixels to account for small lead crops. This fixed all catastrophic detections we observed in the training set, and boosted LB by ~0.2-3. Here are some sample predictions on the training set.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F3e3547c29e7459fd238e239e2262d341%2FIMG_5.jpg?generation=1769130824182840&amp;alt=media\" alt=\"IMG_1\"></p>\n<p>We also added a segmentation branch to predict the pixels corresponding to the 13 different lead classes. We used this predicted segmentation as an input to our digitization models in the next stage.</p>\n<h2>Digitization (Bartley)</h2>\n<p>The first digitization model we used was a <code>maxxvitv2_nano_rw_256.sw_in1k</code> model with a 1D unet decoder. We pooled the encoder features before passing them into the decoder. The model input was a 5-channel image. Three RGB channels, one for the target lead probability, and one for the maximum probability of other leads. The 5-channel input helped generate more robust predictions on crops with overlapping leads.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fd1931b0d7ab3ca66ddd2668d8b13393b%2FIMG_6.jpg?generation=1769130842998048&amp;alt=media\" alt=\"IMG_2\"></p>\n<p>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.</p>\n<p>We used a couple of variations of SNRloss. During the low-resolution stage, we used a variation of SNRloss that forces the model to learn the optimal vertical shift. In the later stages (once vertical shift was learned), we used a variation that used the optimal vertical shift to more closely align with the competition metric. The latter approach was able to achieve higher final scores.</p>\n<p>For the lead II full crops, we used a VIT model with a linear head to go from patch embeddings to pixel-level predictions. This architecture was identical to Harshit's, and we will go into more details in the next section.</p>\n<h2>Digitization (Harshit)</h2>\n<p>For our next set of digitization models, we rely on the same preprocessing steps from Bartley's pipeline. All our models here used a <code>vit_small_patch16_dinov3.lvd1689m</code> encoder with slight differences in the training setup. The architecture was heavily inspired by Harshit’s 1st Place Solution in the Yale Competition <a href=\"https://www.kaggle.com/competitions/waveform-inversion/writeups/harshit-sheoran-1st-place-solution\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F7d1d8f99365f40f5662bb4a367daa825%2FIMG_7.png?generation=1769130914922529&amp;alt=media\" alt=\"IMG_4\"></p>\n<p>We developed three variations of this model to improve diversity. While the architecture and training method remained consistent, we varied the input data sources (crops) and loss functions. We used random padding augmentation during training on crops derived from the sources below.</p>\n<table>\n<thead>\n<tr>\n<th>Model Version</th>\n<th>Input Data Sources (Crops)</th>\n<th>Loss Function</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>V27</strong></td>\n<td><code>train_crops5</code>, <code>train_bartley_crops3</code></td>\n<td>MAE</td>\n</tr>\n<tr>\n<td><strong>V27-SNR</strong></td>\n<td><code>train_crops5</code>, <code>train_bartley_crops3</code></td>\n<td>SNR</td>\n</tr>\n<tr>\n<td><strong>V6</strong></td>\n<td><code>train_gen_crops1</code>, <code>train_crops4</code>, <code>train_bartley_crops2</code>, <code>train_bartley_crops4</code></td>\n<td>MAE</td>\n</tr>\n</tbody>\n</table>\n<p>In addition, we employed a multi-stage training approach with progressive image size scaling to ensure stable convergence. Each stage consisted of <strong>20 epochs</strong>, scaling up the resolution as follows:</p>\n<pre><code>`224x896` → `224x1782` → `224x2688` → `224x3584` → `336x3584`\n</code></pre>\n<p>We intentionally kept augmentations minimal, only using horizontal flips during training. Despite the light augmentation pipeline, the models converged effectively. To test the robustness of this pipeline, we took pictures of ECG plots on different monitors to simulate a distribution shift. Scores were consistent with those on the competition set indicating that the models were robust.</p>\n<h2>Ensemble</h2>\n<p>We expected a large boost when combining our approaches as we used a diverse set of architectures, training pipelines, and loss functions. Harshit’s models excelled on in-distribution samples and when there was a significant drift in the signal. Bartley’s models excelled on out-of-distribution samples and on crops with overlapping leads. A simple mean ensemble of our pipeline scored <strong>22.54/22.10</strong> on the Public/Private LB.</p>\n<h3>Out-of-Distribution (OOD) Detection</h3>\n<p>To further improve ensemble performance, we implemented an out-of-distribution (OOD) detection model. Since Harshit’s models were highly specialized for in-distribution samples, we wanted to detect when to mask his predictions during inference.</p>\n<p>To do this, we trained a feature extractor using <code>tf_efficientnetv2_s</code> and ArcFace Loss. During inference, we calculated the embedding of each test image and compared it to the average embedding of each image type in the training set. If the Cosine Similarity between the test image and any of the average embeddings was &lt;0.5, we masked Harshit's predictions. This strategy boosted our score further to <strong>22.80/22.48</strong>.</p>\n<p>All the models we trained in this competition (Harshit and Bartley) used the Muon optimizer from timm. We found that this significantly outperformed all others and thought it was worth a mention.</p>\n<h2>Final Note</h2>\n<p>Last thing, a quick shout-out to <a href=\"https://www.kaggle.com/TheoViel\" target=\"_blank\">@TheoViel</a>. I recently modified my training pipeline to follow a similar structure to his <a href=\"https://github.com/TheoViel/kaggle_rsna_abdominal_trauma\" target=\"_blank\">RSNA 2023 Solution</a>. It’s an excellent repository that I would recommend checking out.</p>\n<p>Thanks for reading, and as always, Happy Kaggling!</p>\n<p>Code: <a href=\"https://github.com/brendanartley/PhysioNet-Competition\" target=\"_blank\">here</a></p>\n<p>Datasets: <a href=\"https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission---other-data\" target=\"_blank\">here</a></p>\n<p>Inference: <a href=\"https://www.kaggle.com/code/harshitsheoran/physionet-infer-v-final\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Thanks to the PhysioNet hosts and Kaggle for this fun competition. There were many stages to optimize, and we took the opportunity to learn as much as we could from all stages. Congrats to everyone who competed, Harshit and I are excited to read through all the solution write-ups!\n\n## TLDR\n\nOur pipeline uses rotation, lead detection, lead segmentation, digitization, and out-of-distribution detection models. We modified the ECG-image-kit repository to create lead detection training data, relied on the competition data for digitization, and used out-of-distribution detection models to optimize our ensemble. For more details, keep reading!\n\n## Cross Validation\n\nFor validating experiments, we used a k-fold cross-validation scheme across all samples. We used lightweight models on all folds to detect edge cases; for more computationally expensive models, we only validated on 100 samples. We found that 100 samples were enough for a strong CV/LB correlation and increased the speed of experiments.\n\n## Data Generation\n\nNext, due to a lack of lead annotations in the competition data, we modified ECG-image-kit to create data for the rotation, lead detection, and lead segmentation models. We updated the codebase to insert ECG plots into backgrounds, simulate shadows, and use more variable colors and textures in the ECG plots. All modifications we made to the codebase were done in an attempt to make the artifacts more realistic.\n\nWe used the [PTBXL dataset](https://physionet.org/content/ptb-xl/1.0.3/) for the raw ECG values, and the [Describable Textures Dataset (DTD)](https://www.robots.ox.ac.uk/~vgg/data/dtd/) for backgrounds and shadows. Here are a few samples of what we generated with bounding boxes and segmentation labels overlaid.\n\n![IMG_0](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fbeeed8c1741a9b4311dbd1c3b9a2eac3%2FIMG_4.jpg?generation=1769131261528664&alt=media)\n\n## Rotation\n\nThe first model in the pipeline was a simple classification model to predict when an image needed rotation. We used the B4 and B5 variants from the HGNet-V2 model family to predict 4 classes (0,90,180 or 270 degree rotation). We detected 69 images that required rotation in the training dataset, and applied this model first during inference.\n\n## Lead Detection / Segmentation\n\nNext, we trained a set of hybrid lead detection/segmentation models. This was a great learning curve for me (Bartley), as I have always wanted detection models that come without a confusing license. To do this, we designed a model to predict objectiveness, class scores, and offsets. It was important to add sufficient capacity in the bounding box detection head (>=128 channels) for the model to be able to learn the signal. We found that the ConvNeXt model family worked best, though the architecture supports any backbone from the timm library. \n\nWe first ran the detection model to locate the AOI (area of interest). We then cropped the image and re-ran the model for a more precise result. We also added a minimum crop height of 16 pixels and a width of 64 pixels to account for small lead crops. This fixed all catastrophic detections we observed in the training set, and boosted LB by ~0.2-3. Here are some sample predictions on the training set.\n\n![IMG_1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F3e3547c29e7459fd238e239e2262d341%2FIMG_5.jpg?generation=1769130824182840&alt=media)\n\nWe also added a segmentation branch to predict the pixels corresponding to the 13 different lead classes. We used this predicted segmentation as an input to our digitization models in the next stage.\n\n## Digitization (Bartley)\n\nThe first digitization model we used was a `maxxvitv2_nano_rw_256.sw_in1k` model with a 1D unet decoder. We pooled the encoder features before passing them into the decoder. The model input was a 5-channel image. Three RGB channels, one for the target lead probability, and one for the maximum probability of other leads. The 5-channel input helped generate more robust predictions on crops with overlapping leads.\n\n![IMG_2](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fd1931b0d7ab3ca66ddd2668d8b13393b%2FIMG_6.jpg?generation=1769130842998048&alt=media)\n\nDuring training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.\n\nWe used a couple of variations of SNRloss. During the low-resolution stage, we used a variation of SNRloss that forces the model to learn the optimal vertical shift. In the later stages (once vertical shift was learned), we used a variation that used the optimal vertical shift to more closely align with the competition metric. The latter approach was able to achieve higher final scores.\n\nFor the lead II full crops, we used a VIT model with a linear head to go from patch embeddings to pixel-level predictions. This architecture was identical to Harshit's, and we will go into more details in the next section.\n\n## Digitization (Harshit)\n\nFor our next set of digitization models, we rely on the same preprocessing steps from Bartley's pipeline. All our models here used a `vit_small_patch16_dinov3.lvd1689m` encoder with slight differences in the training setup. The architecture was heavily inspired by Harshit’s 1st Place Solution in the Yale Competition [here](https://www.kaggle.com/competitions/waveform-inversion/writeups/harshit-sheoran-1st-place-solution).\n\n![IMG_4](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F7d1d8f99365f40f5662bb4a367daa825%2FIMG_7.png?generation=1769130914922529&alt=media)\n\nWe developed three variations of this model to improve diversity. While the architecture and training method remained consistent, we varied the input data sources (crops) and loss functions. We used random padding augmentation during training on crops derived from the sources below.\n\n| Model Version | Input Data Sources (Crops) | Loss Function |\n| :--- | :--- | :--- |\n| **V27** | `train_crops5`, `train_bartley_crops3` | MAE |\n| **V27-SNR** | `train_crops5`, `train_bartley_crops3` | SNR |\n| **V6** | `train_gen_crops1`, `train_crops4`, `train_bartley_crops2`, `train_bartley_crops4` | MAE |\n\nIn addition, we employed a multi-stage training approach with progressive image size scaling to ensure stable convergence. Each stage consisted of **20 epochs**, scaling up the resolution as follows:\n\n    `224x896` → `224x1782` → `224x2688` → `224x3584` → `336x3584`\n\nWe intentionally kept augmentations minimal, only using horizontal flips during training. Despite the light augmentation pipeline, the models converged effectively. To test the robustness of this pipeline, we took pictures of ECG plots on different monitors to simulate a distribution shift. Scores were consistent with those on the competition set indicating that the models were robust.\n\n## Ensemble\n\nWe expected a large boost when combining our approaches as we used a diverse set of architectures, training pipelines, and loss functions. Harshit’s models excelled on in-distribution samples and when there was a significant drift in the signal. Bartley’s models excelled on out-of-distribution samples and on crops with overlapping leads. A simple mean ensemble of our pipeline scored **22.54/22.10** on the Public/Private LB.\n\n### Out-of-Distribution (OOD) Detection\n\nTo further improve ensemble performance, we implemented an out-of-distribution (OOD) detection model. Since Harshit’s models were highly specialized for in-distribution samples, we wanted to detect when to mask his predictions during inference.\n\nTo do this, we trained a feature extractor using `tf_efficientnetv2_s` and ArcFace Loss. During inference, we calculated the embedding of each test image and compared it to the average embedding of each image type in the training set. If the Cosine Similarity between the test image and any of the average embeddings was <0.5, we masked Harshit's predictions. This strategy boosted our score further to **22.80/22.48**.\n\nAll the models we trained in this competition (Harshit and Bartley) used the Muon optimizer from timm. We found that this significantly outperformed all others and thought it was worth a mention.\n\n## Final Note\n\nLast thing, a quick shout-out to @TheoViel. I recently modified my training pipeline to follow a similar structure to his [RSNA 2023 Solution](https://github.com/TheoViel/kaggle_rsna_abdominal_trauma). It’s an excellent repository that I would recommend checking out.\n\nThanks for reading, and as always, Happy Kaggling!\n\nCode: [here](https://github.com/brendanartley/PhysioNet-Competition)\n\nDatasets: [here](https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission), [here](https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission---other-data)\n\nInference: [here](https://www.kaggle.com/code/harshitsheoran/physionet-infer-v-final)",
      "votes": null
    },
    {
      "id": "3395450",
      "postDate": "01/23/2026 02:20:09",
      "content": "<p>Thanks for sharing.<br>\nI really liked the step-wise task decomposition with minimal interference between stages.<br>\nIt made me realize that properly separating tasks allows us to fully optimize each stage — something I struggled with in my own approach.</p>",
      "rawMarkdown": "Thanks for sharing.<br>\nI really liked the step-wise task decomposition with minimal interference between stages.<br>\nIt made me realize that properly separating tasks allows us to fully optimize each stage — something I struggled with in my own approach.",
      "votes": null
    },
    {
      "id": "3395512",
      "postDate": "01/23/2026 05:20:51",
      "content": "<p>unique and exceptional. While most of the available solutions are somehow based on hengc23's great work. This one goes a different length!</p>",
      "rawMarkdown": "unique and exceptional. While most of the available solutions are somehow based on hengc23's great work. This one goes a different length!",
      "votes": null
    },
    {
      "id": "3395520",
      "postDate": "01/23/2026 05:45:59",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/sanjidh090\" target=\"_blank\">@sanjidh090</a>!</p>",
      "rawMarkdown": "Thanks @sanjidh090!",
      "votes": null
    },
    {
      "id": "3395522",
      "postDate": "01/23/2026 05:54:22",
      "content": "<p>Agreed <a href=\"https://www.kaggle.com/sasaleaf\" target=\"_blank\">@sasaleaf</a>, we learnt a lot by implementing each stage separately. Thanks for your comment :)</p>",
      "rawMarkdown": "Agreed @sasaleaf, we learnt a lot by implementing each stage separately. Thanks for your comment :)",
      "votes": null
    },
    {
      "id": "3396526",
      "postDate": "01/25/2026 09:02:16",
      "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Congratulations. I impressed Bartley's customized YOLO-like model implemented from scratch. Also Harshit's ViT model which predicts 1d output directly from 2d input.</p>\n<blockquote>\n  <p>We intentionally kept augmentations minimal, only using horizontal flips during training.</p>\n</blockquote>\n<p>For training digitization model, I wonder strong augmentations (like geometric transforms) can boost 2nd-stage model. Why did you only apply flip augmentations?</p>",
      "rawMarkdown": "harshitsheoran @brendanartley Congratulations. I impressed Bartley's customized YOLO-like model implemented from scratch. Also Harshit's ViT model which predicts 1d output directly from 2d input.\n\n> We intentionally kept augmentations minimal, only using horizontal flips during training.\n\nFor training digitization model, I wonder strong augmentations (like geometric transforms) can boost 2nd-stage model. Why did you only apply flip augmentations?",
      "votes": null
    },
    {
      "id": "3396629",
      "postDate": "01/25/2026 13:27:30",
      "content": "<p>Hi! <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>,</p>\n<p>Happy to know that you liked our solution,</p>\n<p>My teammate <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> had a great amount of augmentations in his pipeline, </p>\n<blockquote>\n  <p>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.</p>\n</blockquote>\n<p>In my pipeline I was getting great results with or without them,  and training without was certainly faster and easier, and it provides more diversity when ensembling so we continued with both pipelines.</p>",
      "rawMarkdown": "Hi! @tatamikenn,\n\nHappy to know that you liked our solution,\n\nMy teammate @brendanartley had a great amount of augmentations in his pipeline, \n\n>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.\n\nIn my pipeline I was getting great results with or without them,  and training without was certainly faster and easier, and it provides more diversity when ensembling so we continued with both pipelines.",
      "votes": null
    },
    {
      "id": "3494057",
      "postDate": "07/09/2026 04:12:59",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> thx for this amazing solution writeup.\nBTW, how did you handle perspective/skew distortion from angled camera photos? </p>\n<p>I noticed there's no explicit dewarping step (grid-corner → homography rectification) like some other top solutions use — rotate only does discrete 0/90/180/270, and --augment only applies affine (rotation/crop/noise), no true perspective transform. </p>\n<p>Was this intentional (test photos weren't skewed much / small-crop processing absorbs it), or just out of scope? Curious how robust det2.py is on strongly-angled real photos.</p>",
      "rawMarkdown": "brendanartley @harshitsheoran thx for this amazing solution writeup.\nBTW, how did you handle perspective/skew distortion from angled camera photos? \n\nI noticed there's no explicit dewarping step (grid-corner → homography rectification) like some other top solutions use — rotate only does discrete 0/90/180/270, and --augment only applies affine (rotation/crop/noise), no true perspective transform. \n\nWas this intentional (test photos weren't skewed much / small-crop processing absorbs it), or just out of scope? Curious how robust det2.py is on strongly-angled real photos.",
      "votes": null
    },
    {
      "id": "3494058",
      "postDate": "07/09/2026 04:14:31",
      "content": "<blockquote>\n  <p>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks).</p>\n</blockquote>\n<p>Is this the one that actually made it possible to handle distortion/skewed/warped ecg images?</p>",
      "rawMarkdown": ">During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks).\n\nIs this the one that actually made it possible to handle distortion/skewed/warped ecg images?",
      "votes": null
    },
    {
      "id": "3494094",
      "postDate": "07/09/2026 06:28:53",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> \nAlso, what are your thoughts on extending layout domain from '3x4 + 1R only' to adding '6x2', '12x1' layouts etc.\nCause even tho i finetuned this model on '6x2' layout images(about 2000 ecg-image-kit generated), lead detection is not that promising. \nAny suggestion?\nThx again.</p>",
      "rawMarkdown": "brendanartley @harshitsheoran \nAlso, what are your thoughts on extending layout domain from '3x4 + 1R only' to adding '6x2', '12x1' layouts etc.\nCause even tho i finetuned this model on '6x2' layout images(about 2000 ecg-image-kit generated), lead detection is not that promising. \nAny suggestion?\nThx again.",
      "votes": null
    },
    {
      "id": "3494309",
      "postDate": "07/09/2026 16:28:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a>, thanks for the comments. The lead detection model was trained only on  3x4 + 1R layouts, so I would recommend fine-tuning the model on new lead configurations for it to work on new configurations.</p>\n<p>In regards to your comment about augmentations, we intentionally used the heavy augmentations to simulate challenging distorted/skewed/warped images. Although it may not been as strong as other methods, we designed our digitization pipeline with generalization as the main goal! Hope this helps.</p>",
      "rawMarkdown": "Hi @beckpro, thanks for the comments. The lead detection model was trained only on ~~12x1~~ 3x4 + 1R layouts, so I would recommend fine-tuning the model on new lead configurations for it to work on new configurations.\n\nIn regards to your comment about augmentations, we intentionally used the heavy augmentations to simulate challenging distorted/skewed/warped images. Although it may not been as strong as other methods, we designed our digitization pipeline with generalization as the main goal! Hope this helps.",
      "votes": null
    },
    {
      "id": "3494421",
      "postDate": "07/09/2026 23:26:09",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> U meant 'was trained only on 3x4 + 1R' ? Cause those are the competition layouts. Thx anyway</p>",
      "rawMarkdown": "brendanartley U meant 'was trained only on 3x4 + 1R' ? Cause those are the competition layouts. Thx anyway",
      "votes": null
    },
    {
      "id": "3494426",
      "postDate": "07/09/2026 23:35:09",
      "content": "<p>Oops, yes thats correct. Edited above!</p>",
      "rawMarkdown": "Oops, yes thats correct. Edited above!",
      "votes": null
    },
    {
      "id": "3494431",
      "postDate": "07/09/2026 23:46:52",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> oh sorry to bother u again, but\nis the code license free? or like solution itself has license etc. \ni'm trying to mix up a lot of solutions including ur lead detection modules and others and make my own product.</p>",
      "rawMarkdown": "brendanartley oh sorry to bother u again, but\nis the code license free? or like solution itself has license etc. \ni'm trying to mix up a lot of solutions including ur lead detection modules and others and make my own product.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3395450,
      "author_name": "sasaleaf",
      "author_url": "",
      "post_date": "01/23/2026 02:20:09",
      "content": "<p>Thanks for sharing.<br>\nI really liked the step-wise task decomposition with minimal interference between stages.<br>\nIt made me realize that properly separating tasks allows us to fully optimize each stage — something I struggled with in my own approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3395522,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "01/23/2026 05:54:22",
          "content": "<p>Agreed <a href=\"https://www.kaggle.com/sasaleaf\" target=\"_blank\">@sasaleaf</a>, we learnt a lot by implementing each stage separately. Thanks for your comment :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3395512,
      "author_name": "sanjidh090",
      "author_url": "",
      "post_date": "01/23/2026 05:20:51",
      "content": "<p>unique and exceptional. While most of the available solutions are somehow based on hengc23's great work. This one goes a different length!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3395520,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "01/23/2026 05:45:59",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sanjidh090\" target=\"_blank\">@sanjidh090</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3396526,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/25/2026 09:02:16",
      "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Congratulations. I impressed Bartley's customized YOLO-like model implemented from scratch. Also Harshit's ViT model which predicts 1d output directly from 2d input.</p>\n<blockquote>\n  <p>We intentionally kept augmentations minimal, only using horizontal flips during training.</p>\n</blockquote>\n<p>For training digitization model, I wonder strong augmentations (like geometric transforms) can boost 2nd-stage model. Why did you only apply flip augmentations?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3396629,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "01/25/2026 13:27:30",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>,</p>\n<p>Happy to know that you liked our solution,</p>\n<p>My teammate <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> had a great amount of augmentations in his pipeline, </p>\n<blockquote>\n  <p>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.</p>\n</blockquote>\n<p>In my pipeline I was getting great results with or without them,  and training without was certainly faster and easier, and it provides more diversity when ensembling so we continued with both pipelines.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3494057,
      "author_name": "beckpro",
      "author_url": "",
      "post_date": "07/09/2026 04:12:59",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> thx for this amazing solution writeup.\nBTW, how did you handle perspective/skew distortion from angled camera photos? </p>\n<p>I noticed there's no explicit dewarping step (grid-corner → homography rectification) like some other top solutions use — rotate only does discrete 0/90/180/270, and --augment only applies affine (rotation/crop/noise), no true perspective transform. </p>\n<p>Was this intentional (test photos weren't skewed much / small-crop processing absorbs it), or just out of scope? Curious how robust det2.py is on strongly-angled real photos.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3494058,
          "author_name": "beckpro",
          "author_url": "",
          "post_date": "07/09/2026 04:14:31",
          "content": "<blockquote>\n  <p>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks).</p>\n</blockquote>\n<p>Is this the one that actually made it possible to handle distortion/skewed/warped ecg images?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3494094,
      "author_name": "beckpro",
      "author_url": "",
      "post_date": "07/09/2026 06:28:53",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> \nAlso, what are your thoughts on extending layout domain from '3x4 + 1R only' to adding '6x2', '12x1' layouts etc.\nCause even tho i finetuned this model on '6x2' layout images(about 2000 ecg-image-kit generated), lead detection is not that promising. \nAny suggestion?\nThx again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3494309,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "07/09/2026 16:28:37",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a>, thanks for the comments. The lead detection model was trained only on  3x4 + 1R layouts, so I would recommend fine-tuning the model on new lead configurations for it to work on new configurations.</p>\n<p>In regards to your comment about augmentations, we intentionally used the heavy augmentations to simulate challenging distorted/skewed/warped images. Although it may not been as strong as other methods, we designed our digitization pipeline with generalization as the main goal! Hope this helps.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3494421,
              "author_name": "beckpro",
              "author_url": "",
              "post_date": "07/09/2026 23:26:09",
              "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> U meant 'was trained only on 3x4 + 1R' ? Cause those are the competition layouts. Thx anyway</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3494426,
                  "author_name": "brendanartley",
                  "author_url": "",
                  "post_date": "07/09/2026 23:35:09",
                  "content": "<p>Oops, yes thats correct. Edited above!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3494431,
                      "author_name": "beckpro",
                      "author_url": "",
                      "post_date": "07/09/2026 23:46:52",
                      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> oh sorry to bother u again, but\nis the code license free? or like solution itself has license etc. \ni'm trying to mix up a lot of solutions including ur lead detection modules and others and make my own product.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3395437": "Thanks to the PhysioNet hosts and Kaggle for this fun competition. There were many stages to optimize, and we took the opportunity to learn as much as we could from all stages. Congrats to everyone who competed, Harshit and I are excited to read through all the solution write-ups!\n\n## TLDR\n\nOur pipeline uses rotation, lead detection, lead segmentation, digitization, and out-of-distribution detection models. We modified the ECG-image-kit repository to create lead detection training data, relied on the competition data for digitization, and used out-of-distribution detection models to optimize our ensemble. For more details, keep reading!\n\n## Cross Validation\n\nFor validating experiments, we used a k-fold cross-validation scheme across all samples. We used lightweight models on all folds to detect edge cases; for more computationally expensive models, we only validated on 100 samples. We found that 100 samples were enough for a strong CV/LB correlation and increased the speed of experiments.\n\n## Data Generation\n\nNext, due to a lack of lead annotations in the competition data, we modified ECG-image-kit to create data for the rotation, lead detection, and lead segmentation models. We updated the codebase to insert ECG plots into backgrounds, simulate shadows, and use more variable colors and textures in the ECG plots. All modifications we made to the codebase were done in an attempt to make the artifacts more realistic.\n\nWe used the [PTBXL dataset](https://physionet.org/content/ptb-xl/1.0.3/) for the raw ECG values, and the [Describable Textures Dataset (DTD)](https://www.robots.ox.ac.uk/~vgg/data/dtd/) for backgrounds and shadows. Here are a few samples of what we generated with bounding boxes and segmentation labels overlaid.\n\n![IMG_0](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fbeeed8c1741a9b4311dbd1c3b9a2eac3%2FIMG_4.jpg?generation=1769131261528664&alt=media)\n\n## Rotation\n\nThe first model in the pipeline was a simple classification model to predict when an image needed rotation. We used the B4 and B5 variants from the HGNet-V2 model family to predict 4 classes (0,90,180 or 270 degree rotation). We detected 69 images that required rotation in the training dataset, and applied this model first during inference.\n\n## Lead Detection / Segmentation\n\nNext, we trained a set of hybrid lead detection/segmentation models. This was a great learning curve for me (Bartley), as I have always wanted detection models that come without a confusing license. To do this, we designed a model to predict objectiveness, class scores, and offsets. It was important to add sufficient capacity in the bounding box detection head (>=128 channels) for the model to be able to learn the signal. We found that the ConvNeXt model family worked best, though the architecture supports any backbone from the timm library. \n\nWe first ran the detection model to locate the AOI (area of interest). We then cropped the image and re-ran the model for a more precise result. We also added a minimum crop height of 16 pixels and a width of 64 pixels to account for small lead crops. This fixed all catastrophic detections we observed in the training set, and boosted LB by ~0.2-3. Here are some sample predictions on the training set.\n\n![IMG_1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F3e3547c29e7459fd238e239e2262d341%2FIMG_5.jpg?generation=1769130824182840&alt=media)\n\nWe also added a segmentation branch to predict the pixels corresponding to the 13 different lead classes. We used this predicted segmentation as an input to our digitization models in the next stage.\n\n## Digitization (Bartley)\n\nThe first digitization model we used was a `maxxvitv2_nano_rw_256.sw_in1k` model with a 1D unet decoder. We pooled the encoder features before passing them into the decoder. The model input was a 5-channel image. Three RGB channels, one for the target lead probability, and one for the maximum probability of other leads. The 5-channel input helped generate more robust predictions on crops with overlapping leads.\n\n![IMG_2](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Fd1931b0d7ab3ca66ddd2668d8b13393b%2FIMG_6.jpg?generation=1769130842998048&alt=media)\n\nDuring training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.\n\nWe used a couple of variations of SNRloss. During the low-resolution stage, we used a variation of SNRloss that forces the model to learn the optimal vertical shift. In the later stages (once vertical shift was learned), we used a variation that used the optimal vertical shift to more closely align with the competition metric. The latter approach was able to achieve higher final scores.\n\nFor the lead II full crops, we used a VIT model with a linear head to go from patch embeddings to pixel-level predictions. This architecture was identical to Harshit's, and we will go into more details in the next section.\n\n## Digitization (Harshit)\n\nFor our next set of digitization models, we rely on the same preprocessing steps from Bartley's pipeline. All our models here used a `vit_small_patch16_dinov3.lvd1689m` encoder with slight differences in the training setup. The architecture was heavily inspired by Harshit’s 1st Place Solution in the Yale Competition [here](https://www.kaggle.com/competitions/waveform-inversion/writeups/harshit-sheoran-1st-place-solution).\n\n![IMG_4](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2F7d1d8f99365f40f5662bb4a367daa825%2FIMG_7.png?generation=1769130914922529&alt=media)\n\nWe developed three variations of this model to improve diversity. While the architecture and training method remained consistent, we varied the input data sources (crops) and loss functions. We used random padding augmentation during training on crops derived from the sources below.\n\n| Model Version | Input Data Sources (Crops) | Loss Function |\n| :--- | :--- | :--- |\n| **V27** | `train_crops5`, `train_bartley_crops3` | MAE |\n| **V27-SNR** | `train_crops5`, `train_bartley_crops3` | SNR |\n| **V6** | `train_gen_crops1`, `train_crops4`, `train_bartley_crops2`, `train_bartley_crops4` | MAE |\n\nIn addition, we employed a multi-stage training approach with progressive image size scaling to ensure stable convergence. Each stage consisted of **20 epochs**, scaling up the resolution as follows:\n\n    `224x896` → `224x1782` → `224x2688` → `224x3584` → `336x3584`\n\nWe intentionally kept augmentations minimal, only using horizontal flips during training. Despite the light augmentation pipeline, the models converged effectively. To test the robustness of this pipeline, we took pictures of ECG plots on different monitors to simulate a distribution shift. Scores were consistent with those on the competition set indicating that the models were robust.\n\n## Ensemble\n\nWe expected a large boost when combining our approaches as we used a diverse set of architectures, training pipelines, and loss functions. Harshit’s models excelled on in-distribution samples and when there was a significant drift in the signal. Bartley’s models excelled on out-of-distribution samples and on crops with overlapping leads. A simple mean ensemble of our pipeline scored **22.54/22.10** on the Public/Private LB.\n\n### Out-of-Distribution (OOD) Detection\n\nTo further improve ensemble performance, we implemented an out-of-distribution (OOD) detection model. Since Harshit’s models were highly specialized for in-distribution samples, we wanted to detect when to mask his predictions during inference.\n\nTo do this, we trained a feature extractor using `tf_efficientnetv2_s` and ArcFace Loss. During inference, we calculated the embedding of each test image and compared it to the average embedding of each image type in the training set. If the Cosine Similarity between the test image and any of the average embeddings was <0.5, we masked Harshit's predictions. This strategy boosted our score further to **22.80/22.48**.\n\nAll the models we trained in this competition (Harshit and Bartley) used the Muon optimizer from timm. We found that this significantly outperformed all others and thought it was worth a mention.\n\n## Final Note\n\nLast thing, a quick shout-out to @TheoViel. I recently modified my training pipeline to follow a similar structure to his [RSNA 2023 Solution](https://github.com/TheoViel/kaggle_rsna_abdominal_trauma). It’s an excellent repository that I would recommend checking out.\n\nThanks for reading, and as always, Happy Kaggling!\n\nCode: [here](https://github.com/brendanartley/PhysioNet-Competition)\n\nDatasets: [here](https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission), [here](https://www.kaggle.com/datasets/brendanartley/physionet-2025-submission---other-data)\n\nInference: [here](https://www.kaggle.com/code/harshitsheoran/physionet-infer-v-final)",
    "3395450": "Thanks for sharing.<br>\nI really liked the step-wise task decomposition with minimal interference between stages.<br>\nIt made me realize that properly separating tasks allows us to fully optimize each stage — something I struggled with in my own approach.",
    "3395512": "unique and exceptional. While most of the available solutions are somehow based on hengc23's great work. This one goes a different length!",
    "3395520": "Thanks @sanjidh090!",
    "3395522": "Agreed @sasaleaf, we learnt a lot by implementing each stage separately. Thanks for your comment :)",
    "3396526": "harshitsheoran @brendanartley Congratulations. I impressed Bartley's customized YOLO-like model implemented from scratch. Also Harshit's ViT model which predicts 1d output directly from 2d input.\n\n> We intentionally kept augmentations minimal, only using horizontal flips during training.\n\nFor training digitization model, I wonder strong augmentations (like geometric transforms) can boost 2nd-stage model. Why did you only apply flip augmentations?",
    "3396629": "Hi! @tatamikenn,\n\nHappy to know that you liked our solution,\n\nMy teammate @brendanartley had a great amount of augmentations in his pipeline, \n\n>During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks). We also implemented a custom Albumentations module to add ECG-related keywords/phrases to each image, though it was unclear how much this augmentation improved the models. We were unable to converge this model completely, and were still seeing gains at the end of the competition. We believe that more computing power could lead to further performance improvements with this architecture.\n\nIn my pipeline I was getting great results with or without them,  and training without was certainly faster and easier, and it provides more diversity when ensembling so we continued with both pipelines.",
    "3494057": "brendanartley @harshitsheoran thx for this amazing solution writeup.\nBTW, how did you handle perspective/skew distortion from angled camera photos? \n\nI noticed there's no explicit dewarping step (grid-corner → homography rectification) like some other top solutions use — rotate only does discrete 0/90/180/270, and --augment only applies affine (rotation/crop/noise), no true perspective transform. \n\nWas this intentional (test photos weren't skewed much / small-crop processing absorbs it), or just out of scope? Curious how robust det2.py is on strongly-angled real photos.",
    "3494058": ">During training, we applied heavy color, distortion, rotation, horizontal flips, vertical flips, shifts, and thin coarse dropout (to simulate pen marks).\n\nIs this the one that actually made it possible to handle distortion/skewed/warped ecg images?",
    "3494094": "brendanartley @harshitsheoran \nAlso, what are your thoughts on extending layout domain from '3x4 + 1R only' to adding '6x2', '12x1' layouts etc.\nCause even tho i finetuned this model on '6x2' layout images(about 2000 ecg-image-kit generated), lead detection is not that promising. \nAny suggestion?\nThx again.",
    "3494309": "Hi @beckpro, thanks for the comments. The lead detection model was trained only on ~~12x1~~ 3x4 + 1R layouts, so I would recommend fine-tuning the model on new lead configurations for it to work on new configurations.\n\nIn regards to your comment about augmentations, we intentionally used the heavy augmentations to simulate challenging distorted/skewed/warped images. Although it may not been as strong as other methods, we designed our digitization pipeline with generalization as the main goal! Hope this helps.",
    "3494421": "brendanartley U meant 'was trained only on 3x4 + 1R' ? Cause those are the competition layouts. Thx anyway",
    "3494426": "Oops, yes thats correct. Edited above!",
    "3494431": "brendanartley oh sorry to bother u again, but\nis the code license free? or like solution itself has license etc. \ni'm trying to mix up a lot of solutions including ur lead detection modules and others and make my own product."
  },
  "source": "meta"
}