{
  "id": 669887,
  "title": "17th Place Solution: An Alternative Rectification Workflow",
  "url": "/competitions/physionet-ecg-image-digitization/writeups/17th-place-solution",
  "author_name": "",
  "post_date": "2026-01-24T21:35:59.740Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to the organizers and Kaggle staff for this fun competition. I also want to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the multi-stage pipeline idea. My solution follows that framework but uses a different implementation for image rectification, which is the main part I want to share here.</p>\n<h2>Summary of Solution</h2>\n<ul>\n<li>Single-fold model pipeline (no ensemble).</li>\n<li>Final submission takes approximately &lt; 1.5 hour to run.</li>\n<li>Hardware: All models were trained on a single RTX 4080 GPU (16GB).</li>\n</ul>\n<h3>Pipeline Stages:</h3>\n<ol>\n<li><strong>Homography Transform</strong> (Non-ML)</li>\n<li><strong>Grid-Level Rectification</strong></li>\n<li><strong>Line Intensity Segmentation</strong></li>\n<li><strong>Signal Extraction</strong> (from predicted intensity maps)</li>\n</ol>\n<h2>1. Homography Transformation</h2>\n<p>It seems many teams addressed this stage with an ML model, I implemented a solution utilizing <strong>OpenCV</strong> only:</p>\n<ul>\n<li><strong>Keypoint Detection:</strong> I masked a common template and detected <strong>SIFT keypoints</strong> from both the template and the target image. The text headers at the top of the plots are actually useful for alignment.</li>\n<li><strong>Matching:</strong> Keypoints were matched and filtered using <strong>RANSAC</strong> with a large pixel error threshold for warped paper.</li>\n<li><strong>Alignment:</strong> Applied the resulting Homography Transform (cv2.perspectiveTransform) to align the image to the template plane at a consistent scale.</li>\n</ul>\n<div>\n  <p><strong>Template Mask</strong></p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fe52b6ddc5cfcbcf889784608aa5d8fe2%2Ftemplate.png?generation=1769271432651110&amp;alt=media\">\n</div>\n<table>\n<thead>\n<tr>\n<th>Detected Keypoints &amp; Matching</th>\n<th>Homography-rectified image with overlaid template grid points</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fd004b7d7eb1f304e4dab469bc9b73cdd%2Foutput.png?generation=1769271314756734&amp;alt=media\"></td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F2cdd06c301b32255fb4d136901e5f420%2Fplane_transformed_with_grid.png?generation=1769287934983286&amp;alt=media\"></td>\n</tr>\n</tbody>\n</table>\n<h2>2. Grid-Level Rectification</h2>\n<p>This stage focuses on detecting grid points at a sub-pixel level and matching them to the template grid:</p>\n<ul>\n<li><strong>Model &amp; Inference:</strong> I trained a segmentation model to predict grid point probabilities. During inference, I used the monai sliding window function to predict at the original resolution, subsequently deriving coordinates via <strong>weighted local probabilities</strong> to achieve sub-pixel accuracy.</li>\n<li><strong>Generalization:</strong> Strong augmentation on '0001' images enabled the model to generalize well to other types. Consequently, labels were simply extracted as integer grid points from the '0001' set, requiring no further label generation.</li>\n<li><strong>Filtering:</strong> I also predicted a channel for the <strong>line mask probability</strong>. This is used to filter out false-positive grid points that often appear close to the signal lines.</li>\n<li><strong>Grid Assembly:</strong> After Homography Transformation, many points align closely with the template. I used these to build a structured grid and matched the detected points accordingly.</li>\n<li><strong>Rectification:</strong>  PyTorch's grid_sample with bilinear interpolation to perform the final non-rigid warp.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F1043ab97808af33ede4982f38095c931%2Fpredicted_gridpoints.png?generation=1769286441688794&amp;alt=media\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F3f5c63d512aa0040b05bf098ade21004%2Ffilter_gridpoints.png?generation=1769286349015949&amp;alt=media\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F5e60d97902e9c3b0493cbc37f165cb87%2Fgrid_alignment.png?generation=1769286832445608&amp;alt=media\"></p>\n<h2>3. Line Pixel Segmentation</h2>\n<ul>\n<li><strong>Ground Truth:</strong> Masks were generated from a modified <code>ecg-image-kit</code> plotting function, where probability is calculated based on pixel intensity.</li>\n<li><strong>Training:</strong> I trained the segmentation model using a 512-pixel window on the grid rectified images.</li>\n<li><strong>Inference:</strong> Applied <strong>sliding window prediction</strong> to process the full (1700 x 2200) image.</li>\n</ul>\n<h2>4. Signal Extraction</h2>\n<p>I built a regression model to predict sub-pixel level coordinates directly from the segmentation mask.</p>\n<h3>Model Architecture</h3>\n<pre><code>Input: (B, 1440, 1968)                   # [Batch, Height, Width]\n---------------------------------------------------------------------------\n1. Vertical 2D Conv Encoder (Collapse Height)\n   (B, 1, 1440, 1968)  --&gt;  (B, 512, 1968)  # Height is squashed to 1\n2. CNN &amp; Transformer with ROPE (Mix Context)\n   (B, 512, 1968)      --&gt;  (B, 512, 1968)  # Temporal/Lead mixing (no shape change)\n3. Refinement (Skip Connection)\n   (B, 1024, 1968)     --&gt;  (B, 512, 1968)  # Merge original + mixed features\n4. Upsampling (Double Width)\n   (B, 512, 1968)      --&gt;  (B, 256, 3936)  # Pixel Shuffle\n5. Heatmap &amp; Soft-Argmax (Final Projection)\n   (B, 256, 3936)      --&gt;  (B, 4, 256, 3936) # Split into 4 Leads &amp; 256 Bins\n---------------------------------------------------------------------------\nFinal Output: (B, 4, 3936)               # [Batch, Leads, Y-Coordinates]\n</code></pre>\n<h3>Training &amp; Performance</h3>\n<p>Two losses were used: <strong>MSE loss</strong> for the predicted y-coordinates and <strong>SNR loss</strong> for the converted signal values.</p>\n<ul>\n<li><strong>Pre-training:</strong> The model was pre-trained on ground truth masks from the PTB-XL dataset. Validation SNR reached <strong>30dB</strong> when using competition ground truth line intensity as input. This stage took ~10 hours to converge on a single GPU.</li>\n<li><strong>Fine-tuning:</strong> I fine-tuned the model on the predicted line probability masks. Fine-tuning required only a few epochs before overfitting.</li>\n<li><strong>Results:</strong> Local Validation Performance dropped to <strong>~25-26dB</strong> during fine-tuning, primarily limited by the quality of the predicted probability masks from the previous stage.</li>\n</ul>",
  "messages": [
    {
      "id": "3396377",
      "postDate": "01/24/2026 21:09:23",
      "content": "<p>Thanks to the organizers and Kaggle staff for this fun competition. I also want to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the multi-stage pipeline idea. My solution follows that framework but uses a different implementation for image rectification, which is the main part I want to share here.</p>\n<h2>Summary of Solution</h2>\n<ul>\n<li>Single-fold model pipeline (no ensemble).</li>\n<li>Final submission takes approximately &lt; 1.5 hour to run.</li>\n<li>Hardware: All models were trained on a single RTX 4080 GPU (16GB).</li>\n</ul>\n<h3>Pipeline Stages:</h3>\n<ol>\n<li><strong>Homography Transform</strong> (Non-ML)</li>\n<li><strong>Grid-Level Rectification</strong></li>\n<li><strong>Line Intensity Segmentation</strong></li>\n<li><strong>Signal Extraction</strong> (from predicted intensity maps)</li>\n</ol>\n<h2>1. Homography Transformation</h2>\n<p>It seems many teams addressed this stage with an ML model, I implemented a solution utilizing <strong>OpenCV</strong> only:</p>\n<ul>\n<li><strong>Keypoint Detection:</strong> I masked a common template and detected <strong>SIFT keypoints</strong> from both the template and the target image. The text headers at the top of the plots are actually useful for alignment.</li>\n<li><strong>Matching:</strong> Keypoints were matched and filtered using <strong>RANSAC</strong> with a large pixel error threshold for warped paper.</li>\n<li><strong>Alignment:</strong> Applied the resulting Homography Transform (cv2.perspectiveTransform) to align the image to the template plane at a consistent scale.</li>\n</ul>\n<div>\n  <p><strong>Template Mask</strong></p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fe52b6ddc5cfcbcf889784608aa5d8fe2%2Ftemplate.png?generation=1769271432651110&amp;alt=media\">\n</div>\n<table>\n<thead>\n<tr>\n<th>Detected Keypoints &amp; Matching</th>\n<th>Homography-rectified image with overlaid template grid points</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fd004b7d7eb1f304e4dab469bc9b73cdd%2Foutput.png?generation=1769271314756734&amp;alt=media\"></td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F2cdd06c301b32255fb4d136901e5f420%2Fplane_transformed_with_grid.png?generation=1769287934983286&amp;alt=media\"></td>\n</tr>\n</tbody>\n</table>\n<h2>2. Grid-Level Rectification</h2>\n<p>This stage focuses on detecting grid points at a sub-pixel level and matching them to the template grid:</p>\n<ul>\n<li><strong>Model &amp; Inference:</strong> I trained a segmentation model to predict grid point probabilities. During inference, I used the monai sliding window function to predict at the original resolution, subsequently deriving coordinates via <strong>weighted local probabilities</strong> to achieve sub-pixel accuracy.</li>\n<li><strong>Generalization:</strong> Strong augmentation on '0001' images enabled the model to generalize well to other types. Consequently, labels were simply extracted as integer grid points from the '0001' set, requiring no further label generation.</li>\n<li><strong>Filtering:</strong> I also predicted a channel for the <strong>line mask probability</strong>. This is used to filter out false-positive grid points that often appear close to the signal lines.</li>\n<li><strong>Grid Assembly:</strong> After Homography Transformation, many points align closely with the template. I used these to build a structured grid and matched the detected points accordingly.</li>\n<li><strong>Rectification:</strong>  PyTorch's grid_sample with bilinear interpolation to perform the final non-rigid warp.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F1043ab97808af33ede4982f38095c931%2Fpredicted_gridpoints.png?generation=1769286441688794&amp;alt=media\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F3f5c63d512aa0040b05bf098ade21004%2Ffilter_gridpoints.png?generation=1769286349015949&amp;alt=media\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F5e60d97902e9c3b0493cbc37f165cb87%2Fgrid_alignment.png?generation=1769286832445608&amp;alt=media\"></p>\n<h2>3. Line Pixel Segmentation</h2>\n<ul>\n<li><strong>Ground Truth:</strong> Masks were generated from a modified <code>ecg-image-kit</code> plotting function, where probability is calculated based on pixel intensity.</li>\n<li><strong>Training:</strong> I trained the segmentation model using a 512-pixel window on the grid rectified images.</li>\n<li><strong>Inference:</strong> Applied <strong>sliding window prediction</strong> to process the full (1700 x 2200) image.</li>\n</ul>\n<h2>4. Signal Extraction</h2>\n<p>I built a regression model to predict sub-pixel level coordinates directly from the segmentation mask.</p>\n<h3>Model Architecture</h3>\n<pre><code>Input: (B, 1440, 1968)                   # [Batch, Height, Width]\n---------------------------------------------------------------------------\n1. Vertical 2D Conv Encoder (Collapse Height)\n   (B, 1, 1440, 1968)  --&gt;  (B, 512, 1968)  # Height is squashed to 1\n2. CNN &amp; Transformer with ROPE (Mix Context)\n   (B, 512, 1968)      --&gt;  (B, 512, 1968)  # Temporal/Lead mixing (no shape change)\n3. Refinement (Skip Connection)\n   (B, 1024, 1968)     --&gt;  (B, 512, 1968)  # Merge original + mixed features\n4. Upsampling (Double Width)\n   (B, 512, 1968)      --&gt;  (B, 256, 3936)  # Pixel Shuffle\n5. Heatmap &amp; Soft-Argmax (Final Projection)\n   (B, 256, 3936)      --&gt;  (B, 4, 256, 3936) # Split into 4 Leads &amp; 256 Bins\n---------------------------------------------------------------------------\nFinal Output: (B, 4, 3936)               # [Batch, Leads, Y-Coordinates]\n</code></pre>\n<h3>Training &amp; Performance</h3>\n<p>Two losses were used: <strong>MSE loss</strong> for the predicted y-coordinates and <strong>SNR loss</strong> for the converted signal values.</p>\n<ul>\n<li><strong>Pre-training:</strong> The model was pre-trained on ground truth masks from the PTB-XL dataset. Validation SNR reached <strong>30dB</strong> when using competition ground truth line intensity as input. This stage took ~10 hours to converge on a single GPU.</li>\n<li><strong>Fine-tuning:</strong> I fine-tuned the model on the predicted line probability masks. Fine-tuning required only a few epochs before overfitting.</li>\n<li><strong>Results:</strong> Local Validation Performance dropped to <strong>~25-26dB</strong> during fine-tuning, primarily limited by the quality of the predicted probability masks from the previous stage.</li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers and Kaggle staff for this fun competition. I also want to thank @hengck23 for the multi-stage pipeline idea. My solution follows that framework but uses a different implementation for image rectification, which is the main part I want to share here.\n\n## Summary of Solution\n* Single-fold model pipeline (no ensemble).\n* Final submission takes approximately < 1.5 hour to run.\n* Hardware: All models were trained on a single RTX 4080 GPU (16GB).\n\n### Pipeline Stages:\n1.  **Homography Transform** (Non-ML)\n2.  **Grid-Level Rectification**\n3.  **Line Intensity Segmentation**\n4.  **Signal Extraction** (from predicted intensity maps)\n\n\n## 1. Homography Transformation\nIt seems many teams addressed this stage with an ML model, I implemented a solution utilizing **OpenCV** only:\n* **Keypoint Detection:** I masked a common template and detected **SIFT keypoints** from both the template and the target image. The text headers at the top of the plots are actually useful for alignment.\n* **Matching:** Keypoints were matched and filtered using **RANSAC** with a large pixel error threshold for warped paper.\n* **Alignment:** Applied the resulting Homography Transform (cv2.perspectiveTransform) to align the image to the template plane at a consistent scale.\n\n\n<div align=\"center\">\n  <p><strong>Template Mask</strong></p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fe52b6ddc5cfcbcf889784608aa5d8fe2%2Ftemplate.png?generation=1769271432651110&alt=media\" width=\"400\">\n</div>\n\n| Detected Keypoints & Matching | Homography-rectified image with overlaid template grid points |\n| :---: | :---: |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fd004b7d7eb1f304e4dab469bc9b73cdd%2Foutput.png?generation=1769271314756734&alt=media\" width=\"600\"> | <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F2cdd06c301b32255fb4d136901e5f420%2Fplane_transformed_with_grid.png?generation=1769287934983286&alt=media\" width=\"400\"> |\n\n## 2. Grid-Level Rectification\nThis stage focuses on detecting grid points at a sub-pixel level and matching them to the template grid:\n* **Model & Inference:** I trained a segmentation model to predict grid point probabilities. During inference, I used the monai sliding window function to predict at the original resolution, subsequently deriving coordinates via **weighted local probabilities** to achieve sub-pixel accuracy.\n* **Generalization:** Strong augmentation on '0001' images enabled the model to generalize well to other types. Consequently, labels were simply extracted as integer grid points from the '0001' set, requiring no further label generation.\n* **Filtering:** I also predicted a channel for the **line mask probability**. This is used to filter out false-positive grid points that often appear close to the signal lines.\n* **Grid Assembly:** After Homography Transformation, many points align closely with the template. I used these to build a structured grid and matched the detected points accordingly.\n* **Rectification:**  PyTorch's grid_sample with bilinear interpolation to perform the final non-rigid warp.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F1043ab97808af33ede4982f38095c931%2Fpredicted_gridpoints.png?generation=1769286441688794&alt=media\" width=\"800\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F3f5c63d512aa0040b05bf098ade21004%2Ffilter_gridpoints.png?generation=1769286349015949&alt=media\" width=\"800\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F5e60d97902e9c3b0493cbc37f165cb87%2Fgrid_alignment.png?generation=1769286832445608&alt=media\" width=\"800\">\n\n## 3. Line Pixel Segmentation\n* **Ground Truth:** Masks were generated from a modified `ecg-image-kit` plotting function, where probability is calculated based on pixel intensity.\n* **Training:** I trained the segmentation model using a 512-pixel window on the grid rectified images.\n* **Inference:** Applied **sliding window prediction** to process the full (1700 x 2200) image.\n\n\n## 4. Signal Extraction\nI built a regression model to predict sub-pixel level coordinates directly from the segmentation mask.\n\n### Model Architecture\n```\nInput: (B, 1440, 1968)                   # [Batch, Height, Width]\n---------------------------------------------------------------------------\n1. Vertical 2D Conv Encoder (Collapse Height)\n   (B, 1, 1440, 1968)  -->  (B, 512, 1968)  # Height is squashed to 1\n2. CNN & Transformer with ROPE (Mix Context)\n   (B, 512, 1968)      -->  (B, 512, 1968)  # Temporal/Lead mixing (no shape change)\n3. Refinement (Skip Connection)\n   (B, 1024, 1968)     -->  (B, 512, 1968)  # Merge original + mixed features\n4. Upsampling (Double Width)\n   (B, 512, 1968)      -->  (B, 256, 3936)  # Pixel Shuffle\n5. Heatmap & Soft-Argmax (Final Projection)\n   (B, 256, 3936)      -->  (B, 4, 256, 3936) # Split into 4 Leads & 256 Bins\n---------------------------------------------------------------------------\nFinal Output: (B, 4, 3936)               # [Batch, Leads, Y-Coordinates]\n```\n\n### Training & Performance\nTwo losses were used: **MSE loss** for the predicted y-coordinates and **SNR loss** for the converted signal values.\n\n* **Pre-training:** The model was pre-trained on ground truth masks from the PTB-XL dataset. Validation SNR reached **30dB** when using competition ground truth line intensity as input. This stage took ~10 hours to converge on a single GPU.\n* **Fine-tuning:** I fine-tuned the model on the predicted line probability masks. Fine-tuning required only a few epochs before overfitting.\n* **Results:** Local Validation Performance dropped to **~25-26dB** during fine-tuning, primarily limited by the quality of the predicted probability masks from the previous stage.",
      "votes": null
    },
    {
      "id": "3398037",
      "postDate": "01/28/2026 13:43:48",
      "content": "<p>Congrats on the win. Could you provide the mask generation approach a small code snippet will help</p>",
      "rawMarkdown": "Congrats on the win. Could you provide the mask generation approach a small code snippet will help",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3398037,
      "author_name": "cartoonistbeard",
      "author_url": "",
      "post_date": "01/28/2026 13:43:48",
      "content": "<p>Congrats on the win. Could you provide the mask generation approach a small code snippet will help</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3396377": "Thanks to the organizers and Kaggle staff for this fun competition. I also want to thank @hengck23 for the multi-stage pipeline idea. My solution follows that framework but uses a different implementation for image rectification, which is the main part I want to share here.\n\n## Summary of Solution\n* Single-fold model pipeline (no ensemble).\n* Final submission takes approximately < 1.5 hour to run.\n* Hardware: All models were trained on a single RTX 4080 GPU (16GB).\n\n### Pipeline Stages:\n1.  **Homography Transform** (Non-ML)\n2.  **Grid-Level Rectification**\n3.  **Line Intensity Segmentation**\n4.  **Signal Extraction** (from predicted intensity maps)\n\n\n## 1. Homography Transformation\nIt seems many teams addressed this stage with an ML model, I implemented a solution utilizing **OpenCV** only:\n* **Keypoint Detection:** I masked a common template and detected **SIFT keypoints** from both the template and the target image. The text headers at the top of the plots are actually useful for alignment.\n* **Matching:** Keypoints were matched and filtered using **RANSAC** with a large pixel error threshold for warped paper.\n* **Alignment:** Applied the resulting Homography Transform (cv2.perspectiveTransform) to align the image to the template plane at a consistent scale.\n\n\n<div align=\"center\">\n  <p><strong>Template Mask</strong></p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fe52b6ddc5cfcbcf889784608aa5d8fe2%2Ftemplate.png?generation=1769271432651110&alt=media\" width=\"400\">\n</div>\n\n| Detected Keypoints & Matching | Homography-rectified image with overlaid template grid points |\n| :---: | :---: |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2Fd004b7d7eb1f304e4dab469bc9b73cdd%2Foutput.png?generation=1769271314756734&alt=media\" width=\"600\"> | <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F2cdd06c301b32255fb4d136901e5f420%2Fplane_transformed_with_grid.png?generation=1769287934983286&alt=media\" width=\"400\"> |\n\n## 2. Grid-Level Rectification\nThis stage focuses on detecting grid points at a sub-pixel level and matching them to the template grid:\n* **Model & Inference:** I trained a segmentation model to predict grid point probabilities. During inference, I used the monai sliding window function to predict at the original resolution, subsequently deriving coordinates via **weighted local probabilities** to achieve sub-pixel accuracy.\n* **Generalization:** Strong augmentation on '0001' images enabled the model to generalize well to other types. Consequently, labels were simply extracted as integer grid points from the '0001' set, requiring no further label generation.\n* **Filtering:** I also predicted a channel for the **line mask probability**. This is used to filter out false-positive grid points that often appear close to the signal lines.\n* **Grid Assembly:** After Homography Transformation, many points align closely with the template. I used these to build a structured grid and matched the detected points accordingly.\n* **Rectification:**  PyTorch's grid_sample with bilinear interpolation to perform the final non-rigid warp.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F1043ab97808af33ede4982f38095c931%2Fpredicted_gridpoints.png?generation=1769286441688794&alt=media\" width=\"800\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F3f5c63d512aa0040b05bf098ade21004%2Ffilter_gridpoints.png?generation=1769286349015949&alt=media\" width=\"800\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3964695%2F5e60d97902e9c3b0493cbc37f165cb87%2Fgrid_alignment.png?generation=1769286832445608&alt=media\" width=\"800\">\n\n## 3. Line Pixel Segmentation\n* **Ground Truth:** Masks were generated from a modified `ecg-image-kit` plotting function, where probability is calculated based on pixel intensity.\n* **Training:** I trained the segmentation model using a 512-pixel window on the grid rectified images.\n* **Inference:** Applied **sliding window prediction** to process the full (1700 x 2200) image.\n\n\n## 4. Signal Extraction\nI built a regression model to predict sub-pixel level coordinates directly from the segmentation mask.\n\n### Model Architecture\n```\nInput: (B, 1440, 1968)                   # [Batch, Height, Width]\n---------------------------------------------------------------------------\n1. Vertical 2D Conv Encoder (Collapse Height)\n   (B, 1, 1440, 1968)  -->  (B, 512, 1968)  # Height is squashed to 1\n2. CNN & Transformer with ROPE (Mix Context)\n   (B, 512, 1968)      -->  (B, 512, 1968)  # Temporal/Lead mixing (no shape change)\n3. Refinement (Skip Connection)\n   (B, 1024, 1968)     -->  (B, 512, 1968)  # Merge original + mixed features\n4. Upsampling (Double Width)\n   (B, 512, 1968)      -->  (B, 256, 3936)  # Pixel Shuffle\n5. Heatmap & Soft-Argmax (Final Projection)\n   (B, 256, 3936)      -->  (B, 4, 256, 3936) # Split into 4 Leads & 256 Bins\n---------------------------------------------------------------------------\nFinal Output: (B, 4, 3936)               # [Batch, Leads, Y-Coordinates]\n```\n\n### Training & Performance\nTwo losses were used: **MSE loss** for the predicted y-coordinates and **SNR loss** for the converted signal values.\n\n* **Pre-training:** The model was pre-trained on ground truth masks from the PTB-XL dataset. Validation SNR reached **30dB** when using competition ground truth line intensity as input. This stage took ~10 hours to converge on a single GPU.\n* **Fine-tuning:** I fine-tuned the model on the predicted line probability masks. Fine-tuning required only a few epochs before overfitting.\n* **Results:** Local Validation Performance dropped to **~25-26dB** during fine-tuning, primarily limited by the quality of the predicted probability masks from the previous stage.",
    "3398037": "Congrats on the win. Could you provide the mask generation approach a small code snippet will help"
  },
  "source": "meta"
}